Method and apparatus for video conferencing
By recognizing participant actions using image sensors and computer vision technology and generating environmental graphic feedback, this solves the problem of participants having difficulty getting the host's attention in video conferences, achieving non-intrusive and timely feedback and improving interactivity.
Patent Information
- Application Number
- CN202080105144.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-07-16
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2040-07-16
AI Technical Summary
In video conferencing, it is difficult for participants to get the host's attention in a timely manner without interrupting the host's speech. Existing technologies require manually unmute the microphone or rely on real-time video to observe the participants' actions, which affects the meeting flow.
By capturing participants' head or body movements using image sensors, and using computer vision technology to recognize facial expressions and body movements, environmental graphic feedback is generated and displayed on the presenter's monitor without obscuring the original content, thus achieving non-intrusive feedback.
It enhances the interactivity between participants and the host in video conferences, allowing participants to gain the host's attention in a timely manner without interrupting the meeting flow.
Smart Images

Figure CN116210217B_ABST
Abstract
Description
[0001] Cross-application of related applications
[0002] This is the first application filed for the technology that is immediately available. Technical Field
[0003] The present invention relates generally to video conferencing and, more particularly, to systems, methods, and computing devices for video conferencing. Background Art
[0004] Video conferencing systems are commonly used for real-time audio and video communications between multiple client devices located in different locations. Video conferencing systems enable presenters to interact and share information with attendees of online seminars, presentations, webinars, or meetings located in a variety of locations using client devices.
[0005] When a host uses a video conferencing system to conduct an online seminar, presentation, webinar, or meeting (hereinafter referred to as a video conference), a microphone can capture the host's voice during the video conference and generate an audio signal representing the host's voice. A camera can capture the host's digital video and send the host's digital video for display, such as in a window of a graphical user interface, which can be used to display the digital video on a display. The audio signal representing the captured voice and the video showing the host's speech are transmitted to other computing devices used by the video conference participants along with the information to be shared with the video conference participants. A device associated with (e.g., used by) a participant of the video conference can receive the audio signal representing the captured voice and output the voice using a speaker of its device. The device associated with the participant can also receive the video of the host captured by the camera, as well as the information shared by the host, and display them on the display of the electronic device or an external display connected to the electronic device.
[0006] Information is often displayed on a full-screen monitor. Furthermore, attendees often mute their electronic device microphones while listening to the presenter's speech and viewing the information displayed on their electronic device monitors. This makes it difficult for attendees to interrupt the presenter in a timely manner.
[0007] If the moderator's display doesn't indicate any attendee willingness to intervene, the moderator continues speaking. To capture attendee comments or questions, attendees are required to manually unmute their electronic device microphones. This interruption disrupts the speaker and disrupts the natural flow of the presentation.
[0008] Currently available video conferencing systems for video communication can display, in real time, the video captured by the camera of each electronic device used by a participant on the display of the computing device. The numerous live videos or avatars of the participants obscure the original video material displayed on the electronic device displays. Furthermore, in order to promptly identify when one or more participants may want to intervene, the moderator must constantly monitor the live videos of all participants. Summary of the Invention
[0009] The present invention generally provides a video conferencing system that provides non-intrusive feedback to a moderator during a video conference to improve interactivity between video conference participants. When a video conference participant wishes to attract the attention of the video conference moderator, the participant moves their head, body, part of their body (or their entire body). An image sensor, such as a camera, captures a series of images (i.e., video) of the participant as the participant moves their head, body, or part of their body, and sends the image sequence to a video conferencing server. The video conferencing server processes the image sequence using computer vision technology to identify the type of facial expression or body movement in the image sequence and selects an environmental graphic corresponding to the identified facial expression or body movement. The video conferencing server sends the environmental graphic to a client device associated with the moderator. The client device associated with the moderator displays the environmental graphic on a display of the client device without obscuring information displayed on the display of the client device. During the video conference, the moderator can see the environmental graphic without being disturbed.
[0010] According to one aspect of the present invention, a video conferencing server is provided, comprising: a processor, and a non-transitory storage medium storing instructions, wherein the processor can execute the stored instructions to: receive participant video information from a participant client device; perform object recognition on the participant video information received from the participant client to identify the facial expression or body movement of each participant detected in the participant video information; select an environmental graphic based on the facial expression or body movement of each participant detected in the participant video information; and transmit the selected environmental graphic to a host display associated with a host client device, for presenting the environmental graphic on at least a portion of the current content displayed on the host display associated with the host client device without obstructing the current content.
[0011] According to other aspects of the present invention, the video conferencing server further includes transmitting the selected environmental graphic to a participant display associated with the participant client device for displaying the environmental graphic on at least a portion of the current content displayed on the participant display associated with the participant client device.
[0012] According to other aspects of the present invention, the video conferencing server also includes: receiving participant video information from a second participant client device; performing object recognition on the participant video information received from the second participant client device to identify the facial expressions or body movements of each participant detected in the participant video information; selecting environmental graphics based on the facial expressions or body movements of each participant detected in the participant video information; transmitting the selected environmental graphics to a host display associated with the host client device, for presenting the environmental graphics on at least a portion of the current content displayed on the host display associated with the host client device without blocking the current content.
[0013] According to other aspects of the present invention, the video conferencing server also includes transmitting the selected environmental graphic to a participant display associated with the second participant client device, for displaying the environmental graphic on at least a portion of the current content displayed on the participant display associated with the second participant client device.
[0014] According to other aspects of the present invention, the video conferencing server is characterized in that the environmental graphics are semi-transparent. According to other aspects of the present invention, the video conferencing server is characterized in that the participant information includes audio and video information associated with at least one participant.
[0015] According to other aspects of the present invention, the video conferencing server is characterized in that the video information is collected by an image sensor and the audio information is collected by a microphone.
[0016] According to other aspects of the present invention, the video conferencing server is characterized in that facial expressions include at least one of laughing, smiling or nodding, and body movements include at least one of nodding, tilting the head, raising hands, waving, pointing or clapping.
[0017] According to other aspects of the present invention, the video conferencing server is characterized in that selecting the environment graphic also includes generating the environment graphic.
[0018] According to another aspect of the present invention, the video conferencing server further includes determining a number of facial expressions similar to one or more participants in the participant video information.
[0019] According to another aspect of the present invention, the video conferencing server further includes selecting the ambient graphic based on a number of similar facial expressions associated with one or more participants in the participant video information.
[0020] According to other aspects of the present invention, the video conferencing server further includes calculating the number of similar types of body movements.
[0021] According to another aspect of the present invention, the video conference server is characterized in that the environment graphic is selected according to the number of similar types of body movements associated with one or more participants in the participant video information.
[0022] According to another aspect of the present invention, the video conferencing server is characterized in that the current content includes at least one of a digital document stored on a presenter's client device or a digital document accessed online.
[0023] According to another aspect of the present invention, a method is provided, comprising: receiving attendee video information from an attendee client device; performing object recognition on the attendee video information received from the attendee client device to identify facial expressions or body movements of each attendee detected in the attendee video information; selecting environmental graphics based on the facial expressions or body movements of each attendee detected in the attendee video information; and transmitting the selected environmental graphics to a host display associated with a host client device for presenting the environmental graphics on at least a portion of current content displayed on the host display associated with the host client device without obstructing the current content.
[0024] According to other aspects of the invention, the method further transmits the selected ambient graphic to an attendee display associated with the attendee client device for displaying the ambient graphic on at least a portion of current content displayed on the attendee display associated with the attendee client device.
[0025] According to other aspects of the present invention, the method also includes: receiving participant video information from a second participant client device; performing object recognition on the participant video information received from the second participant client device to identify facial expressions or body movements of each participant detected in the participant video information; selecting environmental graphics based on the facial expressions or body movements of each participant detected in the participant video information; transmitting the selected environmental graphics to a host display associated with the host client device, for presenting the environmental graphics on at least a portion of the current content displayed on the host display associated with the host client device without obstructing the current content.
[0026] According to other aspects of the present invention, the method further includes transmitting the selected ambient graphic to an attendee display associated with a second attendee client device for displaying the ambient graphic on at least a portion of current content displayed on the attendee display associated with the second attendee client device.
[0027] According to another aspect of the present invention, the method is characterized in that the environmental graphics are semi-transparent.
[0028] According to another aspect of the present invention, the method is characterized in that the participant information includes audio and video information associated with at least one participant.
[0029] According to other aspects of the present invention, the method is characterized in that the video information is collected by an image sensor and the audio information is collected by a microphone.
[0030] According to other aspects of the present invention, the method is characterized in that the facial expression includes at least one of laughing, smiling or nodding, and the body movement includes at least one of nodding, tilting the head, raising hands, waving, pointing hands or clapping.
[0031] According to another aspect of the present invention, the method is characterized in that selecting the environment graphic further includes generating the environment graphic.
[0032] According to other aspects of the invention, the method further includes determining a number of facial expressions that are similar to one or more conferees in the conferee video information.
[0033] According to other aspects of the present invention, the method further includes selecting the ambient graphic based on a number of similar facial expressions associated with one or more conference participants in the conference participant video information.
[0034] According to other aspects of the present invention, the method further includes calculating the number of similar types of body movements.
[0035] According to another aspect of the present invention, the method is characterized in that the environmental graphic is selected based on the number of similar types of body movements associated with one or more participants in the participant video information.
[0036] According to another aspect of the present invention, the method is characterized in that the current content comprises at least one of a digital document stored on the moderator's client device or a digital document accessed online.
[0037] According to another aspect of the present invention, a computer-readable storage medium is provided, which includes executable instructions that, when executed by a processor, can: receive participant video information from a participant client device; perform object recognition on the participant video information received from the participant client device to identify the facial expression or body movement of each participant detected in the participant video information; select environmental graphics based on the facial expression or body movement of each participant detected in the participant video information; and transmit the selected environmental graphics to a host display associated with a host client device for presenting the environmental graphics on at least a portion of the current content displayed on the host display associated with the host client device without obstructing the current content.
[0038] According to another aspect of the present invention, a video conferencing system is provided, comprising: at least one participant client device, the at least one participant client device being used to: receive participant video information associated with a participant; perform object recognition on the participant video information to identify facial expressions or body movements of each participant detected in the participant video information; send the facial expressions or body movements detected for each participant in the participant video information to a video conferencing server; the video conferencing server being used to: select environmental graphics based on the facial expressions or body movements received for each participant detected in the participant video information; transmit the selected environmental graphics to a host display associated with a host client device, for presenting the environmental graphics on at least a portion of current content displayed on the host display associated with the host client device without obstructing the current content.
[0039] According to another aspect of the present invention, a video conferencing system is provided, comprising: at least one participant client device, the at least one participant client device comprising: an image sensor, the image sensor being used to collect participant video information; the at least one participant client device transmitting the participant video information to a video conferencing server; the video conferencing server being used to: receive participant video information from the at least one participant client device; perform object recognition on the participant video information received from the at least one participant client device to identify facial expressions or body movements of each participant detected in the participant video information; select environmental graphics based on the facial expressions or body movements of each participant detected in the participant video information; transmit the selected environmental graphics to a host display associated with a host client device, for presenting the environmental graphics on at least a portion of current content displayed on the host display associated with the host client device without blocking the current content. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Other features and advantages of the present invention will become apparent from the following detailed description taken in conjunction with the accompanying drawings, in which:
[0041] Figure 1 A video conferencing system for providing audio and video communications suitable for implementing various embodiments of the present invention is described;
[0042] Figure 2A Various embodiments according to the present invention are described. Figure 1 A block diagram of a client device of a video conferencing system;
[0043] Figure 2B Various embodiments according to the present invention are described. Figure 1A block diagram of a video conferencing server of a video conferencing system;
[0044] Figure 3 Described are flowcharts representing processes for methods implemented on a video conferencing system for video conferencing according to various embodiments of the present invention;
[0045] Figures 4A-4D Examples of environmental graphics according to various embodiments of the present invention are described;
[0046] Figures 5A-5B Describes non-limiting examples of environmental graphics displayed on a presenter's display according to various embodiments of the present invention; and
[0047] Figure 6 Described are flowcharts representing processes for methods implemented on a video conferencing system for video conferencing according to various embodiments of the present invention;
[0048] It should be understood that throughout the drawings and corresponding descriptions, like features are identified by like reference numerals. In addition, it should be understood that the drawings and the following descriptions are for illustration purposes only, and such disclosure does not limit the scope of the claims. DETAILED DESCRIPTION
[0049] The present invention is directed to addressing at least some of the deficiencies of the current technology. Specifically, the present invention describes a system and method for video conferencing.
[0050] In the context of this specification, a "server" is a physical machine, virtual machine, or computer program (e.g., software) running on an appropriate physical or virtual machine and capable of receiving requests from "clients" and executing those requests or causing those requests to be executed. A physical machine can be a physical computer or a physical computer system, but neither is required for the purposes of this technology. A virtual machine is a virtual representation of a physical machine or a physical computer system. In the context of this specification, the use of the word "server" does not mean that every task (e.g., instruction or request received) or any particular task will be received, executed, or caused to be executed by the same server (i.e., the same software and / or machine); it is intended to mean that any number of software modules, routines or functions or hardware devices may participate in receiving / sending, executing, or causing the execution of any task or request, or the results of any task or request; all of which software and hardware may be one server or multiple servers, both of which are included in the expression "a server."
[0051] In the context of this specification, a "client device" is any computer capable of running software (e.g., a client application or program) that accesses xx. Thus, some (non-limiting) examples of client devices include personal computers (desktops, laptops, netbooks, etc.), smartphones and tablets, as well as network devices such as routers, switches, and gateways. It should be noted that a device acting as a client device in this context does not exclude acting as a server for other client devices. Use of the expression "client device" does not exclude multiple client devices from being used to receive / send, perform, or cause the performance of any task or request, or the results of any task or request, or the steps of any method described herein.
[0052] In the context of this specification, unless expressly provided otherwise, terms such as "first," "second," and "third" are used solely to distinguish the nouns that they modify from one another, and are not used to describe any particular relationship between those nouns. Thus, for example, it should be understood that the use of the terms "first server" and "third server" is not intended to imply any particular order, type, chronological order, hierarchy, or ranking (for example) of servers, nor does their use (by itself) mean that there must be any "second server" in any given situation. Furthermore, as discussed elsewhere herein, references to a "first" element and a "second" element do not preclude the two elements from being the same actual, real-world element. Thus, for example, in some cases, the "first" server and the "second" server may be the same software and / or hardware, and in other cases, they may be different software and / or hardware.
[0053] In the context of this specification, the expression "information" includes information of any nature or kind. Thus, information includes, but is not limited to, audiovisual works (images, films, recordings, presentations, etc.), data (location data, numerical data, etc.), text (opinions, comments, questions, messages, etc.), documents, spreadsheets, etc.
[0054] In the context of this specification, the expression "document" should be broadly interpreted to include any machine-readable and machine-storable work product. Documents may include emails, websites, files, combinations of files, one or more files with embedded links to other files, newsgroup posts, blogs, online advertisements, etc. In the context of the Internet, a common document is a web page. A web page typically contains text information and may contain embedded information (such as meta information, images, hyperlinks, etc.) and / or embedded instructions (such as JavaScript, etc.). A page may correspond to a document or a part of a document. Therefore, in some cases, the terms "page" and "document" can be used interchangeably. In other cases, a page may refer to a part of a document, such as a subdocument. A page may also correspond to multiple documents.
[0055] In the context of this specification, unless explicitly stated otherwise, a "database" is any structured collection of data, regardless of the specific structure, database management software, or computer hardware used to store, implement, or otherwise present the data. A database may reside on the same hardware as the processes that store or use the information stored in the database, or it may reside on separate hardware, such as a dedicated server or multiple servers.
[0056] Each implementation of the present technology has at least one of the above-mentioned objectives and / or aspects, but does not necessarily have all of these objectives and / or aspects. It should be understood that some aspects of the present technology are intended to achieve the above-mentioned objectives, but may not meet such objectives, and may also meet other objectives not specifically set forth herein.
[0057] The examples and conditional language described herein are intended primarily to help readers understand the principles of the present technology, rather than to limit its scope to these specific examples and conditions. It should be understood that those skilled in the art can design various devices that, although not explicitly described or shown herein, embody the principles of the present technology and are included within its spirit and scope.
[0058] In addition, to facilitate understanding, the following description may describe a relatively simplified implementation of the present technology. Those skilled in the art will appreciate that various implementations of the present technology may have greater complexity.
[0059] In some cases, examples of modifications that are considered useful for the present technology may also be listed. This is done solely to aid understanding and, again, is not intended to define the scope of the present technology or to specify the limits of the present technology. These modifications are merely examples of some of them, and those skilled in the art may make other modifications while remaining within the scope of the present technology. Furthermore, the absence of examples of modifications should not be interpreted as implying that modifications are impossible and / or that the description is the only way to implement that element of the present technology.
[0060] In addition, all descriptions herein describing the principles, aspects, and implementations of the present invention, as well as specific examples thereof, are intended to encompass structural and functional equivalents thereof, whether currently known or developed in the future. Thus, for example, those skilled in the art will understand that any block diagram herein represents a conceptual view of an illustrative circuit that embodies the principles of the present invention. Similarly, it should be understood that any flow charts, flow diagrams, state transition diagrams, pseudocode, and the like represent various processes that can be substantially represented in a computer-readable medium and thus executed by a computer or processor, regardless of whether such computer or processor is explicitly shown.
[0061] The functions of the various elements shown in the figure (including any functional blocks marked as "processors") can be provided by using dedicated hardware and hardware capable of executing software in association with appropriate software. When provided by a processor, these functions can be provided by a single dedicated processor, a single shared processor, or multiple separate processors, some of which can be shared. In some embodiments of the present technology, the processor can be a general-purpose processor, such as a central processing unit (CPU) or a processor dedicated to a specific purpose, such as a graphics processing unit (GPU), a tensor processing unit (TPU), a neural processing unit (NPU). In addition, the explicit use of the term "processor" should not be interpreted as referring specifically to hardware capable of executing software, and may implicitly include but is not limited to arithmetic and logic units, control units, and storage units for storing instructions, data, and intermediate results, as well as hardware accelerators in the form of dedicated integrated circuits or field programmable gate arrays for performing hardware acceleration. Other traditional and / or custom hardware may also be included.
[0062] Software modules, or modules simply implied as software, may be represented herein as any combination of flow chart elements or other elements indicating process steps and / or performance described in the text. Such modules may be executed by hardware, whether explicitly or implicitly shown.
[0063] With these basic elements, the present invention is intended to address at least some of the deficiencies of the current technology. Specifically, the present invention describes a system and method for video conferencing.
[0064] The present invention is directed to addressing at least some of the deficiencies of the current technology. Specifically, the present invention describes a system and method for video conferencing.
[0065] Figure 1 A video conferencing system 100 for real-time audio and video communications is described according to an embodiment of the present invention. The video conferencing system 100 includes a plurality of client devices 112 located at different geographical locations, which are configured to communicate with each other via a communication network 106 and a video conferencing server 250. The plurality of client devices 112 include a first client device 112 associated with (e.g., used by) a host (i.e., moderator 110) of the video conference, a second client device 112 associated with (e.g., used by) a first participant 120 of the video conference, and a third client device 112 associated with (e.g., used by) a second participant 120 of the video conference. The video conferencing system 100 may also include peripheral devices (not shown), such as speakers, microphones, cameras, and display devices, located at different geographical locations, which may communicate with the video conferencing server 250 via the communication network 106. Although Figure 1Two client devices 112 are shown, each associated with one participant 120, but it should be understood that in alternative embodiments, the video conferencing system 100 can include any number of client devices 112. Furthermore, in other alternative embodiments, a client device 112 can be associated with multiple participants 120.
[0066] Figure 2A A block diagram of a client device 112 is depicted, according to an embodiment of the present invention. Client device 112 can be any suitable type of computing device, including a desktop computer, a laptop computer, a tablet computer, a smartphone, a portable electronic device, a mobile computing device, a personal digital assistant, a smartwatch, an e-reader, an internet-enabled application, and the like. Client device 112 has multiple components, including a processor 202 that controls the overall operation of client device 112. Processor 202 is coupled to and interacts with other components of client device 112, including one or more storage units 204, one or more memories 206, a display device 208 (hereinafter referred to as display 208), a network interface 210, a microphone 212, a speaker 214, and a camera 216 (interchangeably referred to as an image sensor 216). In addition to memory 206, display 208, network interface 210, microphone 212, speaker 214, and camera 216, client device 112 also includes a power supply 218 that powers the components of client device 112. The power source 218 may include a battery, a power pack, a micro fuel cell, etc. However, in other embodiments, the power source 218 may include a port (not shown) for connecting to an external power source and a power adapter (not shown), such as an AC-to-DC adapter, to power the components of the client device 112. Optionally, the client device 112 includes one or more input devices 220, one or more output devices 222, and an I / O interface 222.
[0067] The processor 202 of the client device 112 may include one or more of a central processing unit (CPU), an accelerator, a microprocessor, a graphics processing unit (GPU), a tensor processing unit (TPU), a neural processing unit (NPU), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), dedicated logic circuits, a dedicated artificial intelligence processor unit, or a combination thereof.
[0068] Processor 202 is configured to communicate with storage unit 204, which may include a mass storage unit such as a solid-state drive, a hard disk drive, a magnetic disk drive, and / or an optical disk drive. Processor 202 is also configured to communicate with memory 206, which may include volatile memory [e.g., random access memory (RAM)] and non-volatile or non-transitory memory [e.g., flash memory, magnetic storage, and / or read-only memory (ROM)]. Non-transitory memory stores applications or programs including software instructions for execution by processor 202, such as for executing the examples described herein. Non-transitory memory stores video conferencing applications, as described in further detail below. Examples of non-transitory computer-readable media include RAM, ROM, erasable programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), flash memory, CD-ROM, or other portable memory.
[0069] The processor 202 is also used to communicate with the display 208, which may include a flat panel display [e.g., a liquid crystal display, a plasma display, a light emitting diode (LED) display, an organic light emitting diode display (OLED)], a capacitive, resistive, surface acoustic wave (SAW) touch screen display, or any one of an optical touch screen display.
[0070] The processor 202 is further configured to interact with a network interface 210. The network interface 210 may include one or more radios for wireless communication (e.g., cellular or WiFi communication) with the communication network 106, or one or more network adapters for wired communication with the communication network 106. Generally, the network interface 210 is configured to correspond to a network architecture for implementing a communication link between the client device 112 and the communication network 106. The communication network 106 may be the Internet, a local area network, a wide area network, or the like.
[0071] The processor 202 is further configured to interact with a microphone 212, a speaker 214, and a camera 216. The microphone 210 includes any suitable transducer that converts sound into an audio signal and provides the audio signal to the processor 202 for processing and / or transmission to another client device 112. The speaker 214 includes any suitable transducer that receives an audio signal from the processor 202 and converts the audio signal received from the processor 202 into sound waves. The camera 216 is configured to capture video (e.g., a series of digital images) of the field of view of the camera 216 and provide the captured video to the processor 202 for processing. The camera 216 can be any suitable digital camera, such as a high-definition camera, an infrared camera, a stereo camera, etc. In some embodiments, the microphone 210, speaker 214, and camera 216 can be internally integrated into the client device 212. In other embodiments, the microphone 210, speaker 214, and camera 216 can be externally coupled to the client device 112.
[0072] Optionally, the processor 202 can communicate with an input / output (I / O) interface 222 to connect one or more input devices 220 (e.g., a keyboard, a mouse, a joystick, a trackball, a fingerprint sensor, etc.) and / or output devices 222 (e.g., a printer, a peripheral display device, etc.).
[0073] In addition to the processor 202, memory 206, display 208, network interface 210, microphone 212, speaker 212, and camera 214, the client device 112 also includes a bus 226 that provides communication between the components of the client device 112. The bus 226 can be any suitable bus architecture, including, for example, a memory bus, a peripheral bus, or a video bus.
[0074] Figure 2B A block diagram of a video conferencing server 250 is shown in accordance with an embodiment of the present invention. In this embodiment, the video conferencing server is a physical machine (e.g., a physical server) or a virtual machine (e.g., a virtual server) that executes video conferencing system software to enable client devices 112 to participate in video conferences. The video conferencing server 250 includes a processor 252, a memory 254, and a network interface 256.
[0075] The processor 252 of the video conferencing server 250 may include one or more of a central processing unit (CPU), an accelerator, a microprocessor, a graphics processing unit (GPU), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), a dedicated logic circuit, a dedicated artificial intelligence processing unit, or a combination thereof.
[0076] Memory 254 may include volatile memory (e.g., random access memory (RAM)) and non-volatile or non-transitory memory (e.g., flash memory, magnetic storage, and / or read-only memory (ROM)). The non-transitory memory stores platform 258, which controls the overall operation of video conferencing server 250. Platform 258, when executed by processor 252, implements the video conferencing service. Platform 258 stores a unique identifier for each user of the video conferencing service in memory and manages the unique identifier or each user of the video conferencing service. A user's unique identifier may be a user's username or email address. A password may also be associated with the user's unique identifier and stored in memory 254.
[0077] The network interface 256 may include one or more radios for wireless communication with the communication network 106, or one or more network adapters for wired communication with the communication network 106. In general, the network interface 256 is configured to correspond to the network architecture used to implement the communication link between the video conferencing server 250 and the communication network 106.
[0078] It should be noted that server 250 is shown as a standalone computer. However, implementations of various other embodiments of the present invention may include any client-server model in which client devices can run a client version of the video conferencing system software. Other examples of server 250 may include a distributed computing system running a server version of the video conferencing system software, a virtual machine (or multiple virtual machines) instantiated by a public or private cloud infrastructure, or a cloud service provider providing video conferencing system software as a service (SaaS). This implementation or any other similar implementation should not limit the scope of the present invention.
[0079] See again Figure 1 , the client device 112 associated with the host 110 is referred to herein as the host client device 112, and the client device 112 associated with the attendee 120 is referred to herein as the attendee client device 112. In certain non-limiting embodiments, the host client device 112 and each attendee client device 112 can be used to store and implement instructions associated with the video conferencing system software. In other words, the video conferencing system software can be installed on the host client device 112 and the attendee client device 112 to improve the video conferencing experience between the host 112 and the attendee 120. It should be noted that the version of the video conferencing system software may vary from device to device. Specifically, the version of the video conferencing system software may depend on the operating system associated with the host client device 112 and the attendee client device 112. For example, if the operating system of any one of the host client device 112 and the attendee client device 112 is Android TM 、iOSTM , Windows TM etc., you can download and install the video conferencing system software from their respective app stores.
[0080] In other non-limiting embodiments, at least one of the host client device 112 and the attendee client device 112 may use a web browser, such as Chrome. TM Safari TM , Mozilla TM It should be noted that the manner in which the host client device 112 and the participant client device 112 are used to improve the video conference experience should not limit the scope of the present invention by any means.
[0081] In certain non-limiting embodiments, the host client device 112 can be configured to send a video conference invitation associated with a future video conference to an attendee client device 112. Such a video conference invitation can include a time, date, duration, or any other information associated with the future video conference. In some embodiments, the host client device 112 can send the conference invitation using any suitable means, such as electronic mail (email), text message, etc. In some embodiments, the conference invitation can be a password-protected hyperlink, meaning that an attendee client device 112 may require a password to join the video conference. In other embodiments, the conference invitation can be an open hyperlink, meaning that any attendee client device 112 with access to an open network link can join the video conference.
[0082] In certain non-limiting embodiments, the host client device 112 can be located at a first location (e.g., office, home, etc.) associated with the host 110. Each attendee client device 112 can be located at a different location, for example, the first attendee client device 112 can be located at a second location (e.g., office, home, etc.) associated with the first attendee 120, the second attendee client device 112 can be located at a third location (e.g., office, home, etc.) associated with the second attendee 120, and so on.
[0083] In certain non-limiting embodiments, the host client device 112 may be located at a first location (e.g., an office, a home, etc.) associated with the host 110. However, one or more attendee client devices 112 may be located at a common location with more than one attendee 120. For example, at least one attendee client device 112 may be associated with at least two attendees 120. In other examples, one or more attendee client devices 112 may be located at the same location, such as a conference room, a meeting room, etc., and the location may have at least two attendees 120.
[0084] According to a non-limiting embodiment of the present invention, the moderator client device 112 can be used to initiate a video conference between the moderator 110 and the attendees 120 via the communication network 106. After the video conference is initiated, the moderator client device 112 can be used to communicate with the attendee client devices 112.
[0085] In some embodiments, the host client device 112 can share various types of host information 132 with the attendee client devices 112. In some embodiments, the information sharing between the host client device 112 and the attendee client devices 112 can be routed via the video conferencing server 250. Such host information 132 can include, but is not limited to, real-time video of the host 110 captured by a camera 216 associated with the host client device 112 (hereinafter referred to as the host camera 216), audio / sound from the host 110 side captured by a microphone 212 associated with the host client device 112 (hereinafter referred to as the host microphone 212), content displayed in a graphical user interface (GUI) 130 associated with the video conferencing system software (e.g., MS PowerPoint) on a display device 208 associated with the host client device 112 (hereinafter referred to as the host display 208), and the like. TM Presentation slides, MS Word TM document pages, videos, images, pictures, etc.).
[0086] In some embodiments, the attendee client device 112 may receive the presenter information 132 provided by the attendee client device 112. If the presenter information 132 includes a real-time video of the presenter 110 and / or content displayed on the presenter display 208 (e.g., MS PowerPoint TM Presentation slides, MS Word TM If the presenter information 132 includes audio / sound from the presenter 110 side, the audio / sound may be generated by a speaker 214 associated with the conferee client device 112 (hereinafter referred to as the conferee speaker 214).
[0087] It should be noted that the GUI 130 associated with the video conferencing system software can provide various options to the host 110 and attendees 120. In some non-exhaustive examples, the GUI 130 associated with the host client device 112 (hereinafter referred to as the host GUI 130) can provide an option to select specific content from the host client device 112 to be displayed on the host display 208. In another example, the host GUI 130 can provide options to turn on or off various peripherals, such as the host microphone 212, the host camera 216, and so on. In another example, the host GUI 130 can provide an option to add more attendees 120 during an ongoing video conference. In another example, the host GUI 130 can provide an option to record the ongoing video conference, saving the recording to the host client device 112, the video conferencing server 250, or some public or private cloud. In another example, the host GUI 130 can provide an option to schedule a video conference and send invitations to attendees 112. In another example, the host GUI 130 can provide an option to end or leave the video conference.
[0088] In another example, the moderator GUI 130 may provide options for setting various permissions for the attendee client device 112 during the video conference. Such permissions may include whether the GUI 130 associated with the attendee client device 112 (hereinafter referred to as the attendee GUI 130) can share some content, whether more attendees 120 can be added, and whether the video conference can be recorded. It should be noted that the attendee GUI 130 may have similar options as the moderator GUI 130.
[0089] In another example, the moderator GUI 130 may provide a small window (e.g., a window that is smaller than the size of the moderator display 208) to include a list of participants 120 that have joined the video conference, displaying the video of at least one participant 120. It should be noted that in some embodiments, the small window may be hidden by default and may be displayed or popped up on the moderator display 208 based on certain actions performed by the moderator 110, such as selecting the small window by any suitable means. Some non-exhaustive reasons for hiding the small window include that the small window may require some display space on the moderator GUI 130; as the number of participants 120 increases, it may be difficult for the moderator 110 to notice and focus on the small window; in some cases, it may be difficult for the moderator 110 to focus on the presentation during the conference because some moving images may be distracting, etc.
[0090] Thus, the video conferencing system 100 can be used to provide non-intrusive feedback to the presenter client device 112 during a video conference, whenever needed or otherwise.
[0091] During a video conference (ie, when communication between the presenter client device 112 and the attendee client device 112 has been established using the communication network 106 via the video conference server 250), the content to be shared (eg, MS PowerPoint TM Presentation slides, MS Word TM The host 110 may display a page of a document, a video, an image, a picture, etc. on the host display 208. If the host 110 has enabled the host camera 216, the host camera 216 may capture a video (i.e., a series of images) of the host 110 (i.e., a real-time video), and the host microphone 212 may capture the voice of the host 110 and any sounds around the host 110 (hereinafter, the voice of the host 110 and any sounds around the host 110 are collectively referred to as host audio / sound information).
[0092] In some embodiments, the presenter client device 112 can be configured to send presenter information 132 to the video conferencing server 250. The presenter information 132 can include one or more of content displayed on the presenter display 208, a captured series of images of the presenter 110, and presenter audio / sound information. The video conferencing server 250 can be configured to send the presenter information 132 to the attendee client devices 112. Any visual content included in the presenter information 132 (e.g., content displayed on the presenter display 112, video of the presenter 110 (i.e., a series of images)) can be displayed on the attendee display 120, and any audible content in the presenter information 132 can be generated by the attendee speaker 214.
[0093] To analyze the attributes of the participant 120 (e.g., body movements, facial expressions, etc.), the participant camera 216 can capture a series of images of the participant 120 (i.e., real-time video), and the participant microphone 212 can capture the voice of the participant 120 and any sounds surrounding the participant 120 (hereinafter, the voice of the participant 120 and any sounds surrounding the participant 120 are collectively referred to as participant audio / sound information). In some embodiments, the participant client device 112 can be configured to send participant information 134 to the video conferencing server 250. The participant information 134 can include a series of captured images of the participant 120 (hereinafter also referred to as participant video information 134). The participant information 134 can also include content displayed on the participant display 208 and participant audio / sound information.
[0094] It is contemplated that in some embodiments, when all of the attendees 120 are located in different locations, attendee information 134 may be associated with each individual attendee 120. In these embodiments, each attendee client device 112 may be configured to send the corresponding attendee information 134 to the video conferencing server 250. In some embodiments, where multiple attendees 120 may be located in the same location (e.g., a conference room or meeting room, etc.), the associated attendee client device 112 may have one or more attendee displays 208, one or more attendee microphones 212, one or more attendee speakers 214, and one or more cameras 216. In these embodiments, the associated attendee client device 112 may compile the attendee information 134 from the one or more attendee microphones 212 and the one or more cameras 216 and send the corresponding attendee information 134 to the video conferencing server 250.
[0095] As previously described, participant information 134 may include participant audio / sound information and participant video information. Some participant audio / sound information may be useful, while some participant audio / sound information may simply be noise. Useful audio / sound information may include questions, comments, suggestions, requests to start a discussion, applause, etc. associated with one or more participants 120. Useful participant audio / sound information may be directly related to the ongoing video conference. However, there is also participant audio / sound information that is identified as noise, which may include coughing, sneezing, babies crying, dogs barking, traffic sounds, music / TV playing in the background, knocking on the table, phone ringing, talking to other people, or any such sounds associated with one or more participants 120 or generated in the surrounding environment of one or more participants 120 that are not directly related to the ongoing video conference.
[0096] Similarly, some attendee video information may be useful, while some may be just noise. Useful attendee video information may include, for example, body movements, attention-getting body language (raising hands, pointing, etc.), body language expressing agreement or disagreement (e.g., nodding), and facial expressions indicating attendee attention, such as eye gaze, gestures, and unconscious body movements. All of this useful attendee video information may be included in a series of images received by the video conferencing server 250, which are processed by the video conferencing server 250 to determine various attributes of the attendees 120. The video conferencing server 250 may use the attributes to determine indications about the attendees 120. Indications may include, but are not limited to, if one or more attendees 120 want to ask a question, if one or more attendees 120 are focused on and understanding the host information 132, if one or more attendees become sentimental or emotional about the host information 132, if one or more attendees 120 are laughing and enjoying the host information 132, if one or more attendees 120 are indifferent to or have lost interest in the host information 132, and so on.
[0097] However, there is also participant video information that is identified as noise, which may include if one or more participant 120 is eating and / or drinking, someone is moving around, or someone is passing behind one or more participant 120, one or more participant is moving, and the moving background is captured by one or more participant cameras 216. Such participant video information may not provide any useful information that may be directly or indirectly related to the ongoing video conference.
[0098] In order to determine various attributes associated with the attendee 120 in the attendee information 134 (e.g., including body movements and facial expressions such as raising a hand, waving, pointing, and clapping), in some embodiments, the video conferencing server 250 may process the attendee video information 134 using any suitable computer vision technology described below (i.e., performing face detection and body detection on a series of images included in the attendee information 134). In some embodiments, the video conferencing server 250 may be configured to process the attendee audio / sound information included in the attendee information 134.
[0099] Thus, in some embodiments, memory 254 associated with video conferencing server 250 may store instructions for video conferencing software to be executed by processor 252 to perform the methods of the present invention. In some embodiments, the instructions may include training a neural network (i.e., a neural network including parameters learned during a training parameter period) that receives a series of images and performs face detection, person detection, face tracking, and person tracking on the series of images.
[0100] In certain non-limiting embodiments, the video conferencing server 250 may be configured to perform face detection on the participant video information 134 to detect one or more faces in the participant video information 134, where each detected face corresponds to one of the participant 120 in the participant video information 134. Based on each face detected in the participant video information 134, the video conferencing server 250 may generate a bounding box for each corresponding detected face. Furthermore, the video conferencing server 250 may be configured to perform face recognition on each corresponding detected face in the participant video information 134. Performing face recognition on each corresponding detected face includes monitoring changes in the bounding box generated for the corresponding face in the participant video information 134 to determine facial attributes of the corresponding detected face, and analyzing the facial attributes of the corresponding detected face to infer (i.e., predict) a facial expression, emotion, or attention of the corresponding detected face. Examples of facial attributes include head pose, facial landmarks (e.g., forehead, lips, eyes), and eye gaze. Examples of facial expressions inferred (i.e., predicted) for a detected face (i.e., a participant 120 of the video conference) include laughing, smiling, and nodding; examples of attention inferred (i.e., predicted) for a detected face include looking at the participant display; examples of emotions inferred (i.e., predicted) for a detected face include a serious expression.
[0101] In certain non-limiting embodiments, the video conferencing server 250 may also be configured to perform facial landmark recognition to identify facial landmarks of a detected face, such as the forehead, eyes, and lips. Facial landmark recognition includes detecting facial landmarks (e.g., forehead, lips, eyes, etc.) in the detected face, generating sub-bounding boxes for the detected facial landmarks, monitoring changes in the sub-bounding boxes generated for the facial landmarks to determine attributes of the facial landmarks, and analyzing the attributes of the facial landmarks to infer (i.e., predict) the facial landmarks. Facial landmark recognition generates information indicating the type of facial landmarks recognized.
[0102] In certain non-limiting embodiments, before performing facial recognition, the video conferencing server 250 crops the participant video information 134 to generate new participant video information 134 that includes only a portion of the participant video information corresponding to the bounding box generated for the detected face. In other words, each image in the series of images comprising the participant video information 134 is cropped to include a portion of the image corresponding to the bounding box generated for the detected face. In this embodiment, the video conferencing server 250 is configured to perform facial recognition on the new participant video information 134. Performing facial recognition on the new participant video information 134 includes monitoring changes in the new participant video information to determine facial attributes of the detected face, and analyzing the facial attributes of the detected face to infer (i.e., predict) the facial expression, emotion, or attention of the detected face.
[0103] In some non-exhaustive exemplary embodiments, the video conferencing server 250 may be configured to calculate the number of attendees 120 viewing the screen by analyzing the recognized facial attributes of each face detected in the video attendee information 134 (i.e., each attendee 120). In another exemplary embodiment, the video conferencing server 250 may analyze the facial expression inferred for each detected face (i.e., each attendee 120) to determine the overall attention level of the attendees 120. Specifically, in this manner, the video conferencing server 250 may determine the number of attendees 120 with a particular facial expression (e.g., laughing, smiling, serious, bored, etc.).
[0104] In certain non-limiting embodiments, the video conferencing server 250 may be configured to perform human body detection on the participant video information 134 to detect one or more human bodies in the participant video information 134, wherein each detected human body corresponds to one participant 120 in the participant video information 134. Based on each human body detected in the participant video information 134, the video conferencing server 250 may generate a bounding box for each detected human body. The video conferencing server 250 may be configured to perform body movement recognition (also known as body language recognition) on the participant video information 134 to infer (i.e., predict) the body movement (also known as body language) of each detected human body. The body movement recognition for each corresponding detected human body includes monitoring changes in a bounding box generated for the corresponding detected human body (or human body part) in the participant video information 134 to determine the body movement attributes of the corresponding detected human body, and analyzing the body movement attributes of the corresponding detected human body (or human body part) to infer (i.e., predict) the body movement (also known as body language) of the detected human body. Body movement attributes include the speed of a person's movement (or the speed of a person's part ("body part"), such as a hand, arm, or leg), the duration of the body movement (or the duration of the body part movement), the intensity of the body movement (or the intensity of the body part movement), and the relative range of the body movement (or the relative range of the body part movement). Body movements (e.g., body language) inferred (i.e., predicted) by the video conferencing server 250 may include nodding, tilting the head, raising a hand, waving, pointing, clapping, etc.
[0105] In certain non-limiting embodiments, before performing body movement recognition, the video conferencing server 250 crops the participant video information 134 to generate new participant video information 134, the new information including only a portion of the participant video information 134 corresponding to the bounding box generated for the detected human body. In other words, each image in the series of images constituting the participant video information 134 is cropped to include a portion of the image corresponding to the bounding box generated for the detected human body. In this embodiment, the video conferencing server 250 is configured to perform body movement recognition on the new participant video information 134. Performing body movement recognition on the new participant video information 134 includes monitoring changes in the new participant video information 134 to determine the body movement attributes of the detected human body, and analyzing the body movement attributes of the detected human body to infer (i.e., predict) the body movement (body language) of the detected human body.
[0106] In some non-exhaustive examples, the video conferencing server 250 may count the number of participants 120 who raised their hands to ask questions, waved to get the attention of the host 110, or applauded, based on the recognized body movements (e.g., body language). In some embodiments, to correctly infer (e.g., predict) body movements (e.g., body language), the video conferencing server 250 may analyze body movement attributes such as speed, duration, and intensity of the movement. For example, if one of the participants 120 moves their hand to perform some other movement, such as picking up a pen, and quickly returns their hand to its original position or any other position other than raised, the video conferencing server 250 may not recognize the movement as a raised hand by the participant 120. In another example, the video conferencing server 250 may analyze the speed of human body (or human body part) movements, such as the speed at which the participant 120 waves his / her hand.
[0107] In certain non-limiting embodiments, each respective attendee client device 112 can perform face detection, facial landmark detection, and face recognition to identify facial attributes of each detected face in the attendee information 134. In these embodiments, each attendee client device 112 sends the facial expression identified for each detected face to the video conferencing server 250, which analyzes the facial attributes of each respective detected face and infers (i.e., predicts) the facial expression, emotion, or attention of each respective detected face. By performing face detection, facial landmark detection, and face recognition at each attendee client device 112, the amount of data transmitted between each attendee client device 112 and the video conferencing server 250 is significantly reduced because the attendee video information 134 is not transmitted to the video conferencing server 250.
[0108] In certain non-limiting embodiments, each corresponding attendee client device 112 can perform human body detection and body movement recognition (e.g., body language recognition) to identify the body movement attributes of each detected person in the attendee information 134. In these embodiments, each attendee client device 112 sends the inferred (i.e., predicted) body movement for each detected person to the video conferencing server 250, which performs body movement recognition (e.g., body language recognition) using the body movement attributes of each corresponding detected person to determine the body movement (e.g., body language) of each corresponding detected person. By performing human body detection and body movement recognition at each attendee client device 112, the amount of data transmitted between each attendee client device 112 and the video conferencing server 250 is significantly reduced because the attendee video information 134 is not transmitted to the video conferencing server 250.
[0109] In certain non-limiting embodiments, the video conferencing server 250 can be configured to filter out participant video information 134 that is identified as noise within the participant video information 134. For example, if one or more participant 120 is eating and / or drinking, someone is moving around, or someone is passing behind one or more participant 120, one or more participant is moving, and the moving background is captured by one or more participant cameras 216, then this portion of the participant video information may not provide any useful information directly or indirectly related to the ongoing video conference. The video conferencing server 250 can be configured to remove this portion of the participant video information 134.
[0110] In certain non-limiting embodiments, the video conferencing server 250 may be configured to process the participant audio / sound information present in the participant information 134. In some non-exhaustive examples, the video conferencing server 250 may analyze the participant audio / sound information to determine whether the participant 120 is clapping or whether one or more of the participant 120 is asking a question. In certain non-limiting embodiments, the video conferencing server 250 may be configured to filter out some participant audio / sound information identified as noise in the participant information 134. For example, the video conferencing server 250 may filter out a portion of the participant audio / sound information including coughing, sneezing, crying babies, barking dogs, traffic sounds, music / TV playing in the background, banging on tables, ringing phones, conversations with other people, or any other such sounds associated with one or more of the participants 120 or generated in the surrounding environment of one or more of the participants 120 that are not directly related to the ongoing video conference.
[0111] It should be noted that in some embodiments, the video conferencing server 250 may use any suitable audio processing technology to process the audio / sound included in the attendee information 134. How the attendee information 134 is processed should not limit the scope of the present invention. Furthermore, in the above examples, the attendee information 134 is being processed by the video conferencing server 250. However, in some embodiments, the attendee information 134 may be processed locally on the attendee client device 112, and the resulting information may be forwarded to the video conferencing server 250 for further processing.
[0112] After processing the attendee information 134, the video conferencing server 250 can be configured to aggregate the processed attendee information 134. As a non-exhaustive example, during an ongoing video conference, attendees 120 may applaud in response to the moderator 110 presenting the moderator information 132. In another example, in response to the moderator 110 presenting the moderator information 132, one or more of the attendees 120 may raise their hands or wave to ask a question. While aggregating the processed attendee information 134, the video conferencing server 250 can maintain a record of the types of facial expressions or body movements made by the attendees 120. Such a record may include, but is not limited to, the number of attendees 120 who applauded, the number of attendees 120 who raised their hands, the specific attendees 120 who raised their hands, and the like.
[0113] As previously mentioned, during an ongoing video conference, it may be difficult for the moderator 110 to track the responses of the participants 120. This problem becomes more severe as the number of participants 120 increases. To this end, in some embodiments, the memory 254 may store multiple environmental graphics corresponding to the recorded facial expressions or body movements of the participants 120. In some embodiments, the processor 252 may be configured to generate the multiple environmental graphics and store them in a notification in the memory 254.
[0114] As used herein, the term "ambient graphic" may refer to any visual content (e.g., an image, a series of images, a video, an animation, or a combination thereof) that, when displayed on a display (e.g., a presenter display 208 and an attendee display 208), may not completely obscure any portion of the original content displayed on the display. In some embodiments, the ambient graphic may be translucent. As used herein, the term "translucent" refers to partially or somewhat transparent or semi-transparent. In other words, if the ambient graphic is overlaid on some digital content, for example, on content displayed in the presenter GUI 130 on the presenter display 208, both the ambient graphic and the displayed content may be visible to the presenter 110 at the same time.
[0115] Figure 3A flowchart representing a process 300 for implementing a method on a video conferencing system 100 for video conferencing according to various embodiments of the present invention is described. In describing the process 300, reference will also be made to Figure 1 .
[0116] In certain embodiments, process 300 is performed by video conferencing server 250 of video conferencing system 100. Specifically, a non-transitory storage medium associated with video conferencing server 250 (eg, memory 254) stores instructions executable by processor 252 to perform process 300.
[0117] The process 300 begins at step 302, where the video conferencing server 250 receives the participant video information 134. As previously described, the participant video information 134 includes a series of images captured of the participant 120.
[0118] The process 300 proceeds to step 304. In step 304, the video conferencing server 250 performs object detection (e.g., face detection or person detection) on the participant video information 134 to detect each object (e.g., face or person) in the participant video information 134 and generates a bounding box for each object detected in the participant video information 134. Each detected object (e.g., face or person) corresponds to one of the participant 120 in the participant video information 134 and may be tagged with a unique identifier that identifies the detected object (e.g., face or person) of the participant 120.
[0119] Process 300 proceeds to step 306. At step 306, video conferencing server 250 calculates the number of participants 120 participating in the video conference based on the number of objects (eg, faces or bodies) detected in participant video information 134.
[0120] Process 300 proceeds to step 308. At step 308, video conferencing server 250 performs object recognition on each detected object to infer (i.e., predict) a facial expression, a body movement, or both of the detected object. For example, for each detected object, video conferencing server 250 may recognize a facial expression of the detected object. Video conferencing server 250 may also infer (i.e., predict) a body movement of the detected object, such as nodding, tilting the head, raising a hand, waving, pointing, clapping, changing posture when sitting or standing, random hand movements, random shoulder movements, random neck movements, changing position, or any other body movement, for each detected object.
[0121] Process 300 proceeds to step 310. At step 310, the video conference processor 250 determines whether the facial expression or body movement inferred (i.e., predicted) for each detected object in step 308 is a registered facial expression or body movement. In some embodiments, the video conference server 250 may have a record of registered facial expression or body movement types in the memory 254. Such a record may include specific types of facial expressions or body movements that the attendee 120 may perform in some context associated with the video conference. For example, if the attendee 120 wants to get the attention of the host 110, the associated record of registered body movement types may include raising a hand, waving, pointing a hand, etc. On the other hand, if the attendee 120 wants to confirm some context of the video conference, the record of registered body movement types may also include clapping, nodding, etc.
[0122] It should be noted that the above examples included in the registered body movement type record may be non-exhaustive and may include any suitable body movement without limiting the scope of the present invention. If the facial expression or body movement inferred for any detection object in step 308 is a registered facial expression or body movement, the process 300 proceeds to step 312. Otherwise, if the facial expression or body movement of each detection object is not a registered facial expression or body movement, the process 300 returns to step 302, where the video conferencing server 250 receives new participant video information 134. In some embodiments, the facial expression or body movement of the detection object corresponding to a particular participant 120 that is not in the registered facial expression or body movement type record can be filtered out by the video conferencing server 250.
[0123] After determining at step 310 that the facial expressions or body movements of at least some of the detected objects (e.g., one of the conference participants 120) are registered facial expressions or body movements, process 300 proceeds to step 312. At step 312, video conferencing server 250 determines the number of facial expressions or body movements of a particular type recognized for at least some of the detected objects (e.g., one or more of the conference participants 120). For example, video conferencing server 250 may determine the number of conference participants 120 who nodded their heads or the number of conference participants 120 who raised their hands, etc.
[0124] Process 300 proceeds to step 314, where video conferencing server 250 selects an environmental graphic corresponding to the determined facial expression or body movement type. In some embodiments, memory 254 may store a plurality of environmental graphics, each corresponding to a different type of facial expression or body movement.
[0125] Figures 4A-4D Examples of environmental graphics according to various embodiments of the present invention are described. It should be understood that Figures 4A-4DThe examples provided in are non-limiting, and other examples of environmental graphics can be implemented. Specifically, Figure 4A describes the environmental hand profile 402, Figure 4B describes the environmental ripple 404, Figure 4C Describing the flowing smoke image 406 as an environmental graphic, and Figure 4D The semi-transparent pop-up bubble image 408 is depicted as an ambient graphic.
[0126] See again Figure 3 Finally, at step 316, the video conferencing server 250 transmits the selected environment graphics (e.g., 402, 404, etc.) to the moderator client device 112. In some embodiments, the video conferencing server 250 may use the content currently displayed on the moderator display 208 (e.g., MS PowerPoint TM Presentation slides, MSWord TM The video conference server 250 may select the context hand outline 402 and present the context hand outline 402 using the content currently displayed on the moderator display 208.
[0127] In some embodiments, the video conferencing server 250 can present the transmitted ambient graphic using the moderator GUI 130 on the moderator display 208. For example, if one or more attendees 120 are clapping, the video conferencing server 250 can select the ambient ripple 404 and present the ambient ripple 404 using the moderator GUI 130 displayed on the moderator display 208. In some embodiments, the video conferencing server 250 can also transmit the selected ambient graphic to the attendee client device 112 and present it on the attendee display in a manner similar to that presented on the moderator display 208.
[0128] In some embodiments, instead of sending the environment graphics to the host client device 112 and the attendee client devices 112, the video conferencing server 250 may send a notification to the video conferencing system software installed on the host client device 112 and the attendee client devices 112. In doing so, the host client device 112 and the attendee client devices 112 may be configured to locally generate the corresponding environment graphics (e.g., 402, 404, etc.) and present the environment graphics in a similar manner as described above.
[0129] It should be noted that the environmental graphic (e.g., 402) can overlay the current content displayed on the presenter display 208 in such a manner that the environmental graphic (e.g., 402) can completely or partially cover the content being displayed. Furthermore, the environmental graphic (e.g., 402) and the portion of the content covered by the environmental graphic (e.g., 402) are recognizable and visible on the presenter display 208. Thus, the presenter 110 can still see through the environmental graphic (e.g., 402) and identify the content displayed beneath the environmental graphic (e.g., 402).
[0130] It should be noted that the environmental graphics (eg, 402 ) may be overlaid on any portion of the presenter display 208 and the attendee displays 208 .
[0131] Figure 5A A non-limiting example of an environmental graphic 402 displayed on a presenter display 208 according to various embodiments of the present invention is depicted. As shown, the environmental graphic 402 may be overlaid with current content 502 displayed on the presenter display 208. In some embodiments, the current content 502 may include a digital document (e.g., a MS PowerPoint presentation) stored in the memory 204 and / or the storage 206 on the presenter client device 112. TM Presentation slides, MS Word TM At least one of a web page, video, image, picture, etc. of a document or a digital document accessed online. For example, the digital document accessed online may include a web page, a video file, an audio file, etc.
[0132] Both the current content 502 and the features of the ambient graphic 402 are presented simultaneously on the host display 208. In doing so, the host 110 can be aware that at least one of the attendees 120 has raised his / her hand and is willing to ask a question. Because the ambient graphic 402 can be overlaid on the current content 502 in such a manner, the host 110 can still see the current content 502 underneath the ambient graphic 402. It should be noted that the ambient graphic 402 can be displayed anywhere on the host display 208.
[0133] In certain non-limiting embodiments, the moderator GUI 130 may provide an option for the moderator client device 112 to send a notification to the video conferencing server 250 if the moderator 110 wishes to answer a question from a participant who has raised their hand. In certain embodiments, this action may be triggered by a response from the moderator 110. Such a response may include selecting an appropriate option on the moderator GUI 130. The video conferencing server 250 may be configured to notify the attendee client device 112 associated with the participant 120 who has raised their hand that the moderator 110 is ready to answer the question. The notification used by the video conference may be an ambient graphic (e.g., 408). In doing so, the moderator 110 may be notified of the participant's 120 action in such a way that the current flow of the presentation being presented by the moderator 110 is not affected.
[0134] In certain non-limiting embodiments, if the moderator 110 wishes to answer a question from a participant who has raised their hand, and if there is more than one participant 120 at the same location, the video conferencing server 250 may send instructions to the participant client device 112 to select the moderator camera 216 and moderator microphone 212 that are closest to the participant 120 who has raised his / her hand. To this end, the participant client device 112 may be operable to send audio and video information associated with the participant 120 who has raised his / her hand to the participant client device 112 via the video conferencing server 250. Simultaneously, the moderator GUI 130 may provide the moderator 110 with the option to view the video and hear the associated audio associated with the participant 120 who has raised his / her hand.
[0135] Figure 5B A non-limiting example of a side display ambient graphic 404 displayed on the moderator display 208 according to at least one embodiment of the present invention is described. Because one or more of the attendees 120 have applauded, a side display ambient graphic in the form of a ripple 404 may be provided on one side of the moderator display 208. This side display ambient graphic is still semi-transparent and may not obscure the current content 502 displayed on the moderator display 208.
[0136] When displayed on the moderator display 208, the ambient graphics (e.g., 402, 404, 406, 408, etc.) can indicate to the moderator 110 that a facial expression or body movement has been exhibited by one or more of the conference participants 120 of the video conference. The type of facial expression or body movement determined by the video conferencing server 250 can determine the type of ambient graphic to be displayed on the moderator display 208.
[0137] See again Figure 1In some embodiments, the video conferencing server 250 may generate more than one environment graphic at a time, so that the moderator display 208 may display multiple environment graphics simultaneously.
[0138] In some embodiments, both the moderator client device 112 and the attendee client device 112 can receive the environment graphic. In other embodiments, only the moderator client device 112 can receive the environment graphic from the video conferencing server 250.
[0139] In some embodiments, the moderator 110 can become a participant 120, and any one of the participants 120 can become the moderator 110. In doing so, the moderator client device 112 can become a participant client device 112, and the participant client device 112 can also become the moderator client device 112. A non-limiting example of such a scenario can be that during an ongoing video conference, one of the participants 120 may want to share some content (e.g., an MS PowerPoint presentation). TM Presentation slides, MS Word TM document pages, videos, images, pictures, etc.), and the moderator 110 allows the attendees 120 to do so by selecting the appropriate option on the moderator GUI 130.
[0140] The video conferencing system 100 described herein allows attendees 120 to indicate their intent to intervene in the presentation of the presenter 110 by using body language (e.g., facial expressions or body movements), which can be captured by the attendee camera 216. The attendee 120 does not need to physically touch any device, such as by pressing a button, touching a key, clicking a computer mouse, etc.
[0141] It is contemplated that the video conferencing server 250 may determine the type of facial expression of the participant 120 or the physical movement that the participant 120 has made from the participant information 134, rather than transmitting the participant audio / sound information and the participant video information in the participant information 134 to the moderator client device 112. In some embodiments, the audio and video data in the participant information 134 may not be transmitted from the participant client device 112 or from the video conferencing server 250 to the moderator client device 112 unless the moderator client device 112 requires or requests it.
[0142] Furthermore, the video data in the attendee information 134 may not be displayed on the moderator display 208. Instead, the attendee information 134 may be analyzed by the video conferencing server 250 and "converted" into ambient graphics by the video conferencing server 250, the attendee client device 112, the moderator client device 112, or a combination thereof. The ambient graphics may include a summary of the attendee information 134. As a result, such ambient graphics may be less distracting and intrusive to the moderator 110 and attendees 120.
[0143] Furthermore, the video conferencing system 100 allows for a reduction in the amount of data transmitted between the moderator client device 112 and the video conferencing server 250, and between the moderator client device 112 and the attendee client devices 112. Transmitting and displaying the ambient graphics on the moderator display 208 can be performed more quickly than transmitting and displaying actual attendee video from at least some of the attendee client devices 112. Transmitting the ambient graphics between the video conferencing server 250 and the moderator client device 112 can require less bandwidth than transmitting actual attendee video from at least some of the attendee client devices 112.
[0144] Furthermore, when the environmental graphic is semi-transparent, the size of the environmental graphic may be as large as the size of the presenter display 208. The semi-transparent environmental graphic may not obscure any portion of the current content (e.g., 502), and when the current content (e.g., 502) and the environmental graphic (e.g., 402) are presented on the presenter display 208, both the current content (e.g., 502) and the environmental graphic may be recognizable even if the environmental graphic is as large as the presenter display 208.
[0145] Figure 6 A flow chart representing a process 600 for implementing a method on a video conferencing system 100 for video conferencing according to various embodiments of the present invention is described. In describing the process 600, reference will also be made to Figure 1 .
[0146] In certain embodiments, process 600 is performed by video conferencing server 250 of video conferencing system 100. Specifically, a non-transitory storage medium associated with video conferencing server 250 (eg, memory 254) stores instructions executable by processor 252 to perform process 600.
[0147] Process 600 begins at step 602, where the video conferencing server 250 receives participant video information from a participant client device. As previously described, the video conferencing server 250 receives participant information 134 from the participant client device 112. The participant information 134 includes participant video information 134 associated with the participant 120.
[0148] Process 600 proceeds to step 604, where the video conferencing server 250 performs object recognition on the participant video information 134 received from the participant client device 112 to identify a facial expression of each participant 120 or a physical gesture performed by each participant 120 detected in the participant video information 134. As previously described, the video conferencing server 250 may identify at least one facial expression or physical gesture in the participant information 134. Such facial expressions include at least one of laughing, smiling, and nodding, and physical gestures include at least one of nodding, tilting the head, raising a hand, waving, pointing, and clapping.
[0149] Process 600 proceeds to step 606, where video conferencing server 250 selects an ambient graphic based on the facial expression or body movement detected for each participant 112 in participant video information 134. As previously described, video conferencing server 250 may select an ambient graphic (e.g., 402, 404, 406, 408, etc.) in response to recognizing at least one type of facial expression or body movement. In some steps, if the recognized body movement is a raised hand, video conferencing server 250 may select ambient hand silhouette 402 as the ambient graphic.
[0150] Finally, process 600 proceeds to step 608, where the video conferencing server 250 transmits the selected environmental graphic to the host display 208 associated with the host client device 112 for presentation of the environmental graphic over at least a portion of the current content displayed on the host display 208 associated with the host client device 112 without obstructing the current content. As previously described, the video conferencing server 250 transmits the selected environmental graphic (e.g., environmental hand outline 402) to the host display 208 associated with the host client device 112 for presentation of the environmental graphic (e.g., environmental hand outline 402) over at least a portion of the current content (e.g., 502) displayed on the host display 208 associated with the host client device 112 without obstructing the current content.
[0151] It should be understood that the operation and functionality of the video conferencing system 100, its components and associated processes may be implemented by any one or more of hardware-based, software-based and firmware-based elements. Such operational alternatives do not limit the scope of the present invention in any way.
[0152] It should also be understood that although the embodiments set forth herein have been described with reference to specific features and structures, it is apparent that various modifications and combinations can be made without departing from these disclosures. Therefore, the specification and drawings are to be regarded only as illustrative of the implementations or embodiments discussed and their principles as defined by the appended claims, and are intended to cover any and all modifications, variations, combinations or equivalents that fall within the scope of the present invention.
Claims
1. A video conferencing server, characterized in that: include: processor; A non-transitory storage medium storing instructions executable by the processor, wherein when the instructions are executed by the processor, the video conferencing server is configured to: receiving first participant video information from the first participant client device; generating a bounding box for each participant detected from the received first participant video information; For each bounding box corresponding to each participant detected in the first participant video information, perform object recognition on the bounding box to recognize facial expressions and body movements of each participant detected in the first participant video information; determining a number of similar facial expressions and / or body movements among the detected participants; providing an environmental graphic based on the identified attributes of the facial expressions and body movements of each of the detected participants and the number of similar facial expressions and / or body movements among the detected participants; The ambient graphic is transmitted to a moderator display associated with a moderator client device for overlaying and presenting the ambient graphic on at least a portion of current content displayed on the moderator display without obscuring the current content.
2. The video conferencing server according to claim 1, wherein: The method is further configured to transmit the environmental graphic to a participant display associated with the participant client device for displaying the environmental graphic on at least a portion of the current content displayed on the participant display.
3. The video conferencing server according to claim 1, wherein: Also used for: receiving second participant video information from a second participant client device; generating a bounding box for each participant detected from the received second participant video information; For each bounding box corresponding to a participant detected in the second participant video information, perform object recognition on the bounding box to recognize facial expressions and body movements of each participant detected in the second participant video information; providing, for each of the detected participants, an environmental graphic representing attributes of the recognized facial expressions and / or body movements; The ambient graphic is transmitted to a moderator display associated with a moderator client device for overlaying and presenting the ambient graphic on at least a portion of current content displayed on the moderator display without obscuring the current content.
4. The video conferencing server according to claim 3, wherein: The method is further configured to transmit the environment graphic to a participant display associated with the second participant client device, and to display the environment graphic on at least a portion of the current content displayed on the participant display associated with the second participant client device.
5. The video conferencing server according to claim 1, wherein: The environment graphics are semi-transparent.
6. The video conferencing server according to claim 1, wherein: The first participant video information includes audio and video information associated with at least one participant.
7. The video conferencing server according to claim 6, wherein: The video information is collected by an image sensor, and the audio information is collected by a microphone.
8. The video conferencing server according to claim 1, wherein: The facial expression includes at least one of laughing, smiling or nodding, and the body movement includes at least one of nodding, tilting the head, raising hands, waving, pointing or clapping.
9. The video conferencing server according to claim 1, wherein: Providing the environmental graphic also includes generating the environmental graphic.
10. The video conferencing server according to any one of claims 1 to 9, characterized in that: The current content includes at least one of a digital document stored on the moderator client device and a digital document accessed online.
11. A video conferencing method, characterized in that: The method comprises: receiving first participant video information from the first participant client device; generating a bounding box for each participant detected from the received first participant video information; For each bounding box corresponding to each participant detected in the first participant video information, perform object recognition on the bounding box to recognize facial expressions and body movements of each participant detected in the first participant video information; determining a number of similar facial expressions and / or body movements among the detected participants; providing an environmental graphic based on the identified attributes of the facial expressions and body movements of each of the detected participants and the number of similar facial expressions and / or body movements among the detected participants; The ambient graphic is transmitted to a moderator display associated with a moderator client device for overlaying and presenting the ambient graphic on at least a portion of current content displayed on the moderator display without obscuring the current content.
12. The video conferencing method according to claim 11, wherein: The method further includes transmitting the environmental graphic to an attendee display associated with the attendee client device for displaying the environmental graphic on at least a portion of current content displayed on the attendee display.
13. The video conferencing method according to claim 11, wherein: Also includes: receiving second participant video information from a second participant client device; generating a bounding box for each participant detected from the received second participant video information; For each bounding box corresponding to a participant detected in the second participant video information, perform object recognition on the bounding box to recognize facial expressions and body movements of each participant detected in the second participant video information; providing, for each of the detected participants, an environmental graphic representing attributes of the recognized facial expressions and / or body movements; The ambient graphic is transmitted to a moderator display associated with a moderator client device for overlaying and presenting the ambient graphic on at least a portion of current content displayed on the moderator display without obscuring the current content.
14. The video conferencing method according to claim 13, wherein: The method further includes transmitting the environmental graphic to a conferee display associated with the second conferee client device for displaying the environmental graphic on at least a portion of current content displayed on the conferee display associated with the second conferee client device.
15. The video conferencing method according to claim 11, wherein: The environment graphics are semi-transparent.
16. The video conferencing method according to claim 11, wherein: The first participant video information includes audio and video information associated with at least one participant.
17. The video conferencing method according to claim 16, wherein: The video information is collected by an image sensor, and the audio information is collected by a microphone.
18. The video conferencing method according to claim 11, wherein: The facial expression includes at least one of laughing, smiling or nodding, and the body movement includes at least one of nodding, tilting the head, raising hands, waving, pointing or clapping.
19. The video conferencing method according to claim 11, wherein: Providing the environmental graphic also includes generating the environmental graphic.
20. The video conferencing method according to any one of claims 11 to 19, characterized in that: The current content includes at least one of a digital document stored on the moderator client device and a digital document accessed online.
21. A computer-readable storage medium comprising executable instructions, which, when executed by a processor, cause the processor to: receiving first participant video information from the first participant client device; generating a bounding box for each participant detected from the received first participant video information; For each bounding box corresponding to each participant detected in the first participant video information, perform object recognition on the bounding box to recognize facial expressions and body movements of each participant detected in the first participant video information; determining a number of similar facial expressions and / or body movements among the detected participants; providing an environmental graphic based on the identified attributes of the facial expressions and body movements of each of the detected participants and the number of similar facial expressions and / or body movements among the detected participants; The ambient graphic is transmitted to a moderator display associated with a moderator client device for overlaying and presenting the ambient graphic on at least a portion of current content displayed on the moderator display without obscuring the current content.
22. A computer-readable medium comprising instructions which, when executed by a processor of a video conferencing server, cause the video conferencing server to perform the method of any one of claims 11 to 20.
23. A video conferencing system, characterized in that: include: At least one first participant client device, wherein the at least one first participant client device is configured to: receiving video information of a first participant associated with the participant; generating a bounding box for each participant detected from the received first participant video information; For each bounding box corresponding to each participant detected in the first participant video information, perform object recognition on the bounding box to recognize facial expressions and body movements of each participant detected in the first participant video information; Sending the identified facial expressions and body movements of each participant detected in the participant video information to the video conferencing server; The video conferencing server is used for: determining a number of similar facial expressions and / or body movements among the detected participants; providing an environmental graphic based on the received attributes of the detected facial expressions and body movements of each participant and the number of similar facial expressions and / or body movements among the detected participants; as well as The ambient graphic is transmitted to a moderator display associated with a moderator client device for overlaying and presenting the ambient graphic on at least a portion of current content displayed on the moderator display without obscuring the current content.
24. A video conferencing system, characterized in that: include: At least one first participant client device, the at least one first participant client device comprising: An image sensor, configured to capture video information of the first participant; The at least one first participant client device sends the first participant video information to the video conference server; The video conferencing server is used for: receiving first participant video information from the at least one first participant client device; generating a bounding box for each participant detected from the received first participant video information; For each bounding box corresponding to each participant detected in the first participant video information, perform object recognition on the bounding box to recognize facial expressions and body movements of each participant detected in the first participant video information; determining a number of similar facial expressions and / or body movements among the detected participants; providing an environmental graphic based on the identified attributes of the facial expressions and body movements of each of the detected participants and the number of similar facial expressions and / or body movements among the detected participants; The ambient graphic is transmitted to a moderator display associated with a moderator client device for overlaying and presenting the ambient graphic on at least a portion of current content displayed on the moderator display without obscuring the current content.
Citation Information
Patent Citations
Methods and systems for visually chronicling a conference session
US20110029893A1
Visualizing emotions and mood in a collaborative social networking environment
US20130019187A1