Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

805 results about "Videoconferencing" patented technology

Videoconferencing is the conduct of a videoconference by a set of telecommunication technologies which allow two or more locations to communicate by simultaneous two-way video and audio transmissions. It has also been called 'visual collaboration' and is a type of groupware. Videoconferencing differs from videophone calls in that it's designed to serve a conference or multiple locations rather than individuals. It is an intermediate form of videotelephony, first used commercially in Germany during the late-1930s and later in the United States during the early 1970s as part of AT&T's development of Picturephone technology. With the introduction of relatively low cost, high capacity broadband telecommunication services in the late 1990s, coupled with powerful computing processors and video compression techniques, videoconferencing has made significant inroads in business, education, medicine and media. Like all long distance communications technologies, by reducing the need to travel, which is often carried out by aeroplane, to bring people together the technology also contributes to reductions in carbon emissions, thereby helping to reduce global warming.

Face-translator: end-to-end system for speech-translated lip-synchronized and voice preserving video generation

A neural end-to-end system is provided for the face and voice preserving translation of videos. The system is a pipeline of multiple models that produces a video of the original speaker speaking in the target language with modified lip movement to match the target speech, while preserving emphases and prosody of the original speech, and voice characteristics of the original speaker. The pipeline starts with automatic speech recognition including emphasis detection, followed by the translation model. The translated text is then synthesized by a Text-to-Speech model that recreates the original emphases in the target sentence. The resulting synthetic speech is then converted back to the original speakers' voice using a voice conversion model. Finally, to synchronize the lips of the speaker with the translated audio, a generative model generates frames of adapted lip movements which are combined with the audio to produce the final output. The disclosure further describes several use-cases and configurations that apply these techniques to video conferencing, dubbing, low-bandwidth transmission, speech enhancement and assistive technology for the hearing impaired.
Owner:WAIBEL ALEXANDER

Generating A Unified Virtual Background Image For Multiple Video Conference Participants

A unified virtual background image is generated for multiple participants of a video conference to create an immersive conference experience based on its use within video streams of those multiple participants. Generative artificial intelligence software associated with a conferencing system obtains input associated with a video conference. The generative artificial intelligence software generates a virtual background image based on the input. The virtual background image is then for use within multiple participant video streams during the video conference
Owner:ZOOM COMMUNICATIONS INC

Establishing a video conference during a phone call

Some embodiments provide a method for initiating a video conference using a first mobile device. The method presents, during an audio call through a wireless communication network with a second device, a selectable user-interface (UI) item on the first mobile device for switching from the audio call to the video conference. The method receives a selection of the selectable UI item. The method initiates the video conference without terminating the audio call. The method terminates the audio call before allowing the first and second devices to present audio and video data exchanged through the video conference.
Owner:APPLE INC

Real-time document collaborative editing and synchronizing method in video conference

The invention relates to the technical field of video conferences, and discloses a real-time document collaborative editing and synchronizing method in a video conference, which is used for solving the problem of insufficient concurrent editing conflict solving mechanism in the traditional method. The method comprises the following steps: firstly, receiving document editing operations submitted by a plurality of participants through a terminal, and establishing an operation sequence; comparing the timestamps with the position information, identifying concurrent operation groups as conflict sets, and arranging the conflict sets according to a time sequence; types and contents are extracted from the conflict set, semantic matching is carried out, matching operation is converted into a compatible sequence, and subsequent position offset is adjusted; distributing the compatible sequence to a synchronous buffer area, updating a server document state, and generating an incremental update package; a packet is pushed to the online terminal, application changes are buffered locally, and sequence playback is provided for a new participant; the monitoring terminal confirms the signal, re-deduces the unconfirmed change, and circularly checks the Hash consistency; according to the method, conflict resolution is optimized through semantic matching and offset adjustment, and efficient synchronization is achieved.
Owner:WUHAN HONGYANGUO TECH CO LTD

Video conference multi-modal data alignment method and device based on causal mask, equipment and medium

The invention discloses a video conference multi-modal data alignment method and device based on a causal mask, equipment and a medium, and relates to the technical field of computers, and the method comprises the steps: carrying out the feature extraction and fusion of an original audio, an original video stream and an original document in an online video conference, time sequence division is carried out based on the obtained multi-modal fusion features to obtain a triple time sequence window; determining an initial weight value corresponding to the triple time sequence window, and performing normalization adjustment on the initial weight value by using a preset constraint condition to obtain an adjusted weight; indexing a preset time sequence offset matrix by using a speaking identifier of a speaking party, correcting an original time sequence of the triple time sequence window based on an indexing result, and determining a target attention result corresponding to the triple time sequence window by using a preset causal mask mechanism, and performing multi-level alignment fusion on the multi-modal fusion features based on the target attention result to obtain a multi-modal alignment result. The precision of the multi-mode alignment technology is improved, and future information leakage is avoided.
Owner:SHANDONG INSPUR SCI RES INST CO LTD

Universal Identity Verification for Video Conferencing

Systems, methods, and apparatuses are described for verifying a user identity in a video conference. A computing device may receive user data and a plurality of security parameters associated with accessing a video conference based on a confidentiality level of the video conference. The computing device may generate a security code that is encoded with user data. The computing device might cause the security code to be displayed on the mobile device for a predetermined time period. The computing device may receive an indication that the first device scanned the security code by using a camera. To verify the identity of a user, the computing device may decode the security code, compare the decoded user data of the decoded security code and expected user data associated with the video conference. The computing device may determine the authenticity of a user video and allow access to the video conference.
Owner:CAPITAL ONE SERVICES LLC

Video conference echo suppression method based on low-delay adaptive learning model

The invention relates to a video conference echo suppression method based on a low-delay adaptive learning model. The method comprises the following steps: acquiring acoustic scene noise features, and matching a lightweight LSTM neural network model by using the acoustic scene noise features; the method comprises the following steps: collecting real-time sound signals of a video conference, preprocessing and extracting multi-modal features of the real-time sound signals; performing adaptive filtering by using an NLMS algorithm to estimate an echo path between a near-end microphone signal and a far-end reference signal of the real-time sound signal, performing convolution on the echo path and the far-end reference signal to generate an echo estimation signal, and subtracting the echo estimation signal from the near-end microphone signal to obtain a preliminary residual signal; learning a gain by using the lightweight LSTM neural network model, and further suppressing the residual signal by using the gain to obtain a time domain enhanced signal; comfortable noise is introduced, and incremental learning is adaptively triggered. And efficient linear and nonlinear echo cancellation is supported.
Owner:CHINA LIFE INSURANCE CO LTD

Image fuzzy detection method based on fusion of frequency domain analysis and deep learning

PendingCN120807454AImage enhancementImage analysisOptical flowModal method
The invention provides an image fuzzy detection method based on frequency domain analysis and deep learning fusion. The method comprises the following steps: S1, frequency domain feature extraction and quantification; s2, spatial domain feature extraction and modeling; s3, carrying out multi-modal feature fusion; s4, joint optimization and post-treatment are carried out; s5, outputting and verifying; through complementarity design of frequency domain and deep learning, complex fuzzy detection requirements of static images and video streams are covered, high efficiency and reliability are verified in industrial quality inspection, video conferences and other scenes, energy attenuation characteristics caused by global blur are accurately captured through frequency domain analysis, motion blur and out-of-focus blur are effectively distinguished, and the method is suitable for large-scale popularization and application. According to the method, local texture degradation of deep learning network modeling, dynamic track abnormity analysis of an optical flow network and complex scenes covering static images and video streams are realized through a bidirectional feature fusion mechanism, the mAP of mixed fuzzy detection is effectively improved compared with a single-mode method, and dynamic fuzzy and static out-of-focus fuzzy are effectively distinguished.
Owner:YIREN (SHANGHAI) TECH CO LTD

An audio and video signal processing system and a video conference terminal device using the same

The application provides an audio and video signal processing system and a video conference terminal device using the same, and relates to the technical field of Internet.The system comprises an audio and video processing module, a signal preprocessing module, a signal delay measurement module, a signal synchronization processing module and a master control module, the audio and video processing module is used for separating audio signals and video signals in input audio and video data signals, the signal preprocessing module is used for preprocessing the audio signals, the signal delay measurement module is used for measuring relative delay amounts of the audio signals and the video signals, the signal synchronization processing module is used for performing delay compensation according to the relative delay amounts, realizing the synchronization of audio and video display, and the video conference terminal device further comprises a microphone, a camera, a display screen and a sound box.The system and the video conference terminal device can test the relative delay time of audio and video signals, perform delay compensation, ensure the synchronization of output audio and video pictures, and improve the video conference quality and user experience.
Owner:BEIJING ZOBO ELECTRONIC TECH CO LTD

Automatically generating feedback about content shared during a videoconference

Feedback can be automatically generated for visual content presented during a videoconference. For example, a system can receive a user selection of a file containing visual content to be presented to members of a video conference, wherein the visual content includes pages that are to be sequentially presented to the members during the video conference. The system can then facilitate presentation of the visual content to the members of the video conference. The system can also obtain metadata associated with at least one page of the visual content presented during the video conference, determine feedback about the at least one page by analyzing the metadata, and provide the feedback to an editor of the visual content.
Owner:ZOOM COMMUNICATIONS INC

Video conference multi-mode real-time abstract generation method

The invention relates to the technical field of video conference data processing, and discloses a video conference multi-mode real-time abstract generation method, which comprises the following steps of: synchronously acquiring an audio stream, a video stream and a text chat record of a conference, and converting the audio stream, the video stream and the text chat record into a time-aligned text, a key frame sequence and effective chat content through preprocessing; then text semantic features, visual scene features and interactive intention features are extracted, cross-modal correlation analysis is carried out through a multi-modal fusion model, and a fusion feature set is generated; and based on the identified core issue, the key conclusion and the action item, performing structured organization according to the time sequence and the importance degree to form a real-time abstract and performing dynamic updating. According to the method, multi-dimensional information is integrated, the one-sidedness problem of a traditional single-mode abstract is solved, the integrity, accuracy and timeliness of the abstract are improved, participants are assisted in mastering key points of a conference in real time, and the conference efficiency and decision-making quality are improved.
Owner:SHENZHEN ZHONG XUN WANG LIAN SCI & TECH CO LTD

Communication method based on router bridging and router

The invention discloses a communication method based on router bridging and a router, and relates to the technical field of wireless communication. The method and the device are used for solving the problems of link stability and switching efficiency of video conference services in a multi-interference environment. A master control node scans an interference source in an environment, detects the link delay and load of bridging equipment, calculates a resource allocation priority based on interference intensity distribution and node performance, screens low-delay nodes to form a core backhaul candidate set, and marks a continuous high-interference channel as a forbidden frequency band; when the video conference flow exceeds a threshold value, dynamically allocating a DFS channel by utilizing reinforcement learning and combining with a forbidden frequency band, and establishing a pre-synchronization security key data channel; on the basis of a terminal motion state and a service label, a moving track is predicted through return link quality and a motion vector, and target node pre-binding and service data flow mirror image caching of the high-speed mobile terminal are realized; in the switching stage, a target node injects cache data and verifies continuity, rolls back a source link when abnormity occurs, and dynamically adjusts a threshold value and a track parameter.
Owner:HANGZHOU FENGHENG ELECTROMECHANICAL

Systems and methods for automatic speaker tracking for video conferences based on voiceprint, lip and body motion detection

PCT designated stageWO2025213354A1Television conference systemsCharacter and pattern recognitionSpeaker trackingParticipant identifier
The present disclosure provides methods, systems, and mediums for identifying an active speaker within an online conferencing session. The method comprises the steps of receiving an audio / video stream from a client device during an online conferencing session. Upon a participant speaking: identifying, a voiceprint representing the participant, wherein the voiceprint represents one or more unique vocal characteristics of the participant. Detecting a spatial position of the participant based upon movement of one or more markers of interest on the first speaker. Generating a mapping between the voiceprint and the spatial position of the participant. Using the voiceprint and the spatial position of the participant to identify the participant as a first active speaker. The method further comprises generating instructions to adjust positioning of a camera so that the first active speaker is centered within a video stream.
Owner:RINGCENTRAL INC +1

Privacy preserving online video recording

Systems and methods are provided herein for only including portions of a user's environment that have been approved by a user in a video conference while excluding portions that have not been approved. This may be accomplished by a device receiving a policy identifying one or more approved objects of a scene of a video stream. The device may then generate a filtered video stream by only including portions of the scene that comprise the one or more objects that were approved by the policy in the filtered video stream. The filtered video stream may be combined with other video streams to generate a video conference that is transmitted and / or stored by one or more devices participating in the video conference.
Owner:ADEIA GUIDES INC

Accurate camera-display calibration for telepresence videoconferencing

ActiveUS12462427B2Image enhancementImage analysisComputer graphics (images)Checkerboard pattern
Techniques of calibrating telepresence videoconferencing displays and cameras include providing 6 DoF camera locations and orientation vectors in display reference frame based on a plurality of images that indicate specified mirror-plane points of a mirror and specified reflected display plane points. In some implementations, the specified mirror-plane points are located at fiducial markers printed on the mirror. In some implementations, the specified reflected display plane points are located in a checkerboard pattern of fiducial markers on the display.
Owner:GOOGLE LLC

Data processing method and device, video conferencing system, storage medium

A method for data processing and, a device, a video conferencing system, and a computer readable storage medium. The method may include, acquiring audio information, motion feature information of a human face, and a target human face image; adjusting the motion feature information according to the audio information to acquire target motion feature information of the human face; and generating a target human face image sequence corresponding to the audio information according to the target motion feature information and the target human face image.
Owner:ZTE CORP

Video conference system network inspection method and device based on multiple protocols

The invention relates to the technical field of video communication, and discloses a video conference system network inspection method and device based on multiple protocols. The method comprises the following steps: constructing a network topology structure reflecting the health state of a video conference transmission link based on multi-protocol state data and equipment connection information; mapping the screened service flow data to a network topology structure to obtain a service transmission path; threshold judgment is carried out on the service detection data and the multi-protocol state data, and a network stability mode of the video conference system is determined; when the network stability mode is an unstable network, respectively determining an affected service transmission path and an unstable network cause, and obtaining change management data related to the affected service transmission path; and generating a network stability result of the video conference system based on the unstable network cause, the affected service transmission path and the change management data. The evaluation accuracy of the network fault of the video conference system can be improved.
Owner:WENZHOU ELECTRIC POWER BUREAU

Determining prioritized participants during a video conference

A system may receive, during a video conference, video feeds corresponding to conference participants. The system may determine a prioritized participant from the conference participants by using a machine learning model and historical conference participant information. In some implementations, the system may receive input from a conference participant that causes the conference participant to be a prioritized participant (e.g., the conference participant may self-select). The system may emphasize, in a graphical user interface (GUI) during the video conference, the video feed corresponding to the prioritized participant. In some implementations, determining the prioritized participant may include accessing a record including the historical conference participant information. The record may indicate that the prioritized participant and a participant using the GUI have a connection in an organization. In some implementations, the record may indicate that the prioritized participant used sign language during another video conference.
Owner:ZOOM COMMUNICATIONS INC

Machine-learning assisted acoustic echo cancelation

Example methods and systems provide machine-learning assisted acoustic echo cancellation (AEC). The AEC can be used, as an example, to improve the audio quality for online audio and video conferences. A system according to this disclosure includes a pre-trained, machine-learning, AI model designed to detect, in real time, a unitary voice signal, or a signal representing the speech of a single speaker as opposed to that of multiple speakers. A digital signal processing (DSP) algorithm can then detect the echo state, for example, whether distortion results primarily from an echo. Based on these characteristics, the system can, alternatively and automatically apply either a default mode of AEC to the audio signal, or apply a more aggressive mode of AEC.
Owner:ZOOM COMMUNICATIONS INC

E1 / IP video conference protocol conversion device and system thereof

The invention relates to the technical field of protocol conversion, in particular to an E1 / IP video conference protocol conversion device and system, and the system comprises a frame structure mapping module, a signaling adaptation module, a media stream packaging module, a buffer area control module and a synchronous coordination module. According to the invention, the synchronization mark and the load field of the physical layer framing structure are extracted, the maximum transmission unit parameter is combined, the frame packet boundary is accurately identified, the cross-network data positioning accuracy is improved, the call control mapping relation is established based on key field identification, and the session continuity between heterogeneous signaling is enhanced; the audio and video sampling rate and timestamp fundamental frequency are matched and verified, the consistency of media stream transmission time delay is guaranteed, interface throughput and time slot occupation changes are dynamically monitored, a buffer regulation and control instruction is generated, and the data transmission stability under burst flow is improved; the synchronization accuracy of the protocol conversion system in a heterogeneous network environment is realized, and the reliability of real-time services such as video conferences and the like is enhanced.
Owner:BAOSHENG (CHINA) TECH IND CO LTD

Prior model for gaussian splatting - based avatars

A system and method for rendering three-dimensional digital representations of users are disclosed. The described approach addresses challenges in generating photorealistic avatars with minimal input data by utilizing a deep neural network (DNN)-based prior model. The prior model is trained to identify features and generate a canonical template representing average user characteristics. During an enrollment phase, personalized offsets are determined for individual users based on their distinguishing features. These offsets, combined with the canonical template, enable the generation of high-quality, real-time 3D avatars from a single audio or visual input. The avatars can be animated based on user signals, such as expressions or sounds, captured by input devices. Applications include virtual reality, gaming, video conferencing, and entertainment. The system reduces computational resource requirements while improving rendering speed and fidelity, enabling efficient avatar generation and animation in communication sessions.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Gaze-based video conference prompts

Techniques for gaze-based video conference prompts are described that leverage eye-tracking algorithms within a video conference setting to perform a variety of functionality. For instance, a computing device displays a user interface of a video conference session that includes representations of video conference attendees. The computing device receives video content that depicts a user of the computing device and uses eye tracking techniques to determine a gaze location of the user based on the video content. The gaze location corresponds to one of the representations of an attendee. The computing device then communicates a prompt to one or more of the attendees that indicates the gaze location. In another example, the computing device detects a behavior of the user, e.g., a scanning behavior. Responsive to detection of the behavior, the computing device performs an action within the user interface, such as to display a roster view of the attendees.
Owner:MOTOROLA MOBILITY LLC

Video transmission optimization method and system based on display equipment capability

The invention provides a video transmission optimization method and system based on the capability of display equipment. The method comprises the following steps: identifying a maximum display parameter of a display device currently connected with the terminal; obtaining the maximum decoding capability of the terminal according to the maximum display parameter of the display device and the decoding capability of the terminal; and when the terminal and the remote end carry out video conference calling, the two parties determine a video receiving format according to the maximum decoding capability, the coding capability of the remote end and the network bandwidth in a connection establishment process. According to the video transmission optimization method and system based on the display equipment capability, a more accurate and flexible video transmission scheme can be provided by considering the actual capability of the display equipment, and the overall performance and user experience of a video conference are remarkably improved.
Owner:上海赛连信息科技有限公司

Automatic video conference framing using multiple cameras

A method is described in which a first video stream of a scene is received (602) from a first camera (102) and a second video stream of the scene is received (604) from a second camera (104a, 104b). A participant of interest is identified (606) in a first video stream and mapped (608) from the first video stream to a second video stream. A first gesture of the participant of interest relative to the first camera is determined (610), and a second gesture of the participant of interest relative to the second camera is determined (612). The first video stream and the second video stream are rated (614) based on comparing the first gesture and the second gesture. The processor automatically switches between sending one of the first video stream or the second video stream to the display system based on a respective one of the first video stream or the second video stream having a higher rating (616).
Owner:HEWLETT PACKARD DEVELOPMENT COMPANY LP

Video conference system based on augmented reality

The invention relates to the technical field of video conferences, in particular to a video conference system based on augmented reality, which comprises a depth information acquisition and processing unit, a space anchor point synchronization unit, a virtual space mapping unit, a virtual object superposition and rendering unit, a multi-modal interaction unit and a real-time processing and transmission unit. Three-dimensional point cloud data output by the depth information acquisition and processing unit is input to the space anchor point synchronization unit and the virtual space mapping unit, and global coordinate system data output by the space anchor point synchronization unit is input to the virtual space mapping unit and the virtual object superposition and rendering unit. Shared virtual space data output by the virtual space mapping unit is input into the virtual object overlapping and rendering unit, and an interaction command output by the multi-mode interaction unit is input into the virtual object overlapping and rendering unit and the real-time processing and transmitting unit. Through the multi-mode interaction unit, the interaction mechanism becomes visual, the gesture and voice recognition response delay is low, and the accuracy is high.
Owner:BEIJING DEHUI TIANCHEN TECH CO LTD

High-definition video low-delay compression method and system based on adaptive adjustment

The invention discloses a high-definition video low-delay compression method based on self-adaptive adjustment, and aims to solve the problems of insufficient self-adaptive adjustment capability, too high encoding and decoding delay and low code rate and network bandwidth matching degree in the existing high-definition video compression technology. The method comprises the following steps: firstly, acquiring a high-definition video stream and a real-time network bandwidth parameter, extracting intra-frame texture complexity, inter-frame motion intensity and scene switching frequency characteristics, and constructing a two-dimensional adaptive compression strategy in combination with a real-time bandwidth; low-delay coding is executed by adopting parallel frame fragmentation and simplified adaptive arithmetic coding, adaptive code rate control is realized by combining code rate overflow early warning and intra-frame key region priority bit allocation, and finally, fast inverse transformation, quantization recovery and motion compensation optimization are completed at a decoding end based on strategy parameters synchronously transmitted at a coding end so as to reconstruct a high-definition video frame. According to the invention, the adaptive adjustment capability of video compression is improved, low time delay of coding and decoding is realized, and the method is suitable for real-time high-definition video transmission scenes such as remote monitoring and video conferences.
Owner:CHANGSHA CHAOCHUANG ELECTRONICS TECH

Systems and methods for automatic speaker tracking for video conferences based on voiceprint, lip and body motion detection

PendingUS20250317317A1Special service provision for substationCharacter and pattern recognitionSpeaker trackingParticipant identifier
The present disclosure provides methods, systems, and mediums for identifying an active speaker within an online conferencing session. The method comprises the steps of receiving an audio / video stream from a client device during an online conferencing session. Upon a participant speaking: identifying, a voiceprint representing the participant, wherein the voiceprint represents one or more unique vocal characteristics of the participant. Detecting a spatial position of the participant based upon movement of one or more markers of interest on the first speaker. Generating a mapping between the voiceprint and the spatial position of the participant. Using the voiceprint and the spatial position of the participant to identify the participant as a first active speaker. The method further comprises generating instructions to adjust positioning of a camera so that the first active speaker is centered within a video stream.
Owner:RINGCENTRAL INC