Audio or video conferencing end-system

WO2025074099A4PCT designated stage expired Publication Date: 2025-06-05AVOS TECH LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/GB2024/052536
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-10-02
Filing Date
2024-10-02
Publication Date
2025-06-05

AI Technical Summary

Technical Problem

In video conferencing systems, it is challenging to accurately identify the current speaker in multi-participant conferences, especially in environments where multiple people are present in the same room.

Method used

An end-system is developed that uses one or more processors to analyze audio and image data from multiple microphones and imaging devices. The system forms a signature for known participants based on their audio and video characteristics and uses this signature to identify the current speaker in real-time.

Benefits of technology

The system effectively identifies the current speaker in video conferences, even in noisy environments or when participants move within the room, by correlating audio and video data to accurately attribute speech to the correct participant.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure GB2024052536_05062025_PF_FP_ABST
    Figure GB2024052536_05062025_PF_FP_ABST
Patent Text Reader

Abstract

An end-system for conducting a video conference, the end-system being communicatively connectable with one or more microphones configured to capture audio data from multiple participants during the conference and one or more imaging devices configured to capture image data of multiple participants during the conference, the end-system comprising one or more processors configured to: at a first time, receive first data and second data, the first data and the second data being captured by or derived from data captured by one of (i) the one or more microphones and (ii) the one or more imaging devices, the first time corresponding to a time when a participant of the multiple participants is speaking; analyse one of the first data and the second data to detect the participant speaking at the first time to be a known participant; form a signature for the known participant from the other of the first data and the second data; at a second time, receive one or more of third data and fourth data, the third data and the fourth data being captured by or derived from data captured by one of (i) the one or more microphones and (ii) the one or more imaging devices, the second time corresponding to a time when a participant of the multiple participants is speaking; and in dependence on the signature for the known participant, detect the participant speaking at the second time to be the known participant from one or more of the third data and the fourth data.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] AUDIO OR VIDEO CONFERENCING END-SYSTEM

[0002] FIELD OF THE INVENTION

[0003] The present invention relates to audio or video conferencing systems, in particular to the determination of the identity of a participant during a conference.

[0004] BACKGROUND

[0005] Video conferencing systems typically comprise two or more endpoints, each endpoint being configured to communicate with other endpoints via a data or telephony network, or other connections, for implementing a two-way or multiway video and / or audio conference call.

[0006] An endpoint, which may also be referred to as a video end-system, is configured to enable users to communicate with other users over a network by sending data streams, such as audio, video and other content streams (for example, data files, text, or screen sharing) from one endpoint to another in a two-party conference call. For a multiway conference call, the data streams may be transmitted to a centralised or distributed conferencing network where they are switched or transcoded, allowing multiple participants to communicate and share information. A video end-system may comprise a single purpose endpoint device whose main function is to support video and / or audio conference calls. The video end-system may comprise separate video conferencing hardware, which may include one or more video monitors, video cameras, audio microphones, audio speakers and a controller for enabling users to control the video end-system. These devices will typically be connected permanently to the video end-system and remain part of the video end-system’s configuration.

[0007] Video conferencing systems typically comprise two or more endpoints, whereby each endpoint is configured to communicate with other endpoints through a data or telephony network or other connections for implementing a two-way or multiway video and / or audio conferencing call. On some occasions, the audio stream from a conference may be routed to a speech transcription service. There are generally two types of speech transcription service. A first type simply converts any speech that is detected into text in real time and feeds that text back to a user who may provide it as subtitles in the conference or save it for later review. A second type records the audio and then processes it offline using much deeper analysis, such as artificial intelligence techniques, to identify the speaker of each speech segment so that the speakers’ identities may be tagged in the text output. This second type of transcription is generally too complex to be performed in real time and can need guidance to accurately identify speakers.

[0008] In an audio or video conference of individuals, the identity of the current speaker is readily available, since every participant has their own microphone and audio stream. When a conference is held in one or more conference rooms occupied by multiple persons however, the identity of the current speaker cannot be readily established.

[0009] Figure 1 shows a diagram of a typical video conferencing system 100 with audio transcription according to the prior art. A video end-system 101 is in a conference with another video end-system 102 via the internet 109 using a cloud conference controller 107 and a feed to a real-time transcription service 108. In this example, each video endsystem 101 , 102 comprises a respective endpoint 103, 104 and a respective camera 105, 106. Video and audio data are sent between the video end-system 101 and the conference controller 107 and also between the video end-system 102 and the conference controller 107. The conference controller 107 can optionally send a live audio stream to the transcription service 108 which detects speech and transcribes it to text. The transcription service can send this stream back to the conference controller 107 or to another destination 110, which may for example be a meeting archive.

[0010] Traditionally, this text stream contains no information about the identity of the speaker for each speech segment. The resulting transcription may therefore be confusing, as it is not clear which user was speaking at a given time.

[0011] It is desirable to develop an improved mechanism for establishing the identity of a participant in a video conference and optionally applying the identity to operations involving the audio and video data during the conference. SUMMARY OF THE INVENTION

[0012] According to a first aspect, there is provided an end-system for conducting a video conference, the end-system being communicatively connectable with one or more microphones configured to capture audio data from multiple participants during the conference and one or more imaging devices configured to capture image data of multiple participants during the conference, the end-system comprising one or more processors configured to: at a first time, receive first data and second data, the first data and the second data being captured by or derived from data captured by one of (i) the one or more microphones and (ii) the one or more imaging devices, the first time corresponding to a time when a participant of the multiple participants is speaking; analyse one of the first data and the second data to detect the participant speaking at the first time to be a known participant; form a signature for the known participant from the other of the first data and the second data; at a second time, receive one or more of third data and fourth data, the third data and the fourth data being captured by or derived from data captured by one of (i) the one or more microphones and (ii) the one or more imaging devices, the second time corresponding to a time when a participant of the multiple participants is speaking; and in dependence on the signature for the known participant, detect the participant speaking at the second time to be the known participant from one or more of the third data and the fourth data.

[0013] The one or more processors may be configured to analyse one of the first data and the second data to detect the participant speaking at the first time to be the known participant using a first detection method and form the signature for the known participant from the other of the first data and the second data by analysing the other of the first data and the second data using a second detection method different to the first detection method.

[0014] In dependence on the signature for the known participant, the one or more processors may be configured to detect the participant speaking at the second time to be the known participant from the third data or the fourth data. The one of the third data and the fourth data used to detect the participant speaking at the second time to be the known participant may correspond to (e.g. may be the same type of data as) the other of the first data and the second data that was used to form the signature. The one or more processors may use the second detection method to detect the participant speaking at the second time to be the known participant from the third data or the fourth data in dependence on the signature.

[0015] The first detection method and / or the second detection method may be one of a facial recognition method, a motion detection method, a colour detection method, an image segmentation method and a voice recognition method.

[0016] The one or more processors may be configured to form the signature when the detection of the known participant from the first data or the second data is over a threshold certainty.

[0017] The one or more processors may be configured to detect the participant speaking at the first time to be the known participant from image data received from the one or more imaging devices. The image data may be processed to form the first data and / or the second data.

[0018] The one or more processors may be configured to detect the participant speaking at the first time to be the known participant from audio data received from the one or more microphones. The audio data may be processed to form the first data and / or the second data.

[0019] The one or more processors may be configured to detect the participant speaking at the second time to be the known participant from image data received from the one or more imaging devices in dependence on the signature. The image data may be processed to form the third data and / or the fourth data.

[0020] The one or more processors may be configured to detect the participant speaking at the second time to be the known participant from audio data received from the one or more microphones in dependence on the signature. The audio data may be processed to form the third data and / or the fourth data.

[0021] The one or more processors may be configured to annotate one or more segments of a video stream corresponding to the first time and / or the second time with a participant identifier for the known participant. This may be used to identify the known participant, for example to one or more other participants, as the currently speaking participant. The second time may be a time when the known participant is speaking. The first data and the second data may be data captured or derived from data captured by the one or more microphones and the one or more imaging devices respectively. The first data and the second data may be data derived from data captured by one of (i) the one or more microphones and (ii) the one or more imaging devices. For example, data captured by these devices may be processed to form the first data and the second data. The first data and the second data may be data captured by or derived from data captured by the same one of (i) the one or more microphones and (ii) the one or more imaging devices. For example, the first data and the second data may each be data derived from raw data captured by the one or more imaging devices. The first data and the second data may be data derived from raw data captured by the one or more microphones. The first data may be derived from the raw data (for example, the same raw data) using a different data processing method to the second data.

[0022] The second time may be later than the first time. The second time may be within a predefined period (for example, less than one day) of the first time. The first and second times may be time periods. The first and second times may be times in the same conference.

[0023] The one or more processors may be configured to cause a display in communication with the end-system to display a participant identifier for the known participant when the known participant is detected to be speaking.

[0024] The participant identifier may be an arbitrary identifier, or the participant identity may comprise a name of the participant.

[0025] The participant identifier for the known participant may be associated with the formed signature, and optionally stored by the end-system at a memory of or associated with the end-system.

[0026] The one or more processors may be configured to: convey segments of the video stream to a conferencing server; and convey the participant identity to the conferencing server for annotation of one or more of the segments of the video stream corresponding time(s) when the known participant is detected to be speaking with the participant identity.

[0027] The one or more processors may be configured to store the signature at a memory of or associated with the end-system. The one or more processors may be configured to maintain a history of signatures and corresponding known participants.

[0028] If the participant detected to be speaking at the first time is not a known participant, the one or more processors may be configured to form a new signature for the new participant and store an association between the new participant and the new signature.

[0029] The one or more processors may be configured to: convey audio data captured by the one or more microphones to a speech transcription service; and annotate a transcription of the audio data corresponding to the first time and the second time received from the speech transcription service with a participant identifier corresponding to the known participant.

[0030] The one or more processors may be configured to: convey audio data captured by the one or more microphones to a speech transcription service; and convey the participant identity corresponding to the known participant to a conferencing server for annotation of a transcription of the audio data corresponding to the first time and the second time received by the conferencing server from the speech transcription service with the participant identifier.

[0031] The one or more processors may be configured to detect one or more segments of the conference comprising speech by one or more participants (for example, the known participant) in which the speech is irrelevant to a topic of speech preceding and / or following the respective segment, wherein the one or more processors are configured to omit the speech from the audio data conveyed to the video conferencing server and / or speech transcription service.

[0032] The one or more processors may be configured to detect one or more verbal or nonverbal events from the audio data captured by the one or more microphones and / or the image data captured by the one or more imaging devices and produce a textual description of the one or more events for inclusion in a transcript of the conference.

[0033] The one or more processors may be configured to detect one or more participant emotions from the audio data captured by the one or more microphones and / or one or more participant expression from the image data captured by the one or more imaging devices and produce a textual description of the one or more participant emotions and / or expressions for inclusion in a transcript of the conference.

[0034] The one or more processors may be configured to detect one or more participant actions from the image data captured by the one or more imaging devices and produce a textual description of the one or more participant actions for inclusion in a transcript of the conference.

[0035] The one or more processors may be configured to, at the first time, receive first audio data captured by multiple microphones, the first audio data corresponding to speech by a participant; perform a correlation operation between the first audio data received from each of the microphones to generate first signature data representing a first spatial signature for audio signals corresponding to the first audio data; store an association between a participant identity and the first spatial signature; at a second time subsequent to the first time, receive second audio data captured by the microphones, the second audio data corresponding to speech by a participant; perform a correlation operation between the second audio data received from each microphone to generate second signature data representing a second spatial signature for audio signals corresponding to the second audio data; determine that the second spatial signature matches the first spatial signature having the stored association with the participant identity; and perform an operation using the second audio data in dependence on the participant identity.

[0036] According to a second aspect, there is provided a method of conducting a video conference, the method comprising: at a first time, receiving first data and second data, the first data and the second data being captured by or derived from data captured by one of (i) one or more microphones and (ii) one or more imaging devices, the first time corresponding to a time when a participant of multiple participants is speaking; analyse one of the first data and the second data to detect the participant speaking at the first time to be a known participant; form a signature for the known participant from the other of the first data and the second data; at a second time, receive one or more of third data and fourth data, the third data and the fourth data being captured by or derived from data captured by one of (i) the one or more microphones and (ii) the one or more imaging devices, the second time corresponding to a time when a participant of the multiple participants is speaking; and in dependence on the signature for the known participant, detect the participant speaking at the second time to be the known participant from one or more of the third data and the fourth data.

[0037] According to another aspect, there is provided an end-system for conducting an audio or video conference, the end-system being communicatively connectable with multiple microphones configured to capture audio data from multiple participants during the conference and comprising one or more processors configured to: at a first time, receive first audio data captured by the microphones, the first audio data corresponding to speech by a participant; perform a correlation operation between the first audio data received from each of the microphones to generate first signature data representing a first spatial signature for audio signals corresponding to the first audio data; store an association between a participant identity and the first spatial signature; at a second time subsequent to the first time, receive second audio data captured by the microphones, the second audio data corresponding to speech by a participant; perform a correlation operation between the second audio data received from each microphone to generate second signature data representing a second spatial signature for audio signals corresponding to the second audio data; determine that the second spatial signature matches the first spatial signature having the stored association with the participant identity; and perform an operation using the second audio data in dependence on the participant identity.

[0038] The one or more processors may be configured to: convey the first audio data and / or the second audio data to a speech transcription service; and annotate a transcription of the first audio data and / or the second audio data received from the speech transcription service with the participant identity.

[0039] The one or more processors may be configured to: convey the first audio data and / or the second audio data to a speech transcription service; and convey the participant identity to a conferencing server for annotation of a transcription of the first audio data and / or the second audio data received by the conferencing server from the speech transcription service with the participant identity.

[0040] The one or more processors may be configured to store an annotated transcription of the conference for future analysis or reference. The annotated transcription may be annotated with the identities of each participant as they are speaking.

[0041] The conference may be an audio conference. The conference may be a video conference.

[0042] The end-system may comprise one or more cameras configured to capture a video stream of the conference.

[0043] The one or more processors may be configured to determine the participant identity (or a part thereof) from an image or video segment of the video stream corresponding to (e.g. time-aligned with) the first audio data. For example, the one or more processors may be configured to implement a facial recognition algorithm to determine the name of the participant from the image or video segment.

[0044] The one or more processors may be configured to annotate one or more segments of the video stream corresponding to (e.g. time-aligned with) the first audio data and / or the second audio data with the participant identity.

[0045] The one or more processors may be configured to: convey segments of the video stream corresponding to the first audio data and / or the second audio data to a speech to a conferencing server; and convey the participant identity to the conferencing server for annotation of one or more of the segments of the video stream corresponding to the first audio data and / or the second audio data with the participant identity. The conferencing server may then transmit the annotated segment(s) of the video stream to another endsystem participating in a video conference with the video end-system.

[0046] The one or more processors may be configured to store an annotated video stream of the video conference for future analysis or reference. The annotated video stream may be annotated with the identities of each participant as they are speaking. The one or more processors may be configured to implement a face detection method to extract an image or video segment of the participant corresponding to the participant identity from the video stream of the conference.

[0047] The one or more processors may be configured annotate the transcription of the first audio data and / or the second audio data received from the real-time speech transcription service with the image or video segment of the participant corresponding to the participant identity.

[0048] The end-system may be communicatively connected with a display screen. The one or more processors may be configured to cause the display screen to display an image or video segment of the participant on the display screen.

[0049] The one or more processors may be configured to form a mapping between spatial signatures for sound sources within the room containing the end-system and spatial positions within the field of view of one or more of the cameras.

[0050] The one or more processors are configured to correlate each spatial signature with the mapping to identify a suitable camera of the one or more cameras to capture the video stream and / or to determine an image boundary of the video stream captured by the camera for the participant corresponding to the participant identity.

[0051] The one or more processors may be configured to maintain a history of spatial signatures and corresponding participant identities.

[0052] The one or more processors may be configured to compare the second spatial signature with the history of spatial signatures and participant identities to establish whether the participant corresponding to the second audio data is a known participant.

[0053] If the participant corresponding to the second audio data is not identified as a result of the comparison, the one or more processors may be configured to store an association between a new participant identity and the second spatial signature. The participant identity may be an arbitrary identifier. For example, each participant may be assigned a number or a letter.

[0054] The participant identity may comprise a name of the participant.

[0055] The participant identity may comprise information about the relative location of the participant.

[0056] The one or more processors may be configured to analyse the first signature data with respect to known information about a room in which the end-system is located.

[0057] The multiple microphones may be comprised in an array microphone.

[0058] The correlation operations may be performed at a digital signal processor of the array microphone.

[0059] The end-system may comprise the array microphone.

[0060] The correlation operations may comprise beam forming operations. The signature data may comprise beam forming metadata. The signature data may comprise phase and / or amplitude differences for the first and / or second audio data to each microphone.

[0061] According to another aspect, there is provided a method of conducting an audio or video conference, the method comprising: at a first time, receiving first audio data captured by the microphones, the first audio data corresponding to speech by a participant; performing a correlation operation between the first audio data received from each of the microphones to generate first signature data representing a first spatial signature for audio signals corresponding to the first audio data; storing an association between a participant identity and the first spatial signature; at a second time subsequent to the first time, receiving second audio data captured by the microphones, the second audio data corresponding to speech by a participant; performing a correlation operation between the second audio data received from each microphone to generate second signature data representing a second spatial signature for audio signals corresponding to the second audio data; determining that the second spatial signature matches the first spatial signature having the stored association with the participant identity; and performing an operation using the second audio data in dependence on the participant identity.

[0062] According to another aspect, there is provided an end-system for conducting an audio or video conference, the end-system being communicatively connectable with multiple microphones configured to capture audio data from multiple participants during the conference and comprising one or more processors configured to: receive audio data captured by the microphones, the audio data corresponding to speech by a participant; perform a correlation operation between the audio data received from each of the microphones to generate signature data representing a spatial signature for audio signals corresponding to the audio data; store an association between a participant identity and the spatial signature; transmit the audio data to a speech transcription service; and convey the participant identity corresponding to the audio data to a conference server for annotation of an audio transcription of the audio data with the participant identity.

[0063] According to another aspect, there is provided an end-system for conducting an audio or video conference, the end-system being communicatively connectable with multiple microphones configured to capture audio data from multiple participants during the conference and comprising one or more processors configured to: receive audio data captured by the microphones, the audio data corresponding to speech by a participant; perform a correlation operation between the audio data received from each of the microphones to generate signature data representing a spatial signature for audio signals corresponding to the audio data; store an association between a participant identity and the spatial signature; convey the audio data to a speech transcription service; and annotate an audio transcription of an audio transcription received from the speech transcription service with the participant identity.

[0064] According to a further aspect, there is provided a computer program comprising instructions that when executed by a computer cause the computer to perform the methods above.

[0065] According to a further aspect, there is provided a computer-readable storage medium having stored thereon computer readable instructions that when executed at a computer (for example, comprising one or more processors) cause the computer to perform the methods above. The computer-readable storage medium may be a non-transitory computer-readable storage medium. The computer may be implemented as a system of interconnected devices.

[0066] BRIEF DESCRIPTION OF THE FIGURES

[0067] The present invention will now be described by way of example with reference to the accompanying drawings. In the drawings:

[0068] Figure 1 shows a diagram of a typical video conference where a speech transcription service receives audio data from the conference;

[0069] Figure 2 schematically illustrates an example of a video end-system comprising an array microphone;

[0070] Figure 3 schematically illustrates an example of a video conferencing system where two video end-systems are participating in a video conference;

[0071] Figure 4 schematically illustrates a video end-system in communication with a facial recognition system and facial comparison system;

[0072] Figure 5 schematically illustrates the steps of an exemplary method of conducting an audio or video conference;

[0073] Figure 6 schematically illustrates the steps of an exemplary method of conducting a video conference;

[0074] Figure 7 illustrates an exemplary device configured to operate as a processing device of a video end-system.

[0075] DETAILED DESCRIPTION

[0076] The present invention relates to an end-system for a video or audio conference.

[0077] In the following examples, the term video end-system is used to refer to a dedicated video conferencing system. Principles described herein may also be implemented on hardware that is capable of audio and / or video conferencing together with other functions. The video end-system comprises a video conferencing endpoint and may also comprise function specific hardware and software for performing video conferencing calls. The end-system may also be configured to perform audio conferencing calls (without a video component). The endpoint may be a uniquely addressable point on a network that can engage in a call with another video conferencing endpoint. The video end-system may be a shared video end-system that is configurable for use by multiple individual users. For example, the shared video end-system may be located in a corporate meeting room which is available to be reserved and used by multiple users for making conference calls.

[0078] The video end-system comprises one or more processors. The video end-system may comprise a main processor. The main processor may be a processing element of the video end-system which is configured to run system software for operating the video endsystem and perform video conference calls. The main processor may be configured to communicate with video conferencing hardware and receive commands from a controller of the video end-system. The main processor may be a separate endpoint device such as a single-purpose computing device or a local server, which is connected, for example by wired or wireless connection, to the video conferencing hardware and a controller of the video end-system. As such, the device comprising the main processor may also be referred to as an endpoint device. In other examples, the main processor may form part of a single unit which also includes the controller and / or the video conferencing hardware. The main processor may comprise a network interface for sending and receiving a content stream of a video conferencing call.

[0079] The video end-system described herein can therefore be used for performing video or audio conferencing calls.

[0080] One or more processors of the video end-system (for example the main processor) may be configured to implement one or more software entities which are configured to operate the video end-system and enable the performance of a conference call using the video end-system. The video end-system may be operated to conduct a video and / or audio conferencing call by running system software. The system software may be an application or a computer program which includes proprietary software which is provided with and configured to operate with the video end-system. The video end-system may be managed by a standard operating system. For example, the standard operating system may be a mass-market operating system such as Windows, Apple iOS, Linux or Android. One or more processors of each video end-system (for example, the main processor) may be configured to implement a first software entity for supporting the video conference. The first software entity may be an application, computer program or software package running on the video end-system. The first software entity may be capable of receiving media input from any of multiple input channels. The multiple input channels may be, for example, input channels for media output by the video conferencing hardware (such as the camera or microphone), or one or more input channels a port such as a High-Definition Multimedia Interface (HDMI) port to which an additional media source may be connected and media output by the media source may be provided for inclusion in the video conference. The media input to the input channels may be transmitted to the other video end-system during the video conference. The one or more processors may also be configured to implement additional software entities (such as applications), as will be described in more detail below.

[0081] Each video end-system may comprise video conferencing hardware. The video conferencing hardware may comprise one or more devices which are in communication with the one or more processors of the video end-system that are used for video conferencing. For example, the video conferencing hardware may comprise a camera and / or a display screen for sharing images and / or a video stream in a video conferencing call. Specifically, a video stream may be communicated from the camera via one of the media input channels to be included in the video stream of a video conferencing call. One or more additional video streams may be received from the video stream of the video conferencing call and displayed on the display screen. In some examples, the video endsystem may comprise multiple cameras and display screens.

[0082] The video conferencing hardware may also comprise a loudspeaker and / or a microphone for sending and receiving an audio stream of a conference call. Specifically, the video conferencing hardware may comprise one or more audio loudspeakers (e.g. as area broadcasting speakers or as loudspeakers included in personal headphones) for outputting audio content of a conferencing call. The video conferencing hardware may comprise one or more audio microphones from providing an audio stream for the content stream of a conferencing call. In further examples, the control interface and video conferencing hardware may comprise separable units which are in communication with each other (for example via wired or wireless connections).

[0083] Each component of the video conferencing hardware (i.e., cameras, display screens, microphones, loudspeakers, etc) may be provided as one or more integral units, or as separate devices. The video conferencing hardware and the control interface may also be provided as one or more integral units or as separate devices. For example, a controller having a touch screen may serve as a display screen as well as the control interface.

[0084] Each video end-system may also comprise a controller, which may also be referred to as a control interface or a control interface device. The controller may be any device comprising user input controls, such as a user interface, for enabling users to control the video end-system. For example, the control interface device may be a tablet comprising a touch screen. The control interface device may comprise a user input device via which a user may use to issue control signals to the video end-system by engaging with a user input device (for example, by pressing buttons) on the touch screen. Other examples of a control interface device may include a keyboard, a mouse, a control panel with mechanical buttons and a voice control device. For example, a user may use the control interface device to begin and end a conference call, add or remove parties from a conference call, adjust volume levels of speakers in the video end-system and configure display screen settings. Multiple control interface devices may be provided which are connected to (for example via a wired or wireless connection) to the remainder of the video end-system.

[0085] The controller and the video conferencing hardware may also be provided as a single unit. For example, the video end-system may be a single desktop or handheld unit which may also be referred to as a video end-device. In yet further examples, the control interface and video conferencing hardware may comprise separable units which are in communication with each other, for instance via the endpoint (e.g. via wired or wireless connections).

[0086] The media data sent between endpoints in a video conference may include video data, and / or audio data. In some examples the media may include other types of data which may be received from the media source. For example, the media may include computer files, still images or text.

[0087] Each video end-system may also comprise a control device. The control device may comprise a user interface. In some implementations, the user interfaces may each comprise a touch screen on a display screen. The user may interact with one or more user input devices on the user interface (for example, buttons displayed on the touch screen) to control one or more functions of the video end-system.

[0088] Many video conferencing systems make use of array microphones for better speech clarity in a noisy environment. These microphones generally comprise a plurality of microphones and a digital signal processor. The processor can analyse (or forward to a different processor for analysis) the audio from multiple microphones and cross-correlate the audio signals from each microphone to form beams, that is sets of phase differences and amplitude differences for the audio to each microphone in the array that correlate with the physical location of a participant of the video conference who is speaking. By doing this, the processor may be able to reject sound that is coming from other sources. Typically, an array microphone will focus on the loudest current speaker or speakers in the room and will not keep information about historical speakers. Typically, an array microphone will not pass on the metadata indicating the speaker’s beam parameters to other entities, but will simply provide an audio stream to the endpoint.

[0089] Advantageously, as described herein, by cross-correlating audio signals from multiple microphones and utilising beamforming metadata from multiple microphones within audio or video conference, that metadata may be analysed to allow generated spatial signatures, which can provide an indication of the location of the participant relative to the array microphone, to be associated with the identities of participants in the conference. When a subsequently determined spatial signature for subsequent audio data is found to match a previously determined spatial signature, the participant identity associated with the previously determined spatial signature can be attributed to the subsequent audio data. The associated participant identities may be used in further operations, such as the annotation of an audio transcription of the conference or the annotation of a video stream of the conference with one or more participant identifiers for the appropriate periods during the conference that they are speaking. The examples herein are described in the context of video end-systems where there are both audio and video streams as part of the conference. Described features may also be applied to an end-system for an audio-only conference where there is no video stream.

[0090] Figure 2 shows an example of a video end-system 201 comprising an endpoint 202, a camera 203 and an array microphone 204. The array microphone 204 comprises a set of microphones. There are multiple microphones in the set. The camera 203 and the array microphone 204 are communicatively connectable to the endpoint 202.

[0091] In this example, the array microphone 204 comprises a digital signal processor 205 and multiple microphones 206, 207. The room in which the video end-system is situated is occupied by two users 215 and 220. The video end-system 201 can use well-known techniques to cross-correlate the audio data corresponding to speech by user 215 from each of the multiple microphones 206, 207 of the array 204 in order to determine the phase and / or amplitude differences for the audio data for the set of microphones 206, 207 of array 204. This process may be referred to as beam-forming. The sets of phase and / or amplitude differences resulting from this correlation operation are traditionally used to allow the digital signal processor of the array microphone to perform noise rejection. This set of phase and / or amplitude differences collectively form metadata that can be used to indicate the relative location in the room at which user 215 is located (relative to the microphones of the array, or a location between two or more of the microphones) and thus generate signature data representing a spatial signature for audio signals corresponding to the audio data (where the audio data corresponds to speech by user 215).

[0092] In some implementations, the phase and / or amplitude signatures can be determined by the processor 205 of the array microphone 204 and transmitted to the main processor of the endpoint 202. This metadata can be transmitted by the microphone array to the endpoint 202 along with the audio data itself. This correlation operation may therefore in some implementations be performed at the digital signal processor 205 and then the signature data can be sent to the main processor of endpoint 202. The correlation operation may in some implementations be performed at the main processor of endpoint 202. For example, the array microphone 204 may send the raw audio data for each of the microphones in the array to the main processor of the endpoint 202, which may perform the correlation operation between the audio data for each of the microphones to generate the signature data representing the spatial signature for the audio signals corresponding to the audio data.

[0093] A correlation operation can thus be performed between audio data corresponding to speech by a participant received from each of the microphones in the array to generate signature data (such as phase and amplitude signatures determined using beam forming methods) representing a spatial signature for audio signals corresponding to the audio data. The end-system can store an association between a participant identity and the spatial signature.

[0094] The video end-system may keep a record of the historical metadata phase and amplitude signatures for different users in the room. This may allow spatial signatures and associated participant identifies to be stored at the end-system that can then be used to determine when a particular participant is speaking later in the conference.

[0095] The video end-system may be configured to perform correlation operations for the audio data corresponding to each speech segment and determine the identity of the respective participant speaking in a respective segment. For audio data corresponding to speech by a participant at a second time, a correlation operation can be performed to generate signature data representing a spatial signature for the audio signals corresponding to that audio data. The video end-system can then determine whether the spatial signature for the audio data correspond to that speech segment matches any previously determined spatial signatures determined at earlier times. If the spatial signature does match a previously determined spatial signature, the participant identity associated with the previous spatial signature can be retrieved and an operation can be performed using the audio data in dependence on the retrieved participant identity For example, the video stream for the conference captured by camera 203 may be annotated with the participant identity, or an audio transcription of the conference may be annotated with the participant identity. If the spatial signature does not match a previously determined spatial signature, the system can store an association between a new participant identity and the spatial signature which can be applied to subsequent audio data that is analysed and determined as having that spatial signature. The use of historical spatial signatures can conveniently allow the video end-system to perform separation and persistence of users.

[0096] The spatial signatures may define a range of phase and / or amplitude differences that are associated with a particular participant. For example, signature data may be generated from a correlation operation performed on audio signals from each of the microphones at a first time. This signature data may comprise phase and / or amplitude differences for the audio data corresponding to speech by a participant. The spatial signature corresponding to the signature data may encompass phase and / or amplitude differences within a range of the determined signature data, so that small changes in subsequently determined signature data corresponding to speech by the same participant can still be attributed to that participant by matching it to a previously generated spatial signature.

[0097] Put another way, a conferencing end-system may comprise multiple microphones. Optionally it may comprise other audio-visual equipment such as one or more cameras, loudspeakers or video displays. The microphones capture sound and generate data representing the sound. A processor, which may be local to the microphones or remote, can estimate the direction of a sound source by a beamforming technique. The beamforming technique may, for example, involve computing time offsets for each of the microphones except a primary microphone such that when signals from the microphones are delayed by the respective time offset and combined with each other and / or the signal from the primary microphone the best signal is received from a participant. The best signal may, for example, be the signal where there is greatest constructive interference or correlation between the signals or where there is the least interference in the combined signal. That or those time offsets are characteristic of the direction or position of the participant. Data indicative of the direction or position of the participant may then be used as described herein to improve speech-to-text processing of the audio. One way in which the data may be used is to annotate the text. Another way in which the data may be used is to assist in selection of an algorithm or algorithmic parameters for performing speech- to-text conversion.

[0098] Furthermore, the video end-system 201 may be configured to track small changes in the metadata phase and amplitude signatures, which can indicate that the user is moving whilst speaking. This is advantageous in that the video end-system can maintain the correct participant identity if the participant changes position while speaking (for example, because they are walking around the room to use a white board or rolling in a chair).

[0099] Figure 3 shows a video end-system 201 according to figure 2 that has been joined to a video conference with video end-system 208, which comprises an endpoint 209 and a camera 210, via a conference controller 230. In this example, the two video end-systems 201 and 208 are communicatively connectable to each other, allowing a video conference to be established between the two video end-systems.

[0100] In this example, the endpoints 202, 209 of the video end-systems 201 , 208 are communicatively connectable with each other over a communication network 109, which in this example is the internet. The communication network may alternative be, for example, a corporate network. The video end-system 201 may establish a video conference with the endpoint 209 of the other video end-system 208 via the communication network 109 (i.e. by sending media data such as audio data and / or video data over the communication network).

[0101] The video end-systems 201 , 208 may be two of multiple video end-systems in communication with a conference server comprising conference controller 230. That is, there may be further video end-systems participating in the video conference in addition to video end-systems 201 and 208. The video end-system may establish a video conference with one or more other video conferencing endpoints which may be part of respective video end-systems via the server. The video end-system may be communicatively connectable with the one or more other video conferencing endpoints via the network 109. A media stream can be sent between the video end-system 201 and one or more other endpoints over the communication network 109.

[0102] In some implementations, the conference controller 230 can take an audio feed from the conference established between the two video end-systems 201 , 208 and send it to a real-time speech-to-text transcription service 240. The transcription service can return a live stream of text that has been transcribed from the audio stream to the conference controller 230. The microphone array 204 with its digital signal processor 205 can perform the beam-forming to determine the phase and amplitude signature metadata for the currently speaking participant. The video end-system can store an association between the signature metadata and an identifier for the currently speaking participant.

[0103] This participant identifier can be sent to the conference controller 230 which can annotate the text stream from the transcription service with the identity of the current speaker as received from the currently active video end-system 201 or 208. Alternatively, the endpoint 202 may receive the text stream from the transcription service 240 and the text stream may be annotated at the endpoint (for example, by the main processor of the endpoint 202).

[0104] As shown in Figure 3, the video end-system may also comprise a camera 203 configured to capture video data used to form a video stream of the conference. The video stream from camera 203 can be sent via the conference controller 230 to the end-system 208 where it can be displayed on a display screen (not shown in Figure 3) to the participants in a room accommodating end-system 208. The participant identifier may be used to annotate the video stream with the identity of the currently speaking participant. The annotation may be performed at the video end-system 201 (for example by the main processor of endpoint 202) or at the conference controller 230 (the conference server).

[0105] This may be done in real-time so that the identifier can be displayed on a display screen while a participant is speaking. This may be useful when, for example, a meeting is being held between parties for the first time and the parties may not be sure of the names of the other participants. This may allow participants of the conference to become familiar with the identity of each speaker and to be able to subsequently direct questions to particular participants by using their displayed identity.

[0106] The system may therefore be configured to annotate one or more segments of the video stream corresponding to audio data corresponding to speech by a participant with that participant’s identity.

[0107] The participant identity may be an arbitrary identifier. For example, ‘Participant T, ‘Participant A etc. Other participants may be similarly labelled. When that participant is detected to be speaking later in the conference, as a result of generated signature data for an audio segment representing a spatial signature known to the system (i.e. stored previously) with an associated known participant identity, the same identifier is used.

[0108] Alternatively, the participant identity may comprise information about the relative location of the participant. The signature data may be analysed with respect to known information about the room in which the end-system is located. The known information about the room may comprise information indicating the location of seating in the room. The information may comprise an indication of which participants are seated in which locations in the room. Such information may be provided by annotating a map of the room with names (or other identifiers) of the participants and where they are seated in the room. This can be used to associate a particular participant with the location in which they are seated, and thus associated that particular participant (and / or their identity) with a particular spatial signature corresponding to that location in the room.

[0109] The participant identifiers may also be determined using facial recognition methods.

[0110] Figure 4 shows a video end-system 201 according to figure 2 that makes an additional use of the camera 203 connected to the endpoint 202. In this embodiment, the video endsystem 201 is calibrated to have a record of the expected signature data (for example, the beam-forming parameters) for each position in the field of view of the camera 203. In one implementation, this calibration may be performed by a user providing an opportunity for the end-system to detect a single face and a single voice from various positions in the meeting room. By establishing an association between camera position and signature data (e.g. beam-forming parameters) for each participant, the video end-system can locate the face of the current speaker by moving the camera to the associated position for the participant corresponding to the signature data. This is advantageous both from the perspective of providing a good user experience, allowing the video-conferencing system to display a cropped video of the current speaker, or providing a still image that identifies the current speaker that can be sent to the conference controller 230 to be added to the annotated audio transcription.

[0111] The still image or cropped video of the currently-speaking participant can be sent to a facial recognition system 241 to determine the identity of the participant. Once identified, the name of the speaker (or another suitable identifier) can be used as a participant identity and added to, for example, the transcription annotation or video stream. In some implementations, the participant list for the video conference may be passed to the facial recognition system 241 to constrain the facial recognition algorithm to known participants.

[0112] In some implementations, the still image snapshot or a video segment may be submitted to a facial comparison system 242 to correlate with images of other speakers in the conference. If a match is found, then the system may link or merge the identifiers to handle the situation where a participant changes position in the room whilst not speaking and is wrongly identified as a new speaker.

[0113] The transmission of the combined audio for the conference to a speech transcription service and the subsequent merging of annotations received from video end-systems may be performed locally by the video end-system or it may be performed by a centralised or distributed server such as a cloud service or a conference bridge.

[0114] In another implementation, the array microphones and video cameras may be calibrated to map signature data for sound sources within the meeting room space to positions within the field of view of the camera or cameras. Furthermore, live signature data from a conference may be correlated with the calibration map to identify the best camera and image boundary of the video stream for the current speaker. The accuracy of the image boundary may be further enhanced by using facial detection. A still image snapshot of the video stream, a video segment or a cropped video segment (for example, cropped to display a participant’s face) may be added to the speaker’s identifier and used, for example, to annotate a live speech-to-text transcription.

[0115] In another implementation, the still image snapshot or video segment may be submitted to a facial recognition system to determine the human identity of the speaker. This human identification information (such as the name of the participant) may be added to the speaker’s identifier and used, for example, to annotate the live speech-to-text transcription or video stream.

[0116] Therefore, the signature data determined from audio signals gathered by the array microphone may be used to separate speech segments by participant. The end-system can maintain a history of received signature data representing a spatial signature for previous speakers in order to establish persistence of participant identifiers. When signature data is determined from the audio signals acquired by the array microphone, it can be correlated with the history of previous speakers to establish whether the speaker is a known previous speaker. If the correlation is successful and a speaker is identified, then the existing speaker identifier can be used to annotate the transcription service text stream. If the speaker is not found, then the spatial signature can be stored as being associated with a new participant with a new identifier and annotated as such in the transcription service text stream.

[0117] In some implementations, the signature data representing a spatial signature for a participant may comprise the phase differences between audio signals received by each microphone in an array when that participant is speaking. The data may alternatively or additionally comprise amplitude measurements (for example, a time-averaged amplitude) for audio signals received by each of the microphones in the array. The differences in measured amplitude between each microphone in the array may be used to estimate a direction to the currently speaking participant based on the attenuation of sound as it travels further.

[0118] If there are several microphone arrays in a room, per-array amplitude may be used to determine a particular array to be associated with a particular participant. The signature data for that participant may then be determined using audio data acquired by only that array. For example, a respective array which receives audio signals from a participant with the highest average amplitude (average across the microphones in the respective array) may for example indicate that the respective array is the closest array to the participant. This closest array may then be associated with that participant and the signature data representing a spatial signature for that participant may then be determined using audio signals acquired by the microphones of that array only.

[0119] Once the participant identities and their associated spatial signatures have been stored at the end-system, the identities and / or their associated spatial signatures may be sent to a central conference server. The central conference server may receive participant identities and / or their associated spatial signatures from all end-systems that are participating in the conference. From each end-system, the central server may receive the participant identities and / or spatial signatures for the participants located in a room containing that respective end-system. The central conference server may have access to the feed from the audio transcription service and may annotate the audio transcription and / or video stream, if present, with the respective identifier for each participant.

[0120] Alternatively, the annotation task may be performed at an end-system itself (for example by the main processor of the endpoint). The current speaker identifiers (and optionally photos or video segments) may be received by a respective endpoint from other endpoints during the conference.

[0121] In one example, a central server could act as a broker for current-speaker information, receiving identity data from each endpoint in the videoconference and then sending the latest information to any system that is subscribed to the feed, such as a subscription annotation service. In one implementation, the endpoint that performs the transcription / annotation may be a listen-only conference endpoint in the cloud.

[0122] Figure 5 shows the steps of an exemplary computer-implemented method of conducting an audio or video conference. At step 501 , the method comprises at a first time, receiving first audio data captured by the microphones, the first audio data corresponding to speech by a participant. At step 502, the method comprises performing a correlation operation between the first audio data received from each of the microphones to generate first signature data representing a first spatial signature for audio signals corresponding to the first audio data. At step 503, the method comprises storing an association between a participant identity and the first spatial signature. At step 504, the method comprises, at a second time subsequent to the first time, receiving second audio data captured by the microphones, the second audio data corresponding to speech by a participant. At step 505, the method comprises performing a correlation operation between the second audio data received from each microphone to generate second signature data representing a second spatial signature for audio signals corresponding to the second audio data. At step 506, the method comprises determining that the second spatial signature matches the first spatial signature having the stored association with the participant identity. At step 507, the method comprises performing an operation using the second audio data in dependence on the participant identity. For example, the second audio data and the participant identity may be sent to a conference server, where the audio data is transcribed and the formed transcription can be annotated with the participant identity. Advantageously, by cross-correlating audio signals detected by different microphones in an array and utilising signature data, such as beamforming metadata, from array microphones within a conference, that signature data may be analysed to form an association between spatial signatures and participant identities. The approach described herein can be used for identifying a current speaker in an audio or video conference in real time. This may be used, for example, to detect when the same participant is speaking subsequently, to provide annotation for a live speech to text transcription service identifying the speaker for each text segment, or to provide annotation of a video stream during the conference with an identifier for the current speaker, which can be displayed on a screen.

[0123] In some implementations, the system may use information obtained during a period of detection of a known participant (for example, a period of detection of speech by a known participant) from one media type and / or using one detection method or strategy at a first time to aid the detection of the same participant at a later time using a different data type and / or detection method or strategy.

[0124] For example, a participant may be known to the system by their face, such that they can be detected using facial recognition methods from image data acquired by one or more imaging devices, but not known by their voice. As such, they initially cannot be detected by voice recognition methods from audio data acquired by one or more microphones.

[0125] When detecting a speaking participant via facial recognition and / or video detection of them speaking (which may be done when the detection of that participant is over a threshold certainty) from image data during a first time period in which they are speaking, the system may form a signature for this participant from audio data received in that time period or from data derived from such audio data (for example by applying one or more processing techniques). In some implementations, the system may learn the signature. Once formed, the audio signature for the participant could be used to detect that same participant at a later time. At the later time, their face may not be visible or recognisable. They may be detected based solely on the audio data received from one or more microphones in dependence on the signature. They may not be detected based on data captured by one or more imaging devices. Alternatively, the detection methods used at the earlier time and the later time may both use raw data captured by one of (i) the one or more microphones and (ii) the one or mode imaging devices but may process the data differently. Optionally, where data is available, both detection methods may be used at the later time in order to give increased certainty of the identity of the participant.

[0126] As mentioned above, a signature for the participant may be determined from the audio data received from the one or more microphones or the image data received from the one or more imaging devices (or data derived therefrom) during a time when the speaking participant is detected to be the known participant using a different detection method.

[0127] The signature may be stored by the system and used subsequently to correlate data received from a different source, or using the same data type (audio or image data) processed using a different detection method to identify the known participant from the data having the different source or processed using the different detection method.

[0128] For example, at the first time, the known participant may be identified by analysing received audio data from one or more microphones. The known participant may be identified by correlating the received audio data with known audio signatures for participants. The received audio data may match an audio signature corresponding to the known participant.

[0129] At the first time, in addition to the received audio data, the system can also receive image data captured by the one or more imaging devices. The system may form an image signature for the known participant from the image data. For example, the system may detect from the image data that the known participant is wearing a particular colour and / or item of clothing and / or has a particular hair colour.

[0130] At a second time, the system may detect that a participant that is speaking matches the image signature captured during the first time and can detect that the participant is the known participant.

[0131] Similarly, from knowing a participant voice (i.e. having a stored audio signature corresponding to a currently speaking participant) but having no identifying face data or other image signature for that participant, the face information can be learned from image and / or video analysis during a time period when the participant’s voice is detected using voice recognition methods from audio data received from a microphone of the system. A signature can be formed for the participant that allows the participant to be detected from image data, for example from one or more images of their face.

[0132] In the above ways, the use of different data types (for example, audio and image data) and / or detection methods I processing for the same data type (for example, voice recognition using audio data, facial recognition or colour detection using image data) may allow for more accurate detection than using models that are pre-trained prior to a conference.

[0133] For example, a known participant may be detected to be speaking at a first time because their mouth is detected to be moving from image data received from the one or more cameras. The participant may be detected to be the known participant using a facial recognition method. Using a different methodology, for example colour detection and / or matching, the system may be configured to form a clothing colour signature for the participant from the image data. The system may perform colour detecting and / or matching and image segmentation for identifying colour(s) and / or types of clothing. The clothing colour signature may, for example, specify that the participant is wearing a blue shirt and jeans. The system may store the clothing colour signature for the known participant so that it can be used to detect the identity of the participant at a second time later than the first time.

[0134] This above approach may be particularly useful when, at the later time, the known participant cannot be detected using the first detection method due to data quality issues or the absence of data required to detect the participant using the first detection method.

[0135] In the previous example, at the later time, the system may not be able to detect the participant’s face using facial recognition and / or detect that their mouth is moving, for example because the participant is standing up and their face is out of view. The above approach may allow the system to subsequently detect the participant at a later time using information acquired at an earlier time using a different detection method. This may be useful when the method used to detect the known participant at the first time is no longer available or has a certainty of below a predetermined threshold. In another example, in order to more accurately detect a current speaker, one or more processors of the end-system may be configured to detect movement of a user’s mouth. The movement of the user’s mouth may be detected to occur for more than a predetermined time period. This may be performed in addition to facial recognition to identify the participant. An identified participant may be detected as speaking if movement of their mouth is detected to occur for more than a predetermined time period (for example, more than 2 seconds). This may help to distinguish over accidental or incidental mouth movement.

[0136] As described above, the system may be configured to use information identified from historical audio and / or video data to subsequently identify a user. For example, the system may be configured to use information identified for a user previously in the same day to assist detection of that user later in the day. For example, if it was found at an earlier time that a user had a particular characteristic, this characteristic may be detected at a later time and used to identify a participant. For example, the characteristic may be detected for a particular user earlier in a day, such as in the morning, and used to detect the identity of the participant later that day. The characteristic may be detected and used at the later time if the characteristic was detected with a degree of certainty (for example, with a degree of certainty greater than a threshold).

[0137] The signature may specify a visual characteristic of the participant. The visual characteristic may relate to a participant’s appearance. For example, the user’s clothing, in particular the colour(s) of their clothing or the presence of an item of clothing, such as a hat. For example, if it was detected with a degree of certainty that a participant X is wearing a blue shirt and blue jeans in the morning, it is likely they are doing the same in the afternoon. Therefore, a participant may be detected at a later time based on a characteristic detected at an earlier time. This may be used as an additional signal by training a detection model on the video recorded at the earlier time.

[0138] The system may be configured to build participant descriptors based on known recordings of speech by that participant. For example, recordings may be taken from when the user is using their phone or laptop to join meetings. A model to identify the user may be trained on such recordings data. The system may be configured to detect set of users in a room to limit the possible participants to match against. The end-system may utilise "people count" with other signals. Counting the number of people in a room from image and / or video data is common in video conferencing technology. When the people count stays stable, it may be assumed that the set of people in the meeting room is the same set. The known participant may be identified from a set of possible participants, the set of possible participants being determined in dependence on a determined number of people attending the conference. The number of people attending the conference may be determined from image data from the one or more imaging devices.

[0139] For users joining the meeting from a mobile phone application, the system may be configured to use information from the mobile application, such as a public IP address, GPS and wideband audio, to help identify and / or position the user.

[0140] As discussed above, a transcript may be generated for the conference. Based on the detected identifier of the participants speaking during the conference, the transcript may be annotated with the participant identifiers. The following further information may be used to enhance the transcript.

[0141] The transcript may include, for example, a textual description of an event, or an image snapshot of a moment during the conference. One or more processors of the end-system may be configured to detect one or more non-verbal communication cues from image data (e.g. a video) captured by the one or more imaging devices. For example, gestures by the participant (such as thumbs up), nodding of a participant’s head, or shaking of a participant’s head. An indication of such non-verbal cues may be added to the transcript at the point when the non-verbal communication cue occurred. For example, a textual description “User X nodded” may be added to the transcript. One or more processors of the end-system may be configured to detect one or more actions by a participant, such as getting up from a seat, writing on a whiteboard or leaving the room. An indication of such actions may be added to the transcription at the point when the action occurred. For example, a textual description “User X left the room” may be added to the transcript. One or more processors of the end-system may be configured to detect an emotion in a participant’s voice and / or an expression in a participant’s face. An indication of such emotion and / or expression may be added to the transcription at the point when the displayed emotion or expression occurred. For example, a textual description “User X smiled / frowned” may be added to the transcript.

[0142] One or more processors of the end-system may be configured to detect written and / or drawn elements visible to the one or more imaging devices (for example, on a whiteboard) and include the drawings in the transcript, either directly in the transcript alongside the text, or as an attached image.

[0143] One or more processors of the end-system may be configured to detect one or more segments of the conversation that are determined to be irrelevant to the discussion. For example, the one or more processors may be configured to detect sarcasm, jokes, banter, and off-topic conversation. The one or more processors may be configured to reduce such segments determined to be irrelevant to the discussion into a summary and include the summary in the transcript instead of the full segment. For example, the one or more processors may reduce the segment to a short and succinct description, such as “the participants discussed football then moved back to the topic at hand”.

[0144] In embodiments of the present system, the video end-system can operate during a videoconference to use data gathered before the start of the videoconference or gathered outside the videoconference to augment its algorithm for estimating the identity of participants in the conference. When a participant is active in the conference, for example by speaking, writing, or making a gesture or expression, the system can estimate / detect the identity of that participant. The system can then make a record in a journal of the conference. That record can associate the estimated identity with a description of the activity undertaken by the participant. For example, the record might say “Jane said ‘Hello’” or “Roger smiled”. The journal can constitute a record of the events during the conference. Some ways in which the system may augment its algorithm are set out below. Any one or more of these may be combined with any other(s).

[0145] 1 . The system may store data describing the appearance of a participant at a first time during a video conference. Then, at a second time, during a subsequent video conference, or a later part of the same video conference, the system may use that information to help identify the same participant. The system may prioritise or may only use such information that has been gathered within a predetermined period of the second time, for example the later conference or part thereof. This can avoid the system placing too much reliance on data gathered at a time since which the participant’s appearance has changed.

[0146] 2. The system may store data describing the vocal attributes of a participant at a first time during a video conference. Then, at a second time, during a subsequent video conference, or a later part of the same video conference, the system may use that information to help identify the same participant. The system may prioritise or may only use such information that has been gathered within a predetermined period of the second time, for example the later conference or part thereof. This can avoid the system placing too much reliance on data gathered at a time since which the participant’s vocal attributes have changed.

[0147] 3. The system may use data gathered from one or more cameras directed at the location of the end-point to detect motions of participants as they enter the end-point. This data may be used to help estimate the identity of a participant by comparing it with previously captured data for the same participant.

[0148] Figure 6 shows the steps of an exemplary method of conducting a video conference. At step 601 , the method comprises, at a first time, receiving first data and second data, the first data and the second data being captured by or derived from data captured by one of (i) one or more microphones and (ii) one or more imaging devices, the first time corresponding to a time when a participant of multiple participants is speaking. At step 602, the method comprises analysing one of the first data and the second data to detect the participant speaking at the first time to be a known participant. At step 603, the method comprises forming a signature for the known participant from the other of the first data and the second data. At step 604, the method comprises, at a second time, receiving one or more of third data and fourth data, the third data and the fourth data being captured by or derived from data captured by one of (i) the one or more microphones and (ii) the one or more imaging devices, the second time corresponding to a time when a participant of the multiple participants is speaking. At step 605, the method comprises, in dependence on the signature for the known participant, detecting the participant speaking at the second time to be the known participant from one or more of the third data and the fourth data. The device 701 shown in Figure 7 may be configured to operate as a processing element of the video end-system. The device 701 comprises a processor 702, a memory 703 and a transceiver 704. The processor may be implemented as dedicated hardware in the device, such as a processing chip. The memory is arranged to communicate with the respective processor. Memory may be a non-volatile memory. Each device may comprise more than one processor and more than one memory. The memory may store data that is executable by the processor. By executing program code contained in such data, the one or more processors may perform functions as described herein. The memory may store such program code in a non-transitory manner. The processor may be configured to operate in accordance with a computer program stored in non-transitory form on a machine readable storage medium. The computer program may store instructions for causing the processor to perform its methods in the manner described herein. The device 701 also comprises a transceiver 704 for receiving and / or sending data from and / or to one or more of the other entities, such as between video conferencing endpoints. Each device also comprises a power source, such as a battery, or is connectable to mains power.

[0149] The end-system functions to provide an end point to a video conference. The end-system does not need to be located only at the place where the end point is provided. One or more parts of the end-system may be located remotely from that place. The parts may communicate with each other over any suitable network. The video conference may be a one-to-one call between two people or a meeting between numerous people, who may be at two, three or more endpoints, or it may be a call between one or more people and an automated service such as a chat service.

[0150] The applicant hereby discloses in isolation each individual feature described herein and any combination of two or more such features, to the extent that such features or combinations are capable of being carried out based on the present specification as a whole in the light of the common general knowledge of a person skilled in the art, irrespective of whether such features or combinations of features solve any problems disclosed herein. The applicant indicates that aspects of the present invention may consist of any such individual feature or combination of features. In view of the foregoing description, it will be evident to a person skilled in the art that various modifications may be made within the scope of the invention.

Claims

AMENDED CLAIMS received by the International Bureau on 17 of April 2025 (17.04.2025)1. An end-system for conducting a video conference, the end-system being communicatively connectable with one or more microphones configured to capture audio data from multiple participants during the conference and one or more imaging devices configured to capture image data of multiple participants during the conference, the end-system comprising one or more processors configured to: at a first time, receive first data and second data, the first data and the second data being captured by or derived from data captured by one of (i) the one or more microphones and (ii) the one or more imaging devices, the first time corresponding to a time when a participant of the multiple participants is speaking; analyse one of the first data and the second data to detect the participant speaking at the first time to be a known participant; when the detection of the known participant from the one of the first data and the second data is over a threshold certainty, form a signature for the known participant from the other of the first data and the second data; at a second time, receive one or more of third data and fourth data, the third data and the fourth data being captured by or derived from data captured by one of (i) the one or more microphones and (ii) the one or more imaging devices, the second time corresponding to a time when a participant of the multiple participants is speaking; and in dependence on the signature for the known participant, detect the participant speaking at the second time to be the known participant from one or more of the third data and the fourth data.

2. The end-system as claimed in claim 1 , wherein the second time is within a predefined period of the first time.

3. The end-system as claimed in claim 1 or claim 2, wherein the first time is a time period.

4. The end-system as claimed in any preceding claim, wherein the one or more processors are configured to analyse one of the first data and the second data to detect the participant speaking at the first time to be the known participant using a firstdetection method and form the signature for the known participant from the other of the first data and the second data by analysing the other of the first data and the second data using a second detection method different to the first detection method.

5. The end-system as claimed in claim 4, wherein the first detection method and / or the second detection method is one of a facial recognition method, a motion detection method, a colour detection method and a voice recognition method.

6. The end-system as claimed in any preceding claim, wherein the one or more processors are configured to detect the participant speaking at the first time to be the known participant from image data received from the one or more imaging devices.

7. The end-system as claimed in claim any of claims 1 to 5, wherein the one or more processors are configured to detect the participant speaking at the first time to be the known participant from audio data received from the one or more microphones.

8. The end-system as claimed in any preceding claim, wherein the one or more processors are configured to detect the participant speaking at the second time to be the known participant from image data received from the one or more imaging devices in dependence on the signature.

9. The end-system as claimed in claim any of claims 1 to 7, wherein the one or more processors are configured to detect the participant speaking at the second time to be the known participant from audio data received from the one or more microphones in dependence on the signature.

10. The end-system as claimed in any preceding claim, wherein the one or more processors are configured to annotate one or more segments of a video stream corresponding to the first time and / or the second time with a participant identifier for the known participant.11 . The end-system as claimed in claim 10, wherein the one or more processors are configured to cause a display in communication with the end-system to display aparticipant identifier for the known participant when the known participant is detected to be speaking.

12. The end-system as claimed in claim 10 or claim 11 , wherein the participant identifier is an arbitrary identifier, or the participant identity comprises a name of the participant.

13. The end-system as claimed in any of claims 10 to 12, wherein the one or more processors are configured to: convey segments of the video stream to a conferencing server; and convey the participant identity to the conferencing server for annotation of one or more of the segments of the video stream corresponding time(s) when the known participant is detected to be speaking with the participant identity.

14. The end-system as claimed in any preceding claim, wherein the one or more processors are configured to maintain a history of signatures and corresponding known participants.

15. The end-system as claimed in any preceding claim, wherein if the participant detected to be speaking at the first time is not a known participant, the one or more processors are configured to form a new signature for the new participant and store an association between the new participant and the new signature.

16. The end-system as claimed in any preceding claim, wherein the one or more processors are configured to: convey audio data captured by the one or more microphones to a speech transcription service; and annotate a transcription of the audio data corresponding to the first time and the second time received from the speech transcription service with a participant identifier corresponding to the known participant.

17. The end-system as claimed in any of claims 1 to 15, wherein the one or more processors are configured to:convey audio data captured by the one or more microphones to a speech transcription service; and convey the participant identity corresponding to the known participant to a conferencing server for annotation of a transcription of the audio data corresponding to the first time and the second time received by the conferencing server from the speech transcription service with the participant identifier.

18. The end-system a claimed in claim 16 or claim 17, wherein the one or more processors are configured to detect one or more segments of the conference comprising speech by one or more participants in which the speech is irrelevant to a topic of speech preceding and / or following the respective segment, wherein the one or more processors are configured to omit the speech from the audio data conveyed to the video conferencing server and / or speech transcription service.

19. The end-system as claimed in any of claims 16 to 18, wherein the one or more processors are configured to detect one or more verbal or non-verbal events from the audio data captured by the one or more microphones and / or the image data captured by the one or more imaging devices and produce a textual description of the one or more events for inclusion in a transcript of the conference.

20. The end-system as claimed in any of claims 16 to 19, wherein the one or more processors are configured to detect one or more participant emotions from the audio data captured by the one or more microphones and / or the one or more participant expressions from the image data captured by the one or more imaging devices and produce a textual description of the one or more participant emotions and / or expressions for inclusion in a transcript of the conference.

21. The end-system as claimed in any of claims 16 to 20, wherein the one or more processors are configured to detect one or more participant actions from the image data captured by the one or more imaging devices and produce a textual description of the one or more participant actions for inclusion in a transcript of the conference.

22. The end-system as claimed in any preceding claim, wherein the one or more processors are configured to, at the first time, receive first audio data captured by multiple microphones, the first audio data corresponding to speech by a participant; perform a correlation operation between the first audio data received from each of the microphones to generate first signature data representing a first spatial signature for audio signals corresponding to the first audio data; store an association between a participant identity and the first spatial signature; at a second time subsequent to the first time, receive second audio data captured by the microphones, the second audio data corresponding to speech by a participant; perform a correlation operation between the second audio data received from each microphone to generate second signature data representing a second spatial signature for audio signals corresponding to the second audio data; determine that the second spatial signature matches the first spatial signature having the stored association with the participant identity; and perform an operation using the second audio data in dependence on the participant identity.

23. A method of conducting a video conference, the method comprising: at a first time, receiving first data and second data, the first data and the second data being captured by or derived from data captured by one of (i) one or more microphones and (ii) one or more imaging devices, the first time corresponding to a time when a participant of multiple participants is speaking; analysing one of the first data and the second data to detect the participant speaking at the first time to be a known participant; when the detection of the known participant from the one of the first data and the second data is over a threshold certainty, forming a signature for the known participant from the other of the first data and the second data; at a second time within a predefined period of the first time, receiving one or more of third data and fourth data, the third data and the fourth data being captured by or derived from data captured by one of (i) the one or more microphones and (ii) the one or more imaging devices, the second time corresponding to a time when a participant of the multiple participants is speaking; andin dependence on the signature for the known participant, detecting the participant speaking at the second time to be the known participant from one or more of the third data and the fourth data.

24. A computer program comprising instructions that when executed by a computer cause the computer to perform the method of claim 23.

25. A computer-readable storage medium having stored thereon computer readable instructions that when executed by a computer cause the computer to perform the method of claim 23.