Method, video-conferencing endpoint, server
Patent Information
- Application Number
- GB2023019014
- Authority / Receiving Office
- GB · GB
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-13
- Publication Date
- 2025-07-09
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
TECHNICAL FIELD The present disclosure relates to a method, a video-conferencing endpoint, and a server. BACKGROUND In recent years, video conferencing and video calls have gained great popularity, allowing users in different locations to have face-to-face discussions without having to travel to a same single location. Business meetings, remote lessons with students, and informal video calls among friends and family are common uses of video conferencing technology. Video conferencing can be conducted using smartphones or tablets, via desktop computers or via dedicated video conferencing devices. Video conferencing systems enable both video and audio to be transmitted, over a digital network, between two or more participants located at different locations. Video cameras or webcams located at each of the different locations can provide the video input, and microphones provided at each of the different locations can provide the audio input. A screen, display, monitor, television, or projector at each of the different locations can provide the video output, and speakers at each of the different locations can provide the audio output. Hardware or software-based encoder-decoder technology compresses analog video and audio data into digital packets for data transfer over the digital network, and decompresses the data for output at the different locations. The cameras in video conferencing systems map a three-dimensional (3D) scene into a two-dimensional (2D) medium (e.g., an image). This mapping from the 3D coordinates to 2D coordinates is called projection. The most common projection type is rectilinear projection, which corresponds to a pinhole camera model. It transforms straight lines in the 3D scene into straight 2D lines on the image. However, objects and persons at the edge of the camera field of view (FOV) are stretched unevenly and do not keep their shape, whilst objects at the center of the viewpoint keep their shape well. The present disclosure was arrived at in light of the above considerations. SUMMARY Accordingly, in a first aspect, embodiments of the invention provide a method for manipulating a video stream of an object scene captured by a camera, each image frame of the video stream containing one or more regions of interest of the object scene mapped according to a current camera projection type, wherein the method comprises: performing a loop comprising the steps of: (i) analysing the region of interest to determine respective values of one or more predetermined characteristics of the region of interest, (ii) selecting a camera projection type from a set of predetermined camera projection types on the basis of the determined values of the predetermined characteristics, and (iii) switching the current camera projection type to the selected camera projection type to change the mapping of the object scene in the region of interest, thereby providing a remapped video stream. In this way, the switching dynamically changes the current camera projection type of the region of interest in the video stream according to the developments in the predetermined characteristics of the region of interest. Accordingly, a more natural view of the object scene can be presented. The method may include any one, or any combination insofar as they are compatible, of the optional features set out with reference to the first aspect. Where there is one more than one region of interest, the loop may be repeated for each of the region of interest, or step (i) may be repeated for a plurality of regions of interest and step (ii) may select the camera projection type on the basis of the determined values of the predetermined characteristics for all of the regions of interest together. The region of interest may be the entire image frame. Alternatively, the region of interest may comprise only a part of the image frame, for example a region of the image frame containing one or more objects of one or more predefined types. In such examples, different camera projection types can be utilised to ensure that a more natural view of each object within a respective region of interest can be presented. The method may include steps, performed before performing the loop, of: detecting objects of one or more predefined types in a frame of the video stream; selecting a plurality of crop regions from the frame of the initial video stream, each crop region including at least one bounding box, each bounding box including a detected object of a predefined type; and providing the plurality of crop regions of the loop as regions of interest. Steps (i) -(iii) can then be performed on each of the regions of interest. The method may include steps, performed outside of the loop, of: selecting and / or modifying one or more parameters associated with the current camera projection type. The method may include steps, within the loop, of selecting and / or modifying one or more parameters associated with the selected camera projection type. The parameters may include a projection center point and / or a projection orientation. The one or more predetermined characteristics may be any one, two, three, four or all of: a size of the region of interest, a location of the region of interest in the object scene, a camera framing mode, a number of persons included in the region of interest, and a number of predefined objects included in the region of interest. The size of the region of interest may be determined in terms of either a field of view of the region of interest or a number of pixels of the region of interest. The plurality of predetermined camera projection types may include one or more conformal camera projection types. The conformal camera projection types may be selected from the group comprising Mercator and stereographic projection, and combinations of different conform projections (e.g., in a weighted average). Further conformal camera projection types that may be selected include: oblique Mercator projection, and Space-oblique Mercator projection). The plurality of predetermined camera projection types may include one or more non-conformal camera projection types. The non-conformal camera projection types may be selected from the group comprising rectilinear projection, equal-area projection, equidistant projection, orthogonal projection, and one or more cylindrical projections. A plurality of variations of cylindrical projections may be available, for example perspective cylindrical projection, equal-area cylindrical projection, equirectangular projection, Gall stereographic projection, Miller cylindrical projection, and Peters projection. The non-conformal camera projection may be a combination of different non-conformal projections (e.g., in a weighted average). The plurality of predetermined camera projection types may include one or more hybrid camera projection types which each combine characteristics of a conformal and nonconformal camera projection type. The hybrid camera projection type may be a weighted sum of a conformal camera projection type and a non-conformal camera projection type. For example, having a weighted sum of two projection functions / (x) and / 2(x) written as: a ■ f^x) + (1 - a) - / 2(x), where a is a fixed real number between 0 and 1, a is therefore the weight of the function ^(x) and (1 - a) is the weight of the function / 2(x). In one example, the hybrid camera projection type may be a weighted sum of a rectilinear projection and an equidistance projection with a high (a >0.8) weighting on the rectilinear projection so as to give almost rectilinear projection for most angles but with a slight reduction in the stretching at large viewing angles. The video stream may contain plural regions of interest of the object scene each mapped according to a respective current camera projection type, and wherein the steps of (i) analysing, (ii) selecting, and (iii) switching are repeatedly performed for each region of interest. The video stream may be a video stream from a video conferencing endpoint and the regions of interest may correspond to respective crop regions from each image frame of the video stream. The method may be performed on a video-conferencing endpoint and may include a step of providing the remapped video stream to a second video-conferencing endpoint. The method may be performed on a server, connected between a pair of video-conferencing endpoints, wherein the video stream is received from a first video-conferencing endpoint of the pair and the remapped video stream is provide to a second video-conferencing endpoint of the pair. The method may include a step of the first video-conferencing endpoint also providing, to the server, metadata relating to the video stream. The metadata may specify a number or and / or properties relating to the regions of interest (such as the size, number of persons within a given ROI, etc.). Alternatively, the server may generate such metadata itself. The server may select the camera projection type on the basis of the metadata (and so the metadata may include values for the predetermined characteristics). The method may be computer-implemented. In a second aspect, embodiments of the invention provide a video-conferencing endpoint configured to perform the method of the first aspect (including any one or any combination insofar as they are compatible of the optional features set out with reference thereto), and configured to transmit the remapped video stream to a further video-conferencing endpoint (for example via a server). In a third aspect, embodiments of the invention provide a server configured to receive a video stream from a first video-conferencing endpoint, to perform the method of the first aspect (including any one or any combination insofar as they are compatible of the optional features set out with reference thereto), and to transmit the remapped video stream to a further video-conferencing endpoint. In a fourth aspect, embodiments of the invention provide a video-conferencing system including a plurality of the video-conferencing endpoints, at least one of which is a video-conferencing endpoint of the second aspect, connected via a network. The invention includes the combination of the aspects and preferred features described except where such a combination is clearly impermissible or expressly avoided. Further aspects of the present invention provide: a computer program comprising code which, when run on a computer, causes the computer to perform the method of the first aspect; a non-transitory computer readable medium storing a computer program comprising code which, when run on a computer, causes the computer to perform the method of the first aspect; and a computer system programmed to perform the method of the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS Figure 1 shows a schematic drawing of a video conferencing system including multiple video conferencing endpoints; Figure 2 shows a schematic drawing of a data processing device which may be included in a video conferencing endpoint; Figure 3 shows the steps of a method performed by the data processing device of Figure 2; Figure 4 is an example of an image frame mapped according to a Mercator camera projection; Figure 5 is an example of an image frame mapped according to a rectilinear camera projection; Figures 6A and 6B contrast two regions of interest mapped using different camera projections; Figures 7A and 7B contrast two different regions of interest mapped using different camera projections; and Figure 8 shows an example of a remapped video stream where there is more than one region of interest. DETAILED DESCRIPTION Aspects and embodiments of the present invention will now be discussed with reference to the accompanying figures. Further aspects and embodiments will be apparent to those skilled in the art. Figure 1 shows a video conferencing system 1 including a plurality of video conferencing endpoints 10a-10d, each of which may be controlled by a respective user. The video conferencing endpoints 10a-1 Od (and therefore the users) are located at different locations. They comprise a computing device, a mobile device, a tablet, or a dedicated video conferencing device etc. The video conferencing system 1 enables both video and audio signals to be transmitted, via digital network 50, between the video conferencing endpoints 10a-10d. Digital network 50 is preferably wireless, but may also be wired. A video camera (such as video camera 30 in video conferencing endpoint 10a) and microphone (not shown) is positioned at each of the video conferencing endpoints 10a-10d in order to provide a video input and an audio input, respectively. A screen, display, monitor, television, and / or projector, and speakers (not shown) are positioned at each of the video conferencing endpoints 10a-10d in order to provide a video output and an audio output, respectively. In Figure 1, one of the video conferencing endpoints 10a comprises a data processing device 20. In some other example video conferencing systems, each video conferencing endpoint 10a-10d may comprise such a data processing device 20. Camera 30 at video conferencing endpoint 10a is configured to capture an initial video stream (wherein the initial video stream comprises a plurality of video frames, each including one or more users (i.e. people) positioned within the camera’s field of view), and transmit the initial video stream to the data processing device 20 (e.g. via connection 40, which may be wired or wireless). One or more of the cameras 30 of the video conferencing endpoints 10a (and possible all of them) are configured to map the images of their respective object scenes (e.g., including the one or more users) according to one of a plurality of camera projection types. The or each camera is configured to change the camera projection type it is using in response to a command from the respective data processing device 20. Figure 2 schematically illustrates an example data processing device 20 which may be positioned in video conferencing endpoint 10a of video conferencing system 1, shown in Figure 1. The data processing device 20 comprises a receiver 21 which is configured to receive (wirelessly or via a wired connection) the initial video stream from the camera 30. The data processing device 20 also comprises an object detection unit 22, a bounding box setting unit 23, and a cropping unit 24, which together are configured to manipulate the initial video stream received by the receiver 21 into multiple views (or “crop regions”) which correspond to regions of interest in the initial video stream. The regions of interest may correspond to portions of the initial video stream including people, for example. The data processing device 20 also comprises a transmitter 26, which is configured to transmit (wirelessly, or via a wired connection) the crop regions of interest to the video conferencing endpoints 10a-1 Od, for rendering and display thereon. Figure 3 is a flow diagram of a method 100. In a first step S110, a region of interest of an object scene in an image frame is analysed, to determine respective values of one or more predetermined characteristics of the region of interest. For example, the predetermined characteristics may include any of or any combination of: a size of the region of interest, a location of the region of interest in the object scene, a camera framing mode, a number of persons included in the region of interest, and a number of predefined objects included in the region of interest. Next, in step S120, a camera projection type is selected from a set of predetermined camera projection types on the basis of the determined values of the predetermined characteristics. For example, where the region of interest is determined to be of medium sized (e.g., larger than 70°, and / or up to 120° FOV), a rectilinear or close to rectilinear projection may be chosen. If the region of interest, corresponding to a crop of an original image, is small (e.g., less than 45° FOV),, a conformal camera projection type such as Mercator may be chosen. In another example, if the region of interest is determined to be near to the center of the object scene (e.g., it is less than 30° from the center of the region of interest to the center of the object scene captured by the image sensor) then a nonconformal camera projection type such as rectilinear may be chosen. Conversely if the region of interest is determined to be near the edge of the object scene then a conformal camera projection type may be chosen. In a further example, if camera is operating in a framing mode whereby people and / or objects of interest are individually framed (see, e.g., GB2594761Athe contents of which is incorporated herein by reference) then each region of interest may be mapped using a conformal projection mode. In a further example, if it is identified (e.g., by the object detection unit 22) that one person is within the region of interest then a conformal camera projection type may be used. Next, in step S130, the current camera projection type (that is, the one being used to capture images of the object scene) is switched to the selected camera projection type to change the mapping of the object scene in the region of interest, thereby providing a remapped video stream. This remapped video stream can then be provided to any one of the video-conferencing endpoints 10a-10d. The method returns to step S110, whereby the region of interest is again analysed. It should be noted that the result of step (ii) may be a determined that the current camera projection type should be retained. If so, step (Hi) may be omitted. Figure 4 is an example of an image frame mapped according to a Mercator camera projection and Figure 5 is an example of an image frame mapped according to a rectilinear camera projection. As can be seen in Figure 4, the straight lines (pipework, cabling, etc.) towards the edges of the image are distorted and warped relative to the corresponding regions in Figure 5. However, the head and shoulders of the individual in the image appears more natural in the Mercator camera projection than in the rectilinear camera projection, Figures 6A and 6B contrast two regions of interest from image frames of the same object scene but mapped using different camera projections. Figures 7A and 7B contrast two different regions of interest mapped using different camera projections. Figure 8 shows an example of a remapped video stream 800 where there is more than one region of interest. For example, the initial video stream with the object of interest in may undergo the process for manipulating the video stream disclosed in GB 2594761 A, such that a plurality of crop regions (here, regions of interest) are provided. Hence, each region of interest undergoes the steps S110 -S130 discussed above (noting that in some examples the camera projection may remain the same as was originally used to capture it). After each region of interest has undergone steps S110 - S130, the resulting remapped (or originally mapped) regions of interest are utilised to form a composite final video stream (either in the videoconferencing end point capturing the images, in a videoconferencing endpoint receiving the plurality of remapped (or originally mapped) regions of interest, or in a server located between the two endpoints). In other words, the remapped (or originally mapped) regions of interest can either first be compiled into a composite final video stream and transmitted as a single final video stream, or they an be transmitted as individual remapped (or originally mapped) regions of interest from the capturing video conferencing endpoint such that they are compiled into a composite view either by the receiving video conferencing endpoint or an intermediary device (such as a server). The systems and methods of the above embodiments may be implemented in a computer system (in particular in computer hardware or in computer software) in addition to the structural components and user interactions described. The term “computer system” includes the hardware, software and data storage devices for embodying a system or carrying out a method according to the above described embodiments. For example, a computer system may comprise a central processing unit (CPU), input means, output means and data storage. The computer system may have a monitor to provide a visual output display. The data storage may comprise RAM, disk drives or other computer readable media. The computer system may include a plurality of computing devices connected by a network and able to communicate with each other over that network. The methods of the above embodiments may be provided as computer programs or as computer program products or computer readable media carrying a computer program which is arranged, when run on a computer, to perform the method(s) described above. The term “computer readable media” includes, without limitation, any non-transitory medium or media which can be read and accessed directly by a computer or computer system. The media can include, but are not limited to, magnetic storage media such as floppy discs, hard disc storage media and magnetic tape; optical storage media such as optical discs or CD-ROMs; electrical storage media such as memory, including RAM, ROM and flash memory; and hybrids and combinations of the above such as magnetic / optical storage media. While the disclosure has been described in conjunction with the exemplary embodiments described above, many equivalent modifications and variations will be apparent to those skilled in the art when given this disclosure. Accordingly, the exemplary embodiments of the disclosure set forth above are considered to be illustrative and not limiting. Various changes to the described embodiments may be made without departing from the spirit and scope of the disclosure. In particular, although the methods of the above embodiments have been described as being implemented on the systems of the embodiments described, the methods and systems of the present disclosure need not be implemented in conjunction with each other, but can be implemented on alternative systems or using alternative methods respectively. The features disclosed in the description, or in the following claims, or in the accompanying drawings, expressed in their specific forms or in terms of a means for performing the disclosed function, or a method or process for obtaining the disclosed results, as appropriate, may, separately, or in any combination of such features, be utilised for realising the disclosure in diverse forms thereof. While the disclosure has been described in conjunction with the exemplary embodiments described above, many equivalent modifications and variations will be apparent to those skilled in the art when given this disclosure. Accordingly, the exemplary embodiments of the disclosure set forth above are considered to be illustrative and not limiting. Various changes to the described embodiments may be made without departing from the spirit and scope of the disclosure. For the avoidance of any doubt, any theoretical explanations provided herein are provided for the purposes of improving the understanding of a reader. The inventors do not wish to be bound by any of these theoretical explanations. Any section headings used herein are for organizational purposes only and are not to be construed as limiting the subject matter described. Throughout this specification, including the claims which follow, unless the context requires otherwise, the word “comprise” and “include”, and variations such as “comprises”, “comprising”, and “including” will be understood to imply the inclusion of a stated integer or step or group of integers or steps but not the exclusion of any other integer or step or group of integers or steps. It must be noted that, as used in the specification and the appended claims, the singular forms “a,” “an,” and “the” include plural referents unless the context clearly dictates otherwise. Ranges may be expressed herein as from “about” one particular value, and / or to “about” another particular value. When such a range is expressed, another embodiment includes from the one particular value and / or to the other particular value. Similarly, when values are expressed as approximations, by the use of the antecedent “about,” it will be understood that the particular value forms another embodiment. The term “about” in relation to a numerical value is optional and means for example + / -10%.
Claims
1. A method for manipulating a video stream of an object scene captured by a camera, each image frame of the video stream containing one or more regions of interest of the object scene mapped according to a current camera projection type, wherein the method comprises:performing a loop comprising the steps of:(i) analysing the region of interest to determine respective values of one or more predetermined characteristics of the region of interest,(ii) selecting a camera projection type from a set of predetermined camera projection types on the basis of the determined values of the predetermined characteristics, and(iii) switching the current camera projection type to the selected camera projection type to change the mapping of the object scene in the region of interest, thereby providing a remapped video stream.
2. The method according to claim 1, further comprising the steps, outside the loop, of:selecting and / or modifying one or more parameters associated with the current camera projection type.
3. The method according to claim 1 or 2, wherein, within the loop, the method includes steps of selecting and / or modifying one or more parameters associated with the selected camera projection type.
4. The method according to any preceding claim 2 or 3, wherein the parameters include a projection centre point and / or a projection orientation.
5. The method according to any preceding claim, wherein the one or more predetermined characteristics are any one, two, three, four or all of: a size of the region of interest, a location of the region of interest in the object scene, a camera framing mode, a number of persons included in the region of interest, and a number of predefined objects included in the region of interest.
6. The method according to claim 5, wherein the size of the region of interest is determined in terms of either a field of view of the region of interest, or a number of pixels of the region of interest.
7. The method according to any preceding claim, wherein the plurality of predetermined camera projection types includes one or more conformal camera projection types.
8. The method according to claim 7, wherein the conformal camera projection types are selected from the group comprising Mercator projection and stereographic projection.
9. The method according to any preceding claim, wherein the plurality of predetermined camera projection types includes one or more non-conformal camera projection types.
10. The method according to claim 9, wherein the non-conformal camera projection types are selected from the group comprising rectilinear, equidistant projection, equisolid projection, orthogonal projection, and one or more cylindrical projections.
11. The method according to any preceding claim, wherein the plurality of predetermined camera projection types includes one or more hybrid camera projection types which each combines characteristics of a conformal and a nonconformal camera projection type.
12. The method according to claim 11, wherein the hybrid camera projection type is a weighted sum of a conformal camera projection type and a non-conformal camera projection type.
13. The method according to any preceding claim, wherein the video stream contains plural regions of interest of the object scene each mapped according to a respective current camera projection type, and wherein the steps of (i) analysing, (ii) selecting and (iii) switching are repeatedly performed for each region of interest.
14. The method according to claim 13 wherein the video stream is a video stream from a video conferencing endpoint and the regions of interest correspond to respective crop regions from each image frame of the video stream.
15. The method according to any preceding claim, the method performed on a video-conferencing endpoint, and includes a step of providing the remapped video stream to a second video-conferencing endpoint.
16. The method according to any of claims 1-14, wherein the method is performed on a server, connected between a pair of video-conferencing endpoints, wherein the video stream is received from a first video-conferencing endpoint of the pair and the remapped video stream is provided to a second video-conferencing endpoint of the pair.
17. A video-conferencing endpoint configured to perform the method of any of claims 1-14, and configured to transmit the remapped video stream to a further video-conferencing endpoint.
18. A server configured to receive a video stream from a first video-conferencing endpoint, to perform the method of any of claims 1-14, and to transmit the remapped video stream to a further videoconferencing endpoint.
Citation Information
Patent Citations
Video stream manipulation
GB2594761A
Systems and methods for framing videos
US20220210326A1
Distortion correction via modified analytical projection
US20220366547A1