Video Stream Operations

The method enhances video conferencing by detecting and emphasizing predefined objects in the stream, providing equal visibility and optimizing network resources through region-based manipulation.

JP7824286B2Active Publication Date: 2026-03-04NEATFRAME LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2023521714
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-06-16
Filing Date
2021-08-19
Publication Date
2026-03-04
Estimated Expiration
2041-08-19

AI Technical Summary

Technical Problem

Existing video conferencing systems struggle to provide equal visibility of multiple participants in a single camera view, leading to complex and costly installations or ineffective region-of-interest focus that neglects non-audio relevant areas.

Method used

A method to detect predefined objects in a video stream, select crop regions around these objects, and transmit them as separate or combined views, ensuring equal visibility of participants regardless of their number or distance.

Benefits of technology

Enhances visibility of participants by emphasizing relevant regions, reducing unnecessary space, and optimizing network usage while maintaining consistent views.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007824286000001
    Figure 0007824286000001
  • Figure 0007824286000002
    Figure 0007824286000002
  • Figure 0007824286000003
    Figure 0007824286000003
Patent Text Reader

Abstract

Video Stream Manipulation. A method for manipulating an initial video stream captured by a camera in a videoconferencing endpoint into multiple views corresponding to regions of interest. The method includes: detecting objects of one or more predefined types in frames of the initial video stream; selecting multiple crop regions from the frames of the initial video stream; and transmitting the multiple crop regions. Each crop region includes at least one bounding box, and each bounding box includes a detected object of a predefined type.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a method and system for video stream manipulation, and in particular, but not exclusively, to a method and system for region-of-interest based video conferencing video stream manipulation. [Background technology]

[0002] Video conferencing and video calling have gained immense popularity in recent years, allowing users in different locations to hold face-to-face discussions without traveling to the same single location. Business meetings, remote lessons with students, and informal video calls between friends and family are common uses of video conferencing technology. Video conferences can be conducted using smartphones or tablets, via desktop computers, or via dedicated video conferencing devices.

[0003] A videoconferencing system allows both video and audio to be transmitted over a digital network between two or more participants at different locations. Video cameras or webcams at each of the different locations can provide video input, and microphones at each of the different locations can provide audio input. Screens, displays, monitors, televisions, or projectors at each of the different locations can provide video output, and speakers at each of the different locations can provide audio output. Hardware or software-based encoder-decoder technology compresses analog video and audio data into digital packets for data transmission over the digital network and decompresses the data for output at the different locations.

[0004] Often, a video conference may involve multiple users at different locations, e.g., 4, 8, 15, 20, etc. (e.g., each person is in a different location). Here, video and audio streams captured at each location may be transmitted to each of the different locations, allowing each user to see and hear each of the other users. For example, each user's screen may display a video stream from each location, possibly including a video stream from their own location.

[0005] In some cases, there may be only one user at each location, and therefore each video stream contains only a single user. In these situations, each user may position the camera at their location so that they are centered in the camera's field of view.

[0006] Alternatively, in other situations, there may be more than one person at one of the locations, for example, a business meeting held by video conference may have three people at a first location, one person at a second location, and one person at a third location.

[0007] In locations with multiple people, a single camera provides only one view of that location, and this video stream view may appear on each location's screen at a similar size to the other video stream views (e.g., there is only one person at the other location). In these situations, multiple people at a single location may appear much smaller in the video stream than anyone alone at their location.

[0008] In an attempt to address this problem, it is known to install multiple video cameras at a location where multiple people are present, capturing multiple video stream views of the location and covering a larger view than possible with a single camera. In some situations, the multiple video streams may each contain one person. However, this leads to complex systems that are costly and involve complex installation.

[0009] Another approach uses speaker tracking (e.g., audio tracking), whereby the sound from a speaking participant is used as a rule to focus on the speaking participant as a region of interest in the camera's field of view. However, a drawback of this solution is that it excludes areas from the overall picture that may be regions of interest for reasons other than sound. Summary of the Invention [Problem to be solved by the invention]

[0010] The present invention was devised in light of the above circumstances. [Means for solving the problem]

[0011] According to a first aspect of the present invention, there is provided a method for manipulating an initial video stream captured by a camera in a videoconferencing endpoint into multiple views corresponding to regions of interest, the method comprising: detecting objects of one or more predefined types in frames of the initial video stream; selecting a plurality of crop regions from the frames of the initial video stream, each crop region including at least one bounding box, each bounding box including a detected object of a predefined type; and transmitting the plurality of crop regions.

[0012] In this manner, any regions of the frames of the initial video stream that do not contain objects of a predefined type (e.g., not containing people) can be reduced and / or removed, and any regions that contain objects of a predefined type (e.g., containing people) can be enhanced, thereby providing a similar view of each person (e.g., each person can be provided with a similar size) regardless of the number of people at each location or their distance to the camera.

[0013] Optional features are described below.

[0014] The multiple crop regions may be transmitted as or in one or more frames of one or more final video streams, which may be for rendering and display at one or more videoconferencing endpoints.

[0015] As used herein, the term crop region may be understood to be a portion of a frame, and said crop regions from a frame may or may not overlap each other.

[0016] The one or more predefined type objects may include a person or a part / portion of a person. For example, the one or more predefined type objects may include, for example, a person's face, a person's head, and / or a person's head and shoulders. The one or more predefined type objects may additionally / alternatively include a fixed portion of a person's upper body, for example, from the top of the head as a proportion of the size of the person's face.

[0017] Because each crop region includes one or more bounding boxes, each crop region may include one or more people. Each crop region may include the same number of bounding boxes (and thus people) or a different number of bounding boxes (and thus people). For example, one crop region may include one bounding box (and thus a person), while another crop region may include multiple, e.g., two, three, four, or more bounding boxes (and thus two, three, four, or more people).

[0018] Detecting objects of one or more predefined types in a frame includes: Classify objects in the frame; This may include locating objects of said one or more predefined types within the frame.

[0019] In particular, the step of classifying the object in the frame may include predicting a class of the object in the frame in order to identify / recognize said one or more predefined types of object in the frame.

[0020] The method may further comprise setting a bounding box around the extent of each detected object of said one or more predefined types.

[0021] Thus, the / each bounding box set around the extent of each detected object of said one or more predefined types within the frame may contain information regarding the classification (e.g., that the bounding box contains an object of the predefined type) and / or information regarding the location of the object of the predefined type within the frame. The bounding box may contain a label containing the classification information and / or location information.

[0022] The extent of a detected object of a predefined type may be understood as encompassing all or most of the object.

[0023] The / each bounding box may be an imaginary structure superimposed on a frame of the initial video stream, and the / each bounding box contains a single detected object of a predefined type. Preferably, the / each bounding box is rectangular.

[0024] The / each bounding box's position information may include one or more coordinates. For example, the position information may include the coordinates (e.g., horizontal and vertical coordinates) of the four corners of a rectangular bounding box. Alternatively, the position information may include the coordinates (e.g., horizontal and vertical coordinates) of a known point on the rectangular bounding box (e.g., the center or the lower-left corner) and the dimensions (e.g., width and height) of the rectangular bounding box.

[0025] Detecting one or more predefined types of objects in the frame may be performed using a machine learning algorithm, for example using a trained neural network such as a Region-Based Convolutional Neural Network (R-CNN). Detecting one or more predefined types of objects in the frame may be performed using Haar Feature-based Cascade Classifiers, or Histogram of Oriented Gradients (HOG).

[0026] The method may further comprise the step of expanding the plurality of crop regions. The step of expanding the crop regions may be performed before or after the step of transmitting the crop regions.

[0027] Optionally, the / each crop region may be enlarged by scaling it to a size similar to a frame of the initial video stream, so that the / each object of a predefined type in the / each crop region can be observed (when rendered and displayed) as if it were captured using separate cameras, rather than a single camera capturing multiple objects in its field of view.

[0028] The method may include transmitting only the crop region. In this manner, any regions of frames of the initial video stream that do not include objects of a predefined type (e.g., not including people) are removed from frames of the final video stream. Alternatively, the method may include transmitting the crop region in addition to frames of the initial video stream, where the crop region may be transmitted simultaneously with the frames of the initial video stream. In this manner, frames of the initial video stream captured by the camera can still be observed, but objects of a predefined type (e.g., people) are emphasized in the crop region that is also transmitted.

[0029] The method may further include transmitting an alert indicating which of the plurality of crop regions includes the person currently speaking. In particular, the method may include detecting which person within the field of view of the camera is speaking based on audio input from a microphone in the same videoconferencing endpoint as the camera, and transmitting the alert based on the detection.

[0030] The microphones in the videoconferencing endpoint may be a microphone array. Thus, detecting which person within the camera's field of view is speaking may be based on information received from the microphone array regarding the direction of sound detected by the microphone array. This step may also be based on information regarding the camera's field of view and information regarding a selection of multiple crop regions from a frame of the initial video stream. Thus, the information regarding the direction of sound may be mapped to a crop region among the multiple crop regions. An alert may be sent along with the multiple crop regions to indicate to the user which crop region among the multiple crop regions contains the person currently speaking. The alert may cause the crop region containing the person speaking to be highlighted or further emphasized when the crop regions are rendered and displayed on one or more videoconferencing endpoints.

[0031] In some examples, each crop region may be transmitted as a separate frame in a separate final video stream, such that multiple final video streams (or, more specifically, a single frame in each of the multiple final video streams) are transmitted, each containing a single crop region.

[0032] As mentioned above, each crop region contains one or more bounding boxes (and therefore one or more people). So, if a crop region contains one person, the corresponding final video stream contains one person. If a crop region contains multiple people (e.g., two, three, four, etc.), the corresponding final video stream contains multiple people.

[0033] In another example, the crop regions may be combined into a composite view, which may be transmitted as a frame in the final video stream. The composite view may include each / all of the crop regions. Thus, only one final video stream (specifically, a single frame of a single final video stream) is transmitted, and the single frame includes each crop region (each crop region may include one or more people).

[0034] In alternative examples, some (but not all) of the crop regions may be combined into a composite view, and this composite view may be transmitted as a frame in the final video stream. In these examples, the composite view may include some but not all of the crop regions. The remaining crop regions may be combined into another composite view, and this another composite view may be transmitted as another frame in the final video stream. Alternatively, the remaining crop regions may be transmitted individually as separate frames in separate final video streams. Each frame of multiple final video streams (each frame may include a composite view of the crop regions or a single crop region) may be transmitted simultaneously. Compositing multiple or all final video streams (i.e., combining multiple crop regions into a composite view) can be advantageous when the maximum number of video streams is limited. For example, if the maximum number of final video streams is limited and there are more people in the initial frame than the maximum number of final video streams, people may be spaced apart when each crop region is transmitted in a separate final video stream. Compositing helps remove wasted space between people so that each person has more pixel area in the final view stream.

[0035] Combining multiple crop regions into a composite view also has the advantage that the number of separate final video streams is less likely to change with movement, or even loss or addition, of objects of said one or more predefined types in the initial frames (e.g., a person moving or entering or leaving the initial frame). For example, changes in the bounding box or number of crop regions may result in animated transitions between frames in the final video stream.

[0036] In some examples, a composite view of all multiple crop regions may be transmitted as a frame along with frames containing one or more of the multiple crop regions, such that a receiver can receive both a final video stream (as a summary stream) containing the composite view formed by all of the multiple crop regions, and multiple final video streams each containing one or more of the crop regions.

[0037] The number of final video streams to be transmitted may be determined based on the number of objects (e.g., people) of the one or more predefined types identified in the initial video stream. For example, if the number of objects of the one or more predefined types is relatively small (e.g., three or fewer, or four or fewer, or five or fewer), a single final video stream containing a composite view of all crop regions may be transmitted. This reduces network usage and screen real estate at the receiving end while providing the benefits of improved magnification.

[0038] The number of final video streams to be transmitted may be determined so that there are similar or the same number of bounding boxes, and therefore predefined objects (e.g., people), in each final video stream. For example, if four objects of the one or more predefined types are identified, two video streams may be transmitted, each containing a composite frame containing two crop regions (each crop region containing one of the four objects).

[0039] Each final video stream (which may include a single crop area, some but not all crop areas, or all crop areas) may be advertised on the network as a virtual camera and subscribed to. This allows receivers to choose which final video streams to subscribe to and view. This allows higher-level application layers to process and transmit the final video streams in a manner similar to a physical camera.

[0040] If each crop region is transmitted as a separate frame in a separate final video stream, the number of crop regions from the frames of the initial video stream may be based on a predefined maximum number of frames in the final video stream, the number of bounding boxes, and a predefined preferred aspect ratio for each crop region.

[0041] Optionally, the predefined maximum number of final video streams may be determined based on network bandwidth limitations, network capacity, or overall computing system limitations (e.g., limitations of one or more video conferencing endpoints configured to receive the transmitted crop region or limitations of a data processing device configured to perform the method of the first aspect). The method may include receiving information from a remote server or network service provider indicating the predefined maximum number of final video streams. The method may include negotiating the maximum number of final video streams with the network service provider.

[0042] The method may include repeating the step of selecting multiple crop regions from frames of the initial video stream if the predefined maximum number of final video streams changes (e.g., during negotiation with the network service provider). The number of selected crop regions from frames of the initial video stream may be transmitted to the network service provider.

[0043] The number of crop regions selected from the frames of the initial video stream may be less than or equal to a predefined maximum number for the final video stream.

[0044] When each crop region is combined into a composite view and the composite view is transmitted as a frame of a single final video stream, the number of crop regions selected from the frames of the initial video stream may be based on the number of bounding boxes, a predefined preferred aspect ratio of each crop region, and one or more predefined requirements of the composite view of the crop regions.

[0045] For example, the one or more predefined requirements for the composite view of the crop regions may include a predefined total size, a predefined shape, a predefined aspect ratio, and / or a predefined maximum number of rectangles in the layout that form the composite view, where a crop region is provided in each rectangle in the layout when the crop regions are combined into the composite view.

[0046] The layout that forms the composite view may be a grid, for example the composite view includes a grid of rectangles, each rectangle including a crop region.

[0047] The number of crop regions selected from the frames of the initial video stream may be less than or equal to a predefined maximum number of rectangles in the layout that forms the composite view. Limiting the number of rectangles in the composite view layout ensures that each rectangle, and therefore each crop region, is not too small.

[0048] Selecting multiple crop regions from a frame may be performed by reducing the area of ​​the crop regions that do not overlap any bounding box.

[0049] Selecting multiple crop regions from the frame may alternatively / additionally include reducing overlap between the crop regions.

[0050] Optionally, selecting multiple crop regions from the frame includes minimizing the area of ​​the crop regions that do not overlap any bounding box and / or minimizing overlap between the crop regions.

[0051] Each crop region may be rectangular.

[0052] Each crop region may be defined by (e.g., four) parameters. In one example, each crop region may be defined by the coordinates (e.g., horizontal and vertical coordinates) of its four corners. In another example, each crop region may be defined by the coordinates (e.g., horizontal and vertical coordinates) of a known point on the crop region (e.g., its center or lower-left corner) and the dimensions (e.g., width and height) of the crop region.

[0053] An exhaustive search for optimized parameters for each crop region for each frame of the initial video stream may not be ideal due to computational time and space constraints. Therefore, selecting multiple crop regions from the frames of the initial video stream may include parameterizing each crop region by minimizing a cost function (also called a loss function). That is, the parameters (e.g., size, shape, and position) of each crop region may be determined by minimizing the cost function. The cost function may include a weighted sum of (i) a term correlating to the area of ​​the crop region that does not overlap any bounding box; and (ii) a term correlating to the overlap between the crop regions. The cost function may be a weighted sum of (i) the area of ​​the crop region that does not overlap any bounding box; and (ii) the overlap between the crop regions; or the cost function may be, for example, a weighted sum of the squares of these terms.

[0054] Minimizing the cost function is equivalent to maximizing the inverse function, called fitness, and therefore the parameters of each crop region may be determined by maximizing a fitness function that includes a weighted sum of (i) a term that correlates to the area of ​​the crop region that does not overlap any bounding box, and (ii) a term that correlates to the overlap between the crop regions. The coefficient of the weighted sum in the fitness function is negative, while the coefficient of the weighted sum in the cost function is positive. Any term in the cost function that correlates to the two terms is equivalent.

[0055] The step of selecting a plurality of crop regions from frames of the initial video stream may include determining a set of possible candidate crop regions. A cost function may then be minimized with respect to the determined set of possible candidate crop regions. This helps reduce computation time and space by minimizing the cost function only over the determined set of possible candidate crop regions, rather than over every single possible set of crop regions.

[0056] Determining the set of possible candidate crop regions may include one or more of the following: Identifying a set of all possible candidate crop regions that satisfy a predefined preferred aspect ratio for each crop region; restricting the possible candidate set of crop regions to a possible candidate set in which each crop region in the candidate set contains no more than a predefined number of bounding boxes (e.g., no more than four, more preferably no more than three, more preferably no more than two, more preferably one bounding box); Restrict the set of possible candidates for the crop region to the set of possible candidates whose all bounding boxes are contained in the crop region; The possible candidate sets of crop regions are restricted such that the number of crop regions in each possible candidate set of crop regions is less than or equal to a predefined maximum number of crop regions in the final video stream or less than or equal to a predefined maximum number of rectangles in the layout that forms the composite view.

[0057] The cost function can then be minimized over a restricted set of possible candidate crop regions.

[0058] The step of determining the possible candidate sets may further include limiting the possible candidate sets of crop regions such that the number of crop regions in each possible candidate set is less than a first predefined threshold and / or more than a second predefined threshold, such that the potential candidate sets of crop regions are not too large (e.g., do not include too many crop regions) and / or are not too small (e.g., do not include too few crop regions).

[0059] The step of determining the possible candidate set may further include restricting the possible candidate set of crop regions such that the crop regions in each possible candidate set include one or more bounding boxes set around objects of a predefined type from a predefined subset of the predefined types. For example, the predefined subset of the predefined types may include a human head or a human head and shoulders. Thus, the possible candidate set of crop regions is restricted to include only crop regions that include a bounding box around a human head or a human head and shoulders. In this case, the restricted possible candidate set does not include, for example, a bounding box that shows only a human forehead or that shows a human's entire body. This further limits the number of possible candidate sets of crop regions over which the cost function is minimized, thereby further reducing the computational load.

[0060] Alternatively / additionally, the method may further include optimizing a cost function in a cost optimization process. The cost function may be optimized with respect to one or more constraints related to one or more of a predefined maximum number of final video streams, a predefined maximum number of rectangles in the layout forming the composite view, the number of bounding boxes, and a predefined preferred aspect ratio of each crop region. This is called constrained optimization. Minimizing the cost function may be performed using one or more mathematical solvers, such as linear programming, interior point methods, or generalized reduced gradient methods. Alternatively, other mathematical tools, such as genetic algorithms, may be used to minimize the cost function. The cost function may include first and second derivatives (since the geometric area of ​​the frames of the initial video stream involves quadratic transformations of coordinates). The constraints may be expressed in the form of mathematical linear constraints. Linear constraints may be thought of as mathematical expressions to which linear terms (i.e., coefficients multiplied by decision variables) are added or subtracted, forcing the resulting expression to be greater than or equal to, less than or equal to, or exactly equal to the value on the right-hand side. These expressions can help the solver converge to an optimal solution more quickly.

[0061] Constraints may additionally / alternatively relate to predefined types of objects. For example, constraints may relate to the ratio of a fixed part of a person's upper body from the top of the head to the size of the person's face. These constraints may include, for example, maximum and / or minimum ratios.

[0062] When multiple or each crop region is combined into a composite view and the composite view is transmitted as a single frame of a single final video stream, the relative spatial order or relative positions of the objects of the one or more predefined types within the frames of the initial video stream may be maintained in the layout forming the frame of the final video stream. In this way, a person sitting to the left of another person will be observed to be sitting to the left of that person in the layout forming the frame of the final video stream. This can be achieved, for example, by including metadata in the or each final video stream. The metadata may also identify whether an active speaker is located within a given composite view or final video stream and may further identify how many objects of the one or more predefined types are present in any given frame. If the metadata identifies an active speaker, the crop region containing the active speaker may be modified to make the active speaker more prominent than other objects of the one or more predefined types, for example, by enlarging or highlighting them.

[0063] The method may include blurring or replacing a background of each crop region. The background may be any part of the crop region that does not include the or each object of the one or more predefined types. For example, a mask may be applied that blurs or replaces all parts of the crop region that do not correspond to the or each object of the one or more predefined types. The mask may be obtained from a depth sensor or from a pre-trained machine learning model.

[0064] Optionally, the arrangement of the layout (e.g., the size / shape / aspect ratio of the rectangles in the layout, the order of the crop regions in the layout, etc.) may be determined by an optimization process. The optimization process may be based on the aspect ratio of each crop region selected from frames of the initial video stream and / or a predefined preferred aspect ratio for each crop region. As an example, the optimization process may determine that a layout that includes three equally sized crop regions stacked vertically to form a composite view (e.g., a grid) is better than a layout with one large square crop region on the left and two smaller vertically stacked rectangular crop regions on the right. This determination may be based on the aspect ratio of each crop region selected from frames of the initial video stream.

[0065] Due to computer time and space constraints, an exhaustive search for optimized aspect ratios of the rectangles may not be ideal. Thus, determining a layout placement may include determining possible candidates for the layout placement, such as placements of the rectangles in the layout (e.g., possible candidates for the aspect ratio of each rectangle in the layout), before determining (e.g., selecting) the layout placement from the limited candidates for the layout placement.

[0066] The possible candidates for the aspect ratio of each rectangle in the layout may be limited to a set of predefined aspect ratios, such as 16:9, 4:3, 1:1, etc.

[0067] The step of enlarging the crop regions may include scaling each crop region by a scaling factor so that it fits into a respective rectangle in a layout that forms the composite view in a frame of the final video stream.

[0068] Optionally, if the method includes a cost optimization process in which a cost function is minimized using constrained optimization, the constraints may relate to one or more of an optimized aspect ratio of each rectangle in the layout, a predetermined point (e.g., the upper left corner of the layout), and a scaling factor. Thus, the crop region may be sized and shaped to completely cover the layout (e.g., a grid) that forms the composite view within the frame of the final video stream. Thus, there may be boundary conditions for a (linear) combination of (linear) constraints. The step of minimizing the cost function may then be performed using one or more mathematical solvers, such as linear programming, interior point methods, or generalized reduced gradient methods. Alternatively, other mathematical tools, such as genetic algorithms, may be used to minimize the cost function.

[0069] The method may include detecting one or more predefined type objects in subsequent frames of the initial video stream, and setting a bounding box around the extent of each detected object in the subsequent frames.

[0070] The method may include selecting a crop region from a subsequent frame of the initial video stream. This crop region selection for the subsequent frame may be based on the selected crop region from the (initial) frame and the (tracked) movement of the bounding box (and thus the one or more predefined types of detected objects) between the (initial) frame and the subsequent frame. This reduces the computational burden compared to a method in which the crop region in each subsequent frame is selected in the same way as the initial frame. Furthermore, a consistent view of each crop region may be provided even if the detected object moves during the initial video stream.

[0071] Preferably, the same number of crop regions may be selected in the subsequent frame as in the (initial) frame, in this way a consistent view of the crop regions may be provided.

[0072] Preferably, each crop region from a subsequent frame may contain the same detected object(s) as the corresponding crop region from the (initial) frame. In this way, a consistent view of object(s) of a predefined type may be obtained.

[0073] In particular, the method may include tracking movement of the bounding box (and thus the detected object) from an (initial) frame to a subsequent frame, tracking position information of the bounding box (and thus the detected object) in the frame and subsequent frames (e.g., based on position information of a bounding box set around the detected object in the frame and subsequent frames), and selecting a crop region from the subsequent frame based on the position information of the bounding box in the frame and subsequent frames.

[0074] That is, based on position information of the bounding boxes (and thus the one or more objects of the predefined types) in the initial frame and the subsequent frame, each bounding box (and thus each object of the one or more predefined types) for the subsequent frame may be identified as corresponding to a bounding box in the (initial) frame. In particular, each bounding box in the subsequent frame may be identified as corresponding to a bounding box in the (initial) frame based on a difference in position information of the bounding boxes between the frame and the subsequent frame. Here, bounding boxes that are shifted by a minimum distance between the frame and the subsequent frame are identified as a corresponding pair of bounding boxes. Crop regions for the subsequent frame may then be selected based on the corresponding pair of bounding boxes, such that each crop region in the subsequent frame contains the same bounding box (and therefore the same object) as the corresponding crop region in the (initial) frame.

[0075] Alternatively / additionally, tracking the movement of bounding boxes (and thus objects of said one or more predefined types) from a frame to a subsequent frame of the initial video stream can be performed by tracking the movement of pixels within each bounding box using a feature mapping algorithm or using an optical flow estimation algorithm.

[0076] The method may further include transmitting crop regions from the subsequent frame in a similar manner as discussed above with respect to the (initial) frame.

[0077] This method further: This may further include expanding the crop region from the subsequent frame.

[0078] This method further: determining an amount of overlap between a pair of crop regions from subsequent frames; The method may include merging the pair of crop regions from a subsequent frame if the amount of overlap exceeds a predefined overlap threshold.

[0079] In this way, the likelihood of duplication of objects of a predefined type in the crop region from subsequent frames is reduced.

[0080] The method may further include determining a size of one or more crop regions from the subsequent frame, and dividing the crop region from the subsequent frame into two or more resulting crop regions if the size exceeds a predefined size threshold.

[0081] In this way, if a crop region from a subsequent frame contains multiple bounding boxes (and thus multiple objects of a predefined type) that have moved farther from one another in the subsequent frame compared to the (initial) frame, the crop region from the subsequent frame can be split into two or more resulting crop regions to reduce / remove any regions of the subsequent frame of the initial video stream that do not contain objects of one or more predefined types (e.g., people), thereby improving the emphasis of the objects of one or more predefined types from the subsequent frame, even if the objects of one or more predefined types have moved relative to their positions in the (initial) frame.

[0082] Crop regions from subsequent frames may be divided based on the number of bounding boxes they contain. For example, if a particular crop region from a subsequent frame has three bounding boxes (e.g., three people) and the three bounding boxes in the subsequent frame have moved further away from each other compared to the (initial) frame, such that the crop region from the subsequent frame exceeds a size threshold, the crop region is divided into three resulting crop regions, each containing a single bounding box (and therefore a single predefined type of object / person).

[0083] Optionally, the method may further comprise: determining whether the cropped region from the subsequent frame meets a predefined criterion; If the crop regions from subsequent frames do not meet the predefined criteria, reselect crop regions from subsequent frames of the initial video stream, where each reselected crop region contains at least one bounding box.

[0084] In this way, if too much movement of the bounding box (and the one or more predefined types of objects) occurs between the (initial) frame and the subsequent frame, the crop region from the subsequent frame is reselected using the same method as for selecting the crop region from the (initial) frame.

[0085] The step of determining whether the crop region from the subsequent frame satisfies a predefined criterion may be performed according to a metric, where the metric may be a return value of the above cost function, in particular, the predefined criterion may be that the metric is less than a predefined threshold.

[0086] Crop regions from further subsequent frames of the initial video stream (e.g., the third, fourth, fifth, ..., Nth frames) may be selected in a similar manner as for the subsequent frames. For example, the method may include detecting objects of one or more predefined types in the Nth frame of the initial video stream and setting a bounding box around the extent of each detected object in the Nth frame of the initial video stream; and selecting a crop region from the Nth frame of the initial video stream based on the crop region from the (N-1)th frame and movement of the bounding box between the (N-1)th frame and the Nth frame.

[0087] Cropped regions from further subsequent frames (e.g., the Nth frame) of the initial video stream may be transmitted in a similar manner as discussed above for the (initial) frame and subsequent frames, thereby transmitting one or more final video streams (for rendering and display by / at one or more videoconferencing endpoints).

[0088] The initial video stream may be received from a videoconferencing endpoint with a camera over a digital network, which may be wired or wireless.

[0089] A videoconferencing endpoint may be a computing device, a mobile device, or any other data processing device. As used herein, the term videoconferencing may be understood to include video calling.

[0090] The method may include transmitting crop regions from the (initial) frame (and any subsequent frames) to one or more videoconferencing endpoints (e.g., computing devices) for rendering at the one or more videoconferencing endpoints (which may also include the videoconferencing endpoint from which the initial video stream was received). One or more users may then view the crop regions at their respective videoconferencing endpoints (e.g., on the display or screen of their computing device). The crop regions may be transmitted over a digital network, which may be wired or wireless.

[0091] The method of the first aspect may be performed in a videoconferencing endpoint, and in some examples, in a data processing device that may be within or connected to the videoconferencing endpoint having the camera. By manipulating the initial video stream at the videoconferencing endpoint having the camera, the crop region may be of higher quality / resolution because the initial video stream may not have been compressed and / or encoded for prior transmission. Thus, the initial video stream may be a raw data stream from the camera (e.g., a raw, unencoded and / or uncompressed video stream). Alternatively, the method of the first aspect may be performed at a server. The server may be remote from the videoconferencing endpoints (e.g., the videoconferencing endpoint transmitting the initial video stream and one or more videoconferencing endpoints receiving the crop region) and the camera capturing the initial video stream.

[0092] A second aspect of the present invention provides data processing apparatus including means arranged to carry out the method of the first aspect.

[0093] Thus, according to a second aspect of the present invention, there is provided a data processing device for manipulating an initial video stream captured by a camera in a video conferencing endpoint into multiple views corresponding to regions of interest, the data processing device comprising: a receiver configured to receive frames of the initial video stream captured by the camera; an object detection unit configured to detect objects of one or more predefined types in frames of the initial video stream; a cropping unit configured to select a plurality of crop regions from a frame of the initial video stream, each crop region including at least one bounding box, each bounding box including a detected object of a predefined type; A transmitter configured to transmit the plurality of crop regions.

[0094] Optionally, the data processing device may comprise a bounding box setting unit configured to set a bounding box around the extent of each object of said one or more predefined types detected by the object detection unit.

[0095] The data processing device may be configured to receive an initial video stream from a camera or via a video conferencing endpoint having a camera, wirelessly or via a wired connection, and transmit the crop region to one or more video conferencing endpoints (e.g., computing devices), wirelessly or via a wired connection, for rendering and display at / by the one or more video conferencing endpoints.

[0096] The data processing device may be within or connected to a videoconferencing endpoint. In particular, the data processing device may be within or connected to the same videoconferencing endpoint as the camera. The data processing device may be configured to receive an initial video stream from the camera or via a computing device connected to the camera via a wired connection. Alternatively, the data processing device may be a remote server.

[0097] A third aspect of the present invention provides a video conferencing system for manipulating an initial video stream into multiple views corresponding to regions of interest, the system comprising: a data processing device of the second aspect; a camera configured to capture an initial video stream and transmit the initial video stream to a data processing device; and one or more video conferencing endpoints, wherein the data processing device is configured to receive an initial video stream from a camera and transmit a crop region to the one or more video conferencing endpoints for display.

[0098] Thus, according to a third aspect of the present invention, there is provided a video conferencing system for manipulating an initial video stream into multiple views corresponding to regions of interest, the system comprising: With a camera configured to capture the initial video stream; one or more videoconferencing endpoints; a data processing device, the data processing device comprising: a receiver configured to receive frames of the initial video stream captured by the camera; an object detection unit configured to detect objects of the one or more predefined types in frames of the initial video stream; a cropping unit configured to select a plurality of crop regions from a frame of the initial video stream, each crop region including at least one bounding box, each bounding box including a detected object of a predefined type; A transmitter configured to transmit the plurality of crop regions to the one or more videoconferencing endpoints.

[0099] The one or more video conferencing endpoints may be, for example, computing devices or mobile devices, and may be configured to render and display the crop regions. The one or more video conferencing endpoints may be configured to display the crop regions as separate frames of a separate final video stream or as a composite view in a single frame of the final video stream.

[0100] In particular, if each crop region is transmitted from the data processing device as a separate frame in a separate final video stream, multiple final video streams may be rendered and displayed at the videoconferencing endpoint, each final video stream including a single crop region from each frame of the initial video stream. Thus, each object of a predefined type (e.g., each person) is displayed in its own separate video stream as if it were a single person in its own position, regardless of whether multiple people are in the camera's field of view.

[0101] Alternatively, when the magnified crop regions are combined into a composite view, a single final video stream may be rendered and displayed at the one or more videoconferencing endpoints, the single final video stream including a layout (e.g., a grid) that includes each crop region from each frame of the initial video stream.

[0102] A fourth aspect of the present invention provides a computer-readable storage medium having instructions thereon which, when executed by a data processing device, cause the data processing device to carry out the method of the first aspect.

[0103] The present invention includes combinations of the described aspects and preferred features except where such combinations are clearly impermissible or expressly avoided. [Brief explanation of the drawings]

[0104] Illustrative embodiments of the principles of the present invention will now be discussed with reference to the accompanying figures. [Figure 1] 1 shows a schematic diagram of a videoconferencing system including multiple videoconferencing endpoints. [Figure 2] 1 shows a schematic diagram of a data processing device that may be included in a videoconferencing endpoint. [Figure 3] 3 illustrates steps of a method performed by the data processing device of FIG. 2; [Figure 4] 1 illustrates method steps for selecting a crop region from a first frame of an initial video stream captured by a camera in a videoconferencing endpoint. [Figure 5] 10 illustrates steps of a method for selecting a crop region from a second frame of an initial video stream captured by a camera in a videoconferencing endpoint. [Figure 6] 6 illustrates a method including additional steps that can be performed in the method of FIG. 5. [Figure 7a] An exemplary first frame of the initial video stream is shown. [Figure 7b] Figure 7a shows the bounding box in an exemplary first frame of the initial video stream. [Figure 7c] Figure 7a shows the crop region selected and transmitted from an exemplary first frame of the initial video stream. [Figure 8]a and b show how the first frame of the initial video stream is manipulated into crop regions, each of which contains multiple bounding boxes and therefore multiple objects of a predefined type. [Figure 9a] FIG. 1 is a diagram illustrating the advantage of manipulating the initial video stream captured by a camera in a video conferencing endpoint into multiple views corresponding to regions of interest. [Figure 9b] FIG. 1 is a diagram illustrating the advantage of manipulating the initial video stream captured by a camera in a video conferencing endpoint into multiple views corresponding to regions of interest. [Figure 9c] FIG. 1 is a diagram illustrating the advantage of manipulating the initial video stream captured by a camera in a video conferencing endpoint into multiple views corresponding to regions of interest. [Figure 10a] An example frame of the initial video stream is shown. [Figure 10b] An example frame of the final video stream is shown. DETAILED DESCRIPTION OF THE INVENTION

[0105] Aspects and embodiments of the present invention will now be discussed with reference to the accompanying figures. Further aspects and embodiments will be apparent to those skilled in the art.

[0106] FIG. 1 illustrates a videoconferencing system 1 including multiple videoconferencing endpoints 10a-10d, each of which may be controlled by a respective user. The videoconferencing endpoints 10a-10d (and therefore users) are located in different locations and may include computing devices, mobile devices, tablets, or dedicated videoconferencing devices. The videoconferencing system 1 enables both video and audio signals to be transmitted between the videoconferencing endpoints 10a-10d over a digital network 50. The digital network 50 is preferably wireless, but may also be wired.

[0107] A video camera (such as video camera 30 in videoconferencing endpoint 10a) and a microphone (not shown) are disposed in each of videoconferencing endpoints 10a-10d to provide video and audio input, respectively. A screen, display, monitor, television, and / or projector and speakers (not shown) are disposed in each of videoconferencing endpoints 10a-10d to provide video and audio output, respectively.

[0108] 1, one of the videoconferencing endpoints 10a includes a data processing device 20. In some other exemplary videoconferencing systems, each of the videoconferencing endpoints 10a-10d may include such a data processing device 20.

[0109] A camera 30 at the videoconferencing endpoint 10a is configured to capture an initial video stream (the initial video stream including multiple video frames each including one or more users (i.e., people) located within the camera's field of view) and transmit the initial video stream to the data processing device 20 (e.g., via a connection 40, which may be wired or wireless).

[0110] FIG. 2 illustrates schematically an exemplary data processing device 20 that may be located in the videoconferencing endpoint 10a of the videoconferencing system 1 shown in FIG.

[0111] The data processing device 20 comprises a receiver 21 configured to receive (wirelessly or via a wired connection) an initial video stream from a camera 30 .

[0112] The data processing device 20 also has an object detection unit 22, a bounding box setting unit 23, and a cropping unit 24, which together are configured to manipulate the initial video stream received by the receiver 21 into multiple views (or "crop regions") corresponding to regions of interest in the initial video stream. The regions of interest may, for example, correspond to portions of the initial video stream that include people.

[0113] The data processing device 20 also has a transmitter 25 configured to transmit (wirelessly or via a wired connection) the cropped region of interest to the videoconferencing endpoints 10a-10d for rendering and display thereat.

[0114] FIG. 3 illustrates steps of a method 100 performed by the data processing device 20 of FIG.

[0115] First, at S.110, the receiver 21 of the data processing device 20 receives an initial video stream captured by the camera 30 at the videoconferencing endpoint 10a.

[0116] Next, at S.120, object detection unit 22 detects one or more predefined type objects (e.g., the predefined type may be, for example, a person) in a first frame of the initial video stream received by receiver 21. In particular, object detection unit 22 predicts a class of an object in the first frame to identify, recognize, or classify the object in the first frame as a person. Object detection unit 22 may locate (e.g., find the position of) the person in the first frame.

[0117] S.120 may be performed using a machine learning algorithm, for example, using a trained neural network such as a region-based convolutional neural network (R-CNN). Alternatively / additionally, S.120 may be performed using a Haar feature-based cascade classifier, or a histogram of oriented gradients (HOG).

[0118] In S.130, the bounding box setting unit 23 sets a rectangular bounding box around the extent of each detected object such that each bounding box contains a single object of a predefined type. The predefined object types may be, for example, a person, a person's face, a person's head, and / or a person's head and shoulders. The predefined object type may be, for example, a fixed portion of a person's upper body from the top of the head as a proportion of the size of the person's face. Thus, in one example of S.130, a bounding box is set around the extent of each person's head and shoulders in the first frame.

[0119] The / each bounding box set around the extent of each person's head and shoulders in the first frame may include a label containing information about its classification (e.g., that the bounding box contains the person's head and shoulders) and information about the position of the person's head and shoulders (e.g., coordinates).

[0120] In some examples, the position information in the bounding box label may include the coordinates (e.g., horizontal and vertical coordinates) of the four corners of the rectangular bounding box, or the coordinates (e.g., horizontal and vertical coordinates) of a known point on the rectangular bounding box (e.g., the center or the lower-left corner) and the dimensions (e.g., width and height) of the rectangular bounding box.

[0121] At S.140, cropping unit 24 selects crop regions from the first frame. Each crop region may include at least one bounding box (and thus at least one person, e.g., at least one head and shoulders of a person). In some instances, a crop region may include multiple bounding boxes, such as when two bounding boxes are located close to each other in the frame.

[0122] In S.150, transmitter 25 transmits the crop region to videoconferencing endpoints 10a-d for rendering and display thereat. Transmitter 25 may transmit only the crop region, such that any portion of the first frame that does not include the one or more predefined types of objects (e.g., a portion that does not include a person's head and shoulders) is not transmitted. Thus, only the region of interest of the first frame is transmitted. However, in other examples, the first frame of the initial video stream may be transmitted in addition to the crop region, such that the videoconferencing endpoints continue to view the original first frame, but with the one or more predefined types of objects (e.g., people) emphasized in the crop region.

[0123] The cropped regions from the first frame of the initial video stream may be transmitted as the first frame of each of the separate final video streams, i.e., the transmitter 25 may transmit the cropped regions separately.

[0124] Alternatively, crop regions from the first frame of the initial video stream may be combined into a composite view and transmitted as a single first frame of the final video stream. The composite view may include each of the crop regions, such that transmitter 25 transmits the crop regions together in a single frame of a single final video stream.

[0125] 4 illustrates steps of a method 200 for selecting multiple crop regions from a first frame of an initial video stream (these are therefore substeps of S.140 of FIG. 3). The steps of method 200 are performed by cropping unit 24.

[0126] In S.210, a limited number of possible candidate sets of crop regions, including the bounding box, are determined from the first frame, which helps reduce the computation time and space requirements of cropping unit 24 by selecting a preferred (or even optimized) crop region from the limited number of possible candidate sets of crop regions, rather than all possible candidate sets.

[0127] Determining the possible candidate sets of crop regions in S.210 includes determining (e.g., identifying) all possible candidate sets that satisfy a predefined preferred aspect ratio for each crop region, and then limiting the number of possible candidate sets. The predefined preferred aspect ratio for each crop region may be stored in the data processing device 20 or may be received from a remote device (e.g., a remote server) via the receiver 21.

[0128] Restricting the set of possible candidates for crop regions may include one or more of the following: • Restricting the set of possible candidate crop regions to a set of possible candidate crop regions in which each crop region in the candidate set contains no more than a predefined number of bounding boxes; • Restrict the set of possible candidates for the crop region to the set of possible candidates whose all bounding boxes in the first frame are contained in the crop region; restricting the set of possible candidate crop regions such that the number of crop regions in each possible candidate set is less than a first predefined threshold and / or greater than a second predefined threshold; and / or Restrict the set of possible candidate crop regions so that the crop region in each possible candidate set includes one or more bounding boxes set around one of a predefined set of predefined object types (e.g., so that each crop region in each possible candidate set includes a person's head or a person's head and shoulders, but does not include, for example, an entire person).

[0129] In S.150, if the crop regions are transmitted as separate first frames of separate final video streams (i.e., each first frame of each final video stream includes a single crop region), limiting the possible candidate sets of crop regions may include limiting the possible candidate sets of crop regions such that the number of crop regions in each possible candidate set is less than or equal to a predefined maximum number of final video streams. The predefined maximum number of final video streams may be stored in data processing device 20 or received from a remote device (e.g., a remote server). The predefined maximum number of final video streams may be determined based on network bandwidth limitations, network capacity, or limitations of videoconferencing system 1 (e.g., limitations of one or more videoconferencing endpoints 10a-10d or limitations of data processing device 20).

[0130] In S.150, if crop regions are grouped into a composite view and transmitted as a single first frame of the final video stream (i.e., the first frame of the final video stream will include all crop regions), limiting the possible candidate sets of crop regions may include limiting the possible candidate sets of crop regions such that the crop regions in each possible candidate set satisfy one or more predefined requirements of the composite view.

[0131] For example, the one or more predefined requirements for the composite view may include a predefined overall size, a predefined shape, a predefined aspect ratio, and / or a predefined maximum number of rectangles in the layout that form the composite view, where a crop region is provided in each rectangle in the layout when the crop regions are combined into the composite view.

[0132] Thus, in one example, limiting the possible candidate sets of crop regions may include limiting the possible candidate sets of crop regions such that the number of crop regions in each possible candidate set is less than or equal to a predefined maximum number of rectangles in the layout that form the composite view. The predefined maximum number of rectangles in the layout that form the composite view may be stored on data processing device 20 or may be received from a remote device (e.g., a remote server). The predefined maximum number of rectangles in the layout that form the composite view may be determined based on network bandwidth limitations, network capacity, or limitations of videoconferencing system 1 (e.g., limitations of one or more videoconferencing endpoints 10a-10d or limitations of data processing device 20).

[0133] In S.220, parameters of a set of preferred crop regions are determined from the set of possible candidates. In particular, the parameters of each crop region are determined by minimizing a cost function with respect to the determined set of possible candidate crop regions. The cost function includes a weighted sum of (i) a term that correlates to the area of ​​the crop region that does not overlap any bounding box, and (ii) a term that correlates to the overlap between the crop regions.

[0134] Each rectangular crop area may be defined by four parameters, which may be the horizontal and vertical coordinates of its four corners, or the horizontal and vertical coordinates of the center of the crop area and the dimensions of the crop area (e.g., width and height).

[0135] In S.230, crop regions from the first frame are selected based on the determined parameters of each crop region, as shown in FIG. 2, and transmitted in S.150.

[0136] In another method of selecting multiple crop regions from the first frame of the initial video stream (hence, these are substeps of S.140 in FIG. 3), a cost function may be optimized with respect to one or more constraints in a cost optimization process or minimized using constrained optimization. This cost function may be minimized using one or more mathematical solvers, such as linear programming, interior point methods, or generalized reduced gradient methods. Alternatively, other mathematical tools, such as genetic algorithms, may be used to minimize the cost function.

[0137] Constraints may relate to one or more of the following: ● The number of bounding boxes in each crop region (e.g., you want no more than two bounding boxes in each crop region); a predefined preferred aspect ratio for each crop area; and / or Predefined object types, which are parameterized, for example, by a maximum and / or minimum ratio of a fixed part of a person's upper body from the top of the head to the size of the person's face.

[0138] If in S.150 crop regions are transmitted as separate first frames of separate final video streams (i.e., each first frame of each final video stream will contain a single crop region), the constraint may also relate to a predefined maximum number of final video streams.

[0139] If in S.150 crop regions are combined into a composite view and transmitted as a single first frame of the final video stream (i.e. the first frame of the final video stream will contain all crop regions), the constraint may also relate to a predefined maximum number of rectangles in the layout that form the composite view.

[0140] The arrangement of the layout (e.g., grid) that forms the composite view including each crop region may be determined by an optimization process, which may be performed by data processing device 20 or a remote device.

[0141] To reduce computation time and space requirements, possible candidates for the placement of the layout (e.g., placement of rectangles in a grid) may be determined (e.g., identified) before determining the placement of the layout from a limited number of possible candidates for placement for the layout. In particular, possible candidates for aspect ratios of the rectangles in the layout are determined.

[0142] Determining possible candidates for layout placement includes determining (eg, identifying) all possible candidates for layout placement and then limiting the number of possible candidates.

[0143] Restricting the possible candidates for the placement of the layout may include restricting the possible candidates to those with a predefined set of aspect ratios for each rectangle in the layout.

[0144] If the method includes a cost optimization process in which a cost function is minimized using constrained optimization, the constraints may relate to one or more of an optimized aspect ratio of each rectangle in the layout, a predetermined point (e.g., the top left corner of the layout), and a scaling factor, where the scaling factor is the amount by which each crop region is scaled in size to fit into a respective rectangle in the layout that forms the composite view.

[0145] 5 illustrates steps of a method 300 performed by data processing device 20 to select a crop region from a subsequent (e.g., second, third, fourth, ... Nth) frame of the initial video stream. In particular, FIG. 5 illustrates steps for selecting a crop region from a second frame of the initial video stream.

[0146] Similar to S.110 of method 100, at S.310, receiver 21 receives an initial video stream captured by camera 30 at videoconferencing endpoint 10a. At S.320 and S.330, object detection unit 22 detects one or more objects of a predefined type in a second frame of the initial video stream, and bounding box setting unit 23 sets a bounding box around the extent of each detected object in the second frame. Thus, each bounding box in the second frame may contain a single object of the predefined type. S.320 and S.330 may be performed in a manner corresponding to S.120 and S.130, as described above with reference to FIG. 3.

[0147] In contrast to S.140, in S.340, a crop region from a second frame of the initial video stream is selected based on the crop region from the first frame and the tracked movement of the bounding box between the first and second frames, thereby reducing the computational load on cropping unit 24 compared to a method in which crop regions from the second and each subsequent frame are selected in the same manner as the first frame, while still providing a consistent view of objects of the one or more predefined types as the objects move during the initial video stream.

[0148] Preferably, the number of crop regions in the first and second (and any subsequent frames) is consistent, and each crop region from the second frame (and any subsequent frames) contains the same object(s) of the predefined type as the corresponding crop region from the first frame, providing a consistent view of the objects of said one or more predefined types and crop regions.

[0149] In S.140, the movement of bounding boxes between the first and second frames may be tracked using a feature mapping algorithm or by using an optical flow estimation algorithm to track the movement of pixels within each bounding box. Alternatively / additionally, each bounding box in the second frame may be identified as corresponding to a bounding box in the first frame based on the position information of the bounding boxes in each frame, and the bounding box whose position is shifted by the smallest distance between the first and second frames is identified as the corresponding pair of bounding boxes. Thus, crop regions from the second frame may then be selected based on the corresponding pair of bounding boxes such that each crop region in the second frame contains the same bounding box as the corresponding crop region in the first frame.

[0150] Similar to S.150, in S.350 a cropped region from the second frame is transmitted by the transmitter 25.

[0151] Figure 6 shows method 400 including additional steps that may be performed between S.340 and S.350 of Figure 5. Method 400 may be performed by data processing device 20. The steps shown in Figure 6 may be performed in any order or simultaneously (e.g., S.420 may be performed before, simultaneously with, or after S.410).

[0152] At S.410, any substantially overlapping crop regions from the second frame are merged to avoid duplication of the one or more predefined types of objects (e.g., people) in the second frame. In particular, the amount of overlap between pairs of crop regions from the second frame is determined, and if the amount of overlap exceeds a predefined overlap threshold, the pairs of crop regions are merged.

[0153] In S.420, any crop region from the second frame that is too large (e.g., if the crop region includes two people and the two people move further apart between the first and second frames) is split into two or more resulting crop regions. In particular, the size of each crop region from the second frame is determined, and any crop region from the second frame that has a size that exceeds a predefined size threshold is split into two or more resulting crop regions.

[0154] At S.430, it is determined whether the crop region from the second frame satisfies predefined criteria. In particular, if the crop region does not satisfy the predefined criteria, it is determined that too much movement of a predefined type of object has occurred between the first and second frames, and the crop region from the second frame needs to be reselected using one of the methods described above for selecting a crop region from the first frame. Thus, at S.440, if the crop region from the second frame does not satisfy the predefined criteria, the crop region from the second frame is reselected using the method described above that was used when selecting a crop region from the first frame (e.g., by a cost function).

[0155] The step of determining whether the crop region from the second frame satisfies a predefined criterion may be performed according to a metric, where the metric may be a return value of the cost function described above, and in particular, the predefined criterion may be that the metric is less than a predefined threshold.

[0156] As shown in FIG. 6, if the crop regions meet the predefined criteria, they are not reselected.

[0157] Crop regions from further subsequent frames of the initial video stream (e.g., the third, fourth, fifth, ... Nth frames) may be selected in a corresponding manner as the second frame, as described above with reference to Figures 5 and 6.

[0158] A further example of how to allocate several bounding boxes and thus assign people to distinct final video streams, given a maximum number of final video streams NS, is now described.

[0159] In the first step, the maximum number of final video streams is temporarily set to infinity. The above method is then executed to obtain a set of crop regions in the initial frame. This often produces non-overlapping crop regions, where people within the crop regions overlap or are very close to each other, with each crop region closely bounding each person. If the set of crop regions has more than NS crop regions, these crop regions are grouped into N groups, with each group having a similar or equal number of bounding boxes / people within it. The crop regions within each group are selected to be spatially close, and the groups maintain spatial ordering of people between groups, if possible. If the number of people / bounding boxes is NP, the target number of people per group may be NPG = (NP + NS - 1) / NS. In one example, the algorithm may involve first ordering the crop regions in spatial order, such as from left to right, and then iteratively progressing through the remaining crop regions. If a crop region contains exactly NPG people or more, or if there are no pending groups (i.e., groups with less than NPG people), a new group is created containing only all people / bounding boxes within that crop region. Otherwise, all people / bounding boxes within that crop region are added to that pending group, and the group pending state is updated. Each final group contains a collection of people / bounding boxes for each stream.

[0160] After partitioning into a set of distinct streams, each stream contains a set of bounding boxes, and the algorithm described above with reference to Figure 4 can be applied to each stream to obtain a set of crop regions that are combined into a composite view. For each subsequent frame, the method can track whether a distinct stream is created or destroyed when a person enters or leaves the camera's field of view or when movement causes significant overlap in crop regions. The availability of a created final video stream or the unavailability of a destroyed final video stream can be signaled or advertised over the network. Excessive signaling can be avoided by reusing the unique identifiers of now-destroyed streams for newly created streams. In the case of multiple simultaneously created or discarded streams, the method can find the closest pair of reused identifiers to avoid excessive changes in the content of streams with the same identifier. The metric used to determine closeness can be the size of the collection of identical people / bounding boxes.

[0161] The transition in each final video stream may be animated to avoid abrupt changes when the crop region number changes. In some cases, the transition is a fade over a period of time.

[0162] As discussed previously, each final video stream may be exposed as a camera or virtual camera, allowing applications to stream independently using the same standard interface. Metadata may be transmitted along with each final video stream. The metadata may include information about spatial ordering, whether active speakers are located in each stream, and the number of people in a given stream. This allows for more sophisticated layouts on both the sender and receiver sides. Metadata may be associated with each virtual camera and / or the frames generated from these virtual cameras. Most receivers have some static and some dynamic information associated with the subscribed camera. The static information does not change with each generated frame. The static information about the virtual cameras may include an indication of which physical camera each virtual camera is derived from. The per-frame dynamic information may indicate the active speakers present and the number of people in the frame. The per-frame dynamic information may also include information about the spatial ordering of the final video stream, such as the left-to-right order or the crop coordinates used to create each frame. In one example, the receiver may be running the Android operating system, and the static information may be stored in camera characteristic metadata and the dynamic information may be stored in camera capture results metadata.

[0163] The background within a given crop region or frame may be blurred or replaced to avoid clutter caused by cropping to reduce overlap between nearby crop regions. A mask for each person may be obtained from a depth sensor or a pre-trained machine learning model. Enhancement / highlighting of active speakers may be performed in the final video stream, for example, by making them more prominent in the composite view. As previously discussed, active speakers may be detected via a microphone array that calculates the sound direction and matches that sound direction with the identified presence of a person. A separate stream containing that person can then be identified, and a larger layout can be allocated for that person in the final composite view of the separate streams.

[0164] 7a shows a first frame 1000 of an initial video stream captured by a camera, such as camera 30 of FIG. 1. The camera's field of view includes four people. Therefore, if first frame 1000 of the initial video stream is transmitted to multiple video conferencing endpoints, e.g., video conferencing endpoints 10a-10d, the first frame of the final video stream rendered and displayed at the video conferencing endpoints will include each of the four people, but because they are not positioned an optimal distance from the camera, each person will appear smaller than if they each had their own dedicated camera. In particular, the first frame of the final video stream will look like first frame 1000.

[0165] However, if the first frame 1000 of the initial video stream is manipulated to crop regions of interest according to the methods disclosed herein, any regions containing people will be emphasized when rendered and displayed on a videoconferencing endpoint. In particular, Figure 7c shows four crop regions 1200a-d that may be displayed on a videoconferencing endpoint. The four crop regions 1200a-d emphasize the four people, so that any regions of the first frame 1000 that do not contain people will be reduced from the final video stream.

[0166] 7b shows exemplary bounding boxes 1100a-d set around the extent of each object of a predefined type (in this case, a fixed proportion of a person's head and shoulders), which are used in selecting the crop region.

[0167] In Figure 7c, each crop region contains one bounding box and therefore one person.

[0168] As previously mentioned, the crop regions may be transmitted as separate first frames of separate final video streams, or as a composite view that includes each crop region to form a single first frame of a single final video stream.

[0169] The crop regions may be combined into a composite view that is transmitted as a single first frame of a single final video stream. As can be seen with reference to Figures 7a-7c, the relative spatial order or relative positions of people in the first frame 1000 of the initial video stream (shown in Figure 7a) may be maintained in the layout (e.g., grid) that forms the first frame of the final video stream (shown in Figure 7c). In this way, a person sitting to the left of another person in the first frame 1000 of the initial video stream will also be observed to be sitting to the left of that person in the layout that forms the first frame of the final video stream.

[0170] Figures 8a and 8b are similar to Figures 7a and 7c, respectively, except that a first frame 1300 (shown in Figure 8a) of an initial video stream captured by a camera is manipulated into crop regions, each of which contains multiple bounding boxes (rather than just one) and thus multiple objects of a predefined type. In particular, as shown in Figure 8b, two crop regions 1400a, 1400b are selected from the first frame 1300, each of which contains two bounding boxes and thus two people.

[0171] 9a-9c illustrate the benefits of manipulating an initial video stream captured by a camera at a videoconferencing endpoint into multiple views corresponding to regions of interest. In particular, in FIG. 9a, there are three videoconferencing endpoints involved in a videoconferencing call, each represented by a respective initial video stream 2100, 2110, and 2120. Initial video streams 2110 and 2120 each include a single person, such that one person is located within the field of view of the camera at the corresponding videoconferencing endpoint. However, initial video stream 2100 includes three people, because all three people are located within the field of view of the camera at the corresponding videoconferencing endpoint. As can be seen in FIG. 9a, the three people in the initial video stream appear smaller than the people in initial video streams 2110 and 2120.

[0172] According to the disclosed method, crop regions of the initial video stream 2100 are selected to highlight three people in the initial video stream 2100. Figures 9b and 9c show how crop regions of each frame of the initial video stream 2100 are selected to extract crop regions of interest 2200a-c (i.e., to extract three regions of each frame of the initial video stream 2100a that contain people).

[0173] Thus, FIG. 9c shows an exemplary display 2300 of each frame of the final video stream corresponding to initial video streams 2100, 2110, and 2120. Final video streams 2310 and 2320 correspond to initial video streams 2110 and 2120, respectively. However, initial video stream 2100 is partitioned into three final video streams 2301, 2302, and 2303, each of which corresponds to a crop region 2200a-c of initial video stream 2100. Thus, regardless of the number of people at each videoconferencing endpoint and their distance to the camera, each person will appear to be a similar size when displayed on the videoconferencing endpoint.

[0174] 10a and 10b show further example frames of the initial video stream 3100 and final video streams 3200a-3200c, respectively.

[0175] A frame of initial video stream 3100 contains three people, each located at a different distance from a single camera, and therefore the people are different sizes. However, according to the methods disclosed herein, multiple crop regions are selected from a frame of initial video stream 3100 to emphasize the people in the frame of initial video stream 3100, so that when crop regions 3200a-3200c are displayed on a videoconferencing endpoint (as shown in FIG. 10b), the people appear to be similar sizes, regardless of the number of people at each videoconferencing endpoint and their distance to the camera.

[0176] The features disclosed in the above description, or the following claims, or the accompanying drawings, whether individually or expressed as means for performing a disclosed function or a method or process for obtaining a disclosed result, may be utilized, as appropriate, separately or in any combination of such features, to realize the invention in diverse forms thereof.

[0177] While the present invention has been described in conjunction with the exemplary embodiments above, many equivalent modifications and variations will be apparent to those skilled in the art given this disclosure. Accordingly, the exemplary embodiments of the present invention set forth above are considered to be illustrative and not limiting. Various changes may be made to the described embodiments without departing from the spirit and scope of the invention.

[0178] For the avoidance of doubt, any theoretical explanations provided herein are provided for the purpose of enhancing the understanding of the reader, and the inventors do not wish to be bound by any of these theoretical explanations.

[0179] The section headings used herein, if any, are for organizational purposes only and are not to be construed as limiting the subject matter described.

[0180] Throughout this specification, including the claims which follow, unless the context requires otherwise, the words "have" and "include," and variations such as "having," "including," "including," and the like, will be understood to mean the inclusion of a stated integer or step or group of integers or steps, but not to the exclusion of other integers or steps or groups of integers or steps.

[0181] It should be noted that, as used in the specification and the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. Ranges may be expressed herein from "about" one particular value and / or to "about" another particular value. When such a range is expressed, another embodiment includes from the one particular value and / or to the other particular value. Similarly, when values ​​are expressed as approximations, by use of "about," it is understood that the particular value forms another embodiment. The term "about" with respect to numerical values ​​is optional and means, for example, ±10%.

Claims

1. 1. A method for manipulating an initial video stream captured by a camera at a videoconferencing endpoint into multiple views corresponding to regions of interest, the method comprising: detecting objects of one or more predefined types in frames of the initial video stream; selecting a plurality of crop regions from the frames of the initial video stream, each crop region including at least one bounding box, each bounding box including a detected object of a predefined type; transmitting the plurality of crop regions; combining multiple of the crop regions into a composite view and transmitting the composite view as a single frame of a single final video stream, wherein the relative spatial order of the objects of the one or more predefined types in the frame is maintained in a layout forming the frame of the final video stream; or transmitting each of the crop regions as a separate final video stream, each final video stream including metadata containing information about spatial ordering; method.

2. Each crop region is transmitted as a separate frame in a separate final video stream; or combining a plurality of the crop regions into a composite view, and transmitting the composite view as a single frame in a single final video stream; The method of claim 1.

3. detecting one or more predefined types of objects in the frames is performed using a trained neural network, a Haar feature-based cascade classifier, or a histogram of directed gradients; The method according to claim 1 or 2.

4. Selecting a plurality of crop regions from the frame comprises: Reduce the area of ​​the cropped region that does not overlap any bounding box; including reducing overlap between crop regions; The method according to any one of claims 1 to 3.

5. selecting a plurality of crop regions from the frames of the initial video stream includes parameterizing each crop region by minimizing a cost function, the cost function including a weighted sum of (i) a term correlating to the area of ​​the crop region that does not overlap any bounding box, and (ii) a term correlating to the overlap between the crop regions; The method of claim 4.

6. Selecting a plurality of crop regions from the frames of the initial video stream includes determining a possible candidate set of crop regions before minimizing the cost function, and determining the possible candidate set of crop regions comprises: identifying a set of all possible candidate crop regions that satisfy the predefined preferred aspect ratio for each crop region; restricting the possible candidate set of crop regions to a possible candidate set in which each crop region in the candidate set contains no more than a predefined number of bounding boxes; restricting the set of possible candidates for crop regions to a set of possible candidates whose bounding boxes are all contained in the crop region; and The possible candidate sets of crop regions are divided into: No more than a predefined maximum number of final video streams, if each crop region is transmitted as a separate frame in a separate final video stream; or and constraining the crop regions to no more than a predefined maximum number of rectangles in a layout that form the composite view when the crop regions are combined into a composite view and the composite view is transmitted as a single frame in a single final video stream. The method of claim 5.

7. the cost function is optimized in a cost optimization process with respect to one or more constraints related to one or more of a predefined maximum number of final video streams, a predefined maximum number of rectangles in a layout forming a composite view of the crop regions, a number of bounding boxes, and a predefined preferred aspect ratio for each crop region; 7. The method according to claim 5 or 6.

8. detecting objects of the one or more predefined types in subsequent frames of the initial video stream; setting a bounding box around the extent of each of the detected objects of the one or more predefined types in the subsequent frame of the initial video stream; selecting a crop region from the subsequent frame of the initial video based on the crop region from the frame and a movement of the bounding box between the frame and the subsequent frame. The method according to any one of claims 1 to 7.

9. determining an amount of overlap between a pair of crop regions from the subsequent frame; merging the pair of crop regions from the subsequent frame if the amount of overlap exceeds a predefined overlap threshold. The method of claim 8.

10. determining a size of a crop region from the subsequent frame; further comprising dividing the crop region from the subsequent frame into two or more resulting crop regions if the size exceeds a predefined size threshold.

10. The method according to claim 8 or 9.

11. The layout arrangement is determined by an optimization process based on the aspect ratio of each crop region selected from the frames of the initial video stream and / or the predefined preferred aspect ratio of each crop region. The method according to any one of claims 1 to 10.

12. 1. A data processing device for manipulating an initial video stream captured by a camera at a video conferencing endpoint into multiple views corresponding to regions of interest, the data processing device comprising: a receiver configured to receive frames of the initial video stream captured by the camera; an object detection unit configured to detect objects of one or more predefined types in the frames of the initial video stream; a cropping unit configured to select a plurality of crop regions from the frames of the initial video stream, each crop region including at least one bounding box, each bounding box including a detected object of a predefined type; a transmitter configured to transmit the plurality of crop regions; The transmitter combining multiple of the crop regions into a composite view and transmitting the composite view as a single frame of a single final video stream, wherein the relative spatial order of the objects of the one or more predefined types in the frame is maintained in a layout forming the frame of the final video stream; or transmitting each of the crop regions as a separate final video stream, each final video stream configured to include metadata containing information about spatial ordering; Data processing device.

13. the receiver is configured to receive the initial video stream wirelessly or via a wired connection from the videoconferencing endpoint having the camera or directly from the camera; the transmitter is configured to transmit the plurality of crop regions to one or more video conferencing endpoints, wirelessly or via a wired connection, for rendering and display at / by the one or more video conferencing endpoints.

13. A data processing device according to claim 12.

14. A video conferencing system for manipulating an initial video stream into multiple views corresponding to regions of interest, the system comprising: a camera configured to capture an initial video stream; one or more videoconferencing endpoints; and a data processing device configured to carry out the method according to any one of claims 1 to 11, said data processing device comprising: a receiver configured to receive frames of the initial video stream captured by the camera; an object detection unit configured to detect objects of one or more predefined types in the frames of the initial video stream; a cropping unit configured to select a plurality of crop regions from the frames of the initial video stream, each crop region including at least one bounding box, each bounding box including a detected object of a predefined type; a transmitter configured to transmit the plurality of crop regions to the one or more videoconferencing endpoints; The transmitter combining multiple of the crop regions into a composite view and transmitting the composite view as a single frame of a single final video stream, wherein the relative spatial order of the objects of the one or more predefined types in the frame is maintained in a layout forming the frame of the final video stream; or transmitting each of the crop regions as a separate final video stream, each final video stream configured to include metadata containing information about spatial ordering; Video conferencing system.

15. the one or more videoconferencing endpoints are configured to render and display the crop regions as separate frames of a separate final video stream or in a composite view in a single frame of a final video stream; The system of claim 14.

Citation Information

Patent Citations

  • Image processing apparatus, camera apparatus, and image processing method

    JP2020048149A

  • Video conference device and video conference program

    JP2020053741A

  • Systems and methods for decomposing a video stream into face streams

    US20190215464A1

  • TV conference system, TV conference method, and program

    WO2018061173A1