System and method for dynamic natural camera transitions in electronic cameras
By integrating EPTZ camera and microphone arrays in video conferencing devices, combining video processing and audio processing technology, dynamically adjusting the camera view to track speakers, and performing smooth transitions or switching when scene changes, the problem of unclear views in video conferencing is solved, achieving an efficient video conferencing experience.
Patent Information
- Application Number
- CN202080075468.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-09-27
- Filing Date
- 2020-09-24
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2040-09-24
AI Technical Summary
During video conferencing, distal participants may have difficulty seeing the facial expressions of proximal participants and it is difficult to determine who is speaking, resulting in poor results in video conferencing. The prior art requires the user to manually adjust the camera view and the voice tracking camera may fail in the reverberation environment or the speaker turns around.
By integrating EPTZ camera and microphone arrays in a video conferencing device, combining video processing and audio processing techniques, dynamically adjust the camera view to track speakers and perform smooth transitions or switches as scenes change, ensuring that the view in video conferencing is always clear.
This enables scene changes to be completed without user intervention, ensuring that remote participants can always clearly see the facial expressions and speakers of proximal participants, improving the effectiveness and efficiency of video conferencing.
Smart Images

Figure CN114616823B_ABST
Abstract
Description
Background Art
[0001] Typically, cameras in video conferences capture a view that is appropriate for all participants. Unfortunately, far-end participants can lose much of the value of the video because the size of the near-end participant displayed on the far end may be too small. In some cases, far-end participants cannot see the facial expressions of near-end participants and may have difficulty determining who is actually speaking. These issues bring an awkward feel to video conferences and make it difficult for participants to have a productive conversation.
[0002] To deal with poor framing, participants must intervene and perform a series of actions to pan, tilt, and zoom the camera to capture a better view. As expected, manually directing the camera using a remote control can be cumbersome. Sometimes, participants simply don't bother adjusting the camera's view and just use the default wide-angle lens. Of course, when participants do manually frame the camera's view, the procedure must be repeated if the participant changes position during the videoconference or uses a different seating arrangement in a subsequent videoconference.
[0003] Voice tracking cameras with microphone arrays can help direct the camera toward a participant who is speaking during a video conference. While these types of cameras are very useful, they can suffer from some issues. For example, a voice tracking camera can lose track of a speaker when the speaker turns away from the microphone. In very reverberant environments, a voice tracking camera can be pointed at a reflection point instead of the actual sound source. Typical reflections can occur when a speaker turns away from the camera or when the speaker sits at the end of a table. If the reflections are troublesome enough, the voice tracking camera can be directed to point at a wall, table, or other surface instead of the actual speaker.
[0004] One solution as disclosed in U.S. Patent No. 8,248,448, which is incorporated herein by reference, is to use two different cameras, one for the wide angle lens and one for the speaker lens. The speaker view is aimed based on voice tracking, while the wide angle lens remains fixed. The wide angle lens is used when transitioning the speaker view camera between speakers. The speaker view camera image is used when the speaker view camera has been repositioned to a new speaker. This wide view / speaker view arrangement allows the speaker being viewed to be changed without disrupting the motion, but it does require the use of two cameras.
[0005] For these reasons, it is desirable to be able to dynamically customize the views of participants during a video conference based on the context of the conversation, the arrangement of the participants, and who is actually speaking.The disclosed subject matter is directed to overcoming or at least reducing the effects of one or more of the problems set forth above. Summary of the invention
[0006] In an embodiment according to the present invention, scene changes are accomplished pleasingly and without user input or control. Based on the number of speakers and changes in speakers, whether for different individuals or movement of the same speaker, based on the location of the speakers, and based on the overlap of the current scene and the expected scene, a decision is made whether to perform a smooth transition or to make a switch. It has been determined that the decision about switching versus smooth transition is preferably based on the location of the center of the expected new scene and the boundary of the current scene, using a switch if the center is outside the boundary, and using a smooth transition if it is within the boundary. If a smooth transition is to be performed, an easing function (preferably an ease-in and ease-out function) is executed to change the scene. It has also been determined that the preferred value for the smooth transition is to perform the transition at 80 frames assuming operation at 30 frames per second, although values of 60-100 frames are also suitable for providing a pleasing viewing experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] Figure 1 A conference room including several persons and video conferencing endpoints according to the present invention is illustrated.
[0008] Figure 2 yes Figure 1 A first block diagram of a video conferencing endpoint.
[0009] Figure 3 yes Figure 1 A second block diagram of a video conferencing endpoint.
[0010] Figure 4 yes Figure 1 A third diagram of a video conferencing endpoint illustrating various functions performed by the video conferencing endpoint.
[0011] Figure 5 The diagram shows a full scene and a cropped scene from a video conference.
[0012] Figure 6 Illustration of the dimensions of a cropped scene in a video conference relative to the full scene.
[0013] Figure 7 Various easing functions used in transition scenarios according to the present invention are illustrated.
[0014] Fig. 8A A first relationship between two cropping scenes according to the present invention is illustrated.
[0015] Figure 8B A second relationship between two cropping scenes according to the present invention is illustrated.
[0016] Fig. 9 The diagram illustrates the dimensioning between two cropped scenes according to the present invention.
[0017] Fig.10is a flow chart of a view of a video conferencing endpoint according to the present invention. DETAILED DESCRIPTION
[0018] exist Figure 1 In a plan view of FIG. 1 , one arrangement of video conferencing endpoints 10 uses a video conferencing device 80 having a microphone array 60A-B and a camera 50 integrated therewith. A microphone pod 28 can be placed on a table 90, although other types of microphones can be used, such as ceiling microphones, individual table microphones, etc. The microphone pod 28 is communicatively connected to the video conferencing device 80 and captures the audio of the video conference. For its part, the video conferencing device 80 can be incorporated into or mounted on a display and / or video conferencing unit (not shown). Five individuals 92A-92E are seated around the table 90.
[0019] like Figure 2 As shown, Figure 1 A video conferencing device or endpoint 10 in FIG. 1 communicates with one or more remote endpoints 14 via a network 12. Among some common components, the endpoint 10 has an audio module 20 with an audio codec 22 and a video module 30 with a video codec 32. These modules 20 / 30 are operably coupled to a control module 40 and a network module 70.
[0020] During a video conference, camera 50 captures video and provides the captured video to video module 30 and video codec 32 for processing. Preferably, camera 50 is an electronic pan-tilt-zoom (EPTZ) camera. In addition, one or more microphones in microphone pod 28 capture audio and provide the audio to audio module 20 and audio codec 22 for processing. Endpoint 10 primarily uses the audio captured using microphone pod 28 and ceiling mounted microphones, etc., for conference audio.
[0021] Separately, microphone arrays 60A-B having orthogonally arranged microphones 62 also capture audio and provide the audio to audio module 20 for processing. Preferably, microphone arrays 60A-B include vertically and horizontally arranged microphones 62 for determining the location of audio sources during video conferencing. Thus, endpoints 10 primarily use audio from these arrays 60A-B for camera tracking purposes and not for conference audio, although their audio may be used for conferencing.
[0022] After capturing the audio and video, endpoint 10 encodes it using any common encoding standard, such as MPEG-1, MPEG-2, MPEG-4, H.261, H.263, H.264, and H.265. Network module 70 then outputs the encoded audio and video to remote endpoint 14 via network 12 using any appropriate protocol. Similarly, network module 70 receives conference audio and video from remote endpoint 14 via network 12 and sends these to their corresponding codecs 22 / 32 for processing. Finally, speaker 26 outputs the conference audio, and display 34 outputs the conference video. Many of these modules and other components can operate in a conventional manner known in the art, so further details are not provided here.
[0023] Figure 3 is a hardware-centric block diagram of endpoint 10. Exemplary video conferencing device 80 includes a processing unit 502 (such as a DSP or a central processing unit (CPU) or a combination thereof) to perform desired audio and video operations. A memory 504 having both volatile and non-volatile portions includes programs to execute desired modules 506 (such as audio module 20, video module 30, and control module 40, as well as various other audio and video modules) connected to processing unit 502. A network interface 508 (such as an Ethernet interface) is connected to processing unit 502 to allow communication with the far end. An input / output (I / O) interface 510 is connected to processing unit 502 to perform any required I / O operations. An A / D converter block 512 is connected to processing unit 502 and microphone 514. Microphone 514 includes microphone box 28 and one or more directional microphones 60A, 60B. Camera 50 is connected to processing unit 502 to provide near-end video. HDMI interface 518 is connected to processing unit 502 and display 34 to provide video and audio output, display 34 includes speaker 26. It will be appreciated that this is a very simplified schematic diagram of a video conferencing device and that many other designs are possible.
[0024] With the above-described video conferencing endpoints and components understood, the discussion now turns to the operation of the disclosed endpoint 10. First, Figure 4 A control scheme 150 is shown that is used by the disclosed endpoint 10 to conduct a video conference. As previously described, the control scheme 150 uses both video processing 160 and audio processing 170 to control the operation of the camera 50 during a video conference. The video processing 160 and audio processing 170 can be performed individually or combined together to enhance the operation of the endpoint 10. Although briefly described below, several of the various techniques for audio and video processing 160 and 170 are discussed in more detail later. The control scheme 150, video processing 160, and audio processing 170 are preferably programs stored in the module 506 and executed on the processing unit 502.
[0025] In short, the video processing 160 can use the focal length from the camera 50 to determine the distance to the participants, and can use video-based techniques based on color, motion, and facial recognition to track the participants. As shown, the video processing 160 can therefore use motion detection, skin color detection, facial detection, and other algorithms to process the video and control the operation of the camera 50. Historical data of recorded information obtained during the video conference can also be used in the video processing 160.
[0026] For its part, audio processing 170 uses speech tracking using microphone array 60A-B. To improve tracking accuracy, audio processing 170 may use a number of filtering operations known in the art. For example, audio processing 170 preferably performs echo cancellation when performing speech tracking so that the joined sound from the speaker of the endpoint is not picked up as if the speaker from the endpoint is the primary speaker. Audio processing 170 also uses filtering to eliminate non-speech audio from the speech tracking and ignore louder audio that may come from reflections.
[0027] The audio processing 170 may use processing from additional audio cues, such as using a tabletop microphone element or box (28; Figure 1 ). For example, the audio processing 170 can perform speech recognition to identify the speaker's voice and can determine conversation patterns in speech during a video conference. In another example, the audio processing 170 can obtain the direction (i.e., pan) of the source from a separate microphone box (28) and combine it with the position information obtained by the microphone array 60A-B. Because the microphone box (28) can have several microphones positioned in different directions, the orientation of the audio source relative to those directions can be determined.
[0028] When a participant initially speaks, the microphone box (28) can obtain the direction of the participant relative to the microphone box (28). This can be mapped in a mapping table or the like to the position of the participant obtained using the array (60A-B). At some later time, only the microphone box (28) can detect the current speaker, so that only its direction information is obtained. However, based on the mapping table, the endpoint 10 can locate the position of the current speaker (pan, tilt, zoom coordinates), thereby using the mapping information to frame the speaker using the camera 50.
[0029] It should be appreciated that the above is a description of one embodiment of video conferencing device 80 and endpoint 10, and that other configurations of microphones, cameras, processors, etc. may be used to provide speaker location determination and various views.
[0030] Reference now Figure 5 and Figure 6, illustrates the view of a preferred EPTZ camera 50. The resolution of modern electronic cameras is high enough that even a cropped portion of the scene provides enough resolution to provide an enjoyable video conference. The full camera view 602 may contain up to 3840x2160 pixels (referred to as 4K). The cropped scene 604 may then easily have 1920x1080 pixels (referred to as HD). The cropped scene 604 field of view (FOV) has a height h and width w and a center x c ,y c The upper left corner of the cropped scene 604 has coordinate value x 0 ,y 0 , with reference to 0,0 at the upper left corner of the full camera view 602. The lower right corner has coordinates x 1 ,y 1 .
[0031] exist Figure 5 , individual 92C is the speaker, so cropped scene 604 is framed on individual 92C. If individual 92C stops speaking or a different individual starts speaking, cropped scene 604 changes position or uses full camera view 602. However, how the view changes may have an impact on the video conference. Moving a cropped view over a large distance at high speed can be disorienting. Similarly, switching between close cropped views can also be disorienting. In addition, it is well known that changing views too frequently can also be disorienting. According to an embodiment of the present invention, rules are utilized to determine how to move between camera views, such as full camera view to cropped view, cropped view to full camera view, and between two cropped view positions. The rules provide a pleasant experience with the lowest degree of disorientation. The rules of interest in the present disclosure are those related to view movement and view switching, and the rules for changing views are similar to the previous rules.
[0032] Solving for movement first, when considering a transition between two scenes (Scene A and Scene B), an EPTZ transition is created by specifying a different cropped scene or view for each frame of the transition. The variables of each subsequent frame are changed by a specific amount over time to perform a controlled transition. The speed and acceleration of the effective motion are defined by how much change is applied per frame.
[0033] One method for transitioning a variable v from value A to value B at a specific time t is to normalize the range of values of (t) and apply an interpolation function. The normalized output of this function can be applied to the value of each instance of the transition (v i ). The chosen interpolation function (f(t)) will define the characteristics of the perceived motion as the variable (v) changes.
[0034] In the case of EPTZ camera motion, if the technique is applied simultaneously to the center point (x, y) and size (w, h) variables used to describe the two camera scenes (A, B), the perceived motion effect through the transition will be equivalent to the prescribed interpolation function.
[0035] Motion effects commonly used in graphic animations are used to simulate natural camera movements when applied to video output. In an embodiment according to the invention, the function is applied dynamically so that the endpoint selects the appropriate type of motion at runtime and changes properties like a human operator would. Acceleration, deceleration, and speed become inherent properties of the selected function and transition duration, rather than complex input parameters.
[0036] refer to Figure 7 , the simplest function is a linear function f(t)=t, but a linear function transitions between scenes with abrupt starts / stops and even speed, and is therefore not pleasant to experience.
[0037] There are endless polynomial and trigonometric equations that will produce different types of motion with unique accelerations and decelerations. These can be collectively referred to as "easing functions". Figure 7 The various easing functions are illustrated in Figure .
[0038] The main decision in calculating the parameters of the motion effect is to decide how much time the transition should take to complete. Too fast and it will be dizzying, and too slow and it will be boring. The time determines the number of "steps" it takes to iterate through the transition effect. Since this is applied to a camera video stream, the preferred approach is to base the value on the camera's frame rate (fps or frames per second). For example, if a 2 second transition is desired for a camera with a frame rate of 30fps, the number of steps (S) is 60. Once the total number of steps is determined, the easing function is applied to the four variables x, y, h, and w, while determining the bounding box to be used for each frame through the transition.
[0039] The following example takes 60 frames to apply EASE_INf(t)=t from scene A to scene B. 3 Transition. The scenarios are defined in Table 1.
[0040] Table 1
[0041] Scenario A Scenario B <![CDATA[Center point: (x A , y A )]]> <![CDATA[Center point: (x B , y B )]]> <![CDATA[Width: w A > <![CDATA[Width: w B > <![CDATA[Height: h A > <![CDATA[Height: h B >
[0042] The following pseudocode example performs the key calculations:
[0043]
[0044] Once the time parameter S is determined and the easing function is chosen, a brief calculation applied to the four key variables (x, y, w, h) produces the desired result. At each iteration (video frame), the updated cropping parameters are provided to the GPU or video buffer process to scale the video output correctly. Over the selected number of frames, the appropriate transition video effect is created.
[0045] Based on observations of transitions in video conferencing settings, it has been determined that using a transition such as f(t)=3t 2 -2t 3 Or f(t) = 6t 5 -15t 4 An ease-in and ease-out function of +10t at 80 frames (at 30 frames per second (30fps)) provides a pleasing transition. Other frame counts from 60 to 100 provide pleasing transitions, but 80 frames is most preferred. When the frame count exceeds 100 frames, the transition begins to be considered too slow. If it is below 60, the transition may not be considered a transition, but rather a switch. In addition, the number of frames can be changed based on the distance between scenes, but keeping a constant number of frames provides a sense of dynamics to the movement. If 60 frames per second (60fps) is being used, the value is simply doubled. As mentioned above, a variety of other functions can be used for the transition, although functions with abrupt starts or stops are generally considered undesirable. Many changes can be made to the coefficients and polynomials to provide other speed curves that provide pleasing ease-ins and ease-outs.
[0046] Solving the Move vs. Cut Choice In some cases, it may be more appropriate to change the camera view instantly from scene A to scene B. Take some of the following considerations into account when deciding how to decide when to perform a smooth transition or perform a direct cut:
[0047] Will a smooth transition take too long?
[0048] Will the smooth transition go too far?
[0049] Do smooth transitions cause dizziness or disorientation?
[0050] Does direct switching lead to disorientation?
[0051] It has been determined that as the crossover or overlap between the two scenes (A and B) increases, direct switching becomes more disorienting and a smooth transition is preferred. As the crossover decreases and the overlap disappears, a smooth transition becomes more disorienting and a direct switch is preferred.
[0052] It has been determined that in order to balance the comfort level of camera transitions, a simple calculation is applied to decide whether to move smoothly or cut directly between two scenes.
[0053] The center points of scenes A and B are evaluated against the width and height of the current scene (scene A) and are used as an initial calculation to determine the threshold for performing a switch or move operation.
[0054] If the center point of scene B is outside of scene A, a direct transition is selected; otherwise a smooth transition is applied. Fig. 8A and Figure 8B The difference is shown in Fig. 8A In , the cropped region is centered on individual 92C, who is the speaker, and includes individuals 92B and 92D at the edges. Individual 92D becomes the speaker, so the cropped region needs to be moved to position B, where individual 92d is shown in the center. Because the center of scene B is within the boundaries of scene A, a smooth transition using an ease-in / ease-out function is used for the transition. Figure 8B In the example, individual 92A is the speaker, and then individual 92E becomes the speaker. Since scene B is completely outside of scene A, a direct switch from scene A to scene B is used.
[0055] Fig. 9 The variables are illustrated and the decision is determined by the following pseudo code:
[0056]
[0057] The offset of w and h (w o 、h o ) are used to modify the overlap tolerance. If both are set to zero, the effective maximum overlap allowed is basically 1 / 4 of the area of the current field of view. When the offset value is close to the w, h value of scene B (w B 、h B ), the new scene must be completely outside the current scene to trigger a direct cut transition.
[0058] ABS(x A -x B )>(w A / 2)+w o ||ABS(y A -y B )>(h A / 2)+h o
[0059] Another way to calculate the tolerance is to calculate the area of the intersection of the two scenes and base the decision on a value directly related to that value. Since both methods produce equivalent results, the simpler calculation and condition is usually preferred.
[0060] Reference now Fig.10, a flow chart illustrating the operation of determining a particular view is shown. The flow chart illustrates the operation of the control module 40 in cooperation with the audio module 20 and the video module 30 or the control scheme 150 in cooperation with the video processing 160 and the audio processing 170. In step 1002, video is captured from the camera. In step 1004, the received audio is monitored. In step 1006, it is determined whether there is no speaker at the near end. If there is no speaker, the view is zoomed and panned in step 1008 to provide a complete view of the camera. The operation returns to step 1002.
[0061] If there is a speaker in step 1006, then in step 1010 it is determined whether there is only one speaker. If so, then in step 1016 the location of the speaker is determined. In step 1012 it is determined whether it is a different speaker or the speaker has moved. If not, then in step 1014 the current view is output. If it is a new speaker, then in step 1018 a decision is made between a smooth transition or a switch as described above. If a switch is determined to be appropriate, then in step 1020 the switch is made to provide the new view and the operation returns to step 1002. If it is a smooth transition, then in step 1022 an easing function is selected and put into operation to transition to the new speaker or position. The operation returns to step 1002.
[0062] If it is determined in step 1010 that there is not just one speaker, then in step 1024 it is determined whether there are two speakers. If so, then in step 1026 the orientation of the two speakers is determined. In step 1027 it is determined whether there are different speakers or the speakers have moved. If there are no different speakers and no one has moved, then the current view is selected in step 1029. If the speakers are different or have moved, then in step 1028 it is determined whether the two speakers are close together. There are many factors in determining close together. Some factors include avoiding having the same or overlapping background on either side of the split screen in the split screen view, avoiding having the user's outstretched arm appear to need to intrude on the other side of the split screen, and having the speakers separated by more than half of the screen field of view. If they are not close together, then in step 1030 a switch to the split screen view is used to display the two speakers, including adding viewing space if the two speakers are facing each other, rather than just abutting two cropped speaker views. Many factors are used to determine the amount of viewing space added. In one example, the speakers are aligned to the left and right thirds of the screen, leaving 50% to 67% of the screen width as spacing, although speaker size and other adjustments may change the actual amount. Operation returns to step 1002. If the two speakers are close together, the view is zoomed and panned with ease in step 1032 to capture both speakers, with the camera at the center.
[0063] If there are more than two speakers in step 1024, the positions of the speakers are determined in step 1035. In step 1035, it is determined whether there are different speakers or one of the speakers has moved. If so, the view is zoomed and panned with easing to capture all speakers at the near end in step 1036. If there are no different speakers or no one has moved, the current view is selected in step 1038 and the operation returns to step 1002.
[0064] For simplicity, the above operations are just view change logic, and all assume that the change of view is only made after an appropriate waiting period for a specific view, and that the speaker is talking for a period of time sufficient for the view change to occur.
[0065] While the description focuses on the endpoint making the various determinations and transitions, the determinations may also be made in the multipoint control unit (MCU) that is developing views to provide the various endpoints. The MCU receives the full camera view and then develops the various views in a similar manner, particularly if the conference is operating in speaker view mode, but also in continuous presence mode.
[0066] Thus, scene changes (especially using an EPTZ camera) can be accomplished in a pleasing manner and without the need for user input or control. Based on the number of speakers and changes in speakers, whether for different individuals or movement of the same speaker, based on the positions of the speakers and the overlap of the current scene and the expected scene, a decision is made whether to perform a smooth transition or to make a switch. It has been determined that the decision as to whether to switch or smooth transition is preferably based on the position of the center of the expected new scene or the boundary of the current scene, with a switch being used if the center is outside the boundary and a smooth transition being used if it is inside. If a smooth transition is to be performed, an easing function (preferably an ease-in, ease-out function) is performed to change the scene. It has also been determined that, assuming 30fs operation, the preferred value for a smooth transition is to perform the transition at 80 frames, but values of 60-100 frames are also suitable to provide a pleasing viewing experience.
[0067] Various changes may be made to the details of the illustrated method of operation without departing from the scope of the appended claims. For example, the illustrative flow chart steps or process steps may perform the identified steps in a different order than disclosed herein. Alternatively, some embodiments may combine the activities described herein into separate steps. Similarly, one or more of the steps described may be omitted, depending on the specific operating environment in which the method is implemented.
[0068] In addition, the actions according to the flowchart or process steps can be performed by a programmable control device that executes instructions organized into one or more program modules on a non-transitory programmable storage device. The programmable control device can be a single computer processor, a special-purpose processor (e.g., a digital signal processor, "DSP"), multiple processors connected by a communication link, or a custom-designed state machine. The custom-designed state machine can be embodied in a hardware device, such as an integrated circuit, including but not limited to an application-specific integrated circuit ("ASIC") or a field programmable gate array ("FPGA"). Non-transitory programmable storage devices (sometimes referred to as computer-readable media) suitable for tangibly embodying program instructions include but are not limited to: magnetic disks (fixed, floppy and removable) and tapes; optical media, such as CD-ROMs and digital video discs ("DVDs"); and semiconductor memory devices, such as electrically programmable read-only memories ("EPROMs"), electrically erasable programmable read-only memories ("EEPROMs"), programmable gate arrays, and flash memory devices.
[0069] The foregoing description of preferred and other embodiments is not intended to limit or restrict the scope or applicability of the inventive concepts conceived by the applicant. In exchange for disclosing the inventive concepts contained herein, the applicant expects all patent rights provided by the appended claims. Therefore, the appended claims are intended to include all modifications and changes to the full extent that they come within the scope of the appended claims or their equivalents.
Claims
1. A method for operating a video conferencing device to transition between scenes of a room in a video conference, the method comprising: determining a number of speakers in the room; determining the location of any speaker in the room; determining a need to transition a current scene to a new scene based on the determined location of a speaker in the room; determining whether the transition should be a smooth or a switch based on the determined number of speakers in the room, the determined locations of speakers in the room, and the need to transition to a new scene; and performing the transition to the new scene based on the determination of whether to smooth or switch, Among them, it is determined that there is a speaker, wherein it is determined that the speaker is a different speaker or is at a different location, wherein it is determined that a transition to a new scene is to be made, and wherein determining whether the transition should be smooth or switched is based on determining whether a center of the new scene is within the boundaries of the current scene, using smoothing if the center of the new scene is within the boundaries of the current scene, and using switching if the center of the new scene is not within the boundaries of the current scene.
2. The method according to claim 1, wherein: The determination of whether to smooth or switch is smooth, and wherein the transition is performed using an easing function.
3. The method according to claim 2, wherein: The easing function is an ease-in-and-out.
4. The method according to claim 2, wherein: The easing function is executed at 30 frames per second over a range of 60 to 100 frames.
5. The method according to claim 4, wherein: The easing function is executed over 80 frames at 30 frames per second.
6. A method for operating a video conferencing device to transition between scenes of a room in a video conference, the method comprising: determining a number of speakers in the room; determining the location of any speaker in the room; determining a need to transition a current scene to a new scene based on the determined location of a speaker in the room; determining whether the transition should be a smooth or a switch based on the determined number of speakers in the room, the determined locations of speakers in the room, and the need to transition to a new scene; and performing the transition to the new scene based on the determination of whether to smooth or switch, Among them, it is determined that there are two speakers, wherein determining that the speakers are different speakers or in different locations, Among them, it is determined to transition to a new scene. wherein determining whether the transition should be smooth or switched is based on determining whether the two speakers are close together, using smooth if the speakers are close together, and using switched if the speakers are not close together, wherein the smooth transition results in a scene where the camera view is centered between the two speakers, and The switching transition results in a split screen scene of the two speakers, and if the speakers face each other, a viewing space is added to the split screen.
7. The method according to claim 6, wherein: The determination of whether to smooth or switch is smooth, and wherein the transition is performed using an easing function.
8. The method according to claim 7, wherein: The easing function is an ease-in-and-out.
9. The method according to claim 7, wherein: The easing function is executed at 30 frames per second over a range of 60 to 100 frames.
10. The method according to claim 9, wherein: The easing function is executed over 80 frames at 30 frames per second.
11. A non-transitory program storage medium readable by one or more processors in a video conferencing device, and comprising instructions stored thereon, the instructions causing the one or more processors to perform: a method for operating the video conferencing device to transition between scenes of a room in a video conference, the method comprising the steps of: determining a number of speakers in the room; determining the location of any speaker in the room; determining a need to transition a current scene to a new scene based on the determined location of a speaker in the room; determining whether the transition should be a smooth or a switch based on the determined number of speakers in the room, the determined locations of speakers in the room, and the need to transition to a new scene; and The transition to the new scene is performed based on a determination of whether to smooth or switch, wherein it is determined that there is a speaker, wherein it is determined that the speaker is a different speaker or in a different location, wherein it is determined that a transition to a new scene is to be performed, and wherein the determination of whether the transition should be smooth or switch is based on determining whether the center of the new scene is within the boundaries of the current scene, if the center of the new scene is within the boundaries of the current scene, smoothing is used, and if the center of the new scene is not within the boundaries of the current scene, switching is used.
12. The non-transitory program storage medium according to claim 11, wherein: The determination of whether to smooth or switch is smooth, and Therein, the transition is performed using an easing function.
13. The non-transitory program storage medium according to claim 12, wherein: The easing function is an ease-in-and-out.
14. The non-transitory program storage medium according to claim 12, wherein: The easing function is executed at 30 frames per second over a range of 60 to 100 frames.
15. The non-transitory program storage medium according to claim 14, wherein: The easing function is executed over 80 frames at 30 frames per second.
16. A non-transitory program storage medium readable by one or more processors in a video conferencing device, and comprising instructions stored thereon, the instructions causing the one or more processors to perform: a method for operating the video conferencing device to transition between scenes of a room in a video conference, the method comprising the steps of: determining a number of speakers in the room; determining a location of any speakers in the room; determining a need to transition a current scene to a new scene based on the determined locations of the speakers in the room; determining whether the transition should be a smooth or a switch based on the determined number of speakers in the room, the determined locations of the speakers in the room, and the need to transition to the new scene; and performing the transition to the new scene based on the determination of a smooth or switch, wherein it is determined that there are two speakers, wherein determining that the speakers are different speakers or in different locations, Among them, it is determined to transition to a new scene. wherein determining whether the transition should be smooth or switched is based on determining whether the two speakers are close together, using smooth if the speakers are close together, and using switched if the speakers are not close together, wherein the smooth transition results in a scene where the camera view is centered between the two speakers, and The switching transition results in a split screen scene of the two speakers, and if the speakers face each other, a viewing space is added to the split screen.
17. The non-transitory program storage medium according to claim 16, wherein: The determination of whether to smooth or switch is smooth, and wherein the transition is performed using an easing function.
18. The non-transitory program storage medium according to claim 17, wherein: The easing function is an ease-in-and-out.
19. The non-transitory program storage medium according to claim 17, wherein: The easing function is executed at 30 frames per second over a range of 60 to 100 frames.
20. The non-transitory program storage medium of claim 19, wherein: The easing function is executed over 80 frames at 30 frames per second.
21. A video conferencing device for transitioning between scenes of a room in a video conference, the video conferencing device comprising: an electronic pan-tilt-zoom (EPTZ) camera providing a view of the room and having an output; a microphone that provides an output for use in determining a speaker's location; a processor coupled to the EPTZ camera and the microphone and receiving the output from each; A memory coupled to the processor and comprising a program, which when executed causes the processor to perform: a method of operating the video conferencing device to transition between scenes of a room in a video conference, the method comprising the steps of: determining a number of speakers in the room; determining the location of any speaker in the room; determining a need to transition a current scene to a new scene based on the determined location of a speaker in the room; determining whether the transition should be a smooth or a switch based on the determined number of speakers in the room, the determined locations of speakers in the room, and the need to transition to a new scene; and The transition to the new scene is performed based on a determination of whether to smooth or switch, wherein it is determined that there is a speaker, wherein it is determined that the speaker is a different speaker or in a different location, wherein it is determined that a transition to a new scene is to be performed, and wherein the determination of whether the transition should be smooth or switch is based on determining whether the center of the new scene is within the boundaries of the current scene, if the center of the new scene is within the boundaries of the current scene, smoothing is used, and if the center of the new scene is not within the boundaries of the current scene, switching is used.
22. The video conferencing device according to claim 21, wherein: The determination of said smoothing or switching is smooth, and Therein, the transition is performed using an easing function.
23. The video conferencing device of claim 22, wherein: The easing function is an ease-in-and-out.
24. The video conferencing device of claim 22, wherein: The easing function is executed at 30 frames per second over a range of 60 to 100 frames.
25. The video conferencing device of claim 24, wherein: The easing function is executed over 80 frames at 30 frames per second.
26. A video conferencing device for transitioning between scenes of a room in a video conference, the video conferencing device comprising: an electronic pan-tilt-zoom (EPTZ) camera providing a view of the room and having an output; a microphone that provides an output for use in determining a speaker's location; a processor coupled to the EPTZ camera and the microphone and receiving the output from each; A memory coupled to the processor and comprising a program, which when executed causes the processor to perform: a method of operating the video conferencing device to transition between scenes of a room in a video conference, the method comprising the steps of: determining a number of speakers in the room; determining the location of any speaker in the room; determining a need to transition a current scene to a new scene based on the determined location of a speaker in the room; determining whether the transition should be a smooth or a switch based on the determined number of speakers in the room, the determined locations of speakers in the room, and the need to transition to a new scene; and performing the transition to the new scene based on the determination of the smooth or switch, wherein it is determined that there are two speakers, wherein determining that the speakers are different speakers or in different locations, Among them, it is determined to transition to a new scene. wherein determining whether the transition should be smooth or switched is based on determining whether the two speakers are close together, using smooth if the speakers are close together, and using switching if the speakers are not close together, wherein the smooth transition results in a scene where the camera view is centered between the two speakers, and The switching transition results in a split screen scene of the two speakers, and if the speakers face each other, a viewing space is added to the split screen.
27. The video conferencing device of claim 26, wherein: The determination of whether to smooth or switch is smooth, and wherein the transition is performed using an easing function.
28. The video conferencing device of claim 27, wherein: The easing function is an ease-in-and-out.
29. The video conferencing device of claim 27, wherein: The easing function is executed at 30 frames per second over a range of 60 to 100 frames.
30. The video conferencing device of claim 29, wherein: The easing function is executed over 80 frames at 30 frames per second.
31. A computer program product, comprising a computer program, which, when executed by a processor, causes the processor to perform the method according to any one of claims 1 to 5.
32. A computer program product comprising a computer program, which, when executed by a processor, causes the processor to perform the method of any one of claims 6 to 10.
Citation Information
Patent Citations
Automatic camera framing for videoconferencing
US8248448B2
Transition Control in a Videoconference
US20140085404A1
Group and conversational framing for speaker tracking in a video conference system
US20190199967A1
Cited By
System and method for managing and interacting with event information
US12518531B2
System and method for managing and interacting with event information
US20230343097A1