Image compositing system, image compositing method, and computer program

The image synthesis system addresses issues of overexposure and shadow reflection by detecting presenter position and material area, adjusting exposure and composition to ensure clear visibility and appropriate focus on either the presenter or materials.

JP2026066523APending Publication Date: 2026-04-17CANON KK
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
CANON KK
Filing Date
2024-10-07
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing video synthesis systems struggle with overexposure of presenters due to ambient light, reflection of camera shadows, and the difficulty in maintaining focus on either the presenter or presentation materials, especially when they overlap, leading to hidden information and unclear visibility.

Method used

An image synthesis system that detects the presenter's position and area relative to presentation materials, adjusting exposure and composition to ensure clear visibility and appropriate focus on either the presenter or materials based on their positional relationship.

Benefits of technology

The system effectively synthesizes video data to maintain clear visibility of both the presenter and materials by adjusting exposure and composition, addressing issues of overexposure and shadow reflection, ensuring viewers can focus on intended content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026066523000001_ABST
    Figure 2026066523000001_ABST
Patent Text Reader

Abstract

This system provides a video compositing system that can combine data into the area of ​​a document, depending on the positional relationship between the presenter and the document. [Solution] A video synthesis system comprising: a presenter position detection means for detecting the presenter's position and the area of ​​the material based on video footage from a camera unit; and a synthesis means for replacing the video footage of the material area with the original data of the material and synthesizing it when the presenter's position detected by the presenter position detection means does not overlap with the area of ​​the material shown in the video footage.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a video synthesis system, a video synthesis method, a computer program, and the like.

Background Art

[0002] When distributing the captured video in a glass-covered indoor studio or the like, the captured video may be overexposed due to the influence of external light, or the shadow of equipment such as a camera may be reflected. Also, when the presenter moves around and presents, for example, the presenter may overlap the presentation materials displayed on a large display, making the presentation materials difficult to see.

[0003] In such a case, it is difficult to determine whether to create a video that emphasizes the presenter or the presentation materials.

[0004] As a means for solving the above problems, when a caster explains while displaying a weather map in news or the like, when the caster overlaps the weather map, a method such as making the caster semi-transparent to make the weather map easier to see is adopted.

[0005] Also, for example, in Patent Document 1, a system is proposed in which presentation materials are transmitted as image data and the area of the presenter is transmitted as video data, thereby saving the transmission data amount of the area of the presentation materials and transmitting it.

Prior Art Documents

Patent Documents

[0006]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0007] However, in Patent Document 1, if the presenter is positioned in a way that overlaps with the presentation materials, the presenter will be hidden when the presentation materials are combined. As a result, if the presenter is speaking while looking at the camera or holding a product in their hand while explaining it, important parts of that information will be hidden.

[0008] Furthermore, when a presenter moves around while explaining, the position of the presenter, whether they are positioned so that they do not overlap with the presentation materials, can affect whether the audience should focus on the presentation materials or on the presenter's speech.

[0009] Furthermore, the points you want viewers to focus on will change depending on whether or not the presenter is looking at the presentation materials. Also, as mentioned earlier, when streaming video filmed in an indoor studio with glass walls, the shadows of camera equipment may be reflected in the presentation materials due to the influence of ambient light, or the presenter's face may be overexposed and difficult to see.

[0010] Therefore, the present invention aims to provide an image synthesis system that can synthesize data into the area of ​​a document according to the positional relationship between the presenter and the document.

[0011] The image synthesis system of an embodiment of the present invention is A presenter position detection means that detects the presenter's position and the area of ​​the materials based on the video from the camera unit, The presenter's position detected by the presenter position detection means does not overlap with the material area shown in the video, and the presenter's position is characterized by having a synthesis means that replaces the video of the material area with the original data of the material and synthesizes it. [Effects of the Invention]

[0012] According to the present invention, it is possible to provide an image synthesis system that can synthesize data into the area of ​​a document according to the positional relationship between the presenter and the document. [Brief explanation of the drawing]

[0013] [Figure 1] This is a functional block diagram showing an example configuration of the image synthesis system 100 according to an embodiment of the present invention. [Figure 2] This flowchart shows an example of a video synthesis method in the video synthesis system 100 according to an embodiment of the present invention. [Figure 3] This figure shows an example of a scene in an embodiment of the present invention where the shadow of the camera unit 110 is reflected in the presentation material due to the influence of ambient light, and the presenter is outside the frame of the presentation material. [Figure 4] This figure shows an example of a scene in an embodiment of the present invention where the shadow of the camera unit 110 is reflected in the presentation material due to the influence of ambient light, and the presenter is within the frame of the presentation material. [Figure 5] This figure shows an example of a scene in an embodiment of the present invention where the shadow of the camera unit 110 is reflected in the presentation material due to the influence of ambient light, and the presenter is facing the presentation material from outside the frame of the presentation material. [Figure 6] This figure shows an example of a streamed video showing enlarged presentation materials in an embodiment of the present invention. [Modes for carrying out the invention]

[0014] Embodiments of the present invention will be described below with reference to the drawings. However, the present invention is not limited to the following embodiments. In each drawing, the same reference numeral is used for the same member or element, and redundant explanations are omitted or simplified.

[0015] Figure 1 is a functional block diagram showing an example configuration of an image synthesis system 100 according to an embodiment of the present invention. Note that some of the functional blocks shown in Figure 1 are realized by causing a CPU (not shown) or other computer included in the image synthesis system 100 to execute a computer program stored in memory (not shown) as a storage medium.

[0016] However, some or all of them may be implemented in hardware. As hardware, a dedicated circuit (ASIC), a processor (reconfigurable processor, DSP), etc. can be used. Also, each functional block shown in FIG. 1 does not have to be built in the same housing, and may be constituted by separate devices connected to each other via signal paths.

[0017] As shown in FIG. 1, the video from the camera unit 110 is sent to the video synthesis system 100. The video from the camera unit 110 is sent to the position detection unit 101, the gaze detection unit 102, the exposure adjustment unit 103, and the synthesis unit 104 within the video synthesis system 100, respectively. Incidentally, the camera unit 110 may be included in the video synthesis system 100.

[0018] Then, in the synthesis unit 104, a distribution video is created based on the information from the position detection unit 101 and the gaze detection unit 102. The distribution video created by the synthesis unit 104 is distributed to the network 130 via the network IF 120 and is displayed, for example, by the terminal of a viewer (user).

[0019] Incidentally, for example, the synthesis unit 104 incorporates a CPU etc. as a computer and also functions as control means for controlling the operations of each unit such as the camera unit 110 and the video synthesis system 100 based on a computer program stored in a memory as a storage medium.

[0020] Here, the camera unit 110 is a device capable of capturing image information, and has, for example, an optical lens (not shown) and an image pickup device such as a CCD image sensor or a CMOS image sensor. The camera unit 110 converts the optical image of the subject captured by the optical lens into an image signal by the image pickup device, and performs analog-digital conversion, image adjustment processing, etc. to generate image data.

[0021] The position detection unit 101 detects, for example, the position within the viewing angle of the presenter based on the image information (video) from the camera unit 110. Here, the position detection unit 101 functions as presenter position detection means for detecting the position of the presenter and the area of the material based on the video from the camera unit 110.

[0022] It is assumed that the presenter is making a presentation while displaying the original data of the presentation material on a large display. The position detection unit 101 detects the positional relationship between the presentation material displayed on the display and the presenter, and the composition unit 104 determines whether the presenter is outside the area of the presentation material (material area) displayed on the display or within the material area of the presentation material. That is, the composition unit determines whether the presenter overlaps the material area.

[0023] The material area of the presentation material is preset by the user in advance using a device connected to the video composition system 100 based on the image information from the camera unit 110. The device connected to the video composition system 100 may be the camera unit 110, or may be an unillustrated user terminal or the like connected to the network 130 via the network IF 120.

[0024] The gaze detection unit 102 detects whether the direction of the presenter's gaze is directed at the presentation material based on the image information from the camera unit 110. Specifically, for example, when both eyes of the presenter are shown in the image, it can be determined that the direction of the presenter's gaze is approximately the direction of the camera unit 110.

[0025] Also, the approximate direction of the presenter's gaze can be determined by detecting the direction of the presenter's side face or the back of the head. The gaze detection unit 102 functions as gaze direction detection means for detecting the direction of the presenter's gaze.

[0026] The exposure adjustment unit 103 calculates exposure parameters such as imaging time (charge accumulation time), aperture, and ISO sensitivity based on image information from the position detection unit 101 and the camera unit 110, and adjusts the exposure of the camera unit 110 in order to make the brightness of the presenter's image appropriate. The exposure adjustment unit 103 functions as an exposure adjustment means.

[0027] The synthesis unit 104 functions as a synthesis means, and if the presenter's position detected by the presenter position detection means does not overlap with the material area shown in the video, it replaces the video of the material area with the original data of the material and synthesizes it.

[0028] Furthermore, the synthesis unit 104 can acquire the source data of the presentation materials directly from, for example, the presenter's user terminal, or via the network 130 or network IF 120. The video synthesis system 100 may also pre-store the source data of the presentation materials acquired as described above in a memory not shown.

[0029] The video synthesis system 100 creates a streamed video by controlling whether or not to perform a projection transformation and synthesize the image information from the camera unit 110 with the original data of the presentation materials, based on the detection results of the position detection unit 101 and the gaze detection unit 102. The created streamed video is then distributed to, for example, a user terminal via the network IF 120 and network 130.

[0030] Figure 2 is a flowchart showing an example of a video synthesis method in the video synthesis system 100 according to an embodiment of the present invention. The CPU and other components of the computer within the synthesis unit 104 execute a computer program stored in memory, thereby sequentially performing the operations of each step in the flowchart of Figure 2.

[0031] In step S1001, when the camera unit 110 starts shooting and the video synthesis system 100 starts operating, video distribution begins.

[0032] First, in step S1002, the position detection unit 101 detects the positional relationship between the presenter and the presentation materials based on image information (video) from the camera unit 110. The presenter's position is detected, for example, using a facial recognition system. Here, step S1002 functions as a presenter position detection step, detecting the presenter's position and the area of ​​the materials based on the video from the camera unit 110.

[0033] A facial recognition system is a system that can recognize a person by detecting facial features such as eyes, nose, and mouth through image recognition. Furthermore, by capturing feature points such as the relative position and size of each facial part that has been recognized once, it is possible to track a specific person even when there are multiple people or when people move within the frame.

[0034] Furthermore, regarding the area of ​​the presentation materials within the image, when the camera unit 110 is fixed during shooting, it is necessary to pre-set which area within the field of view is the presentation material area before video distribution. In other words, the above-mentioned material area can be pre-set by the user.

[0035] Furthermore, even when the camera unit 110 cannot be panned, tilted, or zoomed, the position of the presentation materials can be recognized by acquiring and tracking feature points in the area of ​​the presentation materials from the image.

[0036] Next, in step S1003, it is determined whether the presenter is outside the document area of ​​the presentation materials. If the result is Yes, the process proceeds to step S1004; otherwise, the process proceeds to step S1005.

[0037] Figure 3 shows an example of a scene in an embodiment of the present invention where the shadow of the camera unit 110 is reflected in the presentation material due to the influence of ambient light, and the presenter is outside the material area of ​​the presentation material. In such a scene, the result is determined to be Yes in step S1003.

[0038] In Figure 3, the presenter is outside the frame of the presentation materials, so the materials are not obstructed by the presenter during filming. However, the shadow of the camera unit 110 is visible on the presentation materials due to the influence of ambient light.

[0039] As shown in Figure 3, if the presenter is outside the frame of the presentation materials, in step S1004, the compositing unit 104 replaces the video within the material area of ​​the presentation materials with the original data of the presentation materials, composites it, and outputs it as a streamed video. Here, step S1004 functions as a compositing step that, if the presenter's position detected in the presenter position detection step does not overlap with the material area shown in the video, replaces the video of the material area with the original data of the material and composites it.

[0040] The source data for the presentation materials used at this time may be obtained, as described above, for example, directly from the presenter's user terminal or via network 130 or network IF 120. Alternatively, the video synthesis system 100 may store the source data for the presentation materials obtained in advance as described above in a memory not shown and use that source data.

[0041] Furthermore, in step S1004, when replacing and compositing with the original data of the presentation material, the original data of the presentation material is projected and composited according to the shape and tilt of the frame of the presentation material detected in step S1002. As a result, the streamed video can be made into an image without shadows projected within the frame of the presentation material.

[0042] Furthermore, when replacing the original data of the presentation materials, it is desirable to perform the compositing in synchronization with the timing of the change in the original data of the presentation materials. Alternatively, the user may be allowed to pre-configure whether or not to perform this type of compositing using the original data of the presentation materials.

[0043] On the other hand, if the result in step S1003 is "No," that is, if it is determined that the presenter is within the frame of the presentation materials, the process proceeds to step S1005. In step S1005, the presentation materials are created using real video, that is, captured images, to produce the video for distribution.

[0044] Thus, in this embodiment, if the presenter's position detected by the presenter position detection means overlaps with the material area shown in the video, the synthesis means does not replace the video of the material area with the original data of the material.

[0045] In step S1005, the exposure adjustment unit 103 calculates exposure parameters such as imaging time (charge accumulation time), aperture, and ISO sensitivity, primarily using the video signal of the presenter's image area based on image information from the camera unit 110. Based on these exposure parameters, it adjusts the exposure of the camera unit 110.

[0046] Specifically, the exposure is adjusted so that the video signal of the presenter's image area and the video signal of the other areas are weighted, for example, and the resulting average is calculated to reach an appropriate level within a predetermined range. Furthermore, the video signal of the presenter's image area is weighted higher than the video signals of other areas to give it more emphasis.

[0047] Furthermore, exposure adjustment may be performed by feedback control of exposure parameters such as imaging time (charge accumulation time), aperture, and ISO sensitivity, so that, for example, the brightness of the video signal in the presenter's image area, or the average brightness, reaches a predetermined appropriate value.

[0048] In this embodiment, if the presenter's position detected by the presenter position detection means overlaps with the material area shown in the video, the exposure adjustment unit 103 adjusts the exposure by focusing on the video signal of the presenter's area. The area to be adjusted for exposure may be pre-configured by the user.

[0049] Figure 4 shows an example of a scene in an embodiment of the present invention where the shadow of the camera unit 110 is reflected in the presentation material due to the influence of ambient light, and the presenter is within the frame of the presentation material. In such a scene, the result is determined to be No in step S1003.

[0050] Then, as shown in Figure 4, in scenes where the presenter is within the frame of the presentation materials, in step S1005, the imaging time (charge accumulation time), aperture, and ISO sensitivity are adjusted so that the image of the presenter's area is properly exposed, as indicated by the dotted line frame. As a result, the presenter's face and hands are adjusted to proper exposure.

[0051] Furthermore, in step S1005, the compositing unit 104 does not replace the video within the frame of the presentation material with the original data of the presentation material as in step S1004, but instead outputs the real video, which is image information from the camera unit 110, as the streamed video.

[0052] Next, in step S1006, the gaze detection unit 102 detects the direction of the presenter's gaze to determine whether the presenter is facing the direction of the presentation materials. For example, in the examples in Figures 3 and 4, the presenter's gaze is directed towards the camera unit 110, so in step S1006 it is determined to be No, and the process proceeds to step S1008, continuing the current video stream.

[0053] On the other hand, Figure 5 shows an example of a scene in an embodiment of the present invention where the shadow of the camera unit 110 is reflected in the presentation material due to the influence of ambient light, and the presenter is facing the presentation material from outside the frame of the presentation material.

[0054] As shown in Figure 5, if the presenter is facing the direction of the presentation materials, step S1006 is determined to be Yes, and the process proceeds to step S1007. In step S1007, the compositing unit 104 composites the presentation materials by expanding the frame of the presentation materials to fill the entire screen and replacing the area within the frame of the presentation materials with the original data. In other words, in step S1007, the material area is expanded if the gaze direction is towards the material area.

[0055] If you have already replaced the original data of the presentation materials in step S1004, you can simply enlarge the frame (material area) of the presentation materials to fill the entire screen. On the other hand, if you have gone through step S1005, you should enlarge the frame (material area) of the presentation materials to fill the entire screen as described above, and then replace the contents within the frame of the presentation materials with the original data of the presentation materials.

[0056] Furthermore, after replacing the video within the presentation frame with the original data of the presentation, the presentation frame (material area) may be enlarged to fill the entire screen. Alternatively, the user may be allowed to pre-configure whether or not to perform the above enlargement using the compositing method.

[0057] Furthermore, the source data of the presentation materials used when replacing them in step S1007 may be obtained, as mentioned above, for example, directly from the presenter's user terminal or via network 130 or network IF 120. Alternatively, the video synthesis system 100 may pre-store the source data of the presentation materials obtained in advance as described above in a memory not shown and use that source data.

[0058] By performing the process in step S1007, the original data of the presentation material is enlarged to fit the frame of the presentation material and fill the entire screen, allowing viewers to focus more on the content of the presentation material.

[0059] Figure 6 shows an example of a streamed video with enlarged presentation materials in an embodiment of the present invention. As shown in Figure 6, in step S1007, the video is enlarged until the presenter is hidden by the presentation materials and output as a streamed video.

[0060] In step S1007, the frame (material area) of the presentation materials is enlarged to fill the entire screen, but it is also acceptable to enlarge it only slightly instead of enlarging it to fill the entire screen. When doing so, you may want to enlarge it so that the presenter is not hidden, as you may want to check the presenter's facial expressions and movements along with the presentation materials.

[0061] The video stream created in the synthesis unit 104 is then streamed to the desired network 130 via the network IF 120 in step S1008 and displayed, for example, on a user terminal. After that, the process returns to step S1002, and video streaming is periodically performed according to the positional relationship between the presenter and the presentation materials.

[0062] As described above, in this embodiment, by switching whether or not to replace the video within the frame of the presentation material with the original data of the presentation material depending on the positional relationship between the presenter and the presentation material, it is possible to create a streamed video that conforms to the presenter's intentions.

[0063] Although the present invention has been described in detail above based on its preferred embodiments, the present invention is not limited to the above embodiments, and various modifications and combinations of the above embodiments are possible in accordance with the spirit of the present invention, and these are not excluded from the scope of the present invention. Furthermore, some of the above embodiments may be combined as appropriate.

[0064] Furthermore, the present invention includes, for example, a system that realizes the functions of the above embodiment using at least one processor such as a CPU, memory, and circuitry (e.g., an ASIC). Alternatively, multiple processors may be used for distributed processing.

[0065] Furthermore, in order to implement some or all of the control in the above embodiment, a computer program that implements the functions of the above embodiment may be supplied to the video synthesis system, etc., via a network or various storage media.

[0066] Furthermore, the computer (or CPU, MPU, etc.) in the image synthesis system may read and execute the program. In that case, the program and the storage medium in which the program is stored constitute the present invention. The present invention includes the following combinations.

[0067] (Configuration 1) A video synthesis system characterized by comprising: a presenter position detection means that detects the position of the presenter and the area of ​​the material based on video footage from a camera unit; and a synthesis means that, if the position of the presenter detected by the presenter position detection means does not overlap with the area of ​​the material shown in the video footage, replaces the video footage of the material area with the original data of the material and synthesizes it.

[0068] (Configuration 2) The video synthesis system according to Configuration 1, characterized in that when the position of the presenter detected by the presenter position detection means overlaps with the material area shown in the video, the video of the material area is not replaced with the original data of the material.

[0069] (Configuration 3) The video synthesis system according to Configuration 1 or 2, characterized in that when the position of the presenter detected by the presenter position detection means overlaps with the material area shown in the video, the exposure adjustment means performs exposure adjustment by focusing on the video signal of the presenter's area.

[0070] (Configuration 4) The video synthesis system according to Configuration 3, characterized in that, when the position of the presenter overlaps with the material area, the exposure adjustment means performs exposure adjustment so that the brightness of the video signal in the area of ​​the presenter becomes a predetermined appropriate value.

[0071] (Configuration 5) The image synthesis system according to Configuration 3 or 4, characterized in that the region for performing exposure adjustment can be set in advance.

[0072] (Configuration 6) A video synthesis system according to any one of Configurations 1 to 5, comprising a gaze direction detection means for detecting the gaze direction of the presenter, and characterized in that the material area is enlarged when the gaze direction is directed toward the material area.

[0073] (Configuration 7) The video synthesis system according to Configuration 6, characterized in that it is possible to set in advance whether or not to perform the enlargement by the synthesis means.

[0074] (Configuration 8) The image synthesis system according to any one of Configurations 1 to 7, characterized in that the synthesis means projects the original data to match the material area and synthesizes it.

[0075] (Configuration 9) The video synthesis system according to any one of Configurations 1 to 8, characterized in that it is possible to set in advance whether or not to perform the synthesis by the synthesis means.

[0076] (Configuration 10) The video synthesis system according to any one of Configurations 1 to 9, characterized in that the synthesis means performs the synthesis in synchronization with the timing at which the source data is switched.

[0077] (Configuration 11) The video synthesis system according to any one of Configurations 1 to 10, characterized in that the material area can be set in advance.

[0078] (Method) A video synthesis method characterized by comprising: a presenter position detection step of detecting the presenter's position and the area of ​​the material based on video footage from a camera unit; and a synthesis step of replacing the video footage of the material area with the original data of the material and synthesizing the video if the presenter's position detected in the presenter position detection step does not overlap with the area of ​​the material shown in the video footage.

[0079] (Program) A computer program for controlling each means of the video synthesis system described in any one of configurations 1 to 11 by computer. [Explanation of Symbols]

[0080] 100: Video Synthesis System 101: Position detection unit 102: Eye-tracking unit 103: Exposure adjustment unit 104: Synthesis section 110: Camera Department 120: Network Interface

Claims

1. A presenter position detection means that detects the presenter's position and the area of ​​the materials based on the video from the camera unit, A video synthesis system characterized by having a synthesis means that, when the position of the presenter detected by the presenter position detection means does not overlap with the material area shown in the video, replaces the video of the material area with the original data of the material and synthesizes it.

2. The video synthesis system according to claim 1, characterized in that the synthesis means does not replace the video of the material area with the original data of the material when the position of the presenter detected by the presenter position detection means overlaps with the material area shown in the video.

3. The video synthesis system according to claim 1, further comprising exposure adjustment means that, when the position of the presenter detected by the presenter position detection means overlaps with the material area shown in the video, performs exposure adjustment by focusing on the video signal of the presenter's area.

4. The exposure adjustment means is The video synthesis system according to claim 3, characterized in that, if the presenter's position overlaps with the material area, exposure adjustment is performed so that the brightness of the video signal in the presenter's area becomes a predetermined appropriate value.

5. The image synthesis system according to claim 3, characterized in that the region for performing the exposure adjustment can be set in advance.

6. The presenter has a gaze direction detection means for detecting the direction of the presenter's gaze, The image synthesis system according to claim 1, characterized in that the material area is enlarged when the line of sight is directed toward the material area.

7. The image synthesis system according to claim 6, characterized in that it is possible to set in advance whether or not to perform the enlargement by the synthesis means.

8. The image synthesis system according to claim 1, characterized in that the synthesis means performs a projection transformation on the original data to match the data area and then synthesizes it.

9. The video synthesis system according to claim 1, characterized in that it is possible to set in advance whether or not to perform the synthesis using the synthesis means.

10. The video synthesis system according to claim 1, characterized in that the synthesis means performs the synthesis in synchronization with the timing at which the source data is switched.

11. The image synthesis system according to claim 1, characterized in that the aforementioned data area can be set in advance.

12. A presenter position detection step that detects the presenter's position and the area of ​​the materials based on the video from the camera unit, A video synthesis method characterized by comprising: a synthesis step of replacing the video of the material area with the original data of the material if the position of the presenter detected by the presenter position detection step does not overlap with the material area shown in the video.

13. A computer program for controlling each means of the video synthesis system described in any one of claims 1 to 11 by computer.

Citation Information

Patent Citations

  • Apparatus and method for transmitting image as well as image transmission program recording computer readable recording medium

    JP2002077844A