Multispectral volume capture

By combining multispectral infrared and color cameras with IR patterns, the problems of lighting constraints and color overflow in traditional green screen capture technology are solved, achieving free lighting and high-quality video capture.

CN114096995BActive Publication Date: 2025-11-25SONY GROUP CORP +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202080050430.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-12-13
Filing Date
2020-12-11
Publication Date
2025-11-25
Estimated Expiration
2040-12-11

AI Technical Summary

Technical Problem

Traditional green screen capture technology imposes constraints on the lighting of the subject and may introduce 'color overflow' in the green background, affecting the video capture effect.

Method used

By employing a combination of multispectral infrared and color cameras, combined with IR patterns, the final captured data is generated through depth solving and color projection, avoiding the use of a green screen.

Benefits of technology

It enables free lighting of the subject without background constraints, improving the quality and flexibility of video capture and reducing color clipping issues.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114096995B_ABST
    Figure CN114096995B_ABST
Patent Text Reader

Abstract

Video capture of a subject, comprising: a first IR camera, a second IR camera, and a color camera to capture video data of a subject; a bar, wherein the first IR camera, the second IR camera, and the color camera are attached to the bar, and wherein the color camera is positioned between the first IR camera and the second IR camera; at least one IR light source to illuminate the subject; and a processor configured to perform the following operations: generate depth-solved data of the subject using data from the first IR camera and the second IR camera; generate projected color data by using data from the color camera to project color onto the depth-solved data; and generate final captured data by merging the depth-solved data and the projected color data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to video capture of a subject, and more specifically, to video capture using a combination of multispectral infrared and color cameras. Background Technology

[0002] Green or blue screens are frequently used to capture the movement of a subject, which is then composited with a custom background using special effects. However, avoiding green screens can be useful because green screen capture imposes constraints on the lighting of the subject and may also introduce "color bleeding" from the green background onto the subject. Summary of the Invention

[0003] This disclosure provides a method for video capture of a subject using a combination of multispectral infrared (IR) and color cameras with IR patterns.

[0004] In one embodiment, a system for video capture of a subject is disclosed. The system includes: a first IR camera for capturing video data of the subject; a second IR camera for capturing video data of the subject; a color camera for capturing video data of the subject; a pole, wherein the first IR camera, the second IR camera, and the color camera are attached to the pole, and wherein the color camera is positioned between the first IR camera and the second IR camera; at least one IR light source for illuminating the subject; and a processor connected to the first IR camera, the second IR camera, and the color camera, wherein the processor is configured to perform the following operations: generate depth solution data of the subject using data from the first IR camera and the second IR camera; generate projected color data using data from the color camera to project color onto the depth solution data; and generate final capture data by merging the depth solution data and the projected color data.

[0005] In another embodiment, a method for video capture of a subject is disclosed. The method includes: generating depth-solved data of the subject using data from a first IR camera and a second IR camera; generating projected color data using data from a color camera to project color onto the depth-solved data; and generating final capture data by merging the depth-solved data and the projected color data, wherein the color camera is positioned between the first IR camera and the second IR camera.

[0006] In another embodiment, a non-transient computer-readable storage medium is disclosed, which stores a computer program for capturing video of a subject. The computer program includes executable instructions that cause the computer to perform the following operations: generate depth-solved data of the subject using data from a first IR camera and a second IR camera; generate projected color data using data from a color camera to project color onto the depth-solved data; and generate final capture data by merging the depth-solved data and the projected color data, wherein the color camera is positioned between the first IR camera and the second IR camera.

[0007] Other features and advantages should be apparent from this description, which illustrates various aspects of this disclosure by way of example. Attached Figure Description

[0008] Details of the structure and operation of this disclosure can be gathered in part by studying the accompanying drawings, wherein similar reference numerals denote similar parts, and wherein:

[0009] Figure 1 This is a block diagram of a video system for video capture according to one embodiment of the present disclosure;

[0010] Figure 2 This is a flowchart of a method for video capture of a subject according to one embodiment of the present disclosure;

[0011] Figure 3A This refers to a computer system and user representation according to embodiments of the present disclosure; and

[0012] Figure 3B This is a functional block diagram illustrating a computer system for a managed video application according to an embodiment of the present disclosure. Detailed Implementation

[0013] As mentioned above, traditional green screen capture imposes constraints on the lighting of the subject and may also introduce "color overflow" from the green background onto the subject.

[0014] Certain embodiments of this disclosure provide systems and methods for processing video data. In one embodiment, the video system uses a combination of multispectral IR and a color camera, combined with IR patterns, to capture video data of the subject, environment, and background. The system captures volumetric data without using a green screen to utilize depth to separate the subject from the background. This allows for free illumination of the subject without considering the background.

[0015] After reading the following description, it will become clear how this disclosure can be implemented in various embodiments and applications. While various embodiments of this disclosure will be described herein, it should be understood that these embodiments are presented by way of example only and are not intended to be limiting. Therefore, the detailed description of the various embodiments should not be construed as limiting the scope or breadth of this disclosure.

[0016] Features provided in the implementation may include, but are not limited to, one or more of the following: (a) point triangulation of depth based on the structure of the action connection point; (b) a combination of infrared and color cameras; (c) cameras placed in a specific order: IR, color, IR, wherein the left and right IR cameras are used for depth solving with contributions from two or more IR cameras; (d) IR LED floodlights are used to illuminate the subject captured by IR depth to create a matte; (e) a central color camera (which captures visible spectral light data) is used to project color onto the IR camera depth solution; (f) a minimum of one group of three cameras, but depending on the subject, can be extended to N groups of three cameras; and (g) IR absorbing pigments are used to block IR light from returning to the sensor.

[0017] In one implementation, the video system is used in a video production or studio environment and includes one or more cameras for image capture, one or more sensors, and one or more computers for processing camera and sensor data. By using an action-oriented architecture utilizing a combination of IR and color cameras, the system captures volumetric data by illuminating and exposing the subject with complete freedom without using a green screen. Avoiding a green screen is useful because green screen capture imposes constraints on the illumination of the subject and may also introduce "color overflow" from a green background onto the subject.

[0018] Figure 1 This is a block diagram of a video system 100 for video capture according to one embodiment of the present disclosure. Figure 1 In the illustrated embodiment, the video system 100 includes: a first IR camera 110, a second IR camera 114, a color camera 112, a rod 140, at least one IR light source 120, 122, and a processor 130.

[0019] In one embodiment, a first IR camera 110, a second IR camera 114, and a color camera 112 are configured to capture video data of the subject 104. Two IR cameras are used for depth calculation. Furthermore, in one embodiment, the first IR camera 110, the second IR camera 114, and the color camera 112 are all attached to a pole 140, such that the color camera 112 is positioned between the first IR camera 110 and the second IR camera 114. In another embodiment, the color camera 112 is positioned between the first IR camera 110 and the second IR camera 114 without necessarily being attached to the pole 140, but rather by other means such as being attached together or positioned on a table. In yet another embodiment, IR light sources 120, 122 are configured to illuminate the subject 104.

[0020] In one embodiment, processor 130 is connected to a first IR camera 110, a second IR camera 114, and a color camera 112. Processor 130 is configured to perform the following operations: (a) generate depth-solved data 111, 115 of a subject using data from the first IR camera 110 and the second IR camera 114; (b) generate projected color data 113 using data from the color camera 112 to project color onto the depth-solved data 111, 115; and (c) generate final capture data 132 by merging the depth-solved data 111, 115, and the projected color data 113. Processor 130 uses data from all cameras to generate the final capture data. Furthermore, processor 130 uses data from all cameras to generate the final capture data.

[0021] In yet another embodiment, the video system 100 includes at least one additional color camera and at least two additional IR cameras to capture video data of the subject 104. In yet another embodiment, the cameras are arranged in groups of three to form a set of nodes, each node having two IR cameras and one color camera.

[0022] In yet another embodiment, the video system 100 also includes a laser 124 attached to the rod 140 for projecting a laser pattern onto the subject 104. In one embodiment, an 830nm laser is projected from each camera node position to project the laser pattern onto the subject, which aids in solving the contrast and depth capture for both IR cameras.

[0023] In yet another embodiment, each of the IR cameras 110, 114 includes a filter that removes visible light. In one embodiment, the filter is a 700-715nm high-cutoff filter. In one embodiment, the video system 100 also includes an IR floodlight (e.g., an 850nm LED floodlight) positioned around the subject 104 to illuminate the subject 104 using IR. This allows for a brightness mask and constant edges around the subject 104 to separate the background 102 from the captured subject 104, thereby aiding in depth generation. In one embodiment, the background 102 is configured with an IR blocking coating to create a “black hole” behind the subject 104, which is captured to maintain a connection point for depth generation (the connection point is a feature identified in two or more images and selected as a reference point). The video system 100 uses an action-dependent structure and relies on the image connection points between each view to measure depth. In one embodiment, IR pigment is applied to the subject 104 to aid in the contrast and depth capture solving of the two or more IR cameras. Pigments reduce or prevent IR light from scattering into the skin.

[0024] In one example of system operation, a video system is used for volumetric capture of a person. Before setting up the cameras, the operator determines the required coverage area of ​​the subject for the final capture, such as 180 degrees. The cameras are arranged as a set of nodes, such as three poles with three rows of three cameras. The poles are positioned around the subject to achieve maximum coverage. All cameras are shutter-synchronized. The video system then uses the cameras to capture data and images of the subject. The system performs depth solving on the resulting IR data and projects the corresponding color camera data onto the solved depth data. The system then merges the resulting depth and projected color data into the final complete capture.

[0025] Figure 2 This is a flowchart of a video capture method 200 for a subject according to one embodiment of the present disclosure. Figure 2 In the illustrated embodiment, at box 210, depth-solving data for the subject is generated using data from the first IR camera and the second IR camera. Furthermore, at box 220, projected color data is generated using data from the color camera to project color onto the depth-solving data. Then, at box 230, final capture data is generated by merging the depth-solving data and the projected color data, wherein the color camera is positioned between the first IR camera and the second IR camera.

[0026] In one embodiment, the subject is illuminated using at least one IR light source. In another embodiment, at least one other color camera and at least two other IR cameras are arranged in groups of three to form a set of nodes, each node having two IR cameras and one color camera. Data from all cameras is then used to generate final capture data. In yet another embodiment, a laser pattern is projected onto the subject. In one embodiment, the laser pattern is projected from each camera node location. In one embodiment, each of the first and second IR cameras includes a filter that removes visible light. In one embodiment, the filter is a high-cutoff filter in the 700-715 nm range. In yet another embodiment, a background plate configured with an IR-blocking coating is provided. In yet another embodiment, IR pigments are applied to the subject.

[0027] Figure 3A This is a representation of a computer system 300 and a user 302 according to an embodiment of the present disclosure. User 302 uses the computer system 300 to implement a video application 390, which is used to implement... Figure 1 Video system 100 and Figure 2 Method 200 is a technique for video capture of a subject.

[0028] Computer system 300 stores and executes Figure 3B The video application 390. Furthermore, the computer system 300 can communicate with the software program 304. The software program 304 may include software code for the video application 390. The software program 304 can be loaded onto external media such as a CD, DVD, or storage drive, as will be explained further below.

[0029] Furthermore, computer system 300 can connect to network 380. Network 380 can be connected in various different architectures, such as client-server architecture, peer-to-peer network architecture, or other types of architecture. For example, network 380 can communicate with server 385, which coordinates data and engines used within video application 390. Moreover, the network can be of different types. For example, network 380 can be the Internet, a local area network (LAN) or any variant of a LAN, a wide area network (WAN), a metropolitan area network (MAN), an intranet or extranet, or a wireless network.

[0030] Figure 3BThis is a functional block diagram illustrating a computer system 300 hosting a video application 390 according to an embodiment of the present disclosure. The controller 310 is a programmable processor and controls the operation of the computer system 300 and its components. The controller 310 loads instructions (e.g., in the form of a computer program) from memory 320 or an embedded controller memory (not shown) and executes these instructions to control the system. In its execution, the controller 310 provides software systems to the video application 390, for example, to enable video capture of a subject. Alternatively, this service can be implemented as a separate hardware component within the controller 310 or the computer system 300.

[0031] The memory 320 temporarily stores data for use by other components of the computer system 300. In one embodiment, the memory 320 is implemented as RAM. In another embodiment, the memory 320 also includes long-term or permanent memory such as flash memory and / or ROM.

[0032] Storage device 330 stores data temporarily or for extended periods for use by other components of computer system 300. For example, storage device 330 stores data used by video application 390.

[0033] In one implementation, memory 330 is a hard disk drive.

[0034] Media device 340 receives removable media and reads and / or writes data to inserted media. For example, in one embodiment, media device 340 is an optical disc drive.

[0035] User interface 350 includes components for accepting user input from a user of computer system 300 and presenting information to user 302. In one embodiment, user interface 350 includes a keyboard, mouse, audio speakers, and a display. Controller 310 uses input from user 302 to adjust the operation of computer system 300.

[0036] I / O interface 360 ​​includes one or more I / O ports for connecting to corresponding I / O devices such as external storage or supplemental devices (e.g., printers or PDAs). In one embodiment, the ports of I / O interface 360 ​​include, for example, USB ports, PCMCIA ports, serial ports, and / or parallel ports. In another embodiment, I / O interface 360 ​​includes a wireless interface for wireless communication with external devices.

[0037] Network interface 370 includes wired and / or wireless network connections such as an RJ-45 interface supporting Ethernet connectivity or a “Wi-Fi” interface (including but not limited to 802.11).

[0038] Computer system 300 includes the typical additional hardware and software of a computer system (e.g., power supply, cooling, operating system), although for simplicity, these components are... Figure 3B Not specifically shown in the text. In other embodiments, different configurations of the computer system (e.g., different bus or storage configurations or multiprocessor configurations) can be used.

[0039] The description of the disclosed embodiments provided herein is intended to enable any person skilled in the art to make or use this disclosure. Numerous modifications to these embodiments will be readily apparent to those skilled in the art, and the principles defined herein can be applied to other embodiments without departing from the spirit or scope of this disclosure. Therefore, this disclosure is not intended to be limited to the embodiments shown herein, but should be followed within the broadest scope consistent with the main and novel features disclosed herein.

[0040] Therefore, additional variations and implementations are also possible. For example, in addition to video production for film or television, implementations of the system and method can also be applied to and adapted to other applications such as virtual production of film, television, and games (e.g., virtual reality environments), other volumetric capture systems and environments, or replacing green screen operations in other capture systems.

[0041] All features of each of the examples described above are not necessarily required in a particular embodiment of this disclosure. Furthermore, it should be understood that the descriptions and drawings presented herein represent a broad range of subjects considered in this disclosure. It is further understood that the scope of this disclosure fully encompasses other embodiments that may become apparent to those skilled in the art, and therefore the scope of this disclosure is not limited by the content other than the appended claims.

Claims

1. A system for video capture of the motion of a subject, comprising: Background plate, configured with an infrared (IR) blocking coating to create a black hole behind the subject to capture and retain the connection point for depth generation, wherein the connection point is a feature identified in two or more images of the captured video data of the subject and is selected as a reference point to measure depth using the structure based on the action connection point; A first IR camera and a second IR camera are used to capture video data of a subject in the presence of a black hole behind the subject, wherein each of the first IR camera and the second IR camera includes an IR filter to remove visible light, and the subject is coated with IR pigment to reduce IR light scattering into the subject's skin; A color camera, used to capture video data of the subject; A pole, wherein a first IR camera, a second IR camera and a color camera are attached to the pole, and wherein the color camera is positioned between the first IR camera and the second IR camera; At least one IR light source for illuminating a subject, wherein the at least one IR light source includes an IR floodlight positioned around the subject to illuminate the subject with IR light and to allow a brightness mask and constant edge around the subject to separate the background from the subject; as well as A processor, connected to a first IR camera, a second IR camera, and a color camera, is configured to perform the following operations: generate depth-solution data for a subject using data from the first and second IR cameras; generate projected color data using data from the color camera to project color onto the depth-solution data; and generate final capture data by merging the depth-solution data and the projected color data. The IR floodlight positioned around the subject, the background plate with an IR blocking coating, the IR pigment applied to the subject, the IR filter in each of the first and second IR cameras, and the processor are combined to provide video capture of the subject's motion by separating the subject from the background.

2. The system according to claim 1, further comprising: At least two other IR cameras; as well as At least one other color camera.

3. The system according to claim 2, wherein, The cameras are arranged in groups of three to form a set of nodes, with each node having two infrared cameras and one color camera.

4. The system according to claim 2, wherein, The processor uses data from all cameras to generate the final captured data.

5. The system according to claim 1, further comprising: A laser attached to a pole is used to project a laser pattern onto the subject.

6. The system according to claim 5, wherein, The lasers are configured to project from each camera node location.

7. The system according to claim 1, wherein, The IR filter is a high cutoff filter in the 700-715 nm range.

8. A method for video capturing the motion of a subject, comprising: Depth-solving data for the subject is generated using video data of the subject captured by a first IR camera and a second IR camera in the presence of a black hole behind the subject. The black hole behind the subject is created by an IR blocking coating configured on a background plate and is captured to maintain a connection point for depth generation. The connection point is a feature identified in two or more images of the captured video data of the subject and is selected as a reference point to measure depth using the structure based on the action connection point. Each of the first IR camera and the second IR camera includes an IR filter that removes visible light, and the subject is treated with IR pigment to reduce IR light scattering into the subject's skin. The subject is illuminated with IR light by IR floodlights positioned around it, and the IR floodlights allow for brightness masking and constant edges around the subject to separate the background from the subject. Projected color data is generated using data from a color camera to project colors onto the depth solution data; and The final capture data is generated by merging depth solution data and projected color data. The color camera is positioned between the first IR camera and the second IR camera, and the IR floodlight positioned around the subject, the background plate configured with an IR blocking coating, the IR pigment applied to the subject, the IR filter in each of the first IR camera and the second IR camera, the generation of depth solution data, the generation of projected color data, and the generation of final capture data are combined to provide video capture of the subject's motion by separating the subject from the background.

9. The method of claim 8, further comprising: Arrange at least two other IR cameras and at least one other color camera in groups of three to form a set of nodes, each node having two IR cameras and one color camera.

10. The method according to claim 9, wherein, The final captured data is generated using data from all cameras.

11. The method of claim 8, further comprising: A laser pattern is projected onto the subject.

12. The method according to claim 11, wherein, A laser pattern is projected from each camera node position.

13. The method according to claim 8, wherein, The IR filter is a high cutoff filter in the 700-715 nm range.

14. A non-transient computer-readable storage medium storing a computer program for capturing video of the motion of a subject, the computer program including executable instructions that cause the computer to perform the following operations: Depth-solving data for the subject is generated using video data of the subject captured by a first IR camera and a second IR camera in the presence of a black hole behind the subject. The black hole behind the subject is created by an IR blocking coating configured on a background plate and is captured to maintain a connection point for depth generation. The connection point is a feature identified in two or more images of the captured video data of the subject and is selected as a reference point to measure depth using the structure based on the action connection point. Each of the first IR camera and the second IR camera includes an IR filter that removes visible light, and the subject is treated with IR pigment to reduce IR light scattering into the subject's skin. The subject is illuminated with IR light by IR floodlights positioned around it, and the IR floodlights allow for brightness masking and constant edges around the subject to separate the background from the subject. Projected color data is generated using data from a color camera to project colors onto the depth solution data; and The final capture data is generated by merging depth solution data and projected color data. in, A color camera is positioned between a first IR camera and a second IR camera, and an IR floodlight positioned around the subject, a background plate configured with an IR blocking coating, IR pigments applied to the subject, IR filters in each of the first and second IR cameras, generation of depth solution data, generation of projected color data, and generation of final capture data are combined to provide video capture of the subject's motion by separating the subject from the background.

Citation Information

Patent Citations

  • Generating depth map

    CN102982530A

  • Determining a segmentation boundary based on images representing an object

    CN105723300A