Information processing apparatus, camera system, information processing method, and program

The camera system synchronizes depth and image data from multiple cameras to address integration challenges, improving 3D model accuracy and efficiency in reconstruction processes.

JP2025107664APending Publication Date: 2025-07-22NIKON CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024000999
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-01-09
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

Existing depth cameras and imaging systems struggle to accurately synchronize and integrate depth and image data from multiple cameras to create coherent three-dimensional models, leading to inconsistencies and inefficiencies in 3D space reconstruction.

Method used

A camera system and information processing apparatus that utilizes a synchronization unit to align frames of depth and image data from multiple cameras based on shape information comparison, ensuring synchronized reference times for accurate 3D model generation.

Benefits of technology

Enhances the accuracy and efficiency of 3D space reconstruction by aligning and integrating depth and image data from multiple cameras, resulting in improved three-dimensional models with enhanced detail and consistency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025107664000001_ABST
    Figure 2025107664000001_ABST
Patent Text Reader

Abstract

To process information acquired by a distance detection apparatus.SOLUTION: An information processing apparatus is an apparatus configured to process information acquired by a distance detection apparatus. The information processing apparatus includes a shape information generation unit which generates shape information representing a subject based on first distance information acquired by a first distance detection apparatus and second distance information acquired by a second distance detection apparatus. The information processing apparatus includes a synchronization unit which synchronizes the first distance information with the second distance information. The shape information generation unit generates first shape information based on a specific frame of the first distance information and a first frame of the second distance information, and second shape information based on the specific frame and a second frame of the second distance information. The synchronization unit synchronizes one of the first and second frames with the specific frame at the same reference time, based on a result of comparing the first shape information with the second shape information.SELECTED DRAWING: Figure 7
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an information processing apparatus, a camera system, an information processing method, and a program.

Background Art

[0002] Unlike a general camera, a depth camera is a camera that can acquire distance information of a subject (see, for example, Patent Document 1).

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

[0004] An information processing apparatus according to an aspect of the present invention is an apparatus that processes information acquired by a distance detection apparatus. The information processing apparatus includes a shape information generation unit that generates shape information representing a subject based on first distance information acquired by a first distance detection apparatus and second distance information acquired by a second distance detection apparatus. The information processing apparatus includes a synchronization unit that synchronizes the first distance information and the second distance information. The shape information generation unit generates first shape information based on a specific frame of the first distance information and a first frame of the second distance information, and second shape information based on the specific frame and a second frame of the second distance information. The synchronization unit aligns either the first frame or the second frame with the same reference time as the specific frame based on a result of comparing the first shape information and the second shape information.

[0005] A camera system according to an aspect of the present invention is a system that acquires information by a distance detection device. The camera system includes a first distance detection device. The camera system includes a second distance detection device. The camera system includes an information processing device that processes distance information of a subject acquired by the distance detection device. The information processing device includes a shape information generation unit that generates shape information representing the subject based on first distance information acquired by the first distance detection device and second distance information acquired by the second distance detection device. The camera system includes a synchronization unit that synchronizes the first distance information and the second distance information. The shape information generation unit generates first shape information based on a specific frame of the first distance information and a first frame of the second distance information, and second shape information based on the specific frame and a second frame of the second distance information. The synchronization unit aligns either the first frame or the second frame with the same reference time as the specific frame based on the result of comparing the first shape information and the second shape information.

[0006] An information processing method according to an aspect of the present invention is a method of processing information acquired by a distance detection device. The information processing method includes generating shape information representing a subject based on first distance information acquired by a first distance detection device and second distance information acquired by a second distance detection device. The information processing method includes synchronizing the first distance information and the second distance information. The information processing method includes generating first shape information based on a specific frame of the first distance information and a first frame of the second distance information, and second shape information based on the specific frame and a second frame of the second distance information. The information processing method includes aligning either the first frame or the second frame with the same reference time as the specific frame based on the result of comparing the first shape information and the second shape information.

[0007] A program according to an aspect of the present invention is a program that causes a computer to function as an information processing device that processes information acquired by a distance detection device. The program causes the computer to function as a shape information generation unit that generates shape information representing a subject based on first distance information acquired by a first distance detection device and second distance information acquired by a second distance detection device. The program causes the computer to function as a synchronization unit that synchronizes the first distance information and the second distance information. The shape information generation unit generates first shape information based on a specific frame of the first distance information and a first frame of the second distance information, and second shape information based on the specific frame and a second frame of the second distance information. The synchronization unit aligns either the first frame or the second frame with the same reference time as the specific frame based on the result of comparing the first shape information and the second shape information.

Brief Description of Drawings

[0008]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Embodiments for Carrying Out the Invention

[0009] The present invention will be described through embodiments of the invention. However, the embodiments do not limit the invention according to the claims. Also, not all combinations of features described in the embodiments are essential for the solution means of the invention. The camera system of the present embodiment uses a plurality of camera units to periodically detect the depth to the subject and the image of the subject. A series of depth data and image data detected by the plurality of camera units are synthesized to generate a series of three-dimensional model data, that is, three-dimensional video data.

[0010] FIG. 1 is a diagram showing an example of the configuration of the camera system 100. The camera system 100 is a system that acquires information by a camera. The camera system 100 includes a first camera unit 110A, a second camera unit 110B, a third camera unit 110C, a fourth camera unit 110D, and an information processing device 120. When the first camera unit 110A, the second camera unit 110B, the third camera unit 110C, and the fourth camera unit 110D are not distinguished, they are collectively referred to as the camera unit 110.

[0011] FIG. 2 is a diagram showing an example of the positional relationship between the camera unit 110 and the subject SU. The camera unit 110 is arranged so as to surround the periphery of the subject SU. For example, the camera unit 110 is arranged at a plurality of locations among the first position P1, the second position P2, the third position P3, the fourth position P4, the fifth position P5, the sixth position P6, the seventh position P7, and the eighth position P8. The first position P1 is a position where the camera unit 110 can detect information of the subject SU when the subject SU is viewed from the 12 o'clock direction with respect to the subject SU. The second position P2 is a position where the camera unit 110 can detect information of the subject SU when the subject SU is viewed from the 1 o'clock direction with respect to the subject SU. The third position P3 is a position where the camera unit 110 can detect information of the subject SU when the subject SU is viewed from the 3 o'clock direction with respect to the subject SU. The fourth position P4 is a position where the camera unit 110 can detect information of the subject SU when the subject SU is viewed from the 5 o'clock direction with respect to the subject SU. The fifth position P5 is a position where the camera unit 110 can detect information of the subject SU when the subject SU is viewed from the 6 o'clock direction with respect to the subject SU. The sixth position P6 is a position where the camera unit 110 can detect information of the subject SU when the subject SU is viewed from the 7 o'clock direction with respect to the subject SU. The seventh position P7 is a position where the camera unit 110 can detect information of the subject SU when the subject SU is viewed from the 9 o'clock direction with respect to the subject SU. The eighth position P8 is a position where the camera unit 110 can detect information of the subject SU when the subject SU is viewed from the 11 o'clock direction with respect to the subject SU.

[0012] The four camera units 110 are arranged in a well-balanced manner so as to surround the subject SU. For example, the four camera units 110 are arranged one by one at the first position P1, the third position P3, the fifth position P5, and the seventh position P7. For example, the four camera units 110 are arranged one by one at the second position P2, the fourth position P4, the sixth position P6, and the eighth position P8.

[0013] For example, the camera units 110 are arranged such that their optical axes are horizontal. Some or all of the four camera units 110 may be arranged so as to look up at the subject SU. Some or all of the four camera units 110 may be arranged so as to look down at the subject SU.

[0014] FIG. 3 is a diagram showing an example of a procedure for calibrating the positions of the camera units 110 with respect to each other. The calibration of the positions of the camera units 110 with respect to each other is a process for accurately aligning the positions and postures of the plurality of camera units 110 when they photograph the same three-dimensional space. By performing this calibration, the information processing apparatus 120 can integrate the images obtained from the plurality of camera units 110 or accurately acquire the information of the 3D space.

[0015] First, the user arranges markers for calibration with respect to the plurality of camera units 110 (S101). These markers are precisely arranged in the 3D space.

[0016] The camera units 110 detect the arranged markers and acquire their position information as 2D images (S102). The information processing apparatus 120 compares the 2D images of the markers obtained from the plurality of camera units 110 and specifies the positions of the markers.

[0017] The information processing device 120 estimates the position and orientation of the camera unit 110 in the three-dimensional space based on the 2D position information of the marker (S103). In S103, the information processing device 120 combines the 2D positions of the markers obtained from the plurality of camera units 110 to calculate the relative information on the position and orientation of the cameras in the 3D space.

[0018] The information processing device 120 combines the information of the plurality of camera units 110 to adjust the positions and orientations of the camera units 110 relative to each other (S104). This process is called bundle adjustment, which adjusts the positions and orientations of the camera units 110 to improve the consistency with the 2D position information of the markers.

[0019] The user checks the final calibration result and verifies that the positions and orientations between the camera units 110 are accurately aligned (S105).

[0020] With accurate calibration, the information processing device 120 can perform advanced video processing and 3D space reconstruction using the plurality of camera units 110.

[0021] Returning to the description of FIG. 1, the first camera unit 110A includes a first depth camera 111A and a first mirrorless single-lens camera 112A. The second camera unit 110B includes a second depth camera 111B and a second mirrorless single-lens camera 112B. The third camera unit 110C includes a third depth camera 111C and a third mirrorless single-lens camera 112C. The fourth camera unit 110D includes a fourth depth camera 111D and a fourth mirrorless single-lens camera 112D. When the first depth camera 111A, the second depth camera 111B, the third depth camera 111C, and the fourth depth camera 111D are not distinguished, they are collectively referred to as the depth camera 111. When the first mirrorless single-lens camera 112A, the second mirrorless single-lens camera 112B, the third mirrorless single-lens camera 112C, and the fourth mirrorless single-lens camera 112D are not distinguished, they are collectively referred to as the mirrorless single-lens camera 112. The depth camera 111 is an example of a distance detection device and a first detection device. The first depth camera 111A is an example of a first distance detection device. The second depth camera 111B, the third depth camera 111C, and the fourth depth camera 111D are examples of a second distance detection device. The mirrorless single-lens camera 112 is an example of a second detection device.

[0022] Unlike a general camera, the depth camera 111 is a camera that can acquire distance information of the subject SU. A conventional camera only acquires a two-dimensional image. However, the depth camera 111 can measure the distance to the subject SU for each pixel. Generally, the depth camera 111 uses a sensor to acquire distance information. Thereby, the depth camera 111 can capture information such as the distance between the subject SU and the depth camera 111, the shape of the subject SU, and the depth in real time.

[0023] FIG. 4 is a diagram showing an example of the camera unit 110. The depth camera 111 is attached to the mirrorless single-lens camera 112. For example, the depth camera 111 is attached to the accessory shoe of the mirrorless single-lens camera 112. The accessory shoe is a small metal attachment portion provided on the upper part of the mirrorless single-lens camera 112 body and is an interface for attaching various accessories to the mirrorless single-lens camera 112. The calibration of the positions of the depth camera 111 and the mirrorless single-lens camera 112 is performed in the same procedure as the procedure shown in FIG. 3.

[0024] Returning to the description of FIG. 1, the depth camera 111 includes a depth sensor S1 and a color sensor S2.

[0025] The depth sensor S1 is a sensor for the depth camera 111 to measure the distance information to the subject SU. The depth sensor S1 is usually used in combination with an infrared light source and a light receiving element. For example, the depth sensor S1 is a structured light sensor, a time-of-flight sensor, etc. When the depth sensor S1 is a structured light sensor, the depth sensor S1 irradiates the subject SU with an infrared light source. The pattern projected from the light source is projected onto the surface of the subject SU. The irradiated structured light is reflected from the surface of the subject SU and returns to the depth sensor S1. The depth sensor S1 receives this reflected light. The depth camera 111 analyzes the received structured light pattern and detects the deformation and distortion of the light on the surface of the subject SU. When the surface of the subject SU is a curved surface, the pattern is distorted, resulting in a change corresponding to the distance to the subject SU. The depth camera 111 measures the distance to the subject SU for each pixel based on the degree of change in the pattern. The depth camera 111 performs distance estimation using calibration information associated with the amount of light deformation and the known distance to the subject SU. The depth camera 111 summarizes the obtained distance information for each pixel and generates depth data D1. The depth data D1 is three-dimensional data indicating the distance to the subject SU for the coordinates of each pixel. The depth data D1 is an example of distance information, first distance information, and second distance information.

[0026] The color sensor S2 is an imaging sensor for capturing RGB (Red - Green - Blue color model) information, similar to a general color camera. The depth camera 111 is equipped with a dedicated depth sensor S1 to provide depth information, and at the same time, it is also equipped with a color sensor S2. The color sensor S2 serves to capture the color information of the scene photographed by the depth camera 111. Generally, the color sensor S2 uses CMOS (Complementary Metal Oxide Semiconductor) or CCD (Charge - Coupled Device) technology as an imaging sensor. The color sensor S2 is controlled by the image processor of the depth camera 111 and captures RGB information as digital data. Each pixel has red, green, and blue values. A color image is generated by combining these values. The color data captured by the color sensor S2 is processed in synchronization with the depth data D1.

[0027] The mirrorless single - lens camera 112 is a camera for acquiring color images. The mirrorless single - lens camera 112 is equipped with an imaging sensor for capturing the upper part of RGB, similar to the color sensor S2 of the depth camera 111. The imaging sensor of the mirrorless single - lens camera 112 has a larger sensor size, more pixels, and larger pixel sizes than the imaging sensor of the depth camera 111. The mirrorless single - lens camera 112 has interchangeable lenses. By replacing the lenses with different focal lengths and characteristics, the user can take photos suitable for various shooting conditions. The mirrorless single - lens camera 112 is equipped with a display for displaying the images acquired by the imaging sensor, and the user can view the captured images in real - time.

[0028] Returning to the description of FIG. 1, the information processing device 120 is a computer that processes the information acquired by the camera unit 110. The information processing device 120 is communicatively connected to the camera unit 110.

[0029] FIG. 5 is a diagram showing an example of the configuration of the information processing apparatus 120. The information processing apparatus 120 includes a CPU 121, a main memory 122, an input / output interface 123, a communication device 124, an input device 125, a display 126, and a storage 127.

[0030] The CPU 121 is a device that controls the main memory 122, the input / output interface 123, the communication device 124, the input device 125, the display 126, and the storage 127, and performs operations on data.

[0031] The main memory 122 is a storage device that stores data and programs inside the information processing apparatus 120, and is connected to the CPU 121 through electrical wiring on the board and the like. The main memory 122 stores, for example, the program code being executed related to the camera system 100, data necessary for current processing, and the like.

[0032] The input / output interface 123 is an interface that connects the communication device 124, the input device 125, the display 126, and the storage 127 to the CPU 121 via a cable or the like, and transmits and receives data and control information.

[0033] The communication device 124 is a device for connecting the information processing apparatus 120 to the camera unit 110. The communication device 124 communicates with the camera unit 110. The communication device 124 is an example of a receiving unit and a transmitting unit.

[0034] The input device 125 is a device for providing data, information, instructions, etc. to the information processing apparatus 120. For example, the input device 125 is used to provide data, information, instructions, etc. to the information processing apparatus 120 when processing the information acquired by the camera unit 110. The input device 125 is an example of an input unit.

[0035] The display 126 is one of the output devices of the information processing apparatus 120 and is a display device that emits light on the screen to project an image. For example, the display 126 is used to display information acquired by the camera unit 110, information processed by the CPU 121, information processed by the CPU 121, and the like. The display 126 is an example of a display unit.

[0036] The storage 127 is a device that permanently stores data. The storage 127 stores depth data D1, low-quality video data D2, high-quality video data D3, mesh data D4, and 3D model data D5.

[0037] The depth data D1 is data representing the distance information of the subject SU in the 3D space. The depth data D1 refers to information indicating the distance to the subject SU captured by the depth sensor S1 of the depth camera 111 for each pixel. The depth data D1 provides information on the depth and arrangement of the subject SU, which cannot be obtained only from the information of the 2D image. The depth data D1 consists of a plurality of frames acquired at different times. The depth data D1 is an example of the information acquired by the distance detection device. The depth data D1 consisting of a plurality of frames is an example of a series of distance information.

[0038] In the following description, the depth data D1 acquired by the depth sensor S1 of the first depth camera 111A is referred to as first depth data D1A. The depth data D1 acquired by the depth sensor S1 of the second depth camera 111B is referred to as second depth data D1B. The depth data D1 acquired by the depth sensor S1 of the third depth camera 111C is referred to as third depth data D1C. The depth data D1 acquired by the depth sensor S1 of the fourth depth camera 111D is referred to as fourth depth data D1D. The first depth data D1A is an example of first distance information. The first depth data D1A consisting of a plurality of frames is an example of a series of first distance information. The second depth data D1B, the third depth data D1C, and the fourth depth data D1D are examples of second distance information. The second depth data D1B, the third depth data D1C, and the fourth depth data D1D consisting of a plurality of frames are examples of a series of second distance information.

[0039] The low-quality video data D2 is video data acquired by the color sensor S2 of the depth camera 111. The low-quality video data D2 is an example of a series of first image information.

[0040] In the following description, the low-quality video data D2 acquired by the color sensor S2 of the first depth camera 111A is referred to as first low-quality video data D2A. The low-quality video data D2 acquired by the color sensor S2 of the second depth camera 111B is referred to as second low-quality video data D2B. The low-quality video data D2 acquired by the color sensor S2 of the third depth camera 111C is referred to as third low-quality video data D2C. The low-quality video data D2 acquired by the color sensor S2 of the fourth depth camera 111D is referred to as fourth low-quality video data D2D.

[0041] The high-quality video data D3 is video data acquired by the mirrorless single-lens camera 112. The image quality of the high-quality video data D3 is higher than the image quality of the low-quality video data D2. The high-quality video data D3 is an example of a series of second image information.

[0042] In the following description, the high-quality video data D3 acquired by the first mirrorless single-lens camera 112A is referred to as the first high-quality video data D3A. The high-quality video data D3 acquired by the second mirrorless single-lens camera 112B is referred to as the second high-quality video data D3B. The high-quality video data D3 acquired by the third mirrorless single-lens camera 112C is referred to as the third high-quality video data D3C. The high-quality video data D3 acquired by the fourth mirrorless single-lens camera 112D is referred to as the fourth high-quality video data D3D.

[0043] Mesh data D4 is a data format for approximately representing the surface of the subject SU. The mesh data D4 stores information on the shape and surface of the subject SU in 3D space and is used for visualization, editing, and processing. A mesh is composed of vertices, edges, polygons, etc., and the shape of the subject SU is defined by combining these elements. A vertex represents a point in 3D space. Each vertex has three-dimensional coordinates and defines the shape and surface characteristics of the subject SU. An edge represents a line segment connecting two vertices. The edge connects the vertices and defines the shape of the side of the subject SU. A polygon is a planar figure created by combining vertices and edges. Triangles and quadrilaterals are common for polygons. A polygon constitutes the surface of an object and can have information such as color and texture assigned to it. The mesh data D4 consists of a plurality of frames. The mesh data D4 is an example of shape information. The mesh data D4 consisting of a plurality of frames is an example of a series of shape information.

[0044] The 3D model data D5 is data for expressing the three-dimensional shape and information of the subject SU. The 3D model data D5 is widely used in fields such as computer graphics, animation, game development, virtual reality, and simulation. The 3D model is used to mathematically approximate the actual subject SU and to endow it with visual properties and physical characteristics. For example, the 3D model data D5 is composed of elements such as mesh data D4, texture information, material information, animation information, and hierarchical structure. The texture information is information that defines the texture applied to the surface of the 3D model represented by the mesh data D4. The appearance and texture of the subject SU are determined by the 3D model data D5 according to the texture information. The material information is information that defines the material of the surface of the subject SU and the light reflection characteristics, etc. The material information includes attributes such as color, reflectivity, transparency, and gloss. The animation information is information for expressing the movement of the subject SU. The animation information includes data such as skeletal animation and morphing. The hierarchical structure is information indicating the combination and hierarchical relationship of partial objects used when expressing a complex subject SU. The 3D model can make the parts operate correctly in conjunction with each other and make the animation operate more naturally due to the hierarchical structure. The 3D model data D5 is saved in various formats and shared among different software and systems. Representative 3D model data D5 includes OBJ, FBX, STL, Collada, GLTF, 3DS, etc. Each format may be suitable for specific applications or platforms. The 3D model data D5 is an important element utilized in a wide variety of applications such as visual representation and physical simulation. The 3D model data D5 consists of multiple frames. The 3D model data D5 consisting of multiple frames is an example of a series of shape information including color information.

[0045] The CPU 121 functions as a determination unit 121A, an alert output unit 121B, a shape information generation unit 121C, a shape information analysis unit 121D, a synchronization unit 121E, an image analysis unit 121F, and a decision unit 121G.

[0046] The determination unit 121A is a software module that determines whether the number of frames is within a specific range. For example, the determination unit 121A determines whether the number of frames of the depth data D1 is within a specific range. For example, the determination unit 121A determines whether the number of frames of the low-quality video data D2 is within a specific range. For example, the determination unit 121A determines whether the number of frames of the high-quality video data D3 is within a specific range. The number of frames of the depth data D1 refers to the number of pieces of depth data D1 acquired continuously. The number of frames of the depth data D1 represents the number of times the depth sensor S1 captures distance information in the 3D space at regular intervals. The number of frames of the depth data D1 varies depending on the application and purpose of use. When using the depth data D1 to generate a 3D model of the subject SU, an appropriate resolution and data density are required. Therefore, the depth camera 111 captures detailed shapes using dozens to hundreds of frames. The number of frames of the depth data D1 varies depending on the performance of the hardware, requirements of the application, processing power, etc. The user can achieve the required accuracy and real-time performance by selecting an appropriate number of frames.

[0047] The alert output unit 121B is a software module that outputs an alert. For example, the alert output unit 121B outputs an alert when the number of frames of the depth data D1 is not within the specific range. For example, the alert output unit 121B outputs an alert when the number of frames of the low-quality video data D2 is not within the specific range. For example, the alert output unit 121B outputs an alert when the number of frames of the high-quality video data D3 is not within the specific range. An alert means a message or the like that is displayed or notified to prompt the user to pay attention or give a warning.

[0048] The shape information generation unit 121C is a software module that generates mesh data D4 and 3D model data D5. For example, the shape information generation unit 121C generates mesh data D4 representing the subject SU based on the depth data D1 acquired by the depth camera 111. For example, the shape information generation unit 121C generates a plurality of mesh data D4 based on each frame of the specific frame of the first depth data D1A and the second depth data D1B. For example, the shape information generation unit 121C generates a plurality of mesh data D4 based on each frame of the specific frame of the first depth data D1A and the third depth data D1C. For example, the shape information generation unit 121C generates a plurality of mesh data D4 based on each frame of the specific frame of the first depth data D1A and the fourth depth data D1D. For example, the shape information generation unit 121C starts generating the depth data D1 when the number of frames of the depth data D1 is within a specific range. For example, the shape information generation unit 121C starts generating the depth data D1 when the number of frames of the low-quality video data D2 is within a specific range. For example, the shape information generation unit 121C starts generating the depth data D1 when the number of frames of the high-quality video data D3 is within a specific range. For example, the shape information generation unit 121C generates the mesh data D4 based on the depth data D1 synchronized by the synchronization unit 121E. For example, the shape information generation unit 121C synthesizes the depth data D1 and the high-quality video data D3 based on the relative relationship determined by the determination unit 121G to generate the 3D model data D5.

[0049] The shape information analysis unit 121D is a software module that analyzes the mesh data D4 generated by the shape information generation unit 121C. For example, the shape information analysis unit 121D calculates the number of polygons in the mesh data D4 generated by the shape information generation unit 121C. The number of polygons in the mesh data D4 refers to the total number of polygons in the 3D model. The number of polygons is an important indicator indicating the level of detail and complexity of the 3D model. The larger the number of polygons, the more detailed the appearance of the 3D model tends to be, and fine shapes such as curves and unevenness are also more accurately represented. On the other hand, when the number of polygons is small, the 3D model has a more simplified shape and requires fewer resources, but detailed expressions and curves may be restricted. The number of polygons has a great influence on the appearance and performance of the model.

[0050] The synchronization unit 121E is a software module that synchronizes the depth data D1.

[0051] For example, the synchronization unit 121E synchronizes the first depth data D1A and the second depth data D1B. For example, the synchronization unit 121E aligns a specific frame of the first depth data D1A with any one of a plurality of frames of the second depth data D1B at the same reference time. For example, the synchronization unit 121E determines a frame to be aligned with the specific frame at the same reference time based on the result of comparing a plurality of mesh data D4 generated for each frame of the specific frame and the second depth data D1B. For example, the synchronization unit 121E determines a frame to be aligned with the specific frame at the same reference time based on the analysis result by the shape information analysis unit 121D. For example, the synchronization unit 121E aligns the frame of the second depth data D1B corresponding to the mesh data D4 with a small number of polygons with the specific frame at the same reference time. For example, the synchronization unit 121E synchronizes the first depth data D1A and the second depth data D1B.

[0052] For example, the synchronization unit 121E synchronizes the first depth data D1A and the third depth data D1C. For example, the synchronization unit 121E aligns a specific frame of the first depth data D1A with any one of a plurality of frames of the third depth data D1C at the same reference time. For example, the synchronization unit 121E determines a frame to be aligned with the specific frame at the same reference time based on the result of comparing a plurality of mesh data D4 generated for each frame of the specific frame and the third depth data D1C. For example, the synchronization unit 121E determines a frame to be aligned with the specific frame at the same reference time based on the analysis result by the shape information analysis unit 121D. For example, the synchronization unit 121E aligns the frame of the third depth data D1C corresponding to the mesh data D4 with a small number of polygons with the specific frame at the same reference time. For example, the synchronization unit 121E synchronizes the first depth data D1A and the third depth data D1C.

[0053] For example, the synchronization unit 121E synchronizes the first depth data D1A and the fourth depth data D1D. For example, the synchronization unit 121E aligns a specific frame of the first depth data D1A with any one of a plurality of frames of the fourth depth data D1D at the same reference time. For example, the synchronization unit 121E determines a frame to be aligned with the specific frame at the same reference time based on the result of comparing a plurality of mesh data D4 generated for each frame of the specific frame and the fourth depth data D1D. For example, the synchronization unit 121E determines a frame to be aligned with the specific frame at the same reference time based on the analysis result by the shape information analysis unit 121D. For example, the synchronization unit 121E aligns the frame of the fourth depth data D1D corresponding to the mesh data D4 with a small number of polygons with the specific frame at the same reference time. For example, the synchronization unit 121E synchronizes the first depth data D1A and the fourth depth data D1D.

[0054] The image analysis unit 121F is a software module that analyzes the high-quality video data D3 acquired by the image sensor of the mirrorless single-lens camera 112. For example, the image analysis unit 121F determines the degree of match between a frame of the high-quality video data D3 and a specific frame of the low-quality video data D2.

[0055] For example, the image analysis unit 121F determines the degree of match by comparing the image features of a frame of the high-quality video data D3 with the image features of a specific frame of the low-quality video data D2. For example, the image analysis unit 121F generates a feature vector of a frame that becomes an image feature using a feature extraction algorithm and determines the degree of match.

[0056] For example, the image analysis unit 121F uses SIFT (Scale-Invariant Feature Transform) as a feature extraction algorithm. SIFT is a computer vision technique for detecting characteristic points from a video and using those points to identify the degree of match. SIFT is an algorithm for extracting features that can be consistently detected at different scales and angles, and is widely used for image match detection and object tracking. When using SIFT, first, the image analysis unit 121F detects the maximum and minimum points in the scale space. By convolving the image at different scales, it is possible to obtain the representations of various features at different scales. SIFT detects characteristic points at different scales. As a result, the image analysis unit 121F can detect features regardless of the size of the subject SU. Next, the image analysis unit 121F refines the local maximum values of the keypoints. After the image analysis unit 121F detects the maximum and minimum points, it evaluates whether these points are actually reliable keypoints. The image analysis unit 121F uses the surrounding luminance information to refine the positions where edges and features appear clearly. Next, the image analysis unit 121F assigns an orientation to the keypoints. The image analysis unit 121F uses the gradient information in the surrounding area of the keypoint to assign a main orientation to each keypoint. As a result, the keypoint becomes invariant to rotation, and the robustness against the rotation of the subject SU is improved. Next, the image analysis unit 121F calculates the feature descriptor. The image analysis unit 121F calculates the feature descriptor of the surrounding area based on the luminance information around each keypoint. As a result, the image analysis unit 121F represents the keypoint as a unique feature vector and uses it as a means to calculate the degree of match between different keypoints. Next, the image analysis unit 121F detects the degree of match of the features. The image analysis unit 121F compares the feature descriptors extracted from different videos, evaluates the similarity, and detects the degree of match. At this stage, the image analysis unit 121F uses the distance or similarity measure between the feature vectors to find pairs of matching features. Since SIFT consistently extracts features that are robust to scale and rotation, it is a useful technique for finding the degree of match between a frame of high-quality video data D3 and a specific frame of low-quality video data D2.The image analysis unit 121F can detect the match between videos through this feature matching.

[0057] For example, the image analysis unit 121F uses SURF (Speeded-Up Robust Features) as the feature extraction algorithm. SURF is a feature extraction method similar to SIFT, but some algorithms are optimized to improve the calculation speed. SURF is often used in scenarios where real-time performance is required. It is a method for extracting features from videos, comparing these features, and identifying the degree of match. SURF improves the calculation speed by adopting acceleration using a hash table and a simplified method in feature calculation. When detecting characteristic points in an image, SURF quickly explores areas where features are likely to exist. As a result, the image analysis unit 121F can quickly detect feature points. SURF creates a feature descriptor by extracting a plurality of local information from the surrounding area of a feature point and combining this information as a vector. This descriptor represents the nature of the feature point and is used for comparison. SURF extracts relatively robust features against rotation, scale changes, and optical distortion. As a result, the image analysis unit 121F can easily and correctly detect the degree of match even in scenes where objects change in the video. SURF compares the feature descriptors extracted from different videos and evaluates the degree of match. The image analysis unit 121F identifies the degree of match between a frame of the high-quality video data D3 and a specific frame of the low-quality video data D2 by matching feature pairs using distance and similarity metrics. Generally speaking, SURF is a feature extraction method that emphasizes high speed while following the idea of SIFT. Therefore, SURF is suitable for applications such as detecting the degree of match between videos where real-time performance is required.

[0058] For example, the image analysis unit 121F uses ORB (Oriented FAST and Rotated BRIEF) as the feature extraction algorithm. ORB is an algorithm that combines fast and robust methods in feature extraction and feature description. ORB is used to extract features from video data and compare those features to identify the degree of match. The first step of ORB is to detect keypoints using FAST (Features from Accelerated Segment Test). FAST is a method for quickly detecting corners and characteristic points in an image. By adopting FAST, ORB realizes fast feature detection. ORB improves the robustness against rotation by assigning an orientation to the keypoints. The orientation assignment is performed using the gradient information in the surrounding area of the keypoints. As a result, the features become invariant to rotation. ORB calculates BRIEF (Binary Robust Independent Elementary Features) descriptors from the surrounding area of the feature points. BRIEF is a method for performing pixel comparison in the surrounding area of the feature points using binary codes and is a fast descriptor. ORB uses Rotated BRIEF, an improved version of BRIEF, to make it robust against rotation. ORB detects the degree of match between a frame of high-quality video data D3 and a specific frame of low-quality video data D2 by comparing the BRIEF descriptors of the feature points. The image analysis unit 121F matches pairs of features using, for example, the Hamming distance between the descriptors. ORB also has a certain degree of robustness with respect to the scale of the feature points. As a result, the image analysis unit 121F can obtain stable results even in detecting the degree of match at different scales. Generally speaking, ORB is a fast and robust feature extraction method and is useful when extracting features from a video and detecting the degree of match. ORB is particularly suitable for applications that require real-time performance and video analysis in constrained environments.

[0059] For example, the image analysis unit 121F determines the degree of coincidence by comparing the movement of the subject SU between consecutive frames of the high-quality video data D3 and the movement of the subject SU between consecutive frames including a specific frame of the low-quality video data D2. For example, the image analysis unit 121F analyzes the change in brightness of each pixel between frames and estimates the movement vector of the pixel corresponding to the subject SU. For example, the image analysis unit 121F determines the degree of coincidence by comparing the movement vector between consecutive frames of the high-quality video data D3 and the movement vector between consecutive frames including a specific frame of the low-quality video data D2. For example, the image analysis unit 121F detects the pixels corresponding to the subject SU using a feature point detection algorithm.

[0060] For example, the image analysis unit 121F uses Harris corner detection as a feature point detection algorithm to detect pixels corresponding to the subject SU. Harris corner detection is a method for detecting corners, which are characteristic points in an image, in computer vision and image processing. Harris corner detection is a method for identifying points where corners and edges intersect within an image and shows a stable response to changes in position and angle within the image. The idea of Harris corner detection is to analyze the change in pixel values within a local region of the image. Near corners, minute changes occur in different directions, so Harris corner detection can find corners by detecting this. First, Harris corner detection selects a small surrounding region for each pixel in the image. Harris corner detection evaluates the change in pixel values within this window. Next, Harris corner detection calculates a gradient representing the pixel value distribution within the window. Usually, Harris corner detection uses an edge detection filter such as Sobel to obtain gradients in the X-axis and Y-axis directions. Next, Harris corner detection calculates a gradient matrix representing the change in gradients within the window. The gradient matrix is a matrix that takes into account the product of the gradient components in the X-direction and the product of the gradient components in the Y-axis direction for each pixel within the window. Next, Harris corner detection calculates a corner response function based on the gradient matrix. The corner response function is an index representing the magnitude and direction of pixel value changes within a local region. Next, Harris corner detection compares the value of the corner response function with a threshold and detects pixels exceeding the threshold as corners.

[0061] For example, the image analysis unit 121F uses FAST as a feature point detection algorithm to detect pixels corresponding to the subject SU. FAST is a fast and effective feature point detection method in computer vision and image processing. FAST is used to detect corners, which are characteristic points in an image, and is suitable for applications in real-time processing and environments with resource constraints. The features of FAST are its high speed and robustness. The basic idea is to compare adjacent pixels existing around a certain pixel to determine whether it is a corner. First, FAST selects a central pixel. Next, FAST selects adjacent pixels existing around the central pixel. These adjacent pixels are arranged in a circular ring shape. Next, FAST subtracts a certain threshold value from the luminance value of the central pixel and calculates its absolute value. Next, FAST compares the luminance value of the adjacent pixel with the luminance value of the central pixel, and determines that there is a change from "dark to bright" or "bright to dark" when the luminance value of the central pixel is equal to or greater than the threshold value. If consecutive pixels among a certain number of adjacent pixels have changes exceeding the threshold value, the central pixel is detected as a feature point. The advantage of FAST is that it has a low computational cost and is suitable for real-time applications.

[0062] For example, the image analysis unit 121F uses SIFT as a feature point detection algorithm to detect pixels corresponding to the subject SU. For example, the image analysis unit 121F uses SURF as a feature point detection algorithm to detect pixels corresponding to the subject SU. For example, the image analysis unit 121F uses ORB as a feature point detection algorithm to detect pixels corresponding to the subject SU.

[0063] For example, the image analysis unit 121F uses an optical flow estimation algorithm to estimate a motion vector. The optical flow estimation algorithm is a method for analyzing the movement of an object between consecutive frames in a video. The image analysis unit 121F can estimate in which direction and by how much each individual pixel in the video has moved by using the optical flow estimation algorithm.

[0064] For example, the image analysis unit 121F estimates the motion vectors using the Lucas-Kanade method as an optical flow estimation algorithm. The Lucas-Kanade method is an optical flow estimation algorithm for estimating the motion vectors of pixels between consecutive frames in a video. The Lucas-Kanade method analyzes the luminance changes in a small area around a specific pixel and estimates the motion vectors based on that change. First, the Lucas-Kanade method selects a window around the pixel to be estimated. It is important that the size of the window is appropriate when the pixel movement is small. Next, the Lucas-Kanade method calculates the gradient of the luminance value of each pixel within the selected window. The luminance gradient is information indicating how much the luminance of the pixel has changed with respect to time. The luminance gradient indicates in which direction and by how much the brightness of the pixel has changed. Next, the Lucas-Kanade method estimates the motion vectors based on the luminance gradients of the pixels within the window and the positions of each pixel within that window. The motion vectors are obtained by analyzing the luminance changes of each pixel within the window using a method such as the least squares method. This motion vector represents the amount and direction of pixel movement. The image analysis unit 121F can visualize the movement pattern of the subject SU within the video based on the estimated motion vectors. Thereby, the image analysis unit 121F can analyze the movement, speed, direction, etc. of the subject SU. The Lucas-Kanade method is particularly effective when the movement of the subject SU is small or when the subject SU is moving against a stationary background.

[0065] For example, the image analysis unit 121F estimates the motion vectors using the Horn-Schunck method as an optical flow estimation algorithm. The Horn-Schunck method is an algorithm for optical flow estimation and is used to estimate the motion vectors of pixels in a video. The Horn-Schunck method assumes the consistency of luminance between consecutive frames and focuses on estimating the overall motion field. The Horn-Schunck method is based on the assumption that the luminance of the subject SU does not change over time. That is, it is assumed that the change in luminance at a certain pixel can be explained by the movement of that pixel. First, the Horn-Schunck method constructs a system of simultaneous equations that represents the consistency of luminance between two consecutive frames. Next, the Horn-Schunck method solves the system of simultaneous equations using a method such as the least squares method to obtain the motion vectors of each pixel. Since the Horn-Schunck method estimates the global motion field across all pixels, it can integrate the information for each pixel to obtain a smooth motion pattern. The image analysis unit 121F smooths the result in order to smooth the obtained motion vectors. As a result, noise and local discontinuities are reduced. The Horn-Schunck method is suitable when the movement of the subject SU is gentle and the luminance consistency is relatively high.

[0066] For example, the image analysis unit 121F estimates the motion vectors using the Farneback method as an optical flow estimation algorithm. The Farneback method is a type of optical flow estimation algorithm and is used to estimate the motion vectors of pixels in a video. The Farneback method uses image convolution and polynomial expansion to estimate the pixel motion patterns between consecutive frames. The Farneback method is characterized by fast operation and is suitable for real-time applications. The Farneback method expresses the change in luminance between images as a polynomial expansion. The Farneback method can express the motion vectors of pixels as the coefficients of the polynomial by this expansion. First, the Farneback method processes the images between two consecutive frames by a convolution operation. The Farneback method calculates the change in luminance for each pixel by this convolution. Also, the Farneback method prepares data for obtaining different polynomial coefficients for each pixel. Next, the Farneback method estimates the coefficients of the polynomial for each pixel based on the convolved image data. Thereby, the image analysis unit 121F obtains the motion vectors for each pixel. The image analysis unit 121F smooths the obtained motion vectors for each pixel to smooth the motion pattern between consecutive pixels. Thereby, noise and discontinuities are reduced. The Farneback method is characterized by fast calculation and smooth results and is suitable for real-time applications.

[0067] For example, the image analysis unit 121F uses a prediction model to simulate and determine the degree of coincidence between the frame of the high-quality video data D3 and a specific frame of the low-quality video data D2. The prediction model is a model that has been machine-learned using, as teacher data, the degree of coincidence between the frame of the high-quality video data D3 acquired in the past and a specific frame of the low-quality video data D2. First, the user prepares, as teacher data, data obtained by measuring the degree of coincidence between the frame of the high-quality video data D3 acquired in the past and a specific frame of the low-quality video data D2. The user quantitatively evaluates the degree of coincidence between the frames using an index based on similarity or image comparison. Next, the user extracts features from the corresponding frames of each video. The features include the luminance and color information of the image, edges, corners, textures, and the like. The feature quantities serve as the input to the model. Next, the user trains a machine learning model using the teacher data. The user inputs the features extracted from the frame of the high-quality video data D3 and a specific frame of the low-quality video data D2 into the model and trains it to predict the degree of coincidence based on the teacher data. The architecture and hyperparameters of the model are selected according to the data and the problem. When the training is completed, the user applies the model to the frame of the high-quality video data D3 and a specific frame of the low-quality video data D2 to make a prediction. The model outputs a predicted value of the degree of coincidence. The user sets a certain threshold value based on the predicted value of the predicted degree of coincidence. As a result, the image analysis unit 121F determines whether the predicted value of the degree of coincidence exceeds the threshold value, and simulates and determines the degree of coincidence between the frame of the high-quality video data D3 and a specific frame of the low-quality video data D2. The user evaluates the performance of the prediction model and adjusts the architecture and hyperparameters of the model as necessary. The evaluation can be performed using the error between the actual degree of coincidence and the predicted degree of coincidence.

[0068] The determination unit 121G is a software module that determines the relative relationship between the information detected by the depth camera 111 and the information detected by the mirrorless single-lens camera 112. For example, the determination unit 121G determines the relative relationship between the depth data D1 and the high-quality video data D3 based on the degree of coincidence between the low-quality video data D2 and the high-quality video data D3. For example, the determination unit 121G aligns the frame with a high degree of coincidence with the specific frame of the low-quality video data D2 among the frames of the high-quality video data D3 and the specific frame of the depth data D1 at the same reference time. The specific frame of the depth data D1 is the frame acquired at the same time as the specific frame of the low-quality video data D2.

[0069] FIG. 6 is a diagram showing an example of the procedure of the process by the information processing apparatus 120. The process shown in FIG. 6 includes the process from starting the shooting by the camera unit 110 to generating the 3D model data D5.

[0070] First, the information processing apparatus 120 transmits a shooting start signal to the camera unit 110 (S201). The shooting start signal is a signal that serves as a trigger to start shooting. When the camera unit 110 receives the shooting start signal, it starts shooting the subject SU. The shooting start signal is an example of the first signal.

[0071] Next, the information processing apparatus 120 determines whether it is the shooting end timing (S202). For example, the shooting time is set in advance by the user. For example, the information processing apparatus 120 determines that it is the shooting end timing when the time elapsed since transmitting the shooting start signal reaches a predetermined shooting time.

[0072] If it is not the shooting end timing (S202; NO), the information processing apparatus 120 repeatedly executes the process of S202 at predetermined intervals.

[0073] When it is the shooting end timing (S202; YES), the information processing apparatus 120 transmits a shooting end signal to the camera unit 110 (S203). The shooting end signal is a signal that serves as a trigger to end the shooting. When the camera unit 110 receives the shooting end signal, it ends the shooting of the subject SU. The shooting end signal is an example of the second signal.

[0074] When the camera unit 110 finishes shooting the subject SU, the information processing apparatus 120 receives data from the camera unit 110 (S204). In S204, the information processing apparatus 120 receives depth data D1 and low-quality video data D2 from the depth camera 111. Also, the information processing apparatus 120 receives high-quality video data D3 from the mirrorless single-lens camera 112.

[0075] Next, the determination unit 121A of the information processing apparatus 120 determines whether the number of frames of the data is within a specific range (S205). In S205, the determination unit 121A determines whether the number of frames of the depth data D1 is within a specific range. Also, the determination unit 121A determines whether the number of frames of the low-quality video data D2 is within a specific range. Also, the determination unit 121A determines whether the number of frames of the high-quality video data D3 is within a specific range. For example, an appropriate number of frames is preset according to the shooting time, required accuracy, real-time performance, etc. For example, the determination unit 121A determines whether the number of frames of the data is within the set appropriate range. For example, the determination unit 121A determines whether the difference in the number of frames of the depth data D1A, depth data D1B, depth data D1C, depth data D1D obtained from each of the camera units 110A, camera unit 110B, camera unit 110C, camera unit 110D is within a predetermined range. Instead of determining whether the number of frames of the data is within a specific range, the number of frames of each type of data obtained from each camera unit 110 may be displayed as a list, and the user may be asked to determine whether the number of frames is within a specific range.

[0076] The cause of the abnormal number of frames of the depth data D1 acquired by the depth camera 111 may be caused by various factors.

[0077] For example, the abnormality in the number of frames of the depth data D1 is due to a hardware failure. When a failure or malfunction occurs in the hardware of the depth camera 111, the depth camera 111 may have difficulty acquiring data. Hardware failures include failures of the depth sensor S1, wiring problems, power supply abnormalities, etc.

[0078] For example, the abnormality in the number of frames of the depth data D1 is due to a data transmission problem. If the depth data D1 acquired by the depth camera 111 cannot be transmitted normally, the depth data D1 may have abnormal frame numbers or data loss. Data transmission problems may be affected by noise during transmission or communication errors.

[0079] For example, the abnormality in the number of frames of the depth data D1 is due to insufficient power supply. If the depth camera 111 is not supplied with sufficient power, the depth camera 111 may not be able to operate normally, and the acquisition of the depth data D1 may be interrupted or abnormal data may be obtained.

[0080] For example, the abnormality in the number of frames of the depth data D1 is due to a sensor calibration problem. If the depth camera 111 is not accurately calibrated, the depth data D1 may not be acquired correctly, and abnormal data may be generated.

[0081] For example, the abnormality in the number of frames of the depth data D1 is due to changes in the surrounding environment. The depth camera 111 may be affected by the surrounding environment. Changes in lighting conditions or changes in the arrangement of the object subject SU may cause abnormalities in the acquisition of the depth data D1.

[0082] For example, the abnormality in the number of frames of depth data D1 is caused by an increase in processing load. When there are many applications that perform processing simultaneously with the acquisition of depth data D1, the depth camera 111 may generate abnormal data by exceeding its processing capacity.

[0083] For example, the abnormality in the number of frames of depth data D1 is caused by a software bug. If there is a bug or defect in the software that controls the depth camera 111, the depth camera 111 may not be able to acquire normal data and may obtain abnormal data.

[0084] The cause of the abnormality in the number of frames of the low-quality video data D2 acquired by the depth camera 111 may be caused by various factors.

[0085] For example, the abnormality in the number of frames of the low-quality video data D2 is caused by a hardware failure. When a failure or defect occurs in the hardware of the depth camera 111, it may become difficult for the depth camera 111 to acquire normal data. Hardware failures include failures of the color sensor S2, wiring problems, abnormal power supply, etc.

[0086] For example, the abnormality in the number of frames of the low-quality video data D2 is caused by a data transmission problem. The acquisition and storage of the low-quality video data D2 involve data transmission. When noise or communication errors occur during data transmission, the low-quality video data D2 may have some frames missing or abnormal data may be recorded.

[0087] For example, the abnormality in the number of frames of the low-quality video data D2 is caused by a problem with the recording medium. If there is a problem with the recording medium for storing the low-quality video data D2, the writing or reading of the low-quality video data D2 may be interrupted, resulting in an abnormal number of frames.

[0088] For example, the abnormality in the number of frames of the low-quality video data D2 is caused by the frame rate setting. Depending on the settings of the depth camera 111, if the frame rate is not set appropriately, the acquisition interval of the low-quality video data D2 becomes abnormal, and the number of frames may not be as expected.

[0089] For example, the abnormality in the number of frames of the low-quality video data D2 is caused by the interference of the application during recording. When another application uses a large amount of system resources during the recording of the low-quality video data D2 by the depth camera 111, the acquisition of the low-quality video data D2 may be hindered, and the number of frames may become abnormal.

[0090] For example, the abnormality in the number of frames of the low-quality video data D2 is caused by the delay in data processing. When data processing is performed simultaneously with the acquisition of the low-quality video data D2, the data processing may be delayed, and the number of frames may not be accurately recorded.

[0091] For example, the abnormality in the number of frames of the low-quality video data D2 is caused by the calibration problem of the depth camera 111. The depth camera 111 requires accurate calibration. If the calibration is not performed correctly, the number of frames of the low-quality video data D2 may become abnormal.

[0092] The cause of the abnormality in the number of frames of the high-quality video data D3 acquired by the mirrorless single-lens camera 112 may be caused by various factors.

[0093] For example, the abnormality in the number of frames of the high-quality video data D3 is caused by hardware problems. When there are obstacles or malfunctions in the camera body or sensor, the mirrorless single-lens camera 112 may be prevented from acquiring normal data. Hardware problems include abnormal power supply, sensor failure, motor malfunction, etc.

[0094] For example, an abnormality in the number of frames of high-quality video data D3 is caused by a problem with the recording medium. If there is a problem with the memory card or storage device that stores the high-quality video data D3, data writing and reading may be interrupted, and the number of frames may become abnormal.

[0095] For example, an abnormality in the number of frames of high-quality video data D3 is caused by a recording time limit. The mirrorless single-lens camera 112 may have a recording time limit. If recording continuously for a long time, the mirrorless single-lens camera 112 may overheat and stop operating, and the number of frames may become abnormal.

[0096] For example, an abnormality in the number of frames of high-quality video data D3 is caused by a battery problem. When the battery is depleted or the power supply is unstable, the acquisition of high-quality video data D3 may be interrupted, or the mirrorless single-lens camera 112 may shut down, and the number of frames may become abnormal.

[0097] For example, an abnormality in the number of frames of high-quality video data D3 is caused by a delay in data processing. When data processing is performed simultaneously with the acquisition of high-quality video data D3, if the data processing is delayed, the mirrorless single-lens camera 112 may not be able to accurately acquire the next frame, and the number of frames may become abnormal.

[0098] For example, an abnormality in the number of frames of high-quality video data D3 is caused by external interference. When electromagnetic waves or wireless communication interfere around the camera, it will affect the acquisition of high-quality video data D3, and the number of frames may become abnormal.

[0099] For example, an abnormality in the number of frames of high-quality video data D3 is caused by a setting error. If the settings of the mirrorless single-lens camera 112 are incorrect, the frame rate and recording mode may not be appropriate, and the number of frames may become abnormal.

[0100] If the number of frames of the data is not within the specific range (S205; NO), the alert output unit 121B of the information processing apparatus 120 outputs an alert (S206). For example, the alert output unit 121B outputs an alert when the number of frames of the depth data D1 is not within the specific range. For example, the alert output unit 121B outputs an alert when the number of frames of the low-quality video data D2 is not within the specific range. For example, the alert output unit 121B outputs an alert when the number of frames of the high-quality video data D3 is not within the specific range. For example, the alert output unit 121B causes a message notifying that the number of frames of the data is not within the specific range to be displayed on the display 126. Then, the information processing apparatus 120 ends the process shown in FIG. 6.

[0101] If the number of frames of the data is within the specific range (S205; YES), the information processing apparatus 120 executes synchronization processing between the depth data D1 (S207).

[0102] FIG. 7 is a diagram showing an example of the procedure of the synchronization process between the depth data D1.

[0103] First, the shape information generation unit 121C of the information processing apparatus 120 reads a specific frame of the reference depth data D1 (S301). For example, the shape information generation unit 121C reads a frame other than the first frame of the reference depth data D1 as the specific frame. For example, the shape information generation unit 121C reads a frame other than the last frame of the reference depth data D1 as the specific frame. For example, the shape information generation unit 121C reads a frame after the first several frames of the reference depth data D1 as the specific frame. For example, the shape information generation unit 121C reads a frame before the last several frames of the reference depth data D1 as the specific frame. In this example, the shape information generation unit 121C reads a specific frame of the first depth data D1A acquired by the first depth camera 111A.

[0104] Next, the shape information generation unit 121C reads out one frame of the depth data D1 to be synchronized (S302). For example, the shape information generation unit 121C reads out a frame at the same timing as or close to the timing of the specific frame read out in S301. In this example, the shape information generation unit 121C reads out one frame of the second depth data D1B acquired by the second depth camera 111B. In the following description, the one frame read out in S302 is referred to as the first frame.

[0105] Next, the shape information generation unit 121C reads out a predetermined frame of the depth data D1 acquired by another depth camera 111 (S303). For example, when the other depth data D1 is not synchronized with the reference depth data D1, the shape information generation unit 121C reads out a frame at the same timing as or close to the timing of the specific frame read out in S301. For example, when the other depth data D1 is synchronized with the reference depth data D1, the shape information generation unit 121C reads out a frame aligned with the same reference time as the specific frame read out in S301. In this example, the shape information generation unit 121C reads out a predetermined frame of the third depth data D1C acquired by the third depth camera 111C. Also, the shape information generation unit 121C reads out a predetermined frame of the fourth depth data D1D acquired by the fourth depth camera 111D.

[0106] Next, the shape information generation unit 121C generates mesh data D4 (S304). When the process of S304 is executed for the first time, the shape information generation unit 121C generates the mesh data D4 using the frames read out in S301, S302, and S303.

[0107] Next, the shape information analysis unit 121D of the information processing apparatus 120 calculates the number of polygons of the mesh data D4 generated in S304 (S305).

[0108] Next, the shape information generation unit 121C determines whether the number of trials for generating the mesh data D4 and calculating the number of polygons has reached a specific number (S306). For example, the specific number that is the threshold for the number of trials is set to a larger number as the frame rate of the depth data D1 is higher. For example, the specific number that is the threshold for the number of trials is set to a smaller number as the frame rate of the depth data D1 is lower. In this example, it is assumed that the specific number that is the threshold for the number of trials is set to 4 times.

[0109] If the number of trials has not reached the specific number (S306; NO), the shape information generation unit 121C changes the frame of the depth data D1 to be synchronized (S307). For example, the shape information generation unit 121C changes the current frame to the frame one after the current frame. In this example, the shape information generation unit 121C changes the first frame of the second depth data D1B to the second frame one after the first frame.

[0110] Next, the shape information generation unit 121C regenerates the mesh data D4 (S304). The shape information generation unit 121C generates the mesh data D4 using the frames read in S301 and S303 and the frame changed in S307.

[0111] Next, the shape information analysis unit 121D recalculates the number of polygons of the mesh data D4 generated in S304 (S305).

[0112] The information processing apparatus 120 repeatedly executes the processes from S304 to S307 until the number of trials reaches the specific number, generates a plurality of mesh data D4, and obtains information on the number of polygons of the mesh data D4.

[0113] FIGs. 8 to 11 are diagrams showing an example of the mesh data D4 generated in S304 and the number of polygons calculated in S305.

[0114] FIG. 8 shows the first mesh data D4A generated using the first frame of the second depth data D1B. The number of polygons in the first mesh data D4A is "66628". The first mesh data D4A is an example of the first shape information.

[0115] FIG. 9 shows the second mesh data D4B generated using the second frame of the second depth data D1B. The 3D model of the second mesh data D4B has fewer irregularities than the 3D model of the first mesh data D4A. Therefore, the number of polygons in the second mesh data D4B is "66422", which is less than the number of polygons in the first mesh data D4A. The second mesh data D4B is an example of the second shape information.

[0116] FIG. 10 shows the third mesh data D4C generated using the third frame of the second depth data D1B. The third frame is the frame one after the second frame. The 3D model of the third mesh data D4C has more irregularities than the 3D model of the first mesh data D4A. Therefore, the number of polygons in the third mesh data D4C is "66636", which is more than the number of polygons in the first mesh data D4A. The third mesh data D4C is an example of the second shape information.

[0117] FIG. 11 shows the fourth mesh data D4D generated using the fourth frame of the second depth data D1B. The fourth frame is the frame one after the third frame. The 3D model of the fourth mesh data D4D has more irregularities than the 3D model of the third mesh data D4C. Therefore, the number of polygons in the fourth mesh data D4D is "66960", which is more than the number of polygons in the third mesh data D4C. The fourth mesh data D4D is an example of the second shape information.

[0118] Returning to the description of FIG. 7, when the number of trials reaches a specific number (S306; YES), the synchronization unit 121E of the information processing apparatus 120 synchronizes the reference depth data D1 and the depth data D1 to be synchronized (S308). For example, the synchronization unit 121E aligns a frame of the depth data D1 to be synchronized corresponding to the mesh data D4 with the smallest number of polygons with the same reference time as a specific frame of the reference depth data D1. In the examples shown in FIGS. 8 to 11, the mesh data D4 with the smallest number of polygons is the second mesh data D4B. Therefore, the synchronization unit 121E aligns the second frame of the second depth data D1B corresponding to the second mesh data D4B with the same reference time as the specific frame of the first depth data D1A.

[0119] Next, the information processing apparatus 120 determines whether all depth cameras 111 are synchronized (S309).

[0120] If all depth cameras 111 are not synchronized (S309; NO), the information processing apparatus 120 changes the depth data D1 to be synchronized (S310) and executes the processes from S301 to S309. In this example, the information processing apparatus 120 changes the depth data D1 to be synchronized from the second depth data D1B to the third depth data D1C and executes the processes from S301 to S309. After that, when executing the process of S310 again, the information processing apparatus 120 changes the depth data D1 to be synchronized from the third depth data D1C to the fourth depth data D1D and executes the processes from S301 to S309.

[0121] When all depth cameras 111 are synchronized (S309; YES), the information processing apparatus 120 ends the synchronization process between the depth cameras shown in FIG. 7.

[0122] Returning to the description of FIG. 6, after executing the process of S207, the information processing apparatus 120 executes a synchronization process between the depth data D1 and the high-quality video data D3 (S208).

[0123] FIG. 12 is a diagram showing an example of a procedure for synchronizing depth data D1 and high-quality video data D3. The information processing apparatus 120 executes the processing shown in FIG. 12 for each of the depth data D1 and the high-quality video data D3 acquired by the common camera unit 110.

[0124] First, the image analysis unit 121F of the information processing apparatus 120 reads a specific frame of the low-quality video data D2 (S401).

[0125] Next, the image analysis unit 121F reads a plurality of frames of the high-quality video data D3 (S402). For example, the image analysis unit 121F reads frames at the same timing or at a timing close to the specific frame read in S401.

[0126] Next, the image analysis unit 121F determines the degree of coincidence between the specific frame read in S401 and each of the plurality of frames read in S402 (S403). For example, the image analysis unit 121F determines the degree of coincidence by comparing the image features of the frames of the high-quality video data D3 with the image features of the specific frame of the low-quality video data D2. For example, the image analysis unit 121F compares the movement of the subject SU between consecutive frames of the high-quality video data D3 with the movement of the subject SU between consecutive frames including the specific frame of the low-quality video data D2 to determine the degree of coincidence. For example, the image analysis unit 121F uses a prediction model to simulate and determine the degree of coincidence between the frames of the high-quality video data D3 and the specific frame of the low-quality video data D2.

[0127] FIG. 13 is a diagram showing an example of a specific frame of the low-quality video data D2. FIG. 14 is a diagram showing an example of the first frame of the high-quality video data D3. FIG. 15 is a diagram showing an example of the second frame of the high-quality video data D3. FIG. 16 is a diagram showing an example of the third frame of the high-quality video data D3. FIG. 17 is a diagram showing an example of the fourth frame of the high-quality video data D3.

[0128] In the example shown in FIGS. 13 to 17, the specific frame of the low-quality video data D2 has the highest degree of coincidence with the second frame of the high-quality video data D3.

[0129] The determination unit 121G of the information processing apparatus 120 refers to the degree of coincidence determined in S403 and synchronizes the depth data D1 and the high-quality video data D3 (S404). For example, the determination unit 121G aligns the frame with the highest degree of coincidence with the specific frame of the low-quality video data D2 among the frames of the high-quality video data D3 and the specific frame of the depth data D1 at the same reference time. The specific frame of the depth data D1 is the frame acquired at the same time as the specific frame of the low-quality video data D2.

[0130] After executing the process of S404, the information processing apparatus 120 ends the synchronization process of the depth data D1 and the high-quality video data D3 shown in FIG. 12. The distance image data generated based on the depth data D1 information and displayed with different colors for each distance to the subject may be used instead of the low-quality data D2.

[0131] Returning to the description of FIG. 6, after executing the process of S208, the shape information generation unit 121C generates 3D model data D5 (S209). When generating the 3D model data D5, the shape information generation unit 121C generates mesh data D4 composed of a plurality of frames based on the correlation relationship between the depth data D1 synchronized in S207. Then, the shape information generation unit 121C generates 3D model data D5 composed of a plurality of frames based on the correlation relationship between the depth data D1 synchronized in S208 and the high-quality video data D3. The texture information and material information of the 3D model data D5 use the information acquired from the high-quality video data D3.

[0132] Although the present invention has been described using embodiments, the technical features of the present invention are not limited to the scope described in the embodiments. It is obvious to those skilled in the art that various changes or improvements can be made to the embodiments. It is obvious from the description of the claims that forms with such changes or improvements can also be included in the technical scope of the present invention.

[0133] The camera system 100 of the embodiment includes four camera units 110. However, the camera system 100 only needs to include a plurality of camera units 110, and is not limited to the configuration including four camera units 110. For example, the camera system 100 may include two camera units 110. For example, the camera system 100 may include three camera units 110. For example, the camera system 100 may include five or more camera units 110.

[0134] The information processing device 120 performs the processing related to calibration in the embodiment. However, the processing related to calibration is not limited to being performed by the information processing device 120. The processing related to calibration may be performed by the camera unit 110. The processing related to calibration may be performed by the depth camera 111. The processing related to calibration may be performed by the mirrorless single-lens camera 112.

[0135] The camera unit 110 of the embodiment acquires distance information of the subject SU by the depth camera 111. However, the camera unit 110 only needs to be able to acquire distance information of the subject SU, and is not limited to the configuration including the depth camera 111. For example, the camera unit 110 may include a configuration that acquires distance information of the subject SU by stereo vision. Stereo vision estimates distance information by measuring the parallax of the subject SU using two cameras. This method mimics the two eyes of a human, and performs triangulation from the distance between the two cameras and the measurement of the parallax to estimate the distance.

[0136] The camera unit 110 of the embodiment acquires the distance information of the subject SU by the depth camera 111. However, the camera unit 110 only needs to be able to acquire the distance information of the subject SU, and is not limited to the configuration including the depth camera 111. For example, the camera unit 110 may be configured to acquire the distance information of the subject SU by an ultrasonic sensor. The ultrasonic sensor generates sound waves and acquires distance information by measuring the time it takes for the sound waves to hit an object and return. The ultrasonic sensor is one of the general-purpose sensors commonly used for distance measurement.

[0137] The camera unit 110 of the embodiment captures high-quality videos by the mirrorless single-lens camera 112. However, the camera unit 110 only needs to be able to capture high-quality videos, and is not limited to the configuration including the mirrorless single-lens camera 112. For example, the camera unit 110 may include a digital cinema camera. The digital cinema camera combines the features of a cinema camera with digital technology. The digital cinema camera supports special shooting modes such as ultra-high resolution and RAW format, and is suitable for movie production and shooting of high-quality content.

[0138] The camera unit 110 of the embodiment captures high-quality videos by the mirrorless single-lens camera 112. However, the camera unit 110 only needs to be able to capture high-quality videos, and is not limited to the configuration including the mirrorless single-lens camera 112. For example, the camera unit 110 may include an action camera. The action camera is a camera that can capture high-resolution videos while being small and lightweight. The action camera is particularly suitable for use in outdoor and sports scenes. Some action cameras support shooting in 4K or high frame rates.

[0139] The camera unit 110 of the embodiment captures high-quality videos with the mirrorless single-lens camera 112. However, the camera unit 110 only needs to be able to capture high-quality videos and is not limited to the configuration with the mirrorless single-lens camera 112. For example, the camera unit 110 may be equipped with a drone camera. The drone camera is specialized for aerial photography and is suitable for capturing high-quality images and beautiful scenery. Some drone cameras are equipped with 4K video shooting and global shutter functions.

[0140] The camera unit 110 of the embodiment captures high-quality videos with the mirrorless single-lens camera 112. However, the camera unit 110 only needs to be able to capture high-quality videos and is not limited to the configuration with the mirrorless single-lens camera 112. For example, the camera unit 110 may be equipped with a video camera. The video camera is a camera designed for general consumers and prosumers and is suitable for high-resolution video shooting. The video camera may also provide functions specialized for video shooting, such as a stabilizer and XLR audio input.

[0141] The depth camera 111 of the embodiment is attached to the mirrorless single-lens camera 112. However, the depth camera 111 only needs to be provided at a position where calibration with the mirrorless single-lens camera 112 is possible and does not necessarily have to be attached to the mirrorless single-lens camera 112. For example, the depth camera 111 may be arranged at a position away from the mirrorless single-lens camera 112.

[0142] The shape information analysis unit 121D may analyze the frequency components of the mesh data D4 generated by the shape information generation unit 121C. The frequency components of the mesh data D4 refer to information indicating the details and changes of different spatial scales regarding the surface shape of the 3D model. The frequency components are useful for capturing factors that affect the appearance and look of the 3D model, such as the smoothness or roughness of the shape and fine features. The frequency components are mainly classified into low-frequency components and high-frequency components. The low-frequency components represent the general shape and curves of the 3D model. The general shape and curves of the 3D model include large curves and smooth surfaces. The low-frequency components are useful for capturing general contours and axial symmetry, etc., in order to indicate the basic shape of the 3D model. The high-frequency components represent the fine details and local changes of the 3D model. The fine details and local changes of the 3D model include acute angles, depressions, and surface characteristics with unevenness. The high-frequency components are used to enhance the reality and detail level of the 3D model. The concept of frequency components is related to techniques such as Fourier transform used in signal processing and image processing. The shape information analysis unit 121D can apply these techniques to convert the mesh data D4 into the frequency domain and analyze different frequency components. In that case, the synchronization unit 121E may align the frame of the depth data D1 of the synchronization target corresponding to the mesh data D4 with a small high-frequency component to the same reference time as a specific frame of the reference depth data D1. The high-frequency component is an example of a predetermined frequency component.

[0143] The display 126 may display the number of frames of the depth data D1. After the number of frames is displayed on the display 126, the image analysis unit 121F may perform analysis based on an instruction input to the input device 125.

[0144] The information processing device 120 may receive some frames of the low-quality video data D2 acquired by the depth camera 111 without receiving all the frames.

[0145] The information processing apparatus 120 may perform the synchronization process between the depth data D1 according to the following procedure instead of the procedure in FIG. 7. First, the first three frames of the respective depth data D acquired by the first depth camera 111A, the second depth camera 111B, the third depth camera 111C, and the fourth depth camera 111D are acquired. One frame is selected from each of the first depth camera 111A, the second depth camera 111B, the third depth camera 111C, and the fourth depth camera 111D1A to generate mesh data. Mesh data is generated for all combinations of the three frames acquired by each depth camera. When three frames are acquired, 81 pieces of mesh data of 3×3×3×3 are generated. Next, the mesh data with the smallest number of polygons among the 81 pieces of mesh data is selected. The frames that are the source of the selected mesh data are set as the frames corresponding to the reference times of the respective depth data.

[0146] The information processing apparatus 120 may perform the synchronization process between the depth data D1 by generating shape information from the depth data D and performing the process based on the volume or surface area of the shape information. In that case, the frame that is the source of the mesh data with the smallest volume or surface area of the shape information becomes the frame corresponding to the reference time.

[0147] Four first camera units arranged one by one at the first position P1, the third position P3, the fifth position P5, and the seventh position P7, and four second camera units arranged one by one at the second position P2, the fourth position P4, the sixth position P6, and the eighth position P8 may be caused to perform imaging at the same frame interval and at different timings from each other. In that case, the synchronization process is performed between the data acquired by the cameras belonging to the first camera unit, and the synchronization process is performed between the data acquired by the cameras belonging to the second camera unit, respectively. Then, by combining the data of the first camera unit and the data of the second camera unit, data with a dense frame rate can be generated.

[0148] In the above-described embodiment, the information processing apparatus 120 includes a shape information generation unit 121C that generates shape information representing a subject based on the first distance information acquired by the first camera unit 110A and the second distance information acquired by the second camera unit 110B, and a synchronization unit 121E that synchronizes the first distance information and the second distance information. The shape information generation unit 121C generates first shape information based on a specific frame of the first distance information and a first frame of the second distance information, and second shape information based on the specific frame and a second frame of the second distance information. The synchronization unit 121E aligns either the first frame or the second frame with the same reference time as the specific frame based on the result of comparing the first shape information and the second shape information. Thereby, even in a system where the camera unit 110A and the camera unit 110B cannot be synchronized with each other, it is possible to select the distance information with the closest detection time to the subject without using the time information of the detection of each frame.

[0149] The information processing apparatus 120 includes a shape information analysis unit 121D that analyzes the shape information generated by the shape information generation unit 121C. The synchronization unit 121E aligns either the first frame or the second frame with the same reference time as the specific frame based on the analysis result by the shape information analysis unit 121D. The analysis of the shape information is to calculate the number of polygons of the shape information. The synchronization unit 121E aligns the frame of the second distance information corresponding to the shape information with fewer polygons and the specific frame with the same reference time. Thereby, the shape information can be accurately synchronized.

[0150] The shape information generation unit 121C generates a mesh as the shape information. Thereby, the data volume can be reduced. Also, it facilitates the transfer of data to other devices.

[0151] The shape information analysis unit 121D analyzes the frequency components of the shape information generated by the shape information generation unit 121C, and the synchronization unit 121F aligns the frame of the second distance information corresponding to the shape information with a small predetermined frequency component among the shape information analyzed by the shape information analysis unit 121D to the same reference time as the specific frame. Thereby, it becomes possible to accurately synchronize the shape information.

[0152] The shape information analysis unit 121D analyzes the volume or surface area of the shape information generated by the shape information generation unit 121C, and the synchronization unit 121F aligns the frame of the second distance information corresponding to the shape information with the smallest volume or surface area among the shape information analyzed by the shape information analysis unit 121D to the same reference time as the specific frame. Thereby, it is possible to accurately synchronize the shape information.

[0153] A determination unit 121G that determines whether the number of frames of the first distance information and the second distance information is within a specific range is provided, and the shape information generation unit 121C starts generating shape information when the number of frames is within the specific range. An alert is output when the number of frames is not within the specific range. When the difference in the number of frames of each depth data is not within a predetermined range, depth data cannot be acquired by some camera units. It is possible to check in advance whether there is an error in the data, and it is possible to re-acquire the data before generating the 3D model data.

[0154] The number of frames of the first distance information and the second distance information is displayed on the display unit, and the shape information generation unit 121C generates shape information based on an instruction input to the input unit after the number of frames is displayed on the display unit. Thereby, before generating 3D model data with a large amount of calculation, the user can check in advance whether there is an error in the original data acquisition.

[0155] A transmission unit that transmits a first signal for starting shooting and a second signal for ending shooting to the first distance detection device and the second distance detection device, and a reception unit that receives distance information from the first distance detection device and the second distance detection device after transmitting a signal for ending shooting. Thereby, the processing of each camera unit can be reduced.

[0156] The shape information generation unit generates a series of shape information based on a series of first distance information and a series of second distance information synchronized by the synchronization unit. Thereby, a 3D video can be generated.

[0157] In the above embodiment, the information processing device 120 includes a reception unit that receives a series of distance information periodically detected and output by the first depth camera 111A and a series of first image information detected and output in synchronization with the distance information, and a series of second image information periodically detected and output by the first mirrorless single-lens camera, and a determination unit 121G that determines the relative relationship between the series of distance information and the series of second image information based on the degree of coincidence between the first image information and the second image information. Thereby, even when the first depth camera 111A and the first mirrorless single-lens camera cannot detect an object synchronously, the distance information of the object detected by the first depth camera 111A and the second image information of the object detected by the first mirrorless single-lens camera can be associated with each other without using information regarding the detection time, and information with close detection times can be associated with each other. Also, when the first depth camera 111A and the first mirrorless single-lens camera can detect an object synchronously, but one of the information cannot be acquired due to an error or the like, or when the detection periods of the first depth camera 111A and the first mirrorless single-lens camera are different, by re-determining the relative relationship of the information at regular intervals in the series of distance information, information with close detection times can always be associated with each other. Also, since it is not necessary to synchronize the detection devices that detect distance information and image information, the degree of freedom of the detection devices increases.

[0158] The information processing apparatus 120 is configured to include a shape information generation unit 121 that synthesizes a series of distance information and a series of second image information based on the relationship determined by the determination unit, and generates a series of shape information including color information. Thereby, color information acquired by the first mirrorless single-lens camera can be added to the shape information generated based on the distance information.

[0159] The information processing apparatus 120 is configured to include a transmission unit that transmits a first signal for starting detection to the first depth camera 111A and the first mirrorless single-lens camera. Thereby, an instruction to start detection can be sent to the first depth camera 111A and the first mirrorless single-lens camera at once, so that detection can be started at a close time.

[0160] In the claims, the description, and the drawings, the execution order of each process such as operations, procedures, steps, and stages in the apparatus, system, program, and method shown is not explicitly stated as "earlier" or "preceding" etc. In addition, it should be noted that the execution order of each process can be realized in any order unless the output of the previous process is used in the subsequent process. Even if the execution order of each process is described using "first," "next," etc. for convenience regarding the operation flow in the claims, the description, and the drawings, it does not mean that it is essential to implement in this order.

Explanation of Reference Numerals

[0161] 100 Camera system 110 Camera unit 110A First camera unit 110B Second camera unit 110C Third camera unit 110D Fourth camera unit 111 Depth camera 111A First depth camera 111B Second depth camera 111C Third depth camera 111D Fourth depth camera 112 Mirrorless single-lens camera 112A First Mirrorless Single-Lens Camera 112B Second Mirrorless Single-Lens Camera 112C Third Mirrorless Single-Lens Camera 112D Fourth Mirrorless Single-Lens Camera 120 Information Processing Device 121 CPU 121A Determination Unit 121B Alert Output Unit 121C Shape Information Generation Unit 121D Shape Information Analysis Unit 121E Synchronization Unit 121F Image Analysis Unit 121G Decision Unit 122 Main Memory 123 Input / Output Interface 124 Communication Device 125 Input Device 126 Display 127 Storage D1 Depth Data D1A First Depth Data D1B Second Depth Data D1C Third Depth Data D1D Fourth Depth Data D2 Low-Quality Video Data D2A First Low-Quality Video Data D2B Second Low-Quality Video Data D2C Third Low-Quality Video Data D2D Fourth Low-Quality Video Data D3 High-Quality Video Data D3A First High-Quality Video Data D3B Second High-Quality Video Data D3C Third High-Quality Video Data D3D Fourth High-Quality Video Data D4 Mesh Data D4A First Mesh Data D4B Second Mesh Data D4C Third Mesh Data D4D Fourth Mesh Data D5 3D Model Data P1 Position 1 P2 Position 2 P3 Position 3 P4 Position 4 P5 Position 5 P6 Position 6 P7 Position 7 P8 Position 8 S1 Depth Sensor S2 Color Sensor SU Subject

Claims

1. An information processing apparatus that processes information acquired by a distance detection device, comprising: a shape information generation unit that generates shape information representing a subject based on first distance information acquired by a first distance detection device and second distance information acquired by a second distance detection device; a synchronization unit that synchronizes the first distance information and the second distance information; the shape information generation unit generates first shape information based on a specific frame of the first distance information and a first frame of the second distance information, and second shape information based on the specific frame and a second frame of the second distance information; the synchronization unit aligns either the first frame or the second frame with the same reference time as the specific frame based on a result of comparing the first shape information and the second shape information.

2. further comprising a shape information analysis unit that analyzes the shape information generated by the shape information generation unit; the synchronization unit aligns either the first frame or the second frame with the same reference time as the specific frame based on an analysis result by the shape information analysis unit. The information processing apparatus according to claim 1.

3. the shape information analysis unit calculates the number of polygons of the shape information generated by the shape information generation unit; the synchronization unit aligns the frame of the second distance information corresponding to the shape information with a smaller number of polygons with the specific frame at the same reference time. The information processing apparatus according to claim 2.

4. the shape information generation unit generates a mesh as the shape information. The information processing apparatus according to any one of claims 1 to 3.

5. the shape information analysis unit analyzes a frequency component of the shape information generated by the shape information generation unit; the synchronization unit aligns the frame of the second distance information corresponding to the shape information with a smaller predetermined frequency component among the shape information analyzed by the shape information analysis unit with the specific frame at the same reference time. The information processing apparatus according to claim 2.

6. the shape information analysis unit calculates the volume of the shape information generated by the shape information generation unit; the synchronization unit aligns the frame of the second distance information corresponding to the shape information with a smaller volume with the specific frame at the same reference time. The information processing apparatus according to claim 2.

7. the shape information analysis unit calculates the surface area of the shape information generated by the shape information generation unit; The synchronization unit aligns the frame of the second distance information corresponding to the shape information with a small surface area and the specific frame at the same reference time, and the information processing apparatus according to claim 2.

8. It includes a determination unit that determines whether the number of frames of the first distance information and the second distance information is within a specific range, The shape information generation unit starts generating shape information when the number of frames is within a specific range, and the information processing apparatus according to any one of claims 1 to 7.

9. It includes an alert output unit that outputs an alert, The alert output unit outputs an alert when the number of frames is not within a specific range, and the information processing apparatus according to claim 8.

10. It includes a display unit that displays information, and An input unit for inputting a user's instruction, and the shape information generation unit generates shape information based on an instruction input to the input unit after the number of frames is displayed on the display unit, and the information processing apparatus according to any one of claims 1 to 9. The display unit displays the number of frames of the first distance information and the second distance information, The shape information generation unit generates shape information based on an instruction input to the input unit after the number of frames is displayed on the display unit, and the information processing apparatus according to any one of claims 1 to 9.

11. A transmission unit that transmits a first signal for starting shooting and a second signal for ending shooting to the first distance detection device and the second distance detection device, A receiving unit that receives the first distance information and the second distance information acquired based on the first signal after transmitting the second signal, and the information processing apparatus according to any one of claims 1 to 10.

12. The synchronization unit synchronizes a series of the first distance information and a series of the second distance information, and the information processing apparatus according to any one of claims 1 to 11.

13. The shape information generation unit generates a series of shape information based on a series of the first distance information and a series of the second distance information synchronized by the synchronization unit, and the information processing apparatus according to claim 12.

14. A camera system that acquires information by a distance detection device, A first distance detection device, A second distance detection device, and An information processing apparatus that processes the distance information of the subject acquired by the distance detection device, and the information processing apparatus includes Based on the first distance information acquired by the first distance detection device and the second distance information acquired by the second distance detection device, a shape information generation unit that generates shape information representing the subject, and A synchronization unit that synchronizes the first distance information and the second distance information. ​ The shape information generation unit generates first shape information based on the specific frame of the first distance information and the first frame of the second distance information, and second shape information based on the specific frame and the second frame of the second distance information. The synchronization unit aligns either the first frame or the second frame with the same reference time as the specific frame based on the result of comparing the first shape information and the second shape information, in a camera system.

15. An information processing method for processing information acquired by a distance detection device, comprising: generating shape information representing a subject based on first distance information acquired by a first distance detection device and second distance information acquired by a second distance detection device; synchronizing the first distance information and the second distance information; generating first shape information based on the specific frame of the first distance information and the first frame of the second distance information, and second shape information based on the specific frame and the second frame of the second distance information; aligning either the first frame or the second frame with the same reference time as the specific frame based on the result of comparing the first shape information and the second shape information, in an information processing method.

16. A program for causing a computer to function as an information processing device that processes information acquired by a distance detection device, comprising: causing the computer to function as a shape information generation unit that generates shape information representing a subject based on first distance information acquired by a first distance detection device and second distance information acquired by a second distance detection device; function as a synchronization unit that synchronizes the first distance information and the second distance information; wherein the shape information generation unit generates first shape information based on the specific frame of the first distance information and the first frame of the second distance information, and second shape information based on the specific frame and the second frame of the second distance information; and the synchronization unit aligns either the first frame or the second frame with the same reference time as the specific frame based on the result of comparing the first shape information and the second shape information, in a program.

Citation Information

Patent Citations

  • Three-dimensional measuring device

    JP2002031513A