Information Processing Apparatus, Information Processing Method, and Computer Program

The information processing apparatus enhances the accuracy of specifying positions in three-dimensional models by using triangulation and epipolar guidance, addressing issues of hidden or missing parts in multi-viewpoint image models.

JP7708089B2Active Publication Date: 2025-07-15SONY GROUP CORP
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2022501829
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-02-19
Filing Date
2021-02-10
Publication Date
2025-07-15
Estimated Expiration
2041-02-10

AI Technical Summary

Technical Problem

Existing methods for generating three-dimensional models from multi-viewpoint images struggle with accurately specifying positions, especially when the model is of low quality or when the desired position is hidden by other parts of the model.

Method used

An information processing apparatus that displays a three-dimensional model based on multiple captured images, allows users to select corresponding positions in two images, and calculates the desired position using triangulation, with optional features like epipolar line guidance and image manipulation.

Benefits of technology

Enables accurate specification of positions within the three-dimensional model, even when parts are missing or hidden, by leveraging triangulation and epipolar constraints to determine precise target points.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007708089000002
    Figure 0007708089000002
  • Figure 0007708089000003
    Figure 0007708089000003
  • Figure 0007708089000004
    Figure 0007708089000004
Patent Text Reader

Abstract

[Problem] To provide an information-processing device, an information-processing method, and a computer program which enable designation of any position of a 3D model with high precision. [Solution] The information processing device according to the present disclosure comprises: a first display control unit that displays, on a display device, a 3D model of an object based on a plurality of captured images obtained by imaging the object from a plurality of points of view; a second display control unit that displays, on the display device, a first captured image from a first point of view and a second captured image from a second point of view, of the plurality of captured images; a position acquisition unit that acquires positional information of a first position included in the object in the first captured image, and obtains positional information of a second position included in the object in the second captured image; a position calculation unit that calculates a third position included in the 3D model on the basis of information about the first point of view and the second point of view, the positional information of the first position, and the positional information of the second position; and a third display control unit that displays, on the display device, positional information of the third position superimposed on the 3D model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to an information processing apparatus, an information processing method, and a computer program.

Background Art

[0002] A technique for generating a three-dimensional model of a subject based on multi-viewpoint images is known. For example, there is a technique for generating a silhouette image by using the difference between a foreground image and a background image, and applying a volume intersection method to the multi-viewpoint silhouette images to calculate an intersection region, thereby generating a three-dimensional model.

[0003] When the three-dimensional model generated in this way is displayed on a display device, there is a demand that a viewer (user) specifies an arbitrary position included in the three-dimensional model and, for example, marks that position. However, when the accuracy of the shape of the generated three-dimensional model is not good (for example, when a part of the three-dimensional model is missing), the position of the three-dimensional model of the subject cannot be accurately specified. Similarly, when the position to be specified is hidden by a part of the body (for example, when the calf is hidden by the hand), the desired position cannot be accurately specified.

[0004] The following Patent Document 1 discloses a method for generating a three-dimensional model with high accuracy. However, it does not describe a method for correctly specifying a desired position in the three-dimensional model when a low-quality three-dimensional model is generated.

Prior Art Documents

Patent Documents

[0005]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0006] The present disclosure provides an information processing apparatus, an information processing method, and a computer program that can specify any position included in a three-dimensional model with high accuracy.

Means for Solving the Problems

[0007] The information processing apparatus of the present disclosure includes: a first display control unit that displays, on a display device, a three-dimensional model of a subject based on a plurality of captured images obtained by capturing the subject from a plurality of viewpoints; a second display control unit that displays, on the display device, a first captured image of a first viewpoint and a second captured image of a second viewpoint among the plurality of captured images; a position acquisition unit that acquires position information of a first position included in the subject in the first captured image and acquires position information of a second position included in the subject in the second captured image; a position calculation unit that calculates a third position included in the three-dimensional model based on information regarding the first viewpoint and the second viewpoint, the position information of the first position, and the position information of the second position; and a third display control unit that superimposes the position information of the third position on the three-dimensional model and displays the superimposed model on the display device.

[0008] The position calculation unit may calculate a position at which a straight line obtained by projecting the first position according to the line-of-sight direction of the first viewpoint intersects a straight line obtained by projecting the second position according to the line-of-sight direction of the second viewpoint. The calculated position is the third position.

[0009] The position calculation unit may calculate the third position based on the principle of triangulation.

[0010] The information processing apparatus includes an image selection unit that selects captured images of two viewpoints that are closest to the viewpoint of a user who views the three-dimensional model from the plurality of captured images. The second display control unit may display the first captured image and the second captured image selected by the image selection unit.

[0011] The information processing apparatus includes an image selection unit that selects the first captured image and the second captured image based on selection information for specifying the first captured image and the second captured image. The second display control unit may display the first captured image and the second captured image selected by the image selection unit.

[0012] The information processing apparatus includes an instruction information acquisition unit that acquires instruction information for instructing at least one of enlargement or reduction of the first captured image and the second captured image. The second display control unit may enlarge or reduce at least one of the first captured image and the second captured image based on the instruction information.

[0013] The position acquisition unit may acquire the position information of the first position from the user's operating device in a state where the first captured image is enlarged, and acquire the position information of the second position from the operating device in a state where the second captured image is enlarged.

[0014] The information processing apparatus includes an instruction information acquisition unit that acquires instruction information for instructing on or off of the display of at least one of the first captured image and the second captured image. The second display control unit may switch on or off the display of at least one of the first captured image and the second captured image based on the instruction information.

[0015] The second display control unit superimposes the position information of the first position on the first captured image and displays it on the display device. The second display control unit may superimpose the position information of the second position on the second captured image and display it on the display device.

[0016] The information processing apparatus includes a guide information generation unit that generates guide information including candidates for the second position based on the information regarding the first viewpoint and the second viewpoint and the position information of the first position, and a fourth display control unit that superimposes the guide information on the second captured image and displays it on the display device. It may be provided with.

[0017] The guide information may be an epipolar line.

[0018] The information processing apparatus includes an instruction information acquisition unit that acquires instruction information for instructing movement of the position information of the third position. The third display control unit moves the position information of the third position based on the instruction information. The second display control unit may move the position information of the first position and the position information of the second position in response to the movement of the position information of the third position.

[0019] The information processing method of the present disclosure displays a three-dimensional model of a subject on a display device based on a plurality of captured images of the subject captured from a plurality of viewpoints. The first captured image of the first viewpoint and the second captured image of the second viewpoint among the plurality of captured images are displayed on the display device. The position information of the first position included in the subject in the first captured image is acquired, and the position information of the second position included in the subject in the second captured image is acquired. Based on the information regarding the first viewpoint and the second viewpoint, the position information of the first position, and the position information of the second position, a third position included in the three-dimensional model is calculated. The position information of the third position is superimposed on the three-dimensional model and displayed on the display device.

[0020] The computer program of the present disclosure includes steps of: displaying a three-dimensional model of a subject on a display device based on a plurality of captured images of the subject captured from a plurality of viewpoints; displaying, on the display device, the first captured image of the first viewpoint and the second captured image of the second viewpoint among the plurality of captured images; acquiring the position information of the first position included in the subject in the first captured image and the position information of the second position included in the subject in the second captured image; calculating a third position included in the three-dimensional model based on the information regarding the first viewpoint and the second viewpoint, the position information of the first position, and the position information of the second position; superimposing the position information of the third position on the three-dimensional model and displaying it on the display device. Execute on a computer.

Brief Description of the Drawings

[0021]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23

Figure 24

Figure 25

Figure 26

Figure 27

Figure 28

Figure 29

Embodiments for Carrying Out the Invention

[0022] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings. In one or more embodiments shown in the present disclosure, elements included in each embodiment can be combined with each other, and the resulting combination is also part of the embodiments shown in the present disclosure.

[0023] (First Embodiment) FIG. 1 is a block diagram of an information processing system 100 including an information processing apparatus according to the first embodiment. The outline of the information processing system 100 in FIG. 1 will be described.

[0024] First, the problems of this embodiment will be described. Based on a plurality of captured images (camera images) of a subject captured from a plurality of viewpoints in advance, three-dimensional model information representing a three-dimensional model of the subject is generated. The information processing apparatus 101 displays a three-dimensional image including the three-dimensional model information of the subject on the display unit 301. A user who is an operator or viewer of the information processing apparatus 101 designates an arbitrary position (desired position) in the displayed three-dimensional model, for example, and there is a request to mark the designated position. However, when a part of the three-dimensional model is missing due to the accuracy of the generation of the three-dimensional model, or when the position to be designated in the three-dimensional model is hidden by another part of the three-dimensional model, the user cannot accurately designate the position to be designated. One of the features of this embodiment is to solve this problem by the following method.

[0025] The information processing apparatus 101 selects and displays two captured images with different viewpoints from the plurality of captured images that are the generation sources of the three-dimensional model. The user specifies positions (first position, second position) corresponding to a desired position in the three-dimensional model in the two displayed captured images, and selects the specified positions as feature points.

[0026] For example, suppose the three-dimensional model represents a human and a specific position on the head of the three-dimensional model is to be designated. In this case, the user specifies the positions corresponding to the specific position on the head of the three-dimensional model in the two captured images respectively, and selects the specified positions (first position, second position) as feature points.

[0027] The information processing apparatus 101 performs a triangulation operation based on the position information of the two selected feature points (the position information of the first position, the position information of the second position) and the information (position, orientation, etc.) regarding the viewpoints of the two captured images respectively. The position (third position) calculated by the triangulation operation is specified as the target point. The information processing apparatus 101 superimposes and displays the other party of the target point (the position information of the third position) on the three-dimensional model. Thereby, the user can correctly designate a desired position in the three-dimensional model. In this embodiment, even in a situation where a desired position cannot be correctly specified in a 3D model, if corresponding feature points are correctly selected in two captured images, the position corresponding to the two feature points can be specified with high accuracy as a target point based on the principle of triangulation. Hereinafter, the information processing system 100 according to this embodiment will be described in detail.

[0028] The information processing system 100 in FIG. 1 includes an information processing device 101, an operation unit 201, and a display unit 301.

[0029] The operation unit 201 is an operation device for a user to input various instructions or data. For example, the operation unit 201 is an input device such as a keyboard, a mouse, a touch panel, or a button. The operation unit 201 may be formed in the same housing as the information processing device 101. For example, the information processing device 101 and the operation unit 201 may be configured to be provided in a single smartphone, tablet terminal, or head-mounted display. Alternatively, the operation unit 201 may be formed as a device independent of the housing of the information processing device 101 and connected to the information processing device 101 by a wireless or wired cable.

[0030] The display unit 301 is a display device for displaying data, such as a liquid crystal display device, an organic EL display device, or a plasma display device. As an example, the display unit 301 is a head-mounted display, a 2D monitor, a 3D monitor, or the like. The display unit 301 may be formed in the same housing as the information processing device 101. For example, the information processing device 101 and the display unit 301 may be configured to be provided in a single smartphone, tablet terminal, or head-mounted display. Alternatively, the display unit 301 may be configured as a display independent of the housing of the information processing device 101 and connected to the information processing device 101 by a wireless or wired cable.

[0031] The information processing apparatus 101 includes an information storage unit 10, a display control unit 20, an interaction detection unit 30, a position storage unit 40, a target point calculation unit (triangulation calculation unit) 50, and an image selection unit 60. Some or all of these elements included in the information processing apparatus 101 are configured by hardware, software, or a combination thereof. The hardware includes, for example, a processor such as a CPU or a dedicated circuit. The information storage unit 10 and the position storage unit 40 are configured by a storage device such as a memory device or a hard disk device. The information storage unit 10 and the position storage unit 40 may be provided as an external device of the information processing apparatus 101 or as a server on a communication network. Further, a clock for counting time may be provided in the information processing apparatus 101.

[0032] The interaction detection unit 30 detects instruction information or data input by the user via the operation unit 201. The interaction detection unit 30 includes an operation instruction acquisition unit 31 and a position acquisition unit 32. The operation instruction acquisition unit 31 detects various instruction information input from the operation unit 201. The detected instruction information is output to the display control unit 20 or the image selection unit 60 according to the type of the instruction information. The position acquisition unit 32 acquires, as feature points, the positions selected by the user in the captured image from the operation unit 201, and stores the position information of the acquired feature points in the position storage unit 40.

[0033] The information storage unit 10 stores, for each of a plurality of frames (a plurality of times), a plurality of captured images (camera images) obtained by capturing a subject with imaging cameras corresponding to a plurality of viewpoints. The plurality of captured images are frame-synchronized. Information on the shooting time is attached to the plurality of captured images, and the plurality of captured images may be frame-synchronized based on the shooting time. The captured image is a color image, and is, for example, an RGB image. The plurality of captured images for a plurality of frames of one viewpoint correspond to a moving image of the viewpoint.

[0034] The information storage unit 10 stores 3D model information representing a 3D model of a subject generated by modeling based on captured images from a plurality of viewpoints. The 3D model information is data representing the 3D shape of the subject. If there are M frames of the captured images for each viewpoint, there are also M pieces of 3D model information accordingly. However, the number of frames and the number of pieces of 3D model information do not have to match. The 3D model can be generated by a modeling method such as Visual Hull based on captured images from a plurality of viewpoints. Hereinafter, an example of generating a 3D model will be described.

[0035] For a plurality of imaging cameras corresponding to a plurality of viewpoints in advance, camera parameters are calculated by calibration. The camera parameters include internal parameters and external parameters. The internal parameters are parameters specific to the imaging camera. As an example, they include the distortion of the camera lens, the inclination of the image sensor and the lens (distortion coefficient), the image center, and the image size. The external parameters include the positional relationship of the plurality of imaging cameras (the position and orientation of the imaging camera, etc.), the center coordinates (Translation) of the lens in the world coordinate system, the direction (Rotation) of the optical axis of the lens, etc. As calibration methods, there are Zhang's method using a chessboard, a method of capturing a 3D object to obtain camera parameters, a method of using a projected image by a projector to obtain camera parameters, etc.

[0036] The captured images corresponding to a plurality of viewpoints are corrected using the internal parameters of the plurality of imaging cameras. One of the plurality of imaging cameras is used as a reference camera, and the others are used as reference cameras. The frames of the captured images of the reference cameras are synchronized with the frames of the captured images of the reference camera.

[0037] For the plurality of captured images, using the difference between the foreground image (image of the subject) and the background image, a silhouette image of the subject is generated by background difference processing. The silhouette image is represented, for example, by binarizing the silhouette indicating the range in which the subject is included in the captured image.

[0038] Modeling of a subject is performed using a plurality of silhouette images and camera parameters. For modeling, for example, the Visual Hull method can be used. In the Visual Hull method, each silhouette image is back-projected onto the original three-dimensional space, and the intersection of the respective visual volumes is obtained as the Visual Hull. By applying a method such as the marching cubes method to the voxel data of the Visual Hull, a plurality of meshes are created.

[0039] The three-dimensional positions of each point (Vertex) constituting the mesh and the geometric information (Geometry) indicating the connection of each point (Polygon) are generated as three-dimensional shape data (polygon model). By configuring the three-dimensional shape data with polygon data, the data volume can be reduced compared to the case of voxel data. By performing a smoothing process on the polygon model, the surface of the polygon may be smoothed.

[0040] Texture mapping is performed to overlay the image of the mesh on each mesh of the three-dimensional shape data, and the three-dimensional shape data after texture mapping is used as three-dimensional model information representing the three-dimensional model of the subject. However, the three-dimensional shape data without texture mapping may also be used as the three-dimensional model information.

[0041] The method for generating the three-dimensional model is not limited to the method described above, and any method may be used as long as it is generated from a plurality of captured images corresponding to a plurality of viewpoints.

[0042] In addition to the three-dimensional model information and the captured images of a plurality of viewpoints, the information storage unit 10 may store other information such as camera parameters. Also, data of background images for displaying the three-dimensional model may be stored for a plurality of viewpoints.

[0043] The display control unit 20 includes a model display control unit (first display control unit) 21, an image display control unit (second display control unit) 22, and a target point display control unit (third display control unit) 23.

[0044] Based on the user's instruction information provided by the operation instruction acquisition unit 31, the model display control unit 21 reads out the three-dimensional model information from the information storage unit 10. The model display control unit 21 displays the three-dimensional model represented by the three-dimensional model information on the display unit 301 from the viewpoint indicated by the instruction information (the viewpoint from which the user views the three-dimensional model). The viewpoint for displaying the three-dimensional model may be predetermined. The viewpoint for viewing the three-dimensional model may be variable freely by 360 degrees in the horizontal direction, for example. A three-dimensional moving image may be reproduced by synthesizing the three-dimensional model with a background image and displaying the three-dimensional image including the three-dimensional model in a time series.

[0045] During the reproduction of the three-dimensional moving image, the model display control unit 21 may temporarily stop the reproduction, fast forward, or rewind based on the user's instruction information. In this case, for example, when the three-dimensional model of the desired frame is being displayed at the position in the three-dimensional model that the user wants to specify, the user can input a pause instruction to display the three-dimensional model in a stationary state at that frame. Alternatively, based on the user's instruction information from the operation unit 201, the three-dimensional model information of a specific frame may be read out and the three-dimensional model represented by the three-dimensional model information may be displayed.

[0046] FIG. 2 shows a display example of a three-dimensional image 70 including a three-dimensional model in a certain frame. The three-dimensional model 71 represents a baseball player. The user is trying to evaluate the pitching form of the baseball player. The background image includes a mound 72 and a pitcher's plate 73. The three-dimensional model 71 includes a glove 71A and a cap 71B.

[0047] The image selection unit 60 selects captured images of two different viewpoints from among a plurality of captured images (a plurality of frame-synchronized captured images that are the basis for generating the three-dimensional model) corresponding to the frames of the three-dimensional image. As a method of selecting two captured images, the captured images of the two viewpoints closest to the viewpoint of the three-dimensional image may be selected, or the captured images of the two viewpoints may be selected arbitrarily or randomly. Alternatively, the user may use the operation unit 201 to input instruction information for selecting the captured images of the two viewpoints, and the image selection unit 60 may select the two captured images based on this instruction information. For example, the user may select a captured image of a viewpoint in which the position to be specified in the three-dimensional model is most clearly reflected.

[0048] Figures 3 and 4 show examples of the captured images 110 and 120 of the two viewpoints selected by the image selection unit 60. The captured images 110 and 120 in Figures 3 and 4 include the baseball player 91 who is the subject. The captured images 110 and 120 also include various objects or people in the actual shooting environment. For example, the captured image 110 in Figure 3 includes imaging cameras 92A, 92B, and 92C, which are part of the imaging cameras of a plurality of viewpoints surrounding the subject. The captured image 120 in Figure 4 includes imaging cameras 92B, 92C, and 92D, which are part of the imaging cameras of a plurality of viewpoints surrounding the subject. The imaging cameras 92B and 92C are commonly included in the captured images 110 and 120 of Figures 3 and 4, but the imaging camera 92A is only included in the captured image 110 in Figure 3, and the imaging camera 92D is only included in the captured image 120 in Figure 4. Elements other than the imaging cameras (for example, photography staff, various equipment in the photography studio, etc.) may also be included in the captured images of Figures 3 and 4, but are omitted from the illustration.

[0049] The image display control unit 22 displays the two selected captured images simultaneously or in sequence. As an example, the two selected captured images are reduced and arranged (composed) side by side within the three-dimensional image 71. As an example, the arrangement position of the two reduced captured images is a position where the three-dimensional model 71 is not displayed. That is, the two captured images are arranged at a position where the three-dimensional model 71 is not hidden by the two captured images.

[0050] FIG. 5 shows an example in which the captured images 110 and 120 of FIGS. 3 and 4 are reduced and combined with the three-dimensional image 70 of FIG. 2. The captured images 110 and 120 are reduced and displayed in a region of the three-dimensional image 70 that does not include the upper subject.

[0051] In the example of FIG. 5, the two selected captured images are combined with the three-dimensional image 71, but these two captured images and the three-dimensional image 71 may be displayed in separate regions obtained by dividing the screen of the display unit 301 into three. In this case, the three-dimensional image 71 is also reduced and displayed.

[0052] Further, the two selected captured images may be displayed as a window different from the three-dimensional image 71.

[0053] Further, instead of displaying the two captured images simultaneously, initially only one of the captured images may be displayed. After feature points are selected for the displayed captured image, the other captured image may be displayed. When the other captured image is displayed, the display of the one captured image may be turned off. There are various variations in the display forms of the two captured images, and other forms are also possible.

[0054] The user uses the operation unit 201 to select positions corresponding to each other as feature points (first feature point, second feature point) for the captured images 110 and 120 of the two viewpoints. A specific example will be described below.

[0055] First, the user inputs enlargement instruction information for the captured image 110 using the operation unit 201. The operation instruction acquisition unit 31 of the interaction detection unit 30 acquires the enlargement instruction information for the captured image 110 and instructs the image display control unit 22 to perform an enlarged display of the captured image 110. The image display control unit 22 performs an enlarged display of the captured image 110. At this time, the three-dimensional image 70 and the captured image 120 may be temporarily hidden by the enlarged captured image 110.

[0056] The user designates, using the operation unit 201, a desired position (first position) on the head of the subject in the enlarged captured image 110 as the first feature point. The position acquisition unit 32 of the interaction detection unit 30 acquires the position information of the feature point from the operation unit 201, and stores the acquired position information in the position storage unit 40. Since the captured image 110 is a two-dimensional image, the first feature point is two-dimensional coordinates. When expressing the two-dimensional image in the uv coordinate system, the first feature point corresponds to the first uv coordinates. The image display control unit 22 superimposes and displays the first feature point on the captured image 110. The first feature point is displayed, for example, by a mark of a predetermined shape.

[0057] FIG. 6 shows a state where the user has selected the first feature point in the enlarged captured image 110. The first feature point 91A is displayed superimposed on the two-dimensional subject.

[0058] After the user selects the first feature point, the user inputs reduction instruction information for the captured image 110. The operation instruction acquisition unit 31 of the interaction detection unit 30 acquires the reduction instruction information for the captured image 110, and instructs the image display control unit 22 to perform a reduced display of the captured image 110. The image display control unit 22 performs a reduced display of the captured image 110. The captured image 110 returns to the original reduced size (see FIG. 5). The display of the first feature point 91A is included in the reduced captured image 110.

[0059] Next, the user inputs enlargement instruction information for the captured image 120 using the operation unit 201. The operation instruction acquisition unit 31 of the interaction detection unit 30 acquires the enlargement instruction information for the captured image 120, and instructs the image display control unit 22 to perform an enlarged display of the captured image 120. The image display control unit 22 performs an enlarged display of the captured image 120. At this time, the three-dimensional image 70 and the captured image 110 may be temporarily hidden by the enlarged captured image 120.

[0060] The user designates, using the operation unit 201, a position (second position) corresponding to the position designated in the captured image 110 as a second feature point in the enlarged captured image 120. That is, the user selects the same position on the subject in the captured images 110 and 120 as the first feature point and the second feature point. If the position is clearly shown in the two captured images 110 and 120, the user can easily select the positions corresponding to each other in the two captured images 110 and 120. The position acquisition unit 32 of the interaction detection unit 30 acquires the position information of the second feature point from the operation unit 201, and stores the acquired position information in the position storage unit 40 as the second feature point. Since the captured image 120 is a two-dimensional image, the second feature point is two-dimensional coordinates. When expressing the two-dimensional image in the uv coordinate system, the second feature point corresponds to the second uv coordinates. The image display control unit 22 superimposes and displays the second feature point (position information of the second position) on the captured image 120.

[0061] FIG. 7 shows a state in which the user has selected a second feature point in the enlarged captured image 120. The second feature point 91B is displayed superimposed on the subject.

[0062] After the user selects the second feature point, the user inputs reduction instruction information for the captured image 120. The operation instruction acquisition unit 31 of the interaction detection unit 30 acquires the reduction instruction information for the captured image 120, and instructs the image display control unit 22 to perform a reduced display of the captured image 120. The image display control unit 22 performs a reduced display of the captured image 120. The captured image 120 returns to the original reduced size (see FIG. 5). The display of the second feature point 91B is included in the reduced captured image 120.

[0063] FIG. 8 shows a display state after the selection of the first feature point and the second feature point is completed in the state of FIG. 5. The first feature point 91A is displayed in the reduced captured image 110. The second feature point 91B is displayed in the reduced captured image 120.

[0064] When the position information of the first feature point and the position information of the second feature point are stored in the position storage unit 40, the target point calculation unit 50 performs a triangulation operation based on the first feature point and the second feature point based on the information regarding the viewpoints of the two selected captured images (110, 120), thereby calculating the position (the third position) in the three-dimensional model. The calculated position is in three-dimensional coordinates and corresponds to the target point.

[0065] FIG. 9 schematically shows an example of the triangulation operation. There is a coordinate system (XYZ coordinate system) representing the space in which the three-dimensional model exists. By triangulation, a point where a straight line projected from the first feature point 91A in the line-of-sight direction 97A of the viewpoint of one captured image 110 intersects with a straight line projected from the second feature point 91B in the line-of-sight direction 97B of the viewpoint of the captured image 120 is calculated. The calculated point is taken as the target point 75. The target point 75 corresponds to the position in the three-dimensional model corresponding to the positions selected by the user in the captured images 110 and 120. That is, if feature points corresponding to each other are selected in two captured images with different viewpoints, based on the principle of triangulation, one point on the three-dimensional model corresponding to these two feature points can be uniquely specified as the target point.

[0066] The target point calculation unit 50 stores the position information of the target point calculated by triangulation in the position storage unit 40. When the position information of the target point is stored in the position storage unit 40, the target point display control unit 23 superimposes and displays the target point on the three-dimensional image 70 according to the position information.

[0067] FIG. 10 shows an example in which the target point 75 is displayed on the three-dimensional image 70. The target point 75 is displayed at the three-dimensional position corresponding to the feature points 91A and 91B selected in the two captured images 110 and 120.

[0068] When the image display control unit 22 receives an instruction from the user to turn off the display of the captured images 110 and 120 in the display state of FIG. 10, the image display control unit 22 may turn off the display of the captured images 110 and 120.

[0069] FIG. 11 shows an example in which the display of the captured images 110 and 120 is turned off in the display state of FIG. 10.

[0070] FIG. 12 is a flowchart of an example of the operation of the information processing apparatus 101 according to the first embodiment. The model display control unit 21 reads the three-dimensional model information from the information storage unit 10 and displays the three-dimensional model (S101). The viewpoint for viewing the three-dimensional model may be changed based on operation information from the user. The operation instruction acquisition unit 31 determines whether it has received instruction information to display a captured image from the user (S102). If it has not received the instruction information, the process returns to step S101. If it has received the instruction information, the image selection unit 60 selects a captured image corresponding to the displayed three-dimensional model from the information storage unit 10 (S103). The image display control unit 22 reduces and displays the two selected captured images on the display unit 301 (S104). As an example, the image selection unit 60 selects the captured images of the two viewpoints closest to the viewpoint of the displayed three-dimensional model. The user may instruct to change one or both of the two captured images displayed on the display unit 301 to another captured image. In this case, the image selection unit 60 selects another captured image according to the instruction information of the user. Alternatively, all or a certain number of captured images corresponding to the displayed three-dimensional model may be displayed, and the user may select two captured images to be used from these captured images. As an example, the user selects a captured image in which a desired position is clearly shown.

[0071] The operation instruction acquisition unit 31 determines whether it has received instruction information to turn off the display of the captured image from the user (S105). If it has received the off instruction information, the image display control unit 22 turns off the display of the two captured images and returns to step S101. If it has not received the off instruction information, the process proceeds to step S106.

[0072] The operation instruction acquisition unit 31 determines whether it has received from the user instruction information to select or enlarge one (working target image) of the two displayed captured images (S106). If it has not received the instruction information to select the working target image, the process returns to step S105. If it has received the instruction information, the image display control unit 22 enlarges and displays the captured image specified by the instruction information (S107).

[0073] The operation instruction acquisition unit 31 determines whether it has received reduction instruction information for the enlarged captured image from the user (S108). If it has received the reduction instruction information, it returns to step S104 and causes the image display control unit 22 to reduce the enlarged captured image to its original state. If it has not received the reduction instruction information, the position acquisition unit 32 determines whether information for the user to select a feature point (first position) has been input (S109). For example, the user selects a feature point by moving a cursor to a desired position in the enlarged captured image and clicking or tapping. If information for selecting a feature point has not been input, it returns to step S108. If information for selecting a feature point has been input, the position acquisition unit 32 stores the position information (coordinates) of the selected feature point in the position storage unit 40 (S110). The image display control unit 22 reads out the position information of the feature point from the position storage unit 40 and displays the feature point on the captured image according to the position information. The feature point is displayed as a mark with an arbitrary shape, arbitrary color, or arbitrary pattern as an example.

[0074] It is determined whether the feature point stored in step S110 is the first feature point or the second feature point (S111). In the case of the first feature point, that is, when the second feature point has not been selected yet, it returns to step S108. In the case of the second feature point, the target point calculation unit 50 calculates a target point (third position) by performing triangulation operations based on the two selected feature points and information (position, orientation, etc.) regarding the viewpoints corresponding to the two captured images in which the two feature points were selected (S112). The target point calculation unit 50 stores the position information of the target point in the position storage unit 40. The target point display control unit 23 reads out the position information of the target point from the position storage unit 40 and superimposes and displays the target point on the 3D model according to the position information (S113). The target point is displayed as a mark with an arbitrary shape, arbitrary color, or arbitrary pattern as an example.

[0075] Through the above processing, the position obtained by triangulation from two feature points selected by the user is acquired as the target point. Then, the acquired target point is superimposed and displayed on the 3D model. Therefore, the user can easily specify an arbitrary position (target point) in the 3D model by selecting the feature points corresponding to each other in the subject included in the two captured images.

[0076] As described above, according to the first embodiment, if the feature points corresponding to each other are selected from two captured images with different viewpoints, the point (position) on the 3D model corresponding to these feature points can be uniquely specified based on the principle of triangulation. Therefore, the user can easily specify an arbitrary position (target point) in the 3D model by selecting the feature points corresponding to each other in the subject included in the two captured images. In particular, even if a part of the 3D model is missing or hidden, it becomes possible to accurately specify the position in that part. Hereinafter, specific examples will be shown.

[0077] FIG. 13 shows an example in which a part (head) of the 3D model is missing. Even if a part of the 3D model is missing in this way, if two captured images in which the head is clearly shown are selected and the corresponding feature points are selected in each of them, the target points included in the missing part can be accurately specified.

[0078] FIG. 14 shows an example in which the target points are specified in the missing part of the 3D model.

[0079] In this way, according to the first embodiment, it becomes possible to specify an arbitrary position included in the 3D model with high accuracy.

[0080] [Modification Example 1] In this embodiment, one target point is calculated. However, two or more target points may be calculated in the same 3D model, and these target points may be superimposed and displayed on the 3D model. In this case, the operation of the flowchart in FIG. 11 may be repeated as many times as the number of target points to be calculated.

[0081] Figure 15 shows an example in which two target points are calculated and the calculated target points are superimposed and displayed on a 3D model. In addition to the target point 75 on the head of the subject, a target point 76 is displayed on the shoulder portion of the subject. A line segment 77 connecting the target point 75 and the target point 76 is drawn by the model display control unit 21 or the target point display control unit 23. The line segment 77 may be drawn based on the operation information of the user using the operation unit 201. The user can evaluate the pitching form, for example, by analyzing the angle or distance of the line segment 77. The evaluation of pitching may be performed by the user or by executing an application that evaluates pitching. In this case, the application may perform the evaluation based on a neural network that takes the angle or distance of the line segment as input and outputs an evaluation value of pitching.

[0082] [Modification Example 2] In Modification Example 1, two target points were calculated for the 3D image in the same frame, but two target points may be calculated for the 3D images in different frames. For example, target points may be calculated for the same position of the subject in the 3D images of each frame, and the trajectory of the movement of the target points between the frames may be analyzed. Hereinafter, a specific example will be described with reference to FIGS. 16 and 17.

[0083] Figure 16 shows an example in which a 3D image 131 corresponding to a certain frame is displayed from a certain user viewpoint. In FIG. 16, the 3D model 71 is viewed from a viewpoint different from the example of the present embodiment described above (see FIG. 2 etc.). The target point 115 is calculated by the processing of the present embodiment described above, and the target point 115 (position information of the third position) is superimposed and displayed on the 3D model 71.

[0084] FIG. 17 shows a display example of a three-dimensional image 132 including a three-dimensional model after a plurality of frames from the display state of FIG. 16. Note that the user's viewpoint has moved slightly due to the user's operation with respect to FIG. 16. A target point 116 is calculated and displayed at the same position corresponding to the target point 115 of the subject in FIG. 16. Also, the first calculated target point 115 is also displayed. That is, the target point display control unit 23 displays the second calculated target point while fixing the display of the first calculated target point. Further, a line segment 117 connecting the target point 115 and the target point 116 is drawn by the model display control unit 21 or the target point display control unit 23. The line segment 117 represents the movement locus of a specific position in the head of the subject during pitching. The line segment 117 may be drawn based on the user operation information using the operation unit 201. By analyzing the angle or distance of the line segment 117, etc., the pitching form of the subject can be evaluated. The evaluation of pitching may be performed by the user or by executing an application that performs the evaluation of pitching. In this case, the application may perform the evaluation based on a neural network that takes the angle or distance of the line segment as an input and outputs an evaluation value of pitching.

[0085] (Second Embodiment) In the first embodiment, when the user selects feature points in two captured images, since the user selects by intuition or visually, there is a possibility that the feature point selected second does not correctly correspond to the feature point selected first. That is, there is a possibility that the first and second selected feature points do not point to the same location on the subject (the possibility that the positions of the feature points are slightly shifted from each other). In this embodiment, when a feature point is selected for a captured image corresponding to one viewpoint, an epipolar line of the imaging camera corresponding to the other viewpoint is calculated. The calculated epipolar line is displayed on the captured image corresponding to the other viewpoint as guide information for selecting the other feature point. The user selects a feature point from on the epipolar line. Thereby, the position corresponding to the first selected feature point can be easily selected.

[0086] FIG. 18 is a block diagram of an information processing system 100 including an information processing apparatus 101 according to the second embodiment. Elements having the same names as those in the information processing system 100 of FIG. 1 are denoted by the same reference numerals, and descriptions thereof are omitted except for the extended or modified processes. In the information processing apparatus 101 of FIG. 18, a guide information calculation unit 80 and a guide information display control unit 24 are added.

[0087] When a user selects a first feature point (a first characteristic point), the guide information calculation unit 80 reads the position information of the first feature point from the position storage unit 40. The guide information calculation unit 80 calculates an epipolar line of an imaging camera corresponding to the other viewpoint based on camera parameters (for example, positions and orientations of two viewpoints) and the first feature point. The guide information display control unit 24 superimposes and displays the calculated epipolar line on an imaging image for selecting a second feature point. The epipolar line serves as a guide for the user to select the second feature point.

[0088] FIG. 19 is an explanatory diagram of an epipolar line. Two viewpoints (imaging cameras) 143 and 144 are shown. Intersection points of a line segment connecting positions (light source positions) 145 and 146 of the imaging cameras 143 and 144 and image planes 141 and 142 of the imaging cameras 143 and 144 correspond to epipoles 147 and 148. It can be said that an epipole is an image obtained by projecting the viewpoint of one imaging camera onto the viewpoint of the other imaging camera. Let an arbitrary point in the three-dimensional space be point X, and projection points when point X is projected onto the image planes 141 and 142 of the imaging cameras 143 and 144 be x1 and x2, respectively. A plane including the three-dimensional point X and the positions 145 and 146 is called an epipolar plane 149. An intersection line of the epipolar plane 149 and the image plane 141 is an epipolar line L1 of the imaging camera 143. An intersection line of the epipolar plane and the image plane 142 is an epipolar line L2 of the imaging camera 144.

[0089] When an arbitrary point x1 on the image plane 141 is given, an epipolar line L2 on the image plane 142 is determined according to the line-of-sight direction of the imaging camera 143. Then, according to the position of the point X, a point corresponding to the point x1 is determined at any position on the epipolar line L2. This determination is called an epipolar constraint. This constraint is expressed by the following formula. [Number] m is the result of converting the point x1 from the normalized image coordinates to the image coordinate system, and m = Ax1. A is the internal parameter matrix of the imaging camera 143. m' is the result of converting the point x2 from the normalized image coordinates to the image coordinate system, and m' = A'x2. A' is the internal parameter matrix of the imaging camera 144. F is called the fundamental matrix. F can be calculated using a method such as the eight-point algorithm from the camera parameters. Therefore, if the point m on the image of the imaging camera 143 is determined, the corresponding point is determined on one of the epipolar lines of the imaging camera 144.

[0090] In this embodiment, based on the epipolar constraint, when a feature point is selected for one of the two selected imaging images, an epipolar line is calculated for the imaging camera of the other imaging image. The calculated epipolar line is superimposed and displayed on the other imaging image. Since the point corresponding to the first feature point exists on the epipolar line, the user can easily select the feature point corresponding to one feature point by selecting the second feature point from the epipolar line. Only a part of the epipolar line, for example, only the part overlapping the three-dimensional model, may be displayed.

[0091] Hereinafter, a specific example of the second embodiment will be described. Assume that the user selects the feature point 91A at the position shown in FIG. 6 described above for the first imaging image of the two imaging images. The guide information calculation unit 80 calculates an epipolar line for the second imaging image based on the position of the feature point 91A and the camera parameters of the imaging cameras that captured the two selected imaging images. The calculated epipolar line is superimposed and displayed on the second imaging image.

[0092] FIG. 20 shows an example in which the epipolar line 151 is superimposed and displayed on the second imaging image. The user may select a point corresponding to the first feature point from the epipolar line 151 as the second feature point. Therefore, the user can select the second feature point with high accuracy.

[0093] FIG. 21 is a flowchart of an example of the operation of the information processing apparatus 101 according to the second embodiment. The description will focus on the differences from the flowchart of FIG. 12 in the first embodiment. The same steps as those in the flowchart of FIG. 12 are denoted by the same reference numerals, and the description thereof will be omitted as appropriate.

[0094] When it is determined in step S111 that the first feature point has been saved, the process proceeds to step S114. In step S114, based on the position information of the first feature point and the camera parameters of the viewpoints (imaging cameras) of the two selected captured images, the epipolar line calculator 80 calculates an epipolar line as guide information for selecting the second feature point. The guide information display controller 24 superimposes and displays the calculated epipolar line on a second captured image (displayed in a reduced size) different from the captured image in which the first feature point was selected. Thereafter, the process returns to step S108. When the user instructs to reduce the display of the first captured image (YES in S108), and by instructing to enlarge the second captured image (after NO in S105 and YES in S106 in S107), the user can select the second feature point while the epipolar line is displayed on the enlarged captured image.

[0095] According to the second embodiment, based on the first feature point and the camera parameters, an epipolar line for the captured image for selecting the second feature point is calculated, and the epipolar line is superimposed on the captured image as guide information. Since the user only needs to select the second feature point from the epipolar line, the feature point corresponding to the first feature point can be easily and accurately selected.

[0096] [Modification Example] The target point display controller 23 may move the target point superimposed on the three-dimensional model based on the user's instruction information. For example, when the calculated target point is deviated from the position desired by the user and the user wants to finely adjust the position of the target point, it is conceivable to move the target point in this way.

[0097] When the user wants to move the target point, the user selects the target point displayed on the screen with a cursor or a touch operation or the like, and inputs instruction information for moving the target point from the operation unit 201. According to the input instruction information, the target point display control unit 23 moves the target point displayed on the screen.

[0098] At this time, the two feature points (the first feature point and the second feature point) that are the basis for calculating the target point are also moved in accordance with the movement of the target point. Specifically, the target point calculation unit 50 calculates two feature points that satisfy the epipolar constraint for the moved target point, and moves the original feature points to the positions of the calculated feature points. Thereby, the user can learn which position should be selected as the feature point.

[0099] Figs. 22 to 24 show specific examples of this modification. Fig. 22 shows a state in which the target point 75 is moved in the direction 161 according to the instruction information of the user from the state of Fig. 11. Fig. 23 shows a state in which the feature point 91A selected in Fig. 6 is moved in the direction 162 in accordance with the movement of the target point 75. Fig. 24 shows a state in which the feature point 91B selected in Fig. 7 is moved in the direction 163 in accordance with the movement of the target point 75. The moved target point 75, the moved feature points 91A and 91B satisfy the epipolar constraint.

[0100] (Hardware Configuration) Fig. 25 shows an example of the hardware configuration of the information processing apparatus 101 of Fig. 1 or Fig. 18. The information processing apparatus 101 of Fig. 1 or Fig. 18 is configured by a computer apparatus 400. The computer apparatus 400 includes a CPU 401, an input interface 402, a display device 403, a communication device 404, a main memory device 405, and an external memory device 406, and these are mutually connected by a bus 407. The computer apparatus 400 is configured as, for example, a smartphone, a tablet, a desktop PC (Personal Computer), or a notebook PC.

[0101] The CPU (Central Processing Unit) 401 executes an information processing program, which is a computer program, on the main memory device 405. The information processing program is a program that realizes each of the above-described functional configurations of the information processing apparatus 101. The information processing program may be realized not by a single program but by a combination of a plurality of programs and scripts. By executing the information processing program, the CPU 401 realizes each functional configuration.

[0102] The input interface 402 is a circuit for inputting operation signals from input devices such as a keyboard, a mouse, and a touch panel into the information processing apparatus 101. The input interface 402 may include an imaging device such as a camera, a TOF (Time Of Flight) sensor, a LiDAR (Light Detection and Ranging), etc.

[0103] The display device 403 displays data output from the information processing apparatus 101. The display device 403 is, for example, an LCD (Liquid Crystal Display), an organic electroluminescence display, a CRT (Cathode Ray Tube), or a PDP (Plasma Display Panel), but is not limited thereto. The data output from the computer device 400 can be displayed on this display device 403.

[0104] The communication device 404 is a circuit for the information processing apparatus 101 to communicate with an external device wirelessly or by wire. Data can be input from an external device via the communication device 404. The data input from the external device can be stored in the main memory device 405 or the external storage device 406.

[0105] The main memory device 405 stores an information processing program, data necessary for the execution of the information processing program, and data generated by the execution of the information processing program, etc. The information processing program is deployed and executed on the main memory device 405. The main memory device 405 is, for example, a RAM, a DRAM, an SRAM, but is not limited thereto. The information storage unit 10 or the position storage unit 40 may be constructed on the main memory device 405.

[0106] The external storage device 406 stores an information processing program, data necessary for the execution of the information processing program, and data generated by the execution of the information processing program. These information processing programs and data are read into the main storage device 405 when the information processing program is executed. The external storage device 406 is, for example, a hard disk, an optical disk, a flash memory, and a magnetic tape, but is not limited thereto. The information storage unit 10 or the position storage unit 40 may be constructed on the external storage device 406.

[0107] Note that the information processing program may be pre-installed in the computer device 400, or may be stored in a storage medium such as a CD-ROM. Also, the information processing program may be uploaded on the Internet.

[0108] Also, the information processing device 101 may be constituted by a single computer device 400, or may be configured as a system including a plurality of computer devices 400 connected to each other.

[0109] (Third Embodiment) FIG. 26 shows an example of a free viewpoint video transmission system according to the third embodiment. The free viewpoint video transmission system includes an encoding system 510, a decoding system 520, and an information processing system 100. The information processing system 100 is the information processing system of FIG. 1 or FIG. 18.

[0110] The free viewpoint video transmission system of FIG. 26 performs three-dimensional modeling of a subject by the encoding system 510, and transmits data related to the three-dimensional model generated by the three-dimensional modeling to the decoding system 520 with a small amount of data. The decoding system 520 decodes the data to obtain three-dimensional model information and the like, and stores the three-dimensional model information and the like in the information storage unit 10 of the information processing system 100.

[0111] The symbolization system 510 and the decoding system 520 are connected via a communication network 530. The communication network 530 is a wired, wireless, or hybrid wired and wireless network. The communication network 530 can be a local area network (LAN) or a wide area network (WAN) such as the Internet. Also, the communication network 530 can be a network of any standard or protocol. For example, the communication network 530 can be a wireless LAN, a 4G or 5G mobile network, etc. Also, the communication network 530 can be a communication cable such as a serial cable.

[0112] FIG. 27 is a block diagram of the symbolization system 510. The symbolization system 510 includes an image acquisition unit 511, a three-dimensional model generation unit 512, a two-dimensional image conversion processing unit 513, a symbolization unit 514, and a transmission unit 515. The symbolization system 510 can be realized with a hardware configuration similar to that of FIG. 25.

[0113] The image acquisition unit 511 captures a subject from a plurality of viewpoints and acquires a plurality of captured images.

[0114] FIG. 28 shows a configuration example of the image acquisition unit 511. The image acquisition unit 511 includes a data acquisition unit 511A and a plurality of imaging cameras 511-1 to 511-N corresponding to a plurality of viewpoints. The imaging cameras 511-1 to 511-N are arranged at positions surrounding the subject 511B. The imaging cameras 511-1 to 511-N are arranged in directions facing the subject 511B. Camera calibration has been performed in advance on the imaging cameras 511-1 to 511-N, and camera parameters have been acquired. The camera parameters include internal parameters and external parameters as described above. The imaging cameras 511-1 to 511-N capture the subject 511B from their respective positions and acquire N captured images (RGB images). Some of the imaging cameras can be stereo cameras. Sensors other than the imaging cameras, such as TOF sensors, LIDAR, etc., can be arranged.

[0115] The 3D model generation unit 512 generates 3D model information representing a 3D model of the subject based on a plurality of captured images acquired by the image acquisition unit 511. The 3D model generation unit 512 includes a calibration unit 551, a frame synchronization unit 552, a background difference generation unit 553, a VH processing unit 554, a mesh creation unit 555, and a texture mapping unit 556.

[0116] The calibration unit 551 corrects the captured images of a plurality of viewpoints using the internal parameters of the plurality of imaging cameras.

[0117] The frame synchronization unit 552 designates one of the plurality of imaging cameras as a reference camera and the others as reference cameras. The frames of the captured images of the reference cameras are synchronized with the frames of the captured images of the reference camera.

[0118] The background difference generation unit 553 generates a silhouette image of the subject by background difference processing using the difference between the foreground image (image of the subject) and the background image for a plurality of captured images. As an example, the silhouette image is represented by binarizing the silhouette indicating the range in which the subject is included in the captured image.

[0119] The VH processing unit 554 performs modeling of the subject using a plurality of silhouette images and camera parameters. For modeling, the Visual Hull method or the like can be used. That is, each silhouette image is back-projected into the original three-dimensional space, and the intersection of the respective viewing volumes is obtained as the Visual Hull.

[0120] The mesh creation unit 555 creates a plurality of meshes by applying the marching cube method or the like to the voxel data of the Visual Hull. The three-dimensional positions of each point (Vertex) constituting the mesh and the geometric information (Geometry) indicating the connection (Polygon) of each point are generated as three-dimensional shape data (polygon model). The surface of the polygon may be smoothed by performing smoothing processing on the polygon model.

[0121] The texture mapping unit 556 performs texture mapping that overlays an image of a mesh on each mesh of the three-dimensional shape data. The three-dimensional shape data after texture mapping is used as three-dimensional model information representing a three-dimensional model of the subject.

[0122] The two-dimensional image conversion processing unit 513 converts the three-dimensional model information into a two-dimensional image. Specifically, based on the camera parameters of a plurality of viewpoints, for each viewpoint, a perspective projection of the three-dimensional model represented by the three-dimensional model information is performed. As a result, a plurality of two-dimensional images in which the three-dimensional model is perspective-projected from each viewpoint are obtained. Further, based on the camera parameters of the plurality of viewpoints, depth information is obtained for each of the plurality of two-dimensional images from these plurality of two-dimensional images, and the obtained depth information is associated with each of the two-dimensional images. By converting the three-dimensional model information into a two-dimensional image in this way, the data amount can be reduced compared to the case of transmitting the three-dimensional model information as it is.

[0123] The transmission unit 515 encodes the transmission data including the plurality of two-dimensional images and depth information, the camera parameters of the plurality of viewpoints, and the captured images of the plurality of viewpoints, and transmits the encoded transmission data to the decoding system 520. A form of transmitting all or part of the transmission data without encoding is not excluded. The transmission data may include other information. As an encoding method, for example, two-dimensional compression techniques such as 3D MVC (Multiview Video Coding), MVC, and AVC (Advanced Video Coding) can be used. By encoding the transmission data, the amount of data to be transmitted can be reduced. Note that a configuration in which the three-dimensional model information is directly encoded and transmitted to the decoding system 520 or the information processing system 100 is also possible.

[0124] FIG. 29 is a block diagram of the decoding system 520. The decoding system 520 includes a receiving unit 521, a decoding unit 522, a three-dimensional data conversion processing unit 523, and an output unit 524. The decoding system 520 can be realized with the same hardware configuration as that in FIG. 25. The receiving unit 521 receives the transmission data transmitted from the encoding system 510 and supplies it to the decoding unit 522.

[0125] The decoding unit 522 decodes the transmission data to obtain a plurality of two-dimensional images and depth information, a plurality of captured images, camera parameters of a plurality of viewpoints, and the like. As a decoding method, the same two-dimensional compression technology as that on the transmission side can be used.

[0126] The three-dimensional data conversion processing unit 523 performs a conversion process of converting a plurality of two-dimensional images into three-dimensional model information representing a three-dimensional model of the subject. For example, modeling is performed by Visual Hull or the like using a plurality of two-dimensional images and depth information, and camera parameters.

[0127] The output unit 524 provides the three-dimensional model information, a plurality of captured images, camera parameters, and the like to the information processing device 101. The information processing device 101 of the information processing system 100 stores the three-dimensional model information, a plurality of captured images, camera parameters, and the like in the information storage unit 10.

[0128] Note that the above-described embodiments are merely examples for embodying the present disclosure, and the present disclosure can be implemented in various other forms. For example, various modifications, substitutions, omissions, or combinations thereof are possible without departing from the gist of the present disclosure. Forms in which such modifications, substitutions, omissions, etc. are made are also included in the scope of the present disclosure, and are also included in the invention described in the claims and the equivalent scope thereof.

[0129] Also, the effects of the present disclosure described in this specification are merely examples, and there may be other effects.

[0130] Note that the present disclosure can also adopt the following configuration. [Item 1] A first display control unit that displays a three-dimensional model of the subject on a display device based on a plurality of captured images obtained by capturing the subject from a plurality of viewpoints; A second display control unit that displays a first captured image of a first viewpoint and a second captured image of a second viewpoint among the plurality of captured images on the display device; A position acquisition unit that acquires position information of a first position included in the subject in the first captured image and acquires position information of a second position included in the subject in the second captured image; and a position calculation unit that calculates a third position included in the three-dimensional model based on information regarding the first viewpoint and the second viewpoint, the position information of the first position, and the position information of the second position; A third display control unit that superimposes the position information of the third position on the three-dimensional model and displays it on the display device; An information processing apparatus comprising the above. [Item 2] The position calculation unit calculates a position where a straight line obtained by projecting the first position according to the line-of-sight direction of the first viewpoint and a straight line obtained by projecting the second position according to the line-of-sight direction of the second viewpoint intersect, and the calculated position is the third position. The information processing apparatus according to Item 1. [Item 3] The position calculation unit calculates the third position based on the principle of triangulation. The information processing apparatus according to Item 1 or 2. [Item 4] An image selection unit that selects captured images of two viewpoints that are closest to the viewpoint of a user who views the three-dimensional model from the plurality of captured images is provided, The second display control unit displays the first captured image and the second captured image selected by the image selection unit. The information processing apparatus according to any one of Items 1 to 3. [Item 5] An image selection unit that selects the first captured image and the second captured image based on selection information for designating the first captured image and the second captured image is provided, The second display control unit displays the first captured image and the second captured image selected by the image selection unit. The information processing apparatus according to any one of Items 1 to 4. [Item 6] Comprising an instruction information acquisition unit that acquires instruction information for instructing at least one of enlargement or reduction of the first captured image and the second captured image. Based on the instruction information, the second display control unit enlarges or reduces at least one of the first captured image and the second captured image. The information processing apparatus according to any one of Items 1 to 5. [Item 7] The position acquisition unit acquires the position information of the first position from the user's operation device in a state where the first captured image is enlarged, and acquires the position information of the second position from the operation device in a state where the second captured image is enlarged. The information processing apparatus according to Item 6. [Item 8] Comprising an instruction information acquisition unit that acquires instruction information for instructing on or off of the display of at least one of the first captured image and the second captured image. Based on the instruction information, the second display control unit switches on or off the display of at least one of the first captured image and the second captured image. The information processing apparatus according to any one of Items 1 to 7. [Item 9] The second display control unit superimposes the position information of the first position on the first captured image and displays it on the display device. The second display control unit superimposes the position information of the second position on the second captured image and displays it on the display device. The information processing apparatus according to any one of Items 1 to 8. [Item 10] A guide information generation unit that generates guide information including candidates for the second position based on the information regarding the first viewpoint and the second viewpoint and the position information of the first position. A fourth display control unit that superimposes the guide information on the second captured image and displays it on the display device. The information processing apparatus according to any one of Items 1 to 9, comprising the above. [Item 11] The guide information is an epipolar line. The information processing apparatus according to item 10. [Item 12] Comprising an instruction information acquisition unit that acquires instruction information for instructing movement of the position information of the third position, and the third display control unit moves the position information of the third position based on the instruction information. The second display control unit moves the position information of the first position and the position information of the second position in response to the movement of the position information of the third position. The information processing apparatus according to item 9. [Item 13] Display a three-dimensional model of the subject based on a plurality of captured images of the subject captured from a plurality of viewpoints on a display device. Display the first captured image of the first viewpoint and the second captured image of the second viewpoint among the plurality of captured images on the display device. Acquire the position information of the first position included in the subject in the first captured image, and acquire the position information of the second position included in the subject in the second captured image. Calculate a third position included in the three-dimensional model based on information regarding the first viewpoint and the second viewpoint, the position information of the first position, and the position information of the second position. Superimpose the position information of the third position on the three-dimensional model and display it on the display device. Information processing method. [Item 14] A step of displaying a three-dimensional model of the subject based on a plurality of captured images of the subject captured from a plurality of viewpoints on a display device. A step of displaying the first captured image of the first viewpoint and the second captured image of the second viewpoint among the plurality of captured images on the display device. A step of acquiring the position information of the first position included in the subject in the first captured image, and acquiring the position information of the second position included in the subject in the second captured image. A step of calculating a third position included in the three-dimensional model based on information regarding the first viewpoint and the second viewpoint, the position information of the first position, and the position information of the second position. A step of superimposing the position information of the third position on the three-dimensional model and displaying it on the display device; A computer program for causing a computer to execute.

Explanation of Signs

[0131] 101: Information processing device, 10: Information storage unit, 20: Display control unit, 30: Interaction detection unit, 40: Position storage unit, 50: Target point calculation unit (triangulation operation unit), 60: Image selection unit, 110, 120: Captured images, 21: Model display control unit, 22: Image display control unit, 23: Target point display control unit, 24: Guide information display control unit, 31: Operation instruction acquisition unit, 32: Position acquisition unit, 70: Three-dimensional image, 71: Three-dimensional model, 71A: Glove, 71B: Cap, 72: Mount, 73: Pitcher's plate, 75, 76: Target points, 77: Line segment 80: Guide information calculation unit, 91: Baseball player, 91A: First feature point, 91B: Second feature point, 92A, 92B, 92C, 92D: Imaging cameras, 115, 116: Target points, 141, 142: Image planes, 143, 144: Two viewpoints (imaging cameras), 145, 146: Positions of imaging cameras (light source positions), 147, 148: Epipolar holes, x1, x2: Positions where point X is projected (two-dimensional coordinates), 149: Epipolar plane, 151: Epipolar line, 400: Computer device, 401: CPU, 402: Input interface, 403: Display device, 404: Communication device, 405: Main memory device, 406: External memory device, 407: Bus, 510: Encoding system, 520: Decoding system, 100: Information processing system, 511: Image acquisition unit, 512: 3D model generation unit, 513: 2D image conversion processing unit, 514: Encoding unit, 515: Transmission unit, 511A: Data acquisition unit, 551: Calibration unit, 552: Frame synchronization unit, 553: Background difference generation unit, 554: VH processing unit, 555: Mesh creation unit, 556: Texture mapping unit, 521: Reception unit, 522: Decoding unit 522, 523: 3D data conversion processing unit, 524: Output unit

Claims

1. A first display control unit that displays a three-dimensional model of the subject on a display device based on a plurality of captured images of the subject captured from a plurality of viewpoints; A second display control unit that displays, on the display device, a first captured image of a first viewpoint and a second captured image of a second viewpoint among the plurality of captured images; A position acquisition unit that acquires position information of a first position included in the subject in the first captured image and acquires position information of a second position corresponding to the same position as the first position included in the subject in the second captured image; A position calculation unit that calculates a third position included in the three-dimensional model based on information regarding the first viewpoint and the second viewpoint, position information of the first position, and position information of the second position; A third display control unit that superimposes position information of the third position on the three-dimensional model and displays it on the display device; An information processing apparatus comprising the above.

2. The position calculation unit calculates a position where a straight line obtained by projecting the first position according to the line-of-sight direction of the first viewpoint and a straight line obtained by projecting the second position according to the line-of-sight direction of the second viewpoint intersect, and the calculated position is the third position. The information processing apparatus according to Claim 1.

3. The position calculation unit calculates the third position based on the principle of triangulation. The information processing apparatus according to Claim 2.

4. An image selection unit that selects captured images of two viewpoints that are closest to the viewpoint of a user who views the three-dimensional model from the plurality of captured images is provided, The second display control unit displays the first captured image and the second captured image selected by the image selection unit. The information processing apparatus according to Claim 1.

5. An image selection unit that selects the first captured image and the second captured image based on selection information for designating the first captured image and the second captured image is provided, The second display control unit displays the first captured image and the second captured image selected by the image selection unit. The information processing apparatus according to Claim 1.

6. An instruction information acquisition unit that acquires instruction information for instructing enlargement or reduction of at least one of the first captured image and the second captured image is provided, The second display control unit enlarges or reduces at least one of the first captured image and the second captured image based on the instruction information. The information processing apparatus according to Claim 1.

7. The position acquisition unit acquires the position information of the first position from the user's operating device in a state where the first captured image is enlarged, and acquires the position information of the second position from the operating device in a state where the second captured image is enlarged. The information processing apparatus according to claim 6.

8. Comprising an instruction information acquisition unit that acquires instruction information for instructing on or off of the display of at least one of the first captured image and the second captured image. The second display control unit switches on or off the display of at least one of the first captured image and the second captured image based on the instruction information. The information processing apparatus according to claim 1.

9. The second display control unit superimposes the position information of the first position on the first captured image and displays it on the display device. The second display control unit superimposes the position information of the second position on the second captured image and displays it on the display device. The information processing apparatus according to claim 1.

10. A guide information generation unit that generates guide information including candidates for the second position based on the information regarding the first viewpoint and the second viewpoint and the position information of the first position. A fourth display control unit that superimposes the guide information on the second captured image and displays it on the display device. The information processing apparatus according to claim 1, comprising the above.

11. The guide information is an epipolar line. The information processing apparatus according to claim 10.

12. Comprising an instruction information acquisition unit that acquires instruction information for instructing the movement of the position information of the third position. The third display control unit moves the position information of the third position based on the instruction information. The second display control unit moves the position information of the first position and the position information of the second position in response to the movement of the position information of the third position. The information processing apparatus according to claim 9.

13. Displays a three-dimensional model of the subject based on a plurality of captured images of the subject captured from a plurality of viewpoints on a display device. Displays a first captured image of a first viewpoint and a second captured image of a second viewpoint among the plurality of captured images on the display device. Acquires the position information of the first position included in the subject in the first captured image, and acquires the position information of the second position corresponding to the same position as the first position included in the subject in the second captured image. Calculates a third position included in the three-dimensional model based on the information regarding the first viewpoint and the second viewpoint, the position information of the first position, and the position information of the second position. Superimpose the position information of the third position on the three-dimensional model and display it on the display device Information processing method.

14. A step of displaying a three-dimensional model of the subject based on a plurality of captured images obtained by capturing the subject from a plurality of viewpoints on a display device; A step of displaying a first captured image of a first viewpoint and a second captured image of a second viewpoint among the plurality of captured images on the display device; A step of obtaining position information of a first position included in the subject in the first captured image, and obtaining position information of a second position corresponding to the same position as the first position included in the subject in the second captured image; A step of calculating a third position included in the three-dimensional model based on information regarding the first viewpoint and the second viewpoint, the position information of the first position, and the position information of the second position; A step of superimposing the position information of the third position on the three-dimensional model and displaying it on the display device; A computer program for causing a computer to execute the above steps.

Citation Information

Patent Citations

  • Device and method for three-dimensional shape input

    JP1996069547A

  • Information processing apparatus, control method of information processing apparatus and program

    JP2018195267A

  • Information processing apparatus, information processing method, and program

    JP2019106617A

  • Controlling presentation of hidden information

    US20200082629A1

  • Devices, Methods, and Graphical User Interfaces for Displaying Applications in Three-Dimensional Environments

    US20210191600A1