Image processing system

By combining SLAM technology and machine learning in a DSLR camera to infer the three-dimensional shape and depth of the photographed object, the problem of insufficient spatial information in existing technologies is solved, the system configuration is simplified and the amount of information is increased.

CN117115256BActive Publication Date: 2026-01-02RAKUTEN GROUP INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311079556.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2017-08-14
Publication Date
2026-01-02
Estimated Expiration
2037-08-14

AI Technical Summary

Technical Problem

Existing technologies cannot adequately increase the amount of information in the observed space using images captured by SLR cameras, and using depth cameras would complicate the system.

Method used

By acquiring photographic images, SLAM technology is used to calculate the three-dimensional coordinates of feature point groups, and machine learning is combined to infer the three-dimensional shape and depth of the photographed object, thus integrating observational spatial information to increase the amount of information.

Benefits of technology

This technology enables the creation of 3D shape and depth images of photographic objects by combining machine learning-derived data without the use of a depth camera, simplifying system configuration and increasing the amount of information in the observed space.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117115256B_ABST
    Figure CN117115256B_ABST
Patent Text Reader

Abstract

The present invention relates to an image processing system. The present invention aims at simplifying the configuration for increasing the amount of information of an observation space. An image processing system (10) acquires a photographic image captured by a photographic mechanism (18) capable of moving in a real space, by a photographic image acquisition mechanism (101). An observation space information acquisition mechanism (102) acquires observation space information including three-dimensional coordinates of a feature point group in an observation space, based on a change in the position of the feature point group in the photographic image. A mechanical learning mechanism (103) acquires additional information related to a feature of a photographic object shown in the photographic image, based on mechanical learning data related to the feature of the object. An integration mechanism (104) integrates the observation space information and the additional information.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Related information of divisional application

[0002] This application is a divisional application. The parent application of this divisional application is the invention patent application with the application date of August 14, 2017, the application number of 201780093930.8, and the invention name of "Image processing system, image processing method, and program". TECHNICAL FIELD

[0003] The present application relates to an image processing system, an image processing method, and a program. BACKGROUND

[0004] In recent years, a technology of analyzing a photographic image captured by a camera and reproducing a situation of a real space in an observation space is being researched. For example, in Non-Patent Literature 1, a technology called SLAM (Simultaneous Localization And Mapping) is described, which generates a 3D map including three-dimensional coordinates of a feature point group in an observation space based on a change in position of the feature point group in a photographic image of an RGB camera (so-called single-lens reflex camera) that does not include a depth camera. In addition, for example, in Non-Patent Literature 2, a technology of generating a 3D map based on a photographic image of an RGB-D camera including an RGB camera and a depth camera is described.

[0005] PRIOR ART DOCUMENT

[0006] NON-PATENT LITERATURE

[0007] Non-Patent Literature 1: Andrew J. Davison, "Real-Time Simultaneous Localization and Mapping with a Single Camera", Proceedings of the 9th IEEE International Conference on Computer Vision Volume 2, 2003, pp. 1403-1410

[0008] Non-Patent Literature 2: Real-time 3D visual SLAM with a hand-held camera (N. Engelhard, F. Endres, J. Hess, J. Sturm, W. Burgard), In Proc. of the RGB-D Workshop on 3D Perception in Robotics at the European Robotics Forum, 2011 SUMMARY

[0009] [Problems to be Solved by the Invention]

[0010] However, in the technology of Non-Patent Literature 1, only the three-dimensional coordinates of the feature point group extracted from the photographic image are represented on the 3D map, and the information amount of the observation space cannot be sufficiently increased. In this regard, in the technology of Non-Patent Literature 2, the depth of the surface of the photographic object can be measured using the depth camera function, and the three-dimensional shape of the photographic object can be represented, so the information amount of the observation space can be increased, but the depth camera needs to be prepared, leading to a complicated configuration.

[0011] The present application was made in view of the above problems, and aims at simplifying the configuration for increasing the information amount of the observation space.

[0012] [Technical Means to Solve the Problems]

[0013] To solve the above problems, the image processing system of the present application is characterized by comprising: a photographic image acquisition mechanism that acquires a photographic image captured by a photographic mechanism that can move in a real space; an observation space information acquisition mechanism that acquires observation space information including three-dimensional coordinates of a feature point group in an observation space, based on a positional change of the feature point group in the photographic image; a mechanical learning mechanism that acquires additional information related to a feature of a photographic object shown in the photographic image, based on mechanical learning data related to the feature of the object; and a consolidation mechanism that consolidates the observation space information and the additional information.

[0014] The image processing method of the present application is characterized by comprising the steps of: a photographic image acquisition step of acquiring a photographic image captured by a photographic mechanism that can move in a real space; an observation space information acquisition step of acquiring observation space information including three-dimensional coordinates of a feature point group in an observation space, based on a positional change of the feature point group in the photographic image; a mechanical learning step of acquiring additional information related to a feature of a photographic object shown in the photographic image, based on mechanical learning data related to the feature of the object; and a consolidation step of consolidating the observation space information and the additional information.

[0015] The program of the present application causes a computer to function as: a photographic image acquisition mechanism that acquires a photographic image captured by a photographic mechanism that can move in a real space; an observation space information acquisition mechanism that acquires observation space information including three-dimensional coordinates of a feature point group in an observation space, based on a positional change of the feature point group in the photographic image; a mechanical learning mechanism that acquires additional information related to a feature of a photographic object shown in the photographic image, based on mechanical learning data related to the feature of the object; and a consolidation mechanism that consolidates the observation space information and the additional information.

[0016] In one aspect of the present application, the additional information is two-dimensional feature quantity information obtained by associating a position of the photographed object in the photographed image and a feature quantity related to the photographed object, the observation space information acquisition mechanism estimates a position of the photographing mechanism based on a change in the position of the feature point group, sets an observation viewpoint in the observation space based on the estimation result, and the integration mechanism performs processing based on a comparison result of the two-dimensional observation information representing a situation in which the observation space is observed from the observation viewpoint and the two-dimensional feature quantity information.

[0017] In one aspect of the present application, the feature quantity is a depth of the photographed object estimated based on the machine learning data, in the two-dimensional observation information, a position of the feature point group in a two-dimensional space is associated with a depth of the feature point group in the observation space, and the integration mechanism sets a mesh of the photographed object in the observation space based on the two-dimensional feature quantity information and changes a scale of the mesh based on a comparison result of the two-dimensional observation information and the two-dimensional feature quantity information.

[0018] In one aspect of the present application, the integration mechanism changes the mesh locally after changing a scale of the mesh based on a comparison result of the two-dimensional observation information and the two-dimensional feature quantity information.

[0019] In one aspect of the present application, the additional information is information related to a three-dimensional shape of the photographed object estimated based on the machine learning data.

[0020] In one aspect of the present application, the additional information is information related to a mesh of the photographed object.

[0021] In one aspect of the present application, the integration mechanism sets the mesh in the observation space based on the additional information and changes the mesh based on the observation space information.

[0022] In one aspect of the present application, the integration mechanism changes a mesh portion around the mesh portion corresponding to a three-dimensional coordinate of the feature point group indicated by the observation space information in the mesh after changing the mesh portion.

[0023] In one aspect of the present application, the observation space information acquisition mechanism estimates a position of the photographing mechanism based on a change in the position of the feature point group, sets an observation viewpoint in the observation space based on the estimation result, and the integration mechanism changes each mesh portion with respect to a direction of the mesh portion with respect to the observation viewpoint.

[0024] In one aspect of the present application, the additional information is information related to a normal line of the photographed object.

[0025] In one aspect of the present application, the additional information is information related to a classification of the photographed object.

[0026] In one aspect of the present application, the photographing mechanism photographs the real space based on a predetermined frame rate, and the observation space information acquisition mechanism and the mechanical learning mechanism perform processing based on the photographed images taken at the same frame.

[0027] [Effects of the Invention]

[0028] According to the present application, it is possible to simplify the configuration for increasing the amount of information of the observation space. BRIEF DESCRIPTION OF DRAWINGS

[0029] Figure 1 is a diagram showing a hardware configuration of an image processing apparatus.

[0030] Figure 2 is a diagram showing a case where a photographing section photographs a real space.

[0031] Figure 3 is a diagram showing an example of a photographed image.

[0032] Figure 4 is a diagram showing an example of three-dimensional coordinates of a feature point group.

[0033] Figure 5 is a diagram showing an example of a depth image.

[0034] Figure 6 is a diagram showing an example of a normal line image generated from a photographed image.

[0035] Figure 7 is a diagram showing an example of an integrated observation space.

[0036] Figure 8 is a functional block diagram showing an example of functions implemented in an image processing apparatus.

[0037] Figure 9 is a diagram showing an example of an observation space image.

[0038] Figure 10 is a diagram showing an example of processing performed by an integration section.

[0039] Figure 11 is an explanatory diagram of a process of changing a mesh by extending an ARAP technique.

[0040] Figure 12is an explanatory diagram of a process of changing a mesh by extending the ARAP method.

[0041] Figure 13 is a flowchart showing an example of a process executed in the image processing apparatus.

[0042] Figure 14 is a flowchart showing an example of a mapping process.

[0043] Figure 15 is a flowchart showing an example of a decryption process.

[0044] Figure 16 is a flowchart showing an example of an integration process.

[0045] Figure 17 is a diagram showing an example of an execution interval of each process.

[0046] Figure 18 is a diagram showing an example of a classified image.

[0047] Figure 19 is a diagram showing an example of a process executed by the integration section.

[0048] Figure 20 is a diagram showing an example of an image processing system in a variation. DETAILED DESCRIPTION

[0049] [1. Hardware configuration of image processing system]

[0050] Hereinafter, an example of an embodiment of an image processing system related to the present application will be described. In the present embodiment, a case where the image processing system is implemented by one computer will be described, but the image processing system can also be implemented by a plurality of computers, like in the following variations.

[0051] Figure 1 is a diagram showing a hardware configuration of the image processing apparatus. The image processing apparatus 10 is a computer that executes image processing, and is, for example, a mobile phone (including a smartphone), a mobile information terminal (including a tablet computer), a personal computer, or a server computer, and the like. As shown in the diagram, the image processing apparatus 10 includes a control section 11, a storage section 12, a communication section 13, an operation section 14, a display section 15, an input / output section 16, a reading section 17, and a photographing section 18. Figure 1

[0052] ​The control section 11 includes, for example, at least one microprocessor. The control section 11 performs processing in accordance with a program or data stored in the storage section 12. The storage section 12 includes a main storage section and an auxiliary storage section. The main storage section is, for example, a volatile memory such as a RAM (Random Access Memory), and the auxiliary storage section is a non-volatile memory such as a hard disk or a flash memory. The communication section 13 is a communication interface for wired or wireless communication, and performs data communication via a network. The operation section 14 is an input device for a user to perform operations, and includes, for example, a pointing device such as a touch panel or a mouse, or a keyboard. The operation section 14 transmits the contents of operations by the user to the control section 11.

[0053] The display section 15 is, for example, a liquid crystal display section or an organic EL (Electroluminescence) display section. The display section 15 displays a screen in accordance with an instruction from the control section 11. The input / output section 16 is an input / output interface, and includes, for example, a USB (Universal Serial Bus) port. The input / output section 16 is used for data communication with an external device. The reading section 17 reads an information storage medium that can be read by a computer, and includes, for example, an optical disk drive or a memory card slot. The imaging section 18 includes at least one camera that captures a still image or a moving image, and includes, for example, an image pickup element such as a CMOS (complementary metal oxide semiconductor) image sensor or a CCD (charge-coupled device) image sensor. The imaging section 18 can continuously capture a real space. The imaging section 18 can capture an image at a predetermined frame rate, or can capture an image at an unspecified frame rate without being particularly specified.

[0054] Furthermore, the program and data stored in the storage section 12 can be supplied from another computer via a network, or can be supplied from an information storage medium (for example, a USB memory, an SD card, or an optical disk) that can be read by a computer via the input / output section 16 or the reading section 17. In addition, the display section 15 and the imaging section 18 are not necessarily assembled inside the image processing apparatus 10, but can be connected via the input / output section 16 outside the image processing apparatus 10. In addition, the hardware configuration of the image processing apparatus 10 is not limited to the example described above, and various hardware configurations can be applied.

[0055] [2. Outline of processing performed by image processing apparatus]

[0056] The image processing apparatus 10 generates an observation space that reproduces a situation of a real space based on a photographed image photographed by the photographing section 18. The real space is a physical space photographed by the photographing section 18. The observation space is a virtual three-dimensional space and is a space defined inside the image processing apparatus 10. The observation space contains a point group that represents a photographed object. The photographed object is an object of the real space appearing in the photographed image and is also called a subject. In other words, the photographed object is a part of the real space appearing in the photographed image.

[0057] The point group of the observation space is information for expressing a three-dimensional shape of the photographed object in the observation space and is a vertex group that constitutes a mesh. The mesh is information also called a polygon and is a constituent element of a three-dimensional model (3D model) that represents the photographed object. The photographing section 18 can photograph an arbitrary place, but in the present embodiment, a case where the photographing section 18 photographs a situation of an indoor space is described.

[0058] Figure 2 is a view that represents a case where the photographing section 18 photographs the real space. As shown in Figure 2 , in the present embodiment, the photographing section 18 photographs the inside of a room surrounded by a plurality of surfaces (a floor, walls, and a ceiling, etc.). In Figure 2 , a bed and a painting are arranged in the real space RS. A user photographs an arbitrary place while moving the image processing apparatus 10 by hand. For example, the photographing section 18 generates a photographed image by continuously photographing the real space RS based on a prescribed frame rate.

[0059] Figure 3 is a view that represents an example of the photographed image. As shown in Figure 3 , in the photographed image G1, the walls, the floor, the bed, and the painting that are within the photographing range of the photographing section 18 are photographed as photographed objects. In addition, in the present embodiment, a screen coordinate axis (Xs axis-Ys axis) is set with the upper left of the photographed image G1 as an origin Os, and a position within the photographed image G1 is represented by a two-dimensional coordinate of a screen coordinate system.

[0060] For example, the image processing apparatus 10 extracts a feature point group from the photographed image G1 and calculates three-dimensional coordinates of the feature point group in the observation space using an SLAM technique. The feature point is a point that represents a characteristic part within an image, for example, a part that represents a contour of a photographed object or a part that represents a color change of a photographed object. The feature point group is a collection of a plurality of feature points.

[0061] Figure 4 is a view that represents an example of the three-dimensional coordinates of the feature point group. Figure 4 shown in is a feature point extracted from the photographed image G1. Hereinafter, the feature point is not distinguished from the photographed object. These feature points are collectively denoted as a feature point group P when distinguished in particular. Further, in the present embodiment, a world coordinate axis (Xw axis - Yw axis - Zw axis) is set with a prescribed position in the observation space OS as an origin Ow, and a position in the observation space OS is expressed by three-dimensional coordinates of the world coordinate system.

[0062] In the present embodiment, the image processing apparatus 10 not only calculates the three-dimensional coordinates of the feature point group P using the SLAM technique, but also estimates the position and direction of the photographing section 18 in the real space RS. The image processing apparatus 10 sets the three-dimensional coordinates of the feature point group P in the observation space OS, and sets the observation view point OV in the observation space OS in a manner corresponding to the position and direction of the photographing section 18. The observation view point OV is also referred to as a virtual camera, and is a view point in the observation space OS.

[0063] Since the feature point group P is nothing but a collection of feature points representing a part of the outline or the like of the photographed object, the density of the feature point group P is insufficient to represent the surface of the photographed object, as shown in FIG. 2. That is, the observation space OS in which the three-dimensional coordinates of the feature point group P are set is sparse point group data, for example, and does not become information of a degree that can represent the surface of the photographed object in detail. Figure 4

[0064] Therefore, the image processing apparatus 10 of the present embodiment estimates the three-dimensional shape of the photographed object using machine learning (deep learning), and integrates the estimated three-dimensional shape with the three-dimensional coordinates of the feature point group P, to increase the information amount of the observation space OS. Specifically, the image processing apparatus 10 roughly estimates the three-dimensional shape of the photographed object using machine learning, and corrects the estimated three-dimensional shape in a manner consistent with the three-dimensional coordinates of the feature point group P as measured values. For example, the image processing apparatus 10 acquires two images of a depth image and a normal image as an estimation result of the three-dimensional shape of the photographed object. Further, the estimation result can be information expressed in two dimensions, and does not necessarily have to be in the form of an image. For example, the estimation result can be data representing a combination of two-dimensional coordinates and information related to depth or normal, such as data in a table form or a tabular form.

[0065] Figure 5 is a diagram that represents an example of a depth image. The depth image G2 is the same size (the same number of pixels in the vertical and horizontal directions) as the photographed image Gl, and is an image that represents the depth of the photographed object. The depth is the distance of the photographed object, that is, the distance of the photographed section 18 from the photographed object. The pixel value of each pixel of the depth image G2 represents the depth of that pixel. That is, the pixel value of each pixel of the depth image G2 represents the distance of the photographed object appearing in that pixel from the photographed section 18. Further, the pixel value is a numerical value assigned to each pixel, and is information also referred to as color, luminance, or brightness.

[0066] ​The depth image G2 can be either a color image or a gray scale image. In the example of Fig. 2, the pixel value of the depth image G2 is schematically represented by the density of a halftone dot, indicating that the denser the halftone dot, the lower the depth (the shorter the distance), and the more transparent the halftone dot, the deeper the depth (the longer the distance). That is, when viewing the photographic object indicated by the pixel of the denser halftone dot from the photographing section 18, the photographic object is on the near side, and when viewing the photographic object indicated by the pixel of the more transparent halftone dot from the photographing section 18, the photographic object is on the far side. For example, the halftone dot of a part of a bed or the like close to the photographing section 18 is denser, and the halftone dot of a part of a wall or the like far from the photographing section 18 is more transparent. Figure 5

[0067] Figure 6 Fig. 3 is an example of a normal image G3 generated from the photographic image G1. The normal image G3 is an image of the normal of the photographic object, and is the same size as the photographic image G1 (the number of pixels in the vertical and horizontal directions is the same). The pixel value of each pixel of the normal image G3 indicates the direction of the normal (vector information) of the pixel. That is, the pixel value of each pixel of the normal image G3 indicates the direction of the normal of the photographic object photographed by the pixel.

[0068] The normal image G3 can be either a color image or a gray scale image. In the example of Fig. 3, the pixel value of the normal image G3 is schematically represented by the density of a halftone dot, indicating that the denser the halftone dot, the more the normal is directed toward the vertical direction (Zw axis direction), and the more transparent the halftone dot, the more the normal is directed toward the horizontal direction (Xw axis direction or Yw axis direction). That is, the photographic object indicated by the pixel of the denser halftone dot has a surface directed toward the vertical direction, and the photographic object indicated by the pixel of the more transparent halftone dot has a surface directed toward the horizontal direction. Figure 6

[0069] For example, the halftone dot of a part of a floor or the like having a surface directed toward the vertical direction is denser, and the halftone dot of a part of a wall or the like having a surface directed toward the horizontal direction is more transparent. Further, in the example of Fig. 3, the halftone dot is more densely represented in the Xw axis direction than in the Yw axis direction. Therefore, for example, the surface of the wall on the right side (the normal is the Xw axis direction) is more densely represented by the halftone dot than the surface of the wall on the left side (the normal is the Yw axis direction) when viewed from the photographing section 18. Figure 6

[0070] The depth image G2 and the normal image G3 are information indicating the three-dimensional shape of the photographic object, and the image processing apparatus 10 can estimate the mesh of the photographic object based on these images. However, the depth image G2 and the normal image G3 are information obtained by machine learning, and although they have a certain degree of accuracy, they are not measured values measured on site by the image processing apparatus 10, and thus are not so accurate. ​​​

[0071] Therefore, even if the mesh estimated from the depth image G2 and the normal image G3 is directly set in the observation space OS to increase the amount of information, there are cases where the scales are inconsistent or the details of the mesh are different, and it is not possible to improve the accuracy of the observation space OS. Therefore, the image processing apparatus 10 improves the accuracy of the three-dimensional shape by integrating the three-dimensional coordinates of the feature point group P, which is a measured value, with the depth image G2 and the normal image G3, and increases the amount of information of the observation space OS.

[0072] Figure 7 is a diagram showing an example of the integrated observation space OS. In Figure 7 , the set of point groups in the observation space OS is schematically shown by solid lines. As Figure 7 indicated, the density of the point groups in the observation space OS can be improved by using machine learning, and the density of the point groups is high enough to represent the surface of the photographed object. That is, the integrated observation space OS is dense point group data, and becomes an amount of information that can represent the surface of the photographed object in detail, for example.

[0073] Further, since what can be reproduced in the observation space OS is only within the photographing range of the photographing section 18, cases outside the photographing range (for example, a dead angle such as the back of the photographing section 18) cannot be reproduced. Therefore, in order to reproduce the entire room, the user moves while holding the image processing apparatus 10 and photographs the room everywhere, and the image processing apparatus 10 repeatedly performs the processing described above to reproduce the entire room.

[0074] As described above, the image processing apparatus 10 of the present embodiment integrates the three-dimensional coordinates of the feature point group P, which is a measured value, with the depth image G2 and the normal image G3 acquired by using machine learning, and can increase the amount of information of the observation space OS even without using a depth camera or the like. Hereinafter, the details of the image processing apparatus 10 will be described.

[0075] [3. Functions implemented in the image processing apparatus]

[0076] Figure 8 is a functional block diagram showing an example of the functions implemented in the image processing apparatus 10. As Figure 8 indicated, in the present embodiment, the case where the data storage section 100, the photographed image acquisition section 101, the observation space information acquisition section 102, the machine learning section 103, and the integration section 104 are implemented will be described.

[0077] [3-1. Data storage section]

[0078] The data storage section 100 is mainly implemented by the storage section 12. The data storage section 100 stores data required to generate the observation space OS that reproduces the case of the real space RS.

[0079] For example, the data storage section 100 stores machine learning data used in machine learning. The machine learning data is data related to various object features. For example, the machine learning data is data indicating appearance features of an object, and can also indicate various features such as a three-dimensional shape, a contour, a size, a color, or a pattern of an object. Further, the three-dimensional shape here refers to a surface concave-convex or a direction.

[0080] In the machine learning data, feature information related to the features of each object is stored. In addition, because even the same object has different features such as a three-dimensional shape, a size, a contour, a color, or a pattern, the machine learning data can be prepared in a manner that includes various features.

[0081] If a bed is described as an example of an object, there are various types of bed frames such as a pipe bed or a two-stage bed, and there are various types of bed three-dimensional shapes or contours. In addition, there are various types of beds such as a single bed size or a double bed size, and there are various types of bed sizes. Also, there are various types of bed colors or patterns, and therefore the feature information is stored in the machine learning data in a manner that includes known beds.

[0082] Further, even the same bed is observed differently depending on the angle, and therefore the feature information in a case where the bed is observed from various angles is stored in the machine learning data. Here, the bed is described as an example, but the same applies to objects other than beds (for example, furniture, home appliances, clothes, vehicles, groceries, and the like), and the feature information in a case where various types of objects are observed from various angles is stored in the machine learning data.

[0083] In the present embodiment, because the depth image G2 and the normal image G3 are acquired by machine learning, the depth and the normal of the object are stored as feature information. Therefore, as an example of the machine learning data, depth learning data related to the depth of the object and normal learning data related to the normal of the object are described.

[0084] For example, the depth learning data and the normal learning data are generated by capturing an object using an RGB-D camera. The RGB-D camera can measure the depth of an object arranged in the real space RS, and therefore the depth learning data is generated based on the depth information as a measured value. In addition, because the depth of the object is information that can specify a three-dimensional shape (a surface concave-convex), the normal direction of the object surface can also be acquired based on the depth information measured by the RGB-D camera. Therefore, the normal learning data is also generated based on the normal direction as a measured value.

[0085] Further, the machine learning data and the algorithm of the machine learning itself can use known data and algorithms, and for example, data and algorithms in a so-called CNN (Convolutional Neural Network) described in "Predicting Depth, Surface Normals and Semantic Labels with a Common Multi-Scale Convolutional Architecture" (http: / / www.cs.nyu.edu / ~deigen / dnl / , https: / / arxiv.org / pdf / 1411.4734v4.pdf) can be used. Further, the feature information stored in the machine learning data can be information indicating the features of the object, and is not limited to the depth and the normal. For example, the feature information can indicate the outline, the size, the color, or the pattern of the object.

[0086] Further, for example, the data storage 100 stores observation space information indicating the state of the observation space OS. For example, in the observation space information, information related to the photographed object and observation viewpoint parameters related to the observation viewpoint OV are stored. The information related to the photographed object is a point group corresponding to the photographed object, and for example, includes the three-dimensional coordinates of the feature point group P and the vertex coordinates of the mesh (indicating the three-dimensional target of the photographed object). The observation viewpoint parameters are, for example, the position, the direction, and the angle of view of the observation viewpoint OV. Further, the direction of the observation viewpoint OV can be indicated by the three-dimensional coordinates of the gaze point or by the vector information indicating the line-of-sight direction.

[0087] Further, the data stored in the data storage 100 is not limited to the example described above. For example, the data storage 100 can store the photographed image G1 in chronological order. Further, for example, the data storage 100 can store the two-dimensional coordinates of the feature point group P extracted from the photographed image G1 in chronological order, and can store the vector information indicating the change in the position of the feature point group P in chronological order. Further, for example, in the case of providing the user with an extended reality, the data storage 100 can store information related to a three-dimensional target indicating an object to be a synthetic object. The object to be a synthetic object is a fictitious object displayed together with the photographed image G1, and for example, a fictitious animal (including a character imitating a human), furniture, a home appliance, clothes, a vehicle, a toy, or a miscellaneous goods. The object to be a synthetic object can move in the observation space OS, or can be stationary on the spot without moving in particular.

[0088] [3-2. Photographed image acquisition unit]

[0089] The photographic image acquisition section 101 is mainly realized by the control section 11. The photographic image acquisition section 101 acquires a photographic image Gl taken by the photographic section 18 that can move in the real space.

[0090] By the term "movable in the real space RS", it means that the position and direction of the photographic section 18 can be changed, for example, it means that the housing containing the photographic section 18 can be moved, or the posture of the housing can be changed, or the housing can be rotated. In other words, it means that the photographic range (field of view) of the photographic section 18 can be changed. In addition, the photographic section 18 does not necessarily have to move continuously, and can temporarily stay at the current location without changing the position and direction.

[0091] In the present embodiment, since the photographic section 18 takes a photograph of the real space RS based on a prescribed frame rate, the photographic image acquisition section 101 acquires a photographic image Gl taken by the photographic section 18 at a prescribed frame rate.

[0092] The frame rate is the number of processes per unit time, and is the number of still images (the number of frames) per unit time in a moving image. The frame rate can be a fixed value, or can be specified by a user. For example, if the frame rate is set to N fps (N: a natural number, fps: Frames Per Second), the length of each frame is 1 / N seconds, the photographic section 18 takes a photograph of the real space RS for a frame as a unit of processing to generate a photographic image Gl, and the photographic image acquisition section 101 continuously acquires the photographic image Gl generated by the photographic section 18.

[0093] In the present embodiment, the photographic image acquisition section 101 acquires a photographic image Gl taken by the photographic section 18 in real time. That is, the photographic image acquisition section 101 acquires the photographic image Gl immediately after the photographic section 18 generates the photographic image Gl. The photographic image acquisition section 101 acquires the photographic image Gl within a prescribed time from the point in time at which the photographic section 18 generates the photographic image Gl.

[0094] In addition, the photographic image Gl can be acquired without being particularly in real time, and in this case, the photographic image acquisition section 101 can acquire image data (that is, still image data or moving image data after photographing is completed) stored in the data storage section 100. In addition, when the image data is stored in a computer or an information storage medium other than the image processing apparatus 10, the photographic image acquisition section 101 can acquire the image data from the computer or the information storage medium.

[0095] In addition, the frame rate can not be set in the imaging section 18, and in the case where imaging is performed irregularly, the imaging image acquisition section 101 can acquire the imaging image Gl every time the imaging section 18 performs imaging. For example, the user can manually perform an imaging instruction from the operation section 14, and in this case, the imaging section 18 generates the imaging image Gl every time the user performs the imaging instruction, and the imaging image acquisition section 101 can acquire the imaging image Gl generated every time the user performs the imaging instruction.

[0096] [3-3. Observation space information acquisition section]

[0097] The observation space information acquisition section 102 is mainly realized by the control section 11. The observation space information acquisition section 102 acquires observation space information including three-dimensional coordinates of the feature point group P in the observation space OS based on a change in the position of the feature point group P in the imaging image Gl.

[0098] The change in the position of the feature point group P refers to a change in the position on the image, and is a change in two-dimensional coordinates. The change in the position of the feature point group P is represented by vector information (two-dimensional vector information) of the screen coordinate system. That is, the observation space information acquisition section 102 acquires vector information representing a change in the position of each feature point included in the feature point group P.

[0099] The observation space information acquired by the observation space information acquisition section 102 is information representing the distribution of the feature point group P in the observation space OS, and is a so-called 3D map of the feature point group P. The observation space information at this stage is described with reference to FIG. 6. Figure 4 As described above, only the three-dimensional coordinates of the feature point group P are stored, and become sparse point group data that cannot represent the surface shape of the imaged object.

[0100] The observation space information acquisition section 102 extracts the feature point group P from the imaging image Gl and tracks the extracted feature point group P. In addition, the feature point can be a point representing a feature of the imaged object imaged by the imaging image Gl, and for example, can be a point representing a part of the outline of the imaged object, or a point inside the imaged object (for example, a center point). The extraction method of the feature point itself can be performed based on a known feature point extraction algorithm, and for example, can set a point on the outline of the imaged object detected by the outline extraction process as the feature point, can set a point at which outline lines intersect each other at a prescribed angle or more as the feature point, or can set an edge portion in the image as the feature point.

[0101] Further, for example, the observation space information acquisition section 102 can also extract feature points based on an algorithm called SIFT (Scale-Invariant Feature Transform: https: / / en.wikipedia.org / wiki / Scale-invariant_feature_transform) and can also extract feature points based on an algorithm called ORB (Oriented fast and Rotated Brief: http: / / www.willowgarage.com / sites / default / files / orb_final.pdf). According to these algorithms, there are cases in which portions other than corners or edges of a photographed object are extracted as feature points.

[0102] The relationship between the positional changes of the feature point group P and the three-dimensional coordinates is stored in advance in the data storage section 100 as a part of a formula, a table, or a program code. Since the positional changes of the feature point group P are two-dimensional information, this relationship can also be called a conversion rule for converting two-dimensional information into three-dimensional information. The observation space information acquisition section 102 acquires the three-dimensional coordinates in association with the positional changes of the feature point group P.

[0103] In the present embodiment, the observation space information acquisition section 102 acquires observation space information using SLAM technology. The feature points move on the image in the opposite direction to the direction in which the photographing section 18 moves with respect to the photographed object in the real space RS. Furthermore, the more distant the photographed object, the smaller the amount of movement of the feature points on the image. In the SLAM technology, the three-dimensional coordinates of the feature point group P are calculated based on these tendencies using the principle of triangulation. That is, the observation space information acquisition section 102 tracks the feature point group P and calculates the three-dimensional coordinates of the feature point group P based on the SLAM technology using the principle of triangulation.

[0104] Further, the observation space information acquisition section 102 estimates the position of the photographing section 18 based on the positional changes of the feature point group P and sets the observation viewpoint OV in the observation space OS based on the estimation result. For example, the observation space information acquisition section 102 estimates the current position and direction of the photographing section 18 and reflects the estimation result in the position and direction of the observation viewpoint OV.

[0105] The relationship between the position change of the feature point group P and the position and direction of the photographing section 18 is stored in advance in the data storage section 100 as a numerical expression, a table, or a part of a program code. The relationship can also show the relationship between two-dimensional vector information indicating the change of the feature point group P and three-dimensional coordinate and direction information indicating the position of the observation viewpoint OV. The observation space information acquisition section 102 acquires the three-dimensional coordinate and vector information associated with the position change of the feature point group P.

[0106] The observation viewpoint OV is set by the observation space information acquisition section 102, and when the photographing section 18 moves in the real space RS, the observation viewpoint OV moves in the observation space OS in the same manner as the photographing section 18. That is, the position and direction of the observation viewpoint OV in the observation space OS change in the same manner as the position and direction of the photographing section 18 in the real space RS. The estimation method of the position and direction of the photographing section 18 itself can apply a known viewpoint estimation method, and for example, the SLAM technique can be used.

[0107] [3-4. Mechanical learning section]

[0108] The mechanical learning section 103 is mainly realized by the control section 11. The mechanical learning section 103 acquires additional information related to the feature of the photographed object shown in the photographed image Gl, based on mechanical learning data related to the feature of the object.

[0109] The additional information indicates the feature of the appearance of the photographed object, and for example, can be information such as the three-dimensional shape, classification (kind), color, or pattern of the photographed object. In the present embodiment, information related to the three-dimensional shape of the photographed object estimated based on the mechanical learning data is described as an example of the additional information. The information related to the three-dimensional shape of the photographed object is information that can specify the concave-convex or direction of the surface of the photographed object in three dimensions, and for example, is information related to the mesh of the photographed object or information related to the normal line of the photographed object. In other words, the information related to the three-dimensional shape of the photographed object is surface information indicating the surface of the photographed object.

[0110] The information related to the mesh of the photographed object can be any information that can express the mesh in the observation space OS, for example, it can be dense point group data, it can be the vertex coordinates that constitute the mesh itself, or it can be a specific depth of the vertex coordinates. In addition, the "dense" here means a density (density of a fixed value or more) that can express the surface shape of the photographed object, for example, it has the same degree of density as the vertex of a general mesh in computer graphics technology. The depth is the depth of the mesh in the case of viewing from the observation viewpoint OV, and is the distance from the observation viewpoint OV to each vertex of the mesh. On the other hand, the information related to the normal of the photographed object can be any information that can specify the normal of the surface of the photographed object, for example, it can be vector information of the normal, or it can be the intersection angle between a specified plane (for example, the Xw-Yw plane) in the observation space OS and the normal.

[0111] The additional information can be in any data form, but in the present embodiment, a case in which two-dimensional feature quantity information is obtained by associating the position (two-dimensional coordinates in the screen coordinate system) of the photographed object in the photographed image G1 and the feature quantity related to the photographed object is described. Furthermore, as an example of the two-dimensional feature quantity information, a feature quantity image in which the feature quantity related to the photographed object is associated with each pixel is described. The feature quantity of each pixel of the feature quantity image is a value that represents the feature of the pixel, for example, it is the depth of the photographed object estimated based on the mechanical learning data. That is, the depth image G2 is an example of the feature quantity image. In addition, the feature quantity is not limited to the depth. For example, the feature quantity of the feature quantity image can be the normal of the photographed object estimated based on the mechanical learning data. That is, the normal image G3 is also an example of the feature quantity image.

[0112] The mechanical learning unit 103 specifies an object similar to the photographed object from among the objects shown in the mechanical learning data. The "similar" means similar in appearance, for example, it can mean similar in shape, or it can mean similar in both shape and color. The mechanical learning unit 103 calculates the degree of similarity of the object shown in the mechanical learning data to the photographed object, and in the case where the degree of similarity is a threshold value or more, it is determined that the object is similar to the photographed object. The degree of similarity can be calculated based on the difference in shape or the difference in color.

[0113] In the mechanical learning data, the objects are associated with the feature information, so the mechanical learning unit 103 acquires the additional information based on the feature information associated with the object similar to the photographed object. For example, the mechanical learning unit 103 acquires additional information that includes a plurality of feature information corresponding to a plurality of similar objects respectively in the case where a plurality of similar objects are specified from among the photographed image G1.

[0114] For example, the mechanical learning section 103 specifies an object similar to the photographed object from among the objects shown in the depth learning data. Then, the mechanical learning section 103 sets a pixel value indicating a depth associated with the specified object to a pixel of the photographed object in the photographed image G1, thereby generating a depth image G2. That is, the mechanical learning section 103 sets, for each region in which the photographed object appears in the photographed image G1, vector information associated with an object similar to the photographed object.

[0115] In addition, for example, the mechanical learning section 103 specifies an object similar to the photographed object from among the objects shown in the normal line learning data. Then, the mechanical learning section 103 sets a pixel value indicating a normal line associated with the specified object to a pixel of the photographed object in the photographed image G1, thereby generating a normal line image G3. That is, the mechanical learning section 103 sets, for each region in which the photographed object appears in the photographed image G1, vector information associated with an object similar to the photographed object.

[0116] Further, the observation space information acquisition section 102 and the mechanical learning section 103 can perform processing based on photographed images G1 photographed in mutually different frames, but in the present embodiment, a case where processing is performed based on photographed images G1 photographed in mutually identical frames will be described. That is, the photographed image G1 referred to by the observation space information acquisition section 102 for acquisition of observation space information is identical to the photographed image G1 referred to by the mechanical learning section 103 for acquisition of additional information, and is photographed from the same viewpoint (position and direction of the photographing section 18).

[0117] [3-5. Integration section]

[0118] The integration section 104 is mainly realized by the control section 11. The integration section 104 integrates the observation space information and the additional information. By integration, it means that the amount of information of the observation space OS is increased based on the observation space information and the additional information. For example, the following corresponds to integration: compared with the observation space OS indicating the three-dimensional coordinates of the feature point group P, the number of point groups is further increased, information other than the three-dimensional coordinates (for example, normal line information) is added to the three-dimensional coordinates of the feature point group P, or these are combined to increase the point groups and the additional information.

[0119] The integration section 104 can generate new information based on the observation space information and the additional information, or can not generate new information and add the additional information to the observation space information. For example, the integration section 104 makes the number of point groups shown in the observation space information increase to be dense point group data, or adds information such as normal information to the three-dimensional coordinates of the feature point group P shown in the observation space information, or makes a combination of these and makes the observation space information dense point group data and adds information such as normal information. In the present embodiment, a case where the additional information indicates a three-dimensional shape of the photographed object, and the integration section 104 adds information related to the three-dimensional shape based on the additional information to the observation space information (sparse point group data) indicating the three-dimensional coordinates of the feature point group P is described.

[0120] In addition, in the present embodiment, since the two-dimensional feature quantity information is used as the additional information, the integration section 104 performs processing based on a result of comparison between the two-dimensional observation information indicating a case where the observation space OS is observed from the observation viewpoint OV and the two-dimensional feature quantity information. The two-dimensional observation information is information in which the observation space OS which is a three-dimensional space is projected to a two-dimensional space, and is information in which information expressed in three dimensions is converted to two dimensions. For example, in the two-dimensional observation information, positions (two-dimensional coordinates) of the feature point group in the two-dimensional space are associated with depths of the feature point group in the observation space OS. Further, the two-dimensional coordinates of the feature point group can be expressed by real numbers. That is, the two-dimensional coordinates of the feature point group do not necessarily have to be expressed by integers, and can be expressed by values including decimals.

[0121] Further, in the present embodiment, a case where the feature quantity image (for example, the depth image G2 and the normal image G3) is used as the two-dimensional feature quantity information is described, and for example, the integration section 104 performs processing based on a result of comparison between the observation space image indicating a case where the observation space OS is observed from the observation viewpoint OV and the feature quantity image. That is, since the observation space information which is information in three dimensions and the feature quantity image which is information in two dimensions differ in dimension, the integration section 104 performs processing after making these dimensions uniform. Further, the integration section 104 can perform processing after projecting the feature quantity image to the observation space OS to be information in three dimensions, like the following variation.

[0122] Figure 9 is a diagram indicating an example of the observation space image. In Figure 9 , a case where the observation space OS is observed from the observation viewpoint OV is indicated. Figure 4In the case of the observation space OS, the feature point group P appearing in the observation space image G4 is schematically represented by a circle of a fixed size, and in fact, each feature point can be represented by only one or several pixels. In addition, as described above, the position of the feature point is not represented by an integer value indicating the position of the pixel, but can be represented by a float value capable of representing a decimal point.

[0123] The integration section 104 generates the observation space image G4 by converting the three-dimensional coordinates of the feature point group P into two-dimensional coordinates of the screen coordinate system. Therefore, the observation space image G4 can be said to be a 2D projection map that projects the observation space OS, which is information in three dimensions, to information in two dimensions. This conversion process itself can apply a known coordinate conversion process (geometric shape process).

[0124] For example, the observation space image G4 indicates the depth of the feature point group P in the observation space OS. That is, the pixel value of the observation space image G4 indicates the depth as with the depth image G2. In addition, with respect to a portion in which the feature point group P does not appear in the observation space image G4, either no pixel value can be particularly set or a prescribed value indicating that the feature point group P does not appear can be set.

[0125] The observation space image G4 is the same size (the same number of pixels in the vertical and horizontal directions) as the photographic image G1, and can be either a color image or a grayscale image. In the case of a grayscale image, the pixel value of the observation space image G4 is represented by a grayscale value, and the grayscale value of the pixel in which the feature point group P appears in the observation space image G4 is set to a value indicating the depth of the feature point group P. Figure 9 In the case of a color image, the pixel value of the observation space image G4 is schematically represented by the density of a halftone dot, and it is indicated that the deeper the halftone dot, the lower the depth (the shorter the distance), and it is indicated that the lighter the halftone dot, the higher the depth (the longer the distance). For example, the halftone dot of the pixel of the feature point P11 close to the observation view point OV is dense, the halftone dot of the pixel of the feature point P12 not so far from the observation view point OV is moderately dense, the halftone dot of the pixel of the feature point P13 far from the observation view point OV is light, and the halftone dot of the pixel of the feature point P14 farthest from the observation view point OV is light.

[0126] The integration section 104 specifies the pixel in which the feature point group P appears in the observation space image G4, and performs processing based on the pixel value of the pixel of the feature amount image (for example, the depth image G2 and the normal image G3). If it is the case of a grayscale image, the integration section 104 specifies the pixel value of the pixel in which the feature point group P appears in the observation space image G4, and performs processing based on the pixel value of the pixel of the feature amount image. Figure 9 If it is the case of a color image, the integration section 104 specifies the two-dimensional coordinates of the pixel in which the feature point group P appears in the observation space image G4, and performs processing based on the pixel value of the pixel of the feature amount image.

[0127] Figure 10 is a diagram indicating an example of the processing performed by the integration section 104. As Figure 10 ​​​​As shown, firstly, the integration unit 104 sets a grid M for the observation space OS based on the depth image G2. For example, the integration unit 104 projects the depth of each pixel shown in the depth image G2 onto the observation space OS, and sets a temporary grid M (the initial grid M) in such a way that the vertex coordinates are the locations that are only at that depth away from the observation viewpoint OV. In other words, the integration unit 104 converts the depth of each pixel in the depth image G2 into three-dimensional coordinates, and sets these three-dimensional coordinates as the vertex coordinates of the grid M.

[0128] Furthermore, the method for converting a set of points in three-dimensional space into a grid based on depth information can itself employ various known techniques. In other words, the method for converting depth information, which is considered 2.5-dimensional information, into three-dimensional point group data can itself employ various known techniques. For example, the techniques described in "On Fast Surface Reconstruction Methods for Large and Noisy Point Clouds" (http: / / ias.informatik.tu-muenchen.de / _media / spezial / bib / marton09icra.pdf) can also be used to define a grid M for the observation space OS.

[0129] like Figure 10 As shown, since the grid M defined from the depth image G2 does not have a scale, it is not limited to the feature point group P, which is the measured value, being located at the same position as the grid M. Therefore, the integration unit 104 changes the scale of the grid M based on the comparison result between the observation spatial image G4 and the depth image G2. That is, the integration unit 104 specifies the portion of the grid M corresponding to the feature point group P, and changes the scale of the grid M in such a way that the specified portion is close to the feature point group P.

[0130] Scale is a parameter that affects the position or size of grid M. Changing the scale changes the spacing of the point groups that make up grid M, or the distance between grid M and the observation viewpoint OV. For example, increasing the scale increases the overall spacing of the point groups and makes grid M larger, or increases the distance between grid M and the observation viewpoint OV. Conversely, decreasing the scale decreases the overall spacing of the point groups and makes grid M smaller, and decreases the distance between grid M and the observation viewpoint OV.

[0131] For example, the integration unit 104 calculates the scale by ensuring that the index value representing the offset between the feature point group P and the grid M is less than a threshold. This index value is calculated based on the distance between the feature point group P and the grid M. For example, the index value can be calculated using a formula that sets the distance between each feature point and the grid M as an argument; for example, it can be either the sum of the distances between the feature point group P and the grid M or the average of those distances.

[0132] For example, the integration section 104 calculates the index value while changing the scale, and determines whether the index value is less than a threshold value. The integration section 104 changes the scale again and re-performs the determination process in a case where the index value is equal to or greater than the threshold value. On the other hand, the integration section 104 determines the current scale in a case where the index value is less than the threshold value. The integration section 104 changes the mesh M in such a manner that the shift of the feature point group P from the mesh M as a whole becomes small, by thus determining the scale.

[0133] In addition, as shown in FIG. 10, the integration section 104 can change the mesh M locally based on the changed mesh M and the feature point group P after changing the mesh M as a whole by changing the scale. For example, the integration section 104 determines whether the distance from each feature point to the mesh M is equal to or greater than a threshold value. If the distance is equal to or greater than the threshold value, the integration section 104 changes the mesh M corresponding to the feature point in such a manner that the mesh M approaches the feature point. The local change of the mesh M is performed by changing the three-dimensional coordinates of a part of the vertices (vertices in the vicinity of the feature point that becomes the object). Figure 10

[0134] Further, the process performed by the integration section 104 is not limited to the above-described example. For example, the integration section 104 can change the mesh M again based on the normal image G3 after changing the mesh M based on the depth image G2. In this case, the integration section 104 acquires the normal information of the mesh M changed based on the depth image G2, and compares the normal information with the normal information shown in the normal image G3. Then, the integration section 104 changes the mesh M locally in such a manner that the difference between the two becomes small. Further, the integration section 104 compares the observation space image G4 with the normal image G3 by using the same process as the depth image G2, whereby the correspondence relationship between the mesh M and the normal information shown in the normal image G3 is specified.

[0135] As described above, the integration section 104 of the present embodiment sets the mesh M of the photographic object in the observation space OS based on the two-dimensional feature quantity information, and changes the scale of the mesh M based on the comparison result of the two-dimensional observation information and the two-dimensional feature quantity information. For example, the integration section 104 sets the mesh in the observation space OS based on the additional information, and changes the mesh based on the observation space information.

[0136] ​For example, the integration section 104 changes the mesh M locally after changing the scale of the mesh M based on the comparison result of the two-dimensional observation information and the two-dimensional feature quantity information. Also, for example, the integration section 104 sets the mesh M of the photographic object in the observation space OS based on the depth image G2, and changes the scale of the mesh M based on the comparison result of the observation space image G4 and the depth image G2. Further, the integration section 104 changes the mesh M locally after changing the scale of the mesh M based on the comparison result of the observation space image and the feature quantity image (for example, the depth image G2 and the normal image G3).

[0137] Further, the integration section 104 can change the mesh portion around the mesh portion corresponding to the three-dimensional coordinates of the feature point group shown in the observation space information after changing the mesh portion. The "around" means a portion within a prescribed distance. For example, the integration section 104 changes the mesh portion in such a manner that the mesh portion between the feature points is smoothed after changing the mesh M temporarily set in such a manner as to coincide with the three-dimensional coordinates of the feature point group. The "smoothed" means, for example, that the change in the concave-convex is not excessively sharp, and the change in position is smaller than a threshold value. For example, the integration section 104 changes the mesh portion in such a manner that the change in the concave-convex of the mesh M is smaller than a threshold value.

[0138] Further, the method of changing the mesh portion itself can also use a known technique, and for example, the method called ARAP described in "As-Rigid-As-Possible Surface Modeling" (http: / / igl.ethz.ch / projects / ARAP / arap_web.pdf) can also be used. By changing the mesh portion around the mesh portion coinciding with the feature point group, it is possible to make each mesh portion coincide with the surrounding order, and it is possible to set a natural mesh more smoothly.

[0139] It is also possible to directly use the ARAP method, but in the present embodiment, a case where the ARAP method is extended and the mesh M is changed based on the reliability of the mesh estimation will be described.

[0140] For example, since the mesh M is estimated using the machine learning, there are portions where the reliability of the mesh estimation is high and portions where the reliability is low in the mesh M. Therefore, the integration section 104 can maintain the shape of the portion where the reliability is high without making a large change, and change the shape of the portion where the reliability is low to some extent. Further, the "reliability" means the degree of similarity of the estimated accuracy of the shape to the surface shape of the photographic object.

[0141] For example, in a case where the subject is oriented toward the photographing section 18, its surface is clearly imprinted in the photographed image G1, so the estimation accuracy of the mesh M is high in many cases. On the other hand, in a case where the subject is oriented aside with respect to the photographing section 18, its surface is not so imprinted in the photographed image G1, so there are cases where the estimation accuracy of the mesh M is low. Therefore, in the present embodiment, it is assumed that the portion of the mesh M that is oriented toward the observation view point OV is highly reliable, and the portion that is not oriented toward the observation view point OV (the portion that is oriented aside with respect to the observation view point OV) is less reliable.

[0142] Figure 11 and Figure 12 An explanatory diagram of a process of changing the mesh M by extending the ARAP method. As shown in Figure 11 In the present embodiment, it is assumed that the closer the angle θ between the normal vector n of the vertex of the mesh M and the vector d connecting the observation view point OV and the vertex is to 180°, the higher the reliability, and the closer the angle θ is to 90°, the lower the reliability. Furthermore, in the present embodiment, it is assumed that there is no case where the mesh M is oriented in the opposite direction of the observation view point OV, and in principle, it is assumed that there is no case where the angle θ is less than 90°.

[0143] For example, the integration section 104 changes the mesh portion based on the direction (angle θ) of the mesh portion with respect to the observation view point OV. That is, the integration section 104 determines the amount of change of the mesh portion based on the direction of the mesh portion with respect to the observation view point OV. The amount of change of the mesh portion refers to how the shape is deformed, and is the amount of change (amount of movement) of the three-dimensional coordinates of the vertex.

[0144] Furthermore, the relationship between the direction with respect to the observation view point OV and the amount of change of the mesh portion is stored in the data storage section 100 in advance. The relationship can be stored as data in the form of a formula or a table, or can be described as a part of the program code. The integration section 104 changes based on the amount of change that associates each mesh portion of the mesh M with the direction of the mesh portion with respect to the observation view point OV.

[0145] For example, the integration section 104 makes the amount of change of the mesh portion smaller the more the mesh portion is oriented toward the observation view point OV (the closer the angle θ is to 180°), and makes the amount of change of the mesh portion larger the less the mesh portion is oriented toward the observation view point OV (the closer the angle θ is to 90°). In other words, the integration section 104 makes the rigidity of the mesh portion higher the more the mesh portion is oriented toward the observation view point OV, and makes the rigidity of the mesh portion lower the less the mesh portion is oriented toward the observation view point OV. Furthermore, the mesh portion not being oriented toward the observation view point OV means that the mesh portion is oriented aside with respect to the observation view point OV.

[0146] Suppose that the rigidity is not changed according to the reliability of each portion of the mesh M as described above, then as Figure 12As shown, there exists a situation where the mesh M deforms unnaturally by being stretched by the feature point P. To address this, by deforming while maintaining the rigidity of the high-reliability portion (the portion facing the viewing point OV), the shape of the high-reliability portion is preserved, thus preventing the unnatural deformation described above and forming a more natural mesh M.

[0147] Furthermore, in the following description, the vertex of the mesh M corresponding to the feature point P will be denoted as v. i For example, vertex v i It is the straight line that most closely connects the observation viewpoint OV and the feature point P. Figure 11 The vertex of the intersection point of the vector d (the point line) and the grid M. For example, the integration unit 104 may also change the grid M based on the following formulas 1-7. For example, formulas 1-7 (especially formulas 3-4) are examples of the relationship between the direction relative to the observation viewpoint OV and the amount of change of the grid portion as described above.

[0148] First, the integration unit 104 targets each vertex v. i Calculate the value of the energy function shown on the left side of the following equation 1.

[0149] [Number 1]

[0150]

[0151] In equation 1, it will be related to vertex v i The corresponding nearest neighbor record is C. i Record each of the nearest neighbor vertices as v j Furthermore, the term "nearest neighbor" refers to vertex v. i The surrounding vertices are defined here as adjacent vertices (one-ring neighborhood), but two or more distant vertices can also be considered as nearest neighbors. Additionally, the modified vertex is recorded as v'. i The changed nearest neighbor is recorded as C' i Record the changed adjacent vertices as v' j .

[0152] N(v) on the right side of equation 1 i ) is vertex v i C, the neighbor i The adjacent vertices v contained in j The set of . The right-hand side of expression 1, R. i It is a 3×3 rotated matrix. As shown in Equation 1, the energy function E(C') i ) is relative to vertex v i adjacent vertices v j The relative positional change multiplied by the weighting coefficient ω ij The sum of the obtained values. Even relative to vertex vi and the adjacent vertex v j moves greatly, the weighted coefficient ω ij is large, and the value of the energy function E(C i ) becomes small. On the contrary, even if the vertex v i moves slightly, the weighted coefficient ω j is small, and the value of the energy function E(C ij ) becomes large. i

[0153] The weighted coefficient ω ij is determined using a combination of the vertex v i and the adjacent vertex v j . For example, the integrating section 104 calculates the weighted coefficient ω ij based on the following mathematical expression 2. α ij and β ij on the right side of the mathematical expression 2 are angles on the opposite side of the edge (i, j) of the mesh M.

[0154] [Equation 2]

[0155]

[0156] For example, the integrating section 104 calculates the total value of the energy function E(C i ) calculated for each vertex v i based on the following mathematical expression 3.

[0157] [Equation 3]

[0158]

[0159] In the mathematical expression 3, the changed mesh M is written as M'. As shown on the right side of the mathematical expression 3, the integrating section 104 calculates the total value of the value obtained by multiplying the value of the energy function E(C i ) for each vertex v i by the weighted coefficient ω i . The weighted coefficient ω i is determined, for example, using an S-shaped function or the like. For example, the integrating section 104 calculates the weighted coefficient ω i based on the following mathematical expression 4.

[0160] [Equation 4]

[0161]

[0162] a and b on the right side of the mathematical expression 4 are coefficients, and are fixed values. For example, the more the angle θ approaches 180°, the larger the weighted coefficient ω i ​The larger the grid portion is, the greater the influence of the change of the grid portion on the total value of the energy function (left side of Equation 3) becomes. Therefore, the total value of the energy function greatly increases by making the grid portion change only slightly. On the other hand, the closer the angle θ is to 90°, the smaller the weighting coefficient ω i The smaller the grid portion is, the smaller the influence of the change of the grid portion on the total value of the energy function becomes. Therefore, the total value of the energy function does not so increase even if the grid portion is greatly changed. By thus setting the weighting coefficient ω i , the rigidity can be changed in accordance with the reliability of the mesh M.

[0163] Further, the integration section 104 can also change the mesh M in a manner in which the total value of the energy function E(C i ) calculated using Equation 3 becomes smaller, but the integration section 104 can further take into account a bending coefficient. The bending coefficient is a value indicating to what extent the surface of the mesh M is bent (deformed), and is calculated based on the following Equation 5, for example, as described in "Z. Levi and C. Gotsman. Smooth rotation enhanced as-rigid-as-possible mesh animation. IEEE Transactions on Visualization and Computer Graphics, 21:264-277, 2015."

[0164] [Equation 5]

[0165] B ij = αA||R i -R j ||

[0166] α on the right side of Equation 5 is a weighting factor, and A is a surface whose characteristics do not change even if the scale is changed. R i , R j on the right side of Equation 1 is a 3 x 3 rotation matrix. For example, the integration section 104 calculates the bending coefficient B i for each combination of the vertex v j and the adjacent vertex v ij , and can reflect this in the total value of the energy function E(C i ) based on the following Equation 6.

[0167] [Equation 6]

[0168]

[0169] Furthermore, since the photographic images G1 are repeatedly acquired according to a prescribed frame rate, and the integration unit 104 repeatedly performs the processing described above, the integration unit 104 can also consider the previously calculated scale and calculate the absolute scale s of the observation space OS at time t based on the following formula 7. w t Furthermore, the s on the right side of the equation 7... c t It is the scale set for grid M.

[0170] [Number 7]

[0171]

[0172] [4. Processing performed in this embodiment]

[0173] Figure 13 This is a flowchart illustrating an example of a process performed in the image processing apparatus 10. Figure 13 The processing shown is executed by the control unit 11 according to the program actions stored in the storage unit 12. Figure 13 The process shown is using Figure 8 The illustrated function block performs one example of processing, and is executed for each frame captured by the camera unit 18.

[0174] In addition, in execution Figure 13 During the processing shown, the initialization of the following mapping process has been completed, and the observation space OS (a 3D map of the feature point group P) has been generated. That is, the control unit 11 tracks the feature point group P extracted from the photographic image G1, and uses SLAM technology to set the three-dimensional coordinates of the feature point group P and the observation viewpoint OV in the observation space OS.

[0175] like Figure 13 As shown, firstly, the control unit 11 performs the photographic image acquisition process (S1). In S1, the control unit 11 acquires the photographic image G1 generated by the photographic unit 18 in the current frame. Furthermore, the control unit 11 can also record the photographic image G1 in the storage unit 12 in a time sequence. That is, the control unit 11 can also record the history of the photographic image G1 in the storage unit 12.

[0176] The control section 11 performs 2D tracking processing (S2) based on the captured image G1 acquired in S1. The 2D tracking processing is processing to track a change in position on an image of the feature point group P. In S2, first, the control section 11 acquires the feature point group P from the captured image G1 acquired in S1. Then, the control section 11 specifies a correspondence relation between the feature point group P and the feature point group P of the captured image G1 acquired in the most recent frame (the previous frame), and acquires vector information indicating a difference in two-dimensional coordinates of the feature point group P. Further, the control section 11 records the two-dimensional coordinates of the feature point group P extracted in S2 in association with the captured image G1 in the storage section 12. In addition, the control section 11 can record the vector information of the feature point group P in the storage section 12 in chronological order.

[0177] The control section 11 determines whether to start mapping processing (S3). The mapping processing is processing to update the observation space information (three-dimensional coordinates of the feature point group P). The mapping processing can be performed every frame or once for a plurality of frames. In the case where the mapping processing is performed once for a plurality of frames, the interval of the mapping processing can be a fixed value or a variable value.

[0178] Further, here, a case where the mapping processing is started again in the next frame after the above previous mapping processing ends is described. Therefore, in S3, it is determined whether the previous mapping processing ends, and if the previous mapping processing ends, it is determined to start the mapping processing, and if the previous mapping processing does not end, it is not determined to start the mapping processing.

[0179] In the case where it is determined to start the mapping processing (S3; Y), the control section 11 starts the mapping processing (S4) based on the captured image G1 acquired in S1. The mapping processing started in S4 is executed in parallel with (or in the background of) the main routine processing shown in FIG. 8. Figure 13

[0180] Figure 14 is a flowchart showing an example of the mapping processing. As shown in Figure 14 , the control section 11 calculates the three-dimensional coordinates of the feature point group P based on the execution result of the 2D tracking processing executed in S2 (S41). In S41, the control section 11 calculates a cumulative amount of movement of the feature point group P from the previous mapping processing, and calculates the three-dimensional coordinates of the feature point group P using the SLAM technique.

[0181] The control section 11 estimates the position of the photographing section 18 based on the execution result of the 2D tracking processing executed in S2 (S42). In S42, the control section 11 calculates a cumulative amount of movement of the feature point group P from the previous mapping processing, and calculates the position and direction of the photographing section 18 using the SLAM technique.

[0182] ​The control section 11 updates the observation space information based on the results of S41 and S42 (S43). In S43, the control section 11 updates the three-dimensional coordinates of the feature point group P and the observation viewpoint parameters based on the three-dimensional coordinates of the feature point group P calculated in S41 and the position and direction calculated in S42.

[0183] Returning to Figure 13 In a case where it is not determined that the mapping processing is started (S3; N), or in a case where the mapping processing is started in S4, the control section 11 determines whether or not the decryption processing is started (S5). The decryption processing is processing of estimating the three-dimensional shape of the photographed object by using machine learning, and in the present embodiment, is processing of acquiring the depth image G2 and the normal image G3. The decryption processing can be executed every frame, or can be executed once for a plurality of frames. In a case where the decryption processing is executed once for a plurality of frames, the execution interval of the decryption processing can be a fixed value, or can be a variable value.

[0184] Further, since the decryption processing is more computationally intensive (higher load) than the mapping processing, in this case, the execution interval of the decryption processing can also be longer than that of the mapping processing. For example, the mapping processing can be executed once for two frames, and the decryption processing can be executed once for three frames.

[0185] In addition, here, a case where the decryption processing is started again for the next frame after the decryption processing of the above previous frame ends is described. Therefore, in S5, it is determined whether or not the decryption processing of the previous time ends, and if the decryption processing of the previous time ends, it is determined that the decryption processing is started, and if the decryption processing of the previous time does not end, it is not determined that the decryption processing is started.

[0186] In a case where it is determined that the decryption processing is started (S5; Y), the control section 11 starts the decryption processing based on the photographed image G1 which is the same as the mapping processing being executed (S6). The decryption processing started in S6 is executed in parallel with (or in the background of) the main routine processing shown in FIG. 8. Figure 13

[0187] Figure 15 is a flowchart showing an example of the decryption processing. As shown in Figure 15

[0188] ​​The control section 11 acquires the normal line image G3 based on the photographic image G1 and the normal line learning data (S62). In S62, the control section 11 specifies a portion similar to the object shown in the normal line learning data from among the photographic image G1. Then, the control section 11 generates the normal line image G3 by setting the vector information of the normal line of the object shown in the normal line learning data as the pixel value of each pixel of the portion.

[0189] Returning to Figure 13 In a case where the decryption processing is not determined to be started (S5; N), or in a case where the decryption processing is started in S6, the control section 11 determines whether the integration processing is started (S7). The integration processing is a processing of setting a mesh of the photographic object in the observation space OS. The integration processing can be executed every frame, or can be executed once for a plurality of frames. In a case where the integration processing is executed once for a plurality of frames, the execution interval of the integration processing can be a fixed value, or can be a variable value.

[0190] Further, here, a case where the integration processing is started in a case where both the mapping processing and the decryption processing are completed is described. Therefore, in S7, it is determined whether the mapping processing and the decryption processing in execution are ended, and if both are ended, it is determined that the integration processing is started, and if either is not ended, it is not determined that the integration processing is started.

[0191] In a case where the integration processing is determined to be started (S7; Y), the control section 11 starts the integration processing (S8). The integration processing started in S8 is executed in parallel with (or in the background of) the main routine processing shown in Figure 13

[0192] Figure 16 is a flowchart showing an example of the integration processing. As shown in Figure 16 The control section 11 generates an observation space image G4 showing a case where the feature point group P in the observation space OS is observed from the observation viewpoint OV (S81). The observation space image G4 is the same image as the depth image G2, and each pixel shows the depth of the feature point group P. In S81, the control section 11 generates the observation space image G4 by calculating the distance from the observation viewpoint OV to the feature point group P.

[0193] ​The control section 11 corrects the mesh shown in the depth image G2 based on the observation space image G4 generated in S81 (S82). In S82, the control section 11 specifies the position of the mesh corresponding to the feature point group P based on the observation space image G4 and the depth image G2, and corrects the scale of the mesh in such a manner that the difference in depth becomes small. Further, the control section 11 locally corrects the mesh in such a manner that the distance between the feature point and the mesh is greater than a threshold value. In addition, the control section 11 corrects the mesh in such a manner that the mesh portion around the mesh portion coinciding with the feature point group P is smoothed. Further, the control section 11 can change the mesh portion based on the direction of the mesh portion with respect to the observation viewpoint OV.

[0194] The control section 11 further corrects the mesh corrected in S82 based on the normal image G3 (S83). In S83, the control section 11 specifies the normal direction corresponding to the feature point group P based on the observation space image G4 and the depth image G2, and corrects the mesh in such a manner that the difference between the normal of the mesh corrected in S82 (the normal of the portion of the mesh corresponding to the feature point group P) and the normal shown in the normal image G3 becomes small.

[0195] The control section 11 updates the observation space OS based on the mesh corrected in S83 (S84). In S84, the control section 11 stores the vertex coordinates of the mesh corrected in S83 in the observation space information. Thus, the observation space information for the sparse point group data in the mapping process becomes dense point group data by the integration process.

[0196] Returning to Figure 13 , in a case where it is not determined to start the integration process (S7; N), or in a case where the integration process is started in S8, the control section 11 ends the present process. Hereinafter, the process of Figure 13 is executed again every time the frame is accessed.

[0197] Further, in a case where the extended reality is provided in real time, the control section 11 can also arrange a three-dimensional target representing a fictitious object in the observation space OS before ending the present process, generate a virtual image representing a case where the observation space OS is observed from the observation viewpoint OV, and display the virtual image after being synthesized with the photographic image Gl on the display section 15. As the photographic image Gl synthesized at this time, either the image acquired in S1 of the present frame or the photographic image Gl referred to in the mapping process and the decryption process can be used. Further, in the extended reality, a target representing a mobile body such as a ball or a vehicle can also be synthesized. In this case, a hit determination between the mesh of the observation space OS and the target representing the mobile body can also be performed, and the mobile body can bounce back or climb a wall.

[0198] Furthermore, as mentioned above, mapping and decryption processes do not need to be executed every frame; they can be executed once across multiple frames. Consequently, since decryption requires more computation than mapping, its execution interval can also be longer than that of mapping.

[0199] Figure 17 This is a diagram illustrating an example of the execution intervals for each process. In Figure 17 In the example shown, the photographic image acquisition process (S1) and the 2D tracking process (S2) are executed every frame. On the other hand, the mapping process ( Figure 14 Frame n (where n is an integer greater than 2) is executed once, and frame m (where m is an integer greater than 2, and m > n) is decrypted once. The integration process is executed after the decryption process is complete. For example... Figure 17 As shown, the photographic image G1 referenced in the mapping and decryption processes becomes the photographic image G1 acquired with the same frame, and the mapping and decryption processes are performed based on the photographic image G1 acquired from the same viewpoint.

[0200] According to the image processing apparatus 10 described above, by integrating the photographic image G1 captured by the camera unit 18 with additional information obtained through machine learning, the composition of information used to improve the observation space OS can be simplified. For example, even without using special sensors such as depth cameras, information other than the three-dimensional coordinates of the feature point group P can be added to the observation space OS. Therefore, even terminals such as smartphones that do not have special sensors can generate a high-precision observation space OS.

[0201] Furthermore, when using feature images (e.g., depth image G2 or normal image G3) as supplementary information, the image processing apparatus 10 can compare images observed from the same viewpoint with each other by comparing the observation space image G4 with the feature images. In other words, in conventional techniques using two RGB-D cameras arranged side-by-side, errors are generated in the observation space OS due to differences in viewpoint positions. However, since the image processing apparatus 10 uses the same viewpoint, it can prevent the generation of errors and improve the reproducibility of the observation space OS.

[0202] Furthermore, by changing the scale of the mesh based on the comparison results between the observation space image G4 and the depth image G2, the image processing device 10 can make the mesh obtained in machine learning as a whole approximate the measured value, thus improving the reproducibility of the observation space OS with simplified processing. For example, since the mesh vertices are not changed individually, but the mesh as a whole approximates the measured value by changing the scale, the processing is simplified (the amount of computation is reduced), the processing load of the image processing device 10 is reduced, and the processing speed is improved.

[0203] In addition, since the mesh is adjusted locally after the scale of the mesh is changed, the degree of reproduction of the observation space OS can be improved more efficiently. In this case, the mesh portion is not changed individually with respect to all the feature point groups P, but only a portion having a large difference is set as a target, whereby simplification of the process for improving the degree of reproduction of the observation space OS can be achieved, so that the processing load of the image processing apparatus 10 can be reduced more efficiently, and improvement of the processing speed is sought.

[0204] In addition, by using the three-dimensional shape of the photographed object as the additional information, the three-dimensional shape of the real space RS can be reproduced in the observation space OS, so that the configuration for reproducing the three-dimensional shape of the real space RS in the observation space OS in detail can be simplified.

[0205] In addition, by using the information related to the mesh of the photographed object as the additional information, the mesh representing the photographed object can be arranged in the observation space OS, so that the configuration for arranging the mesh representing the object of the real space RS in the observation space OS can be simplified. In addition, the observation space OS is sparse and highly accurate based on the observation data, and the additional information has a low correct answer rate since it is a predicted value using machine learning, but by integrating the feature point group of the sparse and accurate observation space OS and the mesh of the additional information which is close and has a low accuracy, the accuracy can be ensured, and close data can be obtained.

[0206] In addition, in the case where the information related to the mesh of the photographed object is used as the additional information, by changing the mesh based on the observation space information which is a measured value, the degree of reproduction of the observation space OS can be improved efficiently.

[0207] In addition, by changing the mesh portion around the mesh portion corresponding to the three-dimensional coordinates of the feature point group after the mesh portion is changed, the shape of the mesh can be smoothed. That is, improvement of the data accuracy as data for storing between the feature points can be sought, and the degree of reproduction of the observation space OS can be improved efficiently.

[0208] In addition, by changing the mesh portion based on the direction of each mesh portion with respect to the observation viewpoint OV, the mesh portion having a high reliability can be integrated while its shape is maintained as much as possible, and the mesh portion having a low reliability can be integrated after its shape is changed, so that the degree of reproduction of the observation space OS can be improved efficiently.

[0209] In addition, by using the information related to the normal line of the photographed object as the additional information, the normal line can be set in the observation space OS, and the three-dimensional shape of the photographed object can be represented, so that the configuration for reproducing the direction of the surface of the object of the real space RS in the observation space OS can be simplified.

[0210] Further, by generating the observation space information and the additional information from the photographic image G1 of the same frame, the correspondence relationship of the image pairs at the same viewpoint can be specified, and thus the generation of errors due to the difference in the viewpoint position as described above can be prevented, and the accuracy of the observation space OS can be more effectively improved.

[0211] [5. Modification]

[0212] Further, the present application is not limited to the embodiments described above. Appropriate changes can be made within the scope of the gist of the present application.

[0213] (1) For example, in the embodiment, the depth or the normal of the photographic object was described as an example of the additional information, but the additional information can also be information related to the classification of the photographic object. That is, the additional information can also be information that groups the pixels of the photographic image G1 for each photographic object. In this modification, as in the embodiment, a case where a feature amount image is used will be described, and a classification image that classifies the pixels of the photographic image G1 will be described as an example of the feature amount image.

[0214] Figure 18 is a drawing that shows an example of the classification image. As shown in Figure 18 , the classification image G5 is an image that is the same size as the photographic image G1 (the number of pixels in the vertical and horizontal directions is the same), and groups the regions within the image for each photographic object. The classification image G5 assigns a pixel value for each photographic object. That is, the classification image G5 is a label image that gives each pixel information that identifies the photographic object. Pixels with the same pixel value indicate the same photographic object.

[0215] The classification image G5 can be either a color image or a grayscale image. In Figure 18 , the pixel values of the classification image G5 are schematically represented by the density of halftone dots, and pixels with the same density of halftone dots indicate the same object. Thus, the pixels that represent the bed are the first pixel value. Similarly, the pixels that represent the wall are the second pixel value, the pixels that represent the floor are the third pixel value, and the pixels that represent the painting are the fourth pixel value.

[0216] For example, the integration unit 104 groups the feature point group P shown by the observation space information based on the classification image G5. For example, the integration unit 104 generates the observation space image G4, and specifies the pixels within the classification image G5 that correspond to the feature point group P, in the same manner as the method described in the embodiment. Then, the integration unit 104 specifies the pixel values of the pixels in the classification image G5, and classifies the feature points that represent the same value as the same group. That is, the integration unit 104 gives the three-dimensional coordinates of the feature point group P information that identifies the group.

[0217] According to the modification (1), by using the information related to the classification of the photographed object as the additional information, it is possible to group the point clouds of the observation space OS.

[0218] (2) In addition, for example, in the embodiment, the case where the normal image G3 is used in order to finely adjust the mesh M changed based on the depth image G2 is described, but the method of using the normal image G3 is not limited to the described example. For example, the integration section 104 can also attach normal information to the three-dimensional coordinates of the point cloud P.

[0219] Figure 19 is a diagram showing an example of the processing performed by the integration section 104. As shown in Figure 19 , the integration section 104 attaches normal information corresponding to each feature point to the feature point. As described in the embodiment, the integration section 104 can specify the correspondence between the feature point and the normal information by comparing the observation space image G4 and the normal image G3. For example, the integration section 104 can also increase the amount of information of the observation space OS by mapping the normal information on the straight line connecting the observation point OV and the feature point, that is, the normal information on the same pixel on the image, to the feature point.

[0220] In this case, the number of point clouds of the observation space OS does not increase, but the normal information is added, so the integration section 104 can generate a mesh showing the surface shape of the photographed object. Furthermore, in combination with the method described in the embodiment, the integration section 104 can also make the observation space OS tight point cloud data and attach normal information to the point cloud P. By so doing, it is possible to further increase the amount of information of the observation space OS.

[0221] In addition, for example, the case where the higher the pixel value of the depth image G2, the higher the depth is described, but the relationship between the pixel value and the depth can also be reversed, and it can also be expressed that the lower the pixel value, the higher the depth. Similarly, as long as the pixel value of the normal image G3 and the normal are based on a fixed rule and there is a correlation between them.

[0222] In addition, for example, in the embodiment, the case where the observation space information as three-dimensional information is converted into the observation space image G4 as two-dimensional information and then compared with the depth image G2 and the normal image G3 as two-dimensional information is described, but the depth image G2 and the normal image G3 can also be converted into three-dimensional information and then compared with the observation space information. That is, as long as the integration section 104 makes the observation space information and the additional information consistent in dimension and then specifies the correspondence between them, it can perform the processing of integrating these.

[0223] In addition, for example, a case where the additional information is image-form information is described, but the additional information can be any data form, and can be numerical group data other than image form, tabular data, or various data forms. In a case where information other than image form is used as the additional information, a process of particularly comparing images with each other can not be performed. Furthermore, the mechanical learning data can be caused to learn vertex coordinates of a mesh, and the additional information can be three-dimensional information other than two-dimensional information like an image. In this case, a process of making the dimension consistent with the observation space information can not be performed.

[0224] In addition, for example, a case where furniture and the like are arranged in the indoor space is described, but the furniture and the like can not be particularly arranged in the indoor space. In addition, for example, the indoor space is described as an example of the real space RS, but the real space RS can be an outdoor space, for example, a road, a parking lot, an event venue, or the like. In addition, for example, a case where the observation space OS reproduced by the image processing apparatus 10 is used for extended reality is described, but the observation space OS can be used in any scene, and can be used for mobile control of a robot.

[0225] (3) In addition, for example, a case where the image processing system is implemented by one image processing apparatus 10 is described, but the image processing system can include a plurality of computers.

[0226] Figure 20 is a diagram illustrating an example of the image processing system in a variation. As illustrated in Figure 20 , the image processing system S of the variation includes the image processing apparatus 10 and the server 20. The image processing apparatus 10 and the server 20 are connected to a network such as the Internet.

[0227] The server 20 is a server computer, and includes, for example, a control section 21, a storage section 22, and a communication section 23. The hardware configurations of the control section 21, the storage section 22, and the communication section 23 are the same as those of the control section 11, the storage section 12, and the communication section 13, respectively, and thus the description thereof is omitted.

[0228] The processing described in the embodiments and variations (1)-(2) can also be shared by the image processing apparatus 10 and the server 20. For example, the image processing apparatus 10 may implement the photographic image acquisition unit 101 and the observation space information acquisition unit 102, while the server 20 may implement the data storage unit 100, the machine learning unit 103, and the integration unit 104. In this case, the data storage unit 100 is mainly implemented by the storage unit 22, and the machine learning unit 103 and the integration unit 104 are mainly implemented by the control unit 21. The server 20 receives the photographic image G1 from the image processing apparatus 10. Moreover, similar to the method described in the embodiments, the machine learning unit 103 acquires additional information, and the integration unit 104 performs integration processing. Furthermore, the image processing apparatus 10 only needs to receive the result of the integration processing performed by the integration unit 104 from the server 20.

[0229] Alternatively, for example, the image processing apparatus 10 may include a photographic image acquisition unit 101, an observation spatial information acquisition unit 102, and a machine learning unit 103, while the server 20 may include an integration unit 104.

[0230] Alternatively, for example, all the functions of the data storage unit 100, the photographic image acquisition unit 101, the observation space information acquisition unit 102, the machine learning unit 103, and the integration unit 104 can be implemented in the server 20. In this case, the server 20 can also send the observation space information to the image processing device 10.

[0231] In addition, Figure 20 In this description, one image processing device 10 and one server 20 are each represented, and the image processing system S includes two computers. However, the image processing system S may also include three or more computers. In this case, the processing can be shared by three or more computers. Furthermore, for example, it may not be necessary to include a camera unit 18 in the image processing device 10; the image acquisition unit 101 may acquire the image G1 captured by the camera unit 18, which is not included in the image processing device 10. Moreover, the data storage unit 100 may be implemented by a server computer or the like located outside the image processing system.

Claims

1. An image processing system, characterized in that... Include: Photographic image acquisition agencies acquire photographic images taken by mobile photographic equipment in real space. The observation space information acquisition mechanism acquires observation space information containing the three-dimensional coordinates of the feature point group in the observation space based on the positional changes of the feature point group in the photographic image. A machine learning mechanism acquires additional information related to the features of the photographic object shown in the photographic image, based on machine learning data related to object features. as well as The integration mechanism integrates the observed spatial information with the additional information; The additional information is information related to the three-dimensional shape of the photographed object inferred from the machine learning data, and is information related to the mesh of the photographed object; The integration mechanism sets the grid in the observation space based on the additional information, and changes the grid based on the observation space information.

2. The image processing system according to claim 1, characterized in that: After changing the portion of the grid corresponding to the three-dimensional coordinates of the feature point group shown in the observation spatial information, the integration mechanism changes the surrounding portion of the grid.

3. The image processing system according to claim 1, characterized in that: The observation space information acquisition mechanism infers the position of the photography mechanism based on the positional changes of the feature point group, and sets the observation viewpoint in the observation space based on the inference result. The integration mechanism changes the grid portion based on the orientation of each grid portion relative to the observation viewpoint.

Citation Information

Patent Citations

  • Position measuring device, position measuring method and position measuring program

    CN101149252A

  • Method for measuring gangue dump surface temperature field through close-range photogrammetry and thermal infrared imager

    CN102927971A