Image processing device, image processing method, and program

The image processing device and method address the challenge of mapping high-resolution three-dimensional data to two-dimensional images by converting and interpolating pixel positions, resulting in high-definition image generation.

JP7815244B2Active Publication Date: 2026-02-17FUJIFILM CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2023531736
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-06-29
Filing Date
2022-06-06
Publication Date
2026-02-17
Estimated Expiration
2042-06-06

AI Technical Summary

Technical Problem

Existing image processing technologies struggle to create high-definition images by accurately mapping three-dimensional data from three-dimensional ranging sensors to two-dimensional images, particularly when the sampling resolution of imaging devices is higher than the ranging sensors, leading to challenges in precise pixel assignment and image representation.

Method used

An image processing device and method that utilizes a processor to convert three-dimensional coordinates from a three-dimensional ranging sensor into two-dimensional coordinates of an imaging device, using interpolation methods to assign pixels accurately within a display system, thereby generating high-definition images by aligning three-dimensional data with two-dimensional image frames.

Benefits of technology

This approach enables the creation of high-definition images by precisely mapping three-dimensional data onto two-dimensional displays, enhancing image clarity and accuracy through pixel interpolation based on known positional relationships between sensors and imaging devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007815244000020
    Figure 0007815244000020
  • Figure 0007815244000021
    Figure 0007815244000021
  • Figure 0007815244000022
    Figure 0007815244000022
Patent Text Reader

Abstract

A processor of this image processing device acquires three-dimensional ranging sensor system three-dimensional coordinates with which the positions of a plurality of measurement points can be specified and which are defined in a three-dimensional coordinate system applied to a three-dimensional ranging sensor, acquires imaging device system three-dimensional coordinates on the basis of the three-dimensional ranging sensor system three-dimensional coordinates, converts the imaging device system three-dimensional coordinates to imaging device system two-dimensional coordinates with which a position in a captured image can be specified, converts the imaging device system three-dimensional coordinates to display system two-dimensional coordinates with which a position in a screen image can be specified, and allocates pixels constituting the captured image to an interpolation position in the screen image, the interpolation position being specified according to an interpolation method using the imaging device system two-dimensional coordinates and the display system two-dimensional coordinates.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to an image processing device, an image processing method, and a program. [Background technology]

[0002] Japanese Patent Laid-Open Publication No. 2012-220471 discloses a development drawing generating device that includes: a storage means for storing an image obtained by photographing a tunnel wall surface; and measurement point data having coordinate values ​​and reflection intensity values ​​of a plurality of measurement points on the wall surface obtained by laser scanning the wall surface; a conversion means for performing coordinate conversion for arranging the plurality of measurement points on the wall surface on a development drawing of the wall surface; a comparison means for aligning the image with the coordinate-converted plurality of measurement points based on the reflection intensity values; a displaced shape generating means for generating a displaced shape obtained by adding a concave-convex shape reflecting coordinate values ​​in a direction perpendicular to the development plane of the wall surface based on the coordinates of the coordinate-converted plurality of measurement points; and a drawing means for drawing a pattern of the aligned image on the displaced shape.

[0003] Japanese Patent Application Laid-Open Publication No. 2017-106749 discloses a point cloud data acquisition system that acquires point cloud data having depth information to each point on the surface of a subject in pixel units, and that includes a first electronic device group having a first electronic device and one or more movable second electronic devices.

[0004] In the point cloud data acquisition system described in JP 2017-106749 A, the second electronic device has a second coordinate system and includes a plurality of predetermined landmarks having visual characteristics, each of the plurality of landmarks being provided at linearly independent positions in the second coordinate system, and a three-dimensional measurement unit that measures point cloud data of the subject based on the second coordinate system. Also, in the point cloud data acquisition system described in JP 2017-106749 A, the first electronic device has a first coordinate system and includes a position where the subject is measured based on the first coordinate system. informationThe point cloud data acquisition system includes a measurement unit, and further includes a marker position information calculation unit having a reference coordinate system and calculating second electronic device position information, which is the coordinate values ​​in the first coordinate system of multiple markers included in the second electronic device, based on data measured by the position information measurement unit; a point cloud data coordinate value conversion unit that converts the coordinate values ​​in the second coordinate system of each point data in the point cloud data measured by the three-dimensional measurement unit included in the second electronic device, into coordinate values ​​in the reference coordinate system, based on the second electronic device position information; and a composite point cloud data creation unit that creates one composite point cloud data from the multiple point cloud data measured by the three-dimensional measurement unit included in the second electronic device, based on the coordinate values ​​in the reference coordinate system converted by the point cloud data coordinate value conversion unit.

[0005] Japanese Patent Application Laid-Open Publication No. 2018-025551 discloses a technology that includes a first electronic device equipped with a depth camera that measures point cloud data of a subject based on a depth camera coordinate system, and a second electronic device equipped with a non-depth camera that acquires two-dimensional image data of the subject based on a non-depth camera coordinate system, and converts the two-dimensional image data of the subject into point cloud data by associating the two-dimensional image data of the subject acquired by the non-depth camera with point cloud data in which the coordinate values ​​in the depth camera coordinate system of each point on the surface of the subject measured by the depth camera are converted into coordinate values ​​in the non-depth camera coordinate system.

[0006] Patent No. 4543820 discloses a three-dimensional data processing device equipped with grid coordinate calculation means comprising: intermediate interpolation means for interpolating auxiliary intermediate points for each line data of three-dimensional vector data including a plurality of line data; TIN generation means for forming an irregular triangulated network from the three-dimensional coordinate values ​​of the points describing each line data of the three-dimensional vector data and the auxiliary intermediate points, and generating TIN (Triangulated Irregular Network) data that defines each triangle; grid coordinate calculation means for applying a grid with a predetermined grid spacing to the TIN data to calculate the coordinate values ​​of each grid point from the TIN data and outputting grid data indicating the three-dimensional coordinates of each grid point; and grid coordinate calculation means comprising means for searching for a TIN included in the TIN data and coordinate value calculation means for calculating the coordinate values ​​of the grid points included in the searched TIN. In the three-dimensional data processing device described in Patent No. 4543820, the coordinate value calculation means determines the maximum grid range that can be accommodated in a rectangle circumscribing the TIN, searches for grid points contained in the TIN searched within the maximum grid range, and calculates the coordinate values ​​of each of the searched grid points from the coordinate values ​​of the three vertices of the TIN.

[0007] "As-is 3D Thermal Modeling for Existing Building Envelopes Using a Hybrid LIDAR System," Chao Wang, Yong K. Cho, Journal of Computing in Civil Engineering, 2013, 27, 645-656 (hereinafter referred to as "Non-Patent Document 1") discloses a method for combining measured temperature data with the three-dimensional geometry of a building. The method described in Non-Patent Document 1 uses a function for virtually representing the energy efficiency and environmental impact of an existing building. The method described in Non-Patent Document 1 also provides visual information by using a hybrid LiDAR (Light Detection and Ranging) system, which encourages building renovations.

[0008] In the method described in Non-Patent Document 1, a 3D temperature model stores geometric point cloud data of a building together with temperature data for each point (temperature and temperature color information generated based on the temperature). Because LiDAR cannot collect geometric data from transparent objects, it is necessary to create virtual temperature vertices for window glass. Therefore, the method described in Non-Patent Document 1 uses an algorithm to identify windows. The service framework proposed in Non-Patent Document 1 consists of the following elements: (1) A hybrid 3D LiDAR system that simultaneously collects point clouds and temperatures from the exterior walls of real buildings. (2) Automatic synthesis of the collected point clouds and temperature data. (3) A window detection algorithm to compensate for the inability of LiDAR to detect window glass. (4) Graphical user interface (GUI) rendering. (5) A web-based floor plan program for deciding on repairs, etc.

[0009] Non-Patent Document 1 describes that the LiDAR and IR (Infrared Rays) camera are fixed to the base of a PTU (Pan and Tilt Unit) (see Fig. 5), and that the distortion of the IR camera is pre-calibrated in the imaging system. Non-Patent Document 1 also describes the use of a black and white inspection board (see Fig. 7). Non-Patent Document 1 also states that calibration of intrinsic and extrinsic parameters is required to combine the point cloud from the LiDAR and the image acquired by the IR camera (see Fig. 6). The intrinsic parameters refer to the focal length, lens principal point, skew distortion coefficient, and distortion coefficient, while the extrinsic parameters refer to the rotation and translation matrices. Summary of the Invention

[0010] One embodiment of the technique of the present disclosure provides an image processing device, an image processing method, and a program that can contribute to creating high-definition images on a screen. [Means for solving the problem]

[0011] A first aspect of the technology disclosed herein is an image processing device that includes a processor and a memory connected to or built into the processor, and in which the positional relationship between a three-dimensional ranging sensor and an imaging device having a higher sampling resolution than the three-dimensional ranging sensor is known, wherein the processor acquires three-dimensional ranging sensor system coordinates that can identify the positions of multiple measurement points based on ranging results by the three-dimensional ranging sensor for multiple measurement points, the three-dimensional ranging sensor system coordinates being defined in a three-dimensional coordinate system applied to the three-dimensional ranging sensor, acquires imaging device system coordinates defined in a three-dimensional coordinate system applied to the imaging device based on the three-dimensional ranging sensor system coordinates, converts the imaging device system three-dimensional coordinates into imaging device system two-dimensional coordinates that can identify a position within an captured image obtained by capturing an image with the imaging device, converts the imaging device system three-dimensional coordinates into display system two-dimensional coordinates that can identify a position within a screen, and assigns pixels that constitute the captured image to interpolated positions within the screen identified by an interpolation method using the imaging device system two-dimensional coordinates and the display system two-dimensional coordinates.

[0012] A second aspect of the technology disclosed herein is an image processing device according to the first aspect, in which the processor generates a polygonal patch based on the three-dimensional coordinates of the imaging device system, and the three-dimensional coordinates of the imaging device system define the positions of the intersection points of the polygonal patch.

[0013] A third aspect of the technology of the present disclosure is the image processing device according to the second aspect, in which the interpolation position is a position within the screen that corresponds to a position other than an intersection point of the polygonal patches.

[0014] A fourth aspect of the technology of the present disclosure is an image processing device according to the second or third aspect, in which a processor creates a three-dimensional image by further assigning pixels at positions corresponding to the positions of intersections of polygonal patches among a plurality of pixels included in a captured image to intersection positions within the screen that correspond to the positions of the intersections of the polygonal patches.

[0015] A fifth aspect of the technology of the present disclosure is the image processing device according to any one of the second to fourth aspects, in which the polygonal patch is defined by a triangular mesh or a quadrilateral mesh.

[0016] A sixth aspect of the technology disclosed herein is an image processing device according to any one of the first to fifth aspects, in which a processor acquires three-dimensional coordinates of an imaging device system by converting three-dimensional coordinates of a three-dimensional ranging sensor system into three-dimensional coordinates of an imaging device system.

[0017] A seventh aspect of the technology of the present disclosure is an image processing device according to any one of the first to sixth aspects, in which a processor obtains three-dimensional coordinates of an imaging device system by calculating the three-dimensional coordinates of the imaging device system based on feature points of the subject contained as images between multiple frame images obtained by imaging the subject from different positions using an imaging device.

[0018] An eighth aspect of the technique of the present disclosure is the image processing device according to any one of the first to seventh aspects, in which the position within the screen is a position within the screen of a display.

[0019] A ninth aspect of the technology of the present disclosure is an image processing device according to any one of the first to eighth aspects, in which the memory holds correspondence information that associates three-dimensional coordinates of a three-dimensional ranging sensor system with two-dimensional coordinates of an imaging device system.

[0020] A tenth aspect according to the technique of the present disclosure is the image processing device according to the ninth aspect, in which the processor refers to the association information and allocates pixels that make up the captured image to the interpolation positions.

[0021] An eleventh aspect of the technology of the present disclosure is an image processing method including: acquiring three-dimensional coordinates of a three-dimensional ranging sensor system that can identify the positions of multiple measurement points based on ranging results by the three-dimensional ranging sensor for multiple measurement points, under the condition that the positional relationship between the three-dimensional ranging sensor and an imaging device having a higher sampling resolution than the three-dimensional ranging sensor is known; acquiring three-dimensional coordinates of an imaging device system that are defined in a three-dimensional coordinate system applied to the three-dimensional ranging sensor, based on the three-dimensional coordinates of the three-dimensional ranging sensor system; converting the three-dimensional coordinates of the imaging device system into two-dimensional coordinates of an imaging device system that can identify a position within an captured image obtained by capturing an image by the imaging device; converting the three-dimensional coordinates of the imaging device system into two-dimensional coordinates of a display system that can identify a position within a screen; and assigning pixels that constitute the captured image to interpolated positions within the screen that are identified by an interpolation method based on the two-dimensional coordinates of the imaging device system and the two-dimensional coordinates of the display system.

[0022] A twelfth aspect of the technology disclosed herein is a program for causing a computer to execute processing including: acquiring three-dimensional coordinates of a three-dimensional ranging sensor system that can identify the positions of multiple measurement points based on the ranging results of the three-dimensional ranging sensor for multiple measurement points, under the condition that the positional relationship between the three-dimensional ranging sensor and an imaging device having a higher sampling resolution than the three-dimensional ranging sensor is known; acquiring three-dimensional coordinates of an imaging device system that are defined in a three-dimensional coordinate system applied to the three-dimensional ranging sensor based on the three-dimensional coordinates of the three-dimensional ranging sensor system; converting the three-dimensional coordinates of the imaging device system into two-dimensional coordinates of an imaging device system that can identify a position within an captured image obtained by capturing an image with the imaging device; converting the three-dimensional coordinates of the imaging device system into two-dimensional coordinates of a display system that can identify a position within a screen; and assigning pixels that constitute the captured image to interpolated positions within the screen that are identified by an interpolation method using the two-dimensional coordinates of the imaging device system and the two-dimensional coordinates of the display system. [Brief explanation of the drawings]

[0023] [Figure 1] 1 is a schematic diagram illustrating an example of the overall configuration of a mobile system. [Figure 2] FIG. 2 is a schematic perspective view showing an example of detection axes of an acceleration sensor and an angular velocity sensor. [Figure 3] FIG. 2 is a block diagram illustrating an example of a hardware configuration of the information processing system. [Figure 4] FIG. 2 is a block diagram showing an example of main functions of a processor. [Figure 5] FIG. 10 is a conceptual diagram illustrating an example of processing content of an acquisition unit. [Figure 6] FIG. 1 is a conceptual diagram showing an example of how a LiDAR coordinate system is transformed into a camera three-dimensional coordinate system. [Figure 7] FIG. 10 is a conceptual diagram illustrating an example of processing content of an acquisition unit. [Figure 8] FIG. 10 is a conceptual diagram illustrating an example of processing content of a conversion unit. [Figure 9] FIG. 10 is a conceptual diagram showing an example of a manner in which camera three-dimensional coordinates are perspectively projected onto an xy plane and a uv plane. [Figure 10] FIG. 10 is a conceptual diagram illustrating an example of processing content of a pixel allocation unit. [Figure 11] 10 is a flowchart showing an example of the flow of a texture mapping process. [Figure 12] This is a comparison diagram comparing an example in which pixels at positions corresponding only to TIN intersections are assigned to the screen, and an example in which pixels at positions corresponding to TIN intersections and pixels at positions other than TIN intersections are assigned to the screen. [Figure 13] 10 is a conceptual diagram showing an example of how pixels at the center of gravity or the center of gravity part of the camera two-dimensional coordinate system are assigned to the screen two-dimensional coordinate system. FIG. [Figure 14] FIG. 10 is a conceptual diagram showing an example of how pixels on a side of a camera two-dimensional coordinate system are assigned to a screen two-dimensional coordinate system. [Figure 15] FIG. 10 is a block diagram showing an example of main functions of a processor according to a first modified example. [Figure 16]10 is a flowchart showing an example of the flow of texture mapping processing according to a first modified example. [Figure 17] FIG. 1 is a conceptual diagram showing an example of an aspect in which the same subject is imaged by a camera on a moving body while changing the image capturing position. [Figure 18] FIG. 10 is a conceptual diagram showing an example of processing contents of a processor according to a second modified example. [Figure 19] FIG. 1 is a conceptual diagram illustrating epipolar geometry. [Figure 20] FIG. 10 is a conceptual diagram showing an example of how texture mapping processing is installed in a computer from a storage medium. DETAILED DESCRIPTION OF THE INVENTION

[0024] Hereinafter, an example of an embodiment of an image processing device, an image processing method, and a program according to the technique of the present disclosure will be described with reference to the accompanying drawings.

[0025] First, the terms used in the following description will be explained.

[0026] CPU is an abbreviation for "Central Processing Unit". GPU is an abbreviation for "Graphics Processing Unit". RAM is an abbreviation for "Random Access Memory". IC is an abbreviation for "Integrated Circuit". ASIC is an abbreviation for "Application Specific Integrated Circuit". PLD is an abbreviation for "Programmable Logic Device". FPGA is an abbreviation for "Field-Programmable Gate Array". SoC is an abbreviation for "System-on-a-chip". SSD is an abbreviation for "Solid State Drive". USB is an abbreviation for "Universal Serial Bus". HDD is an abbreviation for "Hard Disk Drive". EL is an abbreviation for "Electro-Luminescence". I / F is an abbreviation for "Interface". UI is an abbreviation for "User Interface". CMOS is an abbreviation for "Complementary Metal Oxide Semiconductor". CCD is an abbreviation for "Charge Coupled Device." LiDAR is an abbreviation for "Light Detection and Ranging." TIN is an abbreviation for "Triangulated Irregular Network." In this specification, an intersection refers to a point where two adjacent sides of a polygon intersect (i.e., a vertex). Also, in this specification, a position other than an intersection of a polygon or polygon patch refers to a position inside the polygon or polygon patch (within the polygon). The inside of a polygon (within the polygon) also includes the sides of the polygon.

[0027] As an example, as shown in Fig. 1, an information processing system 2 includes a mobile body 10 and an information processing device 20. The mobile body 10 is equipped with a sensor unit 30. An example of the mobile body 10 is an unmanned mobile body. In the example shown in Fig. 1, an unmanned aerial vehicle (e.g., a drone) is shown as an example of the mobile body 10.

[0028] The mobile object 10 is used for surveying and / or inspecting land and / or infrastructure, etc. Examples of infrastructure include road facilities (e.g., bridges, road surfaces, tunnels, guardrails, traffic lights, and / or windbreak fences), waterway facilities, airport facilities, port facilities, water storage facilities, gas facilities, power supply facilities, medical facilities, and / or firefighting facilities.

[0029] Here, an unmanned aerial vehicle is given as an example of the mobile body 10, but the technology of the present disclosure is not limited thereto. For example, the mobile body 10 may be a vehicle. Examples of vehicles include a vehicle with a gondola, a vehicle for working at height, or a bridge inspection vehicle. The mobile body 10 may also be a slider or a dolly on which the sensor unit 30 can be mounted. The mobile body 10 may also be a person. Here, the person refers to, for example, a worker who carries the sensor unit 30 and operates the sensor unit 30 to survey and / or inspect land and / or infrastructure, etc.

[0030] The information processing device 20 is a notebook personal computer. While a notebook personal computer is illustrated here, this is merely an example and the information processing device 20 may be a desktop personal computer. Furthermore, the information processing device 20 is not limited to a personal computer and may be a server. The server may be a mainframe used on-premise together with the mobile object 10, or an external server realized by cloud computing. Furthermore, the server may be an external server realized by network computing such as fog computing, edge computing, or grid computing.

[0031] The information processing device 20 includes a reception device 22 and a display 24. The reception device 22 has a keyboard, a mouse, a touch panel, etc., and receives instructions from a user. The display 24 displays various information (e.g., images, text, etc.). The display 24 may be, for example, an EL display (e.g., an organic EL display or an inorganic EL display). Note that the display is not limited to an EL display, and may be another type of display such as a liquid crystal display.

[0032] The mobile object 10 is connected to an information processing device 20 so as to be able to communicate wirelessly, and various information is exchanged between the mobile object 10 and the information processing device 20 wirelessly.

[0033] The moving body 10 includes a main body 12 and a plurality of propellers 14 (four propellers in the example shown in FIG. 1). By controlling the rotation of each of the plurality of propellers 14, the moving body 10 flies or hovers in three-dimensional space.

[0034] A sensor unit 30 is attached to the main body 12. In the example shown in Fig. 1, the sensor unit 30 is attached to the upper part of the main body 12. However, this is merely an example, and the sensor unit 30 may be attached to a location other than the upper part of the main body 12 (for example, the lower part of the main body 12).

[0035] The sensor unit 30 includes an external sensor 32 and an internal sensor 34. The external sensor 32 senses the external environment of the moving body 10. The external sensor 32 includes a LiDAR 32A and a camera 32B. The LiDAR 32A is an example of a "three-dimensional ranging sensor" according to the technology of the present disclosure, and the camera 32B is an example of an "imaging device" according to the technology of the present disclosure.

[0036] The LiDAR 32A scans the surrounding space by emitting a pulsed laser beam L. The laser beam L is, for example, visible light or infrared light. The LiDAR 32A receives the reflected light of the laser beam L reflected from an object (e.g., a natural object and / or an artificial object) present in the surrounding space, and calculates the distance to a measurement point on the object by measuring the time from emitting the laser beam L to receiving the reflected light. Here, the measurement point refers to the reflection point of the laser beam L on the object. Furthermore, each time the LiDAR 32A scans the surrounding space, it outputs point cloud data representing multiple three-dimensional coordinates as position information that can identify the positions of multiple measurement points. The point cloud data is also called a point cloud. The point cloud data is, for example, data expressed in three-dimensional Cartesian coordinates.

[0037] The LiDAR 32A emits a laser beam L in a field of view S1 that is, for example, 135 degrees left and right and 15 degrees up and down based on the traveling direction of the moving object 10. The LiDAR 32A emits the laser beam L over the entire field of view S1 while changing the angle, for example, by 0.25 degrees in either the left and right or up and down directions. The LiDAR 32A repeatedly scans the field of view S1 and outputs point cloud data for each scan. For ease of explanation, the point cloud data output by the LiDAR 32A for each scan will be referred to as segmented point cloud data PG (see FIG. 3 ).

[0038] Camera 32B is a digital camera having an image sensor. Here, the image sensor refers to, for example, a CMOS image sensor or a CCD image sensor. Camera 32B captures an image of field of view S2 in accordance with instructions given from the outside (for example, information processing device 20). Examples of image capture by camera 32B include image capture for moving images (for example, image capture performed at a predetermined frame rate such as 30 frames / second or 60 frames / second) and image capture for still images. Image capture for moving images and image capture for still images by camera 32B are selectively performed in accordance with instructions given to camera 32B from the outside.

[0039] The sampling resolution of the camera 32B (for example, the density of pixels per unit area of ​​a captured image obtained by capturing an image with the camera 32B) is higher than the sampling resolution of the LiDAR 32A (for example, the density of ranging points per unit area of ​​point cloud data). In addition, in the information processing device 20, the positional relationship between the LiDAR 32A and the camera 32B is known.

[0040] For ease of explanation, the following description will be given on the assumption that the LiDAR 32A and the camera 32B are fixed relative to the moving body 10. For ease of explanation, the following description will be given on the assumption that the field of view S2 is included in the field of view S1 (i.e., the field of view S2 is part of the field of view S1), the position of the field of view S2 relative to the field of view S1 is fixed, and the position of the field of view S2 relative to the field of view S1 is known. For ease of explanation, the field of view S2 is included in the field of view S1, but this is merely an example. For example, the field of view S1 may be included in the field of view S2 (i.e., the field of view S1 may be part of the field of view S2). Specifically, the field of view S1 and the field of view S2 only need to overlap to such an extent that a plurality of segmented point cloud data PG and a captured image that can be used to perform the texture mapping process (see FIG. 11) described below can be obtained.

[0041] The internal sensor 34 has an acceleration sensor 36 and an angular velocity sensor 38. The internal sensor 34 detects physical quantities required to identify the moving direction, moving distance, attitude, etc. of the mobile object 10, and outputs detection data indicating the detected physical quantities. The detection data by the internal sensor 34 (hereinafter also simply referred to as "detection data") is, for example, acceleration data indicating the acceleration detected by the acceleration sensor 36 and acceleration data indicating the angular velocity detected by the angular velocity sensor 38.

[0042] The moving object 10 has a LiDAR coordinate system and a camera three-dimensional coordinate system. The LiDAR coordinate system is a three-dimensional coordinate system (here, as an example, a Cartesian coordinate system in three-dimensional space) applied to the LiDAR 32A. The camera three-dimensional coordinate system is a three-dimensional coordinate system (here, as an example, a Cartesian coordinate system in three-dimensional space) applied to the camera 32B.

[0043] As an example, as shown in FIG. 2, a three-dimensional Cartesian coordinate system applied to the LiDAR 32A, i.e., the LiDAR coordinate system, is a system of X coordinates that are orthogonal to each other. L Axis, Y L axis and Z L The point cloud data is defined by the LiDAR coordinate system.

[0044] The acceleration sensor 36 (see Figure 1) L Axis, Y L axis and Z L The angular velocity sensor 38 (see Figure 1) detects the acceleration in each direction of the axis. L Axis, Y L axis and Z L The internal sensor 34 detects angular velocities around each axis (i.e., in the roll direction, pitch direction, and yaw direction). That is, the internal sensor 34 is a six-axis inertial measurement sensor.

[0045] As an example, as shown in Fig. 3, the main body 12 of the moving body 10 is provided with a controller 16, a communication I / F 18, and a motor 14A. The controller 16 is realized by, for example, an IC chip. The main body 12 is provided with a plurality of motors 14A. The plurality of motors 14A are connected to a plurality of propellers 14. The controller 16 controls the plurality of motors 14A to control the flight of the moving body 10. The controller 16 also controls the scanning operation of the laser beam L by the LiDAR 32A.

[0046] As a result of sensing an object 50, which is an example of the external environment, the external sensor 32 outputs to the controller 16 segmented point cloud data PG obtained by the LiDAR 32A scanning a laser beam L over a field of view S1, and a captured image PD obtained by the camera 32B capturing an image of subject light indicating the field of view S2 (i.e., reflected light indicating a portion of the object 50 included in the field of view S2). The controller 16 receives the segmented point cloud data PG and the captured image PD from the external sensor 32.

[0047] The internal sensor 34 outputs detection data obtained by performing sensing (for example, acceleration data from the acceleration sensor 36 and angular velocity data from the angular velocity sensor 38) to the controller 16. The controller 16 receives the detection data from the internal sensor 34.

[0048] The controller 16 wirelessly transmits the received segmented point cloud data PG, captured image PD, and detection data to the information processing device 20 via the communication I / F 18.

[0049] The information processing device 20 includes a computer 39 and a communication I / F 46 in addition to the reception device 22 and the display 24. The computer 39 has a processor 40, a storage 42, and a RAM 44. The reception device 22, the display 24, the processor 40, the storage 42, the RAM 44, and the communication I / F 46 are connected to a bus 48. The information processing device 20 is an example of an "image processing device" according to the technology of the present disclosure. The computer 39 is an example of a "computer" according to the technology of the present disclosure. The processor 40 is an example of a "processor" according to the technology of the present disclosure. The storage 42 and the RAM 44 are an example of a "memory" according to the technology of the present disclosure.

[0050] The processor 40 includes, for example, a CPU and a GPU, and controls the entire information processing device 20. The GPU operates under the control of the CPU and is responsible for screen display and / or image processing, etc. The processor 40 may be one or more CPUs that have integrated GPU functionality, or one or more CPUs that do not have integrated GPU functionality.

[0051] The storage 42 is a non-volatile storage device that stores various programs, various parameters, etc. Examples of the storage 42 include an HDD and an SSD. Note that the HDD and SSD are merely examples, and a flash memory, a magnetoresistive memory, and / or a ferroelectric memory may be used instead of the HDD and / or SSD, or together with the HDD and / or SSD.

[0052] The RAM 44 is a memory that temporarily stores information and is used as a work memory by the processor 40. The RAM 44 may be, for example, a DRAM and / or an SRAM.

[0053] The communication I / F 46 receives a plurality of segmented point cloud data PG from the mobile object 10 by wirelessly communicating with the communication I / F 18 of the mobile object 10. The plurality of segmented point cloud data PG received by the communication I / F 46 refers to a plurality of segmented point cloud data PG acquired by the LiDAR 32A at different times (i.e., a plurality of segmented point cloud data PG obtained by a plurality of scans). Furthermore, the communication I / F 46 receives the captured image PD and detection data acquired by the internal sensor 34 for each time each of the plurality of segmented point cloud data PG is acquired by wirelessly communicating with the communication I / F 18 of the mobile object 10. The segmented point cloud data PG, captured image PD, and detection data received by the communication I / F 46 in this manner are acquired and processed by the processor 40.

[0054] The processor 40 acquires composite point cloud data SG based on the plurality of segmented point cloud data PG received from the moving object 10. Specifically, the processor 40 generates the composite point cloud data SG by performing a synthesis process to synthesize the plurality of segmented point cloud data PG received from the moving object 10. The composite point cloud data SG is a collection of the plurality of segmented point cloud data PG obtained by scanning the field of view range S1, and is stored in the storage 42 by the processor 40. The plurality of segmented point cloud data PG is an example of "ranging results of a plurality of measurement points by a 3D ranging sensor" according to the technology of the present disclosure. The composite point cloud data SG is also an example of "3D coordinates of a 3D ranging sensor system" according to the technology of the present disclosure.

[0055] As an example, as shown in FIG. 4, a texture mapping processing program 52 is stored in the storage 42. The texture mapping processing program 52 is an example of a "program" according to the technology of the present disclosure. The processor 40 reads the texture mapping processing program 52 from the storage 42 and executes the read texture mapping processing program 52 on the RAM 44. The processor 40 performs texture mapping processing in accordance with the texture mapping processing program 52 executed on the RAM 44 (see FIG. 11). By executing the texture mapping processing program 52, the processor 40 operates as an acquisition unit 40A, a conversion unit 40B, and a pixel allocation unit 40C.

[0056] As an example, as shown in FIG. 5, the acquisition unit 40A acquires composite point cloud data SG from the storage 42. Then, the acquisition unit 40A acquires camera three-dimensional coordinates defined in a camera three-dimensional coordinate system based on the composite point cloud data SG. Here, the acquisition of the camera three-dimensional coordinates by the acquisition unit 40A is realized by converting the composite point cloud data SG into the camera three-dimensional coordinates by the acquisition unit 40A. Here, the camera three-dimensional coordinates are an example of "imaging device system three-dimensional coordinates" according to the technology of the present disclosure.

[0057] As an example, as shown in Fig. 6, the transformation of the composite point cloud data SG into the camera three-dimensional coordinates is realized by transforming the LiDAR coordinate system into the camera three-dimensional coordinate system according to a rotation matrix and a translation vector. The example shown in Fig. 6 shows an example of how three-dimensional coordinates (hereinafter also referred to as "LiDAR coordinates"), which are data specifying the position of one measurement point P in the composite point cloud data SG, are transformed into the camera three-dimensional coordinates.

[0058] Here, the three axes of the LiDAR coordinate system are X L Axis, Y L axis and Z L axis, and the rotation matrix required to transform the LiDAR coordinate system into the camera 3D coordinate system is C L R, the moving object 10 is located in the LiDAR coordinate system along with the X L Around the axis, Y in the LiDAR coordinate system L axis and Z of the LiDAR coordinate system L Rotation matrices representing the transformation of the position of the moving body 10 when rotated by angles φ, θ, and ψ around the axes, respectively. C L R is expressed by the following matrix (1): Note that the angles φ, θ, and ψ are calculated based on the angular velocity data included in the detection data.

[0059]

number

[0060] Then, set the origin of the LiDAR coordinate system as O L The origin of the camera's 3D coordinate system is O C The three axes of the camera's three-dimensional coordinate system are X C Axis, Y C axis and Z C axis, and the 3D coordinates of the measurement point P in the LiDAR coordinate system (i.e., the LiDAR 3D coordinates) are L Let P be the 3D coordinate of the measurement point P in the camera 3D coordinate system (i.e., the camera 3D coordinate). C Let P be the translation vector required to transform the LiDAR coordinate system into the camera 3D coordinate system.C L T and the origin O L The position of in the LiDAR coordinate system L O C When expressed as: camera 3D coordinates C P is expressed by the following equation (2). C L T is calculated based on the acceleration data included in the detection data.

[0061]

number

[0062] where: C L T, C L R, and L O C The relationship is expressed by the following formula (3), so by substituting formula (3) into formula (2), the camera's three-dimensional coordinates are C P is expressed by the following formula (4): The synthesized point cloud data SG including the LiDAR coordinates for the measurement point P is converted into three-dimensional coordinates of multiple cameras by using formula (4).

[0063]

number

[0064]

number

[0065] As an example, as shown in FIG. 7, the acquisition unit 40A generates a TIN 54 based on the three-dimensional coordinates of the multiple cameras, in order to represent the three-dimensional coordinates of the multiple cameras obtained by converting the composite point cloud data SG using Equation (4) in a digital data structure. The TIN 54 is a digital data structure representing a collection of triangular patches 54A defined by a triangular mesh (e.g., an irregular triangular network). The three-dimensional coordinates of the multiple cameras define the positions of the intersections of the multiple triangular patches 54A included in the TIN 54 (in other words, the vertices of each triangular patch 54A). Note that the triangular patch 54A is an example of a "polygonal patch" according to the technology disclosed herein. While the triangular patch 54A is illustrated here, this is merely an example, and patches of polygonal shapes other than triangles may also be used. In other words, instead of the TIN 54, a polygonal set patch having a planar structure in which multiple polygonal patches are set together may be used.

[0066] As an example, as shown in FIG. 8, the conversion unit 40B converts the multiple camera three-dimensional coordinates used in the TIN 54 into multiple two-dimensional coordinates (hereinafter also referred to as "camera two-dimensional coordinates") that can identify a position within the captured image PD. The conversion unit 40B also converts the multiple camera three-dimensional coordinates used in the TIN 54 into multiple two-dimensional coordinates (hereinafter also referred to as "screen two-dimensional coordinates") that can identify a position within the screen of the display 24. Here, the camera two-dimensional coordinates are an example of "imaging device system two-dimensional coordinates" according to the technology of the present disclosure, and the screen two-dimensional coordinates are an example of "display system two-dimensional coordinates" according to the technology of the present disclosure.

[0067] As an example, as shown in FIG. 9, the positions of the pixels constituting the captured image PD are determined by the origin O in the xy plane PD0 (in the example shown in FIG. 9) corresponding to the imaging surface of the image sensor of the camera 32B (see FIGS. 1 and 3). C From Z CThe xy plane PD0 is a plane that can be expressed in a two-dimensional coordinate system defined by the x-axis and y-axis (hereinafter also referred to as the "camera two-dimensional coordinate system"), and the position of each pixel that makes up the captured image PD is specified by the camera two-dimensional coordinates (x, y). Here, when i=0, 1, 2, . . . , n, the position A of a pixel in the xy plane PD0 is i , A i+1 and A i+2 are the positions P of the three vertices of the triangular patch 54A. i , P i+1 and P i+2 Position A corresponds to i , A i+1 and A i+2 The camera's two-dimensional coordinates (x0, y0), (x1, y1) and (x2, y2) are the coordinates of the position P i , P i+1 and P i+2 The 3D coordinates of each camera are obtained by perspective projection onto the xy plane PD0.

[0068] In the example shown in Figure 9, the origin is O d and X d Axis, Y d axis and Z d A display three-dimensional coordinate system having axes is shown. The display three-dimensional coordinates are three-dimensional coordinates applied to the display 24 (see FIGS. 1 and 3). A screen 24A is set in the display three-dimensional coordinate system. The screen 24A is the screen of the display 24, and an image showing an object 50 is displayed on the screen 24A. For example, the image displayed on the screen 24A is a texture image. Here, the texture image refers to an image obtained by texture mapping a captured image PD obtained by capturing an image with the camera 32B onto the screen 24A.

[0069] The position and orientation of the screen 24A set with respect to the display three-dimensional coordinate system are changed, for example, in accordance with an instruction received by the reception device 22 (see FIG. 3). The position of each pixel constituting the screen 24A is determined by the uv plane 24A1 (in the example shown in FIG. 9, the origin O d From Zd The uv plane 24A1 is a plane that can be expressed in a two-dimensional coordinate system defined by the u axis and the v axis (hereinafter also referred to as the "screen two-dimensional coordinate system"). The position of each pixel that makes up the screen 24A, i.e., its position within the screen of the display 24, is specified by the screen two-dimensional coordinates (u, v).

[0070] In the example shown in FIG. 10, the pixel position B i , B i+1 and B i+2 The position P of the three vertices of the triangular patch 54A i , P i+1 and P i+2 Position B corresponds to i , B i+1 and B i+2 The two-dimensional coordinates (u0, v0), (u1, v1) and (u2, v2) of the screen are the positions P i , P i+1 and P i+2 The three-dimensional coordinates of each camera are obtained by perspectively projecting them onto the uv plane 24A1.

[0071] As an example, as shown in FIG. 10, the pixel allocation unit 40C creates a three-dimensional image (i.e., a texture image perceived three-dimensionally by the user through the screen 24A) by allocating pixels at positions corresponding to the positions of intersections of a plurality of triangular patches 54A (see FIGS. 7 to 9) included in the TIN 54, among the plurality of pixels (all pixels, for example), included in the captured image PD, to the on-screen intersection positions. The on-screen intersection positions refer to positions corresponding to the positions of intersections of a plurality of triangular patches 54A (see FIGS. 7 to 9) included in the TIN 54, among the plurality of pixels (all pixels, for example), included in the screen 24A. In the example shown in FIG. 10, among the plurality of pixels included in the captured image PD, the pixels at position A i , A i+1 and A i+2 (i.e., the positions specified by the camera two-dimensional coordinates (x0, y0), (x1, y1) and (x2, y2)) are located at position B i , B i+1 and Bi+2 (i.e., the positions specified by the two-dimensional screen coordinates (u0, v0), (u1, v1) and (u2, v2)).

[0072] However, even if, among the multiple pixels contained in the captured image PD, pixels at positions corresponding to the positions of the intersections of the multiple triangular patches 54A contained in TIN54 are assigned to corresponding positions in the screen 24A, pixels are not assigned to positions other than the intersections of the multiple triangular patches 54A (in other words, the vertices of the triangular patches 54A), and therefore the pixel density of the image displayed on the screen 24A is reduced accordingly.

[0073] Therefore, the pixel allocation unit 40C allocates pixels constituting the captured image PD to interpolation positions within the screen 24A that are specified by an interpolation method using the camera two-dimensional coordinates and the screen two-dimensional coordinates. Here, the interpolation positions are positions corresponding to positions other than the intersections of the triangular patches 54A (in the example shown in FIG. 10, the second interpolation positions D 0). An example of an interpolation method is linear interpolation. Examples of interpolation methods other than linear interpolation include polynomial interpolation and spline interpolation.

[0074] 10, a first interpolation position C0 is shown as a position corresponding to a position other than the intersection points of the triangular patch 54A among the positions of a plurality of pixels (for example, all pixels) included in the captured image PD. Also, in the example shown in FIG. 10, a second interpolation position D0 is shown as a position corresponding to a position other than the intersection points of the triangular patch 54A among the positions of a plurality of pixels (for example, all pixels) included in the screen 24A. The second interpolation position D0 is a position corresponding to the first interpolation position C0. Also, in the example shown in FIG. 10, (x, y) is shown as the camera two-dimensional coordinates of the first interpolation position C0, and (u, v) is shown as the screen two-dimensional coordinates of the second interpolation position D0.

[0075] The first interpolation position C0 is located at position A i , A i+1 and A i+2The first interpolation position C0 is a position that exists inside a triangle with three vertices, i.e., inside a triangular patch formed in the captured image PD. For example, the first interpolation position C0 is a position that exists between the positions C1 and C2. The position C1 is a position that exists between the position A and the position C2 in the captured image PD. i , A i+1 and A i+2 Position A is one of the three sides of the triangle with three vertices. i and A i+2 Position C2 is a position on the side connecting position A and position C1 in the captured image PD. i , A i+1 and A i+2 Position A is one of the three sides of the triangle with three vertices. i and A i+1 The positions C1 and C2 may be determined according to an instruction received by the reception device 22 (see FIG. 3), or may be determined according to an instruction received by the reception device 22 (see FIG. 3). i and A i+2 The edge connecting and position A i and A i+1 Alternatively, the distance may be determined according to a certain internal division ratio applied to the side connecting the two points.

[0076] The camera's 2D coordinates at position C1 are (x 02 ,y 02 ), and the camera's two-dimensional coordinates at position C2 are (x 01 ,y 01 The first interpolation position C0 is the camera two-dimensional coordinates (x0, y0), (x1, y1), (x2, y2), (x 01 ,y 01 ) and (x 02 ,y 02 ) is determined from the camera's 2D coordinates (x,y) interpolated using

[0077] The second interpolation position D0 is the position B i , B i+1 and B i+2The second interpolation position D0 is a position that exists inside a triangle with three vertices, i.e., inside a triangular patch formed in the screen 24A. For example, the second interpolation position D0 is a position that exists between the positions D1 and D2. The position D1 is a position that exists between the position B i , B i+1 and B i+2 Position B of the three sides of the triangle with three vertices i and B i+2 Position D2 is a position on the side connecting position B i , B i+1 and B i+2 Position B of the three sides of the triangle with three vertices i and B i+1 It should be noted that position D1 corresponds to position C1, and position D2 corresponds to position C2.

[0078] The two-dimensional screen coordinates of position D1 are (u 02 ,v 02 ), and the camera's two-dimensional coordinates at position C2 are (u 01 ,v 01 The second interpolation position D0 is the two-dimensional coordinates (u0, v0), (u1, v1), (u2, v2), (u 01 ,v 01 ) and (u 02 ,v 02 ) is specified from the two-dimensional screen coordinates (u,v) interpolated using

[0079] Of the multiple pixels included in the captured image PD, the pixel at the first interpolation position C0 specified by the camera two-dimensional coordinates (x, y) is assigned to the second interpolation position D0 specified by the screen two-dimensional coordinates (u, v) by the pixel assignment unit 40C. The pixel at the first interpolation position C0 may be the pixel itself that constitutes the captured image PD, but if there is no pixel at the position specified by the camera two-dimensional coordinates (x, y), multiple pixels (e.g., position A) adjacent to the position specified by the camera two-dimensional coordinates (x, y) may be assigned to the second interpolation position D0. i , A i+1 and A i+2 A pixel generated by interpolating the pixels at the three vertices may be used as the pixel at the first interpolation position C0.

[0080] The resolution of the captured image PD (i.e., the resolution of sampling by the camera 32B) is preferably higher than the resolution of the composite point cloud data SG (i.e., the resolution of sampling by the LiDAR 32A) to the extent that a triangle in the captured image PD (i.e., a triangle corresponding to the triangular patch 54A) contains at least a number of pixels determined according to the resolution of the interpolation process (e.g., the process of identifying the first interpolation position C0 by interpolation). Here, the number determined according to the resolution of the interpolation process refers to a number determined in advance so that the higher the resolution of the interpolation process, the greater the number of pixels. The resolution of the interpolation process may be a fixed value determined according to the resolution of the captured image PD and / or the screen 24A, or a variable value that changes according to an external instruction (e.g., an instruction accepted by the accepting device 22). Alternatively, the resolution may be a variable value that changes according to the size (e.g., the average size) of the triangular patch 54A, the triangle obtained by perspectively projecting the triangular patch 54A onto the camera two-dimensional coordinate system, and / or the triangle obtained by perspectively projecting the triangular patch 54A onto the screen two-dimensional coordinate system. When the resolution of the interpolation process depends on the triangular patches 54A, etc., for example, the larger the triangular patches 54A, etc., the higher the resolution of the interpolation process, and the smaller the triangular patches 54A, etc., the lower the resolution of the interpolation process.

[0081] In order to assign the pixel at the first interpolation position C0 (i.e., the pixel at the first interpolation position C0 identified from the camera two-dimensional coordinates (x, y)) to the second interpolation position D0 (i.e., the second interpolation position D0 identified from the screen two-dimensional coordinates (u, v)), the pixel allocation unit 40C calculates the screen two-dimensional coordinates (u, v) of the second interpolation position D0 by an interpolation method using the following equations (5) to (12).

[0082]

number

[0083]

number

[0084]

number

[0085]

number

[0086]

number

[0087]

number

[0088]

number

[0089]

number

[0090] Next, as an operation of the information processing system 2, the texture mapping process performed by the processor 40 of the information processing device 20 will be described with reference to FIG.

[0091] Fig. 11 shows an example of the flow of texture mapping processing performed by processor 40. The flow of texture mapping processing shown in Fig. 11 is an example of an "image processing method" according to the technique of the present disclosure. Note that the following description will be given on the assumption that composite point cloud data SG has already been stored in storage 42.

[0092] 11, first, in step ST100, the acquisition unit 40A acquires the composite point cloud data SG from the storage 42 (see FIG. 5). After the processing of step ST100 is executed, the texture mapping processing proceeds to step ST102.

[0093] In step ST102, the acquisition unit 40A converts the composite point cloud data SG into a plurality of camera three-dimensional coordinates (see FIGS. 5 and 6). After the process of step ST102 is executed, the texture mapping process proceeds to step ST104.

[0094] In step ST104, the acquisition unit 40A generates the TIN 54 based on the three-dimensional coordinates of the multiple cameras (see FIG. 7). After the process of step ST104 is executed, the texture mapping process proceeds to step ST106.

[0095] In step ST106, the conversion unit 40B converts the multiple camera three-dimensional coordinates used in the TIN 54 into multiple camera two-dimensional coordinates (see FIGS. 8 and 9). After the processing of step ST106 is executed, the texture mapping processing proceeds to step ST108.

[0096] In step ST108, the conversion unit 40B converts the multiple camera three-dimensional coordinates used in the TIN 54 into multiple screen two-dimensional coordinates (see FIGS. 8 and 9). After the processing of step ST108 is executed, the texture mapping processing proceeds to step ST110.

[0097] In step ST110, the pixel allocation unit 40C creates a texture image by allocating, to the on-screen intersection positions, pixels among the plurality of pixels included in the captured image PD that correspond to the positions of the intersections of the plurality of triangular patches 54A (see FIGS. 7 to 9) included in the TIN 54. The texture image obtained by executing the processing of step ST110 is displayed on the screen 24A.

[0098] Next, in order to refine the texture image obtained by executing the process of step ST110, the pixel allocation unit 40C performs the process of step ST112 and the process of step ST114.

[0099] In step ST112, the pixel allocation unit 40C calculates an interpolation position in the screen 24A, i.e., the screen two-dimensional coordinates (u, v) of the second interpolation position D0, by using the camera two-dimensional coordinates obtained in step ST106, the screen two-dimensional coordinates obtained in step ST108, and an interpolation method using formulas (5) to (12) (see FIG. 10). After the process of step ST112 is executed, the texture mapping process proceeds to step ST114.

[0100] In step ST114, the pixel allocation unit 40C allocates, to an interpolation position on the screen 24A, i.e., the screen two-dimensional coordinates (u, v) of the second interpolation position D0, a pixel at a camera two-dimensional coordinate (in the example shown in FIG. 10, the camera two-dimensional coordinate (x, y)) corresponding to the screen two-dimensional coordinate (u, v) from among the multiple pixels constituting the captured image PD. As a result, a texture image that is more detailed than the texture image obtained by executing the processing of step ST110 is displayed on the screen 24A. After the processing of step ST114 is executed, the texture mapping processing ends.

[0101] As described above, in the information processing system 2, the composite point cloud data SG is converted into a plurality of camera three-dimensional coordinates (see FIGS. 5 to 7), and the plurality of camera three-dimensional coordinates are converted into a plurality of camera two-dimensional coordinates and a plurality of screen two-dimensional coordinates (see FIGS. 8 and 9). Then, within the screen 24A, pixels constituting the captured image PD are assigned to interpolation positions (in the example shown in FIG. 10, the second interpolation positions D0) identified by an interpolation method using the plurality of camera two-dimensional coordinates and the plurality of screen two-dimensional coordinates (see FIG. 10). This allows a three-dimensional image (i.e., a texture image perceived three-dimensionally by the user) to be displayed within the screen 24A. Therefore, this configuration contributes to the creation of a higher-resolution image within the screen 24A than when the image within the screen is created using only the composite point cloud data SG. Furthermore, since the screen 24A is the screen of the display 24, the texture image can be viewed by the user through the display 24.

[0102] Furthermore, in the information processing system 2, among the plurality of pixels included in the captured image PD, pixels at positions corresponding to the positions of the intersections of the plurality of triangular patches 54A included in the TIN 54 (see FIGS. 7 to 9) are assigned to the on-screen intersection positions, thereby creating a texture image. Then, among the plurality of pixels constituting the captured image PD, pixels at camera two-dimensional coordinates (x, y in the example shown in FIG. 10) corresponding to the screen two-dimensional coordinates (u, v) of the interpolation position on the screen 24A, i.e., the second interpolation position D0, are assigned. As a result, as shown in FIG. 12 as an example, when comparing an example in which pixels at positions corresponding only to the positions of the intersections of the plurality of triangular patches 54A included in the TIN 54 are assigned to the on-screen intersection positions with an example in which pixels at positions corresponding to the positions of the intersections of the plurality of triangular patches 54A included in the TIN 54 and pixels at positions other than the intersections of the plurality of triangular patches 54A included in the TIN 54 are assigned to the screen 24A, a texture image with a higher pixel density is obtained in the latter than in the former. In other words, with this configuration, it is possible to display on the screen 24A a texture image that is more detailed than a texture image that is created simply by assigning pixels at positions corresponding to the intersection positions of multiple triangular patches 54A included in the TIN 54 to the intersection positions within the screen.

[0103] Furthermore, in the information processing system 2, a TIN 54 defined by a plurality of triangular patches 54A is generated based on the three-dimensional coordinates of the plurality of cameras. The positions of the intersections of the plurality of triangular patches 54A included in the TIN 54 are perspectively projected onto the xy plane PD0 and the uv plane 24A1. Then, a position A obtained by perspectively projecting the positions of the intersections of the plurality of triangular patches 54A onto the xy plane PD0 is i , A i+1 and A i+2 Each pixel in the figure is a position B obtained by perspectively projecting the intersection points of the multiple triangular patches 54A onto the uv plane 24A1. i , B i+1 and B i+2As a result, a texture image is created within the screen 24A. Therefore, with this configuration, it is possible to create a texture image more easily than when creating a texture image without using polygonal patches such as the triangular patch 54A.

[0104] Furthermore, in the information processing system 2, a position on the screen 24A that corresponds to a position other than the intersection of the triangular patch 54A is applied as the second interpolation position D0. That is, the pixels that make up the captured image PD are assigned not only to the intersection positions on the screen but also to the second interpolation position D0. Therefore, with this configuration, a more detailed texture image can be displayed on the screen 24A than when the pixels that make up the captured image PD are assigned only to the intersection positions on the screen.

[0105] Furthermore, in the information processing system 2, the composite point cloud data SG is converted into a plurality of camera three-dimensional coordinates, and the plurality of camera three-dimensional coordinates are converted into a plurality of camera two-dimensional coordinates and a plurality of screen two-dimensional coordinates. to Then, texture mapping is realized by allocating pixels at positions specified by the camera's two-dimensional coordinates to positions specified by the screen's two-dimensional coordinates. Therefore, with this configuration, a highly accurate texture image can be obtained compared to when texture mapping is performed without converting the composite point cloud data SG into multiple camera three-dimensional coordinates.

[0106] In the above embodiment, one second interpolation position D0 is applied to each triangular patch 54A. However, the technology of the present disclosure is not limited to this. For example, multiple second interpolation positions D0 may be applied to each triangular patch 54A. In this case, multiple first interpolation positions C0 are assigned to each triangular patch 54A. Then, the pixels at the multiple corresponding first interpolation positions C0 are assigned to the multiple second interpolation positions D0. This allows for a more detailed texture image to be created as the image displayed on the screen 24A, compared to when only one second interpolation position D0 is applied to one triangular patch 54A.

[0107] 13, at least one second interpolation position D0 may be applied to a location in the camera's two-dimensional coordinate system corresponding to the center of gravity CG1 and / or center of gravity portion CG2 of the triangular patch 54A (e.g., a circular portion centered on the center of gravity CG1 within an area inside the triangular patch 54A). In this case, at least one first interpolation position C0 is assigned to a location in the camera's two-dimensional coordinate system corresponding to each center of gravity CG1 and / or center of gravity portion CG2 of the triangular patch 54A. Then, in a manner similar to the above embodiment, a pixel at the corresponding first interpolation position C0 is assigned to at least one second interpolation position D0. This can reduce pixel bias in the image displayed on the screen 24A compared to when the second interpolation position D0 is applied to a location closer to the intersection of the triangular patch 54A than to a location corresponding to the center of gravity CG1 of the triangular patch 54A in the camera's two-dimensional coordinate system. 14, at least one second interpolation position D0 may be applied to a location corresponding to a side 54A1 of a triangular patch 54A. In this case, for example, at least one first interpolation position C0 is assigned to a location corresponding to a side 54A1 of the triangular patch 54A (positions C1 and / or C2 in the example shown in FIG. 10). Then, in the same manner as in the above embodiment, the pixel of the corresponding first interpolation position C0 is assigned to at least one second interpolation position D0.

[0108] Furthermore, the number of interpolation positions applied to the set of triangles corresponding to the TIN 54 in the camera two-dimensional coordinate system and the screen two-dimensional coordinate system (for example, the average number of interpolation positions applied to the triangular patch 54A) may be changed in response to an external instruction (for example, an instruction received by the reception device 22). In this case, to further refine the image displayed on the screen 24A, the number of interpolation positions applied to the triangles corresponding to the triangular patch 54A in the camera two-dimensional coordinate system and the screen two-dimensional coordinate system may be increased, and to reduce the computational load required for the texture mapping process, the number of interpolation positions may be reduced.

[0109] [First Modification] In the above embodiment, an example was given in which the processor 40 operates as an acquisition unit 40A, a conversion unit 40B, and a pixel allocation unit 40C by executing the texture mapping processing program 52, but the technology of the present disclosure is not limited to this. For example, as shown in FIG. 15, the processor 40 may further operate as a control unit 40D by executing the texture mapping processing program 52.

[0110] As an example, as shown in FIG. 15, the control unit 40D acquires a plurality of camera three-dimensional coordinates by the acquisition unit 40A. fart The control unit 40D acquires the composite point cloud data SG used in the conversion of the above, and also acquires the three-dimensional coordinates of the multiple cameras acquired by the acquisition unit 40A. The control unit 40D generates a data set 56, which is an example of "association information" according to the technology of the present disclosure, based on the composite point cloud data SG and the three-dimensional coordinates of the multiple cameras. The data set 56 is information that associates the composite point cloud data SG before it is converted into the three-dimensional coordinates of the multiple cameras by the acquisition unit 40A with the three-dimensional coordinates of the multiple cameras into which the composite point cloud data SG is converted by the acquisition unit 40A.

[0111] The control unit 40D stores the data set 56 in the storage 42. As a result, the data set 56 is held by the storage 42.

[0112] As an example, as shown in Fig. 16, the texture mapping process according to the first modification example differs from the texture mapping process shown in Fig. 11 in that it includes processing of step ST200 and processing of steps ST202 to ST206. The processing of step ST200 is executed between the processing of step ST102 and the processing of step ST104. The processing of steps ST202 to ST206 is executed after the processing of step ST114 is executed.

[0113] In the texture mapping process shown in Figure 16, in step ST200, the control unit 40D generates a dataset 56 by associating the composite point cloud data SG acquired in step ST100 with the multiple camera three-dimensional coordinates obtained in step ST102, and stores the generated dataset 56 in the storage 42.

[0114] In step ST202, the control unit 40D determines whether or not the synthesized point cloud data SG included in the dataset 56 in the storage 42 has been selected. The selection of the synthesized point cloud data SG is realized, for example, in accordance with an instruction received by the reception device 22 (see FIG. 3). In step ST202, if the synthesized point cloud data SG included in the dataset 56 in the storage 42 has not been selected, the determination is negative, and the texture mapping process proceeds to step ST206. In step ST202, if the synthesized point cloud data SG included in the dataset 56 in the storage 42 has been selected, the determination is positive, and the texture mapping process proceeds to step ST204.

[0115] In step ST204, the acquisition unit 40A acquires a plurality of camera three-dimensional coordinates associated with the composite point cloud data SG selected in step ST202 from the data set 56 in the storage 42, and generates a TIN 54 based on the acquired plurality of camera three-dimensional coordinates. After the process of step ST204 is executed, the texture mapping process proceeds to step ST106, and in steps ST106 to ST110, processes using the TIN 54 generated in step ST204 are executed. Then, after the process of step ST110 is executed, the processes of steps ST112 and ST114 are executed.

[0116] In step ST206, the control unit 40D determines whether or not a condition for terminating the texture mapping process (hereinafter referred to as the "termination condition") has been satisfied. A first example of the termination condition is that an instruction to terminate the texture mapping process has been accepted by the accepting device 22 (see FIG. 3). A second example of the termination condition is that a predetermined time (e.g., 15 minutes) has elapsed since the start of execution of the texture mapping process without the judgment in step ST202 being affirmative.

[0117] In step ST206, if the termination condition is not satisfied, the determination is negative and the texture mapping process proceeds to step ST202. In step ST206, if the termination condition is satisfied, the determination is positive and the texture mapping process ends.

[0118] As described above, in the first modified example, the storage 42 stores the data set 56 as information associating the composite point cloud data SG with the three-dimensional coordinates of a plurality of cameras. Therefore, according to this configuration, for a portion of the object 50 for which the composite point cloud data SG has already been obtained, texture mapping processing can be performed based on the three-dimensional coordinates of the plurality of cameras associated with the already obtained composite point cloud data SG, without the portion being scanned again by LiDAR.

[0119] Furthermore, in the first modified example, the processor 40 generates the TIN 54 with reference to the data set 56 (see step ST204 in FIG. 16 ), and executes the processes from step ST106 onward. As a result, pixels constituting the captured image PD are assigned to interpolation positions (second interpolation positions D0 in the example shown in FIG. 10 ) within the screen 24A, where multiple camera three-dimensional coordinates are identified by an interpolation method using multiple camera two-dimensional coordinates and multiple screen two-dimensional coordinates. Therefore, with this configuration, for a portion of the object 50 for which composite point cloud data SG has already been obtained, a three-dimensional image (i.e., a texture image perceived three-dimensionally by the user) to be displayed on the screen 24A can be created without being scanned again by LiDAR.

[0120] [Second Modification] In the above embodiment, an example was given in which the composite point cloud data SG is converted into multiple camera 3D coordinates, but the technology disclosed herein is not limited to this, and multiple camera 3D coordinates may be calculated based on multiple captured images PD.

[0121] 17, for example, moving body 10 is moved to a plurality of positions, and camera 32B is caused to capture images of subject 58 from each position. For example, moving body 10 is moved to a first imaging position, a second imaging position, and a third imaging position in this order, and camera 32B is caused to capture images of subject 58 at each of the first imaging position, the second imaging position, and the third imaging position.

[0122] An image of subject 58 is formed on imaging plane 32B1 at a position away from the lens center of camera 32B by the focal length, and the subject image is captured by camera 32B. In the example shown in Fig. 17, a subject image OI1 corresponding to subject 58 is formed on imaging plane 32B1 at the first imaging position. Furthermore, a subject image OI2 corresponding to subject 58 is formed on imaging plane 32B1 at the second imaging position. Furthermore, a subject image OI3 corresponding to subject 58 is formed on imaging plane 32B1 at the third imaging position.

[0123] The captured image PD obtained by capturing an image by the camera 32B at the first imaging position contains an electronic image corresponding to the subject image OI1, the captured image PD obtained by capturing an image by the camera 32B at the second imaging position contains an electronic image corresponding to the subject image OI2, and the captured image PD obtained by capturing an image by the camera 32B at the third imaging position contains an electronic image corresponding to the subject image OI3.

[0124] 18, as an example, the processor 40 acquires a plurality of frames of captured images PD from the camera 32B. Here, the plurality of frames of captured images PD refers to a plurality of frames of captured images PD obtained by the camera 32B capturing images of the subject 58 from at least two different positions, such as a first imaging position, a second imaging position, and a third imaging position. The processor 40 acquires the camera three-dimensional coordinates by calculating the camera three-dimensional coordinates that can identify the positions of the feature points of the subject 58, based on the feature points of the subject 58 included as images (electronic images, as an example) between the plurality of frames of captured images PD obtained by the camera 32B capturing images of the subject 58 from the different positions.

[0125] 18 shows a feature point Q included in the subject 58. A first captured image PD1 obtained by capturing an image of the subject 58 with the camera 32B at the first imaging position includes a feature point q1 in the first captured image that corresponds to the feature point Q. Furthermore, a second captured image PD2 obtained by capturing an image of the subject 58 with the camera 32B at the second imaging position includes a feature point q2 in the second captured image that corresponds to the feature point Q. The camera's three-dimensional coordinates (X, Y, Z) of the feature point Q are calculated based on the two-dimensional coordinates (x3, y3) that are the coordinates in the camera's field of view of the feature point q1 in the first captured image and the two-dimensional coordinates (x4, y4) that are the coordinates in the camera's field of view of the feature point q2 in the second captured image. That is, the processor 40 geometrically calculates the position and orientation of the camera 32B based on the two-dimensional coordinates (x3, y3) of the feature point q1 in the first captured image and the two-dimensional coordinates (x4, y4) of the feature point q2 in the second captured image, and calculates the camera's three-dimensional coordinates (X, Y, Z) of the feature point Q based on the calculation results.

[0126] As an example, as shown in FIG. 19, the position and orientation of the camera 32B are determined based on an epipolar plane 60, which is a triangular plane with three vertices: a feature point Q, a camera center O1, and a camera center O2. The camera center O1 is located on a line passing from the feature point Q to the feature point q1 in the first captured image, and the camera center O2 is located on a line passing from the feature point Q to the feature point q2 in the second captured image. The epipolar plane 60 has a baseline BL. The baseline BL is a line segment connecting the camera centers O1 and O2. The epipolar plane 60 also has epipoles e1 and e2. The epipole e1 is the point where the baseline BL intersects with the plane of the first captured image PD1, and the epipole e2 is the point where the baseline BL intersects with the plane of the second captured image PD2. The first captured image PD1 has an epipolar line EP1. The epipolar line EP1 is a straight line that passes through the epipole e1 and the feature point q1 in the first captured image. The second captured image PD2 has an epipolar line EP2. The epipolar line EP2 is a straight line that passes through the epipole e2 and the feature point q2 in the second captured image.

[0127] In this way, when the subject 58 existing in three-dimensional space is captured by the camera 32B at multiple capturing positions (here, as an example, the first capturing position and the second capturing position) (i.e., when the feature point Q is projected onto the plane of the first captured image PD1 as the feature point q1 in the first captured image and when the feature point Q is projected onto the plane of the second captured image PD2 as the feature point q2 in the second captured image), epipolar geometry appears as a unique geometry based on the epipolar plane 60 between the first captured image PD1 and the second captured image PD2. The epipolar geometry indicates the correlation between the feature point Q, the feature point q1 in the first captured image, and the feature point q2 in the second captured image. For example, if the position of the feature point Q changes, the feature point q1 in the first captured image and the feature point q2 in the second captured image change, and the epipolar plane 60 also changes. If the epipolar plane 60 changes, the epipoles e1 and e2 also change, and therefore the epipolar lines EP1 and EP2 also change. That is, the epipolar geometry includes information required to identify the position and orientation of the camera 32B at a plurality of imaging positions (here, as an example, the first imaging position and the second imaging position).

[0128] Therefore, the processor 40 uses the two-dimensional coordinates (x3, y3) of the feature point q1 in the first captured image, the two-dimensional coordinates (x4, y4) of the feature point q2 in the second captured image, and epipolar geometry to calculate a fundamental matrix E from the following formulas (13) to (19), and estimates a rotation matrix R and a translation vector T from the calculated fundamental matrix E. Then, the processor 40 calculates the camera three-dimensional coordinates (X, Y, Z) of the feature point Q based on the estimated rotation matrix R and translation vector T.

[0129] Here, we will explain an example of a method for calculating the fundamental matrix E. The relationship between the camera three-dimensional coordinates (X1, Y1, Z1) of the feature point Q expressed in the coordinate system of the camera 32B at the first imaging position and the camera three-dimensional coordinates (X2, Y2, Z2) of the feature point Q expressed in the coordinate system of the camera 32B at the second imaging position is expressed by the following formula (13). In equation (13), Q1 is the camera three-dimensional coordinate (X1, Y1, Z1) of feature point Q expressed in the coordinate system of camera 32B at the first imaging position, Q2 is the camera three-dimensional coordinate (X2, Y2, Z2) of feature point Q expressed in the coordinate system of camera 32B at the second imaging position, R is a rotation matrix (i.e., a rotation matrix required for transformation from the coordinate system of camera 32B at the first imaging position to the coordinate system of camera 32B at the second imaging position), and T is a translation vector (i.e., a translation vector required for transformation from the coordinate system of camera 32B at the first imaging position to the coordinate system of camera 32B at the second imaging position).

[0130]

number

[0131] The projection of feature point Q from the camera's three-dimensional coordinates (X1, Y1, Z1) onto two-dimensional coordinates (x3, y3) and (x4, y4) is expressed by the following formula (14): In formula (14), λ1 is a value indicating the depth lost when feature point Q is projected onto the first captured image PD1, and λ2 is a value indicating the depth lost when feature point Q is projected onto the second captured image PD2.

[0132]

number

[0133] Equation (14) can be replaced by the following equation (15): When equation (15) is substituted into equation (13), the following equation (16) is obtained.

[0134]

number

[0135]

number

[0136] The translation vector T is a vector that indicates the distance and direction from the camera center O2 to the camera center O1. In other words, the translation vector T corresponds to a vector extending from the camera center O2 to the camera center O1, and coincides with the baseline BL. Meanwhile, "λ2q2" on the left side of Equation (16) (i.e., "Q2" in Equation (15)) corresponds to a line extending from the camera center O2 to the feature point Q, and coincides with one side of the epipolar plane 60. Therefore, the cross product of the translation vector T and "λ2q2" is a vector W (see FIG. 19) that is perpendicular to the epipolar plane 60. Here, when the cross product of the right side of Equation (16) and the translation vector T is also taken, it is expressed as the following Equation (17). Note that in Equation (17), the translation vector T is expressed as a skew-symmetric matrix [T].

[0137]

number

[0138] Since the vector W (see FIG. 19) and the epipolar plane 60 are orthogonal to each other, when the inner product of both sides of equation (17) and the two-dimensional coordinates (x4, y4) of the feature point q2 in the second captured image is taken, the result is "0" as shown in the following equation (18). λ1 and λ2 are constant terms that do not affect the estimation of the rotation matrix R and the translation vector T, and therefore have been eliminated from equation (18). If "E = [T] × R", equation (19), which indicates the so-called epipolar constraint equation, can be derived from equation (18).

[0139]

number

[0140]

number

[0141] The fundamental matrix E is calculated by simultaneously solving a plurality of epipolar constraint equations derived for each of a plurality of feature points including the feature point Q. Since the fundamental matrix E includes only the rotation matrix R and the translation vector T, the rotation matrix R and the translation vector T can also be obtained by calculating the fundamental matrix E. The rotation matrix R and the translation vector T are singularly calculated based on the fundamental matrix E. value Disassembly and other well-known methods (e.g., https: / / ir.lib.hiroshima-u.ac.jp / 00027688, 20091125note_reconstruction.pdf_Image Engineering Special Lecture Notes: 3D Reconstruction with Linear Algebra_Tamaki Toru_November 25, 2009_See pages 59 "1.4.2_Epipolar Geometry Using Projective Geometry" to 69 "1.4.4_Summary of 3D Reconstruction").

[0142] The camera three-dimensional coordinates (X, Y, Z) of the feature point Q are calculated by substituting the rotation matrix R and translation vector T estimated based on the fundamental matrix E into Equation (13).

[0143] As described above, according to the second modified example, the three-dimensional camera coordinates of the feature point Q of the subject 58 are calculated based on the feature point q1 in the first captured image and the feature point q2 in the second captured image of the subject 58 that are included as images between captured images PD of multiple frames obtained by capturing images of the subject 58 by the camera 32B from different positions. Therefore, according to this configuration, the three-dimensional camera coordinates of the feature point Q of the subject 58 can be calculated without using the LiDAR 32A.

[0144] [Other variations] In the above embodiment, the triangular patch 54A is exemplified, but the technology of the present disclosure is not limited to this. Instead of the triangular patch 54A, a quadrilateral patch defined by a quadrilateral mesh may be used. Also, instead of the triangular patch 54A, a patch of a polygon other than a triangle or a quadrilateral may be used. In this way, since the technology of the present disclosure uses a polygonal patch, texture mapping can be performed more easily than when a polygonal patch is not used.

[0145] The polygonal patch may be a planar patch or a curved patch, but since a planar patch places a smaller load on the texture mapping calculations than a curved patch, it is preferable that the polygonal patch be a planar patch.

[0146] In the above embodiment, an example has been described in which the texture mapping processing program 52 is stored in the storage 42, but the technology of the present disclosure is not limited to this. For example, the texture mapping processing program 52 may be stored in a portable storage medium 100 such as an SSD or a USB memory. The storage medium 100 is a non-transitory computer-readable storage medium. The texture mapping processing program 52 stored in the storage medium 100 is installed in the computer 39 of the information processing device 20. The processor 40 executes texture mapping processing in accordance with the texture mapping processing program 52.

[0147] In addition, the texture mapping processing program 52 may be stored in a storage device such as another computer or server device connected to the information processing device 20 via a network, and the texture mapping processing program 52 may be downloaded and installed in the computer 39 in response to a request from the information processing device 20.

[0148] It is not necessary to store the entire texture mapping processing program 52 in a storage device such as another computer or server device connected to the information processing device 20, or in the storage 42; only a part of the texture mapping processing program 52 may be stored therein.

[0149] Furthermore, although the information processing device 20 shown in FIG. 3 has a built-in computer 39, the technology of the present disclosure is not limited to this. For example, the computer 39 may be provided outside the information processing device 20.

[0150] In the above embodiment, the computer 39 is exemplified, but the technology of the present disclosure is not limited to this, and a device including an ASIC, an FPGA, and / or a PLD may be applied instead of the computer 39. Furthermore, instead of the computer 39, a combination of a hardware configuration and a software configuration may be used.

[0151] The hardware resources for executing the texture mapping process described in the above embodiments can be various processors, as listed below. Examples of processors include a CPU, which is a general-purpose processor that functions as a hardware resource for executing texture mapping by executing software, i.e., a program. Examples of processors include dedicated electrical circuits, such as FPGAs, PLDs, or ASICs, which are processors with a circuit configuration designed specifically for executing specific processes. Each processor has a built-in or connected memory, and uses the memory to execute the texture mapping process.

[0152] The hardware resource that executes the texture mapping process may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the texture mapping process may be a single processor.

[0153] As an example of a system configured with a single processor, there is a first form in which one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes texture mapping processing. A second form is a form in which a processor is used that realizes the functions of the entire system, including multiple hardware resources that execute texture mapping processing, on a single IC chip, as typified by SoCs. In this way, texture mapping processing is realized using one or more of the above-mentioned various processors as hardware resources.

[0154] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The texture mapping process described above is merely an example. It goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the process.

[0155] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[0156] In this specification, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0157] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

Claims

1. a processor; a memory connected to or embedded in the processor; An image processing device in which a positional relationship between a three-dimensional distance measuring sensor and an imaging device having a sampling resolution higher than that of the three-dimensional distance measuring sensor is known, The processor: acquiring three-dimensional coordinates of a three-dimensional distance measuring sensor system that can identify the positions of the plurality of measurement points based on distance measurement results by the three-dimensional distance measuring sensor, the three-dimensional coordinates being defined in a three-dimensional coordinate system applied to the three-dimensional distance measuring sensor; acquiring, based on the three-dimensional coordinates of the three-dimensional distance measuring sensor system, three-dimensional coordinates of the image capturing device system defined in a three-dimensional coordinate system applied to the image capturing device; converting the three-dimensional coordinates of the imaging device system into two-dimensional coordinates of the imaging device system that can identify a position within a captured image obtained by capturing an image by the imaging device; converting the imaging device system three-dimensional coordinates into display system two-dimensional coordinates that can identify a position within a screen; allocating pixels constituting the captured image to interpolation positions within the screen that are specified by an interpolation method using the imaging device system two-dimensional coordinates and the display system two-dimensional coordinates; The processor: converting the three-dimensional coordinates of the three-dimensional distance measuring sensor system into three-dimensional coordinates of the image pickup device system, thereby obtaining the three-dimensional coordinates of the image pickup device system; generating a polygonal patch based on the three-dimensional coordinates of the imaging device system; the imaging device system three-dimensional coordinates define the positions of the intersections of the polygonal patches; the interpolation position is a position within the screen that corresponds to a position other than an intersection point of the polygonal patch, The processor selects a front pixel from among a plurality of pixels included in the captured image. A three-dimensional image is created by further allocating pixels at positions corresponding to the intersection points of the polygonal patches to intersection points within the screen corresponding to the intersection points of the polygonal patches. Image processing device.

2. The polygonal patch is defined by a triangular mesh or a quadrilateral mesh. The image processing device according to claim 1 .

3. The processor acquires the three-dimensional coordinates of the image capturing device system by calculating the three-dimensional coordinates of the image capturing device system based on feature points of the object included as images among a plurality of frame images obtained by capturing images of the object from different positions by the image capturing device.

3. The image processing device according to claim 1.

4. The position within the screen is a position within the screen of the display.

3. The image processing device according to claim 1.

5. The position within the screen is a position within the screen of a display. The image processing device according to claim 3 .

6. The memory stores correspondence information that associates the three-dimensional coordinates of the three-dimensional distance measuring sensor system with the two-dimensional coordinates of the imaging device system.

3. The image processing device according to claim 1.

7. The memory stores correspondence information that associates the three-dimensional coordinates of the three-dimensional ranging sensor system with the two-dimensional coordinates of the imaging device system. The image processing device according to claim 3 .

8. The memory stores correspondence information that associates the three-dimensional coordinates of the three-dimensional ranging sensor system with the two-dimensional coordinates of the imaging device system. The image processing device according to claim 4 .

9. The processor refers to the association information and assigns pixels that make up the captured image to the interpolation positions. The image processing device according to claim 6 .

10. The processor refers to the correspondence information and assigns pixels that make up the captured image to the interpolation positions. The image processing device according to claim 7 .

11. The processor refers to the correspondence information and assigns pixels that make up the captured image to the interpolation positions. The image processing device according to claim 8 .

12. Under the condition that the positional relationship between the three-dimensional ranging sensor and an imaging device having a higher sampling resolution than the three-dimensional ranging sensor is known, based on the ranging results of the three-dimensional ranging sensor for the multiple measurement points, three-dimensional coordinates of the three-dimensional ranging sensor system that can identify the positions of the multiple measurement points are acquired, the three-dimensional coordinates of the three-dimensional ranging sensor system being defined in a three-dimensional coordinate system applied to the three-dimensional ranging sensor; Based on the three-dimensional coordinates of the three-dimensional distance measuring sensor system, three-dimensional coordinates of the imaging device system defined in the three-dimensional coordinate system applied to the imaging device To obtain converting the three-dimensional coordinates of the imaging device system into two-dimensional coordinates of the imaging device system that can identify a position within a captured image obtained by capturing an image by the imaging device; converting the imaging device system three-dimensional coordinates into display system two-dimensional coordinates that can identify a position within a screen; allocating pixels constituting the captured image to interpolation positions within the screen that are specified by an interpolation method based on the two-dimensional coordinates of the image capture device system and the two-dimensional coordinates of the display system; Acquiring the three-dimensional coordinates of the imaging device system by converting the three-dimensional coordinates of the three-dimensional distance measuring sensor system into three-dimensional coordinates of the imaging device system; and generating a polygonal patch based on the three-dimensional coordinates of the image capture device; the imaging device system three-dimensional coordinates define the positions of the intersections of the polygonal patches; the interpolation position is a position within the screen that corresponds to a position other than an intersection point of the polygonal patch, and creating a three-dimensional image by further allocating pixels at positions corresponding to the positions of the intersections of the polygonal patches among the plurality of pixels included in the captured image to positions of intersections within the screen that correspond to the positions of the intersections of the polygonal patches. Image processing methods.

13. A program for causing a computer to execute a process, The process comprises: A three-dimensional distance measuring sensor system three-dimensional coordinate system that can identify the positions of a plurality of measurement points based on the distance measurement results of the three-dimensional distance measuring sensor for a plurality of measurement points under the condition that the positional relationship between the three-dimensional distance measuring sensor and an imaging device having a higher sampling resolution than the three-dimensional distance measuring sensor is known, and the three-dimensional distance measuring sensor system three-dimensional coordinate system that can identify the positions of the plurality of measurement points based on the distance measurement results of the three-dimensional distance measuring sensor for a plurality of measurement points ... Acquiring three-dimensional coordinates of a three-dimensional ranging sensor system defined in a coordinate system; acquiring, based on the three-dimensional coordinates of the three-dimensional distance measuring sensor system, three-dimensional coordinates of the image capturing device system defined in a three-dimensional coordinate system applied to the image capturing device; converting the three-dimensional coordinates of the imaging device system into two-dimensional coordinates of the imaging device system that can identify a position within a captured image obtained by capturing an image by the imaging device; converting the imaging device system three-dimensional coordinates into display system two-dimensional coordinates that can identify a position within a screen; allocating pixels constituting the captured image to interpolation positions within the screen that are specified by an interpolation method using the imaging device system two-dimensional coordinates and the display system two-dimensional coordinates; Acquiring the three-dimensional coordinates of the imaging device system by converting the three-dimensional coordinates of the three-dimensional distance measuring sensor system into three-dimensional coordinates of the imaging device system; and generating a polygonal patch based on the three-dimensional coordinates of the image capture device; the imaging device system three-dimensional coordinates define the positions of the intersections of the polygonal patches; the interpolation position is a position within the screen that corresponds to a position other than an intersection point of the polygonal patch, and creating a three-dimensional image by further allocating pixels at positions corresponding to the positions of the intersections of the polygonal patches among the plurality of pixels included in the captured image to positions of intersections within the screen that correspond to the positions of the intersections of the polygonal patches. program.

Citation Information

Patent Citations

  • Three-dimensional shape measuring system and color information attachment method

    JP2015064755A

  • Systems and methods to transform a colored point cloud to a 3D textured mesh

    US8948498B1