Method and system for generating three-dimensional plant point cloud having multimodal attributes

The use of a multimodal camera system with AI processing generates low-cost, accurate 3D plant point clouds, addressing environmental noise and cost issues to analyze plant state comprehensively.

WO2026063566A1PCT designated stage Publication Date: 2026-03-26KOREA INST OF SCI & TECH
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-11-25
Publication Date
2026-03-26

AI Technical Summary

Technical Problem

Existing methods for generating 3D images of plants are expensive and prone to noise due to environmental interference, especially in outdoor conditions, making it difficult to accurately capture the 3D structure and spatial distribution of plants in dense clusters.

Method used

A method and system using a multimodal camera, comprising an RGB camera and a thermal imaging camera, to capture multiple images and depth information, combined with an artificial intelligence model, to generate accurate 3D plant point clouds from various angles, overcoming environmental noise and cost limitations.

Benefits of technology

Enables the generation of low-cost, accurate 3D plant point clouds that allow comprehensive analysis of plant state from multiple perspectives, including growth status, disease presence, nutritional status, and physiological stress, with precise measurement of structural and spatial features.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024018771_26032026_PF_FP_ABST
    Figure KR2024018771_26032026_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to a method and a system for generating a three-dimensional plant point cloud having multimodal attributes, the method comprising: acquiring, for a plant subject, a plurality of RGB images and a plurality of multimodal images captured at a plurality of capture viewpoints by a multimodal camera; acquiring a plurality of depth images for the plant subject from the plurality of RGB images by using an artificial intelligence model; and generating, using the plurality of depth images for the plant subject, an RGB-type three-dimensional plant point cloud based on the plurality of RGB images and a multimodal-type three-dimensional plant point cloud based on the plurality of multimodal images.
Need to check novelty before this filing date? Find Prior Art

Description

Method and system for generating a 3D plant point cloud with multimodal attributes

[0001] This relates to a method and system for generating a three-dimensional point cloud for plants.

[0002] It is difficult to obtain phenotypic information about plants, such as their spatial distribution pattern, volume, leaf area, or leaf curvature, using 2D image analysis. To obtain this phenotypic information without distortion, it is necessary to generate 3D images that represent the plant's accurate 3D structure. Currently, LiDAR sensors, laser scanners, and depth cameras are used to generate 3D images of plants. However, these devices are very expensive, and there is a problem in that noise is generated due to mutual interference between the light emitted by the devices and sunlight in external environments where sunlight is exposed.

[0003] LiDAR sensors provide relatively high-accuracy 3D information but are expensive, and laser scanners are sensitive to environmental conditions, limiting their use in outdoor environments. Depth cameras are relatively inexpensive, but noise occurs when exposed to sunlight, making it difficult to obtain accurate data. In particular, in environments such as greenhouses or open fields, various plants are generally cultivated in dense clusters. Due to this density of plants, it has been difficult to obtain shooting or measurement data for each plant from various angles, which has resulted in the problem of being unable to generate 3D images that accurately represent the 3D structure of the plants.

[0004] The invention aims to provide a method for generating a 3D plant point cloud and a system for generating a 3D plant point cloud that can easily generate a point cloud showing the accurate 3D structural features and spatial distribution patterns of plants at a low cost using only a multimodal camera, and can comprehensively analyze the current state of plants from various perspectives. The invention is not limited to the technical challenges described above, and other technical challenges may be derived from the following description.

[0005] A method for generating a three-dimensional plant point cloud according to one aspect of the present invention comprises: a step of acquiring a plurality of RGB images and a plurality of multimodal images captured at a plurality of capture viewpoints by a multimodal camera for a plant subject; a step of acquiring a plurality of depth images for the plant subject from the acquired plurality of RGB images using an artificial intelligence model; a step of generating an RGB-type three-dimensional plant point cloud based on the acquired plurality of RGB images using the plurality of depth images for the plant subject; and a step of generating a multimodal-type three-dimensional plant point cloud based on the acquired plurality of multimodal images using the plurality of depth images for the plant subject, wherein each of the plurality of multimodal images is an image having image attributes different from the image attributes of each of the plurality of RGB images.

[0006] The above method for generating a three-dimensional plant point cloud further includes the step of acquiring multiple pose information of a multimodal camera corresponding to the multiple shooting times, and the step of acquiring the multiple depth images can be acquired from the artificial intelligence model by inputting the acquired multiple RGB images and the acquired multiple pose information into the artificial intelligence model.

[0007] Each pose information of the multimodal camera may include the 3D position coordinates (x, y, z) of the multimodal camera, the rotation angle in the pan direction, and the rotation angle in the tilt direction.

[0008] The above method for generating a three-dimensional plant point cloud further includes the step of generating multiple image-pose pairs for the plurality of shooting points by grouping the RGB images and pose information of the same shooting point among the acquired plurality of RGB images and the acquired plurality of pose information, thereby generating image-pose pairs for each shooting point for the plurality of shooting points, and the step of acquiring the plurality of depth images can be acquired from the artificial intelligence model by inputting the generated multiple image-pose pairs into the artificial intelligence model.

[0009] The step of acquiring the plurality of RGB images and the plurality of multimodal images can be achieved by repeating the process of acquiring the RGB image and the multimodal image captured at each shooting point by a multimodal camera moved to the pose represented by each pose information for the plurality of shooting points.

[0010] The step of acquiring multiple pose information of the multimodal camera can be performed by using a checkerboard to acquire multiple pose information of the multimodal camera.

[0011] The step of acquiring the plurality of depth images involves acquiring a plurality of RGB images and a plurality of depth images at a plurality of reconstruction viewpoints of the artificial intelligence model from the artificial intelligence model, and the method for generating a three-dimensional plant point cloud further includes the step of constructing a set of RGB images composed of a plurality of RGB images and a plurality of depth images at the plurality of reconstruction viewpoints, and the step of generating a three-dimensional plant point cloud of the RGB type can generate the three-dimensional plant point cloud of the RGB type by repeating the process of generating points of the point cloud at each reconstruction viewpoint using the RGB images and depth images for each viewpoint from the set of RGB images for the plurality of reconstruction viewpoints.

[0012] The above method for generating a three-dimensional plant point cloud further includes the step of acquiring multiple pose information of a multimodal camera corresponding to the plurality of shooting points, and the step of acquiring the plurality of depth images can acquire multiple RGB images and multiple depth images at the plurality of reconstruction points by inputting the acquired multiple RGB images and the acquired multiple pose information into the artificial intelligence model.

[0013] The step of acquiring multiple poses of the multimodal camera may further include the step of acquiring a scale ratio between the coordinate system of the actual space where the plant subject is located and the pixel coordinate system of the multiple RGB images, along with the multiple pose information of the multimodal camera, by performing calibration of the multimodal camera using a checkerboard, and the method for generating a 3D plant point cloud may further include the step of setting the pixel scale of each of the multiple RGB images and the multiple depth images at the multiple reconstruction points according to the acquired scale ratio.

[0014] The above method for generating a three-dimensional plant point cloud further includes the step of obtaining a plurality of multimodal images and a plurality of depth images at a plurality of reconstruction points from the artificial intelligence model; and the step of constructing a set of multimodal images composed of a plurality of multimodal images and a plurality of depth images at a plurality of reconstruction points, wherein the step of generating a three-dimensional plant point cloud of the multimodal type can generate the three-dimensional plant point cloud of the multimodal type by repeating the process of generating points of the point cloud at each reconstruction point using the multimodal image and depth image for each time point in the set of multimodal images for the plurality of reconstruction points.

[0015] The above method for generating a three-dimensional plant point cloud further includes the step of acquiring multiple pose information of a multimodal camera corresponding to the plurality of shooting points; and the step of correcting the plurality of multimodal images based on the plurality of RGB images so that the acquired plurality of multimodal images are aligned with the plurality of RGB images on a pixel-by-pixel basis, and the step of acquiring the plurality of depth images can acquire the plurality of multimodal images and the plurality of depth images at the plurality of reconstruction points by inputting the corrected plurality of multimodal images and the acquired multiple pose information into the artificial intelligence model.

[0016] The step of acquiring multiple poses of the multimodal camera may further include the step of acquiring a scale ratio between the coordinate system of the actual space where the plant subject is located and the pixel coordinate system of the multiple multimodal images, along with the multiple pose information of the multimodal camera, by performing calibration of the multimodal camera using a checkerboard, and the 3D plant point cloud generation method may further include the step of setting the pixel scale of each of the multiple multimodal images and the multiple depth images at the multiple reconstruction points according to the acquired scale ratio.

[0017] The above artificial intelligence model may be NeRF (Neural Radiance Fields).

[0018] Each of the above multiple multimodal images may be a thermal image having thermal properties different from the visual properties of each of the above multiple RGB images.

[0019] The multimodal camera comprises an RGB camera that generates a plurality of RGB images by photographing the plant subject at the plurality of shooting points, and a thermal imaging camera that generates a plurality of thermal images as the plurality of multimodal images by photographing the plant subject at the plurality of shooting points, and the step of acquiring the plurality of RGB images and the plurality of multimodal images may involve acquiring the plurality of RGB images from the RGB camera and acquiring the plurality of thermal images from the thermal imaging camera.

[0020] According to another aspect of the present invention, a computer-readable recording medium is provided that records a program for executing the three-dimensional plant point cloud generation method on a computer.

[0021] A three-dimensional plant point cloud generation system according to another aspect of the present invention comprises: an image acquisition module that acquires a plurality of RGB images and a plurality of multimodal images captured at a plurality of capture viewpoints by a multimodal camera for a plant subject; a reconstruction module that acquires a plurality of depth images for the plant subject from the acquired plurality of RGB images using an artificial intelligence model; and a point cloud generation module that generates an RGB type three-dimensional plant point cloud based on the acquired plurality of RGB images and a multimodal type three-dimensional plant point cloud based on the acquired plurality of multimodal images using the plurality of depth images for the plant subject, wherein each of the plurality of multimodal images is an image having image attributes different from the image attributes of each of the plurality of RGB images.

[0022] By acquiring multiple RGB images and multiple multimodal images captured at multiple shooting points on a plant subject by a multimodal camera, acquiring multiple depth images of the plant subject from the multiple RGB images using an artificial intelligence model, and generating an RGB-type three-dimensional plant point cloud based on multiple RGB images and a multimodal-type three-dimensional plant point cloud based on multiple multimodal images using the multiple depth images of the plant subject, a point cloud showing the accurate three-dimensional structural features and spatial distribution pattern of the plant can be easily generated at a low cost using only a multimodal camera.

[0023] In particular, as a three-dimensional plant point cloud with multimodal properties can be generated, the current state of the plant can be comprehensively analyzed from various aspects. For example, the growth status, presence of disease, nutritional status, and physiological stress of the plant can be analyzed based on the color of the RGB type three-dimensional plant point cloud, for example, the color of the leaves, and the water stress and transpiration rate of the plant can be analyzed based on the color of the multimodal type three-dimensional plant point cloud corresponding to a thermal image.

[0024] The effects are not limited to those described above, and other effects may be derived from the following description.

[0025] FIG. 1 is a configuration diagram of a three-dimensional plant point cloud generation system according to one embodiment of the present invention.

[0026] FIG. 2 is a flowchart of a method for generating a three-dimensional plant point cloud according to an embodiment of the present invention.

[0027] FIGS. 3 and 4 are exemplary diagrams of the multimodal camera (10) and camera movement module (20) shown in FIG. 1.

[0028] FIGS. 5 and 6 are example diagrams of pose movement of a multimodal camera (10) by a camera movement module (20) shown in FIG. 1.

[0029] FIGS. 7 and 8 are example diagrams of the calibration process of the multimodal camera (10) shown in FIG. 1.

[0030] FIG. 9 is an example of an image-pose pair database stored in the storage (80) shown in FIG. 1.

[0031] FIG. 10 is an example of the process in which an RGB type three-dimensional plant point cloud is generated by the reconstruction module (50) and the point cloud generation module (70) shown in FIG. 1.

[0032] FIG. 11 is an example of the process in which a multimodal type 3D plant point cloud is generated by the reconstruction module (50) and the point cloud generation module (70) shown in FIG. 1.

[0033] FIG. 12 is an example of a three-dimensional plant point cloud of RGB type generated by the point cloud generation module (70) shown in FIG. 1.

[0034] FIG. 13 is an example of a multimodal type three-dimensional plant point cloud generated by the point cloud generation module (70) shown in FIG. 1.

[0035] FIGS. 14 and 15 are examples of post-processing of an RGB type 3D plant point cloud generated by the point cloud generation module (70) shown in FIG. 1.

[0036] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings. The embodiments of the present invention described below relate to a method for generating a three-dimensional plant point cloud and a system for generating a three-dimensional plant point cloud that can easily generate a point cloud showing the accurate three-dimensional structural features and spatial distribution pattern of a plant at a low cost using only a multimodal camera, and can comprehensively analyze the current state of the plant from various aspects. Hereinafter, this method and system will be briefly referred to as the “method for generating a three-dimensional plant point cloud” and the “system for generating a three-dimensional plant point cloud.”

[0037] FIG. 1 is a configuration diagram of a three-dimensional plant point cloud generation system according to an embodiment of the present invention. Referring to FIG. 1, the three-dimensional plant point cloud generation system according to the present embodiment comprises a multimodal camera (10), a camera movement module (20), a control module (30), an image acquisition module (40), a reconstruction module (50), an image matching module (60), a point cloud generation module (70), a storage (80), and a user interface (90). FIG. 1 illustrates the essential configuration of the present embodiment to ensure that the present embodiment is easily understood while preventing the features of the present embodiment from being obscured. It can be understood that other configurations may be added in addition to the configurations shown in FIG. 1.

[0038] A multimodal camera (10) generates multiple images having multimodal properties by using various types of sensors, such as a visible light sensor, an infrared sensor, a chlorophyll fluorescence sensor, etc. For example, the multimodal camera (10) may be composed of an RGB (Red Green Blue) camera and a thermal imaging camera, and generates RGB images and thermal images having multimodal properties of visual information and thermal information. The multimodal camera (10) may be composed of an RGB camera and a chlorophyll fluorescence camera, and generates RGB images and chlorophyll fluorescence images having multimodal properties of visual information and chlorophyll fluorescence information.

[0039] The camera movement module (20) serves to move the multimodal camera (10) to a pose indicated by pose information selected by the control module (30) under the control of the control module (30). Each pose information of the multimodal camera (10) consists of the 3D position coordinates (x, y, z) of the multimodal camera (10), a rotation angle (θ) in the pan direction, and a rotation angle (_) in the tilt direction. The camera movement module (20) may be implemented as a robot arm, or it may be implemented as a combination of an XYZ linear motion module and a pan-tilt module. The camera movement module (20) repeats the process of moving the multimodal camera (10) to a certain pose under the control of the control module (30), waiting for a certain period of time, and then moving the multimodal camera (10) to the next pose.

[0040] The control module (30) selects one of the pose information among the multiple pose information acquired by the image acquisition module (40) and controls the camera movement module (20) so that the multimodal camera (10) moves to the pose indicated by the selected pose information. Subsequently, the control module (30) controls the multimodal camera (10) so that an RGB image and a multimodal image are captured by the multimodal camera (10) during the waiting time. The above-described process is repeated for all of the multiple pose information acquired by the image acquisition module (40).

[0041] Additionally, the control module (30) reads out a three-dimensional plant point cloud stored in the storage (80), generates a two-dimensional plant image from the three-dimensional plant point cloud according to a projection viewpoint input by the user through the user interface (90), and outputs the two-dimensional plant image thus generated to the user interface (90). Through the operation of the control module (30), the three-dimensional plant point cloud stored in the storage (80) is converted into a two-dimensional image of the projection viewpoint input by the user and displayed on the user interface (90).

[0042] Storage (80) stores a plurality of image-pose pairs generated by an image acquisition module (40), a plurality of image-pose pairs replaced with a plurality of multimodal images by an image matching module (60), a set of RGB images and a set of multimodal images constructed by a reconstruction module (50), and an RGB type 3D plant point cloud and a multimodal type 3D plant point cloud generated by a point cloud generation module (70). Storage (80) may store a program for executing the 3D plant point cloud generation method illustrated in FIG. 2 on a computer. In this case, the control module (30), image acquisition module (40), reconstruction module (50), image matching module (60), and point cloud generation module (70) may be implemented as a combination of a processor and a computer program.

[0043] The user interface (90) receives from the user an execution command for the method of generating a three-dimensional plant point cloud shown in FIG. 2, a three-dimensional plant point cloud stored in storage (80), etc., or outputs a two-dimensional plant image according to the control of the control module (30). The user interface (90) can be implemented as a combination of a display panel and a touchscreen panel.

[0044] FIG. 2 is a flowchart of a method for generating a three-dimensional plant point cloud according to an embodiment of the present invention. Referring to FIG. 2, the method for generating a three-dimensional plant point cloud according to the present embodiment consists of the following steps executed by the three-dimensional plant point cloud generation system illustrated in FIG. 1. Hereinafter, with reference to FIG. 2, the operation of the image acquisition module (40), the reconstruction module (50), the image matching module (60), and the point cloud generation module (70) will be described in detail.

[0045] In step 21, the image acquisition module (40) acquires multiple pose information of the multimodal camera (10) corresponding to multiple capture viewpoints pre-set by the user. The image acquisition module (40) acquires multiple pose information of the multimodal camera (10) by performing calibration of the multimodal camera (10) using a checkerboard installed at the location of the plant subject. Each pose information of the multimodal camera (10) consists of the 3D position coordinates (x, y, z) of the multimodal camera (10), a rotation angle (θ) in the pan direction, and a rotation angle (_) in the tilt direction. The multimodal camera (10) moved to the pose indicated by each pose information of the multimodal camera (10) has a capture viewpoint corresponding to each pose information.

[0046] In step 21, the image acquisition module (40) can obtain the accumulation ratio between the coordinate system of the actual space where the plant subject is located and the pixel coordinate system of the RGB image captured by the multimodal camera (10) along with multiple pose information of the multimodal camera (10) by performing calibration of the multimodal camera (10) using a checkerboard installed at the location of the plant subject so that the scale of the 3D plant point cloud matches the scale of the coordinate system of the actual space where the plant subject is located.

[0047] For example, if the unit length of a 3D plant point cloud is 1 mm, and that unit length of 1 mm represents a length of 1 mm in actual space, then the scale of the 3D plant point cloud corresponds to the scale of the coordinate system in actual space where the plant subject is located. Since the 3D plant point cloud can be enlarged or reduced and displayed on the user interface (90), the unit length of the 3D plant point cloud is not a physical measured length, but a length corresponding to a pre-set unit for measuring the distance between points of the 3D plant point cloud.

[0048] In step 22, the control module (30) selects one of the pose information among the multiple pose information acquired by the image acquisition module (40) in step 21, and the camera movement module (20) moves the multimodal camera (10) to the pose indicated by the pose information selected by the control module (30). In this embodiment, the movement of the multimodal camera (10) to the pose indicated by a certain pose information means that, in order for the multimodal camera (10) to be captured at the shooting point corresponding to the pose information, the multimodal camera (10) is positioned at the 3D position coordinates (x, y, z) of the pose information, and the multimodal camera (10) is rotated to the left or right by the rotation angle (θ) of the pan direction of the pose information, and rotated upward or downward by the rotation angle (_) of the tilt direction of the pose information.

[0049] In step 23, the image acquisition module (40) acquires an RGB image and a multimodal image taken at the shooting point corresponding to the pose information selected in step 22 by the multimodal camera (10) moved by the camera movement module (20) in step 22. In step 24, the image acquisition module (40) checks whether the shooting of the multimodal camera (10) is completed for all of the multiple pose information acquired in step 21. If, as a result of the check in step 24, the shooting of the multimodal camera (10) is completed for all of the multiple pose information acquired in step 21, the process proceeds to step 25. Otherwise, it returns to step 23.

[0050] When returning to step 23, the control module (30) selects pose information different from the already selected pose information among the multiple pose information obtained in step 21, and the camera movement module (20) moves the multimodal camera (10) to the pose indicated by the selected pose information. In this way, as steps 22 to 24 are repeated, in step 23, the image acquisition module (40) obtains multiple RGB images and multiple multimodal images by repeating the process of obtaining RGB images and multimodal images taken at each shooting point by the multimodal camera (10) moved to the pose indicated by each pose information obtained in step 21 for multiple shooting points. That is, in step 23, the image acquisition module (40) obtains multiple RGB images and multiple multimodal images taken at multiple shooting points corresponding to the multiple pose information obtained in step 21 by the multimodal camera (10) for the plant subject.

[0051] In this embodiment, each of the plurality of multimodal images is an image having image attributes different from the image attributes of each of the plurality of RGB images. As described above, each of the plurality of multimodal images may be a thermal image having thermal attributes different from the visual attributes of each of the plurality of RGB images. In this case, the multimodal camera (10) may be composed of an RGB camera that generates a plurality of RGB images by photographing a plant subject at a plurality of shooting points, and a thermal camera that generates a plurality of thermal images as a plurality of multimodal images by photographing a plant subject at a plurality of shooting points. In step 23, the image acquisition module (40) acquires a plurality of RGB images from the RGB camera and acquires a plurality of thermal images from the thermal camera.

[0052] FIGS. 3 and 4 are exemplary diagrams of the multimodal camera (10) and camera movement module (20) illustrated in FIG. 1. FIGS. 3 and 4 show the multimodal camera (10) mounted on a robot arm corresponding to the camera movement module (20). FIG. 3 shows the overall view of the robot arm corresponding to the camera movement module (20), and FIG. 4 shows an enlarged view of the part where the multimodal camera (10) is mounted on the camera movement module (20).

[0053] FIGS. 5 and 6 are example diagrams of pose movement of a multimodal camera (10) by a camera movement module (20) illustrated in FIG. 1. FIG. 5 shows a plurality of poses of the multimodal camera (10) that are sequentially moved by the camera movement module (20), represented by a plurality of wireframes. As shown in FIG. 6, the wireframes representing each pose of the multimodal camera (10) are a type of 5-dimensional vector representing the 3-dimensional position coordinates (x, y, z) of the multimodal camera (10) in a 3-dimensional coordinate system, the rotation angle (θ) in the pan direction, and the rotation angle (_) in the tilt direction.

[0054] FIGS. 7 and 8 are exemplary diagrams of the calibration process of the multimodal camera (10) illustrated in FIG. 1. FIG. 7 illustrates a scene in which an image acquisition module (40) performs calibration of the multimodal camera (10) based on SFM (Structure From Motion) using a checkerboard installed at a location where a plant subject is to be placed. By performing calibration of the multimodal camera (10) using a checkerboard, the internal and external parameters of the multimodal camera (10) can be accurately estimated.

[0055] Examples of internal parameters of the multimodal camera (10) include focal length, principal point, radial distortion coefficient, and tangential distortion coefficient. Examples of external parameters of the multimodal camera (10) include the three-dimensional position coordinates (x, y, z) of the multimodal camera (10), the rotation angle in the pan direction (θ), and the rotation angle in the tilt direction (_).

[0056] When calibration of the multimodal camera (10) is performed using a checkerboard, in addition to the internal and external parameters of the multimodal camera (10), the scale ratio between the coordinate system of the actual space where the plant subject is located and the pixel coordinate system of the RGB image captured by the multimodal camera (10) can be accurately estimated. For example, if the distance between two adjacent points on the checkerboard is 5 cm and the distance between two points corresponding to these two points in the pixel coordinate system is 500 pixels, the scale ratio is 5 cm / 500 pixels, so it becomes 0.1 mm / pixel.

[0057] In step 25, the image acquisition module (40) generates multiple image-pose pairs for multiple shooting points by combining the RGB images and pose information of the same shooting point among the multiple RGB images acquired in step 23 and the multiple pose information acquired in step 21, thereby generating image-pose pairs for each shooting point for all multiple shooting points. The image acquisition module (40) stores the multiple image-pose pairs generated in this manner in the storage (80). Each image-pose pair for each shooting point consists of an RGB image taken at each shooting point and pose information corresponding to each shooting point.

[0058] FIG. 9 is an example of an image-pose pair database stored in the storage (80) illustrated in FIG. 1. Referring to FIG. 9, the image-pose pair database records the size of the image (w, h), focal length (fl_x, fl_y), principal point (cx, cy), radial distortion coefficients (k1, k2), and tangential distortion coefficients (p1, p2) as internal parameters of the multimodal camera (10). Additionally, the image-pose pair database records the file path of the RGB image and the file path of the multimodal image each time it is captured at each pose. Furthermore, the image-pose pair database records the external parameters of the multimodal camera (10) in the form of a transform matrix.

[0059] In step 26, the reconstruction module (50) uses an artificial intelligence model to obtain multiple RGB images and multiple depth images of a plant subject at multiple reconstruction viewpoints of the artificial intelligence model from multiple RGB images obtained by the image acquisition module (40) in step 23 and multiple pose information obtained by the image acquisition module (40) in step 21. The reconstruction module (50) obtains multiple RGB images and multiple depth images at multiple reconstruction viewpoints from the artificial intelligence model by inputting multiple RGB images obtained by the image acquisition module (40) in step 23 and multiple pose information obtained by the image acquisition module (40) in step 21 into the artificial intelligence model, that is, by inputting multiple image-pose pairs generated by the image acquisition module (40) in step 25 and stored in storage (80) into the artificial intelligence model.

[0060] The multiple reconstruction points of an AI model consist of a much larger number of points than the number of multiple shooting points. Accordingly, a 3D point cloud for a plant subject can be formed with very high density. This AI model is trained using a training input dataset consisting of large-scale image-pose pairs and a label set consisting of large-scale RGB images and depth images. When multiple image-pose pairs from the training input dataset are input to the AI ​​model, the AI ​​model is trained by repeating the process of updating the AI ​​model so that multiple RGB images and multiple depth images corresponding to the labels of the input multiple image-pose pairs can be output from the AI ​​model.

[0061] The multiple reconstruction points of an AI model are determined by the number of multiple RGB images and multiple depth images constituting each label in the label set used for training the AI ​​model. The multiple RGB images and multiple depth images constituting each label can be considered as a set of pairs of RGB images and depth images at the same reconstruction point. An example of such an AI model is Neural Radiance Fields (NeRF), a type of Neural Scene Representation (NSR).

[0062] The reconstruction module (50) can set the pixel scale of each of the multiple RGB images and multiple depth images at the reconstruction point obtained from the artificial intelligence model according to the accumulation ratio obtained by the image acquisition module (40) in step 21, so that each of the multiple RGB images and multiple depth images obtained from the artificial intelligence model has a scale of the coordinate system of the actual space where the plant subject is located. Taking the accumulation ratio of the example shown in FIGS. 7 and 8 as an example, the length of 1 pixel in the multiple RGB images and multiple depth images having a scale of the coordinate system of the actual space means 0.1 mm in the actual space where the plant subject is located.

[0063] In step 27, the reconstruction module (50) constructs a set of RGB images consisting of multiple RGB images and multiple depth images at multiple reconstruction points obtained in step 26. The reconstruction module (50) stores the set of RGB images thus constructed in storage (80). The set of RGB images thus constructed consists of multiple pairs of RGB images and depth images, and each pair of RGB images and depth images has the same reconstruction point.

[0064] In step 28, the point cloud generation module (70) generates a three-dimensional plant point cloud of RGB type based on the multiple RGB images obtained by the reconstruction module (50) in step 26 using the multiple RGB images and multiple depth images obtained by the reconstruction module (50) in step 26. In step 27, the point cloud generation module (70) generates a three-dimensional plant point cloud of RGB type according to the accumulation ratio obtained by the image acquisition module (40) in step 21 by repeating the process of generating points of the point cloud at each reconstruction point using the RGB images and depth images at each reconstruction point from the set of RGB images constructed by the reconstruction module (50) for all multiple reconstruction points of the artificial intelligence model. The point cloud generation module (70) stores the three-dimensional plant point cloud of RGB type generated in this way in the storage (80).

[0065] In step 29, the image matching module (60) corrects the multiple multimodal images based on the multiple RGB images so that the multiple multimodal images acquired by the image acquisition module (40) in step 23 are aligned pixel by pixel with the multiple RGB images acquired by the image acquisition module (40) in step 23, and replaces the multiple RGB images of the multiple image-pose pairs generated by the image acquisition module (40) in step 25 and stored in the storage (80) with the multiple multimodal images corrected in this way. When the multiple RGB images of the multiple image-pose pairs are replaced with the multiple multimodal images, the image-pose pairs for each shooting time point consist of the multimodal image at each shooting time point and pose information corresponding to each shooting time point. The multiple image-pose pairs replaced with the multiple multimodal images in this way are stored in the storage (80).

[0066] For example, if each of the multiple multimodal images obtained in step 23 is a thermal image, the resolution of each thermal image is lower than the resolution of each of the multiple RGB images obtained in step 23. In this case, the image matching module (60) can correct the multiple multimodal images based on the multiple RGB images so that the multiple multimodal images obtained in step 23 are matched pixel by pixel to the multiple RGB images obtained in step 23 by interpolating each of the multiple multimodal images obtained in step 23 so that the resolution of each of the multiple multimodal images obtained in step 23 matches the resolution of each of the multiple RGB images obtained in step 23.

[0067] In step 210, the reconstruction module (50) uses an artificial intelligence model to obtain multiple multimodal images and multiple depth images of a plant subject at multiple reconstruction points of the artificial intelligence model from multiple multimodal images corrected by the image matching module (60) in step 29 and multiple pose information obtained by the image acquisition module (40) in step 21. The reconstruction module (50) obtains multiple multimodal images and multiple depth images at multiple reconstruction points from the artificial intelligence model by inputting multiple multimodal images corrected by the image matching module (60) in step 29 and multiple pose information obtained by the image acquisition module (40) in step 21 into the artificial intelligence model, that is, by inputting multiple image-pose pairs replaced by multiple multimodal images by the image matching module (60) in step 29 into the artificial intelligence model.

[0068] Since there is only a difference in pixel values ​​between the RGB image obtained in step 23 and the multimodal image corrected in step 29, for example, the thermal image, the artificial intelligence model trained using the RGB image as described above can be used as is in step 210. Since the multiple depth images obtained in step 210 are identical to the multiple depth images obtained in step 26, the reconstruction module (50) in step 210 may obtain multiple depth images by reusing the multiple depth images obtained in step 26 instead of obtaining multiple depth images from the artificial intelligence model.

[0069] The reconstruction module (50) can set the pixel scale of each of the multiple multimodal images and multiple depth images at the reconstruction point obtained from the artificial intelligence model according to the accumulation ratio obtained by the image acquisition module (40) in step 21, so that each of the multiple multimodal images and multiple depth images obtained from the artificial intelligence model has a scale of the coordinate system of the actual space where the plant subject is located. Taking the accumulation ratio of the example shown in FIGS. 7 and 8 as an example, a 1 pixel length in the multiple multimodal images and multiple depth images having a scale of the coordinate system of the actual space means 0.1 mm in the actual space where the plant subject is located.

[0070] In step 211, the reconstruction module (50) constructs a set of multimodal images consisting of multiple multimodal images and multiple depth images at multiple reconstruction points obtained in step 26. The reconstruction module (50) stores the multimodal image set thus constructed in storage (80). The multimodal image set thus constructed consists of multiple pairs of multimodal images and depth images, and each pair of multimodal images and depth images has the same reconstruction point.

[0071] In step 212, the point cloud generation module (70) generates a multimodal type 3D plant point cloud based on multiple multimodal images obtained by the image acquisition module (40) in step 23, using multiple RGB images and multiple depth images obtained by the reconstruction module (50) in step 210. The point cloud generation module (70) generates a multimodal type 3D plant point cloud according to the accumulation ratio obtained by the image acquisition module (40) in step 21 by repeating the process of generating points of the point cloud at each reconstruction point using the multimodal images and depth images at each reconstruction point from the set of multimodal images constructed by the reconstruction module (50) in step 211 for all multiple reconstruction points of the artificial intelligence model. The point cloud generation module (70) stores the multimodal type 3D plant point cloud generated in this way in the storage (80).

[0072] FIG. 10 is an example of the process of generating an RGB type three-dimensional plant point cloud by the reconstruction module (50) and the point cloud generation module (70) shown in FIG. 1. A process of displaying a color according to the value of a pixel at a point in three-dimensional space specified by the x and y coordinates of a pixel in the RGB image and the z coordinate of a corresponding pixel in the depth image at each reconstruction point in the RGB image and the depth image is repeated for all of the multiple reconstruction points, thereby generating an RGB type three-dimensional plant point cloud as shown in FIG. 10.

[0073] FIG. 11 is an example of the process of generating a multimodal type 3D plant point cloud by the reconstruction module (50) and the point cloud generation module (70) shown in FIG. 1. A multimodal type 3D plant point cloud as shown in FIG. 11 can be generated by repeating the process of displaying a color according to the value of a pixel at a point in 3D space specified by the x, y coordinates of a pixel in an RGB image and the z coordinate of a corresponding pixel in a depth image at each reconstruction point in a multimodal image and a depth image, for all of the multiple reconstruction points.

[0074] FIG. 12 is an example of an RGB type 3D plant point cloud generated by the point cloud generation module (70) shown in FIG. 1, and FIG. 13 is an example of a multimodal type 3D plant point cloud generated by the point cloud generation module (70) shown in FIG. 1. As shown in FIG. 12 and 13, the 3D plant point cloud generated by the point cloud generation module (70) shows the accurate 3D structural features and spatial distribution pattern of the plant. In particular, as the number of reconstruction points increases, the 3D plant point cloud can express the 3D structural features and spatial distribution pattern of the plant more precisely.

[0075] According to the present embodiment, the number of reconstruction points can be adjusted through an artificial intelligence model, so a 3D plant point cloud with the precision desired by the user can be easily generated. Since the 3D plant point cloud of the present embodiment has multimodal attributes such as visual information and thermal intensity, the current state of the plant can be comprehensively analyzed from various aspects using the 3D plant point cloud of the present embodiment. For example, the growth status, presence of disease, nutritional status, and physiological stress of the plant can be analyzed based on the color of the RGB type 3D plant point cloud, for example, the color of the leaves, and the water stress and transpiration rate of the plant can be analyzed based on the color of the multimodal type 3D plant point cloud corresponding to a thermal image.

[0076] FIGS. 14 and 15 are examples of post-processing of an RGB type 3D plant point cloud generated by the point cloud generation module (70) shown in FIG. 1. FIG. 14 illustrates the process of representing geometric information of a plant as a graph through skeletonization processing of an RGB type 3D plant point cloud. Through this skeletonization processing, information regarding the 3D structural features of the plant, such as the length between nodes, can be easily extracted. FIG. 15 illustrates the process of extracting the area of ​​a plant leaf through surface reconstruction processing of an RGB type 3D plant point cloud. Through this surface reconstruction processing, information regarding the external features of the plant, such as the area of ​​a plant leaf, can be easily extracted.

[0077] In particular, since each of the multiple multimodal images and multiple depth images obtained from the artificial intelligence model by the reconstruction module (50) has a scale of the coordinate system of the real space where the plant subject is located, the RGB type 3D plant point cloud or the multimodal type 3D plant point cloud generated by the point cloud generation module (70) also has a scale of the coordinate system of the real space. Accordingly, information on the 3D structural features of the plant, such as the length between nodes of the plant, and information on the external features of the plant, such as the area of ​​the plant leaves, match the measurement information of the real world immediately without a separate calculation processing process. As a result, the user can obtain measurement information of the real world simply by measuring the length or area of ​​the 3D plant point cloud.

[0078] Meanwhile, the method for generating a three-dimensional plant point cloud according to one embodiment of the present invention as described above can be written as a program executable on a computer processor and can be implemented on a computer that executes the program by recording it on a computer-readable recording medium. The computer includes all types of computers capable of executing the program, such as desktop computers, laptop computers, smartphones, and embedded type computers. In addition, the data structure used in the one embodiment of the present invention described above can be recorded on a computer-readable recording medium through various means. A computer-readable recording medium includes storage media such as RAM, ROM, SSD (Solid State Drive), magnetic storage media (e.g., floppy disk, hard disk, etc.), and optical reading media (e.g., CD-ROM, DVD, etc.).

[0079] The present invention has been described above with reference to its preferred embodiments. Those skilled in the art will understand that the present invention may be embodied in modified forms without departing from the essential characteristics of the invention. Therefore, the disclosed embodiments should be considered in an illustrative rather than a restrictive sense. The scope of the invention is defined by the claims, not by the foregoing description, and all variations within the scope of the claims should be interpreted as being included in the invention.

Claims

1. A step of acquiring multiple RGB images and multiple multimodal images captured at multiple capture viewpoints by a multimodal camera for a plant subject; A step of obtaining a plurality of depth images of the plant subject from the plurality of RGB images obtained using an artificial intelligence model; A step of generating an RGB type three-dimensional plant point cloud based on the acquired plurality of RGB images using a plurality of depth images of the plant subject; and The method includes the step of generating a multimodal type three-dimensional plant point cloud based on the acquired plurality of multimodal images using a plurality of depth images of the plant subject, A method for generating a three-dimensional plant point cloud, characterized in that each of the plurality of multimodal images is an image having image properties different from the image properties of each of the plurality of RGB images.

2. In Paragraph 1, The method further includes the step of acquiring multiple pose information of a multimodal camera corresponding to the above multiple shooting points, A method for generating a three-dimensional plant point cloud, characterized in that the step of acquiring the plurality of depth images involves acquiring the plurality of depth images from the artificial intelligence model by inputting the acquired plurality of RGB images and the acquired plurality of pose information into the artificial intelligence model.

3. In Paragraph 2, A method for generating a three-dimensional plant point cloud, characterized in that each pose information of the multimodal camera includes the three-dimensional position coordinates (x, y, z) of the multimodal camera, a rotation angle in the pan direction, and a rotation angle in the tilt direction.

4. In Paragraph 2, The method further includes the step of generating multiple image-pose pairs for the plurality of shooting times by grouping RGB images and pose information of the same shooting time among the plurality of obtained RGB images and the plurality of obtained pose information, thereby generating image-pose pairs for each shooting time for the plurality of shooting times. A method for generating a three-dimensional plant point cloud, characterized in that the step of acquiring the plurality of depth images involves acquiring the plurality of depth images from the artificial intelligence model by inputting the generated plurality of image-pose pairs into the artificial intelligence model.

5. In Paragraph 2, A method for generating a three-dimensional plant point cloud, characterized in that the step of acquiring the plurality of RGB images and the plurality of multimodal images involves acquiring the plurality of RGB images and the plurality of multimodal images by repeating the process of acquiring the RGB images and the multimodal images captured at each shooting point by a multimodal camera moved to the pose indicated by each pose information for the plurality of shooting points.

6. In Paragraph 2, A method for generating a three-dimensional plant point cloud, characterized in that the step of acquiring multiple pose information of the multimodal camera is to acquire multiple pose information of the multimodal camera by performing calibration of the multimodal camera using a checkerboard.

7. In Paragraph 1, The step of acquiring the plurality of depth images involves acquiring a plurality of RGB images and a plurality of depth images at a plurality of reconstruction viewpoints of the artificial intelligence model from the artificial intelligence model, and The method further includes the step of constructing a set of RGB images composed of multiple RGB images and multiple depth images at the aforementioned multiple reconstruction points. A method for generating a three-dimensional plant point cloud of the RGB type, characterized in that the step of generating the three-dimensional plant point cloud of the RGB type involves generating the point cloud points at each reconstruction point using the RGB image and depth image for each reconstruction point in the set of RGB images, thereby generating the three-dimensional plant point cloud of the RGB type.

8. In Paragraph 7, The method further includes the step of acquiring multiple pose information of a multimodal camera corresponding to the above multiple shooting points, A method for generating a three-dimensional plant point cloud, characterized in that the step of acquiring the plurality of depth images involves acquiring the plurality of RGB images and the plurality of depth images at the plurality of reconstruction points by inputting the acquired plurality of RGB images and the acquired plurality of pose information into the artificial intelligence model.

9. In Paragraph 8, The step of acquiring multiple poses of the multimodal camera comprises performing calibration of the multimodal camera using a checkerboard to acquire the accumulation ratio between the coordinate system of the actual space where the plant subject is located and the pixel coordinate system of the multiple RGB images, along with the multiple pose information of the multimodal camera. A method for generating a three-dimensional plant point cloud, characterized by further including the step of setting the pixel scale of each of the plurality of RGB images and the plurality of depth images at the plurality of reconstruction points according to the accumulated ratio obtained above.

10. In Paragraph 7, A step of obtaining a plurality of multimodal images and a plurality of depth images at the plurality of reconstruction points from the artificial intelligence model; and The method further includes the step of constructing a set of multimodal images composed of multiple multimodal images and multiple depth images at the aforementioned multiple reconstruction points, A method for generating a three-dimensional plant point cloud of the multimodal type, characterized in that the step of generating the three-dimensional plant point cloud of the multimodal type involves generating the point cloud at each reconstruction point using a multimodal image and a depth image for each reconstruction point in the set of multimodal images, thereby generating the three-dimensional plant point cloud of the multimodal type by repeating the process for the plurality of reconstruction points.

11. In Paragraph 10, A step of acquiring multiple pose information of a multimodal camera corresponding to the above multiple shooting points; and The method further includes a step of correcting the plurality of multimodal images based on the plurality of RGB images so that the plurality of acquired multimodal images are aligned with the plurality of RGB images on a pixel-by-pixel basis, and A method for generating a three-dimensional plant point cloud, characterized in that the step of acquiring the plurality of depth images involves acquiring the plurality of multimodal images and the plurality of depth images at the plurality of reconstruction points by inputting the corrected plurality of multimodal images and the acquired plurality of pose information into the artificial intelligence model.

12. In Paragraph 11, The step of acquiring multiple poses of the multimodal camera comprises performing calibration of the multimodal camera using a checkerboard to obtain the accumulation ratio between the coordinate system of the actual space where the plant subject is located and the pixel coordinate system of the multiple multimodal images, along with the multiple pose information of the multimodal camera. A method for generating a three-dimensional plant point cloud, characterized by further including the step of setting the pixel scale of each of the plurality of multimodal images and the plurality of depth images at the plurality of reconstruction points according to the accumulated ratio obtained above.

13. In Paragraph 1, A method for generating a three-dimensional plant point cloud characterized by the above artificial intelligence model being NeRF (Neural Radiance Fields).

14. In Paragraph 1, A method for generating a three-dimensional plant point cloud, characterized in that each of the plurality of multimodal images is a thermal image having thermal properties different from the visual properties of each of the plurality of RGB images.

15. In Paragraph 14, The multimodal camera includes an RGB camera that generates a plurality of RGB images by photographing the plant subject at the plurality of shooting points, and a thermal imaging camera that generates a plurality of thermal images as the plurality of multimodal images by photographing the plant subject at the plurality of shooting points. A method for generating a three-dimensional plant point cloud, characterized in that the step of acquiring the plurality of RGB images and the plurality of multimodal images comprises acquiring the plurality of RGB images from the RGB camera and acquiring the plurality of thermal images from the thermal imaging camera.

16. A computer-readable recording medium storing a program for executing the method of claim 1 on a computer.

17. An image acquisition module that acquires multiple RGB images and multiple multimodal images captured at multiple capture viewpoints by a multimodal camera for a plant subject; A reconstruction module that acquires multiple depth images of the plant subject from the acquired multiple RGB images using an artificial intelligence model; and It includes a point cloud generation module that generates an RGB type three-dimensional plant point cloud based on the acquired plurality of RGB images and a multimodal type three-dimensional plant point cloud based on the acquired plurality of multimodal images using a plurality of depth images of the plant subject, A three-dimensional plant point cloud generation system characterized in that each of the plurality of multimodal images is an image having image properties different from the image properties of each of the plurality of RGB images.

Citation Information

Patent Citations

  • Deep learning-based plant rapid three-dimensional rendering and representation extraction tool

    CN117078821A

  • Measurement system

    JP2024118043A

  • Under water surface washing machine for small-sized boat using dry ice

    KR1020210108731A

  • Methods and systems for automated micro farming

    US20180373937A1

  • Image processing

    WO2023021303A1