Object shape or height recognition device, program, and method
A two-dimensional image-based system creates point cloud data to estimate object shape and height using polynomial approximation and clustering, addressing the challenge of accurate estimation without a three-dimensional camera.
Patent Information
- Application Number
- JP2024000510
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-01-05
- Publication Date
- 2025-07-17
AI Technical Summary
Existing technologies struggle to accurately estimate the shape or height of an object using only two-dimensional photography without a three-dimensional camera.
A device and method that utilizes a two-dimensional moving image capturing device to acquire data, extracts image data, creates point cloud data, and estimates the shape or height of an object using polynomial approximation and clustering techniques.
Enables accurate estimation of object shape and height from two-dimensional images, overcoming the limitations of three-dimensional camera dependency.
Smart Images

Figure 2025106912000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to an object shape or height recognition device, program, and method.
Background Art
[0002] Conventionally, the development of technologies for recognizing the shape of nails by images has been underway. For example, Patent Document 1 (Japanese Unexamined Patent Application Publication No. 2017-018158) discloses a nail shape recognition technology for identifying the shape (outline) of nails with a three-dimensional camera.
[0003] Also, conventionally, the development of technologies for estimating the height of an object by images has been underway. For example, Patent Document 2 (Japanese Unexamined Patent Application Publication No. 2020-178659) discloses a device for estimating the height of agricultural crops with a stereo camera.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Patent Document 2
Summary of the Invention
Problems to be Solved by the Invention
[0005] When there is no three-dimensional camera and only two-dimensional photography is possible, it is required to easily and highly accurately read the shape or height of an object.
Means for Solving the Problems
[0006] The invention according to the first aspect includes an input unit that acquires a moving image of an object and its surroundings from a two-dimensional moving image capturing device or a terminal equipped with a two-dimensional moving image capturing device, a storage unit that stores the moving image acquired by the input unit, an image data extraction unit that extracts a plurality of pieces of image data from the moving image stored in the storage unit, a shape estimation unit that creates point cloud data from the image data, reads the object from the point cloud data, and estimates the shape or height of the object from the read object, and an output unit that outputs the estimated shape or height.
Brief Description of the Drawings
[0007]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Embodiments for Carrying Out the Invention
[0008] Hereinafter, embodiments of an object shape or height recognition device according to the present disclosure will be described together with the drawings.
[0009] (First Embodiment) <1. Configuration of the Nail Shape Recognition Device> FIG. 1 is a schematic diagram showing the configuration of a nail shape recognition device 30 according to the present embodiment. As its configuration, it includes a camera or a terminal device 10 with a camera, and a processing unit 20. Here, the camera or the terminal device 10 with a camera may be a nail shape recognition device 30 that is completely independent of the processing unit 20 (in the following embodiments, the processing unit 20 is read as the nail shape recognition device 30), or the camera or the terminal device 10 with a camera may be a nail shape recognition device 30 integrated with the processing unit 20. In the present embodiment, the camera or the terminal device 10 with a camera will be described for a nail shape recognition device 30 that is completely independent of the processing unit 20.
[0010] The camera or the terminal device 10 with a camera is a camera capable of shooting a two-dimensional video or a terminal device with a camera capable of shooting a two-dimensional video. Here, taking the terminal device 10 as an example, it is a smartphone with a camera capable of shooting a two-dimensional video.
[0011] The processing unit 20 can be realized by an arbitrary computer, and includes a storage unit 21, an input unit 22, an image data extraction unit 23, a shape estimation unit 24, and an output unit 25.
[0012] The storage unit 21 stores various information and is realized by an arbitrary storage device such as a memory and a hard disk. Here, the storage unit 21 stores the back of the hand including the nails photographed above and the video around it.
[0013] The input unit 22 is realized by any input device such as a keyboard, a mouse, or a touch panel, and inputs various types of information into the computer. Here, the input unit may be an external terminal instead of being integrated with the computer.
[0014] The image data extraction unit 23 extracts a plurality of pieces of image data from a video, and is created by software such as a Python library for image processing, for example.
[0015] The shape estimation unit 24 creates point cloud data from a plurality of pieces of image data and estimates the shape of the nail. Here, the point cloud data is created by software such as a Python library for image processing, for example.
[0016] The output unit 25 is realized by any output device such as a display, a touch panel, or a speaker, and outputs various types of information from the computer. Here, the output unit may be an external terminal instead of being integrated with the computer.
[0017] <2. Operation of the Nail Shape Recognition Device> The operation of the nail shape recognition device 30 according to the present embodiment will be described. Three-dimensional point cloud data is created from a video of the back of a hand with nails placed on a plane taken by the camera of the smartphone of the terminal device 10, and the shape (outline) of the nail is approximated by a curved surface expressed by a polynomial.
[0018] (1) First, the terminal device 10 captures video data showing the back of the hand (Fig. 2) and inputs the video data into the processing unit 20 via the input unit 22. Here, the connection between the terminal device 10 and the processing unit 20 may be wireless or wired.
[0019] (2) Next, the image data extraction unit 23 creates a plurality of pieces of image data from the video data (Fig. 3). At this time, images that are out of focus are automatically excluded. Also, in order to create an accurate point cloud, ensure that there is a sufficient amount of overlapping area among the data shown in consecutive photos.
[0020] (3) Extract pixels representing the nail region from the image data (Fig. 4). Since there is a group of photos for each nail, if the pixels are extracted, it will be possible to know which part is the nail even when it becomes three-dimensional point cloud data.
[0021] (4) In the shape estimation unit 24, first, create point cloud data from the group of photos created in (2) (Fig. 5). At this time, using the data created in (3), mark the points representing the nails in the point cloud data.
[0022] (5) Remove the noise included in the data by finding outliers from the point cloud data (Fig. 5).
[0023] (6) Extract the plane on which the hand is placed from the point cloud data, and estimate the normal vector perpendicular to the plane (Fig. 5). Also, estimate the upper side of the data from the positional relationship between the hand and the plane. Here, in the three-dimensional point cloud data itself, the up and down of the data are not known, so the upper side of the data is estimated by the plane to confirm the presence of the nails. Also, separate each finger (nail) by performing clustering on the points with each marking.
[0024] (7) For each nail, perform approximation with a curved surface (Fig. 6). (a) Specify the finger for which the curved surface approximation is to be performed. (i) Perform principal component analysis on the point cloud representing the nail of the specified finger. (u) Consider the eigenvector related to the smallest eigenvalue (the direction with the smallest variance) as the normal vector to the nail, and consider the direction orthogonal to it as the average plane of the nail. This is because the normal vector of the nail is not necessarily in the same direction as the normal vector estimated in (6). (e) Rotate the orientation of the image so that the normal vector specified in (u) faces upward (faces the positive direction of the Z-axis in the XYZ coordinate axes). (o) Draw a straight line from the lower side of the nail, connect each point representing the nail, and remove the noise so that only the points outside the surface of the nail remain. (a) The reference point is the point P(xavg, yavg, z) represented by using the Z coordinate z = 2*zmin - zmax calculated from the average value (xavg, yavg) of the XY coordinates of the point group representing the claw and the maximum value zmax and the minimum value zmin of the Z coordinate of the point group representing the claw. Here, z = zmin - (zmax – zmin) where the value is lower than the minimum value of the claw by the width of the claw is used as the Z coordinate z. (b) Generate random points uniformly distributed on the upper hemisphere centered at P. Let this point group be the random point group Qi (i = 1, 2, …, n). (c) Let each point of the point group representing the claw be Xk (k = 1, 2, ..., N). For each Xk, find the point in the random point group Qi (i = 1, 2, …, n) that is at the closest distance to Xk. When the point at the closest distance is Qj, the point Xk is considered to belong to the point Qj. (d) For each i = 1, 2, …, n, let the maximum value of the distance from P among the point group belonging to Qi be Ri. Among the points belonging to Qi, those with a distance from P less than 0.9 * Ri are regarded as noise and deleted. (e) Use the function expressed as z = f(x, y) to obtain the function of the claw surface. For example, a two-variable three-dimensional polynomial (Equation 1)
[0025] f(x,y)=ax 3 +bx 2 +cx+dy 3 +ey 2 +fy+gx 2 y+hxy 2 +kxy+m a~m: real numbers This can be achieved by performing Ridge regression using the above. Since the rotation is performed in (e) so that the surface of the claw faces upward, it is possible to use such a functional form except for the case of extremely curled claws. The specific method of Ridge regression is as follows. (A) The coordinates of each point of the point group representing the surface of the claw are (x i , y i , z i(Let \(i = 1, 2, \cdots, N\).) (B) (Equation 2)
[0026] TIFF2025106912000002.tif9160 Select \(a, b, c, d, e, f, g, h, k, m\) to minimize the value of \(\cdots\). Here, \(\lambda\) is a constant of the regularization term to prevent \(a, b, c, d, e, f, g, h, k, m\) from becoming too large.) (C) The solution to this problem is (Equation 3)
[0027] TIFF2025106912000003.tif29170 is given by (Equation 4)
[0028] TIFF2025106912000004.tif2766 as (E) By considering the convex hull including projecting the function of the curved surface of the rotated claw onto the XY plane, estimate the two-dimensional contour of the claw part (Figure 7).)
[0029] <3. Effects> It becomes possible to estimate the shape considering the height and curvature of the claw only by taking a video.)
[0030] (Second Embodiment) <1. Configuration of the Object Height Estimation Device> FIG. 8 is a schematic diagram showing the configuration of the object height estimation device 60 according to the present embodiment. As its configuration, it includes a camera or a terminal device 40 with a camera, and a processing unit 50. Here, the camera or the terminal device 40 with a camera may be an object height estimation device 60 that is completely independent of the processing unit 50 (in the following embodiments, the processing unit 50 is read as the object height estimation device 60), or the camera or the terminal device 40 with a camera may be an object height estimation device 60 integrated with the processing unit 50. In the present embodiment, the camera or the terminal device 40 with a camera will be described as an object height estimation device 60 that is completely independent of the processing unit 50.
[0031] The camera or the terminal device 40 with a camera is a camera capable of shooting a two-dimensional video or a terminal device with a camera capable of shooting a two-dimensional video. Here, as an example of the terminal device 40, it is a smartphone with a camera capable of shooting a two-dimensional video.
[0032] The processing unit 50 can be realized by an arbitrary computer, and includes a storage unit 51, an input unit 52, an image data extraction unit 53, a height estimation unit 54, and an output unit 55.
[0033] The storage unit 51 stores various information and is realized by an arbitrary storage device such as a memory and a hard disk. Here, the storage unit 51 stores the table including the plastic bottle photographed above and the video around it.
[0034] The input unit 52 is realized by an arbitrary input device such as a keyboard, a mouse, and a touch panel, and inputs various information to the computer. Here, the input unit may be an external terminal instead of being integrated with the computer.
[0035] The image data extraction unit 53 extracts a plurality of pieces of image data from the video and is created by software such as a Python library for image processing, for example.
[0036] The height estimation unit 54 creates point cloud data from a plurality of image data and estimates the height of an object. Here, for example, the point cloud data is created using software such as a Python library for image processing.
[0037] The output unit 55 is realized by an arbitrary output device such as a display, a touch panel, a speaker, etc., and outputs various information from the computer. Here, the output unit may be an external terminal instead of being integrated with the computer.
[0038] <2. Operation of the Object Height Estimation Device> The operation of the object height estimation device 60 according to this embodiment will be described. Create 3D point cloud data from a video in which an object whose height is to be measured is shown by the camera of the terminal device 40, and estimate the height of each point included in the point cloud based on a base whose height is known in advance and shown in the video.
[0039] (1) First, measure the height of the base A, which is the height H of the plastic bottle in this embodiment. The terminal device 40 captures video data in which a plastic bottle and an object whose height is to be measured are shown (Fig. 9), and inputs the video data to the processing unit 50 via the input unit 52. Here, the connection between the terminal device 40 and the processing unit 50 may be wireless or wired.
[0040] (2) Next, the image data extraction unit 53 creates a plurality of image data from the video data (Fig. 10). At this time, images that are out of focus are automatically excluded. Also, in order to create accurate point cloud data, make sure that there is a sufficient amount of overlapping area among the data shown in consecutive photos.
[0041] (3) Extract the pixels representing the plastic bottle from the image data (Fig. 11). Since there is a group of photos of the plastic bottle, if the pixels are extracted, it will be known which part is the plastic bottle even when the 3D point cloud data is formed.
[0042] (4) The height estimation unit 54 first creates point cloud data from the group of photos created in (2) (Fig. 12). At this time, using the data created in (3), the points representing the PET bottle in the point cloud data are marked.
[0043] (5) By searching for isolated points from the point cloud data, the noise contained in the data is removed (Fig. 12).
[0044] (6) A plane and a normal vector perpendicular to the plane are estimated from the point cloud data. Specifically, as follows. (a) For each point in the point cloud, the point cloud in the vicinity of each point is extracted. (i) Perform principal component analysis on the vicinity of each point, and set the direction of the point cloud with the smallest spread as the normal direction of that point. (u) Clustering is performed regarding the direction of the normal vector. (e) The points included in the cluster containing the most points are regarded as points on the plane. (o) The average value of the normal vectors at the points regarded as being on the plane is set as the height direction (the normal vector of the plane).
[0045] (7) By calculating the inner product of the normal vector calculated in (6) and each point in the point cloud, the height component of each point on the point cloud data coordinates is calculated.
[0046] (8) By performing clustering on the height components of each point on the plane, planes with different heights are separated. In Fig. 13, the plane is divided into three clusters by this clustering. In this embodiment, the points in the cluster containing the most points are regarded as the floor. Because the floor is the plane that appears most in the photo.
[0047] (9) Calculate the difference between the maximum value and the minimum value of the height component from the point cloud data corresponding to the PET bottle marked in (4). The magnification ratio between the point cloud data and the actual object is calculated from this value. (Magnification ratio) = (Actual height H of the PET bottle) ÷ (Maximum value of the height component of the PET bottle - Minimum value of the height component of the PET bottle)
[0048] (10) Calculate the height of each point by comparing the height component of the bed coordinates with the height component of the coordinates of each point. (Height of each point) = (Height component of each point - Average value of the height components of the bed) × (Magnification ratio)
[0049] (11) Mark the points whose height falls within the threshold range.
[0050] <3. Effects> It becomes possible to estimate the height of an object just by taking a two-dimensional video.
[0051] (Other Embodiments) The present disclosure is not limited to the above-described embodiments as they are. The present disclosure can be embodied by modifying the components without departing from the gist thereof at the implementation stage. Also, the present disclosure can form various disclosures by appropriately combining a plurality of components disclosed in the above-described embodiments. For example, some components may be deleted from all the components shown in the embodiments. Furthermore, components may be appropriately combined from different embodiments.
Description of Reference Numerals
[0052] 10 Terminal device 20 Processing unit 21 Storage unit 22 Input unit 23 Image data extraction unit 24 Shape estimation unit 25 Output unit 40 Terminal device 50 Processing unit 51 Storage unit 52 Input unit 53 Image data extraction unit 54 Height estimation unit 55 Output unit
Claims
1. An input unit that acquires a video of an object and its surroundings from a two-dimensional video shooting device or a terminal equipped with a two-dimensional video shooting device; A storage unit that stores the video acquired by the input unit; An image data extraction unit that extracts a plurality of pieces of image data from the video stored in the storage unit; A shape estimation unit that creates point cloud data from the image data, reads an object from the point cloud data, and estimates the shape or height of the object from the read object; An output unit that outputs the estimated shape or height; An object shape or height recognition device comprising the above.
2. In the object shape or height recognition device according to Claim 1, An input unit that acquires a video of the back of a hand including nails and its surroundings from a two-dimensional video shooting device or a terminal equipped with a two-dimensional video shooting device; A storage unit that stores the video acquired by the input unit; An image data extraction unit that extracts a plurality of pieces of image data from the video stored in the storage unit; A shape estimation unit that creates point cloud data from the image data, reads the back of the hand from the point cloud data, approximates the nail with a curved surface from the read back of the hand, and estimates the nail shape; An output unit that outputs the estimated nail shape; A nail shape recognition device comprising the above.
3. The nail shape recognition device according to Claim 2, wherein the two-dimensional video shooting device or the terminal equipped with the two-dimensional video shooting device is integrated with the nail shape recognition device.
4. The nail recognition device according to Claim 2, wherein the input unit and the output unit are outside the nail shape recognition device.
5. The nail shape recognition device according to Claim 2, wherein the periphery of the back of the hand is a plane and the back of the hand is on the plane.
6. The nail shape recognition device according to Claim 2, wherein Ridge regression is performed when approximating the nail with a curved surface.
7. A program for shooting a video of the back of a hand including nails and its surroundings with a two-dimensional shooting device or a terminal equipped with a two-dimensional shooting device, extracting a plurality of pieces of image data from the shot video, creating point cloud data from the image data, reading the back of the hand from the point cloud data, approximating the nail with a curved surface from the read back of the hand, and estimating the nail shape.
8. A nail shape recognition method for a two-dimensional imaging device or a terminal equipped with a two-dimensional imaging device, which captures a video of the back of a hand including nails and its surrounding area, extracts a plurality of image data from the captured video, creates point cloud data from the image data, reads the back of the hand from the point cloud data, approximates the nails with a curved surface from the read back of the hand, and estimates the nail shape.
9. In the object shape or height recognition device according to Claim 1, an input unit that acquires a video in which a base body of a known height and an object whose height is to be estimated are shown from a two-dimensional video imaging device or a terminal equipped with a two-dimensional video imaging device; a storage unit that stores the video acquired by the input unit; an image data extraction unit that extracts a plurality of image data from the video stored by the storage unit; a height estimation unit that creates point cloud data from the image data, reads the height of the base body in the point cloud data, and estimates the actual height of the object shown in the video from the read height of the base body; an output unit that outputs the estimated actual height of the object; An object height estimation device comprising:
10. The object height estimation device according to Claim 9, wherein the two-dimensional video imaging device or the terminal equipped with the two-dimensional video imaging device is integrated with the object height estimation device.
11. The object height estimation device according to Claim 9, wherein the input unit and the output unit are outside the object height estimation device.
12. A program for a two-dimensional imaging device or a terminal equipped with a two-dimensional imaging device, which captures a video in which a base body of a known height and an object whose height is to be estimated are shown, extracts a plurality of image data from the captured video, creates point cloud data from the image data, reads the height of the base body in the point cloud data, and estimates the actual height of the object shown in the video from the read height of the base body.
13. A method for a two-dimensional imaging device or a terminal equipped with a two-dimensional imaging device, which captures a video in which a base body of a known height and an object whose height is to be estimated are shown, extracts a plurality of image data from the captured video, creates point cloud data from the image data, reads the height of the base body in the point cloud data, and estimates the actual height of the object shown in the video from the read height of the base body.
Citation Information
Patent Citations
Three-dimensional nail arm modelling method
JP2017018158A
Harvester
JP2020178659A