Information processing device, information processing method, and program

By generating images where objects are colored based on their distance from a reference point, the system addresses misidentification issues in conventional distance imaging, enabling accurate object recognition.

JP2026082322APending Publication Date: 2026-05-19UNIVERSITY OF MIYAZAKI
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
UNIVERSITY OF MIYAZAKI
Filing Date
2024-11-07
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Conventional distance imaging techniques display objects in different colors based on their shooting distances, leading to misidentification of common objects as separate entities, especially when they are captured from varying distances.

Method used

Generate a specific image where objects are displayed in colors corresponding to their distance from a reference object, such as the floor, using position information relative to a shooting position, allowing consistent color representation for objects at similar distances.

Benefits of technology

Facilitates accurate identification of objects by maintaining consistent colors for objects at similar distances, reducing misidentification and simplifying object recognition processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026082322000001_ABST
    Figure 2026082322000001_ABST
Patent Text Reader

Abstract

The present invention provides an information processing device that makes it easier to accurately identify each object in a distance image in which each object is displayed with a color corresponding to the shooting distance. [Solution] The system comprises a generation unit that generates second information in which the positions of each object are represented relative to the position of a reference object, using first information in which the positions of each object photographed from the shooting position are represented relative to the shooting position, and a creation unit that can create a specific image in which each object is displayed in a color corresponding to the distance from the reference object, using the second information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an information processing apparatus, an information processing method, and a program.

Background Art

[0002] Conventionally, a technique has been proposed for capturing a distance image in which an object is displayed in a color corresponding to the shooting distance from the shooting position (the position of the camera) to the object. In the above distance image, for example, an object with a longer shooting distance is displayed in a brighter color, and an object with a shorter shooting distance is displayed in a darker color.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In the above conventional technology, even when a common object is captured, if the shooting distances are different, the colors of the object in the distance image are different. However, even if each object with a different color is actually a common object, there is a situation where it is likely to be identified as a different type of object. Therefore, in the conventional technology, there may be a disadvantage that each common object is identified as a separate object. In view of the above circumstances, an object of the present invention is to make it easier to accurately identify each object in a distance image.

Means for Solving the Problems

[0005] To solve the above problems, the system comprises a generation unit that generates second information in which the positions of each object are represented relative to the position of a reference object, using first information in which the positions of each object photographed from the shooting position are represented relative to the shooting position, and a creation unit that can create a specific image in which each object is displayed in a color corresponding to the distance from the reference object, using the second information.

[0006] In the above configuration, if each common object (e.g., a pig) is located in an area where the distance from a reference object (e.g., the floor) is approximately constant, then each object will be displayed in a common color, even if the shooting positions differ. Therefore, compared to a configuration where the color of a common object changes depending on the shooting distance, it becomes easier to accurately identify each object in a specific image. [Effects of the Invention]

[0007] According to the present invention, it becomes easier to accurately identify each object in a distance image. [Brief explanation of the drawing]

[0008] [Figure 1] This is a diagram illustrating the various components of an information processing system. [Figure 2] This is a hardware configuration diagram of an information processing system. [Figure 3] This is a functional block diagram of an information processing system. [Figure 4] This diagram illustrates the overview of each process. [Figure 5] This is a diagram illustrating a specific example of a generated image. [Figure 6] This is a diagram to illustrate other specific examples of generated images. [Figure 7] This is a diagram to explain the generation process. [Figure 8] This is a flowchart of the processes performed by an information processing device. [Modes for carrying out the invention]

[0009] Figure 1 is a diagram illustrating the various components of the information processing system 1000 in this embodiment. As shown in Figure 1, the information processing system 1000 includes a computer 100 and a camera 200. These components are connected in a communicative manner. The camera 200 outputs captured images Gx to the computer 100. The captured images Gx include various objects (livestock A, floor F, etc.) captured by the camera 200.

[0010] The captured image Gx in this embodiment is a depth image that displays each object in grayscale (black and white). Specifically, each object in the captured image Gx is displayed in a brighter (closer to white) color the further it is from the camera 200, and in a darker (closer to black) color the closer it is to the camera 200. For example, the captured image Gx is captured using LIDAR (Light Detection and Ranging, Laser Imaging Detection and Ranging) technology. The camera 200 is also equipped with a tilt sensor. This tilt sensor detects the magnitude of the tilt of the shooting direction S relative to the vertical direction.

[0011] In this embodiment, when livestock A is photographed by camera 200, a generated image Gy is generated by computer 100 from the captured image Gx representing livestock A. Computer 100 also estimates the weight (volume) of livestock A using the generated image Gy. As a technique for estimating the weight of a photographed object, for example, the technique described in Japanese Patent Application Publication No. 2023-127801 can be employed. In the following embodiments, livestock A (for example, a pig) is used as an example of an object whose weight is estimated (hereinafter referred to as the "target object"), but the target object is not limited to the above example.

[0012] Camera 200 can capture a color image (see Gc in Figure 5 below) simultaneously with the image Gx (distance image). Typically, this color image includes objects other than the target object (e.g., livestock A), such as the floor F. In the conventional techniques described above, it is necessary to identify the target object from each object included in the color image in order to estimate its weight.

[0013] However, depending on the type of object included in the color image, it may not be possible to accurately identify the target object. For example, consider the case where the target object, livestock A, is a so-called "black pig." Livestock A has a body surface that is mostly black in color. Also, bedding material may be used on the floor F where livestock A is raised. This bedding material is prone to fermentation and discoloration to black.

[0014] As shown in Figure 1, when livestock A is photographed from above, the background of livestock A becomes the floor F. In other words, in the specific example in Figure 1, a color image is taken in which the black floor F (bedding) is displayed against the background of the black livestock A (black pig). As will be explained in detail using Figure 5, in the above color image, it is difficult to accurately identify the boundary between livestock A and the floor F (the outer edge of livestock A), and the problem of livestock A not being identifiable is likely to occur.

[0015] Considering the above circumstances, this embodiment employs a configuration that identifies the target object (livestock A) from a generated image Gy, which is a distance image, instead of a color image. As will be described in detail later, each object in the generated image Gy is displayed in a color corresponding to its distance from the floor F (ground level). Therefore, regardless of the actual color of each object, livestock A and the floor F are displayed in different colors in the generated image Gy (distance image). With the generated image Gy described above, even if the background of livestock A is the floor F, the boundary between livestock A and the floor F can be easily and accurately identified, thus suppressing the aforementioned disadvantages that occur when using a color image.

[0016] Incidentally, the above-described captured image Gx is a distance image, just like the generated image Gy. Each object in the above captured image Gx is displayed in a color corresponding to the distance from the camera 200. Also, the distance from the camera 200 to the livestock A and the distance from the camera 200 to the floor F are usually different. Therefore, even in the captured image Gx, the livestock A and the floor F are displayed in different colors, and the outer edge of the livestock A is easy to identify.

[0017] However, in the captured image Gx, even objects that are supposed to display originally common objects (for example, two livestock A) will be displayed in different colors when the distance to the camera 200 is not constant. For example, when the livestock A is photographed from a distance, the livestock A in the captured image Gx is displayed in a color close to white. On the other hand, when the same livestock A is photographed from close by, the livestock A in the captured image Gx is displayed in a color close to black.

[0018] However, when objects that are supposed to be identified as originally common objects are displayed in different colors, there is a situation where it is easy to identify the objects as different objects. Due to the above situation, in the configuration for identifying the target object from the captured image Gx, a new problem may arise that it is difficult to accurately identify the target object from the color of the object.

[0019] In this embodiment, in order to solve the above new problem, a configuration is adopted to identify the target object from the generated image Gy in which each object is displayed in a color corresponding to the distance (height above the ground) from the floor F. The above configuration will be described in detail later.

[0020] FIG. 2 is a hardware configuration diagram of the information processing system 1000. As described above, the information processing system 1000 includes a computer 100 and a camera 200. As shown in FIG. 2, the computer 100 includes a communication device 101, a processing device 102, and a storage device 103. Each of the above components is communicably connected via a system bus.

[0021] The processing unit 102 controls the entire computer 100. The processing unit 102 may consist of one or more processors. Specifically, the processing unit 102 includes a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), an ISP (Image Signal Processor), an NPU (Neural Network Processing Unit), and a DSP (Digital Signal Processor). Alternatively, the processing unit 102 may consist of one or more types of processors, such as an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit).

[0022] The storage device 103 stores various programs, including the program PG. As the storage device 103, known recording media such as semiconductor recording media and magnetic recording media may be used. Furthermore, the storage device 103 may consist of one recording media or multiple recording media.

[0023] The communication device 101 communicates various types of information with external devices. For example, the computer 100 receives the captured image Gx from the camera 200 at predetermined triggers. For example, when a predetermined shooting operation is received, the camera 200 generates (captures) the captured image Gx, which is a distance image, and transmits it to the computer 100. However, the triggers for generating the captured image Gx are not limited to the above examples. For example, the system may be configured to automatically generate the captured image Gx at predetermined time intervals.

[0024] Figure 3 is a functional block diagram of the information processing device 10. For example, the computer 100 functions as the information processing device 10 when the processing device 102 described above executes the program PG. As shown in Figure 3, the information processing device 10 is composed of a generation unit 11, a creation unit 12, and an identification unit 13. In addition, the camera 200 described above functions as the imaging device 20.

[0025] The generation unit 11 uses first information (point cloud information Dc, point cloud information Dw, described later) in which the positions of each object (livestock A, floor F, etc.) captured from the shooting position Ps are represented relative to the shooting position Ps, to generate second information (point cloud information Dg, described later) in which the positions of each object are represented relative to the position of a reference object (for example, floor F). Each of the above point cloud information D(c, w, g) represents the position (coordinates) of each object relative to a predetermined position (shooting position, reference position).

[0026] The creation unit 12 can use the second information to create specific images (generated images Gy1 and Gy2 shown in Figure 6 below) in which each object is displayed in a color corresponding to its distance from the reference object. In this embodiment, the floor F was used as the reference object. The identification unit 13 identifies the target object from each object in the specific image (generated image Gy). Specifically, the identification unit 13 uses a machine learning model to identify the target object (livestock A).

[0027] In the generated image Gy, each livestock A, which is at the same distance from the floor F, is displayed in the same color regardless of its distance to the camera 20 (see Figure 6 for details). Therefore, compared to a distance image (e.g., captured image Gx) where the color of each object changes depending on the distance to the camera 20, the generated image Gy has the advantage of making it easier to accurately identify the target object (livestock A). In addition, the machine learning model (classifier) ​​used by the identification unit 13 is trained using a specific image (generated image Gy) as training data.

[0028] Incidentally, when machine learning models identify (classify) each object, they consider not only the shape of the object but also its color. Therefore, if distance images (for example, captured image Gx) where the same livestock A objects may have different colors are used as training data, a problem may arise where a huge amount of training data is required before the target objects can be accurately identified.

[0029] In the generated image Gy of this embodiment, each target object that commonly represents livestock A is displayed in a common color regardless of the distance from livestock A to the camera 20. Therefore, using the generated image Gy as training data has the advantage of suppressing the aforementioned problems.

[0030] The specific configuration for identifying each object in the generated image Gy may employ any appropriate technology. In this embodiment, the YOLO (You Only Look Once) object detection model and the GrabCut algorithm are used to identify the target object from the generated image Gy.

[0031] Specifically, the information processing device 10 identifies a rectangular region surrounding the target object in the generated image Gy using YOLO technology and sets the center of the rectangular region as a marking point. This marking point can be estimated to be located on the target object. Therefore, in this embodiment, the GrabCut algorithm is used to identify the object containing the marking point as the target object.

[0032] The configuration for identifying target objects from the generated image Gy is not limited to the examples above. For example, a configuration in which each object in the generated image Gy is segmented (divided) may be adopted. Artificial intelligence (AI) technology may be used for the above segmentation. For example, a technology in which each object in the generated image Gy is segmented using a trained FCN (Fully Convolutional Network) (for example, the technology described in Japanese Patent Publication No. 2022-29169) may be adopted.

[0033] Figure 4 is a diagram illustrating the processes (S1 to S4 in Figure 4) for generating a generated image Gy from a captured image Gx acquired from the imaging device 20. As will be explained in detail below, in this embodiment, the captured image Gx is converted into point cloud information Dc. Then, the point cloud information Dc is converted into point cloud information Dw, and the point cloud information Dw is converted into point cloud information Dg. Furthermore, the generated image Gy is created from the point cloud information Dg (i.e., it is generated in the order of Gx → Dc → Dw → Dg → Gy).

[0034] Each of the point cloud pieces D(c, w, g) described above represents the position (coordinates) of the captured object in a different coordinate system (c, w, g). Prior to explaining the processes for generating the generated image Gy from the captured image Gx acquired from the imaging device 20, the coordinate systems (c, w, g) in this embodiment will be explained below.

[0035] The upper part of Figure 4 is a diagram illustrating each coordinate system (c, w, g). The upper part of Figure 4 shows the shooting position Ps (position of the shooting device 20), the shooting direction S, the vertical direction V, and the floor F on which the livestock A is located.

[0036] The first point cloud information Dc generated among the point cloud information D(c,w,g) represents the position of each object in coordinate system c. This coordinate system c uses the shooting position Ps as the reference position (origin), and represents the coordinates of each object (each point constituting the point cloud) using the position on the Zc axis pointing in the shooting direction S (Zc coordinate), the position on the Xc axis perpendicular to the Zc axis (Xc coordinate), and the position on the Yc axis perpendicular to the Zc-Xc plane (Yc coordinate).

[0037] Figure 4 shows the Zc axis of coordinate system c, which points in the direction of photography S, among the coordinate axes (Xc, Yc, Zc). For the purposes of this explanation, the absolute value of the coordinate (zc) on the Zc axis of position P (i.e., the distance from the Xc-Yc plane to position P) may be referred to as the photography distance (zc).

[0038] The point cloud information Dw generated from the point cloud information Dc represents the position of each object in coordinate system w. This coordinate system w uses the shooting position Ps as the reference position (origin) and represents the coordinates of each object (each point constituting the point cloud) using the position on the Zw axis parallel to the vertical direction V (Zw coordinate), the position on the Xw axis perpendicular to the Zw axis (Xw coordinate), and the position on the Yw axis perpendicular to the Zw-Xw plane (Yw coordinate). Figure 4 shows the Zw axis among the coordinate axes (Xw, Yw, Zw) of the coordinate system w.

[0039] The point cloud information Dg generated from the point cloud information Dw represents the position of each object in coordinate system g. This coordinate system g uses the floor F directly below the shooting position Ps as the reference position (origin), and represents the position of each object by its position on the Zg axis parallel to the vertical direction V (Zg coordinate), its position on the Xg axis perpendicular to the Zg axis (Xg coordinate), and its position on the Yg axis perpendicular to the Zg-Xg plane (Yg coordinate). In other words, when converted from coordinate system w to coordinate system g, the reference position representing the position of each object changes from the shooting position Ps to the floor F. Figure 4 shows the Zg axis among the coordinate axes (Xg, Yg, Zg) of coordinate system g.

[0040] The lower part of Figure 4 shows each process (S1 to S4) performed in the process from the captured image Gx to the generated image Gy, and the point cloud information D(c,w,g) created in each process.

[0041] As shown in Figure 4, point cloud information Dc is created from the captured image Gx through a calculation process (S1 in Figure 4). The information processing device 10 in this embodiment uses a so-called pinhole camera model to calculate the coordinates (xc, yc, zc) corresponding to each point (point cloud) in each object in the three-dimensional real space from the captured image Gx (two-dimensional image). The information processing device 10 also creates point cloud information Dc that shows the calculation results.

[0042] Specifically, each pixel in the captured image Gx is assigned a pixel value (small (dark) to large (bright)). As mentioned above, the pixel value (color of the object) of a pixel in the captured image Gx changes according to the shooting distance to the object (position P) in real space that corresponds to that pixel. Therefore, the coordinate (zc) on the Zc axis of each position P in the object in real space can be calculated from the pixel value of the pixel representing that position P. Furthermore, the coordinate (xc) on the Xc axis and the coordinate (yc) on the Yc axis of each position P (1, 2...n...) in the object are calculated using the coordinate (zc) on the Zc axis of that position P from equation (1) in Equation 1 below.

[0043] TIFF2026082322000002.tif2762

[0044] In equation (1), "uc" represents the position in the horizontal axis direction in the captured image Gx. Also, "vc" in equation (1) represents the position in the vertical axis direction in the captured image Gx. "f" in equation (1) represents the focal length f and is a constant determined for each imaging device 20. In the calculation process, the information processing device 10 calculates the coordinates (P1(xc,yc,zc), P2(xc,yc,zc)...Pn(xc,yc,zc)...) of each position P(1, 2...n...) represented by each pixel in the captured image Gx from the pixel value of that pixel, and generates point cloud information Dc. In the point cloud information Dc described above, each coordinate of each position P(1, 2...n...) is represented in coordinate system c (see the upper part of Figure 4).

[0045] The information processing device 10 performs a rotation process (S2 in Figure 4) following the calculation process. In the above rotation process, point cloud information Dw (coordinate system w) is generated from point cloud information Dc (coordinate system c) using the inclination θ of the shooting direction S with respect to the vertical direction V. Specifically, the above inclination θ can also be rephrased as the inclination of coordinate system c in the Zc direction with respect to the Zw direction of coordinate system w (see the upper part of Figure 4). In the rotation process, each coordinate (P1(xc,yc,zc), P2(xc,yc,zc)...Pn(xc,yc,zc)...) indicated by point cloud information Dc is transformed (rotated) into each coordinate (P1(xw,yw,zw), P2(xw,yw,zw)...Pn(xw,yw,zw)...) of coordinate system w by equation (2) in the following equation 2.

[0046] TIFF2026082322000003.tif2470

[0047] In equation (2), "Ry" represents the rotation matrix around the yaw axis. Similarly, "Rp" represents the rotation matrix around the pitch axis, and "Rr" represents the rotation matrix around the low axis. Each of these rotation matrices is calculated from the tilt θ measured by the tilt sensor. Note that well-known rotation matrices can be used as appropriate.

[0048] The information processing device 10 executes a generation process (S3 in Figure 4) to generate point cloud information Dg from point cloud information Dw. Each coordinate of the point cloud information Dg (P1(xg,yg,zg), P2(xg,yg,zg)...Pn(xg,yg,zg)...) represents the position Pn on the object in coordinate system g. The reference position (origin) of coordinate system g is located on the floor F. In other words, in the generation process, the coordinate system representing each position P of the object is transformed from a coordinate system w with the origin at the shooting position Ps to a coordinate system g with the origin located on the floor F. The generation process described above will be explained in detail later using Figure 7.

[0049] The information processing device 10 executes the creation process (S4 in Figure 4) to generate a generated image Gy from the point cloud information Dg. Specifically, the pixel value (color intensity) of each pixel (vertical axis coordinate = ug, horizontal axis coordinate = vg) in the generated image Gy is determined according to the Zg coordinate of the position P corresponding to that pixel (zg (ground height from floor F to position P)). Furthermore, the horizontal axis coordinate (ug) and vertical axis coordinate (vg) in the generated image Gy are calculated using a pinhole camera model by equation (3) shown in Equation 3 below. Note that "f" in equation (3) means the focal length of the imaging device 20, similar to "f" in equation (1) above.

[0050] TIFF2026082322000004.tif3348

[0051] The lower part of Figure 5 is a diagram illustrating a specific example of the generated image Gy produced in this embodiment. On the other hand, the upper part of Figure 5 is a simulated image of the color image Gc. As described above, in the prior art, livestock A was identified from the color image Gc. In the specific example in Figure 5, we assume a color image Gc in which a black livestock A (black pig) is photographed against a black bedding background. In the above color image Gc, a target object Gca representing the black livestock A is displayed against a floor object Gcf representing the black bedding background.

[0052] In the following explanation, the outer edge of the target object may be referred to as "outer edge Ed". As can be seen from Figure 5, in the color image Gc, the color of the target object Gca and the color of the floor object Gcf are almost identical (black), which makes it difficult to distinguish the outer edge Ed of the target object Gca.

[0053] The generated image Gy shown in the lower part of Figure 5, like the color image Gc shown in the upper part of Figure 5 above, assumes a photograph taken from above of a black livestock A (black pig) on ​​black bedding. In this case, the floor object Gyf representing the bedding is displayed in the background, and the target object Gya representing livestock A is displayed in the generated image Gy.

[0054] However, the generated image Gy is a distance image in which each object is displayed in a color corresponding to its distance from the floor F (ground height zg). Also, the ground height (zg) of the bedding is approximately "0 centimeters," while the ground height (zg) of livestock A is higher than "0 centimeters." Therefore, as shown in Figure 5, the color of the floor object Gyf and the color of the target object Gya are different. With the generated image Gy described above, the aforementioned inconvenience of the outer edge Ed of the target object becoming difficult to distinguish is suppressed.

[0055] Figure 6 is a diagram illustrating another specific example of the generated image Gy(1, 2) produced in this embodiment. As described above, the generated image Gy is generated using the captured image Gx received from the imaging device 20. In the specific example in Figure 6, it is assumed that the generated image Gy2 is generated from the captured image Gx1, and the generated image Gy2 is generated from the captured image Gx2.

[0056] In the specific example shown in Figure 6 above, we assume that the target object (livestock A), which has roughly the same height, is photographed at different shooting distances (zc). Specifically, we assume that the ground clearance from the floor F to the position Pn of the target object (distance zg from the floor F to Pn in the Zg axis direction) at the time each image Gx(1, 2) was taken is approximately 70 centimeters. Furthermore, for image Gx1, we assume that the shooting distance (zc) to the position Pn of the target object is approximately 100 centimeters. On the other hand, for image Gx2, we assume that the shooting distance (zc) to the position Pn is approximately 150 centimeters.

[0057] As described above, when the shooting distance (zc) differs, the color of the target object Gxa(1, 2) representing livestock A in the captured image Gx will differ. Specifically, the farther the shooting distance (zc) of an object, the brighter (closer to white) it will appear in the captured image Gx (distance image). Therefore, as shown in Figure 6, compared to the target object Gxa1 in captured image Gx1, where the shooting distance (zc) to livestock A (position Pn) is approximately 100 centimeters, the target object Gxa2 in captured image Gx2, where the shooting distance (zc) to livestock A is approximately 150 centimeters, will appear in a brighter color.

[0058] Furthermore, in the specific example shown in Figure 6, the shooting distance (zc) to the floor F is greater in image Gx2 than in image Gx1. Therefore, the object Gxf2 representing the floor F in image Gx2 is displayed in a brighter color compared to the object Gxf1 representing the floor F in image Gx1.

[0059] On the other hand, as mentioned above, the generated image Gy displays each object in a color corresponding to the height above ground (zg), regardless of the shooting distance (zc). Therefore, as shown in Figure 6, the color of the object Gyf(1, 2) representing the floor F (ground surface) is approximately the same in each generated image Gy(1, 2) regardless of the shooting distance (zc). Also, the color of the target object Gya(1, 2) representing livestock A, whose height above ground (zg) is approximately the same (about 70 centimeters in the example in Figure 6), is approximately the same regardless of the shooting distance (zc).

[0060] As shown in Figure 6, each generated image Gy(1, 2) produced from each captured image Gx(1, 2) is used for identification processing. In the above identification processing, the target object Gya is identified from the generated image Gy using the machine learning model described above. In the generated image Gy of this embodiment, the common livestock A is displayed with the same color regardless of the shooting distance (zc). Therefore, the target object in the generated image Gy is easily and accurately identified.

[0061] In this embodiment, the results of the identification process are used to estimate the weight of the target object (livestock A). For example, a 3D image representing livestock A is generated from the target object Gya (2D image) identified in the identification process, and the weight of livestock A is estimated from the 3D image. As a configuration for generating a 3D image representing livestock A from the target object Gya, a configuration using the pinhole camera model described above (see equation (3) above) can be adopted. Furthermore, as a specific configuration for estimating the weight of the target object from the 3D image, for example, the technology described in Japanese Patent Publication No. 7210862 can be adopted.

[0062] As shown in Figure 6, each generated image Gy, created from each captured image Gx, is also used in the training process. In the training process described above, the machine learning model is trained using the generated images Gy as training data. As mentioned above, in the generated images Gy, the common livestock A is displayed as the target object Gya with the same color. Therefore, compared to a configuration that uses images (e.g., captured images Gx) in which the common livestock A is displayed with different colors as training data, there is an advantage in that the number of training data required to complete the training of the machine learning model can be reduced.

[0063] Figure 7 is a diagram illustrating a specific example of the generation process. As described above, the generation process generates point cloud information Dg from point cloud information Dw. Point cloud information Dw represents each coordinate of each position P of the object in coordinate system w. Point cloud information Dg represents each coordinate of each position P of the object in coordinate system g. Furthermore, the reference position (origin) of coordinate system w is the shooting position Ps, and the reference position of coordinate system g is located at the floor F. Therefore, the generation process can also be described as the process of converting the reference position of each coordinate of point cloud information D from the shooting position Ps to the floor F (ground surface H, described later).

[0064] In the specific example shown in Figure 7, we assume that the distance from the floor F to the shooting position Ps is "Lg," as shown in the left portion. Figure 7 illustrates a horizontal plane hx that passes through a part of the livestock A and is parallel to the floor F. As shown in Figure 7, the distance from the shooting position Ps to the horizontal plane hx is "Lw."

[0065] The central part of Figure 7 is a diagram illustrating the point cloud information Dw generated when livestock A is photographed in the above specific example. As mentioned above, in the point cloud information Dw, each coordinate of each position P(1, 2…n…) is represented in coordinate system w. For the purposes of the following explanation, we will assume that points Pw are placed at each coordinate of the point cloud information Dw in coordinate system w space (XwYwZw space). In addition, the origin of coordinate system w space may be referred to as "origin Pow". The origin Pow of coordinate system w space corresponds to the photographing position Ps in real space.

[0066] Furthermore, the total number of points Pw at each position along the Zw axis may be expressed as "number N". For example, the vertical axis shown in the central part of Figure 7 is the Zw axis. The central part of Figure 7 shows the number N at each position along the Zw axis. As shown in Figure 7, the number N of points Pw located on the horizontal plane hx with Zw coordinate "Lw" is "nx".

[0067] In the generation process of this embodiment, each coordinate on the floor F is identified from each coordinate of the point cloud information Dw, and the ground surface H passing through the floor F is calculated. For example, when livestock A is photographed with the floor F as the background, as shown in Figure 7, each coordinate of point Pw in group ga corresponding to livestock A and group gf corresponding to floor F is included in the point cloud information Dw.

[0068] The information processing device 10 estimates the group gf corresponding to the floor F from each of the above groups g. Specifically, the group furthest from the origin Pow (imaging position Ps) is identified as group gf. The information processing device 10 calculates the ground surface H passing through the floor F using the coordinates included in group gf. Specifically, several points are selected from each point Pw included in group gf, and the ground surface H passing through each selected point Pw is calculated using the least squares method. Note that the configuration for calculating the ground surface H is not limited to the above example.

[0069] The information processing device 10 converts point cloud information Dw into point cloud information Dg using the ground surface H. As described above, point cloud information Dg is generated by representing each coordinate of each position P in point cloud information Dw, which is represented in coordinate system w, in coordinate system g. Specifically, the Zw coordinate (zw) of each position P in point cloud information Dw represents the height (distance in the Zw axis direction) from that position P to the shooting position Ps. In the generation process, the Zw coordinate (zw) of each coordinate of each position P in point cloud information Dw is converted into the height (ground height (zg)) from the ground surface H to that position P.

[0070] For example, the distance from the horizontal plane hx to the shooting position Ps (corresponding to the origin Pow) is "Lw". Therefore, the Zw coordinate of each position P in the horizontal plane hx in coordinate system w is "Lw". On the other hand, the distance from the floor F (corresponding to the ground surface H) to the horizontal plane hx is "Lg-Lw". In this case, the Zg coordinate of each position P in the horizontal plane hx is transformed to "Lg-Lw". As shown in Figure 7, the origin Pog in coordinate system g is located at the ground surface H (floor F).

[0071] Figure 8 is a flowchart of the process executed by the information processing device 10. For example, the information processing device 10 executes the process shown in Figure 8 when a predetermined shooting operation is received.

[0072] As shown in Figure 8, when a shooting operation is received, the information processing device 10 executes an acquisition process (S101). In the acquisition process, the information processing device 10 acquires the captured image Gx taken by the shooting device 20 in response to the shooting operation. After executing the acquisition process, the information processing device 10 executes a calculation process (S102). In the calculation process, point cloud information Dc is generated from the captured image Gx acquired in the previous acquisition process (the coordinates (xc, yc, zc) of each position P are calculated).

[0073] When the information processing device 10 generates point cloud information Dc in the calculation process, it executes a rotation process (S103) to generate point cloud information Dw from the point cloud information Dc. Subsequently, the information processing device 10 executes a generation process (S104) to generate point cloud information Dg from the point cloud information Dw generated in the previous rotation process. The information processing device 10 also creates a generated image Gy from the point cloud information Dg generated in the previous generation process, and then terminates the process shown in Figure 8.

[0074] <Variation> Each of the above forms can be modified in various ways. Specific examples of these modifications are given below. Two or more forms can be arbitrarily selected from the following examples and combined as appropriate.

[0075] In each configuration, the "target object" identified from the generated image Gy is not limited to "livestock A". For example, an item other than a living organism may be used as the "target object". Also, the "reference object" that serves as the reference position for the generated image Gy is not limited to "floor F". For example, a wall photographed in the background of the item may be used as the "reference object". Furthermore, the present invention may be used for purposes other than weight estimation of the "target object". For example, the configuration may be such that a generated image Gy including the item is generated for the purpose of counting the number of items or detecting deformation of the item.

[0076] <Summary of the operation and effects of the embodiment> <First aspect> The information processing device (10) in this embodiment includes a generation unit (11) that generates second information (point cloud information Dg) in which the positions of each object, among those objects, are represented relative to the position of a reference object (floor F), using first information (point cloud information Dc, point cloud information Dw) in which the positions of each object, which are photographed from the shooting position (Ps), are represented relative to the shooting position, and a creation unit that can create a specific image (generated image) in which each object is displayed in a color corresponding to the distance from the reference object, using the second information. According to this embodiment, each object in the distance image in which each object is displayed in a color corresponding to the distance from the reference position becomes easier to accurately identify.

[0077] <Second and Third Embodiments> The information processing device (10) in this embodiment includes an identification unit (13) that identifies a target object from each object in a specific image. The identification unit identifies the target object using a machine learning model, and the machine learning model is trained using the specific image as training data.

[0078] <Fourth aspect> The information processing method of this embodiment comprises the steps of generating second information (S104 in Figure 8) in which the positions of each object photographed from the shooting position are represented with respect to the shooting position, using first information in which the positions of each object are represented with respect to the shooting position, and using second information in which a specific image can be created (S105) in which each object is displayed in a color corresponding to the distance from the reference object. According to this embodiment, the same effects as the first embodiment described above can be achieved.

[0079] <Fifth aspect> The program of this embodiment causes the computer to execute each step of the fourth embodiment. According to this embodiment, the same effects as those of the first embodiment described above are achieved. [Explanation of Symbols]

[0080] 10... Information processing device, 11... Generation unit, 12... Creation unit, 13... Identification unit, 20... Imaging device.

Claims

1. A generation unit generates second information in which the positions of each object, each of which has been photographed from the shooting position, are represented relative to the shooting position, using first information in which the positions of each object are represented relative to the position of a reference object among the respective objects. A creation unit capable of creating a specific image in which each of the objects is displayed in a color corresponding to the distance from the reference object, using the second information described above. An information processing device equipped with the following.

2. Identification unit that identifies the target object from each object in the specified image. The information processing apparatus according to claim 1, comprising:

3. The identification unit identifies the target object using a machine learning model, The aforementioned machine learning model is trained using the aforementioned specific image as training data. The information processing apparatus according to claim 2.

4. The process involves generating second information using first information in which the positions of each object photographed from the shooting position are represented relative to the shooting position, and second information in which the positions of each object are represented relative to the position of a reference object among the objects. A step of creating a specific image in which each of the objects is displayed in a color corresponding to the distance from the reference object, using the second information described above. An information processing method comprising the following.

5. A program that causes a computer to perform each of the steps described in claim 4.