Object learning detection system, object learning system, object detection system, object learning method, and computer program

The object learning detection system addresses the challenge of detecting individual trees in densely packed forests by training a model on specific tree regions, enhancing detection accuracy and reliability.

JP2025110811APending Publication Date: 2025-07-29KYOTO UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024004869
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-01-16
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

Existing AI-based instance segmentation methods struggle to reliably detect individual trees in densely packed forests due to the high density of trees of the same type, which complicates accurate identification and classification.

Method used

An object learning detection system that acquires learning data including regions with predetermined distance between adjacent points around each tree, trains an instance segmentation model, and detects tree regions and types using a learned model.

Benefits of technology

Enhances the reliability of AI-based detection of individual trees in densely forested areas by accurately distinguishing and classifying tree types, improving detection accuracy and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025110811000001_ABST
    Figure 2025110811000001_ABST
Patent Text Reader

Abstract

To provide an object learning detection system that uses an AI to more reliably detect individual objects from photographs of places where same type of objects are densely packed, such as forest trees, than a conventional system.SOLUTION: An object learning detection system comprises steps of: acquiring, by a computer 1, learning data including a learning image, an area surrounding each learning target object appearing in the learning image, a distance between two adjacent points among a plurality of designated points being within a predetermined range, and a type of the learning object; generating a learned model by training, on the basis of the learning data, an instance segmentation model; and detecting, on the basis of the learned model, an area and a type of each detection target object appearing in an inference image.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to AI (Artificial Intelligence) technology for detecting an object shown in a photograph.

Background Art

[0002] Conventionally, a method has been proposed in which information on a forest is acquired by special hardware such as a hyperspectral camera, a multispectral camera, or a LiDAR (Light Detection and Ranging) sensor, and the types of trees growing in the forest are specified based on the acquired information.

[0003] However, since these hardware devices are expensive, this method incurs high costs. Therefore, a method of inferring the types of trees by AI has been proposed.

[0004] For example, the inventions described in Patent Documents 1 and 2 detect trees by instance segmentation as follows.

[0005] The invention described in Patent Document 1 is a method for segmenting an example of a street tree image based on Mask R-CNN, and includes the following steps. First, a plurality of street tree images are selected as samples and labeled to obtain a street tree mask image. The original street tree images and street tree mask images of all samples, including size conversion and data augmentation, are preprocessed, and all processed images are obtained as a sample set. All samples in the sample set are input into a street tree example segmentation model based on Mask R-CNN for learning to obtain a learned model. Then, a street tree image is acquired in real time, size conversion is performed on the original street tree image so that the size of the original street tree image matches the size to be size-converted, the learned model is input, and a segmentation result is obtained.

[0006] The invention described in Patent Document 2 includes an agricultural machine that moves in a field and is equipped with an image sensor that captures images of the plants in the field. A control system accesses the captured images and applies those images to a machine-learned plant identification model. The plant identification model identifies the pixels representing the plants and classifies the plants into plant groups (such as plant species). The identified pixels are labeled as a plant group, and the positions of the pixels are determined. The control system activates a processing mechanism based on the identified plant group and position. The plant identification model uses an instance segmentation method.

Prior Art Documents

Patent Documents

[0007]

Patent Document 1

Patent Document 2

Summary of the Invention

Problems to be Solved by the Invention

[0008] Instance segmentation is an AI technology that detects each individual from a photo in which a number of individuals are shown and determines their types.

[0009] However, in forests, there are often many trees of the same type densely packed. Therefore, simply adopting instance segmentation alone makes it difficult to reliably detect each individual tree in the forest. This is one of the problems of instance segmentation.

[0010] The invention described in Patent Document 1 is for detecting street trees shown in a photograph taken horizontally of a road on the ground. In such a photograph, street trees are rarely shown densely and are shown at intervals. The invention described in Patent Document 2 is for detecting plants shown in a photograph of a field. In such a photograph, the same type of plants are rarely shown densely either.

[0011] That is, the inventions described in Patent Documents 1 and 2 both adopt instance segmentation, but do not aim to detect each tree from a photograph of a place where the same type of trees are densely located, such as a forest. Also, there is no means for solving the above problems.

[0012] In view of such circumstances, the object of the present invention is to more reliably perform AI-based detection of each object from a photograph of a place where the same type of objects, such as forest trees, are densely located than in the prior art.

Means for Solving the Problems

[0013] An object learning detection system according to one embodiment of the present invention includes acquisition means for acquiring learning data including a learning image, a region in which the distance between two adjacent points among a plurality of designated points surrounding each learning target object shown in the learning image is within a predetermined range, and the type of the learning target object, learning means for generating a learned model by training a model of instance segmentation based on the learning data, and detection means for detecting the region and type of each detection target object shown in an inference image based on the learned model.

[0014] An object learning detection system according to another aspect of the present invention includes acquisition means for acquiring learning data including a learning image and a region in which the distance between two adjacent points among a plurality of designated points surrounding each of a predetermined type of learning target object shown in the learning image is within a predetermined range, learning means for generating a learned model by training a model of instance segmentation based on the learning data, and detection means for detecting a region of each of the predetermined type of detection target objects shown in an inference image based on the learned model.

[0015] An object learning detection system according to another aspect of the present invention includes learning image acquisition means for acquiring a learning image in which a plurality of learning target objects belonging to any one of a plurality of types are shown, first discrimination means for discriminating a first distribution region of the learning target objects belonging to each of the plurality of types by discriminating the type of the learning target object shown in each pixel or pixel block of the learning image by semantic segmentation, acquisition means for acquiring learning data including, for each of the plurality of types, an image of the first distribution region of that type in the learning image and a region in which the distance between two adjacent points among a plurality of designated points surrounding each of the learning target objects shown in the image is within a predetermined range, learning means for generating a learned model for each of the plurality of types by training a model of instance segmentation based on the learning data of that type, second discrimination means for discriminating a second distribution region of the detection target objects belonging to each of the plurality of types by discriminating the type of the detection target object shown in each pixel or pixel block of the inference image by semantic segmentation, and detection means for detecting the detection target objects shown in the image of the second distribution region of that type in the inference image for each of the plurality of types based on the learned model of that type.

Advantages of the Invention

[0016] According to the present invention, the AI detection of each object can be performed more reliably than before from a photograph of a place where the same type of objects such as forest trees are concentrated.

Brief Description of the Drawings

[0017]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Embodiments for Carrying Out the Invention

[0018] 〔1. Overall Configuration〕 FIG. 1 is a diagram showing an example of the overall configuration of the tree learning detection system 3. FIG. 2 is a diagram showing an example of the hardware configuration of the computer 1. FIG. 3 is a diagram showing an example of the functional configuration of the computer 1.

[0019] As shown in FIG. 1, the tree learning detection system 3 is composed of a computer 1, a drone 2, etc. The tree learning detection system 3 provides a service for detecting trees in a forest shown in a photo taken from above by AI (Artificial Intelligence) and estimating the number of trees, etc.

[0020] The drone 2 is a UAV (Unmanned Aerial Vehicle) equipped with a digital camera, and is used to acquire color photos that serve as the basis for learning data by photographing the forest from above. It is also used to acquire color photos of the forest that are the objects for inferring the type and number of trees, etc. The drone 2 can be a commercially available one. For example, PHANTOM 4 of DJI is used.

[0021] As shown in FIG. 2, the computer 1 is composed of a processor 10, a RAM (Random Access Memory) 11, a ROM (Read Only Memory) 12, an auxiliary storage device 13, a network adapter 14, a keyboard 15, a pointing device 16, an input / output board 17, a touch panel display 18, and an audio output unit 19, etc.

[0022] In addition to the operating system, various programs are installed in the ROM 12 or the auxiliary storage device 13. In particular, in this embodiment, a tree detection program 40 (see FIG. 3) is installed. As the auxiliary storage device 13, an SSD (Solid State Drive) or a hard disk, etc. is used.

[0023] RAM 11 is the main memory of computer 1. Programs such as the tree detection program 40 in addition to the operating system are loaded into RAM 11.

[0024] The processor 10 executes the programs loaded into RAM 11. As the processor 10, a GPU (Graphics Processing Unit) or a CPU (Central Processing Unit) etc. is used.

[0025] The network adapter 14 is a device for communicating with other devices such as drone 2 using a protocol such as TCP / IP (Transmission Control Protocol / Internet Protocol).

[0026] The keyboard 15 and the pointing device 16 are input devices for the operator to input commands or data etc.

[0027] The input / output board 17 communicates with drone 2 via wire or wirelessly. As the input / output board 17, for example, an input / output board compliant with USB (Universal Serial Bus) or Bluetooth is used.

[0028] The touch panel display 18 displays a screen for inputting commands or data or a map etc. generated by the processor 10.

[0029] The voice output unit 19 is composed of a voice board and a speaker etc. and outputs voices such as warning sounds.

[0030] The tree detection program 40 is a computer program for realizing an annotation reception unit 401, a learning data acquisition unit 402, a learning data storage unit 403, a model storage unit 404, a machine learning unit 405, a tree detection unit 406, a tree count estimation unit 407, and a result display unit 408, etc., as shown in FIG. 3. According to the tree detection program 40, a learned model for detecting trees can be generated, and based on this learned model, trees can be detected from aerial photos of a forest. Furthermore, based on the detection results, the number, density, and volume of trees in the forest are estimated.

[0031] Hereinafter, taking the case of tuning and using the tree learning detection system 3 so that trees of several specific types (for example, cedar, cypress, pine, elm, hemlock, etc.) can be detected as an example, the processing of each part shown in the drone 2 and FIG. 3 will be roughly classified into a learning phase and an inference phase and described. A unique label 42 is pre-associated with each of these specific types.

[0032] 〔2. Learning phase〕 〔2.1 Network and data〕 FIG. 4 is a diagram showing an example of an aerial photo 60. FIG. 5 is a diagram showing an example of an overall photo 6. FIG. 6 is a diagram showing an example of the designation of an area 43 for the overall photo 6. FIG. 7 is a diagram showing an example of an overall photo 6A.

[0033] In the computer 1, a tree detection network 51 is pre-stored in the model storage unit 404 (see FIG. 3). The tree detection network 51 is a neural network for detecting trees shown in the input photo. In this embodiment, YOLO (You Only Look Once) v8 of Ultralytics is used as the tree detection network 51. YOLOv8 is an instance segmentation model that can detect a large number of objects shown in the input photo at once.

[0034] In addition to the tree detection network 51, some functions of the annotation reception unit 401 to the result display unit 408 described later can be realized by incorporating or calling tools such as YOLOv8 provided by Ultralytics via GitHub.

[0035] The operator flies the drone 2 over the forest 80 and takes an aerial photo 60 as shown in FIG. 4 by shooting the forest 80 from a predetermined altitude. Note that the aerial photo 60 is an RGB color photo. When the forest 80 is large, a plurality of aerial photos 60 are obtained by shooting while moving the drone 2 horizontally little by little. Then, these aerial photos 60 are aligned and arranged as shown in FIG. 5 to generate a photo of the entire forest 80 (hereinafter referred to as "entire photo 6") and input it to the computer 1.

[0036] More specifically, a plurality of aerial photos 60 are obtained by shooting the forest 80 a plurality of times while moving the drone 2 from an altitude of about 40 m. The aerial photo 60 shows a portion of the forest 80 that is approximately 20 m × 20 m. Then, these aerial photos 60 are aligned and arranged to generate the entire photo 6 and input it to the computer 1. One pixel of the entire photo 6 corresponds to a portion of the forest 80 that is approximately 0.63 cm × 0.63 cm.

[0037] Furthermore, the operator performs the same operation for other forests to generate respective entire photos 6 and inputs these entire photos 6 to the computer 1.

[0038] Then, in the computer 1, the annotation reception unit 401 and the learning data acquisition unit 402 execute the process for acquiring learning data as follows.

[0039] The annotation reception unit 401 displays the first one of the input plurality of entire photos 6 on the touch panel display 18.

[0040] Here, the operator annotates the area 43 and the type label 42 of each tree shown in the entire photo 6. Annotation generally means "annotation", but particularly in the field of AI, it means associating information such as tags or metadata with various data such as text, audio, images, or videos in order to generate teacher data. In the present embodiment, the area 43 of each tree is specified by tracing the contour of each tree crown so as to be surrounded by a plurality of sides as shown in FIG. 6. Tracing in this way is equivalent to specifying the end points of each side. It is also equivalent to specifying a plurality of points representing the area 43 and two adjacent points. Although it is desirable to trace accurately, trace so that one side has a length within a predetermined range (for example, 8 pixels or more and 12 pixels or less). Thereby, the area 43 of each tree is specified by a polygon in which one side has a length within a predetermined range. Further, the label 42 of each tree is specified by selecting the label 42 associated with the type of that tree.

[0041] Thereby, the position and image of the area 43 of each tree are associated with the entire photo 6, and further, the label 42 of the type of that tree is associated with the area 43 of each tree. Note that each side of the area 43 does not have to be a strict straight line. That is, it may be a straight curve like traced freehand, or a straight line between adjacent inflection points.

[0042] Note that the number of sides shown in FIG. 6 is less than that in the actual tracing due to the resolution allowed in the drawings attached to the application. Actually, one tree crown is traced by about 2 to 3 times the number of sides shown in FIG. 6. The same applies to FIG. 7.

[0043] By the way, at the edge of the overall photo 6, there may be trees in which only a part of the tree crown is shown. The operator designates an image of such a tree as a part to be trimmed (cut) from the overall photo 6. However, images of objects that are not the tree crown (for example, the ground, roads, buildings, automobiles, grass, fallen trees, etc.) may be left without being trimmed. In this case, a label 42 corresponding to the object that is not the tree crown is designated.

[0044] Note that the trees in which only a part of the tree crown is shown may be trimmed first, and then the tree crowns of the remaining trees may be traced.

[0045] Then, the annotation reception unit 401 receives the content of the annotation and the part to be trimmed.

[0046] Then, the learning data acquisition unit 402 generates an overall photo 6A as shown in FIG. 7 by trimming (cutting) the received part from the overall photo 6, and generates learning data 41 that sets the overall photo 6A and the content of the annotation of each tree (such as the label 42 and the position and shape of the region 43), and stores it in the learning data storage unit 403.

[0047] The annotation reception unit 401, the learning data acquisition unit 402, and the operator perform the same processing and operations on the remaining overall photo 6. As a result, each learning data 41 is stored in the learning data storage unit 403.

[0048] Note that the computer 1 may perform the process of merging the aerial photo 60 to generate the overall photo 6.

[0049] 〔2.2 Learning〕 The machine learning unit 405 uses a part (e.g., 70%) of the plurality of learning data 41 stored (saved) in the learning data storage unit 403 to train the tree detection network 51 stored in the model storage unit 404 by the YOLOv8 algorithm. Further, it is verified whether the trained tree detection network 51 has a certain accuracy using the remaining learning data 41. And if it can be confirmed that it has a certain accuracy, the training (learning) is terminated. Hereinafter, the tree detection network 51 trained in this way is described as the "trained model 52".

[0050] Note that the tree detection network 51 may be trained by inputting the overall photo 6 instead of the overall photo 6A. That is, it may be trained without trimming. Or, even for a tree in which only a part of the tree crown is shown, if the tree crown is shown at a predetermined ratio (e.g., 80%) or more, it may be left in the overall photo 6A without trimming.

[0051] 〔3. Inference phase〕 〔3.1 Detection of tree position and type〕 FIG. 8 is a diagram showing an example of the detection results of the regions and types of individual trees.

[0052] The operator prepares the overall photo 6B of the forest 81 to be inferred in the same way as the overall photo 6 (untrimmed photo) or the overall photo 6A (trimmed photo) was obtained in the learning phase. And the overall photo 6B is input to the computer 1.

[0053] Then, the tree detection unit 406 appropriately processes the overall photo 6B such as size adjustment and inputs it to the trained model 52 to detect the regions of the respective trees shown in the overall photo 6B. Further, for each tree, the probability of belonging to each specific type is inferred, and it is determined that it belongs to the type with the highest probability.

[0054] As a result, the area (tree crown) and type of each tree shown in the overall photograph 6B are detected as shown in FIG. 8. Also, the area of the region is calculated by obtaining the product of the actual area corresponding to one pixel of the overall photograph 6B (hereinafter referred to as the "unit pixel area") and the number of pixels occupied by the region.

[0055] [3.2 Estimation of the Number, Density, and Volume of Trees] FIG. 9 is a diagram showing an example of the relationship between the crown area and the diameter at breast height.

[0056] The tree number estimator 407 estimates the number, density, and volume of trees of each specific type in the forest 81 as follows.

[0057] The number of trees of each specific type is obtained from the discrimination result by the tree detection unit 406. That is, for each specific type, it is obtained by counting the number of trees determined to belong to it.

[0058] The density of trees of each specific type is obtained by the following formula (1). Di = Ni / Sa...... (1) "Di" is the density of the trees of the i-th label 42 (specific type). "Ni" is the number of trees determined to belong to the i-th label 42. "Sa" is the area of the entire land shown in the overall photograph 6B. The unit of area is, for example, ha (hectare). Sa can be calculated by obtaining the product of the unit pixel area and the number of pixels of the overall photograph 6B. The area of objects that are not trees (roads, etc.) may be subtracted from this product.

[0059] The volume of a single tree is generally calculated based on its height and breast height diameter. However, according to the method described in Prior Art 1 ("Knowing the volume from the diameter", Hideichi Yokoi, "Forestry in Gifu Prefecture" No. 537, issued in June 1998, https: / / www.forest.rd.pref.gifu.lg.jp / rd / ikurin / 9806gr.html), the volume can be calculated based on the breast height diameter without relying on the tree height. Specifically, for each tree species, a function with only the breast height diameter as a parameter, such as the formula (2), is prepared in advance as a function for calculating the volume (hereinafter referred to as the "VL function"). Vi = αi·Li βi ·γi Li …… (2) "Li" is a parameter and is the breast height diameter of the i-th specific type of tree for which the volume is to be calculated. "Vi" is the volume of the target tree. "αi", "βi", and "γi" are all constants for the i-th specific type and are obtained in advance by the method described in Prior Art 1.

[0060] However, the information obtained by the tree detection unit 406 does not directly include the breast height diameter. Therefore, it is considered to define a function (hereinafter referred to as the "LS function") representing the relationship between the crown area and the breast height diameter, such as the following formula (3). Li = fi(x) …… (3) fi(x) is an increasing function representing the relationship between the crown area of the i-th specific type of tree and the breast height diameter, and is defined by the following method.

[0061] The developer of the tree learning detection system 3 obtains an overall photo of a forest by photographing it from above with a drone 2, and discriminates the type and crown area of each tree shown in the overall photo. At this time, the tree detection unit 406 may be made to estimate (discriminate) the type and crown area of each tree. Furthermore, the breast height diameter of each tree shown in the overall photo is actually measured.

[0062] The developer selects, from the discrimination results, the pair of the area of the first specific type of tree crown and the diameter at breast height. When each pair is plotted on the xy plane, it is distributed as shown in Fig. 9. Then, f1(x) is generated by a machine learning algorithm such as known linear regression or non-linear regression. This is the LS function of the first specific type. For the second and subsequent specific types, the LS function is generated in the same way.

[0063] Then, by incorporating Equation (3) into Equation (2), the following function of Equation (4) (hereinafter referred to as the "VS function") is obtained. Vi = αi·(f(x)) βi ·γi f(x) …… (4) By the above method, a VS function for estimating the volume of each specific type of tree (i = 1, 2, …) is obtained. Then, each VS function is prepared in the tree number estimating unit 407.

[0064] When the tree type and the area of the tree crown of each tree shown in the overall photo 6B are obtained by the tree detection unit 406, the tree number estimating unit 407 calculates the volume of each tree by substituting the area of each tree crown into "x" of the VS function corresponding to the type of each tree.

[0065] That is, for example, when it is estimated by the tree detection unit 406 that a certain tree is of the third specific type and the area of the tree crown is Sb, the tree number estimating unit 407 calculates the volume V3 by substituting Sb into "x" of the VS function of the third specific type.

[0066] Then, the tree number estimating unit 407 calculates the total sum of the volumes of the trees belonging to each specific type.

[0067] 〔3.3 Output of Results〕 Fig. 10 is a diagram showing an example of the detection result screen 7. Fig. 11 is a diagram showing a modified example of the distribution window 7A.

[0068] The result display unit 408 causes the touch panel display 18 to display a detection result screen 7 as shown in FIG. 10. The detection result screen 7 is provided with a distribution window 7A and an estimation result window 7B.

[0069] In the distribution window 7A, in addition to all or a part of the entire photograph 6B, a bounding box 6C surrounding the image of the detected tree is arranged. Further, at the upper left of the bounding box 6C, a tag 6D indicating the identified specific type of the tree corresponding to the bounding box 6C and the probability belonging thereto is arranged.

[0070] When the operator can only see a part of the entire photograph 6B, the operator may perform a scrolling operation. Then, the result display unit 408 causes the distribution window 7A to display other parts of the entire photograph 6B according to the operation. Further, accordingly, the bounding boxes 6C and tags 6D of the respective trees shown in the displayed part are arranged in the distribution window 7A.

[0071] Also, the operator can enlarge or reduce the entire photograph 6B. Then, the result display unit 408 arranges the bounding boxes 6C and tags 6D of the respective trees shown in the part shown in the distribution window 7A in the distribution window 7A according to the enlargement or reduction.

[0072] Alternatively, as shown in FIG. 11, the result display unit 408 colors the area of each tree in a color corresponding to the specific type to which it belongs, colors the area where no tree exists in a specific color (for example, white), and further displays a distribution diagram in which the bounding boxes 6C and tags 6D are arranged in the distribution window 7A. Alternatively, as shown in FIG. 8, it may be displayed in a form in which the bounding box 6C is not arranged and the content of the tag 6D is described in the area.

[0073] The estimation result window 7B shows the sum of the number, density, and volume of the trees of each label 42 (specific type) estimated by the tree number estimation unit 407.

[0074] Instead of the overall photo 6B, the aerial photo 60 taken by the drone 2 in real time may be placed in the distribution window 7A.

[0075] That is, while flying over the forest 81, the drone 2 takes a photo of the forest 81 every predetermined time (for example, 10 seconds) and transmits the aerial photo 60 to the computer 1.

[0076] In the computer 1, every time the tree detection unit 406 receives the aerial photo 60, it appropriately processes the received aerial photo 60 and inputs it to the learned model 52, thereby detecting individual trees shown in the aerial photo 60 and inferring the probability that each tree belongs to a specific type. At this time, it is determined that the tree belongs to the type with the highest probability.

[0077] Then, the result display unit 408 regenerates the detection result screen 7 based on the aerial photo 60, the detection result, the inference result, and the determination result, and displays it on the touch panel display 18.

[0078] 〔3.4 Length of one side of the region 43〕 As described above, the region 43 (see FIGS. 6 and 7) is designated so that the length of one side falls within a predetermined range. The shorter one side is, the more precisely the tree crown of the tree can be traced. However, compared to the appearance of objects such as automobiles and buildings, even trees of the same type have various appearances of their respective tree crowns. Therefore, the more precisely traced, the more prominent the characteristics of the regions 43 of a plurality of trees of the same type become, and it is difficult to find commonality. Therefore, more learning data is required to generate a highly accurate learned model.

[0079] On the other hand, the longer one side is, the rougher the region 43 becomes. Then, it becomes difficult to distinguish between the regions 43 of a plurality of trees of different types. Therefore, when machine learning is performed based on such a region 43, the accuracy of the generated learned model becomes low.

[0080] Therefore, it is important to determine how to define the above-mentioned predetermined range. The inventor conducted experiments and verifications under the following conditions. · Location of the forest: Kyoto Prefecture · Usage model: YOLOv8 · Number of regions 43 per type: 50 to 300 · Length of one side of region 43: Three types, approximately 5 pixels, approximately 10 pixels, and approximately 20 pixels Then, the trained model generated by specifying region 43 to have a side length of approximately 10 pixels was able to detect trees and discriminate their types more accurately than the trained models generated by specifying a side length of approximately 5 pixels and the trained model generated by specifying a side length of approximately 20 pixels.

[0081] In this embodiment, a side of 10 pixels corresponds to approximately 6.3 cm in actuality. The average length of the leaves of broad-leaved trees growing in Japanese forests is approximately 6.3 cm, and generally belongs to the range of approximately 3.15 cm to 12.6 cm. Also, the length of the bundles of needle-shaped leaves (sometimes called "leaf bundles" or "needles") of coniferous trees growing in Japanese forests is similarly approximately 6.3 cm on average and generally belongs to the range of approximately 3.15 cm to 12.6 cm.

[0082] As a result of the inventor's consideration from various perspectives, it was found that when humans identify the type of a tree by looking at the tree crown from a certain distance, the coarseness of the tree crown's contour being about the length of the leaves of broad-leaved trees or the leaf bundles of coniferous trees as the side length is sufficient.

[0083] Therefore, in this embodiment, by designating, as region 43, a polygon having a side length corresponding to the average length of the leaves of broad-leaved trees or the leaf bundles of coniferous trees, that is, approximately 10 pixels as the side length in this embodiment, the generation of a trained model 52 of a certain quality can be realized even with relatively little learning data 41.

[0084] The length of one side does not have to be exactly 10 pixels. For example, it may fall within the range of 0.5 times to 2 times the average length of the leaves or leaf clusters of the target forest. That is, in this embodiment, it may fall within the range of 5 to 20 pixels. This range can be arbitrarily changed. For example, it may be within the range of 0.75 times to 1.5 times, or it may be within the range of 0.4 times to 2.2 times.

[0085] Also, the length of one side and the range are changed according to the average length of the leaves or leaf clusters of the trees growing in the target forest. For example, even in forests with the same vegetation, the average length may vary depending on the region or season. Therefore, the approximate length of one side of region 43 may be changed according to the region or season.

[0086] Alternatively, although the types of trees in the tropical rainforest are diverse, the leaves of the tall trees visible from above are generally large, with a length of about 30 to 100 cm. Therefore, a polygon with the number of pixels corresponding to 30 to 100 cm as the approximate length of one side may be designated as region 43.

[0087] [4. Overall processing flow and effects of this embodiment] FIG. 12 is a flowchart showing an example of the overall processing flow by the tree detection program 40.

[0088] Next, the overall processing flow by the computer 1 will be described with reference to the flowchart of FIG. 12. The computer 1 executes processing according to the procedure shown in FIG. 12 based on the tree detection program 40.

[0089] When the overall photo 6 is input, the computer 1 displays it on the touch panel display 18 (step #101 in FIG. 12). The operator designates the region 43 and label 42 (type) of each tree shown in the overall photo 6, such as by tracing. Further, an image of a tree with only a part of its crown shown is trimmed (cut).

[0090] Then, the computer 1 receives the specified content (#102), generates and stores, as learning data 41, the trimmed entire photo 6 (entire photo 6A) and the data indicating the specified content (annotation content) (#103). Since a plurality of entire photos 6 are input to the computer 1, for each entire photo 6, the processes of steps #101 to #103 and the operations of the operator are executed.

[0091] Then, the computer 1 generates a learned model 52 by training the tree detection network 51 using the learning data 41 (#104). Thus, the learning phase is completed.

[0092] In the inference phase, when the entire photo 6B is input, the computer 1 detects each tree shown in the entire photo 6B (#121), discriminates the label 42 (type) of each tree (#122), and estimates the number, density, and volume of each type of tree (#123). Then, based on the results of steps #121 to #123, the detection result screen 7 is displayed (#124).

[0093] According to the present embodiment, each tree can be more reliably detected by AI from a photo of a forest where trees are dense than in the past.

[0094] 〔5. Modification〕 In order to generate a learned model with high versatility for the environment as the learned model 52, as the entire photo 6, the entire photos for several time zones (dawn, around 10 to 12 am, around 1 to 3 pm, evening) may be used, or the entire photos for several weather conditions (sunny, cloudy, rainy) may be used. Or, the entire photos for each combination of these (dawn sunny, dawn cloudy,...) may be used.

[0095] Or, dedicated learned models 52 may be generated for each of several time zones, each of several weather conditions, or each combination of these.

[0096] In this embodiment, the computer 1 acquired the polygon designated by the operator as the area 43, but it may be acquired by other methods. For example, even when the operator designates an area such as an ellipse in outline, the computer 1 may provide a plurality of inflection points so that the distance between two adjacent points becomes a predetermined distance (for example, 8 pixels or more and 12 pixels or less) based on the designated area and the image of the area, and recognize and process it as a polygon or a substantially polygon.

[0097] In this embodiment, the computer 1 was used to classify trees, but it may be used to classify other objects. For example, it may be used to classify flowers, shrubs, animal habitats (animal nests such as ant mounds and mole nests), or terrain distributed in grasslands. Or, it may be used to classify rocks (volcanic rocks, hypabyssal rocks, plutonic rocks, sedimentary rocks, etc.).

[0098] In any case of classification, the computer 1 and the drone 2 may basically perform the same processing as in this embodiment.

[0099] However, when classifying flowers or shrubs, it is desirable that the drone 2 collects the aerial photograph 60 by taking pictures at a lower altitude than when classifying trees.

[0100] On the other hand, when classifying terrain (lakes, plains, hills, rivers), the drone 2 may collect the aerial photograph 60 by taking pictures at a higher altitude than when classifying trees.

[0101] In this embodiment, the case of detecting trees of a plurality of types has been described as an example, but the computer 1 can also generate a model for detecting only one type of tree as the learned model 52. In this case, it is not necessary for the operator to specify the label 42.

[0102] In this embodiment, the case of using YOLOv8 as a machine learning and inference tool has been described as an example. However, if it is an instance segmentation tool, other versions of YOLO may be used, or tools other than the YOLO series, for example, Mask R-CNN (Mask Regional Convolutional Neural Network) may be used.

[0103] FIG. 13 is a diagram showing a modified example of the functional configuration of the computer 1. FIG. 14 is a flowchart showing a modified example of the overall processing flow by the tree detection program 40B. FIG. 15 is a diagram showing an example of a method for generating the training data 46. FIG. 16 is a diagram showing a modified example of the inference method.

[0104] CPM (chopped picture method) may be combined with this embodiment. In this method, a process of acquiring a learning image in which the identification targets included in the object to be identified are imaged integrally, a process of specifying a plurality of pixel blocks composed of a predetermined number of pixels from the image, and a label corresponding to the type of the identification target are associated with each of the specified plurality of pixel blocks to generate training data, and a process of training a learning model for identifying the identification target based on the generated training data are executed. Then, the identification target included in the input image is identified using this learning model. The details of this method are disclosed in Japanese Patent Laid-Open No. 2019-128842 or https: / / arxiv.org / ftp / arxiv / papers / 1708 / 1708.01986.pdf.

[0105] According to this method, individuals cannot be detected independently one by one, but the identification targets (types of individuals) shown in each pixel or pixel block of the input image can be identified with high accuracy and high speed.

[0106] Hereinafter, embodiments in the case of combining CPM will be described with reference to FIGS. 13 to 16 and the like.

[0107] In this modification example, a tree detection program 40B is installed in the computer 1 and executed instead of the tree detection program 40 (see FIG. 3). According to the tree detection program 40B, a first tree species identification unit 420, an annotation reception unit 421, a learning data acquisition unit 422, a learning data storage unit 423, a model storage unit 424, a machine learning unit 425, a tree detection unit 426, a tree number estimation unit 427, a result display unit 428, and a second tree species identification unit 429 shown in FIG. 13 are realized.

[0108] A learning model 58 for identifying trees grown in forests 80, 81 (see FIG. 1) is generated in advance by the method described in Japanese Patent Application Laid-Open No. 2019-128842 and stored in the model storage unit 424. Each part in FIG. 13 describes the processing in the procedure shown in the flowchart of FIG. 14.

[0109] In the machine learning phase, similar to the annotation reception unit 401 (see FIG. 3), the annotation reception unit 421 displays the first one of the input plurality of overall photos 6 on the touch panel display 18 (#131 in FIG. 14).

[0110] Here, the operator designates the area 43 for each tree by tracing as shown in FIG. 15(A) and trims the trees in which only a part of the tree crown is shown, using the method described with reference to FIGS. 6 to 7. Since the type of the tree is identified by the first tree species identification unit 420 as described later, it is not necessary for the operator to perform labeling. In FIGS. 15 to 16, the photos of the tree crowns are omitted and the shape of the area 43 is simplified.

[0111] Then, the annotation reception unit 421 receives the content of the annotation (the traced area 43 of each tree) and the part to be trimmed (cut) by the operator (#132).

[0112] The first tree species identification unit 420 identifies the type of tree shown in each pixel block or pixel of the overall photo 6 using the CPM and the learning model 58 (#133). As a result, a distribution map of the types of trees in the forest 80 as shown in FIG. 15(B) is obtained.

[0113] The learning data acquisition unit 422 generates the overall photo 6A by trimming the received portion from the overall photo 6 (#134), and generates learning data 46 for each specific type as shown in FIG. 15(C) (#135). The learning data 46 of the i-th (i = 1, 2,...) specific type includes an image of the area in the overall photo 6A where the i-th specific type of tree is shown (hereinafter referred to as "tree species image 6H"), and the content of the annotation of each tree of the i-th specific type (such as the position and shape of the area 43).

[0114] The annotation reception unit 421, the learning data acquisition unit 422, and the operator perform the same processing and operations on the remaining overall photo 6. As a result, each learning data 46 is stored in the learning data storage unit 423.

[0115] The model storage unit 424 further stores one tree detection network 56 for each specific type in advance. The tree detection network 56 is a neural network that detects the trees shown in the input photo, similar to the tree detection network 51.

[0116] The machine learning unit 425 uses a part (for example, 70%) of the learning data 46 of the i-th specific type to train the i-th specific type of tree detection network 56 stored in the model storage unit 424 using the YOLOv8 algorithm (#136). Further, it is verified whether the trained tree detection network 56 has a certain accuracy using the remaining learning data 46 of the i-th specific type. And if it can be confirmed that it has a certain accuracy, the training (learning) is terminated. Hereinafter, the tree detection network 56 trained in this way is referred to as "trained model 57".

[0117] Through the above processes and operations, trained models 57 of multiple specific types are obtained respectively.

[0118] In the inference phase, when the overall photo 6B of the forest 81 is input into the computer 1 by the operator, the second tree species identification unit 429 appropriately processes the overall photo 6B, such as resizing, and for each pixel block or pixel of the overall photo 6B, it identifies the type of tree shown therein by the CPM and the learning model 58 (#141). Thereby, the distribution of specific types of trees in the forest 81 as shown in FIG. 16 is obtained.

[0119] The tree detection unit 426 inputs the image of the i-th specific type of part in the overall photo 6B into the i-th trained model 57 of the specific type, thereby detecting the individual trees shown therein (#142).

[0120] Thereby, the type of each tree and the crown area shown in the overall photo 6B are specified. Also, the area of each crown area is calculated.

[0121] The tree number and other estimation unit 427 estimates the number, density, and volume of each specific type of tree in the forest 81 in the same manner as the tree number and other estimation unit 407 (#143).

[0122] The result display unit 428 causes the detection result screen 7 as shown in FIG. 10 to be displayed on the touch panel display 18 in the same manner as the result display unit 408 (#144). However, the content of the detection result screen 7 is based on the results obtained by the tree detection unit 426 and the tree number and other estimation unit 427 respectively.

[0123] Note that the first tree species identification unit 420 or the second tree species identification unit 429 may obtain the distribution of each specific type by using the algorithms and models of semantic segmentation methods other than the CPM.

[0124] In this embodiment, all the functions shown in FIG. 3 are aggregated in the computer 1, but they may be distributed among a plurality of devices. For example, the learning data acquisition unit 402 may be provided in a first computer, the learning data storage unit 403, the model storage unit 404, and the machine learning unit 405 may be provided in a second computer, the tree detection unit 406 and the tree number estimation unit 407 may be provided in a third computer, and the result display unit 408 may be provided in a fourth computer. These computers may be interconnected so that each piece of information can be transmitted and received. Similarly, each function shown in FIG. 13 may also be distributed among a plurality of devices.

[0125] In addition, the configuration of the entire tree learning detection system 3, the computer 1 or each part thereof, the content of the processing, the order of the processing, the configuration of the screen, the configuration of the data, etc. can be appropriately changed in accordance with the gist of the present invention.

Explanation of Reference Numerals

[0126] 1 Computer 3 Tree learning detection system (object learning detection system, object learning system, object detection system) 41 Learning data 43 Region 46 Learning data 51 Tree detection network (model) 52 Learned model 56 Tree detection network (model) 57 Learned model 402 Learning data acquisition unit (acquisition means) 405 Machine learning unit (learning means) 406 Tree detection unit (detection means) 407 Tree number estimation unit (number etc. estimation means) 420 First tree species identification unit (first discrimination means) 422 Learning data acquisition unit (learning image acquisition means, acquisition means) 425 Machine learning unit (learning means) 427 Tree number estimation unit (detection means) 429 Second tree species identification unit (second discrimination means) 6A Overall photograph (learning image) Full photo of 6B (inference image)

Claims

1. An acquisition means for acquiring learning data including a learning image, a region in which the distance between two adjacent points among a plurality of designated points surrounding each learning target object shown in the learning image is within a predetermined range, and the type of the learning target object; A learning means for generating a learned model by training an instance segmentation model based on the learning data; A detection means for detecting the region and type of each detection target object shown in the inference image based on the learned model; An object learning detection system, characterized by comprising the above.

2. An acquisition means for acquiring learning data including a learning image and a region in which the distance between two adjacent points among a plurality of designated points surrounding each learning target object of a predetermined type shown in the learning image is within a predetermined range; A learning means for generating a learned model by training an instance segmentation model based on the learning data; A detection means for detecting the region of each detection target object of the predetermined type shown in the inference image based on the learned model; An object learning detection system, characterized by comprising the above.

3. A learning image acquisition means for acquiring a learning image in which a plurality of learning target objects belonging to any one of a plurality of types are shown; A first discrimination means for discriminating a first distribution region of the learning target objects belonging to each of the plurality of types by discriminating the type of the learning target object shown in each pixel or pixel block of the learning image by semantic segmentation; For each of the plurality of types, an acquisition means for acquiring learning data including an image of the first distribution region of the type in the learning image and a region in which the distance between two adjacent points among a plurality of designated points surrounding each learning target object shown in the image is within a predetermined range; A learning means for generating respective learned models by training instance segmentation models of the plurality of types based on the learning data of the type; A second discrimination means for discriminating a second distribution region of the detection target objects belonging to each of the plurality of types by discriminating the type of the detection target object shown in each pixel or pixel block of the inference image by semantic segmentation; For each of the plurality of types, detection means for detecting a detection target object shown in an image of the second distribution region of that type among the inference images based on the learned model of that type; An object learning detection system characterized by having the above.

4. An acquisition means for acquiring learning data including a learning image, a region in which the distance between two adjacent points among a plurality of designated points surrounding each learning target object shown in the learning image is within a predetermined range, and the type of the learning target object; A learning means for generating a learned model by training a model of instance segmentation based on the learning data; An object learning system characterized by having the above.

5. An acquisition means for acquiring learning data including a learning image and a region in which the distance between two adjacent points among a plurality of designated points surrounding each learning target object of a predetermined type shown in the learning image is within a predetermined range; A learning means for generating a learned model by training a model of instance segmentation based on the learning data; An object learning system characterized by having the above.

6. A learning image acquisition means for acquiring a learning image in which a plurality of learning target objects belonging to any one of a plurality of types are shown; A discrimination means for discriminating the distribution region of the learning target objects belonging to each of the plurality of types by discriminating the type of the learning target object shown in each pixel or pixel block of the learning image by semantic segmentation; For each of the plurality of types, an acquisition means for acquiring learning data including an image of the distribution region of that type in the learning image and a region in which the distance between two adjacent points among a plurality of designated points surrounding each learning target object shown in the image is within a predetermined range; A learning means for generating respective learned models by training models of instance segmentation for each of the plurality of types based on the learning data of that type; An object learning system characterized by having the above.

7. The discrimination means discriminates the distribution region by using the chopped picture method as the semantic segmentation. The object learning system according to Claim 6.

8. The learning image is a photograph obtained by photographing a forest from above, The object to be learned is a tree, The predetermined range is a range from R times (0.5 ≤ R < 1) to S times (1 < S ≤ 2) the average length of the leaves or leaf clusters of the tree, The object learning system according to any one of claims 4 to 7.

9. Detection means for detecting each detection target object shown in the inference image based on the learned model generated by the object learning system described in claim 1, An object detection system characterized by comprising:

10. Means for estimating the number, etc., Comprising: The inference image is a photograph obtained by photographing a forest from above, The detection target object is a tree, Based on the detection result by the detection means, the number of detection target objects belonging to the type shown in the inference image, the number per unit area of the detection target object, and the sum of the detection target objects are estimated for each of the plurality of types by the number etc. estimation means, The object detection system according to claim 9.

11. Acquiring learning data including a learning image, a region in which the distance between two adjacent points among a plurality of designated points surrounding each learning target object shown in the learning image is within a predetermined range, and the type of the learning target object, Generating a learned model by training a model of instance segmentation based on the learning data, An object learning method characterized by the above

12. A process of acquiring learning data including a learning image, a region in which the distance between two adjacent points among a plurality of designated points surrounding each learning target object shown in the learning image is within a predetermined range, and the type of the learning target object, A process of generating a learned model by training a model of instance segmentation based on the learning data, A computer program characterized by causing a computer to execute the above.

Citation Information

Patent Citations

  • Mask R-CNN-based street tree image instance segmentation method

    CN112116612A

  • Plant Group Identification

    US20230177697A1