Information processing device and information processing method

The system addresses the challenge of predicting agricultural yields and detecting non-productive areas in new fields by employing a selection mechanism for learning models based on environmental conditions and an annotation process, achieving efficient and accurate predictions without extensive historical data.

JP7681975B2Active Publication Date: 2025-05-23CANON KK
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2021000560
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-01-05
Publication Date
2025-05-23
Estimated Expiration
2041-01-05

AI Technical Summary

Technical Problem

Existing methods for predicting agricultural yields and detecting non-productive areas in fields face challenges when introduced to new fields, as they require sufficient historical data and manual adjustments, which are time-consuming and costly.

Method used

A system that uses a plurality of learning models trained for object detection in images, with a selection mechanism to choose the most appropriate model based on environmental conditions, and an annotation process to refine the models, allowing for accurate predictions even in new fields without extensive historical data.

Benefits of technology

Enables efficient and accurate processing of images from new fields, reducing the need for extensive historical data and manual adjustments, thereby lowering costs and improving prediction accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007681975000006
    Figure 0007681975000006
  • Figure 0007681975000007
    Figure 0007681975000007
  • Figure 0007681975000008
    Figure 0007681975000008
Patent Text Reader

Abstract

To provide an information processing device and an information processing method capable of performing processing by a learning model according to a situation when it is difficult to perform the processing based only on information collected in the past or even when there is no information collected in the past.SOLUTION: An information processing method selects one or more learning models from a part or all of learning models based on the result of object detection process for plural captured images by a part or all of the multiple learning models learned in learning environments differing from one another. The information processing method obtains the score for each of the multiple captured images based on the result of the object detection process for the plural captured images by the selected learning model. The method identifies the captured image subjected to annotation works out of the multiple captured images based on the score for each of the multiple captured images.SELECTED DRAWING: Figure 2A
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to a technique for prediction based on captured images. [Background technology]

[0002] In recent years, there have been many efforts to use IT to solve a variety of problems in agriculture, such as predicting yields, predicting optimal harvest times, controlling the amount of pesticide sprayed, and planning field restoration.

[0003] For example, Patent Document 1 discloses a method for early detection and response to abnormal growth conditions by appropriately referring to sensor information acquired from the field where agricultural crops are grown and a database that stores such information, thereby grasping the growth conditions and harvest predictions at an early stage.

[0004] Furthermore, Patent Document 2 discloses a method for field management that reduces variation in the quality and yield of agricultural crops by making arbitrary inferences based on information acquired from a wide variety of agricultural crop-related sensors and referring to registered information. [Prior art documents] [Patent documents]

[0005] [Patent Document 1] JP 2005-137209 A [Patent Document 2] JP 2016-49102 A Summary of the Invention [Problem to be solved by the invention]

[0006] However, the methods proposed so far have been based on the premise that a sufficient number of previously acquired cases for the field for which predictions are to be performed are stored and that adjustments have been made so that predictions can be made with high accuracy based on information about those cases.

[0007] On the other hand, the success or failure of agricultural crops is generally greatly affected by environmental fluctuations such as weather and climate, and also varies greatly depending on the spraying of fertilizers and pesticides by workers. If all conditions due to external factors remained constant from year to year, there would be no need to predict yields or harvest times, but unlike industry, agriculture has many external factors that workers cannot control, making predictions extremely difficult. Also, when predicting yields when unprecedented weather conditions continue, it is difficult to make accurate predictions using the estimation system adjusted from the above-mentioned past cases.

[0008] The most difficult case to predict is when the above prediction system is newly introduced to a field. For example, consider the case of predicting yield in a specific field or detecting non-productive areas for the purpose of repairing poorly growing areas (dead branches, diseased areas). In such a task, images and parameters of crops collected in the past in the field are usually stored in a database. Then, when actually carrying out predictions on the field, images taken in the current field and data related to growth information obtained from other sensors are cross-referenced and adjusted to make a highly accurate prediction. However, as mentioned above, when these prediction systems and non-productive area detectors are introduced to a different new field, they cannot be applied immediately because the conditions (of the field) often do not match. In such a case, it was necessary to collect a sufficient amount of data in the new field and make adjustments.

[0009] In addition, when the above-mentioned prediction system or non-productive area detector is adjusted manually, it takes a lot of time because the parameters related to crop growth are high-dimensional. Even when deep learning or similar machine learning methods are used, manual labeling (annotation) is usually required to achieve good performance for new input, which results in high work costs.

[0010] Ideally, when introducing a new forecasting system or in the case of a natural disaster or weather that has never occurred before, it would be desirable to be able to make good forecasts and estimates with simple settings that place a low burden on the user.

[0011] The present invention provides technology that enables processing using a learning model appropriate to the situation, even when processing is difficult using only previously collected information or when no previously collected information is available. [Means for solving the problem]

[0012] One aspect of the present invention is A plurality of learning models that are trained to output a result of an object detection process as an image area related to an object in a photographed image when the photographed image is input, Different from each other With parameter set Multiple learning models learned home part Learning model Or all Learning model Based on the results of object detection processing on multiple captured images by According to the specified conditions , the part Learning model Or all Learning model from A plurality of the above-mentioned predetermined conditions are satisfied A selection means for selecting a learning model; The selection means selects Multiple Based on a result of the object detection process for the plurality of photographed images by the learning model, Based on the difference in the results of the object detection processing using the multiple learning models selected by the selection means A means for obtaining a score; Among the plurality of captured images Applicable Score Images with a score above the threshold are used as the target images for annotation, which is the process of assigning labels as training data indicating the correct answer. Identifying means and 、 and a means for performing additional learning using the captured image identified by the identification means and a captured image used in the initial learning of the learning model selected by the selection means, so as to output a result of an object detection process that is a result of an image area related to an object in the captured image when the captured image is input, for the learning model selected by the selection means, The acquisition means obtains the score for each of the plurality of captured images, the score being defined so as to be larger as the results of the object detection process between the learning models selected by the selection means differ more significantly. It is characterized by: Effect of the Invention

[0013] According to the configuration of the present invention, it is possible to provide technology that enables processing using a learning model according to the situation, even when processing is difficult using only information collected in the past, or when there is no information collected in the past. [Brief description of the drawings]

[0014] [Figure 1] FIG. 1 is a diagram showing an example of a system configuration. [Figure 2A]A flowchart of a series of processes for identifying a captured image that requires annotation work, accepting annotation work for the captured image, and performing additional learning of a learning model using the captured image that has been annotated. [Figure 2B] 10 is a flowchart showing details of the process in step S23. [Figure 2C] 10 is a flowchart showing details of the process in step S234. [Diagram 3] 1 is a diagram showing an example of a method for photographing a farm field using a camera 10. FIG. [Figure 4] Diagram showing a difficult case. [Diagram 5] FIG. 13 is a diagram showing the results of annotation work performed on a captured image. [Figure 6] FIG. 13 is a diagram showing an example of a GUI display. [Figure 7] FIG. 13 is a diagram showing an example of a GUI display. [Figure 8A] 4 is a flowchart of a setting process (setting process for visual inspection) of the inspection device. [Figure 8B] 10 is a flowchart showing details of the process in step S83. [Figure 8C] 10 is a flowchart showing details of the process in step S833. [Figure 9] FIG. 4 is a diagram showing an example of a detection region. [Figure 10] FIG. 13 is a diagram showing an example of a GUI display. [Figure 11] 1A is a diagram showing an example of the configuration of a query parameter, FIG. 1B is a diagram showing an example of the configuration of a parameter set of a learning model, and FIG. 1C is a diagram showing an example of the configuration of a query parameter. [Figure 12] Venn diagram. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0015] Hereinafter, the embodiments will be described in detail with reference to the attached drawings. Note that the following embodiments do not limit the invention according to the claims. Although the embodiments describe a number of features, not all of these features are essential to the invention, and the features may be combined in any manner. Furthermore, in the attached drawings, the same reference numbers are used for the same or similar configurations, and duplicated descriptions are omitted.

[0016] [First embodiment] In this embodiment, a system is described that performs analysis processing of a farm field, such as predicting the yield of agricultural crops in the field and detecting areas that need repair, based on an image of the field captured by a camera.

[0017] First, an example of the configuration of a system according to this embodiment will be described with reference to Fig. 1. As shown in Fig. 1, the system according to this embodiment includes a camera 10, a cloud server 12, and an information processing device 13.

[0018] First, the camera 10 will be described. The camera 10 captures a moving image of the field, and outputs the image of each frame in the moving image as a "captured image of the field". Alternatively, the camera 10 captures still images of the field periodically or irregularly, and outputs the captured still images as a "captured image of the field". In order to perform accurate predictions, which will be described later, from the captured images, it is desirable that images captured in the same field are captured in the same environment and conditions as much as possible. The captured images output from the camera 10 are transmitted to a cloud server 12 or an information processing device 13 via a communication network 11, such as a LAN or the Internet.

[0019] The method of photographing the field by the camera 10 is not limited to a specific photographing method. An example of the method of photographing the field by the camera 10 will be described with reference to FIG. 3(A). In FIG. 3(A), cameras 33 and 34 are used as the camera 10. In a typical field, rows of crop trees are planted by farmers in a planned manner. For example, as shown in FIG. 3(A), rows of crop trees are planted side by side, such as a row of crop trees 30 and a row of crop trees 31. The farm tractor 32 is provided with a camera 34 that photographs the row of crop trees 31 on the left side in the traveling direction indicated by the arrow, and a camera 33 that photographs the row of crop trees 30 on the right side. Therefore, when the farm tractor 32 moves between the rows 30 and 31 in the traveling direction indicated by the arrow, the camera 34 will photograph multiple images of the crop trees in the row 31, and the camera 33 will photograph multiple images of the crop trees in the row 30.

[0020] In many farm fields designed for agricultural tractors 32 to enter and work in, where agricultural trees are planted at equal intervals, it is relatively easy to photograph a larger number of agricultural trees at a constant height and at a constant distance from the agricultural trees by photographing the agricultural trees with cameras 33, 34 installed on the agricultural tractor 32 as shown in Fig. 3(A). This makes it possible to photograph the entire target farm field under almost the same conditions, and image capture under desirable conditions can be easily achieved.

[0021] However, other photographing methods may be used as long as the photographing of the field can be performed under roughly the same conditions. An example of a method of photographing a field using the camera 10 will be described with reference to FIG. 3(B). In FIG. 3(B), cameras 38 and 39 are used as the camera 10. As shown in FIG. 3(B), in a field or the like where the interval between the row 35 of the crop trees and the row 36 of the crop trees is narrow and travel by a tractor is impossible, photographing may be performed using the camera 38 and the camera 39 attached to the drone 37. The drone 37 is provided with a camera 39 that photographs the row 36 of the crop trees on the left side in the traveling direction indicated by the arrow, and a camera 38 that photographs the row 35 of the crop trees on the right side. Therefore, when the drone 37 moves between the rows 35 and 36 in the traveling direction indicated by the arrow, the camera 39 will take multiple photographed images of the crop trees in the row 36, and the camera 38 will take multiple photographed images of the crop trees in the row 35.

[0022] In addition, the images of the crop trees may be taken by a camera installed on the self-propelled robot. Although the number of cameras used for taking the images is two in Figures 3(A) and (B), the number is not limited to a specific number.

[0023] Regardless of the method used to capture an image of a crop tree, camera 10 outputs the image together with the shooting information at the time the image was captured (Exif information recording the shooting location (e.g., the shooting location measured by GPS), the shooting date and time, information related to camera 10, etc.).

[0024] Next, the cloud server 12 will be described. In the cloud server 12, captured images (captured images with Exif information attached) transmitted from the camera 10 are registered. In addition, in the cloud server 12, a plurality of learning models (detectors / settings) for detecting image areas relating to agricultural products (objects) from the captured images (object detection) are registered, and each learning model is a model that has been trained in a mutually different learning environment. In addition, the cloud server 12: SelfAmong the multiple learning models held by the cloud server 12, a learning model that is relatively robust in terms of detection accuracy when detecting an image region related to agricultural products from a photographed image is selected. Then, the cloud server 12 uses the photographed image in which the deviation in detection results between the selected learning models is relatively large for additional learning of the selected learning model.

[0025] CPU 191 executes various processes using computer programs and data stored in RAM 192 and ROM 193. In this way, CPU 191 controls the overall operation of cloud server 12, and executes or controls various processes that will be described as being performed by cloud server 12.

[0026] The RAM 192 has an area for storing computer programs and data loaded from the ROM 193 or the external storage device 196, and an area for storing data received from the outside via the I / F 197. The RAM 192 further has a work area used when the CPU 191 executes various processes. In this way, the RAM 192 can provide various areas as needed.

[0027] ROM 193 stores setting data for cloud server 12, computer programs and data related to the startup of cloud server 12, computer programs and data related to the basic operation of cloud server 12, and the like.

[0028] The operation unit 194 is a user interface such as a keyboard, a mouse, a touch panel, etc., and the user can input various instructions to the CPU 191 by operating it.

[0029] Display unit 195 has a screen such as a liquid crystal screen or a touch panel screen, and can display the results of processing by CPU 191 as images, characters, etc. Note that display unit 195 may be a projection device such as a projector that projects images and characters.

[0030] The external storage device 196 is a large-capacity information storage device such as a hard disk drive device. The external storage device 196 stores an OS (operating system) and computer programs and data for causing the CPU 191 to execute or control various processes described as being performed by the cloud server 12. The data stored in the external storage device 196 includes data related to the above-mentioned learning model. The computer programs and data stored in the external storage device 196 are loaded into the RAM 192 as appropriate under the control of the CPU 191, and become targets for processing by the CPU 191.

[0031] I / F 197 is a communication interface for performing data communication with the outside, and cloud server 12 transmits and receives data to and from the outside via I / F 197. CPU 191, RAM 192, ROM 193, operation unit 194, display unit 195, external storage device 196, and I / F 197 are all connected to system bus 198. Note that the configuration of cloud server 12 is not limited to the configuration shown in FIG.

[0032] In addition, the captured image output from the camera 10 may be temporarily stored in the memory of another device, and the captured image may be transferred from the memory to the cloud server 12 via the communication network 11.

[0033] Next, the information processing device 13 will be described. The information processing device 13 is a computer device such as a PC (personal computer), a smartphone, or a tablet terminal device. The information processing device 13 accepts annotation work for a captured image identified by the cloud server 12 as a "captured image requiring manual labeling (teaching data (GT: Ground Truth) indicating a correct answer) (annotation work)." The cloud server 12 performs additional learning of a "learning model that is relatively robust in terms of detection accuracy when detecting an image area related to agricultural products from a captured image" using a plurality of captured images including a captured image on which annotation work has been performed by a user, and updates the learning model. The cloud server 12 performs the above analysis process by detecting an image area related to agricultural products from an image captured by the camera 10 using the learning model held by itself.

[0034] The CPU 131 performs various processes using computer programs and data stored in the RAM 132 and the ROM 133. As a result, the CPU 131 controls the operation of the entire information processing device 13, and also executes or controls various processes that will be described as being performed by the information processing device 13.

[0035] The RAM 132 has an area for storing computer programs and data loaded from the ROM 133, and an area for storing data received from the camera 10 and the cloud server 12 via the input I / F 135. The RAM 132 further has a work area used when the CPU 131 executes various processes. In this way, the RAM 132 can provide various areas as needed.

[0036] The ROM 133 stores setting data for the information processing device 13, computer programs and data related to the startup of the information processing device 13, computer programs and data related to the basic operation of the information processing device 13, and the like.

[0037] The output I / F 134 is an interface used by the information processing device 13 to output / transmit various types of information to the outside.

[0038] The input I / F 135 is an interface used by the information processing device 13 to input / receive various types of information from the outside.

[0039] The display device 14 has a liquid crystal screen or a touch panel screen, and can display the results of processing by the CPU 131 as images, characters, etc. The display device 14 may be a projection device such as a projector that projects images and characters.

[0040] The user interface 15 includes a keyboard and a mouse, and can be operated by the user to input various instructions to the CPU 131. Note that the configuration of the information processing device 13 is not limited to the configuration shown in Fig. 1, and may include, for example, a large-capacity information storage device such as a hard disk drive device, in which computer programs such as a GUI (to be described later) and data are stored. The user interface 15 may also include a touch sensor such as a touch panel.

[0041] Next, a task flow for predicting the yield of agricultural products harvested in a farm field at an earlier stage than the harvest time from an image of the farm field captured by the camera 10 will be described. When predicting the yield by simply counting the fruits to be harvested at the harvest time, the purpose can be achieved by simply detecting the target fruits from the captured image using a classifier by a method called specific object detection. Since the fruits themselves have a very distinctive appearance, this method detects them using a classifier that has learned this distinctive appearance.

[0042] In this embodiment, when the crop is a fruit, the yield of the fruit is predicted at an earlier stage than the harvest time, in addition to counting the fruits after they have ripened. For example, the yield is predicted from the number of inflorescences that will later become fruits, the yield is predicted by detecting dead branches or diseased areas that are unlikely to bear fruit, or the yield is predicted from the state of leaf growth. In order to make such predictions, a prediction method that can handle the fact that the growth conditions of the crops differ depending on the shooting time and climate is required. In other words, it is necessary to select a prediction method with good prediction performance depending on the condition of the crops. In this case, it is expected that the above predictions will be made appropriately using a learning model that matches the field to be predicted.

[0043] Here, various objects captured in captured images are classified into classes such as the trunk class, branch class, dead branch class, and support class of agricultural crops, and the yield is predicted according to the class. Since the appearance of objects belonging to classes such as the trunk class and branch class changes depending on the time of shooting, it is difficult to make a universal prediction. Figure 4 shows such a difficult example.

[0044] 4(A) and 4(B) show examples of images captured by the camera 10. These captured images show crop trees at approximately equal intervals, but the fruits to be harvested have not yet borne fruit, so the task of detecting fruits cannot be performed from the captured images. The tree in the captured image of FIG. 4(A) is a crop tree captured at a relatively early stage of the season, and the tree in the captured image of FIG. 4(B) is a tree captured at a stage when the leaves have grown to a certain extent. In the captured image of FIG. 4(A), all the branches have the same amount of leaves, so it can be determined that there are no areas of poor growth, and all areas can be determined as harvestable. On the other hand, in the captured image of FIG. 4(B), the leaves of the branches near the central region 41 in the captured image are clearly different from the others, so it is easy to determine that they are poorly grown. However, the appearance of the central region 41 (an area with few leaves) can also be found as a similar pattern near the region 40 in the captured image of FIG. 4(A). These two examples show that abnormal regions in agricultural trees cannot be identified by local patterns. In other words, it is not possible to make a judgment based on only local patterns, as in the specific object detection method described above, but it is necessary to make a judgment by reflecting the context obtained from the entire image.

[0045] In other words, unless the above specific object detection is performed using a learning model that has been trained using images of crops in similar growth conditions photographed in the past, sufficient performance will not be achieved.

[0046] In order to handle all kinds of cases, such as when an image is input that was taken in a new field where no images have been taken before, when an image is input that was taken under conditions different from those previously taken due to some external factor such as prolonged drought or extremely heavy rainfall, or when an image is input that was taken at a time convenient for the user, it is necessary to obtain a learning model that is trained under conditions close to those of the input image each time.

[0047] Here, we will explain what kind of annotation work is required when annotation work and deep learning are performed every time a farm field is photographed. For example, the results of annotation work performed on the photographed images in Figures 4(A) and 4(B) are shown in Figures 5(A) and 5(B), respectively.

[0048] Rectangular regions 500-504 in the captured image in FIG. 5(A) are image regions designated by annotation work. Rectangular region 500 is an image region designated as a region of normal branches, and rectangular regions 501-504 are image regions designated as a region of the tree trunk. Rectangular region 500 is an image region that represents a normal state of tree growth, and therefore this image region is an area that is highly relevant to yield prediction. Hereinafter, an area that represents a normal state of tree growth, such as rectangular region 500, and an area of ​​a part where fruit or the like can be harvested, will be referred to as a production area.

[0049] Rectangular regions 505-507 and 511-514 in the captured image in FIG. 5(B) are image regions designated by annotation work. Rectangular regions 505 and 507 are image regions designated as normal branch regions, and rectangular region 506 is an image region designated as an abnormal dead branch region. Regions that indicate abnormal conditions, such as rectangular region 506, and regions where fruit or the like cannot be harvested are referred to as non-productive regions. Rectangular regions 511-514 are image regions designated as tree trunk regions. The image regions that are determined to be regions where fruit or the like can be harvested (productive regions) are rectangular regions 505 and 507, and therefore rectangular regions 505 and 507 are regions that are closely related to yield prediction.

[0050] To carry out such annotation work on a large number of captured images (e.g., hundreds to thousands of images) each time a farm field is photographed is extremely costly (costs in time, labor, etc.). Therefore, in this embodiment, the captured images to be subject to such annotation work are narrowed down, thereby reducing the cost associated with the annotation work.

[0051] Next, a series of processes, which includes identifying captured images that require annotation work, accepting annotation work for the captured images, and performing additional learning of the learning model using the annotated captured images, will be described with reference to the flowchart in Fig. 2A. This additional learning makes it possible to perform additional learning with a relatively smaller number of captured images than learning using randomly selected captured images of farm fields. Therefore, it is possible to obtain good prediction results while keeping the cost of cumbersome manual annotation work to a minimum.

[0052] In step S20, the camera 10 captures an image of the farm field while a mobile object such as a farm tractor 32 or a drone 37 is moving, thereby generating a captured image of the farm field.

[0053] In step S21, the camera 10 attaches Exif information to the captured image generated in step S20, and transmits the captured image with the Exif information attached via the communication network 11 to the cloud server 12 and the information processing device 13.

[0054] In step S22, the CPU 131 of the information processing device 13 acquires information on the field and the crops photographed by the camera 10 (such as the variety and age of the crops, and the cultivation and pruning methods of the crops) as photographed field parameters. For example, the CPU 131 displays a GUI (Graphical User Interface) shown in Fig. 6(A) on the display device 14 and accepts input of the photographed field parameters from the user.

[0055] In the GUI of Fig. 6(A), a map of the entire farm field is displayed in area 600. The map of the farm field displayed in area 600 is divided into a number of sections, and each section is displayed with its own unique identifier (ID). The user operates user interface 15 to specify a location in area 600 that corresponds to the section photographed by camera 10 (i.e., the section on which the above-mentioned analysis processing is to be performed), or inputs the identifier of the section in area 601. When the user operates user interface 15 to specify a location in area 600 that corresponds to the section photographed by camera 10, the identifier of the section is displayed in area 601.

[0056] The user can operate the user interface 15 to input a crop name (the name of the agricultural product) in an area 602. The user can also operate the user interface 15 to input a variety of the agricultural product in an area 603. The user can also operate the user interface 15 to input a Trellis in an area 604. For example, if the agricultural product is grapes, a Trellis is a method of designing grape vines for growing grapes in a grape field. The user can also operate the user interface 15 to input a Planted Year in an area 605. For example, if the agricultural product is grapes, the Planted Year indicates the time when the grape vines were planted. Note that it is not essential to input the photography field parameters for all of these items.

[0057] When the user operates the user interface 15 to select a registration button 606, the CPU 131 of the information processing device 13 transmits the photographed farm field parameters of each item input in the GUI of Fig. 6(A) to the cloud server 12. The CPU 191 of the cloud server 12 stores (registers) the photographed farm field parameters transmitted from the information processing device 13 in the external storage device 196.

[0058] Furthermore, when the user operates the user interface 15 to instruct a correction button 607, the CPU 131 of the information processing device 13 enables correction of the photographed field parameters that have already been input in the GUI of FIG. 6(A).

[0059] The GUI in Fig. 6(A) is a GUI for inputting photographed field parameters, particularly assuming management of a grape field, but even if the purpose is the same, the photographed field parameters input by the user are not limited to those shown in Fig. 6(A). Similarly, even if the crop is not grapes, the photographed field parameters input by the user are not limited to those shown in Fig. 6(A). For example, when the crop name input in the area 602 is changed, the titles of the areas 603 to 605 and the photographed field parameters to be input may be changed.

[0060] Since the photographed field parameters entered in the GUI in Fig. 6(A) can basically be used as fixed parameters once they have been determined, for example, when photographing a field every year to predict yields, the already registered photographed field parameters can be called and used. If the photographed field parameters have already been registered for a desired division, the next time, as shown in Fig. 6(B), the photographed field parameters corresponding to the desired division can be displayed in regions 609 to 613 by pointing to a location in region 600 that corresponds to the desired division.

[0061] Here, it is desirable to input all correct photographed field parameters in order to select a learning model at a later stage, but even if there are photographed field parameters that could not be input because they are unknown to the user, subsequent processing can be carried out while they remain unknown.

[0062] Note that even if the photographed field parameters are not acquired in step S22, the process of step S22 is not essential since it simply means that "selection of a candidate learning model using the photographed field parameters" described below is not performed. The photographed field parameters do not need to be acquired, for example, when information about the field or crops (such as the variety or age of the crops, or the cultivation or pruning method of the crops) is not known. Note that if the photographed field parameters are not acquired, in the subsequent processes, N candidate learning models will be selected from "all learning models" rather than from "the selected M candidate learning models".

[0063] In step S23, a process is performed to select captured images that are learning data used for additional learning of the learning model. Details of the process in step S23 will be described with reference to the flowchart in FIG. 2B.

[0064] In step S230, the CPU 191 of the cloud server 12 determines whether or not the photographed field parameters have been acquired from the information processing device 13. If the result of this determination is that the photographed field parameters have been acquired from the information processing device 13, the process proceeds to step S231, and if the photographed field parameters have not been acquired from the information processing device 13, the process proceeds to step S234.

[0065] In step S231, the CPU 191 of the cloud server 12 generates query parameters from the Exif information attached to each captured image obtained from the camera 10 and the captured field parameters (the captured field parameters of the section corresponding to the captured image) obtained from the information processing device 13 and registered in the external storage device 196.

[0066] An example of the configuration of the query parameters is shown in Fig. 11(A). The query parameters in Fig. 11(A) are generated when the photographed farm field parameters in Fig. 6(B) are input.

[0067] The "query name" is set to "F5" input in the area 609. The "variety" is set to "Shiraz" input in the area 611. The "Trellis" is set to "Scott-Henry" input in the area 612. The "age" is set to the number of years elapsed from "2001" input in the area 613 to the shooting date and time (year) included in the Exif information as the age of the tree "19". The "shooting date" is set to "Oct 20", the shooting date and time (month and day) included in the Exif information. The "shooting time zone" is set to "12:00-14:00", the time zone between the oldest shooting date and time (time) and the most recent shooting date and time (time) among the shooting dates and times (times) in the Exif information attached to each captured image received from the camera 10. The "latitude, longitude" is set to the shooting position "35°28'S, 149°12'E" included in the Exif information.

[0068] The method of generating the query parameters is not limited to the above method. For example, data that a farmer of the crop has already used for field management may be read, and a set of parameters that match the above items may be used as the query parameters.

[0069] In some cases, information about some items may be unknown. For example, if information about the planted year or variety is unknown, it is not possible to fill in all items as shown in Figure 11(A). In this case, some of the query parameters will be left blank, as shown in Figure 11(C).

[0070] Next, in step S232, the CPU 191 of the cloud server 12 selects (narrows down) M (1 ≦ M < E) candidate learning models (candidate learning models) from among the E learning models (where E is an integer of 2 or more) stored in the external storage device 196. In this selection, a learning model trained based on an environment similar to the environment indicated by the query parameter is selected as a candidate learning model. In the external storage device 196, for each of the E learning models, a parameter set indicating the environment based on which the learning model was trained is stored. A configuration example of the parameter set of each learning model in the external storage device 196 is shown in FIG. 11(B).

[0071] "Model Name" is the name of the learning model, "Variety" is the variety of the agricultural crop on which the learning model was trained, and "Trellis" is the "method for designing grapevines for growing grapes in a vineyard" on which the learning model was trained. "Tree Age" is the age of the agricultural crop on which the learning model was trained, and "Shooting Date" is the date and time of the captured image of the agricultural crop used for learning by the learning model. "Shooting Time Zone" is the period from the oldest shooting date and time to the most recent shooting date and time among the captured images of the agricultural crop used for learning by the learning model. "Latitude, Longitude" is the shooting position "35°28’S, 149°12’E" of the captured image of the agricultural crop used for learning by the learning model.

[0072] Depending on the learning model, there are also those that learn by mixing datasets collected from blocks in multiple fields. Therefore, for example, there may be cases where the parameter set is set to include multiple settings (such as variety and tree age) like the learning models with model names "M004" and "M005".

[0073] Therefore, the CPU 191 of the cloud server 12 calculates the similarity between the query parameter and the parameter set for each learning model shown in FIG. 11(B), and selects the top M learning models in descending order of the similarity as candidate learning models.

[0074] The parameter sets of each learning model with model names M001, M002, ... are 1 ,M 2 , ..., the CPU 191 receives a query parameter Q and a parameter set M x The similarity D(Q,M x ) is calculated using the following formula (1).

[0075]

number

[0076] Here, q k represents the k-th element from the top in the query parameter Q. In the case of FIG. 11(A), the query parameter Q includes six elements, namely, “variety”, “Trellis”, “age”, “photographed date”, “photographed time zone”, and “latitude, longitude”, so k=1 to 6.

[0077] m x、k is the parameter set M x 11B, the parameter set includes six elements: "variety", "Trellis", "age", "photographed date", "photographed time zone", and "latitude, longitude", so k=1 to 6.

[0078] f k (a k , b k ) is element a k and b k This is a function for calculating the distance between k (a k , b k ) may be carefully set in advance through experiments, but the distance definition in the above formula (1) should basically be set to a larger value for learning models with different characteristics, so it can be simply set as follows.

[0079] In other words, elements are basically divided into two types: classification elements (variety, Trellis) and continuous value elements (age of tree, date of photography, etc.). Therefore, the function that specifies the distance between classification elements is defined as in the following formula (2), and the function that specifies the distance between continuous value elements is defined as in the following formula (3).

[0080]

number

[0081]

number

[0082] The functions for all elements (k) are implemented in advance as a rule-based function. In addition, α is set according to the influence of each element on the final inter-model distance. k For example, the difference due to the "variety" (k=1) does not appear much in the image, so α 1 As α approaches 0, the difference in "Trellis" (k=2) has a large effect. 2 Adjust it in advance, for example, by setting it large.

[0083] In addition, in the case of a learning model in which multiple settings are registered for "variety" or "age of tree," such as the learning models with model names "M004" and "M005" in Fig. 11(B), for example, in the case of "variety," distances are calculated for each setting registered for "variety," and the average distance is regarded as the distance corresponding to "variety." Similarly, in the case of "age of tree," distances are calculated for each setting registered for "age of tree," and the average distance is regarded as the distance corresponding to "age of tree."

[0084] In addition, the selection method is not limited to a specific one, so long as the CPU 191 of the cloud server 12 selects M learning models as candidate learning models based on the above similarity. For example, the CPU 191 of the cloud server 12 may select M learning models having a similarity equal to or greater than a threshold. In addition, the smaller the value of the similarity D calculated using the above formulas (1) to (3), the higher the similarity, and the larger the value of the similarity D calculated using the above formulas (1) to (3), the lower the similarity.

[0085] However, if all elements in the query parameters are empty, the process in step S231 is not performed, and as a result, the subsequent processes are performed with all learning models as candidate learning models.

[0086] The effects of selecting candidate learning models are manifold. First, by eliminating learning models with low probability as prior knowledge in this step, the processing time for subsequent steps such as ranking by scoring learning models can be significantly reduced. Also, even in rule-based scoring of learning models, if learning models that do not need to be compared are included as candidates, there is a possibility that the accuracy of the learning model selection will decrease, but this possibility can be minimized.

[0087] On the other hand, in step S233, the CPU 191 of the cloud server 12 selects P (P is an integer equal to or greater than 2) captured images from the captured images received from the camera 10 as model selection target images. The method of selecting P captured images from the captured images received from the camera 10 is not limited to a specific selection method. For example, the CPU 191 may randomly select P captured images from the captured images received from the camera 10, or may select them according to some criteria.

[0088] In step S234, the M candidate learning models (or all learning models) selected in step S232 and the P captured images selected in step S233 are used to select captured images with GT (learning data with GT) and captured images without GT (learning data without GT).

[0089] The captured image with GT (learning data with GT) is a captured image in which the detection of the image area related to the crop is performed relatively accurately. The captured image without GT (learning data without GT) is a captured image in which the detection of the image area related to the crop is not performed very accurately. Details of the process in step S234 will be described with reference to the flowchart in FIG. 2C.

[0090] In step S2340, the CPU 191 of the cloud server 12 performs, for each of the M candidate learning models, "an object detection process that is a process of detecting an object from each of the P captured images using the candidate learning model."

[0091] As a result, for each of the P captured images, the "result of the object detection process for the captured image" of each of the M candidate learning models is obtained. In this embodiment, the "result of the object detection process for the captured image" is position information of the image area (rectangular area, detection area) of the object detected from the captured image.

[0092] In step S2341, the CPU 191 obtains a score for the "result of the object detection process for each of the P captured images" for each of the M candidate learning models. Then, the CPU 191 ranks the M candidate learning models based on the scores, and selects N (N≦M) candidate learning models from the M candidate learning models.

[0093] In this case, since the captured images do not have labels (annotation information), accurate evaluation of detection accuracy is not possible. However, for objects that are systematically designed and maintained, such as farms, it is possible to predict and evaluate the accuracy of object detection processing using the following rules. The score for the result of object detection processing by the candidate learning model is calculated, for example, as follows.

[0094] In a typical farm field, crops are planted at equal intervals, as shown in Fig. 3(A) and (B). Therefore, it is difficult to detect objects such as the annotations (rectangular regions) shown in Fig. 5(A) and (B). vinegar When detecting an image, the detection is normal if rectangular areas are always detected continuously and equally from the left edge to the right edge of the image.

[0095] For example, in the case where the entire area from the left end to the right end of the photographed image is detected as an area where fruits or the like can be harvested, as in FIG. 5(A), a production area should be detected as rectangular area 500. Also, in the case where a rectangular area 506 that is a non-production area is present in the photographed image, as in FIG. 5(B), rectangular areas 505, 506, and 507 should be detected from the left end to the right end of the photographed image. If an object detection process is performed on the photographed image using a learning model that does not match the conditions of the photographed image, there is a possibility that some of the above rectangular areas will not be detected. The more the learning model corresponds to conditions that are farther from the conditions of the photographed image, the higher the possibility of this happening. Therefore, the simplest scoring method for evaluating candidate learning models can be, for example, the following method.

[0096] The target candidate learning model detects detection areas of a plurality of objects from the target photographed image. Therefore, the target photographed image is searched for a detection area in the vertical direction to count the number of pixels Cp in the area where the detection area is not present, and the ratio of the number of pixels Cp to the number of pixels in the width of the target photographed image is set as the penalty score of the target photographed image. In this way, a penalty score is obtained for each of the P photographed images on which the target candidate learning model has performed object detection processing, and the total value of the obtained penalty scores is set as the score of the target candidate learning model. By performing such processing for each of the M candidate learning models, the score of each candidate learning model is determined. Then, the M candidate learning models are ranked in order of decreasing score, and the top N candidate learning models are selected in order of decreasing score. When making this selection, a condition that "the score is less than a threshold value" may be added.

[0097] Also, as the score of the candidate learning model, a score inferred from the detection area of ​​the trunk of a tree, which is usually planted at equal intervals, may be obtained. As shown in Fig. 5(A), tree trunks should be detected at approximately equal intervals, such as rectangular areas 501, 502, 503, and 504, so the expected number of "detected tree trunk areas" for the width of the captured image is fixed. Images with fewer / more than the expected number are likely to have detection errors, so this detection number may be reflected in the score.

[0098] The N candidate learning models selected from the M candidate learning models (hereinafter simply referred to as "N candidate learning models") are learning models trained based on images captured in a shooting environment similar to the shooting environment of the captured image acquired in step S20. In other words, the N candidate learning models are learning models trained based on an environment similar to the environment indicated by the query parameters. The N candidate learning models are learning models that are relatively robust in terms of detection accuracy when detecting image areas related to agricultural crops from captured images.

[0099] Therefore, in step S2342, the CPU 191 acquires, from the group of captured images stored in the external storage device 196, the captured images used in learning each of the N candidate learning models as "GT-attached captured images".

[0100] In the steps up to this point, the learning models have been narrowed down by a predetermined scoring method. The tendency of the object detection results by the learning models selected in this step is generally similar for the majority of cases, but the object detection results often differ significantly for some cases. For example, for photographic images corresponding to learned events common to many learning models, or photographic images corresponding to simple cases that no learning model can mistake, all of the above N candidate learning models produce almost the same detection results. However, for cases that do not occur often in the photographic images that have been learned up to now, a phenomenon is observed in which the object detection results by each learning model are different.

[0101] Therefore, in step S2343, CPU 191 determines the captured image corresponding to an important event that has hardly been learned as the captured image for additional learning. More specifically, in step S2343, information of different parts in the object detection results by N candidate learning models is evaluated to determine the priority of the captured image to be additionally learned. Here, an example of this determination method will be described.

[0102] In step S2343, the CPU 191 determines a larger score for each of the P captured images as the arrangement patterns of the detection regions are more different among the N candidate learning models. Such a score can be obtained, for example, by calculating the following formula (4).

[0103]

number

[0104] Here, Score(z) is the score for the captured image Iz. Iz (Ma, Mb) is a function for calculating a score based on the difference between the result of the object detection process (arrangement pattern of the detection area) performed by the candidate learning model Ma on the photographed image Iz and the result of the object detection process (arrangement pattern of the detection area) performed by the candidate learning model Mb on the photographed image Iz. Various functions can be applied to such a function, and it is not limited to a specific function. For example, for each detection area Ra detected by the candidate learning model Ma from the photographed image Iz, a function is defined as T that calculates the difference between the position (e.g., the position of the upper left corner and the position of the lower right corner) of the detection area Rb detected by the candidate learning model Mb from the photographed image Iz that is closest to the detection area Ra and the position of the detection area Ra (e.g., the position of the upper left corner and the position of the lower right corner), and returns the sum of the calculated differences. Iz (Ma, Mb) may also be used.

[0105] Then, the CPU 191 specifies, from among the P captured images, captured images for which a score (a score calculated according to the formula (4)) is less than the threshold value, as GT-added captured images (GT-added learning data).

[0106] On the one hand, the CPU 191 identifies, as "photographed images requiring annotation work" (photographed images without GT (learning data without GT)), the photographed images among the P photographed images for which scores equal to or higher than the threshold value (scores obtained according to Equation (4)) are obtained. Then, the CPU 191 transmits the photographed images (photographed images without GT) identified as "photographed images requiring annotation work" to the information processing apparatus 13.

[0107] In step S24, the CPU 131 of the information processing apparatus 13 receives the photographed images without GT transmitted from the cloud server 12, and stores the received photographed images without GT in the RAM 132. Note that the CPU 131 of the information processing apparatus 13 may display the photographed images without GT received from the cloud server 12 on the display device 14 to present the photographed images without GT to the user.

[0108] In step S25, since the user of the information processing apparatus 13 operates the user interface 15 to perform annotation work on the photographed images without GT received from the cloud server 12, the CPU 131 accepts the annotation work. Then, the CPU 131 assigns the label input in the annotation work for the photographed images without GT to the photographed images without GT, so that the photographed images without GT become photographed images with GT.

[0109] Here, not only the photographed images without GT received from the cloud server 12, but also, for example, the photographed images identified as follows may be identified as targets for the user to perform annotation work.

[0110] The CPU 191 of the cloud server 12 identifies the top Q (Q < P) captured images (or other captured image groups) in descending order of the scores (scores obtained according to Equation (4)) from the P captured images. Then, the CPU 191 transmits the Q captured images, the scores of each of the Q captured images, the "results of object detection processing for the Q captured images" of each of the N candidate learning models, information about the N candidate learning models (such as model names), etc. to the information processing device 13. As described above, in this embodiment, the "results of object detection processing for the captured image" are the position information of the image regions (rectangular regions, detection regions) of the objects detected from the captured image. Such position information is transmitted to the information processing device 13 as data in a file format such as the json format or the txt format, for example.

[0111] For each of the above N candidate learning models, the CPU 131 of the information processing device 13 causes the display device 14 to display the Q captured images received from the cloud server 12 and the results of object detection processing for the captured images received from the cloud server 12. At this time, the Q captured images are arranged and displayed from left to right in descending order of the scores.

[0112] A display example of the GUI displaying the captured images and the results of object detection processing for each candidate learning model is shown in FIG. 7(A). FIG. 7(A) shows the case where N = 3 and Q = 4.

[0113] In the top row, the model name "M002" of the candidate learning model with the highest score is displayed. On the right side, the top 4 captured images in descending order of the scores are arranged and displayed in order from left to right together with the check boxes 70. A frame indicating the detection region of the object detected from the captured image by the candidate learning model with the model name "M002" is superimposed and displayed on the captured image.

[0114] The middle row displays the model name "M011" of the candidate learning model with the second highest score, and to the right of that, the top four captured images with the highest scores are displayed in order from left to right along with check boxes 70. A frame indicating the detection area of ​​the object detected from the captured image by the candidate learning model with the model name "M011" is superimposed on the captured image.

[0115] The bottom row displays the model name "M009" of the candidate learning model with the third highest score, and to the right of that, the top four captured images with the highest scores are displayed in order from left to right along with check boxes 70. A frame indicating the detection area of ​​the object detected from the captured image by the candidate learning model with the model name "M009" is superimposed on the captured image.

[0116] In this GUI, captured images arranged in the same row are displayed as the same captured images, so that the results of the object detection processing using each candidate learning model can be easily compared at a glance.

[0117] In the example of Fig. 7(A), in the subsequent additional learning, the annotated photographed images used in the learning of the three candidate learning models are used by the three candidate learning models, and an additional "photographed image that is likely to express a phenomenon not learned in the annotated photographed image" is identified.

[0118] Here, the relationship between a set of captured images and the results of object detection processing by three candidate learning models for each captured image belonging to the set will be described using the Venn diagram in FIG. 12. In the Venn diagram in FIG. 12, the results of object detection processing by three candidate learning models (each of which has a model name of "M002", "M009", and "M011") are expressed as a binary value. The inside of each circle, which corresponds to "M002", the circle which corresponds to "M009", and the circle which corresponds to "M011", represent a set of captured images for which a correct object detection processing result has been obtained. Also, the outside of each circle, which corresponds to "M002", the circle which corresponds to "M009", and the circle which corresponds to "M011", represent a set of captured images for which an incorrect object detection processing result has been obtained.

[0119] The set of captured images included in area 127, i.e., the set of captured images for which correct object detection processing results were obtained using all three candidate learning models, are captured images that are considered to have already been learned using the three candidate learning models, and therefore are of little value in being added as targets for additional learning.

[0120] The set of captured images included in the region 128, i.e., the set of captured images for which incorrect object detection processing results were obtained in all of the three candidate learning models, are considered to be captured images that have not been learned in the three candidate learning models or images that represent insufficiently learned events. Therefore, the captured images included in the region 128 are highly likely to be captured images that should be actively added to the targets of additional learning.

[0121] The captured images displayed in the GUI of Fig. 7(A) are likely to include not only the captured image corresponding to area 128, but also captured images (captured images included in area 121) in which only the candidate learning model with the model name "M002" has obtained a correct object detection processing result, captured images (captured images included in area 122) in which only the candidate learning model with the model name "M009" has obtained a correct object detection processing result, and captured images (captured images included in area 123) in which only the candidate learning model with the model name "M011" has obtained a correct object detection processing result. In addition, depending on the difference in the arrangement pattern of the detection areas, captured images corresponding to areas 124, 125, and 126 may also be included in the captured images displayed in the GUI of Fig. 7(A).

[0122] Therefore, a system that does not know the true answer simply displays the captured images determined based on the score (score based on the difference in the results of the object detection process) calculated according to formula (4) as "candidate captured images to be additionally learned." Therefore, the missing captured images that are not yet included in the learned captured images must be determined by instruction from the user.

[0123] Therefore, the CPU 131 of the information processing device 13 accepts a user's operation to specify "a captured image to be subjected to annotation work." In the case of FIG. 7(A), the user checks the results of the object detection process for each of the candidate learning models with the model name "M002," the candidate learning model with the model name "M011," and the candidate learning model with the model name "M009." Then, the user operates the user interface 15 to specify and turn on (put a check mark on) the check box 70 of the captured image for which the user has determined that the result of the object detection process is favorable.

[0124] In the example of Fig. 7(A), the check boxes 70 of the top and middle images in the first column from the left are checked, and none of the check boxes 70 of the images in the second column from the left are checked. Also, the check box 70 of the middle image in the third column from the left is checked, and the check boxes 70 of all of the images in the fourth column from the left are checked.

[0125] When the user operates the user interface 15 to select the decision button 71, the CPU 131 of the information processing device 13 counts the number of captured images with check marks for each column of captured images. The CPU 131 of the information processing device 13 then identifies the captured images corresponding to the columns whose scores based on the counted numbers are equal to or greater than a threshold as "captured images to be additionally studied (captured images to be subjected to annotation work for that purpose)."

[0126] For images corresponding to columns without check marks, the object detection process results in "fail" in all three candidate learning models, so the images are judged to be images included in area 128 and are selected as images with high importance for further learning.

[0127] On the other hand, the captured images corresponding to the columns with all check marks have the object detection process result of "success" in all three candidate learning models, and therefore are selected as captured images with low importance for additional learning.

[0128] In many cases, the captured images for which similar object detection processing results are obtained in all candidate learning models according to the score obtained in accordance with formula (4) should not be displayed in the above GUI. However, in cases where the detection area arrangement patterns are different but have the same meaning, or in cases where the detection area arrangement patterns are different but both are correct depending on the use case, the check boxes 70 of all captured images in a vertical column may be checked. Therefore, in the above GUI, a score is obtained for the captured images in a column with fewer check marks, so that the importance of additional learning is higher, and the captured images to be annotation-processed are identified from the Q captured images based on the score. Such a score can be obtained, for example, according to the following formula (5).

[0129]

number

[0130] Here, Score(Iz) is the score for the captured image Iz. Iz w indicates the number of captured images in the column of the captured image Iz for which the check box 70 is turned on (the number of check marks in the column). z is a weighting value proportional to the score of the captured image Iz obtained according to the above formula (4).

[0131] Then, the CPU 131 of the information processing apparatus 13 identifies, as "the captured image to be subject to annotation work", the captured image among the Q captured images whose score obtained by the formula (5) is equal to or greater than the threshold value. For example, the captured image corresponding to the column having no check mark may be identified as "the captured image to be subject to annotation work". Thus, when the "captured image to be subject to annotation work" is identified as described above by operating the GUI in Fig. 7(A), the user of the information processing apparatus 13 operates the user interface 15 to perform annotation work on the "captured image to be subject to annotation work". Therefore, in step S25, the CPU 131 accepts the annotation work and assigns the label input in the annotation work for the captured image to the captured image.

[0132] Also, the result of the object detection process displayed in the GUI of Fig. 7(A) for the captured image in which the check box 70 is on may be used as the label for the captured image, and the labeled captured image may be included in the target of the additional learning described above.

[0133] Note that for a user who understands the criteria for identifying the above "captured image to be subject to annotation work", directly selecting the "captured image to be subject to annotation work" may facilitate the input work. In that case, the "captured image to be subject to annotation work" may be identified according to the user operation via the GUI as shown in Fig. 7(B).

[0134] In the GUI of FIG. 7(B), a radio button 72 is provided for each column of captured images. When the user operates the user interface 15 to turn on the radio button 72 corresponding to the first column from the left, the captured image corresponding to that column is specified as a "captured image to be subjected to annotation work". When the user operates the user interface 15 to turn on the radio button 72 corresponding to the second column from the left, the captured image corresponding to that column is specified as a "captured image to be subjected to annotation work". When the user operates the user interface 15 to turn on the radio button 72 corresponding to the third column from the left, the captured image corresponding to that column is specified as a "captured image to be subjected to annotation work". When the user operates the user interface 15 to turn on the radio button 72 corresponding to the fourth column from the left, the captured image corresponding to that column is specified as a "captured image to be subjected to annotation work".

[0135] When using such a GUI to specify "a captured image that is the subject of annotation work," it is sufficient to turn on the radio button 72 corresponding to the captured image in which an object is likely to be erroneously detected.

[0136] In this way, when the "captured image to be the subject of annotation work" is identified as described above by operating the GUI of Figure 7 (B), the user of the information processing device 13 operates the user interface 15 to perform annotation work on the "captured image to be the subject of annotation work", so in step S25, the CPU 131 accepts the annotation work and assigns the label inputted in the annotation work for the captured image to the captured image.

[0137] Then, the CPU 131 of the information processing device 13 transmits, to the cloud server 12, the captured image on which the annotation work has been performed by the user (the captured image with the GT).

[0138] In step S26, the CPU 191 of the cloud server 12 performs additional learning of the N candidate learning models using the captured image (captured image with GT) to which the label was added in step S25 and the "captured image (captured image with GT) used in learning each of the N candidate learning models" acquired in step S2342. Then, the CPU 191 of the cloud server 12 stores the N candidate learning models that have undergone additional learning again in the external storage device 196.

[0139] As an example of the learning and inference method used here, a region-based CNN technique such as Fater RCNN is used. With this method, learning is possible if there is a set of annotation information of the rectangular coordinates and labels used in this embodiment and an image.

[0140] In this way, according to this embodiment, even if an image taken in an unknown field is input, non-productive areas and the like can be detected with high accuracy for each image. In particular, it is possible to predict the yield of crops harvested from a target field by integrating the ratio obtained by subtracting the proportion of non-productive areas estimated by this method from the yield per unit area when 100% harvest is possible.

[0141] If you want to repair an area where the width of a rectangular area detected as a non-productive area exceeds a user-defined value, you can specify the target image based on the width of the detected rectangular area and xif The system uses information such as location on a map to show the user where the tree in need of repair is located.

[0142] Note that the learning model in this embodiment is a model learned by deep learning, but various object detection techniques such as a rule-based detector defined by various parameters, fuzzy inference, genetic algorithms, etc. may also be used as the learning model.

[0143] [Second embodiment] In the following embodiments, differences from the first embodiment will be described, and unless otherwise specified below, it is assumed that the present embodiment is the same as the first embodiment. In the present embodiment, a system for performing visual inspection in a factory production line will be described as an example. The system according to the present embodiment detects abnormal areas in industrial products that are the objects of inspection.

[0144] Traditionally, in visual inspections on factory production lines, the imaging conditions of the inspection equipment (equipment that photographs and inspects the appearance of products) were carefully adjusted for each production line, and it was common to spend time adjusting the settings of the inspection equipment each time a production line was started up. However, in recent years, manufacturing sites are expected to respond immediately to diversifying customer needs and changes in the market. There is an increasing need for speedy responses, such as setting up a line in a short period of time even for small lots, producing a quantity that matches demand, and then immediately dismantling the line once sufficient supply has been completed to prepare for the next production line.

[0145] In this case, if the settings for visual inspection are set every time based on the experience and intuition of a manufacturing site expert, as in the past, it will not be possible to respond to a rapid start-up.If similar products have been inspected in the past, and the related setting parameters can be stored and the past setting parameters can be called up when a similar inspection is to be performed, anyone can set up the inspection device without relying on the experience of an expert.

[0146] As in the first embodiment, the above object can be achieved by assigning a learning model already held to an image of a new product to be inspected. Therefore, the above system can be applied to the second embodiment.

[0147] The setting process for the inspection device by the system according to this embodiment (setting process for visual inspection) will be described with reference to the flowchart in Fig. 8A. Note that the setting process for visual inspection is assumed to be performed at the start of an inspection step in a manufacturing line.

[0148] The external storage device 196 of the cloud server 12 has multiple learning models (appearance inspection models / settings) registered for performing appearance inspection on captured images, and each learning model is a model that has been learned in a mutually different learning environment.

[0149] The camera 10 is a camera for photographing a product that is the subject of visual inspection (product to be inspected). As in the first embodiment, the camera 10 may be a camera that photographs periodically or irregularly, or may be a camera that photographs moving images. In order to accurately detect an abnormal area in the product to be inspected from the photographed image, when the product to be inspected that includes an abnormal area enters the inspection process, it is desirable to photograph the product under conditions that emphasize the abnormal area as much as possible. The camera 10 may be a multi-camera if the product to be inspected is photographed under multiple conditions.

[0150] In step S80, the camera 10 captures an image of the product to be inspected to generate a captured image of the product to be inspected. In step S81, the camera 10 transmits the captured image generated in step S80 to the cloud server 12 and the information processing device 13 via the communication network 11.

[0151] In step S82, the CPU 131 of the information processing device 13 acquires information relating to the inspection target product photographed by the camera 10 (part name and material of the inspection target product, manufacturing date, photographing system parameters at the time of photographing, lot number, temperature, humidity, etc.) as inspection target product parameters. For example, the CPU 131 displays a GUI on the display device 14 to accept input of the inspection target product parameters from the user. Then, when the user operates the user interface 15 to input a registration instruction, the CPU 131 of the information processing device 13 transmits the inspection target product parameters of each of the above items input in the GUI to the cloud server 12. The CPU 191 of the cloud server 12 saves (registers) the inspection target product parameters transmitted from the information processing device 13 in the external storage device 196.

[0152] Note that even if the parameters of the product to be inspected are not acquired in step S82, the process of step S82 is not essential since the "selection of a candidate learning model using the parameters of the product to be inspected" described below is not performed. The parameters of the product to be inspected do not need to be acquired, for example, when information about the product to be inspected photographed by the camera 10 (part name and material of the product to be inspected, manufacturing date, photographing system parameters at the time of photographing, lot number, temperature, humidity, etc.) is not known. Note that if the parameters of the product to be inspected are not acquired, in the subsequent process, N candidate learning models will be selected from "all learning models" rather than from "selected M candidate learning models".

[0153] In step S83, a process for selecting captured images to be used for learning the learning model is performed. Details of the process in step S83 will be described with reference to the flowchart in FIG. 8B.

[0154] In step S830, the CPU 191 of the cloud server 12 determines whether or not the inspection target product parameters have been acquired from the information processing device 13. If the result of this determination is that the inspection target product parameters have been acquired from the information processing device 13, the process proceeds to step S831, and if the inspection target product parameters have not been acquired from the information processing device 13, the process proceeds to step S833.

[0155] In step S831, the CPU 191 of the cloud server 12 selects M candidate learning models (candidate learning models) from the E learning models stored in the external storage device 196. The CPU 191 generates query parameters from the inspection target product parameters and Exif information registered in the external storage device 196 in the same manner as in the first embodiment, and selects a learning model that has learned about an environment similar to the environment indicated by the query parameters (a learning model used in a similar past inspection).

[0156] If the query parameter contains "PCB" as the "part name", a learning model that was used in past PCB inspections is more likely to be selected. Furthermore, if the query parameter contains "glass epoxy" as the "material", a learning model that was used in the inspection of glass epoxy PCBs is more likely to be selected.

[0157] In step S831, as in the first embodiment, M candidate learning models are selected using the parameter set of the learning model and the query parameters, and in this case, the above formula (1) is used as in the first embodiment.

[0158] Next, in step S832, the CPU 191 of the cloud server 12 selects P images from the images received from the camera 10. For example, products flowing through the main inspection process of the main production line are randomly selected, and P images are acquired from images captured by the camera 10 with the same settings as those during actual operation. Since the number of abnormal products that occur on a production line is usually small, the processing in the following steps does not function well when the number of products photographed in the process is small. Therefore, as a guideline, it is desirable to photograph several hundred or more products.

[0159] In step S833, the M candidate learning models (or all learning models) selected in step S831 and the P captured images selected in step S832 are used to select captured images with GT (learning data with GT) and captured images without GT (learning data without GT).

[0160] The captured image with GT (learning data with GT) according to this embodiment is a captured image in which detection of abnormal areas of the industrial product to be inspected is performed relatively accurately. The captured image without GT (learning data without GT) is a captured image in which detection of abnormal areas of the industrial product to be inspected is not performed very accurately. Details of the process in step S833 will be described with reference to the flowchart in FIG. 8C.

[0161] In step S8330, the CPU 191 of the cloud server 12 performs "object detection processing, which is processing for detecting an object from each of the P photographed images using the candidate learning model for each of the M candidate learning models" for each of the M photographed images. In this embodiment, too, the result of the object detection processing for the photographed image is position information of the image area (rectangular area, detection area) of the object detected from the photographed image.

[0162] In step S8331, the CPU 191 obtains a score for the "result of the object detection process for each of the P captured images" for each of the M candidate learning models. The CPU 191 then ranks the M candidate learning models based on the scores, and selects N (N≦M) candidate learning models from the M candidate learning models. The score for the result of the object detection process using the candidate learning models is obtained, for example, as follows.

[0163] For example, in a task of detecting anomalies on a printed circuit board, object detection processing is performed on various specific local patterns on a fixed print pattern. Here, it is assumed that detection areas 901 to 906 as shown in FIG. 9(A) are obtained from a captured image of a normal product in a specific learning model. Since the frequency of occurrence of anomalies in products produced on a manufacturing line is extremely low, a good learning model for performing the above task is one that can output stable results against expected variations in captured images. For example, a slight change in the appearance of an image of a product due to a change in the environment on the area sensor side may make it impossible to detect detection area 906 out of detection areas 901 to 906 as shown in FIG. 9(B). In such a case, a penalty should be imposed on the evaluation score of a learning model in which the detection area changes for an input with only a slight difference.

[0164] Therefore, for example, the CPU 191 of the cloud server 12 determines a higher score for each of the M candidate learning models, the greater the difference in the arrangement pattern of the detection area by the candidate learning model between the P captured images. Then, the CPU 191 ranks the M candidate learning models in ascending order of score, and selects the top N candidate learning models in ascending order of score. When making this selection, a condition that "the score is less than a threshold value" may be added.

[0165] In step S8332, the CPU 191 acquires, from the group of captured images stored in the external storage device 196, the captured images used in learning each of the N candidate learning models as "GT-attached captured images."

[0166] In step S8333, a captured image corresponding to an important event that has hardly been learned is determined as a captured image for additional learning. More specifically, in step S8333, information of different parts in the object detection results by N candidate learning models is evaluated to determine the priority of the captured image to be additionally learned. Here, an example of this determination method will be described.

[0167] In step S8333, the CPU 191, similarly to step S2343 described above, specifies captured images having scores (scores calculated according to formula (4)) less than a threshold value among the P captured images as GT-added captured images (GT-added learning data).

[0168] On the other hand, the CPU 191 specifies, among the P captured images, a captured image having a score equal to or higher than a threshold (a score calculated according to formula (4)) as a "captured image requiring annotation work" (a captured image without GT (learning data without GT)). Then, the CPU 191 transmits the captured image specified as the "captured image requiring annotation work" (a captured image without GT) to the information processing device 13.

[0169] In step S84, the CPU 131 of the information processing apparatus 13 receives the captured image without GT transmitted from the cloud server 12, and stores the received captured image without GT in the RAM 132.

[0170] In step S85, since the user of the information processing apparatus 13 operates the user interface 15 to perform an annotation operation on the captured image without GT received from the cloud server 12, the CPU 131 accepts the annotation operation. Then, the CPU 131 assigns the label input in the annotation operation for the captured image without GT to the captured image without GT, and the captured image without GT becomes a captured image with GT.

[0171] Here, not only the captured image without GT received from the cloud server 12, but also, for example, the captured images specified as follows may be specified as the objects for which the user performs the annotation operation.

[0172] The CPU 191 of the cloud server 12 specifies the top Q (Q < P) captured images in descending order of the scores (scores obtained according to Equation (4)) from P captured images (or other groups of captured images). Then, the CPU 191 transmits the Q captured images, the scores of each of the Q captured images, the "results of object detection processing for the Q captured images" of each of the N candidate learning models, information (such as model names) regarding the N candidate learning models, etc. to the information processing apparatus 13.

[0173] The CPU 131 of the information processing apparatus 13 causes the display device 14 to display the Q captured images received from the cloud server 12 and the results of object detection processing for the captured images received from the cloud server 12 for each of the above N candidate learning models. At that time, the Q captured images are arranged and displayed from left to right in descending order of the scores.

[0174] An example of the display of the GUI showing the captured images and the results of object detection processing for each candidate learning model is shown in FIG. 10(A). FIG. 10(A) shows the case where N = 3 and Q = 4.

[0175] The top row displays the model name "M005" of the candidate learning model with the highest score, and to the right of that, the top four captured images with the highest scores are displayed in order from left to right along with check boxes 100. A frame indicating the detection area of ​​an object detected from the captured image by the candidate learning model with the model name "M005" is superimposed on the captured image.

[0176] The middle row displays the model name "M023" of the candidate learning model with the second highest score, and to the right of that, the top four captured images with the highest scores are displayed from left to right along with check boxes 100. A frame indicating the detection area of ​​the object detected from the captured image by the candidate learning model with the model name "M023" is superimposed on the captured image.

[0177] The bottom row displays the model name "M014" of the candidate learning model with the third highest score, and to the right of that, the top four captured images with the highest scores are displayed in order from left to right along with check boxes 100. A frame indicating the detection area of ​​the object detected from the captured image by the candidate learning model with the model name "M014" is superimposed on the captured image.

[0178] In this GUI, captured images arranged in the same row are displayed as the same captured images so that the results of the object detection process by each candidate learning model can be easily compared at a glance. The user operates the user interface 15 to specify and turn on (put a check mark in) the check box 100 of the captured image that the user judges to have a favorable result of the object detection process.

[0179] When the user operates the user interface 15 to select the decision button 101, the CPU 131 of the information processing device 13 counts the number of captured images with check marks for each column of captured images. The CPU 131 of the information processing device 13 then identifies the captured images corresponding to the columns for which the score based on the counted numbers is equal to or greater than a threshold as "captured images to be additionally studied (captured images to be subjected to annotation work)." In this way, the series of processes for identifying "captured images to be subjected to annotation work" is the same as in the first embodiment.

[0180] In this way, when the "captured image to be the subject of annotation work" is identified as described above by operating the GUI of Figure 10 (A), the user of the information processing device 13 operates the user interface 15 to perform annotation work on the "captured image to be the subject of annotation work", so in step S85, the CPU 131 accepts the annotation work and assigns the label inputted in the annotation work for the captured image to the captured image.

[0181] In addition, the result of the object detection process displayed in the GUI of Figure 10 (A) for a captured image for which check box 100 is turned on may be used as a label for the captured image, and the captured image with the label may be included in the target of the above-mentioned additional learning.

[0182] For a user who understands the criteria for identifying the above-mentioned "captured image to be annotation-worked", it may be easier to input the information by directly selecting the "captured image to be annotation-worked". In this case, the "captured image to be annotation-worked" may be identified according to a user operation via a GUI as shown in Fig. 10(B).

[0183] The method for specifying the “captured image to be the subject of annotation work” using the GUI in Figure 10(B) is similar to the method for specifying the “captured image to be the subject of annotation work” using the GUI in Figure 7(B), so explanation will be omitted.

[0184] Then, the CPU 131 of the information processing device 13 transmits, to the cloud server 12, the captured image on which the annotation work has been performed by the user (the captured image with the GT).

[0185] In step S86, the CPU 191 of the cloud server 12 performs additional learning of the N candidate learning models using the captured image (captured image with GT) to which the label was added in step S85 and the "captured image (captured image with GT) used in learning each of the N candidate learning models" acquired in step S8332. Then, the CPU 191 of the cloud server 12 stores the N candidate learning models that have undergone additional learning again in the external storage device 196.

[0186] <Modification> Each of the above embodiments is an example of a technology for reducing the cost of learning a learning model and adjusting the settings each time a detection / identification process is performed on a new target in a task for performing a target detection / identification process. Therefore, the technology described in each of the above embodiments is not limited to application to prediction of crop yields, detection of repair areas, detection of abnormal areas in industrial products to be inspected, etc., but is also applicable to agriculture, industry, fisheries, and a wide range of other fields.

[0187] Furthermore, the above radio buttons and check boxes are displayed as examples of selection sections for the user to select an object, and other display items may be displayed instead as long as they can achieve similar functionality.

[0188] Also, the subject of each process in the above description is just an example. For example, a part or all of the processes described as being performed by the CPU 191 of the cloud server 12 may be performed by the CPU 131 of the information processing device 13. Also, a part or all of the processes described as being performed by the CPU 131 of the information processing device 13 may be performed by the CPU 191 of the cloud server 12.

[0189] In the above description, the system of each embodiment is described as performing the analysis process. However, the subject of the analysis process is not limited to the system of each embodiment, and for example, the analysis process may be performed by another device / system.

[0190] In addition, the various functions described above may be implemented by the information processing device 13 as the functions of the cloud server 12. In that case, the system does not need to include the cloud server 12. In addition, the method of acquiring the learning model is not limited to a specific method. In addition, various object detectors may be applied instead of the learning model.

[0191] Furthermore, the numerical values, processing timing, processing order, processing subject, data (information) structure / destination / source / storage location, etc. used in each of the above embodiments and variant examples are given as examples to provide a concrete explanation, and are not intended to be limited to such examples.

[0192] In addition, some or all of the embodiments and modified examples described above may be used in appropriate combination. In addition, some or all of the embodiments and modified examples described above may be used selectively.

[0193] (Other embodiments) The present invention can also be realized by a process in which a program for implementing one or more of the functions of the above-described embodiments is supplied to a system or device via a network or a storage medium, and one or more processors in a computer of the system or device read and execute the program. The present invention can also be realized by a circuit (e.g., ASIC) that implements one or more of the functions.

[0194] The invention is not limited to the above-described embodiments, and various modifications and variations are possible without departing from the spirit and scope of the invention. Accordingly, the following claims are appended to apprise the public of the scope of the invention. [Explanation of symbols]

[0195] 10: Camera 11: Communication network 12: Cloud server 13: Information processing device 14: Display device 15: User interface 131: CPU 132: RAM 133: ROM 134: Output I / F 135: Input I / F 191: CPU 192: RAM 193: ROM 194: Operation unit 195: Display unit 196: External storage device 197: I / F 198: System bus

Claims

1. A plurality of learning models trained to output the results of object detection processing as image areas relating to objects in a photographed image when the photographed image is input, comprising: a selection means for selecting from among some or all of the plurality of learning models which satisfy a predetermined condition based on the results of object detection processing on a plurality of photographed images using some or all of the plurality of learning models trained with mutually different parameter sets; and an acquisition means for obtaining a score based on a difference between results of object detection processing for the plurality of photographed images using the plurality of learning models selected by the selection means, for each of the plurality of photographed images, based on results of object detection processing for the plurality of photographed images using the plurality of learning models selected by the selection means; A specification means for specifying, among the plurality of photographed images, photographed images having a score equal to or greater than a threshold value as photographed images to be subjected to an annotation work, which is a work of assigning a label as teacher data indicating a correct answer; and a means for performing additional learning using the captured image identified by the identification means and a captured image used in the initial learning of the learning model selected by the selection means, so as to output a result of an object detection process that is a result of an image area related to an object in the captured image when the captured image is input, for the learning model selected by the selection means, The acquisition means obtains the score for each of the plurality of captured images, the score being defined so as to be larger as the results of the object detection process between the learning models selected by the selection means differ more significantly.

23. An information processing apparatus comprising:

2. The information processing device according to claim 1, characterized in that the selection means generates a query parameter set which is a set of parameters related to the type of object and the shooting environment of the object, compares a plurality of parameters included in the query parameter set with a plurality of learning parameters included in a parameter set used by the plurality of learning models during learning, calculates a similarity between the plurality of parameters and the plurality of learning parameters, selects the plurality of learning models whose similarity is equal to or greater than a threshold, and selects a plurality of learning models from the plurality of learning models or all of the learning models that satisfy the predetermined condition according to a predetermined condition based on a result of object detection processing on the plurality of captured images by some or all of the learning models among the selected plurality of learning models.

3. Furthermore, The information processing device according to claim 1 or 2, further comprising a display means for displaying the results of object detection processing using the learning model selected by the selection means for each of a specified number of captured images from the top of the plurality of captured images in order of the largest score.

4. the display means displays a graphical user interface including a selection section for selecting a result of the object detection process performed by the learning model selected by the selection means for each of the specified number of captured images; The identification means counts the number of selection parts selected in response to a user operation for each of the specified number of captured images in addition to the captured images identified based on a score based on a difference between results of object detection processing by the multiple learning models selected by the selection means, and obtains a higher score as the counted number becomes smaller. The identified captured images having a score equal to or greater than a threshold value among the specified number of captured images are identified as captured images to be subjected to annotation work.

4. The information processing apparatus according to claim 3.

5. the display means displays a graphical user interface including a selection section for each of the specified number of captured images; The specifying means specifies, as a target image for annotation work, a captured image corresponding to a selection portion selected in response to a user operation among the specified number of captured images, in addition to the captured image specified based on a score based on a difference between results of object detection processing by the multiple learning models selected by the selecting means.

4. The information processing apparatus according to claim 3.

6. Furthermore, The information processing device according to any one of claims 1 to 5, characterized in that it is equipped with a means for predicting the yield of a crop based on the number of the crops obtained from the detection area of ​​the crop obtained as a result of the object detection process, and for detecting areas of poor crop growth obtained from the detection area of ​​repair areas in the field in which the crop is produced as repair areas in the field.

7. Furthermore, 6. The information processing device according to claim 1, further comprising: a means for performing an appearance inspection of an object in a captured image based on an abnormal area of ​​the object obtained from a detection area of ​​the object obtained as a result of the object detection process.

8. The information processing apparatus according to claim 2 , wherein the parameters include photographing information at the time of photographing the photographed image, photographed field parameters of a section corresponding to the photographed image, and information relating to a type of object included in the photographed image.

9. The information processing apparatus according to claim 2 , wherein the parameters include information relating to a type of an object included in the photographed image.

10. An information processing method performed by an information processing device, a selection step in which a selection means of the information processing device selects a plurality of learning models that satisfy a predetermined condition from among a part or all of the learning models, the learning models being trained to output a result of an object detection process as an image area related to an object in a captured image when the captured image is input, according to a predetermined condition based on a result of an object detection process performed on a plurality of captured images by a part or all of the learning models among the plurality of learning models trained with mutually different parameter sets; an acquisition step in which an acquisition means of the information processing device obtains a score based on a difference between results of object detection processing for the plurality of photographed images using the plurality of learning models selected in the selection step, for each of the plurality of photographed images, based on results of object detection processing for the plurality of photographed images using the plurality of learning models selected in the selection step; a step of identifying, by a identifying means of the information processing device, a captured image having a score equal to or greater than a threshold among the plurality of captured images as a target image for annotation work, which is a work of assigning a label as teacher data indicating a correct answer; and a step of performing additional learning using the captured image identified in the identification step, to which a label has been added as teacher data indicating a correct answer as a result of annotation work performed on the captured image, and the captured image used in the initial learning of the learning model selected in the selection step, so that the learning model selected in the selection step outputs a result of an object detection process that is a result of an image area related to an object in the captured image when the captured image is input, In the obtaining step, the score is obtained for each of the plurality of captured images, the score being defined so as to be larger as the results of the object detection processing between the learning models selected in the selecting step are significantly different.

23. An information processing method comprising:

11. A computer program for causing a computer to function as each of the means of the information processing device according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Plant maintenance system

    JP2005137209A

  • Farm field management system, farm field management method, and program

    JP2016049102A

  • Information processing device, control method of information processing device and program

    JP2019087229A

  • Learning apparatus, learning system and learning method

    JP2020008905A

  • Image recognition system and image recognition method

    JP2020160966A