Information Processing Apparatus, Information Processing Method, and Recording Medium

By using multiple learned models in the information processing device to detect different types of objects separately and integrate the results, the problem of insufficient detection accuracy in the prior art is solved, and object recognition with higher accuracy is achieved, especially in effectively detecting improper objects in waste incineration facilities.

CN112651281BActive Publication Date: 2025-07-11科纳维株式会社 +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202011063846.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-10-11
Filing Date
2020-09-30
Publication Date
2025-07-11
Estimated Expiration
2040-09-30

AI Technical Summary

Technical Problem

现有技术中,机器学习模型在物体检测时存在检测精度不足的问题,尤其在行人检测和招牌检测时容易出现漏检或误检,且无法有效补救。

Method used

Multiple learned models are used to detect different types of objects, and the final detection results are determined by integrating the detection results of these models, including the first detection unit and the second detection unit, respectively, and the detection of multiple detection objects and parts thereof or different types of objects are carried out, and image recognition is performed using the neural network model constructed by deep learning.

Benefits of technology

It improves detection accuracy, reduces the possibility of false detection and missed detection, and can more accurately identify multiple detection objects, especially when detecting improper objects in waste incineration facilities, reducing misjudgment and misjudgment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112651281B_ABST
    Figure CN112651281B_ABST
Patent Text Reader

Abstract

Technical problem: To improve the detection accuracy of detection using a machine learning model that has been trained. Solution: The information processing apparatus (1) includes: a first detection unit (101) that inputs input data to a first learned model capable of detecting multiple detection targets and detects the detection targets; and a second detection unit (102) that inputs the input data to a second learned model capable of detecting at least a part of the multiple detection targets and detects a second detection target, and the information processing apparatus (1) determines a final detection result based on the detection results of the two detection units.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an information processing device and the like that uses a learned model constructed by machine learning to detect an object to be detected. Background Art

[0002] In recent years, due to the development of machine learning such as deep learning, the recognition / detection accuracy of objects in images has been improving. Moreover, as a result, the uses of image recognition technology are expanding. However, since the current detection accuracy is not 100%, further improvement is required to further expand the uses.

[0003] For example, the image recognition device described in Patent Document 1 below performs pattern matching on an image of an object using a plurality of templates. Moreover, when it is determined that the patterns match a plurality of templates and the degree of overlap between the templates is equal to or greater than a threshold value, it is recognized as at least one recognition object related to the templates. Thus, compared with the case of performing pattern matching using one template, the recognition accuracy can be improved.

[0004] Prior Art Documents

[0005] Patent Documents

[0006] Patent Document 1: Japanese Patent Application Laid-Open Gazette "Japanese Patent Application Laid-Open No. 2008-165394 (published on July 17, 2008)" Summary of the Invention

[0007] (1) Technical Problem to be Solved

[0008] However, in the prior art as described above, there is room for improving the detection accuracy. For example, in the case of using a template for pedestrian detection and a template for sign detection, when a pedestrian is missed in the template for pedestrian detection, since a pedestrian is not detected in the template for sign detection, there is no method for remedying the missed detection of the pedestrian. In addition, when a false detection occurs in the template for pedestrians (for example, when a roadside tree is misidentified as a pedestrian), there is also no method for remedying the false detection. In addition, when the object to be detected is not an object (for example, when sound data is input to a learned model to detect a specified sound component), even when using a plurality of templates such as those in Patent Document 1, there is still room for improving the detection accuracy.

[0009] An object of one aspect of the present invention is to realize an information processing device and the like that can improve the detection accuracy of detection using a machine learning model that has been performed.

[0010] (2) Technical Solution

[0011] To solve the above technical problems, an information processing apparatus according to one aspect of the present invention includes: a first detection unit that inputs input data to a first learned model that has been machine-learned in a manner capable of detecting a plurality of first detection objects and detects the first detection objects; and a second detection unit that inputs the input data to a second learned model that has been machine-learned in a manner capable of detecting at least a part of the plurality of first detection objects, i.e., second detection objects, and detects the second detection objects, or inputs the input data to a third learned model that has been machine-learned in a manner capable of detecting third detection objects different from the first detection objects and detects the third detection objects, and the information processing apparatus determines a final detection result based on the detection result of the first detection unit and the detection result of the second detection unit.

[0012] To solve the above technical problems, an information processing method according to one aspect of the present invention is executed by one or more information processing apparatuses, and includes: a first detection step of inputting input data to a first learned model that has been machine-learned in a manner capable of detecting a plurality of detection objects and detecting the detection objects based on the input data; a second detection step of inputting the input data to a second learned model that has been machine-learned in a manner capable of detecting at least a part of the plurality of detection objects, or inputting the input data to a third learned model that has been machine-learned in a manner capable of detecting detection objects different from the first learned model and detecting the detection objects based on the input data, and the information processing method further includes: a determination step of determining a final detection result based on the detection result of the first detection step and the detection result of the second detection step.

[0013] (III) Advantageous Effects

[0014] According to one aspect of the present invention, it is possible to improve the detection accuracy of detection using a machine-learned model. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 FIG. is an example of a functional block diagram of a control unit of the information processing apparatus according to Embodiment 1 of the present invention.

[0016] Figure 2 FIG. is a block diagram showing a structural example of an improper object detection system including the above information processing apparatus.

[0017] Figure 3 FIG. is a view showing a situation where a garbage collection vehicle drops garbage into a garbage pit in a garbage incineration facility.

[0018] Figure 4 FIG. is a view showing inside the garbage pit.

[0019] Figure 5It is a diagram showing the process of the processing executed by the above information processing apparatus.

[0020] Figure 6 It is a diagram showing the construction and re-learning of the learned model.

[0021] Figure 7 It is a block diagram showing a structural example of the control unit included in the information processing apparatus according to Embodiment 2 of the present invention.

[0022] Figure 8 It is a diagram showing the process of the processing executed by the above information processing apparatus.

[0023] Figure 9 It is a block diagram showing a structural example of the control unit included in the information processing apparatus according to Embodiment 3 of the present invention.

[0024] Figure 10 It is a diagram showing the process of the processing executed by the above information processing apparatus.

[0025] Explanation of reference numerals

[0026] 1 - Information processing apparatus; 101, 201, 302 - First detection unit; 102, 202A, 202B, 303 - Second detection unit; 304 - Detection result integration unit. Detailed description of the embodiments

[0027] 〔Embodiment 1〕

[0028] In recent years, there has been a problem of throwing improper waste (hereinafter simply referred to as improper waste) into waste incineration facilities. Due to the throwing of improper waste into the incinerator, the combustion in the incinerator deteriorates, and blockages occur at the ash removal equipment of the incinerator. In some cases, the incinerator may be stopped urgently. Currently, the staff of waste incineration facilities randomly select the collected waste and manually confirm whether the selected waste contains improper waste, which places a heavy burden on the operators.

[0029] In addition, in order to reduce the improper waste transported to waste incineration facilities, a system is required that detects improper waste from the transported waste and prompts the detected improper waste to the collection person in a situation where the person in charge of collecting waste needs to be alerted. In this case, it is not desirable to prompt waste that is not actually improper waste as improper waste. In addition, when directly showing the captured image to the person in charge, it is not preferable because it is difficult to grasp at which moment and at which position the improper waste was photographed.

[0030] An information processing apparatus 1 according to an embodiment of the present invention can solve the above problems. The information processing apparatus 1 has a function of detecting inappropriate objects from the garbage transported into the waste incineration facility. Specifically, the information processing apparatus 1 detects inappropriate objects by using an image of the garbage being thrown into the garbage pit. In addition, regarding the garbage pit, it will be described based on Figure 4 which will be described below. In addition, inappropriate objects can be detected after the garbage is thrown. In addition, the so-called inappropriate objects are objects that should not be incinerated in the incinerator provided in the waste incineration facility. Specific examples of inappropriate objects will be described below.

[0031] 〔System Structure〕

[0032] Based on Figure 2 The structure of the inappropriate object detection system of this embodiment will be described. Figure 2 It is a block diagram showing a structural example of the inappropriate object detection system 100. The inappropriate object detection system 100 includes an information processing apparatus 1, a garbage imaging apparatus 2, a vehicle information collection apparatus 3, a selection display apparatus 4, and an inappropriate object display apparatus 5.

[0033] In addition, in Figure 2 an example of the hardware structure of the information processing apparatus 1 is also shown. As shown in the figure, the information processing apparatus 1 includes: a control unit 10, a high-speed storage unit 11, a large-capacity storage unit 12, an image IF (interface) unit 13, a vehicle information IF unit 14, a selection display IF unit 15, and an inappropriate object display IF unit 16. The information processing apparatus 1 can be, for example, a personal computer, a server, or a workstation as an example.

[0034] The control unit 10 overall controls each part of the information processing apparatus 1. Based on Figure 1 The functions of each part of the control unit 10 described below can be implemented by a logic circuit (hardware) formed in an integrated circuit (IC chip) or the like, or can be implemented by software. In this software, it may also include an information processing program that causes a computer to function as each part of the control unit 10 included in the Figure 1 , 7 , 9 described below. In the case of implementation by software, the control unit 10 may be constituted by, for example, a CPU (Central Processing Unit), or may be constituted by a GPU (Graphics Processing Unit), or may be constituted by a combination thereof. In addition, in this case, the above software is pre-stored in the large-capacity storage unit 12. Moreover, the control unit 10 reads the above software into the high-speed storage unit 11 and executes it.

[0035] Both the high-speed storage unit 11 and the large-capacity storage unit 12 are storage devices that store various data used by the information processing device 1. The high-speed storage unit 11 is a storage device that can write and read data at high speed compared to the large-capacity storage unit 12. The large-capacity storage unit 12 has a larger data storage capacity compared to the high-speed storage unit 11. As the high-speed storage unit 11, it is also possible to apply, for example, a high-speed access memory such as SDRAM (Synchronous Dynamic Random-Access Memory). In addition, as the large-capacity storage unit 12, it is also possible to apply, for example, an HDD (Hard Disk Drive), an SSD (Solid-State Drive), an SD (Secure Digital) card, or an eMMC (embedded Multi-Media Controller).

[0036] The image IF unit 13 is an interface for communicatively connecting the trash shooting device 2 to the information processing device 1. In addition, the vehicle information IF unit 14 is an interface for communicatively connecting the vehicle information collection device 3 to the information processing device 1. These IF units can be either IF units for wired communication or IF units for wireless communication. For example, as these IF units, it is also possible to apply USB (Universal Serial Bus), LAN (Local-Area Network), wireless LAN, etc.

[0037] The selection display IF unit 15 is an interface for communicatively connecting the selection display device 4 to the information processing device 1. In addition, the inappropriate object display IF unit 16 is an interface for communicatively connecting the inappropriate object display device 5 to the information processing device 1. These IF units can also be either IF units for wired communication or IF units for wireless communication. For example, as these IF units, it is also possible to apply HDMI (High-Definition Multimedia Interface, registered trademark), DisplayPort, DVI (Digital Visual Interface), VGA (Video Graphics Array) terminals, S terminals, or RCA terminals, etc.

[0038] The garbage photographing device 2 photographs the garbage on its way to being dropped into the garbage pit and sends the photographed image to the information processing device 1. Hereinafter, this photographed image will be referred to as a garbage image. As an example, the garbage photographing device 2 may also be a high-speed shutter camera that photographs a moving image. In addition, the garbage image may be a moving image or a series of still images taken continuously. The garbage image is input to the information processing device 1 via the image IF unit 13. Moreover, the input garbage image may be directly processed by the control unit 10 or may be stored in the high-speed storage unit 11 or the large-capacity storage unit 12 and then processed by the control unit 10.

[0039] The vehicle information collection device 3 transports the garbage and collects the identification information of the vehicle (so-called garbage collection vehicle) that drops the garbage into the garbage pit and sends it to the information processing device 1. In addition, hereinafter, based on Figure 4 The process of dropping garbage into the garbage pit using a garbage collection vehicle will be described. This identification information is used by the incoming vehicle determination unit 105 to determine the incoming entity of the garbage. The above-mentioned identification information may be, for example, information such as a license plate number. In this case, the vehicle information collection device 3 may photograph the license plate and send the photographed image to the information processing device 1 as the identification information. In addition, the vehicle information collection device 3 may accept the input of the identification information of the garbage collection vehicle and send it to the information processing device 1.

[0040] The selection display device 4 displays the image of the improper object detected by the information processing device 1. In the improper object detection system 100, considering that the information processing device 1 may misjudge garbage that is not an improper object as an improper object, the selection display device 4 displays the image of the improper object detected by the information processing device 1, and visually confirms whether the object presented in the image is an improper object. Then, the person in charge of visual confirmation selects the image presenting the improper object from the images displayed on the selection display device 4.

[0041] The improper object display device 5 displays the image of the improper object selected via the selection display device 4 from the images of the improper objects detected by the information processing device 1, that is, the image visually confirmed to present the improper object. The improper object display device 5 displays the above image in order to attract the attention of the person in charge of transporting the above improper object, the operator, etc.

[0042] 〔Photographing of Garbage Image〕

[0043] Figure 3 It is a diagram showing the situation where the garbage collection vehicle 200 drops garbage into the garbage pit in the garbage incineration facility. Figure 4 It is a diagram showing the inside of the garbage pit. The garbage pit is a place for temporarily storing the garbage collected in the garbage incineration facility, and the garbage in the garbage pit is successively sent into the incinerator for incineration. As Figure 3As shown, a plurality of doors such as doors 300A and 300B are provided on the waste incineration facility (hereinafter, when it is not necessary to distinguish each door, they are collectively referred to as door 300). In addition, as Figure 4 shown, a waste pit is provided in front of the door 300. That is, by opening the door 300, a dropping port for dropping waste into the waste pit appears. As Figure 3 shown, the waste collection vehicle 200 drops waste into the waste pit from the dropping port.

[0044] The waste photographing device 2 is installed at a position where it can photograph the waste flowing along the Figure 4 inclined surface 600. For example, the waste photographing device 2 can be installed at the installation position 400 as Figure 3 and Figure 4 shown. Since the installation position 400 is located on the surface of each door 300, when the waste photographing device 2 is installed at the installation position 400, the waste photographing device 2 is located above the inclined surface 600 when the door 300 is opened, and this position is suitable for photographing the waste. Of course, the installation position of the waste photographing device 2 can be set to any position where it can photograph the waste flowing along the inclined surface 600.

[0045] In addition, when the vehicle information collection device 3 is a photographing device, the vehicle information collection device 3 can also be installed at the installation position 400. When the waste collection vehicle 200 approaches the door 300, since the door 300 is closed, the license plate of the waste collection vehicle 200 can be photographed from the vehicle information collection device 3 installed at the installation position 400. Of course, the installation position of the vehicle information collection device 3 can be set to any position where it can photograph the waste collection vehicle 200, and it can also be installed at a position different from the waste photographing device 2. In addition, the vehicle information collection device 3 can be, for example, an information input device. In this case, the structure can be such that the vehicle information collection device 3 is installed in the operation room and accepts the input of the identification information of the waste collection vehicle 200 by the operator.

[0046] 〔Device Structure〕

[0047] Based on Figure 1 the structure of the information processing device 1 will be described. Figure 1 is an example of the functional block diagram of the control unit 10 of the information processing device 1. In the Figure 1 shown control unit 10, it includes a first detection unit 101, a second detection unit 102, a learning unit 103, a selection display control unit 104, an incoming vehicle determination unit 105, and an improper object display control unit 106. In addition, in the Figure 1 shown large-capacity storage unit 12, it includes an input data storage unit 121, a detection result storage unit 122, a learned model storage unit 123, and a teaching data storage unit 124.

[0048] The first detection unit 101 inputs input data into a first learned model that has been machine-learned in such a way as to be able to detect a plurality of first detection objects, and detects the first detection objects. In addition, detecting a plurality of detection objects means that there are a plurality of classifications (also referred to as classes) learned in the first learned model.

[0049] In addition, the second detection unit 102 inputs the input data into a second learned model that has been machine-learned in such a way as to be able to detect at least a part of the plurality of first detection objects, that is, second detection objects, and detects the second detection objects. In the present embodiment, an example will be described in which the plurality of first detection objects are various defective items, and the second detection objects are also defective items. In addition, among the first detection objects and the second detection objects, there may also be included objects that, although similar in appearance to defective items, are not defective items.

[0050] The first learned model and the second learned model are read out from the learned model storage unit 123, and the input data is read out from the input data storage unit 121. As will be described in detail below, the first learned model and the second learned model are constructed by machine learning that uses the original teaching data 124a stored in the teaching data storage unit 124. In addition, the first detection objects and the second detection results are stored in the detection result storage unit 122. These detection results are used as additional teaching data 122a. The additional teaching data 124b is obtained by copying the additional teaching data 122a to the teaching data storage unit 124. The relearning of the first learned model and the second learned model is performed using the original teaching data 124a and the additional teaching data 124b stored in the teaching data storage unit 124.

[0051] The first learned model and the second learned model may be any models constructed by machine learning. In the present embodiment, an example will be described in which the first learned model and the second learned model are learned models of neural networks constructed using deep learning. More specifically, these learned models take an image as input data and output object information of the detection object presented in the image. The object information may include an identifier indicating the classification of the object, information indicating the position, size, shape, etc. In addition, in the object information, there may also be included a probability value indicating the certainty of the detection result. This probability value may be, for example, a numerical value from 0 to 1.

[0052] In addition, the first detection unit 101 and the second detection unit 102 can perform object detection based on the above probability values. In this case, a detection threshold can be preset in advance, and an object with a probability value larger than the threshold among the above probability values is used as the detected object. Although the larger the detection threshold, the higher the detection accuracy, the number of missed detections increases. Although the smaller the detection threshold, the fewer the missed detections, the false detections increase. Therefore, it is only necessary to set an appropriate detection threshold according to the required detection accuracy, etc. In addition, regarding the construction of the learned model, the following will be based on Figure 6 for explanation.

[0053] The learning unit 103 performs relearning of the first learned model and the second learned model. In addition, the structure may also be that the learning unit 103 also constructs the first learned model and the second learned model. Regarding the construction and relearning of the learned model, the following will be based on Figure 6 for explanation. In addition, the learning unit 103 for the first learned model and the learning unit 103 for the second learned model may be separately provided.

[0054] The selection display control unit 104 causes the selection display device 4 to display the detection result (for example, an image determined to show an inappropriate object) determined based on the detection results of the first detection unit 101 and the second detection unit 102. The person in charge of visual confirmation confirms whether an inappropriate object is presented in the displayed image and selects the image showing the inappropriate object. Moreover, the selection display control unit 104 accepts the image selection performed by the person in charge of visual confirmation. Thereby, false detections can be basically reliably avoided.

[0055] The incoming vehicle determination unit 105 determines the incoming vehicle of the garbage (for example Figure 3 the garbage collection vehicle 200) using the identification information received from the vehicle information collection device 3. Moreover, when the inappropriate object display control unit 106 detects an inappropriate object in the garbage previously brought in by the incoming vehicle determined by the information processing device 1 from the incoming vehicle determination unit 105, the inappropriate object display device 5 is caused to display an image of the above inappropriate object. Thereby, it is possible to draw the attention of the person in charge who brought in the garbage by this incoming vehicle by presenting an image of the inappropriate object.

[0056] As described above, the information processing apparatus 1 includes: a first detection unit 101 that inputs input data to a first learned model that has been machine-learned to be able to detect a plurality of first detection objects and detects the first detection objects; and a second detection unit 102 that inputs the input data to a second learned model that has been machine-learned to be able to detect a second detection object that is at least a part of the plurality of first detection objects and detects the second detection object. Further, the information processing apparatus 1 determines a final detection result based on the detection results of the first detection unit 101 and the second detection unit 102. Specifically, in the information processing apparatus 1, the first detection unit 101 and the second detection unit 102 output detection results to a common output destination, and thus the detection result output to the common output destination is set as the final detection result.

[0057] According to the above structure, since at least a part of the detection objects of the first learned model and the second learned model overlap, the possibility of false detection of the overlapping part can be reduced. For example, the appearance of a board that is an improper object and corrugated paper that is not an improper object is similar. In this case, it is possible to falsely detect the board as corrugated paper, or falsely detect the corrugated paper as a board. According to the above structure, even if false detection such as the above board and corrugated paper occurs in one of the first learned model and the second learned model, as long as the other can accurately detect the object, the correct detection result of the object can finally be output. Therefore, the detection accuracy of the detection using the machine-learned model can be improved.

[0058] In addition, according to the above structure, since the first detection unit 101 uses the first learned model that has been machine-learned to be able to detect a plurality of first detection objects, a plurality of first detection objects can be effectively detected together.

[0059] Furthermore, the method for determining the final detection result is not limited to the method of making the output destinations of the detection results of the first detection unit 101 and the second detection unit 102 common. For example, a block for synthesizing detection results can be added to the control unit 10, and the detection results of the first detection unit 101 and the second detection unit 102 can be synthesized using this block and used as the final detection result. As a synthesis method, the following methods can be cited, for example.

[0060] (1) The detection result of the second detection unit 102 is added to the detection result of the first detection unit 101 as the final detection result (for example, when the detection unit 101 detects improper objects A and B from a certain image, and the second detection unit 102 detects improper objects B and C from the same image, the final detection result from this image is set as improper objects A, B, C, etc.).

[0061] (2) Take the common part of the detection results of the first detection unit 101 and the second detection unit 102 as the final detection result (for example, when the first detection unit 101 detects improper objects A, B, and C from a certain image, and the second detection unit 102 detects improper object B from the same image, set the final detection result from this image as improper object B, etc.).

[0062] (3) Add the detection result of the second detection unit 102 to the detection result of the first detection unit 101, but exclude the part where the two detection results do not match from the final detection result (for example, regarding an image presenting object X and object Y, when the first detection unit 101 detects that object X is improper object A and object Y is improper object B, and the second detection unit 102 detects that object X is improper object A and object Y is improper object C, set the final detection result as improper object A, etc.).

[0063] 〔Processing flow〕

[0064] Figure 5 It is a diagram showing the processing flow executed by the information processing apparatus 1 of the present embodiment. The processing executed by the information processing apparatus 1 of the present embodiment and its execution order can be defined by, for example Figure 5 the setting file F1 shown, and can also be represented by the flowchart of this diagram.

[0065] (Regarding the setting file)

[0066] Figure 5 The setting file F1 shown is a segmented data structure. One segment starts from a segment name. In Figure 5 the example, the string enclosed by "[" and "]" is the segment name. Specifically, [EX1] and [EX2] are segment names. One segment ends with the start of the next segment or the end of the setting file. The content to be executed at each stage is defined in one segment.

[0067] The execution of the segment is set to be in the segment order in this example, but it is not limited to this example. For example, the execution order can also be defined in a part of the segment name. In addition, different algorithms can be executed for each segment. For example, an algorithm that processes using neural networks with different numbers of intermediate layers can be executed for each segment, or an algorithm that performs different internal processing can be executed. Furthermore, the above definitions can be made by other methods (such as arguments when executing Figure 5 the flowchart) instead of the setting file.

[0068] In the setting file F1, for each variable as follows, namely <key>Define it. <key>The value is <value> 。

[0069] <key> = <value>

[0070] In this example <key>In this case, three items, namely "script", "src", and "dst", are defined for each segment. Among them, "script" represents the file name of the script to be executed. The script file is, for example, an algorithm for detecting improper items. In this algorithm, processing other than improper item detection may be included.

[0071] In addition, "src" represents the input data used to execute the script. For example, in the case of a script for object detection on an image, "src" may represent the storage location of the image to be processed in the mass storage unit 12, or a list of images to be processed. Moreover, "dst" represents the output destination of the script execution result. For example, "dst" may represent a location on the mass storage unit 12. In addition, other parameters used when executing the script can be defined in, for example, "script".

[0072] The control unit 10 reads the setting file F1 and executes the script file named "ex1" defined in [EX1] of the setting file F1, whereby the control unit 10 functions as the first detection unit 101. Moreover, the first detection unit 101 determines the image to be processed with reference to "ex_src", performs object detection on the determined image, and records the result in "ex_dst". Next, the control unit 10 executes the script file named "ex2" defined in [EX2], whereby the control unit 10 functions as the second detection unit 102. Moreover, the second detection unit 102 determines the image to be processed with reference to "ex_src", performs object detection on the determined image, and records the result in "ex_dst".

[0073] "ex1" in [EX1] of the setting file F1 is a script for performing object detection on an image as input data using the above-mentioned first learned model. In addition, "ex2" in [EX2] is a script for performing object detection on an image as input data using the above-mentioned second learned model.

[0074] In addition, the values of "src" and "dst" in [EX1] and [EX2] in the setting file F1 are common. That is to say, the images to be processed by the first detection unit 101 and the second detection unit 102 are common, and the output destinations of the object detection results for this image are also common.

[0075] As [EX1] in the setting file F1, for example, the following can also be applied Figure 5 The script SC11 shown (script file name: ex1). This script file is an example of a shell script in the UNIX (registered trademark) form including Linux (registered trademark), but it can also be in other forms such as the Batch form of DOS (Disk Operating System), for example.

[0076] When using this script file, as Figure 5 shown, (1) based on the set file items, (2) execute the script. Here, ex1 is the script file name executed as described above, and ex_src and ex_dst are passed to the script as arguments. The ex1 script (SC11) has a one-line structure. Here, the execution file (first detection instruction) of the first detection unit 101 is executed. In this example, this instruction uses three arguments. The argument ex_src of ex1 is passed to the first argument $1, the argument ex_dst of ex1 is passed to the second argument $2, and the third argument is the file name of the learned model (first learned model). As the first learned model, for example, a learned model that has performed machine learning on all detection target objects is used. The detection target objects can be, for example, corrugated paper, boards, wood, mats, and long objects. Among these detection target objects, boards, wood, mats, and long objects are inappropriate objects. Corrugated paper is not an inappropriate object, but its appearance is similar to that of a board, so it is included in the detection target objects. In this way, by including garbage with an appearance similar to that of inappropriate objects in the detection target objects, the possibility of missed detection and false detection of inappropriate objects can be reduced.

[0077] In addition, the execution file in this example uses three arguments, but other setting items (such as detection thresholds) can be used as arguments as needed. Also, the order of the arguments can be changed as needed. In addition, although SC11 has a one-line structure, instructions before and after the processing performed by the first detection unit 101 (such as pre-processing and post-processing instructions) can also be added as needed. For example, an instruction for organizing input data, an instruction for extracting necessary information from the executed log data, etc. can be added.

[0078] When using the script SC11, for example, the file of the moving image (hereinafter referred to as the moving image file) or multiple still image files captured by the garbage photographing device 2 can be specified by src = ex_src. As a result, the detection results (such as the image of the detected object and the position and size information of the object, etc.) of this moving image file are saved in "ex_dst".

[0079] [EX2] in the setting file F1 can apply the script SC12. The difference between SC11 and SC12 lies in the execution file name and the learned model used. Additionally, [EX2] can set, among the detection target objects of [EX1], the objects for which false detection is particularly desired to be avoided as detection target objects. For example, when it is desired to avoid false detection of wood, [EX2] can set wood as the detection target object. Thus, even if wood is falsely detected in [EX1], as long as wood can be detected by [EX2], overall, false detection of wood will not occur. Additionally, when [EX1] is set to detect a part of all detection target objects, [EX2] can be set to detect another part of all detection target objects. In this case, at least one of the detection target objects of [EX1] and the detection target objects of [EX2] can also be made to overlap in advance. Furthermore, although different execution files are used in SC11 and SC12, when only the learned models are different, the first detection instruction and the second detection instruction can be the same. That is, the first detection unit 101 and the second detection unit 102 can be the same.

[0080] (Regarding the flowchart)

[0081] For Figure 5 the processing (information processing method) of the flowchart shown is described. This flowchart represents the processing flow of the setting file F1 shown in this figure. The moving image file captured by the garbage shooting device 2 before the processing of this flowchart is stored in "ex_src" of the input data storage unit 121. Additionally, instead of the moving image file, a plurality of frame images extracted from this moving image file or a plurality of still image files sequentially captured by the garbage shooting device 2 can be stored.

[0082] In S11 (the first detection step), the first detection unit 101 performs object detection on the images stored in the input data storage unit 121 using the first learned model. Specifically, the first detection unit 101 sets the frame image extracted from the moving image file stored in "ex_src" of the input data storage unit 121 as the input data, inputs this frame image to the learned model with the script name "ex1" and makes it output object information. Moreover, the first detection unit 101 determines whether an object is detected based on the object information. When an object is detected, the frame image is associated with the object information as the detection result and recorded in "ex_dst" of the detection result storage unit 122. These processes are respectively performed on the frame images extracted from the moving image file stored in "ex_src".

[0083] In S12 (second detection step, determination step), the second detection unit 102 performs object detection on the same frame image as in S11 using the second learned model. In this way, by making the input data input to the first learned model and the second learned model the same data, it is possible to suppress the occurrence of missed detections. This is because, even if a missed detection of an improper object occurs in the detection by one learned model, as long as the improper object can be detected in the detection by the other, no missed detection will occur overall.

[0084] In S12, specifically, the second detection unit 102 uses the frame image extracted from the moving image file of "ex_src" stored in the input data storage unit 121 as the input data. In addition, the second detection unit 102 inputs the frame image to the learned model with the script name "ex2" and makes it output object information. Moreover, the second detection unit 102 determines whether an object is detected based on the object information. If an object is detected, the frame image is associated with the object information as the detection result. The second detection unit 102 records this detection result in "ex_dst" of the detection result storage unit 122, which is the same output destination as the detection result of the first detection unit 101. These processes are performed for each frame image extracted from the moving image file stored in "ex_src". At the end of the processing in S12, the data recorded in "ex_dst" of the detection result storage unit 122 is the final detection result. That is, through the processing in S12, the final detection result is determined.

[0085] In S13, the detection result is output, and thus the processing ends. For example, the selection display control unit 104 can output the detection result. In this case, the selection display control unit 104 can cause the selection display device 4 to display the frame image and object information recorded in "ex_dst" through the processing in S11 and S12. Thereby, the user of the selection display control unit 104 can visually confirm whether the detection result of the information processing device 1 is correct and can input the confirmation result to the information processing device 1. In addition, the selection display control unit 104 can determine the image presenting an improper object based on the input confirmation result. In addition, the selection display control unit 104 can correct the object information based on the visual confirmation result. Moreover, the image presenting an improper object confirmed visually is displayed on the improper object display device 5 by the improper object display control unit 106 when, for example, the transporter of the improper object transports garbage again.

[0086] In addition, the execution order of the processing in S11 and the processing in S12 is not limited to Figure 5 the example. The processing in S11 can be executed after the processing in S12, or these processes can be performed in parallel. In any case, the final detection result is determined at the end of both the processing in S11 and the processing in S12. In addition, in Figure 5 In the example, two learned models are used, but more than three learned models can also be used. In this case, it is only necessary to append a segment corresponding to the learned models after the third one in the setting file F1.

[0087] 〔Regarding the resolution of input data〕

[0088] In the case where the input data for the learned model is image data as in the present embodiment, the first detection unit 101 can reduce the resolution of the image data and input it into the first learned model. Alternatively, the second detection unit 102 can reduce the resolution of the image data and input it into the second learned model. By reducing the resolution of the input image data, the amount of computation for the object detection process using the learned model is reduced, and the required time can be shortened.

[0089] In addition, regarding whether to reduce the resolution of the input data for which learned model, it only needs to be determined in advance according to the detection object of each learned model. For example, it is easy to detect large-sized objects such as wood and boards using low-resolution image data, but it is preferable to use high-resolution image data for detecting small-sized objects such as cans. Therefore, for the learned model among the used learned models with a large detection object, the result of reducing the resolution of the captured garbage image can be used as the input data. Thereby, for large-sized objects, high-speed detection processing can be achieved without reducing the detection accuracy. In addition, the learned model using low-resolution image data can pre-construct the same low-resolution image data as the input data as teaching data. Additionally, the structure can be such that after object detection using low-resolution image data as the input data, the resolution of the original image data is returned for output. The structure can be such that the first detection unit 101 or the second detection unit 102 performs the resolution change process, or a resolution change module can be additionally appended to the control unit 10.

[0090] In addition, the number of intermediate layers of each learned model can be different according to the detection object. For example, the number of intermediate layers of the learned model for detecting large-sized objects can be less than the number of intermediate layers of the learned model for detecting smaller-sized objects. In such a structure, similarly to the above example, high-speed detection processing can be achieved for large-sized objects without reducing the detection accuracy.

[0091] 〔Construction and re-learning of learned models〕

[0092] Based on Figure 6 The construction and re-learning of the above-mentioned first learned model and second learned model will be described. Figure 6 This is a diagram illustrating the construction and relearning of a learned model. Here, an example of constructing a learned model of a neural network will be described. When using a neural network, multiple intermediate layers can be provided, and in this case, machine learning is deep learning. Of course, the number of intermediate layers can be one, and machine learning algorithms other than neural networks can also be applied.

[0093] As shown in the figure, a learned model is constructed through initial learning. Moreover, object detection is performed using the learned model constructed through initial learning, and relearning is performed using the object detection results, and the learned model is updated.

[0094] In initial learning, an image presenting a detection object is used as a teaching image, and object information of the detection object presented in the teaching image (for example, an identifier indicating the classification of the object, information indicating position, size, shape, etc.) is used as correct data. It is assumed that the teaching data is stored in Figure 1 the teaching data storage unit 124. The teaching data used in initial learning is only the original teaching data 124a. In machine learning, the learning unit 103 inputs the teaching image into the neural network, and while changing the teaching image, repeats the process of updating the weight values so that the output value of the neural network approaches the correct data.

[0095] In machine learning, basically, the more the number of repetitions, the closer the weight value is to the optimal value. However, due to overlearning or other reasons, the weight value may deviate from the optimal value after repetition. In addition, when constructing a learned model for detecting multiple detection objects, when a certain weight value is applied, the detection accuracy of a certain detection object is high, but the detection accuracy of other detection objects may become low.

[0096] Therefore, in Figure 6 the example, multiple learned models 1 to I with different weight values are generated. Moreover, a model to be applied as the first learned model and a model to be applied as the second learned model are selected from these learned models. For example, an image presenting a detection object can be input into each of the learned models 1 to I as test data, the detection accuracy of each detection object can be calculated based on the output value, and the above selection can be made based on the calculated detection accuracy. In addition, these learned models are stored in the learned model storage unit 123.

[0097] In addition, in this selection, it is preferable not to generate an object to be detected with low detection accuracies for both the first learned model and the second learned model. For example, when the detection accuracy of a long object in the first learned model is low, it is preferable to use, as the second learned model, a model with a high detection accuracy for long objects. In addition, it is not necessary to select both the first learned model and the second learned model from the learned models 1 to I. For example, when selecting the first learned model from the learned models 1 to I, the second learned model can be selected from a plurality of separately constructed learned models.

[0098] In addition, the selection of the learned model can be made manually or by the information processing apparatus 1. In the latter case, criteria for selecting the learned model are set in advance, and it is only necessary to input to the information processing apparatus 1 or have it calculate the information required to determine whether the selection criteria are satisfied (for example, information indicating the detection accuracy of each object to be detected in each learned model).

[0099] By performing object detection on the detection target image using the first learned model and the second learned model thus selected, as described based on Figure 5 the image of the object and the object information are recorded in the detection result storage unit 122. The learning unit 103 uses this image as a teaching image and performs re-learning using the teaching data obtained by adding the object information of this image as correct data and the original teaching data 124a for initial learning. In addition, preferably, the image and the object information used as the teaching data are images and object information that have been visually confirmed to be correct via the selection display device 4. The teaching data selected here is set as additional teaching data 122a and can be copied as additional teaching data 124b to the teaching data storage unit 124 for re-learning. In addition, here, the additional teaching data 124b is a copy of the additional teaching data 122a, but it is also possible to set 124b to be the same as 122a without copying.

[0100] In addition, although the image and the object information are recorded in the detection result storage unit 122, when the input data is a still image, it is not necessary to save the image, and only the image file name of the input data can be recorded.

[0101] In re-learning, the learning unit 103 learns using the original teaching data 124a and the additional teaching data 124b stored in the teaching data storage unit 124, and constructs a plurality of re-learned models 1 to J with different weight values. Moreover, a first learned model and a second learned model are selected from among them. The difference from the initial learning is that the additional teaching data 124b is added to the teaching data for machine learning. By adding the additional teaching data, an improvement in the detection accuracy of the learned model can be expected. In addition, in re-learning, the learning unit 103 does not need to use all the object information recorded as the result of object detection, and can select a part of the object information for use. In addition, the learning unit 103 can repeat re-learning multiple times. In addition, each learned model after Embodiment 2 can be constructed in the same manner as described above, and re-learning can also be performed.

[0102] 〔Embodiment 2〕

[0103] Another embodiment of the present invention will be described below. In addition, for convenience of explanation, components having the same functions as those described in the above embodiment are denoted by the same reference numerals, and their descriptions are not repeated. The same applies to Embodiment 3 described below.

[0104] 〔Structural example of the control unit〕

[0105] Based on Figure 7 A structural example of the control unit 10 of the information processing apparatus 1 according to the present embodiment will be described. Figure 7 is a block diagram showing a structural example of the control unit 10 included in the information processing apparatus 1 according to Embodiment 2. In addition, the mass storage unit 12 is also shown in Figure 7 .

[0106] As Figure 7 shown, the control unit 10 includes a first detection unit 201, a second detection unit 202A, and a second detection unit 202B. In addition, the learning unit 103, the selection display control unit 104, the incoming vehicle determination unit 105, and the improper object display control unit 106 are the same as those in Embodiment 1, and thus their illustrations are omitted.

[0107] The first detection unit 201 inputs input data to a first learned model that has been machine-learned to be able to detect a plurality of first detection objects and detects the first detection objects. The first detection unit 201 has the same function as the first detection unit 101 in Embodiment 1, but is different from the first detection unit 101 in terms of determining the input data used by the second detection unit 202A and the second detection unit 202B based on the detection result of the first detection unit 201.

[0108] The second detection unit 202A inputs the above input data (the detection result of the first detection unit 201) into the second learned model A that has been machine-learned to be able to detect at least a part of a plurality of first detection objects, i.e., the second detection object A, and detects the second detection object A. For the second detection object A, the detection result of the second detection unit 202A is the final detection result. The second detection unit 202A uses, as the input data for the second learned model A, the input data in the input data of the first detection unit 201 in which the first detection unit 201 has detected the detection object, which is different from the second detection unit 102 in Embodiment 1 in this regard.

[0109] Similarly to the second detection unit 202A, the second detection unit 202B inputs the input data in which the first detection unit 201 has detected the detection object into the second learned model B and detects at least a part of a plurality of first detection objects, i.e., the second detection object B, based on the input data (the detection result of the first detection unit 201). For the second detection object B, the detection result of the second detection unit 202B is the final detection result. In addition, the second detection object A and the second detection object B are different objects.

[0110] As described above, the information processing apparatus 1 includes: a first detection unit 201 that inputs input data into a first learned model that has been machine-learned to be able to detect a plurality of first detection objects and detects the first detection objects; and a second detection unit 202A that inputs the input data into a second learned model that has been machine-learned to be able to detect at least a part of the plurality of first detection objects, i.e., the second detection object, to detect the second detection object. Moreover, the information processing apparatus 1 determines the final detection result based on the detection result of the first detection unit 201 and the detection result of the second detection unit 202A. Specifically, in the information processing apparatus 1 of the present embodiment, the second detection unit 202A inputs the input data in which the first detection unit 201 has detected the first detection object among the above plurality of input data into the second learned model and detects the second detection object A, and for the second detection object A, sets the detection result of the second detection unit 202A as the final detection result. In addition, the second detection unit 202B inputs the input data in which the first detection unit 201 has detected the first detection object among the above plurality of input data into the second learned model and detects the second detection object B, and for the second detection object B, sets the detection result of the second detection unit 202B as the final detection result.

[0111] According to the above structure, since at least a part of the detection objects of the first learned model and the second learned model A overlap, the possibility of false detection of this overlapping part can be reduced. Similarly, since at least a part of the detection objects of the first learned model and the second learned model B also overlap, the possibility of false detection of this overlapping part can be reduced. Therefore, the detection accuracy of the detection using the machine-learned model can be improved. In addition, according to the above structure, since the first detection unit 201 uses the first learned model that has been machine-learned in such a way as to be able to detect a plurality of first detection objects, a plurality of first detection objects can be effectively detected together.

[0112] In addition, the method for determining the final detection result is not limited to the above method. For example, as described in Embodiment 1, a module for synthesizing detection results can be added to the control unit 10, and the detection results of the first detection unit 201 and the second detection unit 202A can be synthesized using this module and used as the final detection result. The same applies to the synthesis of the detection results of the first detection unit 201 and the second detection unit 202B. In addition, similar to the example described in "Regarding the Resolution of Input Data" in Embodiment 1, as the input data of the first learned model, or any one or both of the second learned model A and the second learned model B, image data with reduced resolution can be used.

[0113] In addition, in Figure 7 the example, two second detection units 202A and 202B are described as modules for detecting a part of the detection object of the first detection unit 201. However, the module for detecting a part of the detection object of the first detection unit 201 can be only one, or three or more.

[0114] [Process Flow]

[0115] Figure 8 is a diagram for explaining the process flow executed by the information processing apparatus 1 of the present embodiment. The process executed by the information processing apparatus 1 of the present embodiment and its execution order can be defined by, for example, Figure 8 the setting file F2 shown, and can also be represented by the flowchart of this figure.

[0116] (Regarding the Setting File)

[0117] In the setting file F2, three sections of [EX_all], [EX_goza], and [EX_tree] are defined in this order. Although the detailed content of each script is omitted, it is the same as that of SC11 and SC12. The learned model used in [EX_all] is the same as that of SC11, and it is a section for detecting all objects to be detected (e.g., corrugated paper, board, wood, mat, and strip). [EX_goza] is a section dedicated to the detection of mats among the objects to be detected, and [EX_tree] is a section dedicated to the detection of wood among the objects to be detected.

[0118] The src of [EX_goza] and [EX_tree] is "all_res" which is the dst of [EX_all]. That is, [EX_goza] detects mats from an image in which at least any one object to be detected is detected by [EX_all], and [EX_tree] detects wood from an image in which at least any one object to be detected is detected by [EX_all]. Moreover, the detection result of the mat based on [EX_goza] is output to [goza_res], and the detection result of the wood based on [EX_tree] is output to [tree_res]. These detection results are the final detection results.

[0119] (Regarding the flowchart)

[0120] The control unit 10 functions as the first detection unit 201 by executing the script file named "all" defined in [EX_all]. In addition, after the execution of the above script file is completed, the control unit 10 functions as the second detection unit 202A by executing the script file named "goza" defined in [EX_goza]. Moreover, after the execution of the above script file is completed, the control unit 10 functions as the second detection unit 202B by executing the script file named "tree" defined in [EX_tree].

[0121] The following describes the processing (information processing method) performed by these processing units based on the flowchart. It is assumed that before the processing of this flowchart, the moving image file captured by the garbage shooting device 2 is stored in "all_src" of the input data storage unit 121. In addition, instead of the moving image file, a plurality of frame images extracted from the moving image file or a plurality of still image files sequentially captured by the garbage shooting device 2 may be stored.

[0122] In S21 (the first detection step), the first detection unit 201 performs object detection on all detection target objects from the processing target image stored in the input data storage unit 121 using the first learned model. Specifically, the first detection unit 201 sets the frame image extracted from the moving image file of "all_src" stored in the input data storage unit 121 as the input data, inputs this frame image into the learned model with the script name "all", and makes it output object information. Moreover, when the first detection unit 201 detects an object from the frame image, it associates this frame image with the object information as the detection result and records it in "all_res" of the detection result storage unit 122. These processes are performed respectively for the frame images extracted from the above-mentioned moving image file.

[0123] As described above, the frame images recorded in "all_res" become the input data for [EX_goza] and [EX_tree], and an attempt is made to detect the mat and the wood again. Therefore, it is preferable that the first detection unit 201, although increasing the false detection of the mat and the wood, avoids overlooking the mat and the wood, that is, avoids being unable to detect the mat and the wood presented in the frame image. Thus, in the object detection based on the output value of the first learned model, the detection threshold for comparing with the probability value included in this output value can be set lower.

[0124] In S22 (the second detection step, determination step), the second detection unit 202A uses the second learned model A to detect the second detection target A in this example, that is, the mat, from the image in which an object is detected in S21. Specifically, the second detection unit 202A sets the frame image recorded in "all_res" of the detection result storage unit 122 as the input data, inputs this frame image into the learned model with the script name "goza", and makes it output object information. Moreover, when the second detection unit 202A determines that a mat has been detected, it associates this frame image with the object information as the detection result and records it in [goza_res] of the detection result storage unit 122. These processes are performed respectively for the frame images stored in "all_res". The data recorded in [goza_res] of the detection result storage unit 122 at the end of the processing in S22 is the final detection result for the mat. That is to say, through the processing in S22, the final detection result of the mat is determined.

[0125] In S23, the second detection unit 202B detects the second detection target B in this example, i.e., wood, from the image of the object detected in S21 by using the second learned model B. Specifically, the second detection unit 202B sets the frame image of "all_res" recorded in the detection result storage unit 122 as the input data, inputs the frame image into the learned model with the script name "tree", and makes it output object information. Moreover, when the second detection unit 202B determines that wood has been detected, it associates the frame image with the object information as the detection result and records it in [tree_res] of the detection result storage unit 122. These processes are performed for each frame image stored in "all_res". The data recorded in [tree_res] of the detection result storage unit 122 at the end of the process in S23 is the final detection result for wood. That is, through the process in S23, the final detection result of wood is determined.

[0126] In S24, the detection result is output in the same manner as Figure 5 S13 of, and thus the process ends. In addition, the detection result of the mat can be read out from [goza_res] of the detection result storage unit 122, and the detection result of wood can be read out from [tree_res] of the detection result storage unit 122. In addition, the detection results of detection target objects other than the mat and wood can be read out from "all_res" of the detection result storage unit 122.

[0127] In the detection result of [EX_all], there may be false detections in which an object that is not wood is detected as wood and an object that is not a mat is detected as a mat. However, according to the above process, for the frame image in which a certain object is detected in [EX_all], since it is provided for object detection based on [EX_tree] and [EX_goza], the mat and wood can be detected with high accuracy.

[0128] In addition, according to the above structure, compared with the case of using [EX_tree] and [EX_goza] for object detection, the processing can be speeded up. For example, in the case of extracting 200 frame images from a single moving image file, if [EX_tree] and [EX_goza] are used, 200 frame images are processed using [EX_tree] and [EX_goza] respectively. In this case, a total of 400 object detection processes are performed. On the other hand, according to the above structure, initially 200 frame images are processed using [EX_all] respectively. Here, when an object is detected in 30 frame images, the number of frame images processed by [EX_tree] and [EX_goza] is 30 respectively, and a total of 260 object detection processes are performed. Therefore, the number of executions of the object detection process can be significantly reduced, and the time required for this process can be significantly reduced.

[0129] 〔Embodiment 3〕

[0130] Based on Figure 9 A structural example of the control unit 10 of the information processing apparatus 1 according to the present embodiment will be described. Figure 9 It is a block diagram showing a structural example of the control unit 10 included in the information processing apparatus 1 according to Embodiment 3. In addition, in Figure 9 the large-capacity storage unit 12 is also shown in the figure.

[0131] As Figure 9 shown, the control unit 10 includes a garbage image extraction unit 301, a first detection unit 302, a second detection unit 303, and a detection result integration unit 304. In addition, the learning unit 103, the selection display control unit 104, the incoming vehicle determination unit 105, and the improper object display control unit 106 are the same as those in Embodiment 1, and thus their illustration is omitted.

[0132] The garbage image extraction unit 301 extracts an image presenting garbage from an image that is an object of object detection (for example, each frame image extracted from a moving image file). Thus, since the images that can be the detection objects of the first detection unit 302 and the second detection unit 303 can be locked to the images presenting garbage, the number of executions of the object detection process can be reduced, and the process can be speeded up. For example, in the case where 3 / 4 of the frame images extracted from a moving image file do not present garbage, the first detection unit 302 and the second detection unit 303 only need to perform object detection processing on the frame images presenting garbage (1 / 4 of the total frame images). Therefore, compared with the case of performing object detection processing on all frame images, the number of executions of the object detection process can be significantly reduced, and the time required for this process can be significantly reduced.

[0133] The garbage image extraction unit 301 can perform the above extraction using, for example, a learned model constructed from teaching data including images of garbage flowing on the inclined surface 600 and images that do not flow. This learned model only needs to be able to identify whether there is garbage on the inclined surface 600 and does not require discrimination of garbage classification, etc. Therefore, when this learned model is a neural network model, compared with the learned models used by the first detection unit 302 and the second detection unit 303 described below, the number of intermediate layers can be reduced.

[0134] In addition, the image used as the input data of the learned model used by the garbage image extraction unit 301 can be set as a low-resolution image compared with the images used as the input data of the learned models used by the first detection unit 302 and the second detection unit 303. In addition, it is preferable to input higher-resolution images to the first detection unit 302 and the second detection unit 303. Therefore, when the garbage image is made low-resolution and used as the input data of the garbage image extraction unit 301, it is preferable that the garbage image extraction unit 301 saves the image with the resolution before low-resolution as the output result.

[0135] The garbage image extraction unit 301 is not an essential structural element of the information processing device 1 and is included to make the object detection performed by the first detection unit 302 and the second detection unit 303 efficient. The garbage image extraction unit 301 can also be applied to the information processing device 1 of Embodiments 1 and 2.

[0136] The first detection unit 302 inputs input data to a first learned model that has been machine-learned to be able to detect a variety of first detection objects and detects the above first detection objects. In addition, the second detection unit 303 inputs the above input data to a third learned model that has been machine-learned to be able to detect a third detection object different from the above first detection object and detects the above third detection object.

[0137] Similar to the above-described embodiments, in the information processing device 1 of this embodiment, the final detection result is also determined based on the detection result of the first detection unit 302 and the detection result of the second detection unit 303. Specifically, the detection result integration unit 304 determines the final detection result based on the detection result of the first detection unit 302 and the detection result of the second detection unit 303.

[0138] More specifically, the detection result integration unit 304 uses the remaining part after removing the objects detected as the third detection object by the second detection unit 303 from the objects detected as the first detection object by the first detection unit 302 as the detection result of the first detection object. In other words, when the detection result based on the first learned model does not match the detection result based on the third learned model, the detection result integration unit 304 invalidates the detection result based on the first learned model.

[0139] According to the above structure, the detection objects of the first learned model and the third learned model are different. Therefore, for a certain detection object, when the first detection unit 302 detects it as the first detection object, for the same detection object, the second detection unit 303 can detect it as the third detection object. In such a case, it can be determined that either the first detection unit 302 or the second detection unit 303 has a false detection. Thus, according to the above structure, false detection can be reduced. This structure uses, as the detection result of the first detection object, the remaining part after removing, from the detection objects detected by the first detection unit 302 as the first detection object, the objects detected by the second detection unit 303 as the third detection object. Thus, according to the above structure, the detection accuracy of the detection using the machine-learned model can be improved. In addition, since the first detection unit 302 uses the first learned model that has been machine-learned in such a way as to be able to detect a variety of first detection objects, a variety of first detection objects can be effectively detected together.

[0140] In addition, similarly to the example described in "Regarding the Resolution of Input Data" in Embodiment 1, in this embodiment, image data with reduced resolution can also be used as the image data input to the first learned model or the third learned model.

[0141] 〔Flow of Processing〕

[0142] Figure 10 FIG. is a diagram for explaining the flow of processing executed by the information processing apparatus 1 of this embodiment. The processing executed by the information processing apparatus 1 of this embodiment and its execution order can be defined by, for example Figure 10 the setting file F3 shown, and can also be represented by the flowchart of this figure.

[0143] (Regarding the Setting File)

[0144] In the setting file F3, four sections, [EX_trash], [EX_all], [EX_bag], and [EX_final], are defined in this order. [EX_all] is the same as the section [EX1] of Figure 5 and is the section for detecting all detection objects (for example, corrugated paper, board, wood, mat, and long objects). [EX_trash] is the section for extracting images showing trash from images showing trash and images not showing trash. In addition, [EX_bag] is the section for using the garbage bag and the inclined surface 600 (refer to Figure 4 ) as detection objects, and [EX_final] is the section for outputting the result after removing the result of [EX_bag] from the result of [EX_all].

[0145] The dst of [EX_trash] is "trash_res", and "trash_res" is the src of [EX_all]. That is, the object to be detected is detected for the image extracted in [EX_trash] using [EX_all].

[0146] In addition, the dst of [EX_all] is "all_res", and "all_res" is the src of [EX_bag]. That is, the garbage bag and the inclined surface 600 are detected for the image in which at least any object to be detected is detected by [EX_all] using [EX_bag].

[0147] Moreover, the src of [EX_final] is "all_res", and the dst is "final_res". The detection results of [EX_all] and [EX_bag] are recorded in "all_res", so [EX_final] outputs the final detection result based on these detection results to "final_res".

[0148] Here, in the present embodiment, the reason for using the combination of [EX_all], [EX_bag], and [EX_final] is explained. The shapes and colors of garbage bags for loading garbage are countless, so the shapes of the parts between the garbage bags presented in the garbage image are also diverse. For example, when the part between the garbage bags is seen as a long bar, it may be misdetected as presenting an improper object based on this part where there is actually no improper object. To avoid such misdetection, it is preferable to also make it learn the garbage bag, but the number of garbage bags presented in the garbage image is, for example, about several tens in one image, which is very large. Therefore, it is very time-consuming and laborious to create correct data for all teaching images. The same applies to the inclined surface 600. Since the appearance of the inclined surface 600 presented in the image changes according to the state of the garbage on the inclined surface 600, it is very time-consuming and laborious to create correct data for the inclined surface for all teaching images.

[0149] Therefore, in the present embodiment, in addition to [EX_all] for detecting improper objects and the like, [EX_bag] dedicated to the detection of garbage bags and the inclined surface 600 is also used. The third learned model used in [EX_bag] can be constructed by machine learning using only the teaching data of garbage bags and the inclined surface 600.

[0150] [EX_final] is a section for confirming whether a garbage bag or a section of the inclined surface 600 is detected at the position where an improper object is detected by [EX_all]. Specifically, [EX_final] is a section that outputs the detection result of the improper object to "final_res" when neither a garbage bag nor the inclined surface 600 is detected at the position where an improper object is detected by [EX_all]. Thereby, it is possible to quickly sort out the detection results of improper objects based on [EX_all], excluding the part that may misdetect the garbage bag or the inclined surface 600 as an improper object according to the detection results of [EX_bag]. That is, it is possible to effectively reduce the misdetection of the garbage bag or the inclined surface 600 as an improper object with a short processing time.

[0151] (Regarding the flowchart)

[0152] The control unit 10 functions as the garbage image extraction unit 301 by executing the script file named "trash" defined in [EX_trash]. In addition, after the execution of the above script file ends, the control unit 10 functions as the first detection unit 302 by executing the script file named

[0153] "all" defined in [EX_all]. Moreover, after the execution of the above script file ends, the control unit 10 functions as the second detection unit 303 by executing the script file named "bag" defined in [EX_bag]. Moreover, after the execution of the above script file ends, the control unit 10 functions as the detection result integration unit 304 by executing the script file named "final" defined in [EX_final].

[0154] The following describes the processing (information processing method) performed by these processing units based on the flowchart. The moving image file captured by the garbage imaging device 2 before the processing of this flowchart is set and stored in "all_src" of the input data storage unit 121. In addition, instead of the moving image file, a plurality of frame images extracted from the moving image file or a plurality of still image files sequentially captured by the garbage imaging device 2 may be stored.

[0155] In S31, the garbage image extraction unit 301 extracts an image presenting garbage from the image to be processed. Specifically, the garbage image extraction unit 301 inputs the frame image extracted from the moving image file of "all_src" stored in the input data storage unit 121 into the learned model with the script name "trash" and makes it output object information. Moreover, the garbage image extraction unit 301 records the frame image determined to have detected garbage based on this object information in "trash_res" of the detection result storage unit 122. These processes are performed respectively for the frame images extracted from the above-mentioned moving image file.

[0156] In S32 (first detection step), the first detection unit 302 performs object detection on all detection target objects from the frame images extracted in S31. Specifically, the first detection unit 302 uses the frame image in "trash_res" stored in the detection result storage unit 122 as input data, inputs this frame image into the learned model with the script name "all" and makes it output object information. Moreover, the first detection unit 302 associates the frame image determined to have detected an object based on this object information with this object information as a detection result, and records it in "all_res" of the detection result storage unit 122. These processes are performed respectively for the frame images stored in "trash_res".

[0157] In S33 (second detection step), the second detection unit 303 detects the garbage bag and the inclined plane 600 from the images extracted in S31 using the third learned model. Specifically, the second detection unit 303 uses the frame image recorded in "trash_res" of the detection result storage unit 122 as input data, inputs this frame image into the learned model with the script name "bag" and makes it output object information. Moreover, the second detection unit 303 determines whether an object is detected based on this object information, that is, whether a garbage bag or an inclined plane is detected. Here, in the case where it is determined that an object is detected, the second detection unit 303 associates this frame image with the object information as a detection result, and records it in "all_res" of the detection result storage unit 122. These processes are performed respectively for the frame images stored in "trash_res".

[0158] In addition, a structure in which the garbage bag and the inclined plane 600 are detected using other learned models respectively may also be adopted. In addition, the detection target of the second detection unit 303 only needs to be an object different from the detection target of the first detection unit 302, and is not limited to the garbage bag and the inclined plane 600. However, it is preferable that the detection target of the second detection unit 303 is an object having an appearance similar to that of the detection target of the first detection unit 302. For example, the detection target of the second detection unit 303 may use an object having an appearance similar to that of an improper object but not being an improper object (such as corrugated paper, etc.) as the detection target.

[0159] In S34 (determination step), the detection result integration unit 304 determines the final detection result based on the detection results of S32 and S33. More specifically, the detection result integration unit 304 takes the remainder obtained by removing the objects detected as garbage bags or slopes by the second detection unit 303 in S33 from the detected objects detected by the first detection unit 302 in S32 as the final detection result of the detection target object.

[0160] Specifically, the detection result integration unit 304 determines the range occupied by the detection target object on the image according to the object information of each detection target object stored in "all_res" by the first detection unit 302. Next, the detection result integration unit 304 determines whether a garbage bag or a slope 600 is detected in the above range based on the object information stored in "all_res" by the second detection unit 303. Here, when the detection result integration unit 304 determines that no garbage bag or slope 600 is detected in the above range, it associates the object information of the detection target object with the frame image as the final detection result and records it in "final_res" of the detection result storage unit 122. On the other hand, when the detection result integration unit 304 determines that a garbage bag or a slope 600 is detected in the above range, it does not record the object information and the frame image of the detection target object. That is, the detection result of the detection target object is invalidated as a false detection.

[0161] In S35, the detection result is output in the same manner as Figure 5 S13 of, and thus the process ends. In addition, the output detection result can be read out from "final_res" of the detection result storage unit 122.

[0162] 〔Modification Example〕

[0163] In object detection, object classification, etc. in the above-described respective embodiments, artificial intelligence / machine learning algorithms other than neural networks (including content that has undergone deep learning) that have undergone machine learning can also be used.

[0164] The execution subject of each process described in the above-described respective embodiments can be appropriately changed. For example, Figure 1 , Figure 8 or Figure 10 at least any one of the respective blocks shown can be omitted, and the omitted processing unit can be provided in one or more other devices. In this case, the processing of the above-described respective embodiments is executed by one or more information processing devices.

[0165] In addition, examples of detecting inappropriate objects and the like from garbage images have been described in the above-described embodiments. However, the object to be detected is arbitrary and is not limited to inappropriate objects and the like. Moreover, the input data for the learned model used by the information processing apparatus 1 is not limited to image data, and may be, for example, sound data. In this case, the information processing apparatus 1 may be configured to detect a component of a specified sound included in the input sound data as an object to be detected.

[0166] The present invention is not limited to the above-described embodiments, and various modifications can be made within the scope of the claims. Embodiments obtained by appropriately combining technical solutions respectively disclosed in different embodiments are also included in the technical scope of the present invention.< / key> < / value> < / key> < / value> < / key> < / key>

Claims

1. An information processing apparatus, characterized in that, Comprising: A first detection unit that inputs input data to a first learned model that has been machine-learned in such a way as to be able to detect a plurality of first detection objects and detects the first detection objects; And A second detection unit that inputs the input data to a third learned model that has been machine-learned in such a way as to be able to detect a third detection object different from the first detection object and detects the third detection object, The input data is image data or sound data, The information processing device determines a final detection result based on the detection result of the first detection unit and the detection result of the second detection unit, The information processing device includes a detection result integration unit, and the detection result integration unit uses the remaining part after removing the detection objects detected as the third detection objects by the second detection unit from the detection objects detected as the first detection objects by the first detection unit as the final detection result of the first detection objects.

2. The information processing device according to claim 1, wherein The input data is image data, The first detection unit inputs data with reduced resolution of the image data to the first learned model, or The second detection unit inputs data with reduced resolution of the image data to the third learned model.

3. An information processing method, which is executed by one or more information processing devices, characterized in that, Including: A first detection step of inputting input data to a first learned model that has been machine-learned in such a way as to be able to detect a plurality of detection objects, and detecting the detection objects based on the input data; A second detection step of inputting the input data to a third learned model that has been machine-learned in such a way as to be able to detect detection objects different from those of the first learned model, and detecting detection objects from the input data; And A determination step of determining a final detection result based on the detection result of the first detection step and the detection result of the second detection step, The input data is image data or sound data, In the determination step, the remaining part after removing the detection objects detected in the second detection step from the detection objects detected in the first detection step is used as the final detection result.

4. A computer-readable recording medium that records an information processing program, the information processing program being used to cause a computer to function as the information processing device according to claim 1, and cause the computer to function as the first detection unit, the second detection unit, and the detection result integration unit.

Citation Information

Patent Citations

  • Image recognition device

    JP2008165394A

  • Image label determination method and device and terminal

    CN108664989A

  • Information processing method and device for image detection and storage medium

    CN109829491A