Inference program, learning program, inference method, and learning method

The inference program enhances self-checkout systems by accurately detecting and counting individual objects in target images through the generation of mask images and estimation of object positions, addressing the challenge of detecting objects not present in background images.

JP7687186B2Active Publication Date: 2025-06-03FUJITSU LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2021174706
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-10-26
Publication Date
2025-06-03
Estimated Expiration
2041-10-26

AI Technical Summary

Technical Problem

Existing self-checkout systems struggle to detect objects that exist in a target image but not in a background image, particularly when multiple objects are close together, leading to inaccurate identification and counting of products.

Method used

An inference program that acquires background and target images, generates intermediate feature amounts, creates a mask image to identify regions of objects not present in the background, and uses an estimation model to specify the position and size of individual objects within the target image.

Benefits of technology

Effectively detects and counts individual objects in the target image, even if they are unknown or closely positioned, improving the accuracy of product identification and fraud detection in self-checkout systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007687186000001
    Figure 0007687186000001
  • Figure 0007687186000002
    Figure 0007687186000002
  • Figure 0007687186000003
    Figure 0007687186000003
Patent Text Reader

Abstract

To detect an object that does not exist in a background image but exists in a target image.SOLUTION: An information processing apparatus is configured to: acquire a background image in which a target area in which an object is arranged is captured, and a target image in which the object and the area are captured; generate an intermediate feature quantity by inputting the background image and the target image to a feature extraction model; generate a mask image that indicates a region of an object that does not exist in the background image but exists in the target image by inputting the intermediate feature quantity to a generation model; and specify the object that does not exist in the background image but exists in the target image by inputting the generated mask image and the intermediate feature quantity to an estimation model.SELECTED DRAWING: Figure 5
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an inference program and the like.

Background Art

[0002] Self-checkout has become widespread in stores such as supermarkets and convenience stores. Self-checkout is a POS (Point Of Sale) cash register system in which the user who purchases the goods performs operations from reading the barcode of the goods to settlement. For example, by introducing self-checkout, it is possible to suppress labor costs and prevent settlement errors by store clerks.

[0003] On the other hand, in self-checkout, it is required to detect user fraud such as not reading the barcode.

[0004] FIG. 21 is a diagram for explaining the prior art. In the example shown in FIG. 21, it is assumed that user 1 picks up product 2 placed on temporary stand 6 and performs an operation of scanning the barcode of product 2 with respect to self-checkout 5 and then packs it. In the prior art, by analyzing the image data of camera 10, object detection of the products placed on temporary stand 6 is performed to identify the number of products. By checking whether the identified number of products matches the actually scanned number, fraud can be detected. In the prior art, when performing object detection as described in FIG. 21, technologies such as Deep Learning (hereinafter referred to as DL) are used.

[0005] When performing object detection using DL, a large amount of labeled data is prepared manually, and machine learning is executed on an object detection model for performing object detection using this labeled data. Here, since the object detection model only detects pre-learned objects, it is not realistic to repeatedly prepare labeled data and re-perform machine learning on the object detection model under the conditions where the types of products are enormous and the products are replaced daily as in the above-mentioned store.

[0006] There is also a prior art that can identify the area of such an object even for an unknown object that has not been previously learned. FIG. 22 is a diagram for explaining the prior art for identifying the area of an unknown object. In the prior art described with reference to FIG. 22, a DNN (Deep Neural Network) 20 that outputs a mask image 16 indicating an area different from the background is obtained by machine learning using a variety of large amounts of images from a background image 15a and a target image 15b.

[0007] The shooting area of the background image 15a and the shooting area of the target image 15b are the same shooting area, but the background image 15a does not include the objects 3a, 3b, and 3c that exist in the target image 15b. The mask image 16 shows areas 4a, 4b, and 4c corresponding to the objects 3a, 3b, and 3c. For example, the pixels in the areas 4a, 4b, and 4c of the mask image 16 are set to "1", and the pixels in the other areas are set to "0".

[0008] In the prior art of FIG. 22, since an area different from the background is identified, the area of an unknown object can also be identified. Therefore, by applying the prior art of FIG. 22 to the prior art of FIG. 21, it is conceivable to identify the area of an unknown product and identify the number of products.

Prior Art Documents

Patent Documents

[0009]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0010] However, in the above-described prior art, there is a problem that an object that exists in the target image but does not exist in the background image cannot be detected.

[0011] In the prior art described with reference to FIG. 22, it is a technique for identifying the entire area of a region different from the background, and does not identify the position, size, etc. of individual objects. Therefore, when a plurality of objects are close to each other, they will form a single mass.

[0012] FIG. 23 is a diagram for explaining the problems of the prior art. In the example shown in FIG. 23, when a background image 17a and a target image 17b are input to a pre-trained DNN 20, a mask image 18 is output. The target image 17b includes products 7a, 7b, and 7c. Since the products 7a to 7c in the target image 17b are close to each other, a single mass region 8 is shown in the mask image 18. Based on the region 8 of the mask image 18, it is difficult to identify the regions corresponding to the products 7a, 7b, and 7c and determine the product count "3".

[0013] On one aspect, an object of the present invention is to provide an inference program, a learning program, an inference method, and a learning method capable of detecting an object that exists in a target image but not in a background image.

Means for Solving the Problem

[0014] In the first aspect, the computer is caused to execute the following processing. The computer acquires a background image of the area where the object is to be placed and a target image of the object and the area. The computer generates intermediate feature amounts by inputting the background image and the target image into a feature extraction model. The computer generates a mask image indicating the region of the object that does not exist in the background image but exists in the target image by inputting the intermediate feature amounts into a generation model. The computer identifies the object that does not exist in the background image but exists in the target image by inputting the generated mask image and the intermediate feature amounts into an estimation model.

Effect of the Invention

[0015] An object that exists in the target image but not in the background image can be detected.

Brief Description of the Drawings

[0016]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23

BEST MODE FOR CARRYING OUT THE INVENTION

[0017] Hereinafter, embodiments of the inference program, learning program, inference method, and learning method disclosed in the present application will be described in detail with reference to the drawings. Note that the present invention is not limited to these embodiments.

EXAMPLE

[0018] An example of the system according to the first embodiment will be described. FIG. 1 is a diagram showing the system according to the first embodiment. As shown in FIG. 1, this system includes a camera 10 and an information processing apparatus 100. The camera 10 and the information processing apparatus 100 are interconnected via a network 11. The camera 10 and the information processing apparatus 100 may be directly connected by wire or wirelessly.

[0019] The camera 10 may be a camera that captures the state inside or outside the store, or may be a camera that captures the temporary placement table 6 on which the products shown in FIG. 21 are placed. The camera 10 transmits the image data within the imaging range to the information processing apparatus 100.

[0020] In the following description, image data captured by the camera 10 that does not contain an object to be detected is referred to as "background image data". For example, the background image data corresponds to the background image 15a described in FIG. 22. Image data captured by the camera 10 that contains an object to be detected is referred to as "target image data". For example, the target image data corresponds to the target image 15b described in FIG. 22. The target image 15b contains objects 3a to 3c. The shooting area of the background image 15a is the same as the shooting area of the target image 15b, but the background image 15a does not contain the objects 3a to 3c.

[0021] The information processing apparatus 100 is an apparatus that infers the regions of individual objects included in the target image data based on the background image data and the target image data. Before starting the inference, the information processing apparatus 100 receives the background image data from the camera 10 in advance, and sequentially receives the target image data from the camera 10 when starting the inference.

[0022] Hereinafter, the processing of the basic part of the information processing apparatus 100 and the characteristic processes 1 and 2 added to the processing of such basic part will be described in order.

[0023] FIG. 2 is a diagram for explaining the processing of the basic part of the information processing apparatus according to the first embodiment. As shown in FIG. 2, the information processing apparatus 100 that executes the processing of the basic part includes feature extraction units 50a and 50b, a synthesis unit 51a, and an estimation unit 52.

[0024] The feature extraction units 50a and 50b correspond to a general convolutional neural network (CNN). When the background image data 25a is input, the feature extraction unit 50a outputs an image feature amount to the synthesis unit 51a based on the parameters trained by machine learning. When the target image data 25b is input, the feature extraction unit 50b outputs an image feature amount based on the parameters trained by machine learning to the synthesis unit 51a.

[0025] The image feature amounts output from the feature extraction units 50a and 50b are the values before being converted into probability values based on a softmax function or the like. In the following description, the image feature amount of the background image data 25a is denoted as the "background image feature amount". The image feature amount of the target image data 25b is denoted as the "target image feature amount". The background image feature amount and the target image feature amount correspond to intermediate feature amounts.

[0026] The same parameters are set for the feature extraction units 50a and 50b. In FIG. 2, for convenience of explanation, the feature extraction units 50a and 50b are separately illustrated, but the feature extraction units 50a and 50b are the same CNN.

[0027] The combining unit 51a combines the background image feature amount and the target image feature, and outputs the combined feature amount to the estimation unit 52.

[0028] The estimation unit 52 corresponds to a general convolutional neural network (CNN). When the feature amount obtained by combining the background image feature amount and the target image feature amount is input, the estimation unit 52 identifies the BBOX of each object based on the parameters trained by machine learning. For example, a BBOX (Bounding Box) is region information surrounding an object and has information on the position and size. In the example shown in FIG. 2, three BBOXes 30a, 30b, and 30c are identified.

[0029] Subsequently, "characteristic process 1" added to the process of the basic part of the information processing apparatus 100 shown in FIG. 2 will be described. FIG. 3 is a diagram for explaining characteristic process 1. In FIG. 3, in addition to the feature extraction units 50a and 50b, the combining unit 51a, and the estimation unit 52 described in FIG. 2, it has a position coordinate feature amount output unit 53 and a combining unit 51b.

[0030] The description regarding the feature extraction units 50a and 50b is the same as the description regarding the feature extraction units 50a and 50b described in FIG. 2.

[0031] The combining unit 51a combines the background image feature amount and the target image feature amount, and outputs the combined feature amount to the combining unit 51b.

[0032] The position coordinate feature amount output unit 53 outputs a plurality of coordinate feature amounts with coordinate values arranged in the image plane. For example, as shown in FIG. 3, the position coordinate feature amount output unit 53 outputs an x coordinate feature amount 53a, a y coordinate feature amount 53b, and a distance feature amount 53c to the synthesis unit 51b.

[0033] For each pixel of the x coordinate feature amount 53a, coordinate values from "-1" to "+1" are set in ascending order in the row direction from left to right. The same coordinate value is set for each pixel in the column direction. For example, "-1" is set for each pixel in the leftmost column of the x coordinate feature amount 53a.

[0034] For each pixel of the y coordinate feature amount 53b, coordinate values from "-1" to "+1" are set in ascending order in the column direction from top to bottom. The same coordinate value is set for each pixel in the row direction. For example, "-1" is set for each pixel in the uppermost row of the y coordinate feature amount 53b.

[0035] For the distance feature amount 53c, coordinate values from "0" to "+1" are set in ascending order in the direction from the central pixel to the outside. For example, "0" is set for the central pixel of the distance feature amount 53c.

[0036] The synthesis unit 51b outputs information obtained by synthesizing the background image feature amount, the target image feature amount, the x coordinate feature amount 53a, the y coordinate feature amount 53b, and the distance feature amount 53c to the estimation unit 52.

[0037] When information obtained by synthesizing the background image feature amount, the target image feature amount, the x coordinate feature amount 53a, the y coordinate feature amount 53b, and the distance feature amount 53c is input, the estimation unit 52 identifies the BBOX of each object based on parameters trained by machine learning.

[0038] FIG. 4 is a diagram for supplementarily explaining characteristic process 1. For example, assume a case where convolution by a neural network is performed on the image 21 shown in FIG. 4. In a normal convolution process, since it is invariant to position, it is difficult to discriminate objects with the same appearance as separate ones. For example, the objects 22 and 23 included in the image 21 have the same appearance. Therefore, the result 22b obtained by performing the convolution process on the region 22a and the result 23b obtained by performing the convolution process on the region 23a will be the same.

[0039] In contrast, in the characteristic process 1 described with reference to FIG. 3, convolution is to be executed on the image feature amount with respect to the information obtained by synthesizing the x - coordinate feature amount 53a, the y - coordinate feature amount 53b, and the distance feature amount 53c. For example, when performing convolution on the region 22a, convolution is also performed on the regions 53a - 1, 53b - 1, and 53c - 1 together. Similarly, when performing convolution on the region 23a, convolution is also performed on the regions 53a - 2, 53b - 2, and 53c - 2 together. As a result, the result 22b obtained by performing the convolution process on the region 22a and the result 23b obtained by performing the convolution process on the region 23a will not be the same, and it becomes possible to discriminate the objects 22 and 23.

[0040] Next, “characteristic process 2” added to the process of the basic part of the information processing apparatus 100 shown in FIG. 2 will be described. FIG. 5 is a diagram for explaining characteristic process 2. In FIG. 5, in addition to the feature extraction units 50a and 50b, the synthesis units 51a and 51b, the estimation unit 52, and the position coordinate feature amount output unit 53 described with reference to FIG. 3, a mask generation unit 54 is provided.

[0041] The explanations regarding the feature extraction units 50a and 50b and the position coordinate feature amount output unit 53 are the same as those given in FIGS. 2 and 3.

[0042] The synthesis unit 51a synthesizes the background image feature amount and the target image feature amount, and outputs the synthesized feature amount to the synthesis unit 51b and the mask generation unit 54.

[0043] The mask generation unit 54 is compatible with a general convolutional neural network (CNN). When the feature quantity synthesized from the background image feature quantity and the target image feature quantity is input, the mask generation unit 54 generates a mask image 40 based on the parameters trained by machine learning. The mask image 40 is information indicating the region of an object that does not exist in the background image data 25a but exists in the target image data 25b. For example, the mask image 40 is a bitmap, where "1" is set for the pixels corresponding to the object region and "0" is set for the pixels corresponding to other regions.

[0044] The synthesis unit 51b outputs the synthesis information 45 synthesized from the background image feature quantity, the target image feature, the x - coordinate feature quantity 53a, the y - coordinate feature quantity 53b, the distance feature quantity 53c, and the mask image 40 to the estimation unit 52.

[0045] When the synthesis information 45 is input, the estimation unit 52 identifies the BBOX of each object based on the parameters trained by machine learning. For example, the synthesis information 45 is information in which the background image feature quantity, the target image feature, the x - coordinate feature quantity 53a, the y - coordinate feature quantity 53b, the distance feature quantity 53c, and the mask image 40 overlap. The estimation unit 52 places a kernel with the parameters set on the synthesis information 45 where each information overlaps, and performs convolution while moving the position of the kernel.

[0046] Here, a supplementary explanation regarding the characteristic process 2 will be given. For example, assuming machine learning without using the mask generation unit 54, machine learning will be performed using the learning background image data and the learning target image data as input data, and the BBOX of the object included in the learning target image data as the correct data (GT: Ground Truth).

[0047] When performing such machine learning, it may be possible to memorize the features of individual objects included in the target image data and estimate the BBOX of the object only from the target image data without using the background image data. That is, it can be said that the objects included in the target image data for learning are memorized as they are and cannot correspond to unknown objects, which is called overfitting (overlearning).

[0048] In order to suppress the above overfitting, machine learning is performed using, as an auxiliary task, a task that cannot be solved without using background image data, so that the NN utilizes the background image. For example, the process of machine learning the mask generation unit 54 shown in FIG. 5 becomes an auxiliary task. For example, the above-mentioned estimation of the BBOX is the main task, and the task of generating a mask image is the auxiliary task.

[0049] Further, the mask image 40 generated by the mask generation unit 54 is input to the estimation unit 52, and machine learning is executed to estimate the BBOX of the object. As a result, it is possible to expect the effect of limiting the object to be detected to the object area of the mask image.

[0050] In FIG. 5, the information processing apparatus 100 inputs input data to the feature extraction units 50a and 50b, and the parameters of the feature extraction units 50a and 50b, the estimation unit 52, and the mask generation unit 54 are trained so that the error between the BBOX output by the estimation unit 52 and the correct data (correct value of the BBOX) and the error between the mask image output from the mask generation unit 54 and the correct data (correct value of the mask image) are reduced.

[0051] Next, an example of the configuration of the information processing apparatus 100 that executes the processing described with reference to FIGS. 2 to 4 will be described. FIG. 6 is a functional block diagram showing the configuration of the information processing apparatus according to the first embodiment. As shown in FIG. 6, the information processing apparatus 100 includes a communication unit 110, an input unit 120, a display unit 130, a storage unit 140, and a control unit 150.

[0052] The communication unit 110 performs data communication with the camera 10 and an external device (not shown). For example, the communication unit 110 receives image data (background image data, target image data) from the camera 10. The communication unit 110 receives learning data 141 and the like used for machine learning from the external device.

[0053] The input unit 120 corresponds to an input device for inputting various types of information into the information processing apparatus 100.

[0054] The display unit 130 displays the output result from the control unit 150.

[0055] The storage unit 140 has learning data 141, an image table 142, a feature extraction model 143, a generation model 144, and an estimation model 145. The storage unit 140 corresponds to a semiconductor memory element such as a RAM (Random Access Memory) or a flash memory, or a storage device such as an HDD (Hard Disk Drive).

[0056] The learning data 141 is data used when performing machine learning. FIG. 7 is a diagram showing an example of the data structure of the learning data according to the first embodiment. As shown in FIG. 7, the learning data 141 holds the item number, the input data, and the correct answer data (GT) in association with each other. The input data includes background image data for learning and target image data for learning. The correct answer data includes the GT of the mask image and the GT of the BBOX (coordinates of the region of the object).

[0057] The image table 142 is a table that holds background image data and target image data used during inference.

[0058] The feature extraction model 143 is a machine learning model (CNN) executed by the feature extraction units 50a and 50b. When image data is input to the feature extraction model 143, an image feature amount is output.

[0059] The generation model 144 is a machine learning model (CNN) executed by the mask generation unit 54. When information synthesized from the background image feature amount and the target image feature is input to the generation model 144, a mask image is output.

[0060] The estimation model 145 is a machine learning model (CNN) executed by the estimation unit 52. When the synthesis information 45 is input to the estimation model 145, a BBOX is output.

[0061] The control unit 150 includes an acquisition unit 151, a learning processing unit 152, and an inference processing unit 153. The control unit 150 corresponds to a CPU (Central Processing Unit) or the like.

[0062] When the acquisition unit 151 acquires the learning data 141 from an external device or the like, the acquired learning data 141 is registered in the storage unit 140.

[0063] The acquisition unit 151 pre-acquires background image data from the camera 10 and registers it in the image table 142. The acquisition unit 151 acquires target image data from the camera 10 and registers it in the image table 142.

[0064] Based on the learning data 141, the learning processing unit 152 executes machine learning for the feature extraction units 50a, 50b (feature extraction model 143), the mask generation unit 54 (generation model 144), and the estimation unit 52 (estimation model 145).

[0065] FIG. 8 is a diagram for explaining the processing of the learning processing unit according to the first embodiment. For example, the learning processing unit 152 includes feature extraction units 50a, 50b, synthesis units 51a, 52b, an estimation unit 52, a mask generation unit 54, and a position coordinate feature amount output unit 53. The learning processing unit 152 also includes error calculation units 60a, 60b, a synthesis unit 61, and a weight update value calculation unit 62. In the following description, the feature extraction units 50a, 50b, the estimation unit 52, the position coordinate feature amount output unit 53, and the mask generation unit 54 are collectively referred to as the "neural network" as appropriate.

[0066] The processing of the feature extraction units 50a and 50b is the same as the description given in FIG. 5. For example, the feature extraction units 50a and 50b read and execute the feature extraction model 143. The feature extraction units 50a and 50b input the image data into the feature extraction model 143, and calculate the image feature amount based on the parameters of the feature extraction model 143.

[0067] The description of the composition units 51a and 51b is the same as the description given in FIG. 5.

[0068] The processing of the position coordinate feature amount output unit 53 is the same as the description given in FIG. 3.

[0069] The processing of the mask generation unit 54 is the same as the description given in FIG. 5. For example, the mask generation unit 54 reads and executes the generation model 144. The mask generation unit 54 inputs the feature amount obtained by synthesizing the background image feature amount and the target image feature amount into the generation model 144, and generates a mask image based on the parameters of the generation model 144. The mask generation unit 54 outputs the mask image to the composition unit 51b and the error calculation unit 60a.

[0070] The processing of the estimation unit 52 is the same as the description given in FIG. 5. For example, the estimation unit 52 reads and executes the estimation model 145. The estimation unit 52 reads and executes the estimation model 145. The estimation unit 52 inputs the synthesis information into the estimation model 145, and specifies the BBOX of each object based on the parameters of the estimation model 145. The estimation model 145 outputs the BBOX to the error calculation unit 60b.

[0071] The learning processing unit 152 acquires the learning background image data 26a from the learning data 141 and inputs it to the feature extraction unit 50a. The learning processing unit 152 acquires the learning target image data 26b from the learning data 141 and inputs it to the feature extraction unit 50b. Further, the learning processing unit 152 acquires the GT of the mask image from the learning data 141 and inputs it to the error calculation unit 60a. The learning processing unit 152 acquires the GT of the BBOX from the learning data 141 and inputs it to the error calculation unit 60b.

[0072] The error calculation unit 60a calculates the error between the mask image 41 output from the mask generation unit 54 and the GT of the mask image in the training data 141. In the following description, the error between the mask image 41 and the GT of the mask image is referred to as the "first error". The error calculation unit 60a outputs the first error to the synthesis unit 61.

[0073] The error calculation unit 60b calculates the error between the BBOX output from the estimation unit 52 and the GT of the BBOX in the training data 141. In the following description, the error between the BBOX output from the estimation unit 52 and the GT of the BBOX in the training data 141 is referred to as the "second error". The error calculation unit 60b outputs the second error to the synthesis unit 61.

[0074] The synthesis unit 61 calculates the sum of the first error and the second error. In the following description, the sum of the first error and the second error is referred to as the "total error". The synthesis unit 61 outputs it to the weight update value calculation unit 62.

[0075] The weight update value calculation unit 62 updates the parameters (weights) of the neural network so that the total error becomes smaller. For example, the weight update value calculation unit 62 uses the error backpropagation method or the like to update the parameters of the feature extraction units 50a, 50b (feature extraction model 143), the mask generation unit 54 (generation model 144), and the estimation unit 52 (estimation model 145).

[0076] The learning processing unit 152 repeatedly executes the above processing using each input data and correct answer data stored in the training data 141. The learning processing unit 152 registers the machine-learned feature extraction model 143, generation model 144, and estimation model 145 in the storage unit 140.

[0077] Returning to the description of FIG. 6, the inference processing unit 153 uses the machine-learned feature extraction units 50a, 50b (feature extraction model 143), the mask generation unit 54 (generation model 144), and the estimation unit 52 (estimation model 145) to identify the region of an object that exists in the target image data but does not exist in the background image data.

[0078] FIG. 9 is a diagram for explaining the processing of the inference processing unit according to the first embodiment. For example, the inference processing unit 153 includes feature extraction units 50a and 50b, composition units 51a and 52b, an estimation unit 52, a mask generation unit 54, and a position coordinate feature amount output unit 53.

[0079] The processing of the feature extraction units 50a and 50b is the same as the description given in FIG. 5. For example, the feature extraction units 50a and 50b read and execute the learned feature extraction model 143. The feature extraction units 50a and 50b input the image data into the feature extraction model 143, and calculate the image feature amount based on the parameters of the feature extraction model 143.

[0080] The description of the composition units 51a and 51b is the same as the description given in FIG. 5.

[0081] The processing of the position coordinate feature amount output unit 53 is the same as the description given in FIG. 3.

[0082] The processing of the mask generation unit 54 is the same as the description given in FIG. 5. For example, the mask generation unit 54 reads and executes the learned generation model 144. The mask generation unit 54 inputs the feature amount obtained by synthesizing the background image feature amount and the target image feature amount into the generation model 144, and generates a mask image based on the parameters of the generation model 144. The mask generation unit 54 outputs the mask image to the composition unit 51b.

[0083] The processing of the estimation unit 52 is the same as the description given in FIG. 5. For example, the estimation unit 52 reads and executes the learned estimation model 145. The estimation unit 52 reads and executes the estimation model 145. The estimation unit 52 inputs the synthesis information 45 into the estimation model 145, and specifies the BBOX of each object based on the parameters of the estimation model 145.

[0084] The inference processing unit 153 acquires the background image data 25a from the image table 142 and inputs it to the feature extraction unit 50a. The inference processing unit 153 acquires the target image data 25b from the image table 142 and inputs it to the feature extraction unit 50b. The inference processing unit 153 may output the information of the BBOX specified by the estimation unit 52 to the display unit 130 or to an external device.

[0085] Next, an example of the processing procedure of the information processing apparatus 100 according to the first embodiment will be described. Hereinafter, the processing procedure of the learning process and the processing procedure of the inference process executed by the information processing apparatus 100 will be described in order.

[0086] The processing procedure of the learning process will be described. FIGS. 10 and 11 are flowcharts showing the processing procedure of the learning process according to the first embodiment. As shown in FIG. 10, the learning processing unit 152 of the information processing apparatus 100 acquires background image data from the learning data 141 (step S101). The feature extraction unit 50a of the learning processing unit 152 extracts background image feature amounts based on the background image data (step S102).

[0087] The learning processing unit 152 acquires target image data from the learning data 141 (step S103). The feature extraction unit 50b of the learning processing unit 152 extracts target image feature amounts based on the target image data (step S104).

[0088] The synthesis unit 51a of the learning processing unit 152 synthesizes the background image feature amounts and the target image feature amounts (step S105). The mask generation unit 54 of the learning processing unit 152 generates a mask image based on the synthesized feature amounts (step S106).

[0089] The position coordinate feature amount output unit 53 of the learning processing unit 152 generates position coordinate feature amounts (step S107). The synthesis unit 51b of the learning processing unit 152 generates synthesis information obtained by synthesizing each feature amount (step S108).

[0090] The estimator 52 of the learning processing unit 152 estimates the BBOX based on the composite information (step S109). The learning processing unit 152 proceeds to step S110 in FIG. 11.

[0091] Proceed to the description of FIG. 11. The learning processing unit 152 obtains the GT of the mask image from the learning data 141 (step S110). The error calculation unit 60a of the learning processing unit 152 calculates the first error based on the mask image and the GT of the mask image (step S111).

[0092] The learning processing unit 152 obtains the GT of the BBOX from the learning data 141 (step S112). The error calculation unit 60b calculates the second error based on the BBOX and the GT of the BBOX (step S113).

[0093] The composite unit 61 of the learning processing unit 152 calculates the total error of the first error and the second error (step S114). The weight update value calculation unit 62 of the learning processing unit 152 calculates the update value of the parameters of the neural network (step S115). The learning processing unit 152 updates the parameters of the neural network (step S116).

[0094] If the learning processing unit 152 continues machine learning (step S117, Yes), it proceeds to step S101 in FIG. 10. If it does not continue machine learning (step S117, No), the machine learning of the neural network ends.

[0095] Subsequently, the processing procedure of the inference process will be described. FIG. 12 is a flowchart showing the processing procedure of the inference process according to the first embodiment. As shown in FIG. 12, the inference processing unit 153 of the information processing apparatus 100 obtains background image data from the image table 142 (step S201). The feature extraction unit 50a of the inference processing unit 153 extracts background image feature amounts based on the background image data (step S202).

[0096] The inference processing unit 153 acquires target image data from the image table 142 (step S203). The feature extraction unit 50b of the inference processing unit 153 extracts target image feature amounts based on the target image data (step S204).

[0097] The composition unit 51a of the inference processing unit 153 composes the background image feature amounts and the target image feature amounts (step S205). The mask generation unit 54 of the inference processing unit 153 generates a mask image based on the composed feature amounts (step S206).

[0098] The position coordinate feature amount output unit 53 of the inference processing unit 153 generates position coordinate feature amounts (step S207). The composition unit 51b of the inference processing unit 153 generates composite information obtained by composing the respective feature amounts (step S208).

[0099] The estimation unit 52 of the inference processing unit 153 estimates a BBOX based on the composite information (step S209).

[0100] Next, the effects of the information processing apparatus 100 according to the first embodiment will be described. The information processing apparatus 100 inputs background image data to the feature extraction unit 50a and inputs target image data to the feature extraction unit 50b, thereby extracting background image feature amounts and target image feature amounts. The information processing apparatus 100 inputs the feature amounts obtained by composing the background image feature amounts and the target image feature amounts to the mask generation unit 54 to generate a mask image. The information processing apparatus 100 inputs the mask image and the information obtained by composing the respective feature amounts to the estimation unit 52 to specify the region of an object. Thereby, even if the object included in the target image data is an unknown object that has not been learned in advance, each object can be discriminated and detected.

[0101] The information processing apparatus 100 inputs the information obtained by composing the background image feature amounts, the target image feature amounts, the mask image, and the coordinate feature amounts to the estimation unit 52 to specify the region of an object. Thereby, even when the target image data includes objects having the same appearance, it is possible to execute a convolution process so that the respective objects can be distinguished.

[0102] The information processing apparatus 100 executes machine learning of the feature extraction units 50a and 50b, the mask generation unit 54, and the estimation unit 52 based on the learning data 141. As a result, it is possible to perform machine learning of a neural network that can discriminate and detect each object even if the object included in the target image data is an unknown object that has not been pre-learned.

[0103] In addition to each feature amount, the information processing apparatus 100 inputs information obtained by further synthesizing coordinate feature amounts to the estimation unit 52 and executes machine learning. As a result, even when the target image data includes objects having the same appearance, it is possible to distinguish each object and perform machine learning of the neural network.

[0104] In addition to each feature amount, the information processing apparatus 100 inputs information obtained by further synthesizing a mask image to the estimation unit 52 and executes machine learning. As a result, it is possible to expect an effect of limiting the object to be detected to the object region of the mask image.

Example

[0105] The configuration of the system according to the second embodiment is the same as the system described in the first embodiment. It is assumed that the information processing apparatus according to the second embodiment is connected to the camera 10 via the network 11 in the same manner as in the first embodiment.

[0106] The information processing apparatus according to the second embodiment performs machine learning on the feature extraction units 50a and 50b, which are the basic parts described in FIG. 2, and the estimation unit 52. The information processing apparatus identifies each object using the feature extraction units 50a and 50b and the estimation unit 52 on which machine learning has been performed.

[0107] FIG. 13 is a functional block diagram showing the configuration of the information processing apparatus according to the second embodiment. As shown in FIG. 13, this information processing apparatus 200 includes a communication unit 210, an input unit 220, a display unit 230, a storage unit 240, and a control unit 250.

[0108] The descriptions of the communication unit 210, the input unit 220, and the display unit 230 are the same as those of the communication unit 110, the input unit 120, and the display unit 130 described in the first embodiment.

[0109] The storage unit 240 has learning data 241, an image table 242, a feature extraction model 243, and an estimation model 244. The storage unit 240 corresponds to a semiconductor memory element such as a RAM or a flash memory, or a storage device such as an HDD.

[0110] The learning data 241 is data used when performing machine learning. FIG. 14 is a diagram showing an example of the data structure of the learning data according to the second embodiment. As shown in FIG. 14, the learning data 241 holds the item number, the input data, and the correct answer data (GT) in association with each other. The input data includes background image data for learning and target image data for learning. The correct answer data includes the GT of the BBOX (coordinates of the object region).

[0111] The image table 242 is a table that holds background image data and target image data used during inference.

[0112] The feature extraction model 243 is a machine learning model (CNN) executed by the feature extraction units 50a and 50b. When image data is input to the feature extraction model 243, an image feature amount is output.

[0113] The estimation model 244 is a machine learning model (CNN) executed by the estimation unit 52. When the background image feature amount and the target image feature amount are input to the estimation model 244, a BBOX is output.

[0114] The control unit 250 has an acquisition unit 251, a learning processing unit 252, and an inference processing unit 253. The control unit 250 corresponds to a CPU or the like.

[0115] When the acquisition unit 251 acquires the learning data 241 from an external device or the like, the acquired learning data 241 is registered in the storage unit 240.

[0116] The acquisition unit 251 pre-acquires background image data from the camera 10 and registers it in the image table 242. The acquisition unit 251 acquires target image data from the camera 10 and registers it in the image table 242.

[0117] The learning processing unit 252 performs machine learning on the feature extraction units 50a, 50b (feature extraction model 243) and the estimation unit 52 (estimation model 244) based on the learning data 241.

[0118] FIG. 15 is a diagram for explaining the processing of the learning processing unit according to the second embodiment. For example, the learning processing unit 252 includes feature extraction units 50a, 50b, a synthesis unit 51a, and an estimation unit 52. The learning processing unit 252 also includes an error calculation unit 80 and a weight update value calculation unit 81. In the following description, the feature extraction units 50a, 50b and the estimation unit 52 are collectively referred to as the "neural network" as appropriate.

[0119] The processing of the feature extraction units 50a, 50b is the same as that described with reference to FIG. 2. For example, the feature extraction units 50a, 50b read and execute the feature extraction model 143. The feature extraction units 50a, 50b input image data to the feature extraction model 243 and calculate image feature amounts based on the parameters of the feature extraction model 243.

[0120] The synthesis unit 51a synthesizes the background image feature amount and the target image feature amount and outputs the synthesized feature amount to the estimation unit 52.

[0121] The estimation unit 52 reads and executes the estimation model 244. The estimation unit 52 reads and executes the estimation model 244. The estimation unit 52 inputs the synthesized feature amount to the estimation model 244 and identifies the BBOX of each object based on the parameters of the estimation model 244. The estimation model 244 outputs the BBOX to the error calculation unit 80.

[0122] The learning processing unit 252 acquires the background image data 26a for learning from the learning data 241 and inputs it to the feature extraction unit 50a. The learning processing unit 252 acquires the target image data 26b for learning from the learning data 241 and inputs it to the feature extraction unit 50b. The learning processing unit 252 acquires the GT of the BBOX from the learning data 241 and inputs it to the error calculation unit 80.

[0123] The error calculation unit 80 calculates the error between the BBOX output from the estimation unit 52 and the GT of the BBOX in the learning data 241. The error calculation unit 80 outputs the calculated error to the weight update value calculation unit 81.

[0124] The weight update value calculation unit 81 updates the parameters (weights) of the neural network so that the error becomes smaller. For example, the weight update value calculation unit 81 uses the error backpropagation method or the like to update the parameters of the feature extraction units 50a and 50b (feature extraction model 243) and the estimation unit 52 (estimation model 244).

[0125] The learning processing unit 252 repeatedly executes the above processing using each input data and correct answer data stored in the learning data 241. The learning processing unit 252 registers the learned feature extraction model 243 and estimation model 244 in the storage unit 240.

[0126] Returning to the description of FIG. 13. The inference processing unit 253 uses the learned feature extraction units 50a and 50b (feature extraction model 243) and the estimation unit 52 (estimation model 244) to identify the region of an object that exists in the target image data but does not exist in the background image data.

[0127] FIG. 16 is a diagram for explaining the processing of the inference processing unit according to the second embodiment. For example, the inference processing unit 253 includes the feature extraction units 50a and 50b, the composition unit 51a, and the estimation unit 52.

[0128] The processing of the feature extraction units 50a and 50b is the same as the description given in FIG. 2. For example, the feature extraction units 50a and 50b read and execute the learned feature extraction model 243. The feature extraction units 50a and 50b input the image data into the feature extraction model 243 and calculate the image feature amounts based on the parameters of the feature extraction model 243.

[0129] The synthesis unit 51a synthesizes the background image feature amount and the target image feature amount, and outputs the synthesized feature amount to the estimation unit 52.

[0130] The processing of the estimation unit 52 is the same as the description given in FIG. 2. For example, the estimation unit 52 reads and executes the learned estimation model 244. The estimation unit 52 reads and executes the estimation model 244. The estimation unit 52 inputs the information obtained by synthesizing the background image feature amount and the target image feature amount into the estimation model 244, and specifies the BBOX of each object based on the parameters of the estimation model 244.

[0131] The inference processing unit 253 acquires the background image data 25a from the image table 242 and inputs it to the feature extraction unit 50a. The inference processing unit 253 acquires the target image data 25b from the image table 242 and inputs it to the feature extraction unit 50b. The inference processing unit 253 may output the information of the BBOX specified by the estimation unit 52 to the display unit 230 or to an external device.

[0132] Next, an example of the processing procedure of the information processing apparatus 200 according to the second embodiment will be described. Below, the processing procedure of the learning process and the processing procedure of the inference process executed by the information processing apparatus 200 will be described in order.

[0133] The processing procedure of the learning process will be described. FIG. 17 is a flowchart showing the processing procedure of the learning process according to the second embodiment. As shown in FIG. 17, the learning processing unit 252 of the information processing apparatus 200 acquires background image data from the learning data 241 (step S301). The feature extraction unit 50a of the learning processing unit 252 extracts a background image feature amount based on the background image data (step S302).

[0134] The learning processing unit 252 acquires target image data from the learning data 241 (step S303). The feature extraction unit 50b of the learning processing unit 252 extracts target image feature amounts based on the target image data (step S304).

[0135] The synthesis unit 51a of the learning processing unit 252 synthesizes the background image feature amounts and the target image feature amounts (step S305). The estimation unit 52 of the learning processing unit 252 estimates the BBOX based on the synthesized feature amounts (step S306).

[0136] The learning processing unit 252 acquires the GT of the BBOX from the learning data 241 (step S307). The error calculation unit 80 calculates an error based on the BBOX and the GT of the BBOX (step S308).

[0137] The weight update value calculation unit 81 of the learning processing unit 252 calculates an update value of the parameters of the neural network (step S309). The learning processing unit 252 updates the parameters of the neural network (step S310).

[0138] If the learning processing unit 252 continues machine learning (step S311, Yes), it proceeds to step S301. If it does not continue machine learning (step S311, No), it ends the machine learning of the neural network.

[0139] Subsequently, the processing procedure of the inference processing will be described. FIG. 18 is a flowchart showing the processing procedure of the inference processing according to the second embodiment. As shown in FIG. 18, the inference processing unit 253 of the information processing apparatus 200 acquires background image data from the image table 242 (step S401). The feature extraction unit 50a of the inference processing unit 253 extracts background image feature amounts based on the background image data (step S402).

[0140] The inference processing unit 253 acquires target image data from the image table 242 (step S403). The feature extraction unit 50b of the inference processing unit 253 extracts target image feature amounts based on the target image data (step S404).

[0141] The composition unit 51a of the inference processing unit 253 composes the background image feature amounts and the target image feature amounts (step S405).

[0142] The estimation unit 52 of the inference processing unit 253 estimates the BBOX based on the composed feature amounts (step S406).

[0143] Next, the effects of the information processing apparatus 200 according to the second embodiment will be described. The information processing apparatus 200 inputs the background image data to the feature extraction unit 50a and inputs the target image data to the feature extraction unit 50b, thereby extracting the background image feature amounts and the target image feature amounts. The information processing apparatus 100 inputs the feature amounts obtained by composing the background image feature amounts and the target image feature amounts to the estimation unit 52, thereby specifying the region of the object. As a result, even if the object included in the target image data is an unknown object that has not been pre-learned, each object can be discriminated and detected.

Embodiment

[0144] Next, an example of the system according to the third embodiment will be described. FIG. 19 is a diagram showing the system according to the third embodiment. As shown in FIG. 19, this system includes a self-checkout 5, a camera 10, and an information processing apparatus 300. The self-checkout 5, the camera 10, and the information processing apparatus 300 are connected by wire or wirelessly.

[0145] The user 1 picks up the product 2 placed on the temporary stand 6 and performs an operation of scanning the barcode of the product 2 with respect to the self-checkout 5, and it is assumed that the product is packaged.

[0146] Self-checkout 5 is a POS (Point Of Sale) register system in which user 1 who purchases a product performs operations from reading the barcode of the product to settlement. For example, when user 1 moves the product to be purchased to the scanning area of self-checkout 5, self-checkout 5 scans the barcode of the product. When the scanning by user 1 is completed, self-checkout 5 notifies information processing device 300 of the information on the number of scanned products. In the following description, the information on the number of scanned products is referred to as "scanning information".

[0147] Camera 10 is a camera that photographs temporary stand 6 of self-checkout 5. Camera 10 transmits the image data of the shooting range to information processing device 300. It is assumed that camera 10 transmits in advance the image data (background image data) of temporary stand 6 on which no product is placed to information processing device 300. When a product to be purchased is placed on temporary stand 6, camera 10 transmits the image data (target image data) of temporary stand 6 to information processing device 300.

[0148] Information processing device 300 performs machine learning of a neural network in the same manner as information processing device 100 described in Example 1. The neural network includes feature extraction units 50a, 50b, synthesis units 51a, 51b, estimation unit 52, position coordinate feature amount output unit 53, and mask generation unit 54.

[0149] Information processing device 300 identifies each object included in the target image data by inputting the background image data and the target image data into the machine-learned neural network. Information processing device 300 counts the identified objects to identify the product number. When the identified product number does not match the product number included in the scanning information, information processing device 300 detects a scanning omission.

[0150] For example, the information processing apparatus 300 outputs, as an output result 70, the result of inputting background image data and target image data into a neural network. Since the output result 70 includes three BBOXes, i.e., BBOX 70a, 70b, and 70c, the information processing apparatus 300 specifies the number of items as "3". When the number of items included in the scan information is less than "3", the information processing apparatus 300 detects a scanning omission. The information processing apparatus 300 may notify a management server (not shown) or the like of the scanning omission.

[0151] As described above, by applying the information processing apparatus 100 (200) described in the first and second embodiments to the system shown in FIG. 19, it is possible to detect user fraud such as not reading a barcode.

[0152] Next, an example of the hardware configuration of a computer that realizes the same functions as the information processing apparatus 100 (200, 300) shown in the above embodiments will be described. FIG. 20 is a diagram showing an example of the hardware configuration of a computer that realizes the same functions as the information processing apparatus.

[0153] As shown in FIG. 20, the computer 400 includes a CPU 401 that executes various arithmetic processes, an input device 402 that receives input of data from a user, and a display 403. The computer 400 also includes a communication device 404 that receives distance image data from the camera 10, and an interface device 405 that connects to various devices. The computer 400 includes a RAM 406 that temporarily stores various information, and a hard disk device 407. Then, each of the devices 401 to 407 is connected to a bus 408.

[0154] The hard disk device 407 includes an acquisition program 407a, a learning process program 407b, and an inference process program 407c. The CPU 401 reads out the acquisition program 407a, the learning process program 407b, and the inference process program 407c and expands them in the RAM 406.

[0155] The acquisition program 407a functions as an acquisition process 406a. The learning process program 407b functions as a learning process 406b. The inference process program 407c functions as an inference process 406c.

[0156] The processing of the acquisition process 406a corresponds to the processing of the acquisition units 151 and 251. The processing of the learning process 406b corresponds to the processing of the learning processing units 152 and 252. The processing of the inference process 406c corresponds to the processing of the inference processing units 153 and 253.

[0157] Note that for each of the programs 407a to 407c, it is not necessarily stored in the hard disk device 407 from the beginning. For example, each program is stored in a "portable physical medium" such as a flexible disk (FD), CD-ROM, DVD disk, magneto-optical disk, or IC card inserted into the computer 400. Then, the computer 400 may read and execute each of the programs 407a to 407c.

[0158] Regarding the embodiments including the above embodiments, the following additional remarks are disclosed.

[0159] (Supplementary Note 1) An inference program characterized by causing a computer to execute a process of acquiring a background image of an area where an object is to be placed and a target image of the object and the area, generating an intermediate feature amount by inputting the background image and the target image into a feature extraction model, generating a mask image showing a region of an object that does not exist in the background image but exists in the target image by inputting the intermediate feature amount into a generation model, and specifying an object that does not exist in the background image but exists in the target image by inputting the generated mask image and the intermediate feature amount into an estimation model.

[0160] (Appendix 2) The process of identifying the object is characterized in that the inference program according to Appendix 1 is obtained by inputting the mask image, the intermediate feature amount, and the coordinate feature amount with coordinate values arranged in an image plane into the estimation model to identify the object.

[0161] (Appendix 3) Using the background image of the area where the object is to be placed and the target image of the object and the area as input data, learning data is obtained with the area of the object that does not exist in the background image but exists in the target image and the position of the object included in the target image as correct data. Based on the learning data, by inputting the background image and the target image, a feature extraction model that outputs an intermediate feature amount, a generation model that outputs a mask image indicating the area of the object that does not exist in the background image but exists in the target image by inputting the intermediate feature amount, and a model that inputs the mask image and the intermediate feature amount and outputs the position of the object that does not exist in the background image but exists in the target image. Machine learning for the estimation model is executed. A learning program characterized by causing a computer to execute a process.

[0162] (Appendix 4) The process of executing the machine learning is characterized in that the parameters of the feature extraction model, the parameters of the generation model, and the parameters of the estimation model are trained so that the difference between the area of the object output from the generation model and the area of the object in the correct data and the difference between the position of the object output from the estimation model and the correct data are reduced. The learning program according to Appendix 3.

[0163] (Appendix 5) The process of executing the machine learning is characterized in that the machine learning is executed by inputting the mask image, the intermediate feature amount, and the coordinate feature amount with coordinate values arranged in an image plane into the estimation model. The learning program according to Appendix 3 or 4.

[0164] (Appendix 6) Obtain the background image of the area where the object is to be placed and the target image of the object and the area. By inputting the background image and the target image into a feature extraction model, an intermediate feature amount is generated, By inputting the intermediate feature amount into a generation model, a mask image indicating a region of an object that exists in the target image but not in the background image is generated, By inputting the generated mask image and the intermediate feature amount into an estimation model, an object that exists in the target image but not in the background image is identified A reasoning method characterized in that a computer executes the process.

[0165] (Appendix 7) The process of identifying the object is characterized in that the object is identified by inputting the mask image, the intermediate feature amount, and a coordinate feature amount in which coordinate values are arranged in an image plane into the estimation model, according to the reasoning method described in Appendix 6.

[0166] (Appendix 8) A background image of an area where an object is to be placed is taken, and a target image of the object and the area is taken as input data, and learning data is obtained with the area of the object that exists in the target image but not in the background image and the position of the object included in the target image as correct answer data, Based on the learning data, by inputting the background image and the target image, a feature extraction model that outputs an intermediate feature amount, a generation model that outputs a mask image indicating a region of an object that exists in the target image but not in the background image by inputting the intermediate feature amount, and a mask image and the intermediate feature amount are input, and machine learning for an estimation model that outputs the position of an object that exists in the target image but not in the background image is executed A learning method characterized in that a computer executes the process.

[0167] (Appendix 9) The process of executing the machine learning trains the parameters of the feature extraction model, the parameters of the generation model, and the parameters of the estimation model so that the difference between the region of the object output from the generation model and the region of the object in the correct data, and the difference between the position of the object output from the estimation model and the correct data are reduced. The learning method according to Appendix 8, characterized in that.

[0168] (Appendix 10) The process of executing the machine learning is characterized in that the mask image, the intermediate feature amount, and the coordinate feature amount in which coordinate values are arranged in an image plane are input to the estimation model to execute the machine learning. The learning method according to Appendix 8 or 9.

Explanation of Signs

[0169] 50a, 50b Feature extraction unit 51a, 51b, 61 Synthesis unit 52 Estimation unit 53 Position coordinate feature amount output unit 54 Mask generation unit 60a, 60b Error calculation unit 62 Weight update value calculation unit 100, 200, 300 Information processing device 110, 210 Communication unit 120, 220 Input unit 130, 230 Display unit 140, 240 Storage unit 141 Learning data 142 Image table 143 Feature extraction model 144 Generation model 145 Estimation model 150, 250 Control unit 151, 251 Acquisition unit 152, 252 Learning processing unit 153, 253 Inference processing unit

Claims

1. Obtain a background image of the area where the object is to be placed and a target image of the object and the area, By inputting the background image and the target image into a feature extraction model, generate intermediate feature quantities, By inputting the intermediate feature quantities into a generation model, generate a mask image indicating the area of the object that does not exist in the background image but exists in the target image, By inputting the generated mask image and the intermediate feature quantities into an estimation model, identify the object that does not exist in the background image but exists in the target image A reasoning program characterized by causing a computer to execute the process.

2. The process of identifying the object is characterized by inputting the mask image, the intermediate feature quantities, and coordinate feature quantities with coordinate values arranged in the image plane into the estimation model to identify the object, as described in Claim 1 of the reasoning program.

3. Use the background image of the area where the object is to be placed and the target image of the object and the area as input data, and obtain training data with the area of the object that does not exist in the background image but exists in the target image and the position of the object included in the target image as correct data, Based on the training data, execute machine learning on a feature extraction model that outputs intermediate feature quantities by inputting the background image and the target image, a generation model that outputs a mask image indicating the area of the object that does not exist in the background image but exists in the target image by inputting the intermediate feature quantities, and an estimation model that outputs the position of the object that does not exist in the background image but exists in the target image by inputting the mask image and the intermediate feature quantities A learning program characterized by causing a computer to execute the process.

4. The process of executing the machine learning is characterized by training the parameters of the feature extraction model, the parameters of the generation model, and the parameters of the estimation model so that the difference between the area of the object output from the generation model and the area of the object in the correct data and the difference between the position of the object output from the estimation model and the correct data become small, as described in Claim 3 of the learning program.

5. The process of executing the machine learning is characterized in that the mask image, the intermediate feature amount, and the coordinate feature amount in which coordinate values are arranged in an image plane are input to the estimation model to execute the machine learning. The learning program according to claim 3 or 4.

6. Obtain a background image of the area where the object is to be placed and a target image of the object and the area being photographed, By inputting the background image and the target image into a feature extraction model, an intermediate feature amount is generated. By inputting the intermediate feature amount into a generation model, a mask image indicating the area of the object that does not exist in the background image but exists in the target image is generated. By inputting the generated mask image and the intermediate feature amount into an estimation model, the object that does not exist in the background image but exists in the target image is identified. A reasoning method characterized in that a computer executes the process.

7. Using the background image of the area where the object is to be placed and the target image of the object and the area being photographed as input data, and using the area of the object that does not exist in the background image but exists in the target image and the position of the object included in the target image as correct answer data, learning data is obtained. Based on the learning data, by inputting the background image and the target image, a feature extraction model that outputs an intermediate feature amount, by inputting the intermediate feature amount, a generation model that outputs a mask image indicating the area of the object that does not exist in the background image but exists in the target image, and by inputting the mask image and the intermediate feature amount, an estimation model that outputs the position of the object that does not exist in the background image but exists in the target image are subjected to machine learning. A learning method characterized in that a computer executes the process.

Citation Information

Patent Citations

  • Image processing apparatus, learning apparatus, image processing method, learning method, image processing program, and learning program

    JP2019153057A

  • Model generating device, estimating device, model generating method, and model generating program

    JP2021082155A

  • Left Behind Object Detection

    JP2022530299A

  • Image survellance apparatus applied with moving-path tracking technique using multi camera

    US20210314531A1

  • Left-behind subject detection

    WO2021189641A1