Processing device, processing method, and program

The processing device accurately recognizes products held by customers by analyzing images from multiple cameras, ensuring consistent position and product type information across different views, thereby addressing the need for accurate product recognition in store systems without checkout processing.

JP7679865B2Active Publication Date: 2025-05-20NEC CORP
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
JP2023201508
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2023-11-29
Publication Date
2025-05-20
Estimated Expiration
2040-05-22

AI Technical Summary

Technical Problem

There is a demand for technology that can accurately recognize products held by customers in store systems that do not require payment processing at checkout counters, and this technology is also useful for customer behavior analysis in stores for marketing research.

Method used

A processing device and method that acquire multiple images of a product from different angles using multiple cameras, detect objects in each image, generate position and product type information, extract sets of objects where position and product type information are consistent across different cameras, and output product recognition results for these sets.

Benefits of technology

This technology enables accurate recognition of products held by customers, reducing erroneous recognition and improving the accuracy of product identification, which is beneficial for both automated payment systems and customer behavior analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007679865000001
    Figure 0007679865000001
  • Figure 0007679865000002
    Figure 0007679865000002
  • Figure 0007679865000003
    Figure 0007679865000003
Patent Text Reader

Abstract

To accurately recognize a commodity.SOLUTION: A processing device (10) is provided which has: an acquisition unit (11) which acquires multiple images generated by multiple cameras imaging a product from different directions; a detection unit (12) which detects an object in each of multiple images; a position information generation unit (13) which generates position information indicating a position of each detected object in the images; a product type-related information generation unit (14) which, on the basis of the images, generates product type-related information for each detected object that identifies a product type; an extraction unit (15) which extracts a set of multiple objects detected from the images generated by the different cameras, wherein pieces of position information of the multiple objects mutually satisfy position conditions and pieces of product type-related information of the multiple objects mutually satisfy product type conditions; and a product recognition result output unit (16) which outputs a product recognition result for each extracted set.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to a processing device, a processing method, and a program. [Background technology]

[0002] Non-Patent Documents 1 and 2 disclose a store system that eliminates payment processing (such as product registration and payment) at the cash register counter. This technology recognizes the products held by customers based on images generated by a camera that photographs the inside of the store, and automatically processes the payment based on the recognition results when the customer leaves the store.

[0003] Patent Document 1 discloses the following device. First, the device detects a first flying object in a first image acquired from a first camera, and obtains an epipolar line indicating the direction of the first flying object as seen from the first camera. Then, the device controls a second camera to capture an image along the epipolar line. Next, the device detects a second flying object in a second image acquired from a second camera, determines whether the first flying object and the second flying object are the same, and calculates the positions of the first flying object and the second flying object.

[0004] Patent document 2 discloses a technology that accurately obtains the three-dimensional position of an object regardless of the number of cameras capturing the object by switching the method of estimating the three-dimensional position of a person depending on the position of the person's head in images obtained from multiple cameras. [Prior art documents] [Patent documents]

[0005] [Patent Document 1] JP 2018-195965 A [Patent Document 2] JP 2017-103602 A [Non-patent literature]

[0006] [Non-Patent Document 1] Takuya Miyata, "Amazon Go's mechanism: a cashierless supermarket realized using cameras and microphones," [online], December 10, 2016, [Retrieved December 6, 2019], Internet<URL:https: / / www.huffingtonpost.jp / tak-miyata / amazon-go_b_13521384.html> [Non-Patent Document 2] "NEC opens cashierless store "NEC SMART STORE" in headquarters--Facial recognition used, payment is made simultaneously when leaving the store", [online], February 28, 2020, [Retrieved March 27, 2020], Internet<URL: https: / / japan.cnet.com / article / 35150024 / > Summary of the Invention [Problem to be solved by the invention]

[0007] There is a demand for technology that can accurately recognize products that customers have picked up. For example, in a store system that does not require payment processing (product registration and payment, etc.) at a checkout counter as described in Non-Patent Documents 1 and 2, technology that can accurately recognize products that customers are holding in their hands is required. In addition, this technology is also useful when investigating customer behavior in a store for the purpose of customer preference research, marketing research, etc.

[0008] An object of the present invention is to provide a technology for accurately recognizing a product held by a customer. [Means for solving the problem]

[0009] According to the present invention, An acquisition means for acquiring a plurality of images generated by photographing a product from different directions using a plurality of cameras; A detection means for detecting an object from each of the plurality of images; a position information generating means for generating position information indicating a position within the image for each of the detected objects; a product type related information generating means for generating product type related information for identifying a product type for each of the detected objects based on the image; an extraction means for extracting a set of a plurality of objects detected from the images generated by the cameras different from each other, the set of the plurality of objects being such that the position information of the plurality of objects mutually satisfies a position condition and the product type related information of the plurality of objects mutually satisfies a product type condition; a product recognition result output means for outputting a product recognition result for each of the extracted sets; A processing apparatus is provided having:

[0010] Further, according to the present invention, The computer Acquire multiple images of the product by using multiple cameras to capture images of the product from different angles, Detecting an object from each of the plurality of images; generating position information for each of the detected objects that indicates a position within the image; generating product type related information for identifying a product type for each of the detected objects based on the image; A set of a plurality of objects detected from the images generated by the cameras different from each other, the position information of the plurality of objects satisfying a position condition with respect to each other and the product type related information of the plurality of objects satisfying a product type condition with respect to each other, is extracted; A processing method for outputting a product recognition result for each of the extracted sets is provided.

[0011] Further, according to the present invention, Computer, An acquisition means for acquiring a plurality of images generated by photographing the product from different directions using a plurality of cameras; a detection means for detecting an object from each of the plurality of images; a position information generating means for generating position information indicating a position within the image for each of the detected objects; a product type related information generating means for generating product type related information for identifying a product type for each of the detected objects based on the image; an extraction means for extracting a set of a plurality of objects detected from the images generated by the cameras different from each other, the set of the position information of the plurality of objects satisfying a position condition with respect to each other and the set of the product type related information of the plurality of objects satisfying a product type condition with respect to each other; a product recognition result output means for outputting a product recognition result for each of the extracted sets; A program is provided to function as a Effect of the Invention

[0012] According to the present invention, a technology is realized that can accurately recognize a product held by a customer. [Brief description of the drawings]

[0013] [Figure 1] FIG. 2 is a diagram illustrating an example of a hardware configuration of a processing device according to the present embodiment. [Diagram 2] FIG. 2 is a functional block diagram of a processing apparatus according to an embodiment of the present invention; [Diagram 3] FIG. 2 is a diagram for explaining an example of camera installation according to the present embodiment. [Figure 4] FIG. 2 is a diagram for explaining an example of camera installation according to the present embodiment. [Diagram 5] FIG. 2 is a diagram showing an example of an image processed by the processing device of the present embodiment. [Figure 6] FIG. 4 is a diagram illustrating an example of information processed by the processing device of the present embodiment. [Figure 7] FIG. 4 is a diagram illustrating an example of information processed by the processing device of the present embodiment. [Figure 8] 4 is a flowchart showing an example of a processing flow of the processing device of the present embodiment. [Figure 9] FIG. 4 is a diagram illustrating an example of information processed by the processing device of the present embodiment. [Figure 10] FIG. 4 is a diagram illustrating an example of information processed by the processing device of the present embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0014] <First embodiment> "overview" In this embodiment, product recognition processing is performed based on images generated by multiple cameras that capture products held by a customer from different positions and directions. Then, product recognition results are output only for objects for which the analysis results of the images generated by the multiple cameras are consistent (positions are consistent, appearances are consistent, etc.), and other product recognition results are, for example, rejected. According to the processing device of this embodiment, by adding such a condition that "if the analysis results of the images generated by the multiple cameras are consistent (positions are consistent, appearances are consistent, etc.), it is considered to be true," it is possible to suppress erroneous recognition and accurately recognize products held by a customer.

[0015] "Hardware Configuration" Next, an example of the hardware configuration of the processing device will be described.

[0016] Each functional part of the processing device is realized by any combination of hardware and software, centered on the CPU (Central Processing Unit) of any computer, memory, programs loaded into the memory, a storage unit such as a hard disk that stores the programs (in addition to programs that are stored before the device is shipped, it can also store programs downloaded from storage media such as CDs (Compact Discs) or servers on the Internet), and a network connection interface.Those skilled in the art will understand that there are various variations in the methods and devices for realizing the above.

[0017] FIG. 1 is a block diagram illustrating a hardware configuration of a processing device. As shown in FIG. 1, the processing device has a processor 1A, a memory 2A, an input / output interface 3A, a peripheral circuit 4A, and a bus 5A. The peripheral circuit 4A includes various modules. The processing device does not have to have the peripheral circuit 4A. The processing device may be composed of multiple devices that are physically and / or logically separated, or may be composed of a single device that is physically and / or logically integrated. When the processing device is composed of multiple devices that are physically and / or logically separated, each of the multiple devices can have the above hardware configuration.

[0018] The bus 5A is a data transmission path for the processor 1A, memory 2A, peripheral circuit 4A, and input / output interface 3A to transmit and receive data to each other. The processor 1A is an arithmetic processing device such as a CPU or a GPU (Graphics Processing Unit). The memory 2A is a memory such as a RAM (Random Access Memory) or a ROM (Read Only Memory). The input / output interface 3A includes an interface for acquiring information from an input device, an external device, an external server, an external sensor, a camera, etc., and an interface for outputting information to an output device, an external device, an external server, etc. Examples of the input device include a keyboard, a mouse, a microphone, a physical button, a touch panel, etc. Examples of the output device include a display, a speaker, a printer, a mailer, etc. The processor 1A can issue commands to each module and perform calculations based on the results of those calculations.

[0019] "Function Configuration" 2 shows an example of a functional block diagram of the processing device 10. As shown in the figure, the processing device 10 includes an acquisition unit 11, a detection unit 12, a position information generation unit 13, a product type related information generation unit 14, an extraction unit 15, and a product recognition result output unit 16.

[0020] The acquisition unit 11 acquires a plurality of images generated by a plurality of cameras capturing images of an item held by a customer from different positions and directions. Images generated by the plurality of cameras at the same time can be identified by a timestamp or the like. Images may be input to the acquisition unit 11 by real-time processing or batch processing. Which processing is used can be determined depending on, for example, the usage of the item recognition result.

[0021] Here, the multiple cameras are described. In this embodiment, multiple cameras (two or more cameras) are installed so that products held by customers can be photographed from multiple positions and multiple directions. For example, multiple cameras may be installed for each product display shelf at positions and orientations that photograph products taken from each shelf. The cameras may be installed on the product display shelves, on the ceiling, on the floor, on the wall, or in other locations. Note that the example of installing a camera on each product display shelf is merely one example, and is not limited to this.

[0022] The camera may capture moving images continuously (e.g., during business hours), or it may continuously capture still images at time intervals greater than the frame interval of the moving images, or it may perform these captures only while a motion sensor or the like detects a person in a specified position (such as in front of a product display shelf).

[0023] Here, an example of camera installation is shown. Note that the camera installation example described here is merely an example, and is not limited to this. In the example shown in Fig. 3, two cameras 2 are installed for each product display shelf 1. Fig. 4 is a diagram extracting the frame 4 from Fig. 3. Each of the two parts that make up the frame 4 is provided with a camera 2 and lighting (not shown).

[0024] The light emission surface of the lighting extends in one direction, and has a light-emitting unit and a cover that covers the light-emitting unit. The lighting mainly emits light in a direction perpendicular to the extension direction of the light-emitting surface. The light-emitting unit has a light-emitting element such as an LED, and emits light in a direction that is not covered by the cover. When the light-emitting element is an LED, multiple LEDs are lined up in the direction in which the lighting extends (up and down in the figure).

[0025] Camera 2 is provided at one end of the linearly extending frame 4 component, and its imaging range is the direction in which the illumination light is emitted. For example, in the frame 4 component on the left side of Fig. 4, camera 2's imaging range is downward and diagonally downward to the right. Also, in the frame 4 component on the right side of Fig. 4, camera 2's imaging range is upward and diagonally upward to the left.

[0026] As shown in FIG. 3, the frame 4 is attached to the front frame (or the front of both side walls) of the product display shelf 1 that constitutes the product placement space. One of the components of the frame 4 is attached to one of the front frames in a direction in which the camera 2 is located downward. The other of the components of the frame 4 is attached to the other of the front frames in a direction in which the camera 2 is located upward. The camera 2 attached to one of the components of the frame 4 takes pictures above and diagonally upward so that the opening of the product display shelf 1 is included in the shooting range. On the other hand, the camera 2 attached to the other of the components of the frame 4 takes pictures below and diagonally downward so that the opening of the product display shelf 1 is included in the shooting range. By configuring in this way, the two cameras 2 can shoot the entire range of the opening of the product display shelf 1. As a result, it is possible to shoot the product (the product held by the customer) being taken out of the product display shelf 1 with the two cameras 2.

[0027] For example, when the configurations shown in Figures 3 and 4 are adopted, an item held by a customer is photographed by two cameras 2 as shown in Figure 5. As a result, two images 7 and 8 of the item are generated, which are photographs of the item from multiple positions and multiple directions.

[0028] In the following, it is assumed that "the product held by the customer is photographed with two cameras." Then, at the end of this embodiment, as a modified example, a configuration in which "the product held by the customer is photographed with three or more cameras" will be described.

[0029] Returning to FIG. 2, the detection unit 12 detects an object by detecting an area where the object exists from each of the multiple images generated by the multiple cameras. The technology for detecting an area where an object exists from an image is widely known, so a detailed description will be omitted here. The detected "area where the object exists" may be a rectangular area including the object and its periphery, or an area with a shape that follows the contour of the object where only the object exists. For example, when a method is adopted for determining whether an object exists for each rectangular area in an image, the detected "area where the object exists" is a rectangular area W including the object and its periphery, as shown in FIG. 5. On the other hand, when a method for detecting a pixel area where a detection target exists, called semantic segmentation or instance segmentation, is adopted, the detected "area where the object exists" is an area with a shape that follows the contour of the object where only the object exists.

[0030] The position information generating unit 13 generates position information indicating a position in the image for each detected object. The position information is indicated, for example, by coordinates in a two-dimensional coordinate system set on the image. The position information may indicate a certain area in the image, or may indicate a point in the image. The position information indicating a certain area in the image may indicate, for example, an area in which the above-mentioned object exists. The position information indicating a point in the image may indicate, for example, a representative point (center point, center of gravity point, etc.) in the area in which the above-mentioned object exists.

[0031] The product type related information generating unit 14 generates product type related information for identifying the product type for each detected object based on the image. The product type related information in this embodiment is product type identification information (product name, product code, etc.) for identifying multiple product types from each other.

[0032] Techniques for recognizing the product type of an object included in an image are widely known, and any technique can be adopted in this embodiment. For example, the product type related information generating unit 14 may recognize the product type of the object based on a classifier previously generated by machine learning or the like and the image of the "area where the object exists". Alternatively, the product type related information generating unit 14 may recognize the product type of the object by pattern matching that compares a template image of the appearance of each product prepared in advance with the image of the "area where the object exists".

[0033] The acquiring unit 11, the detecting unit 12, the position information generating unit 13, and the product type related information generating unit 14 described above generate information as shown in FIG. 6 and FIG.

[0034] The first object information shown in Fig. 6 indicates the position information and product type related information of each of the multiple objects detected from the image generated by the first camera. In the figure, "1-01" and "1-02" are serial numbers for identifying the multiple objects detected from the image.

[0035] The second object information shown in Fig. 7 indicates the position information and product type related information of each of the multiple objects detected from the image generated by the second camera. In the figure, "2-01" and "2-02" are serial numbers for identifying the multiple objects detected from the image.

[0036] Returning to FIG. 2, the extraction unit 15 extracts a set of a plurality of objects detected from images generated by different cameras, the position information of which satisfies a position condition for each object and the product type related information of which satisfies a product type condition for each object. In the case of an example of "taking pictures of a product held by a customer with two cameras", the extraction unit 15 extracts a pair of a first object detected from an image generated by a first camera and a second object detected from an image generated by a second camera, the position information of which satisfies a position condition for each object and the product type related information of which satisfies a product type condition for each object. The extraction unit 15 performs the extraction process based on information as shown in FIG. 6 and FIG. 7.

[0037] The extraction by extraction unit 15 means extraction of an object where the analysis result of the image generated by the first camera and the analysis result of the image generated by the second camera match (match in position, match in appearance, etc.).

[0038] First, the position condition is described below: The position condition is that the position of a first object in an image generated by a first camera and the position of a second object in an image generated by a second camera satisfy the positional relationship when the first object and the second object are the same subject.

[0039] An example of the position condition is that the "position where the first object may exist in the three-dimensional space" estimated from the "setting information of the first camera" and the "position of the first object in the image generated by the first camera" and the "position where the second object may exist in the above three-dimensional space" estimated from the "setting information of the second camera", the "position of the second object in the image generated by the second camera", and the "relative relationship between the first camera and the second camera" are consistent (satisfying the positional relationship when the first object and the second object are the same subject). The details of the method of determining whether or not such a position condition is satisfied are not particularly limited, and any method can be adopted. An example will be described below, but is not limited to this.

[0040] For example, the use of epipolar lines can be considered. First, based on the setting information of the first camera (focal length, angle of view, etc.), the setting information of the second camera (focal length, angle of view, etc.), and the relative relationship between the first camera and the second camera (relative positional relationship, relative orientation relationship, etc.), a line (epipolar line) that passes through the first camera and a predetermined point in the image generated by the first camera and is projected onto the image generated by the second camera can be obtained. By setting the predetermined point based on the position of the first object in the image generated by the first camera, a position in the second image where the first object may exist can be obtained. If the second object exists at a position in the second image where the first object may exist, it can be determined that the first object and the second object satisfy the position condition (their positions in the images are consistent).

[0041] Next, the product type condition will be described. As described above, the product type related information in this embodiment is product type identification information specified based on the feature amount of the appearance of the object. The product type condition in this embodiment is that the product type identification information of the first object and the product type identification information of the second object match (the product type recognition results match).

[0042] Returning to Fig. 2, the product recognition result output unit 16 outputs the product recognition result (product type identification information) of the first object or the second object for each pair (set) extracted by the extraction unit 15. In the case of this embodiment, the pair (set) extracted by the extraction unit 15 satisfies the product type condition that "the product type identification information of the first object and the product type identification information of the second object match", and therefore the product recognition result of the first object and the product recognition result of the second object match.

[0043] In this embodiment, the contents of the subsequent processing for the commodity recognition result output by the commodity recognition result output unit 16 are not particularly limited.

[0044] For example, the product recognition result may be used in payment processing in a store system that does not require payment processing (product registration, payment, etc.) at a checkout counter as disclosed in Non-Patent Documents 1 and 2. An example will be described below.

[0045] First, the store system registers the output product recognition result (product type identification information) by linking it to information that identifies the customer holding the product. For example, a camera that captures the face of a customer holding a product in his / her hand may be installed in the store, and the store system may extract the external features of the customer's face from the image generated by the camera. Then, the store system may register the product type identification information of the product held by the customer and other product information (unit price, product name, etc.) by linking it to the external features of the face (information that identifies the customer). The other product information can be obtained from a product master (information that links the product type identification information with other product information) that is stored in advance in the store system.

[0046] Alternatively, the customer's customer identification information (membership number, name, etc.) may be linked to the features of the customer's facial appearance and registered in advance in any location (store system, center server, etc.). Then, when the store system extracts the features of the customer's facial appearance from an image including the face of the customer holding a product, the store system may identify the customer's customer identification information based on the previously registered information. Then, the store system may register the product type identification information of the product held by the customer and other product information in association with the identified customer identification information.

[0047] The store system also calculates the payment amount based on the registered details and executes the payment process. For example, the payment process is executed when the customer leaves the store through the gate or when the customer leaves the store through the exit. The detection of these timings may be realized by detecting the customer's exit from an image generated by a camera installed at the gate or exit, by inputting customer identification information of the customer leaving the store into an input device (such as a reader that performs short-range wireless communication) installed at the gate or exit, or by other methods. The details of the payment process may be a payment process using a credit card based on pre-registered credit card information, a payment based on pre-charged money, or other methods.

[0048] Other examples of the application of the product recognition results output by the product recognition result output unit 16 include customer preference surveys and marketing surveys. For example, by linking the products picked up by each customer to each customer and registering them, it is possible to analyze the products in which each customer is interested. In addition, by registering that the customer picked up each product, it is possible to analyze which products the customer is interested in. Furthermore, by estimating customer attributes (gender, age, nationality, etc.) using conventional image analysis technology and registering the attributes of the customers who picked up each product, it is possible to analyze what attributes of customers are interested in each product.

[0049] Next, an example of the flow of processing by the processing device 10 will be described with reference to the flowchart of FIG.

[0050] First, the acquisition unit 11 acquires two images captured by the first camera and the second camera at the same time (S10). The first camera and the second camera are installed so as to capture images of a product held by a customer from different positions and directions.

[0051] Next, the detection unit 12 analyzes each of these two images and detects objects from each image (S11). Next, the position information generation unit 13 generates position information indicating a position within the image for each object detected in S11 (S12). In addition, the product type related information generation unit 14 generates product type related information that specifies the product type for each detected object based on the image (S13). Note that the processing order of S12 and S13 is not limited to that shown in the figure.

[0052] By the processing up to this point, information as shown in Figures 6 and 7 is generated. The first object information shown in Figure 6 indicates the position information and product type related information of each of the multiple objects detected from the image generated by the first camera. The second object information shown in Figure 7 indicates the position information and product type related information of each of the multiple objects detected from the image generated by the second camera.

[0053] Next, the extraction unit 15 extracts pairs (sets) of a first object detected from an image generated by a first camera and a second object detected from an image generated by a second camera, where the position information of the first object and the second object satisfy a position condition and the product type related information of the second object satisfy a product type condition (S14).

[0054] Then, the commodity recognition result output unit 16 outputs the commodity recognition result (commodity type identification information) of the first object or the second object for each pair (set) extracted in S14 (S15).

[0055] "Action and effect" According to the processing device 10 of the present embodiment described above, it is possible to execute a commodity recognition process based on images generated by a plurality of cameras that capture a commodity held by a customer from different positions and directions. Then, it is possible to output only the commodity recognition results of objects with which the analysis results of the images generated by the plurality of cameras are consistent (matching in position, matching in appearance, etc.), and to discard, for example, the other commodity recognition results. The other commodity recognition results are the commodity recognition results of the first object and the second object that were not extracted by the extraction unit 15.

[0056] According to the processing device 10 of this embodiment, by adding the condition that "it is considered true if the analysis results of images generated by multiple cameras are consistent (positions are consistent, appearances are consistent, etc.)", it is possible to suppress erroneous recognition and accurately recognize products held by customers.

[0057] "Variations" As described above, in this embodiment, an item held by a customer may be photographed by three or more cameras from different positions and directions.

[0058] In this case, the processing device 10 outputs only the product recognition results of objects for which all the analysis results of images generated by N cameras (N is an integer equal to or greater than 3) are consistent (matching in position, matching in appearance, etc.), and may, for example, discard other product recognition results. In this case, the extraction unit 15 extracts a set of multiple objects detected from N images generated by N cameras, where the position information mutually satisfies a position condition, the product type related information mutually satisfies a product type condition, and the set includes N objects. This condition differs from the above-mentioned condition in that a condition regarding the number of objects (members) belonging to the set is further added.

[0059] Additionally, the processing device 10 may output only the product recognition results of objects for which at least M (M is an integer of 2 or more and M is less than N) analysis results among N analysis results of images generated by N cameras (N is an integer of 3 or more) are consistent (matching in position, matching in appearance, etc.), and may discard other product recognition results, for example. In this case, the extraction unit 15 extracts a set of multiple objects detected from N images generated by N cameras, where the position information mutually satisfies a position condition, the product type related information mutually satisfies a product type condition, and the set includes M or more objects. This condition differs from the above-mentioned condition in that a condition regarding the number of objects (members) belonging to the set is further added.

[0060] Additionally, the processing device 10 may output only the product recognition results of objects with which a predetermined percentage or more of the analysis results of images generated by N cameras (N is an integer equal to or greater than 3) are consistent (matching in position, matching in appearance, etc.), and may discard other product recognition results, for example. In this case, the extraction unit 15 extracts a set of multiple objects detected from N images generated by N cameras, where the position information mutually satisfies a position condition, the product type related information mutually satisfies a product type condition, and the set includes a predetermined percentage or more of N objects. This condition differs from the above-mentioned condition in that a condition regarding the number of objects (members) belonging to the set is further added.

[0061] The above-mentioned effects are also achieved in this modified example. In addition, by increasing the number of cameras and setting the above-mentioned conditions, even if a product is in a blind spot due to a human hand or something and some cameras are unable to capture the product, it can be considered true if the analysis results of the images generated by the other multiple cameras are consistent. As a result, convenience is further improved.

[0062] <Second embodiment> In this embodiment, the product type condition is different from that of the first embodiment. Specifically, the product type condition of this embodiment is that "the product types match" and "the relationship between the characteristic part of the product facing the first camera, which is identified based on the characteristic amount of the appearance of the object extracted from the image generated by the first camera, and the characteristic part of the product facing the second camera, which is identified based on the characteristic amount of the appearance of the object extracted from the image generated by the second camera, satisfies the orientation condition."

[0063] For example, as in the examples of Figures 3 to 5, when the first camera and the second camera capture images of a product sandwiched between them and the capture directions differ by approximately 180°, the orientation condition is "front and back relationship." In other words, the orientation condition is that the characteristic part of the product facing the first camera and the characteristic part of the product facing the second camera are in a front and back relationship on the product.

[0064] For example, feature amounts extracted from images captured from multiple directions are registered for each product type as shown in Fig. 9. Note that, although feature amounts for images captured from six directions (front, back, top, bottom, right, and left) are registered in Fig. 9, the number of image capturing directions is not limited to this.

[0065] In addition, the relationship between the shooting directions of the first camera and the second camera is registered as shown in Fig. 10. This relationship indicates the relationship of "when the first camera shoots a product from a certain direction, from which direction the second camera will shoot the product."

[0066] Then, the extraction unit 15 can determine whether or not the above-mentioned orientation condition is satisfied based on this information.

[0067] Specifically, first, the product type related information generating unit 14 identifies from which direction the characteristic part of the product photographed faces the first camera by comparing the feature amount of the object's appearance extracted from the image generated by the first camera with the feature amount shown in Fig. 9. Also, the product type related information generating unit 14 identifies from which direction the characteristic part of the product photographed faces the second camera by comparing the feature amount of the object's appearance extracted from the image generated by the second camera with the feature amount shown in Fig. 9. These identification processes may be realized by using a classifier generated by machine learning, by pattern matching, or by other methods.

[0068] Then, the extraction unit 15 determines that the above orientation condition is satisfied when the shooting direction in which the characteristic part is photographed facing the first camera and the shooting direction in which the characteristic part is photographed facing the second camera satisfy the relationship shown in Figure 10.

[0069] Other configurations of the processing device 10 of this embodiment are the same as those of the first embodiment. The processing device 10 of this embodiment can also adopt a modified example in which an item held by a customer is photographed from different positions and directions by three or more cameras. For example, if the relationship between the photographing directions of three or more cameras is registered in advance, the same action and effect can be achieved by the same process as above.

[0070] According to the processing device 10 of this embodiment, the same action and effect as in the first embodiment is realized. In addition, the processing device 10 of this embodiment further adds the orientation condition described above, taking into consideration the characteristic that "when a product is photographed with a plurality of cameras from different positions and directions, the characteristic parts of the product that appear in the image may differ depending on the direction from which the product is photographed." By adding the orientation condition, it is possible to further suppress erroneous recognition and more accurately recognize the product held by the customer.

[0071] <Third embodiment> In this embodiment, the product type related information is an external appearance feature of an object extracted from an image, and the product type condition is that the similarity of the external appearance feature is equal to or greater than a reference value.

[0072] Other configurations of the processing apparatus 10 of this embodiment are similar to those of the first embodiment. According to the processing apparatus 10 of this embodiment, the same effects as those of the first embodiment are achieved.

[0073] <Modification> Here, a modified example applicable to all the embodiments will be described. In the above embodiment, the position information generating unit 13 generates position information for each detected object, the product type related information generating unit 14 generates product type related information for each detected object, and then the extraction unit 15 extracts a set of multiple objects that satisfy the position condition and the product type condition.

[0074] In the first modification, after the position information generating unit 13 generates position information for each detected object, the extracting unit 15 extracts a set of multiple objects that satisfy a position condition. Then, the product type related information generating unit 14 judges whether the multiple objects belonging to the extracted set mutually satisfy the product type condition. Then, the extracting unit 15 extracts a set of multiple objects that are judged to satisfy the product type condition.

[0075] In this case, the product type related information generating unit 14 may execute a process of identifying the product type identification information of each of the plurality of objects based on the feature amount of the appearance of each of the objects. Then, the product type related information generating unit 14 may determine that a combination of objects whose identified product type identification information matches each other satisfies the product type condition.

[0076] As another processing example, the product type related information generating unit 14 may determine whether the appearance feature of another object matches the "appearance feature of the product identified by the product identification information of the identified first object" after identifying the product type identification information of the first object based on the appearance feature of the first object. If there is a match, the product type related information generating unit 14 may determine that the product type condition is satisfied. In this processing example, the process of identifying the product type identification information by matching with the feature of each of multiple product types is performed only on the first object, and not on the other objects. This reduces the processing load on the computer.

[0077] In this specification, "acquisition" includes at least one of the following: "the device itself goes to retrieve data stored in another device or storage medium (active acquisition)" based on user input or based on program instructions, such as receiving data by making a request or inquiry to another device, or accessing another device or storage medium and reading it, and "the device itself inputs data output from another device (passive acquisition)" based on user input or based on program instructions, such as receiving data that is distributed (or transmitted, push notification, etc.), and selecting and acquiring data from among the received data or information, and "editing data (converting it to text, rearranging data, extracting some data, changing the file format, etc.) to generate new data and acquiring the new data."

[0078] Although the present invention has been described above with reference to the embodiments (and examples), the present invention is not limited to the above-mentioned embodiments (and examples). Various modifications that can be understood by those skilled in the art can be made to the configuration and details of the present invention within the scope of the present invention.

[0079] A part or all of the above-described embodiments can be described as follows, but is not limited to the following. 1. An acquisition means for acquiring a plurality of images generated by photographing an item held by a customer from different angles using a plurality of cameras; A detection means for detecting an object from each of the plurality of images; a position information generating means for generating position information indicating a position within the image for each of the detected objects; a product type related information generating means for generating product type related information for identifying a product type for each of the detected objects based on the image; an extraction means for extracting a set of a plurality of objects detected from the images generated by the cameras different from each other, the set of the plurality of objects being such that the position information of the plurality of objects mutually satisfies a position condition and the product type related information of the plurality of objects mutually satisfies a product type condition; a product recognition result output means for outputting a product recognition result for each of the extracted sets; A processing device having 2. The processing device according to 1, wherein the position condition is that the position of the object in the image satisfies a positional relationship that would occur if the objects in the image were the same subject. 3. The processing device described in 2, wherein the position condition is that a position where the first object may exist in a three-dimensional space estimated from setting information of a first camera and the position of the first object in the image generated by the first camera, and a position where the second object may exist in the three-dimensional space estimated from setting information of another camera, the position of a second object in the image generated by the other camera, and the relative relationship between the first camera and the other camera, satisfy a positional relationship when the first object and the second object are the same subject. 4. The product type related information is a feature quantity of the appearance of the object extracted from the image, 4. The processing device according to any one of 1 to 3, wherein the product type condition is that the similarity of the appearance feature amount is equal to or greater than a reference value. 5. The product type related information is product type identification information identified based on the feature amount of the appearance of the object extracted from the image, 4. The processing device according to any one of 1 to 3, wherein the product type condition is that the product type identification information matches. 6. The above product type conditions are: The product type matches, and The processing device described in 5, wherein a relationship between a characteristic part of a product facing a first camera identified based on the appearance features of the object extracted from the image generated by the first camera and a characteristic part of a product facing another camera identified based on the appearance features of the object extracted from the image generated by the other camera satisfies an orientation condition. 7. The computer: Acquire multiple images of the product being held by the customer, which are taken from different angles by multiple cameras. Detecting an object from each of the plurality of images; generating position information for each of the detected objects that indicates a position within the image; generating product type related information for identifying a product type for each of the detected objects based on the image; A set of a plurality of objects detected from the images generated by the cameras different from each other, the position information of the plurality of objects satisfying a position condition with respect to each other and the product type related information of the plurality of objects satisfying a product type condition with respect to each other, is extracted; A processing method for outputting a product recognition result for each of the extracted sets. 8. The computer An acquisition means for acquiring a plurality of images generated by photographing an item held by a customer from different directions using a plurality of cameras; a detection means for detecting an object from each of the plurality of images; a position information generating means for generating position information indicating a position within the image for each of the detected objects; a product type related information generating means for generating product type related information for identifying a product type for each of the detected objects based on the image; an extraction means for extracting a set of a plurality of objects detected from the images generated by the cameras different from each other, the set of the position information of the plurality of objects satisfying a position condition with respect to each other and the set of the product type related information of the plurality of objects satisfying a product type condition with respect to each other; a product recognition result output means for outputting a product recognition result for each of the extracted sets; A program that functions as a

Claims

1. An acquisition means for acquiring a plurality of images generated by photographing an item held by a customer from different angles using a plurality of cameras; A detection means for detecting an object from each of the plurality of images; a position information generating means for generating position information indicating a position within the image for each of the detected objects; a product type related information generating means for generating product type related information for identifying a product type for each of the detected objects based on the image; an extraction means for extracting, as a set of the same objects, a set of a plurality of objects detected from the images, the sets of the objects being detected from the images generated by different cameras, the positional relationship between the positions of the objects in the images satisfying a position condition, and the product type condition that the product types of the objects are the same; a product recognition result output means for outputting a product recognition result for each of the extracted sets; A processing device having

2. The processing device according to claim 1 , wherein the position condition is that the position of the object in the image satisfies a positional relationship that would occur if the objects in the image were a single subject.

3. 3. The processing device of claim 2, wherein the position condition is that a position where the first object may exist in a three-dimensional space estimated from setting information of a first camera and the position of the first object in the image generated by the first camera, and a position where the second object may exist in the three-dimensional space estimated from setting information of another camera, the position of a second object in the image generated by the other camera, and the relative relationship between the first camera and the other camera, satisfy a positional relationship when the first object and the second object are the same subject.

4. the product type related information is an external appearance feature of the object extracted from the image, The processing device according to claim 1 , wherein the product type condition is that a similarity of the appearance feature amount is equal to or greater than a reference value.

5. the product type related information is product type identification information identified based on a feature amount of an appearance of the object extracted from the image, The processing device according to claim 1 , wherein the product type condition is that the product type identification information matches.

6. The product type condition is: The product type matches, and The processing device according to claim 5, wherein a relationship between a characteristic part of a product facing a first camera identified based on the appearance features of the object extracted from the image generated by the first camera and a characteristic part of a product facing another camera identified based on the appearance features of the object extracted from the image generated by the other camera satisfies an orientation condition.

7. The computer Acquire multiple images of the product being held by the customer, which are taken from different angles by multiple cameras. Detecting an object from each of the plurality of images; generating position information for each of the detected objects that indicates a position within the image; generating product type related information for identifying a product type for each of the detected objects based on the image; extracting, as a set of the same objects, a set of a plurality of objects detected from the images, the sets of the objects being detected from the images generated by different cameras, the positional relationship between the positions of the objects in the images satisfying a position condition, and the sets of the objects satisfying a product type condition that the product types of the objects are the same; A processing method for outputting a product recognition result for each of the extracted sets.

8. Computer, An acquisition means for acquiring a plurality of images generated by photographing an item held by a customer from different angles using a plurality of cameras; a detection means for detecting an object from each of the plurality of images; a position information generating means for generating position information indicating a position within the image for each of the detected objects; a product type related information generating means for generating product type related information for identifying a product type for each of the detected objects based on the image; an extraction means for extracting, as a set of the same objects, a set of a plurality of objects detected from the images, the sets of the objects being detected from the images generated by different cameras, the positional relationship between the positions of the objects in the images satisfying a position condition, and the product type condition that the product types of the objects are the same; a product recognition result output means for outputting a product recognition result for each of the extracted sets; A program that functions as a

Citation Information

Patent Citations

  • Information processor and program

    JP2013054673A

  • Merchandise identification system

    JP2015210651A

  • Article interaction and movement detection method

    JP2016532932A

  • Position detection device, and position detection method and program

    JP2017103602A

  • Flying object position detection apparatus, flying object position detection system, flying object position detection method, and program

    JP2018195965A