Foreign object detection system, control method for foreign object detection system, electronic equipment, imaging system, imaging method, computer program and storage medium
The foreign object detection system enhances accuracy by capturing images of containers in multiple postures and adjusting training data, effectively detecting foreign objects in diverse bottle shapes.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- CANON MARKETING JAPAN INC
- Filing Date
- 2022-02-28
- Publication Date
- 2026-05-21
AI Technical Summary
Conventional foreign object detection methods in liquid containers suffer from low detection accuracy and are limited to cylindrical shapes, often misidentifying non-foreign objects as foreign objects and failing to detect certain types of foreign objects effectively.
A foreign object detection system that captures images of containers in multiple postures, using machine learning to differentiate between foreign objects and non-foreign objects by adjusting the number of images in training data to enhance detection accuracy, particularly for floating and falling objects.
The system achieves high-accuracy detection of foreign objects in liquid containers, improving differentiation between foreign objects and non-foreign objects, and is applicable to various bottle shapes beyond cylindrical.
Smart Images

Figure 0007863431000012 
Figure 0007863431000013 
Figure 0007863431000014
Abstract
Description
Technical Field
[0001] The present invention relates to a foreign object detection system for detecting foreign objects in a liquid container, a control method of the foreign object detection system, an electronic device, a photographing system, a photographing method, a computer program, a storage medium, and the like.
Background Art
[0002] Conventionally, as a method for inspecting whether there is a foreign object in a container (bottle) filled with a liquid, a method of detecting a foreign object based on an image obtained by imaging the container filled with the liquid is known.
[0003] In Patent Document 1, it is described that using an infrared light, a CCD sensor, etc., it is inspected whether foreign objects (floating foreign objects, precipitating foreign objects) such as metal, cloth, hair, dust, etc. are mixed in a bottle or a PET bottle filled with a liquid, using a partial enlargement function by a conical prism.
[0004] Also, in Patent Document 2, an inspection device having a rotation mechanism for placing a test bottle filled with a beverage liquid on a turntable and rotating it and an inclination mechanism for inclining the bottle is described. When the test bottle is inclined, a transmission image including a shadow image of a heavy foreign object floating in the beverage liquid is obtained, the test bottle is rotated, and when the rotation is stopped thereafter, a transmission image including a shadow image of a light foreign object floating in the beverage liquid is obtained. Further, it is described that the transmission images of the test bottle each including the shadow image of the heavy foreign object and the shadow image of the light foreign object are processed to detect foreign objects in the beverage liquid.
Prior Art Documents
Patent Documents
[0005]
Patent Document 1
Patent Document 2
Summary of the Invention
Problems to be Solved by the Invention
[0006] However, the conventional techniques disclosed in Patent Documents 1 and 2 lack sufficient detection accuracy, sometimes failing to detect foreign objects in containers or misidentifying non-foreign objects as foreign objects. For example, even though bubbles appearing inside a container are not foreign objects, the fine bubbles captured in the image may be mistakenly detected as foreign objects.
[0007] Furthermore, in Patent Document 2, the container is rotated on an upright central axis to levitate foreign matter, which limits the container shape to a cylindrical shape, making it difficult to apply to bottle shapes other than cylindrical.
[0008] Therefore, the present invention aims to provide a foreign matter detection system, etc., capable of detecting foreign matter in a liquid container with high accuracy. [Means for solving the problem]
[0009] A foreign object detection system according to one aspect of the present invention is: The acquisition means captures images of the container when its posture is changed between a first posture in which a specific part of the container containing liquid is facing upward relative to the direction of gravity, and a second posture in which the specific part is facing downward relative to the direction of gravity, as the image to be detected. A control means controls the input of the image acquired by the acquisition means to a trained model, and controls the notification of the foreign object detection result based on the estimation result information obtained as a result of the input. The system includes a generation means for generating a trained model for estimating whether or not there is a foreign object in a container by performing machine learning using a set of images consisting of multiple images obtained by photographing the containers containing the liquid when they are changed between a first and a second posture, as training data. The generating means further includes an adjusting means that adjusts the number of images in a first image group consisting of multiple images obtained by photographing the containers containing the liquid when they are changed between a first posture and a second posture, and generates a second image group in which the number of second type images in which the foreign object is floating or rising in the liquid is greater than the number of first type images in which the foreign object is falling from top to bottom. The generation means performs the machine learning using the second group of images as training data. It is characterized by the following: [Effects of the Invention]
[0010] According to the present invention, it is possible to realize a foreign object detection system that can detect foreign objects in a liquid container with high accuracy. [Brief explanation of the drawing]
[0011] [Figure 1] Figure 1(A) is a conceptual diagram of the learning phase in the foreign object detection system, and Figure 1(B) is a conceptual diagram of the estimation phase in the foreign object detection system. [Figure 2] It is a configuration diagram of the foreign object detection system. [Figure 3] It is a diagram showing an example of the hardware configuration of information processing devices such as the client terminal 201, the data collection server 211, the learning server 212, and the estimation server 213. [Figure 4] It is a flowchart of a part of the learning process performed by the client terminal 201. [Figure 5] It is a flowchart of the remaining part of the learning process performed by the client terminal 201. [Figure 6] Figure 6(A) is a flowchart of the estimated image transmission process executed by the client terminal 201, and Figure 6(B) is a flowchart of the estimation process executed by the estimation server 213. [Figure 7] It is a diagram for explaining an example of the arrangement of the camera 202 and the bottle 203 in the imaging system example 1. [Figure 8] It is a diagram showing the state where the entire jig housing 204 is smoothly rotated 90° clockwise about an axis perpendicular to the paper surface. [Figure 9] It is a diagram showing the state where the entire jig housing 204 is further slowly and smoothly rotated clockwise from the state of Figure 8, for example, by about 20° to 60°. [Figure 10] In Figure 9, the part including the camera 202, the foreign objects 215 and 216, the bottle 203, and the LED illumination 205 is extracted and shown with the bottle facing downward. [Figure 11] It is a diagram for explaining the image 215i of the foreign object 215 reduced and imaged on the sensor surface 217 of the camera. [Figure 12] It is a schematic diagram showing the virtual image of the foreign object when the container is a round bottle. [Figure 13] It is a diagram for explaining the imaging relationship when the bottle 203 is a round bottle. [Figure 14]It is a diagram for explaining the imaging relationship when the bottle 203 in the imaging system example 2 is a square bottle. [Figure 15] (A) to (C) are diagrams for explaining the depth of focus and the allowable blur amount (allowable circle of confusion). [Figure 16] It is a diagram with the horizontal axis representing the diameter of the allowable circle of confusion in units of sensor pixels and the vertical axis representing the contrast in the focused state normalized with 1. [Figure 17] It is a diagram for explaining the relationship between the depth of the field of view, the depth of focus, and the allowable circle of confusion (allowable blur amount) in the case of a round bottle container. [Figure 18] In imaging system example 1, it is a diagram showing Table 1 in which the part where the front depth of the field of view is r (38 mm) or more in radius and the rear depth of the field of view is 2r (76 mm) or more in bottle diameter is brightly displayed and the rest is shown in gray. [Figure 19] In imaging system example 3, it is a diagram showing Table 2 representing the front depth of the field of view and the rear depth of the field of view when the Fno of the lens is changed. [Figure 20] In imaging system example 2, it is a diagram showing Table 3 representing the relationship between the overall depth of the field of view combined with the front depth of the field of view and the rear depth of the field of view and the allowable circle of confusion when the aperture of the lens is changed. [Figure 21] (A) and (B) are diagrams for explaining the range where foreign objects can be observed from outside the container when illuminating a round bottle container with parallel illumination light 221. [Figure 22] (A) and (B) are diagrams showing the range where foreign objects can be imaged when illuminating a round bottle container with diffused light. [Figure 23] (A) and (B) are diagrams for explaining the diffused light illumination for illuminating a round bottle container suitable for imaging system example 1. [Figure 24] It is a flowchart of the imaging process.
Modes for Carrying Out the Invention
[0012] Embodiments of the present invention will be described below with reference to the drawings. However, the present invention is not limited to the following embodiments. In each drawing, the same reference numeral is used for the same member or element, and redundant explanations are omitted or simplified.
[0013] This embodiment describes a method for inspecting whether foreign matter is present in a container containing liquid (such as a bottle or PET bottle; hereinafter simply referred to as "bottle") by photographing the bottle containing the liquid. The liquid in this embodiment includes various forms such as beverages (alcoholic beverages such as wine and whiskey, soft drinks such as fruit juices, etc.), food products (liquid seasonings and soups, etc.), pharmaceuticals (pesticides / alcohol disinfectants, etc.), and hygiene and beauty products (shampoo / liquid soap / lotion, etc.).
[0014] For bottles containing various liquids, quality control is required to ensure that products with foreign objects inside the bottles are not supplied to consumers. If the liquid is manufactured and sealed in-house, strict quality control during the manufacturing process can minimize the number of bottles contaminated with foreign objects. However, for products imported with liquid sealed in bottles, it is necessary to inspect for foreign object contamination after import and before sale.
[0015] Traditionally, such inspections have mainly relied on visual inspection by skilled personnel. However, to address the shortage of skilled personnel and the increasing number of products to be inspected, it is desirable to be able to perform inspections with higher accuracy using equipment without relying on human visual inspection. Therefore, this embodiment describes an apparatus and system for detecting the presence or absence of foreign matter inside a bottle with high accuracy.
[0016] <Conceptual diagram> Figure 1(A) is a conceptual diagram of the learning phase in the foreign object detection system according to this embodiment, and Figure 1(B) is a conceptual diagram of the estimation phase in the foreign object detection system according to this embodiment. In this embodiment, machine learning is performed using images of a bottle containing a foreign object and a bottle without a foreign object, respectively, and the presence or absence of a foreign object in the bottle to be inspected is detected using the trained model generated by machine learning.
[0017] In Figure 1(A), an image set 101 is prepared, which includes multiple images (a group of images) of a bottle without foreign matter and a bottle with foreign matter. All images in image set 101 are images of the bottle taken when the bottle's opening is facing upwards relative to gravity, and the bottle is inverted, or tilted at an angle of less than 90 degrees from the inverted position. The position tilted at an angle of less than 90 degrees from the inverted position is also a position where the opening is below horizontal relative to gravity.
[0018] In this embodiment, the images in image set 101 are saved as a video file, such as in AVI format. The reason for using video is that when the orientation of the bottle opening is changed from, for example, with the opening facing upwards to with the opening facing downwards, if there is a foreign object present, the behavior of the foreign object can be observed mainly in one of the following types. Type (1): A type that falls from top to bottom in a liquid in the direction of gravity. Type (2): A type that floats randomly in a liquid.
[0019] Furthermore, image attributes 102 corresponding to each video file in image set 101 are saved. The image attributes include information for each image file of each bottle, including whether or not there is a foreign object, and if there is a foreign object, which of the above types (1) or (2) the foreign object behaves (movement of the foreign object). In addition, if there is a foreign object, it also includes information including the location of the foreign object (which frame and coordinates it is shown in in the video).
[0020] By aggregating the attributes of each image, we can determine that image set 101 contains A, B, and C images with no foreign objects, A, B, and C images with foreign objects of type (1) and type (2), respectively. Here, one image refers to one video file.
[0021] Next, an adjusted image set 103 is generated from the image set 101 to serve as training data for the learning model (adjustment process). The adjusted image set 103 is a group of images with an adjusted number of images in order to efficiently generate a trained model that enables accurate foreign object detection. The adjusted image set 103 is generated such that C' > B' is satisfied when A' of the images are free of foreign objects, B' of the type (1) images with foreign objects, and C' of the type (2) images with foreign objects.
[0022] The reason for making these adjustments is that type (2) images, which are randomly suspended in a liquid, are difficult to distinguish from images that contain only non-foreign bubbles. Therefore, by increasing the number of type (2) images C', it is possible to generate a trained model that can more accurately distinguish between suspended foreign objects and bubbles.
[0023] Conversely, images of type (1) foreign objects falling from top to bottom in a liquid are easier to distinguish than type (2), so even if the number of type (1) images B' is less than the number of type (2) images C', there will be no significant difference in accuracy. On the other hand, if the amount of training data input to the learning model increases, the processing time increases, and it becomes inefficient as it takes a lot of time to generate a trained model. Therefore, the number of type (1) images B' is reduced compared to the number of type (2) images C'.
[0024] On the other hand, in order to accurately distinguish between floating foreign objects and bubbles, training the model with a considerable number of images containing bubbles but without foreign objects will contribute to improving the accuracy of the trained model. Therefore, in addition to the relationship C'>B', the adjusted image set 103 may be generated to satisfy the relationship A'>B'. Alternatively, the adjusted image set 103 may be generated to satisfy the relationship C'>A'>B'.
[0025] The following selection and supplementation methods are used when generating the adjusted image set 103. (Selection) Among the image groups of each type included in the image set 101, for the types that contain a sufficient number of images, an image group to be included in the adjusted image set 103 is extracted (selected) from the image group of that type. That is, extraction is performed such that at least one of A≧A´, B≧B´, C≧C´ holds.
[0026] For example, when B = C, the number (ratio) of images selected as the images to be included in the adjusted image set 103 from the image group of type (1) included in the image set 101 is less than the ratio of images selected for type (2). Conversely, in this case, the number of images picked up as the images to be included in the adjusted image set 103 from the image group of type (2) included in the image set 101 is more than the number of images picked up for type (1).
[0027] (Supplement) When the image set 101 does not contain a sufficient number of images for any type, the images of the lacking type are processed to supplement the number. That is, supplementation is performed such that at least one of A<A´, B<B´, C<C´ holds. New images are generated by processing the images of the lacking type included in the image set 101, and the generated images are supplemented and included in the adjusted image set 103. As the processing method, at least one of the following processes is performed to generate new images: affine transformation including image enlargement, reduction, rotation, and attribute transformation including brightness, contrast, edge intensity (enhancing process, blurring process).
[0028] Furthermore, a teacher label 104 corresponding to each image in the adjusted image set 103 is generated. The teacher label 104 is generated based on the attributes (image attributes 102) corresponding to each image. The teacher label 104 includes information on whether each image is correct data (positive example data) or incorrect data (negative example data) (either of the two values of having a foreign object or not having a foreign object), and information on the position of the foreign object when there is a foreign object (the position of the foreign object in each frame image of the video). In this embodiment, for example, the correct data is data with a foreign object, and the incorrect data is data without a foreign object. The adjusted image set 103 generated in this way, along with the teacher labels 104, are input into the learning model 105, and training is performed to generate the trained model 106.
[0029] Figure 1(B) shows a conceptual diagram of the estimation phase in the foreign object detection system. An image 110 of a bottle containing liquid, whose presence or absence of foreign objects is unknown, is input to a trained model 106 generated in the learning phase. Here, the image 110 is an image of the bottle taken when the opening is changed to the downward side in the direction of gravity, similar to the images included in the image set 101. In the estimation phase, estimation processing is performed using the trained model 106, and either a result of "no foreign object" or "foreign object present" is output as output data. In this embodiment, the presence or absence of foreign objects in the bottle to be inspected is detected by the learning phase and the estimation phase.
[0030] <System Configuration> Figure 2 is a diagram of the foreign object detection system according to this embodiment, and it executes the processes in the learning phase and estimation phase described in Figures 1(A) and (B). A client terminal 201 and a camera 202, which acts as an imaging device, are connected to the local network (LAN) 200. A bottle 203, which is known to be free of foreign objects and is installed in a jig housing 204, is illuminated by an LED light 205, and an image is acquired by the camera 202.
[0031] In the learning phase described in Figure 1(A), the images acquired by the camera 202 are stored in the storage medium of the client terminal 201 or in the storage medium of the data collection server 211, which can be accessed, for example, via the internet 210. Then, the aforementioned image set 101 is generated. For each image in this image set 101, the user operates the operation unit 306 of the client terminal 201 to input attribute values, thereby generating the aforementioned image attributes 102 in the recording medium of the client terminal 201 or in the data collection server 211.
[0032] The generated image attributes 102 are aggregated and adjusted by the processor (CPU 301, described later) of the client terminal 201 or the data collection server 211 to generate the adjusted image set 103 and teacher labels 104. The generated adjusted image set 103 and teacher labels 104 are input to the learning server 212, which is accessible (communicable) from the client terminal 201 via the internet 210. The learning server 212 has a built-in CPU 301 or GPU 304, described later, and performs training based on the input adjusted image set 103 and teacher labels 104, generating and storing a trained model 105.
[0033] In the estimation phase described in Figure 1(B), a bottle containing a liquid whose presence or absence of foreign matter is unknown is photographed by camera 202, and the image is transmitted to an estimation server 213 accessible via the internet 210. The CPU 301 or GPU 304 of the estimation server 213 performs estimation processing on the received image using a trained model 105 stored in the training server 212, and transmits the estimation result to the client terminal 201. The client terminal displays the received estimation result and records the result.
[0034] In this embodiment, we describe an example where learning is performed on a learning server accessible via the internet and estimation is performed on an estimation server. However, this embodiment can also be applied to a system configuration that is entirely local. Specifically, the image set 101 and image attributes 102 are stored in the storage medium of the client terminal 201 (such as the recording medium 308 or non-volatile memory 303 described later), and the adjusted image set 103 and teacher labels 104 are generated under the control of the client terminal's CPU 301.
[0035] Then, the generated adjusted image set 103 and teacher labels 104 are input to the training model 105 stored in the storage medium of the client terminal 201, and the CPU 301 or GPU 304 of the client terminal 201 performs the training process. The resulting trained model 106 is then stored in the storage medium of the client terminal 201.
[0036] In the estimation phase, the CPU 301 or GPU 304 of the client terminal 201 performs estimation processing by inputting the captured image 110 of a bottle containing liquid whose presence or absence of foreign matter is unknown into the trained model 106. The results of the estimation processing are displayed on the display 305 of the client terminal 201 and stored in the storage medium of the client terminal 201. In other words, the functions of the corner block shown in Figure 2 may be provided by, for example, a single electronic device (client terminal).
[0037] <Hardware Configuration> Figure 3 shows an example of the hardware configuration of information processing devices, including a client terminal 201, a data collection server 211, a learning server 212, and an estimation server 213. The client terminal 201, data collection server 211, learning server 212, and estimation server 213 each have the hardware configuration shown in Figure 3. These can be configured as, for example, a personal computer (PC), a mobile terminal such as a smartphone or tablet, or a server device.
[0038] In Figure 3, the CPU 301, memory 302, non-volatile memory 303, GPU 304, display 305, control unit 306, recording medium interface 307, external interface 309, and communication interface 130 are connected to the internal bus 350. Each component connected to the internal bus 350 is configured to exchange data with each other via the internal bus 350.
[0039] Memory 302 consists of volatile memory, such as RAM, which utilizes semiconductor elements. The CPU 301 executes computer programs stored in non-volatile memory 303, and uses memory 302 as work memory to control various parts of the information processing device. Non-volatile memory 303 stores image data, audio data, other data, and various computer programs for operating the CPU 301. Non-volatile memory 303 is composed of, for example, a hard disk drive (HDD), SSD, or ROM.
[0040] The GPU (Graphics Processing Unit) 304 functions as an image processing unit and is a semiconductor chip (processor) that performs the calculations necessary for image rendering. Based on the control of the CPU 301, the GPU 304 performs image processing on image data stored in the non-volatile memory 303 and recording medium 308, video signals acquired via the external I / F 309, image data acquired via the communication I / F 130, and captured images.
[0041] Image processing performed by GPU304 includes A / D conversion, D / A conversion, image data encoding, compression, decompression, decoding, resizing, noise reduction, and color conversion. GPU304 may also be configured with dedicated circuit blocks for specific image processing tasks. Furthermore, depending on the type of image processing, it is possible to configure the system so that CPU301 performs image processing according to a program without using GPU304.
[0042] Since the GPU304 can perform calculations efficiently by processing more data in parallel, it is desirable to use the GPU304 when performing training multiple times using a learning model such as deep learning. Therefore, in this embodiment, the GPU304 is used in addition to the CPU301 for the training process described later.
[0043] In other words, when executing a learning program that includes a learning model, the CPU 301 and GPU 304 work together to perform calculations and learn. However, the learning process may be performed by either the CPU 301 or the GPU 304 alone. Furthermore, in the estimation process described later, similar to the learning process, the GPU 304 of the information processing device that performs the estimation process may be used, or it may be performed by either the CPU 301 or the GPU 304 alone.
[0044] The display 305 includes, for example, a liquid crystal display for displaying images and GUI (Graphical User Interface) screens, based on the control of the CPU 301. The CPU 301 generates display control signals according to a computer program and controls each part of the information processing device to generate video signals for display on the display 305 and output them to the display 305. The display 305 displays images based on the output video signals.
[0045] Furthermore, the information processing device itself is limited to an interface for outputting video signals to be displayed on the display 305, and the display 305 may be an external monitor (including a television, etc.).
[0046] The operation unit 306 is an input device for receiving user input, including a keyboard and other text information input devices, a mouse and a touch panel, buttons, dials, joysticks, touch sensors, and a touchpad. The touch panel is configured to be superimposed on the display 305 in a planar manner and is configured to output coordinate information corresponding to the position of contact.
[0047] The recording medium interface 307 has a configuration that allows for the insertion and removal of recording media 308 such as memory cards, CDs, and DVDs, and, based on the control of the CPU 301, reads data from the inserted recording media 308 and writes data to the recording media 308.
[0048] External I / F309 is an interface for connecting to external devices via wired or wireless connections and for inputting and outputting video and audio signals. Communication I / F130 is for communicating with external devices via Net 311, such as Internet 210 and Local Network 200, and for sending and receiving various data such as files and commands.
[0049] <Learning Process> Figure 4 is a flowchart of part of the learning process performed on client terminal 201, and Figure 5 is a flowchart of the remaining part of the learning process performed on client terminal 201. Each process in the flowcharts of Figures 4 and 5 is realized when the CPU 301 in the client terminal 201 loads the computer program stored in the non-volatile memory 303 into memory 302 and executes it.
[0050] In step S401, the CPU 301 of the client terminal 201 initializes the variable n, which is stored in memory 302, to 1. n is a counter that indicates which image out of the image set 101 has been retrieved. The image set 101 is assumed to contain a total of N images.
[0051] In step S402, the CPU 301 of the client terminal 201 retrieves the nth image from the image set 101 stored in the client terminal 201 or the data collection server 211.
[0052] In step S403, the CPU 301 of the client terminal 201 accepts a user operation to assign attribute information to the acquired nth image. The user can input attribute information for the nth image by operating the operation unit 306 of the client terminal 201. At this time, the CPU 301 plays and displays the video of the nth image on the display 305, or displays a representative image (for example, the first frame) as a still image. Alternatively, it may simply display an image identifier such as the file name.
[0053] The user looks at the nth image in the video playback and visually identifies the presence or absence of a foreign object, the type of movement of the foreign object, and the frame number and coordinates (location of the foreign object) in which the foreign object is visible, and inputs this information. Alternatively, if the user has already photographed a bottle in which the presence or absence of a foreign object and the type of movement of the foreign object are known, they may input the information based on the information already known. In this case, information can be input without playing the video. Since it is sufficient to identify the frame in which the foreign object is visible, the video's timestamp information or elapsed time information may be used instead of the frame number.
[0054] For example, the location of a foreign object may be determined by the user inputting the location of the foreign object while an image of a predetermined frame is displayed. The CPU 301 can then automatically determine (by performing tracking) whether the same foreign object is visible in other frames and automatically input the information based on the determination.
[0055] In step S404, the CPU 301 of the client terminal 201 determines, based on the operation received in step S403, whether or not input indicating that the nth image was not taken in a properly tilted state. In this embodiment, a properly tilted state is an image of a bottle taken when the bottle is changed from a position where the opening faces upward in the direction of gravity to an upside-down position, or a position tilted at an angle of less than 90 degrees from the upside-down position (i.e., a position where the opening faces downward in the direction of gravity).
[0056] If it is determined that the input is inappropriate, the process proceeds to step S405; otherwise, it proceeds to step S406. Furthermore, if the input indicates that the image is inappropriate for use as training data due to factors other than whether or not it was captured at the correct tilt, the process also proceeds to step S405.
[0057] In step S405, the CPU 301 of the client terminal 201 proceeds to step S411 without recording the nth image in the image attribute 102 (list of attribute information) stored in the non-volatile memory 303. In this way, the nth image, which is not suitable, is excluded from the training data.
[0058] In step S406, the CPU 301 of the client terminal 201 determines whether or not input indicating the presence of a foreign object was received in step S403. If no input indicating the presence of a foreign object was received, or if input indicating the absence of a foreign object was received, the process proceeds to step S407. If input indicating the presence of a foreign object was received, the process proceeds to step S408.
[0059] In step S407, the CPU 301 of the client terminal 201 records in the image attribute 102 of the non-volatile memory 303 that there are no foreign objects in the nth image. For example, the file name in the image attribute 102 of Figure 1 is the image file 0002.avi, indicating that there are no foreign objects.
[0060] In step S408, the CPU 301 of the client terminal 201 determines whether or not it was input in step S403 that the nth image is an image of a foreign object moving in the direction of gravity of type (1) (a type that falls from top to bottom in a liquid). If it is type (1), the process proceeds to step S409; otherwise, the process proceeds to step S410.
[0061] In step S409, the CPU 301 of the client terminal 201 records information about the nth image in the image attribute 102 of the non-volatile memory 303, indicating the presence of a foreign object, the type of motion being type (1), and the position of the foreign object, which shows the frame in which the foreign object is captured and its coordinates. For example, in the image file named:0001.avi in the image attribute 102 of Figure 1, the presence of a foreign object, the type of motion of the foreign object being type (1), and the position of the foreign object are recorded, respectively.
[0062] The movement of foreign matter in a liquid can be broadly classified into the above-mentioned types (1) and (2). Therefore, if there is foreign matter (Yes in step S406) and it is not of type (1) (No in step S408), then in this embodiment, the image will show type (2) foreign matter (the type that floats randomly in the liquid).
[0063] In step S410, the CPU 301 of the client terminal 201 records information about the nth image in the image attribute 102 of the non-volatile memory 303, indicating the presence of a foreign object, the type of motion being type (2), and the position of the foreign object, which shows the frame and coordinates in which the foreign object is captured. For example, as in the image file name:0003.avi in the image attribute 102 of Figure 1, the presence of a foreign object, the type of motion of the foreign object being type (2), and the position of the foreign object are recorded, respectively.
[0064] In step S411, the CPU 301 of the client terminal 201 determines whether n=N or not. That is, it determines whether all images included in the image set 101 have been acquired. If n=N (if there are still images that have not been acquired), the process proceeds to step S412, incrementing n by 1, and proceeding to step S402, where the next image is acquired, and the processing from step S402 to step S411 is repeated. If n=N (if all images included in the image set 101 have been acquired), the process proceeds to step S501 in Figure 5.
[0065] In this way, the image attribute 102 is generated. In this embodiment, an example of generating the image attribute 102 on the client terminal 201 has been described. However, if the image attribute 102 has already been generated by another device, steps S401 to S412, which are the processes for generating the image attribute on the client terminal 201, may be omitted, and processing may start from step S501.
[0066] In step S501, the CPU 301 of the client terminal 201 counts and obtains the number A of images without foreign objects among the images in which information is recorded in image attribute 102 (images included in image set 101). It also counts and obtains the number B of images with foreign objects of type (1) and the number C of images with foreign objects of type (2).
[0067] In step S502, the CPU 301 of the client terminal 201 performs an adjustment process to adjust the total number of images to be used as training data based on the number of images of each type aggregated in step S501. Then, it generates an adjusted image set 103 and training labels 104 and records them in the non-volatile memory 303. The adjustment process is as explained in the conceptual diagram of the learning phase in Figure 1(A). That is, it generates an adjusted image set 103 consisting of a total of M images of types A', B', and C' that satisfy the relationship C'>B'.
[0068] Thus, step S502 adjusts the number of images in an image set 101, which consists of multiple images obtained by photographing the containers when they are changed between a first and a second orientation, for multiple containers containing liquid. This adjustment then functions to generate an adjusted image set 103 in which there are more images of the second type, where the foreign object is floating or rising in the liquid, than images of the first type, where the foreign object is falling from top to bottom.
[0069] Furthermore, in step S502, as described above, the adjusted image set 103 may be generated by adjusting the number of images so that the number of images without foreign objects is greater than the number of images of type 1. Alternatively, the adjusted image set 103 may be generated by adjusting the number of images so that the number of images without foreign objects is greater than the number of images of type 1, and the number of images of type 2 is greater than the number of images without foreign objects.
[0070] In step S503, the CPU 301 of the client terminal 201 initializes the variable m, which is stored in memory 302, to 1. m is a counter that indicates how many images from the adjusted image set 103 have been input into the learning model.
[0071] In step S504, the CPU 301 of the client terminal 201 inputs the m-th image from the adjusted image set 103 and the information (label) indicating the presence or absence and location of a foreign object from the m-th image of the training label 104 into the learning model 105. Specifically, it sends the m-th image and label to the learning server 212. In other words, it controls the machine learning to be performed using the label that is associated with each image in the training data, indicating either the presence or absence of a foreign object.
[0072] In step S505, the CPU 301 of the client terminal 201 performs machine learning processing using the learning model 105 for the m-th image. For example, it receives a response from the learning server 212 indicating whether the m-th image and label were correctly input into the learning model 105.
[0073] Furthermore, when performing the learning process locally, in step S504, the m-th image and label are input to the learning model 105 held on the client terminal 201 without being sent to the learning server 212. Then, in step S505, the CPU 301 and GPU 304 of the client terminal 201 work together to perform the learning process. This learning process is computationally intensive, so the adjustment process in step S502 adjusts the number of images input to the learning model 105 as training data. In other words, by performing machine learning using the aforementioned adjusted image set 103 as training data, unnecessary processing can be omitted and the processing time related to learning can be shortened.
[0074] Specific machine learning algorithms include nearest neighbors, naive Bayes, decision trees, and support vector machines. Deep learning, which uses neural networks to generate features and connection weights for learning, is another example. Appropriately, any of the above algorithms can be applied to this embodiment.
[0075] In step S506, the CPU 301 of the client terminal 201 determines whether m=M or not. If m=M (if not all images included in the adjusted image set 103 have been input to the training model 105), the process proceeds to step S507 to increment m, and then to step S504 to input the next image and label to the training model 105. If m=M (if all images included in the adjusted image set 103 have been input to the training model 105), the training process is terminated. In this way, the trained model 106 is generated.
[0076] Thus, in steps S401 to S507, machine learning is performed using a set of images consisting of multiple images obtained by photographing multiple containers containing liquid in a first and second orientation. This then functions to generate a trained model for estimating whether or not there is a foreign object in the container.
[0077] <Estimation Process> Figures 6(A) and 6(B) are flowcharts of the estimation process, and Figure 6(A) is a flowchart of the estimated image transmission process executed on the client terminal 201. The process in the flowchart of Figure 6(A) is realized on the client terminal 201 by the CPU 301 loading the computer program stored in the non-volatile memory 303 into memory 302 and executing it.
[0078] In step S601, the CPU 301 of the client terminal 201 performs an imaging process on the bottle to be inspected, whose presence or absence of foreign matter is unknown. During the imaging process, the bottle is changed from a position where the opening is facing upwards relative to gravity to an inverted position, or a position tilted at an angle of less than 90 degrees from the inverted position (i.e., a position where the opening is facing downwards relative to gravity). An image of the bottle in this position is then acquired.
[0079] Furthermore, the reference point is not limited to the bottle's opening; other parts such as the bottom or neck may also be used. In other words, any image of the bottle taken when the specific reference point of the container is changed from an upward position relative to gravity to a downward position relative to gravity is sufficient. This change in posture allows for the observation of the movement of a type (1) foreign object falling from top to bottom in the liquid, making it easier to distinguish between a case without a foreign object and a type (2) foreign object.
[0080] Specifically, the system acquires images of the container as detection targets when the container's orientation is changed between a first orientation in which a specific part of the container containing the liquid is facing upwards relative to the direction of gravity, and a second orientation in which a specific part is facing downwards relative to the direction of gravity. This image was taken in the same posture as the images in image set 101, which are collected during the training process. Details of the image capture process will be described later using Figure 24.
[0081] In step S602, the CPU 301 of the client terminal 201 transmits the image (the image to be detected as a foreign object) acquired in the imaging process of step S601 to the estimation server 213 via the communication I / F 310. Note that, in order to improve the accuracy of the estimation process, various image processing (processing) such as cropping, rotation, and color adjustment may be performed on the image before transmission.
[0082] In step S603, the CPU 301 of the client terminal 201 determines whether or not it has received estimation result information (estimation result information) from the estimation server 213. The estimation result received here is the same as the one transmitted in S613 of Figure 6(B), which will be described later. If the estimation result has not been received, the system waits; if the estimation result has been received, the system proceeds to step S604.
[0083] In step S604, the CPU 301 of the client terminal 201 determines whether the estimation result received in step S603 indicates "foreign object present". In this embodiment, the estimation result will either indicate "foreign object present" or "no foreign object present". If it is determined that "foreign object present", the process proceeds to step S606; if it is determined that "no foreign object present", the process proceeds to step S605.
[0084] In step S605, the CPU 301 of the client terminal 201 displays on the display 305 that the estimated result of the estimated image (the image acquired in step S601) is "no foreign objects". Furthermore, at this time, along with the estimation result, at least one of the following may be displayed: information that serves as an identifier for the estimated image (file name, file number, etc.), a video of the estimated image (the image obtained in step S601) or a still image of a representative image, or identification information of the bottle of the estimated target.
[0085] Furthermore, since "no foreign objects" means "no problems" ("good") in terms of the inspection result, the operator does not need to take any special action. Therefore, it is not necessary to display anything in response to the estimated result of "no foreign objects." In addition, information indicating that "no foreign objects" were found is recorded in the non-volatile memory 303, associated with the identifier of the estimated image or the identification information of the bottle. However, this recording process may also be omitted.
[0086] In step S606, the CPU 301 of the client terminal 201 displays on the display 305 that the estimation result for the estimated image (the image acquired in step S601) is "foreign object present". At this time, along with the estimation result, at least one of the following may be displayed: information that serves as an identifier for the estimated image (file name, file number, etc.), a video of the estimated image (the image acquired in step S601) or a still image of a representative image, or identification information of the bottle that was estimated.
[0087] Here, steps S602 to S606 control the system to input the image acquired in S601 into the trained model, and then control the system to notify the foreign object detection result based on the estimated result information obtained as a result of the input.
[0088] Furthermore, "Foreign object present" means "Problem found" ("Defective") in the inspection results, so some action is required, such as a detailed inspection of the affected bottle or removal from shipment. Therefore, it is necessary to ensure that operators are notified, and it is preferable to make the notification more conspicuous than "No foreign object present." For example, it is desirable to display the indicator item showing "Foreign object present" larger than the indicator item showing "No foreign object present," to display it in a conspicuous color such as red, or to make it flash.
[0089] In addition to displaying the information, a warning sound may be emitted from a speaker connected to the external I / F 309 or from a built-in speaker (not shown). Alternatively, an identification item (such as a sticker) corresponding to the estimated "foreign object present" result may be attached to the target bottle using a drive unit (not shown). Furthermore, information indicating that a "foreign object was present" is recorded in the non-volatile memory 303, associated with the identifier of the estimated image or the identification information of the bottle.
[0090] Figure 6(B) is a flowchart of the estimation process performed on the estimation server 213. This process is achieved when the CPU 301 of the estimation server 213 loads the computer program stored in the non-volatile memory 303 into memory 302 and executes it.
[0091] In step S610, the CPU 301 of the estimated server 213 determines whether or not it has received the captured image (the image sent in step S602 above) transmitted from the client terminal 201. If it determines that it has received the captured image, it proceeds to step S611; otherwise, it waits until it receives the image.
[0092] In step S611, the CPU 301 of the estimation server 213 inputs the captured image received in step S610 into the trained model 106. This trained model 106 was generated by the training process described in Figures 4 and 5. Before inputting the captured image into the trained model 106, various image processing (processing) such as cropping, rotation, and color adjustment may be performed on the image to improve the accuracy of the estimation process.
[0093] In step S612, the CPU 301 of the estimation server 213 performs estimation processing using the trained model 106 to estimate whether or not there is a foreign object in the input image (whether or not it is an image of a bottle with a foreign object). This processing is performed in cooperation with the GPU 304 as needed.
[0094] In step S613, the CPU 301 of the estimation server 213 sends information to the client terminal 201 indicating either "foreign object present" or "foreign object absent" as a result of the estimation process in step S612. In this way, steps S610 to S613 perform estimation processing using the trained model.
[0095] In this embodiment, a system has been described in which the client terminal 201 that obtains the image to be inspected and the estimation server 213 that performs estimation processing are separate devices. However, it is also possible to perform these processes within the same device. For example, if the trained model 106 is stored in the non-volatile memory 303 of the client terminal 201, it is possible to perform both the acquisition of the captured image to be inspected and the estimation processing on the client terminal 201 without communicating with the internet.
[0096] In that case, the captured image is not transmitted in step S602, and the processing in steps S611 and S612 is performed on the client terminal 201 side. Then, depending on the estimation result, the processing in step S605 or step S606 can be performed.
[0097] In step S612, the estimation process was described using a pre-trained machine learning model 106, but rule-based processing such as a lookup table (LUT) may also be used. In that case, for example, the relationship between input data and output data is created in advance as a LUT. This created LUT is then stored in the non-volatile memory 303 of the device performing the estimation process. When performing the estimation process, the output data can be obtained by referring to this stored LUT.
[0098] <Example of a shooting system 1> Next, using Figures 7 to 9, we will describe an imaging system suitable for a foreign object detection system for detecting foreign objects mixed into a container holding liquid. Figure 7 is a diagram illustrating an example of the arrangement of the camera 202 and bottle 203 in the imaging system described in Figures 1 to 6.
[0099] In Figure 7, the camera 202 used for imaging is positioned on the jig housing 204 so as to face the body of the bottle 203. The jig housing 204 is configured to allow easy attachment and detachment of the bottle 203 by a mechanism not shown. The image captured by the camera 202 is output to the client terminal 201. The bottle 203 is light-transmitting, and an LED light 205 is positioned on the opposite side of the bottle 203 from the camera 202, which is the imaging means, to illuminate the container.
[0100] Furthermore, when the camera 202 takes a picture, illuminating the inside of the bottle allows the camera 202 to capture the transmitted light from the LED light 205. Here, the camera 202 and the LED light 205 function as means for capturing images of the container. The distance between the camera 202 and the bottle 203 can be adjusted manually or electrically by the user.
[0101] The jig housing 204 is fixed to the jig base 213, and the entire jig housing 204 is configured to rotate around the rotation center 214. Inside the bottle 203, there are foreign objects 215 with a specific gravity greater than the liquid and foreign objects 216 with a specific gravity less than the liquid. When the bottle 203 is placed upright along the vertical direction, the foreign objects 215 with a higher specific gravity sink to the bottom of the bottle, and the foreign objects 216 with a lower specific gravity float to the neck of the bottle. With the bottle in this state, the LED light 205 is turned on and shooting begins.
[0102] Figure 8 shows the jig housing 204 being smoothly rotated 90° clockwise around an axis perpendicular to the plane of the paper. To prevent the generation of a large amount of air bubbles during rotation, the rotation and stopping are controlled slowly (at a low speed) and smoothly. The bottle 203 will be placed horizontally, and in this case, the foreign matter 215 with a high specific gravity will sink to the lower side of the body near the bottom of the bottle, while the foreign matter 216 with a low specific gravity will float to the upper side near the shoulder of the bottle. The state shown in Figure 8 may be left undisturbed for a short time, or it may be slowly and smoothly transitioned to the state shown in Figure 9.
[0103] Figure 9 shows the jig housing 204 rotated slowly and smoothly clockwise from the state shown in Figure 8, for example, by about 20° to 60°. When a horizontally positioned bottle is tilted upside down, the denser foreign matter 215 that had settled on the underside of the bottle body near the bottom sinks along the side of the bottle body on the side furthest from the camera. The same effect can be obtained by tilting the bottle up to a vertical 90 degrees to an inverted state.
[0104] On the other hand, the foreign object 216, which has a low specific gravity, floats along the side of the bottle body closest to the camera. While the general tendency of foreign objects to move is described according to their specific gravity, extremely small foreign objects of 0.2-0.3 mm or less may float randomly inside the bottle. In this example of a shooting system, foreign objects moving inside the bottle can be accurately photographed using the camera 202.
[0105] Figure 10 is a diagram showing the parts of Figure 9, specifically the bottle 203 containing the camera 202, foreign objects 215 and 216, and the LED light 205, with the bottle facing downwards. Camera 202 is equipped with a photographic lens 218 as an optical system, and photographs foreign objects 215 and 216 inside the bottle, which are illuminated by LED lighting 205, through the optical system.
[0106] The image inside the bottle is reduced and projected onto the camera's sensor surface 217 by the imaging lens 218. Here, D1 is the distance from the front main plane of the lens to the center of the bottle, and D2 is the distance from the LED illumination to the center of the bottle. Also, L is the vertical length of the drawing of the shooting range inside the bottle (for example, the transparent area from the bottom of the bottle to near the neck).
[0107] In the first example of the shooting system, the shooting lens 218 and the camera's sensor surface 217 are selected so that the entire shooting range L inside the bottle can be captured at once. For example, let's assume the focal length of the photographic lens 218 is f=25mm, the vertical size of the sensor in Figure 10 is x=14.1mm, and the length of the shooting range of the bottle body is L=234mm. In this case, in order to photograph the entire shooting range L at once, it is desirable that the distance D1 as a shooting parameter satisfies Equation 1, and it can be seen that the distance D1 = 440mm is sufficient.
[0108] D1=(1+L / x)f...(Formula 1) Furthermore, the LED lighting was designed with dimensions of 236mm in height and 182mm in width to illuminate the entire shooting range L and the bottle diameter 2r = 76mm. The illumination wavelength was set to λ = 850nm, which maximizes transmittance regardless of the glass color or liquid color of the bottle. The surface of the lighting was designed as a diffusing surface so that foreign objects could be illuminated at various light angles. On the other hand, D2 = 100mm was set to prevent heat from the lighting from being transferred into the bottle.
[0109] Figure 11 is a diagram illustrating the image 215i of the foreign object 215, which is reduced in size and formed on the sensor surface 217 of the camera. The sensor in imaging system example 1 has a long side dimension of 14.1 mm × short side dimension of 10.35 mm, a sensor pixel 219 pitch p = 3.45 μm, and a total of 12.28 million pixels. A foreign object at a distance D1 is imaged onto the sensor surface 217 based on the following equation 2. Magnification -x / L=-1 / 16.6 (Formula 2)
[0110] Foreign objects can be categorized into two types: those that do not transmit light and appear as black shadows against the lighting, such as metal fragments and cork pieces, and those that refract or partially transmit the lighting, such as glass fragments and plastic pieces. The former have high contrast and can be detected relatively easily, while the latter have lower contrast and are difficult to detect. When various types and sizes of foreign objects were actually photographed, it was found that objects occupying 3x3 pixels on the sensor could be detected with almost certainty. That is, the size a of the foreign object, the sensor pixel pitch p, and the focal length f and distance D1 of the lens as shooting parameters should satisfy the following equation 3.
[0111]
number
[0112]
number
[0113] First, let's explain the case where bottle 203 in this example of the imaging system is a round bottle. Because a round bottle has curvature in the direction of its diameter, it has a lens effect on foreign objects inside the bottle. Figure 12 is a schematic diagram showing a virtual image of a foreign object when the container is a round bottle, and the lens effect of the bottle will be explained using Figure 12. A round bottle can be considered a cylindrical lens because it has curvature in the direction of the bottle's diameter. The lens effect is greatest when the object passes near the side of the bottle body on the side furthest from the camera, as shown by the foreign object 215 in Figure 10.
[0114] When a foreign object, such as 216, passes near the side of the bottle body closer to the camera, the lens effect is small, and there is almost no lens effect in the vertical direction of the bottle. Therefore, when a foreign object, such as 215, passes near the side of the bottle body further from the camera, the image is magnified only in the bottle diameter direction, as shown in the foreign object image 215i in Figure 12. It goes without saying that the lens effect magnifies the foreign object in the bottle diameter direction, making detection easier.
[0115] Figure 13 is a diagram illustrating the imaging relationship when bottle 203 is a round bottle. The side 220 of bottle 203 closest to the camera can be considered a cylindrical lens with curvature 1 / radius r, so the foreign object 215 located near the side of the bottle body farther from the camera appears as a virtual image 215' of the foreign object 215 due to the lens effect. If the refractive index of the liquid is n, the distance from the cylindrical lens to the virtual image 215' of the foreign object 215 and the magnification are expressed by equations 5 and 6, respectively. Here, the focal length of the cylindrical lens is fc and the distance from the cylindrical lens to the foreign object 215 is d.
[0116]
number
number
[0117]
number
number
[0118] If n = 1.333, then from Equation 7, the distance d from the cylindrical lens to the foreign object 215 is 3r. Also, from Equation 8, it can be seen that the magnification is 2 regardless of the bottle diameter. In other words, the foreign object 215 in the liquid near the side of the bottle body of the round bottle furthest from the camera appears magnified twice in the direction of the bottle diameter. Strictly speaking, the refractive index of the bottle glass is about 1.5, which is slightly higher than the refractive index of the liquid, but the difference in refractive index is small, so refraction at the interface, the thickness of the bottle glass, etc. can be ignored.
[0119] On the other hand, the foreign object 216 located near the side of the bottle body on the side closer to the camera is in close proximity to the cylindrical lens, so its lens effect is negligible. Furthermore, in the case of either foreign object, there is almost no lens effect in the vertical direction (longitudinal direction) of the bottle.
[0120] Next, we consider how each foreign object is imaged near the camera's sensor surface 217. The foreign object 215 located near the side of the bottle body on the side furthest from the camera (towards the back of the bottle) appears to be in the virtual image 215' of the foreign object 215 due to the lens effect. Therefore, according to Equation 7, the distance from the imaging lens 218 to the virtual image of the foreign object 215 is: D1-r+2r / (2-n)=D1+nr / (2-n) Therefore, if n=1.333, the result is D1+2r. Although not shown in the diagram, the foreign object 215 is hardly affected by the lens effect in the bottle height direction (longitudinal direction), so the distance appears to be shorter by the refractive index of the liquid.
[0121] That is, the distance from the photographic lens 218 to the virtual image of the foreign object 215 is D1-r+r / n=D1+r(2-n) / n, and if n=1.333, it becomes D1+r / 2. On the other hand, the foreign object 216 located near the side of the bottle body on the side closer to the camera (the front side of the bottle) can have its lens effect ignored, so the distance from the photographic lens 218 to the foreign object 216 is D1-r.
[0122] When the focus position of the imaging lens 218 is adjusted so that the center of the bottle is imaged on the sensor surface 217, the foreign object 215 located at the back of the bottle is imaged as an image 215i of the foreign object 215 at a position closer to the imaging lens 218 than to the sensor surface 217. In other words, the image in the direction of the bottle diameter is imaged at a closer position. The foreign object 216 located at the front of the bottle is imaged as an image 216i of the foreign object 216 at a position further from the imaging lens 218 than to the sensor surface 217.
[0123] Therefore, by optimizing the imaging system so that the images 215i of foreign object 215 and 216i of foreign object 216 fall within a depth of field with a small amount of blur, it is possible to photograph all foreign objects inside the bottle at once. Depth of field refers to the range in which a point image blurs and spreads out to the acceptable circle of confusion, and will be explained in more detail later.
[0124] <Example of a shooting system 2> As an example of a shooting system, we will describe the case where bottle 203 is a square bottle. Since the side of the bottle's body is flat, if this flat surface is positioned directly in front of the camera, the lens effect on foreign objects inside the bottle will be minimal.
[0125] Figure 14 is a diagram illustrating the imaging relationship when bottle 203 is a square bottle in the example imaging system 2. The side of bottle 203 closest to the camera (the front side) can be considered a plane, so the light rays from the foreign object 215 near the side of the bottle body furthest from the camera are refracted by the bottle side and appear to be in the virtual image 215' of the foreign object 215. If the refractive index of the liquid is n, the distance from the bottle side to the virtual image 215' of the foreign object 215 is expressed by Equation 9. However, the distance (depth dimension) from the front side of the bottle to the foreign object 215 is assumed to be 2t.
[0126] 2t / n (Formula 9) If n = 1.333, then from Equation 9, the distance from the side of the bottle to the foreign object 215 is 1.5t. On the other hand, the foreign object 216, which is near the side of the bottle body that is closer to the camera, is in close proximity to the side of the bottle, so the effect of refraction at the side of the bottle can be ignored.
[0127] Next, we consider how each foreign object is imaged near the camera's sensor surface 217. The foreign object 215 located near the side of the bottle body on the side furthest from the camera (towards the back of the bottle) appears as a virtual image 215' of the foreign object 215 in the direction of the bottle's diameter due to refraction by the liquid.
[0128] Therefore, according to equation 9, the distance from the photographic lens 218 to the virtual image 215' of the foreign object 215 is, D1-t+2t / n=D1+t(2-n) / n, and if we set n=1.333, it becomes D1+t / 2. On the other hand, the foreign object 216 located near the side of the bottle body on the side closer to the camera (the front side of the bottle) is affected by refraction at the side of the bottle, so the distance from the photographic lens 218 to the foreign object 216 is D1-t.
[0129] By adjusting the focus position of the imaging lens 218 so that the center of the bottle is imaged on the sensor surface 217, the foreign object 215 located at the back of the bottle is imaged as an image 215i of the foreign object 215 at a position closer to the imaging lens 218 than to the sensor surface 217. In other words, the image in the direction of the bottle diameter is imaged at a closer position. The foreign object 216 located at the front of the bottle is imaged as an image 216i of the foreign object 216 at a position further from the imaging lens 218 than to the sensor surface 217. Therefore, by optimizing the imaging system so that the images 215i of foreign object 215 and 216i of foreign object 216 fall within the depth of field, the foreign objects inside the bottle can be imaged simultaneously.
[0130] Next, Figures 15(A) to (C) are diagrams to explain depth of field and blur tolerance (allowable circle of confusion). Figure 15(A) shows the case where the foreign object image 215i is imaged on the sensor surface 217 to occupy a size of 4x4 pixels, as expressed by Equation 4. The pitch of the sensor pixels 219 is p. The imaging lens 218 is assumed to be an ideal lens that can image the foreign object onto a 4p square pixel.
[0131] Figure 15(B) shows a state where the same foreign object has shifted position in the optical axis direction, resulting in blurring due to defocus on the sensor surface 217. In the focused state, a portion of the foreign object is imaged to the size of one pixel, but in Figure 15(B), it has spread to adjacent pixels, resulting in a width of three pixels. The dashed line in the figure represents the area occupied by the foreign object image in the focused state (A), and the shaded area represents the area of the blurred image. The area occupied by the foreign object image expands to more than twice its original size, and the contrast decreases, but it is still possible to distinguish the presence or absence of the foreign object.
[0132] Figure 15(C) shows a further increase in blurring due to defocusing. In the case of in-focus images, a portion of the foreign object is imaged to the size of one pixel, but in Figure 15(C), it extends to the adjacent pixel, becoming 5 pixels wide. The area occupied by the foreign object image expands fourfold, and the contrast decreases to one-quarter, but the presence or absence of the foreign object is still discernible.
[0133] For foreign objects with high light transmittance and relatively low contrast, such as glass or plastic fragments, detection may fail if the minimum circle of confusion (blur) exceeds 6 pixels. Therefore, it is desirable to perform pre-processing such as contrast enhancement or edge enhancement. Similarly, for foreign objects with relatively high contrast, such as metal or cork fragments, detection may fail if the minimum circle of confusion (blur) exceeds 7 pixels. Therefore, it is desirable to keep the minimum circle of confusion (blur) within 7 pixels and within the depth of field range suitable for foreign object detection.
[0134] Furthermore, if the same foreign object is present inside a round bottle, and its radial dimension is magnified twice due to the lens effect of the bottle's side, resulting in an image occupying 4x8 pixels, the reduction in contrast due to blurring is mitigated. This is illustrated in Figure 16.
[0135] Figure 16 shows the allowable circle of confusion diameter in units of sensor pixels on the horizontal axis, and the normalized contrast of the focused state set to 1 on the vertical axis. As mentioned above, in the case of a square bottle, for example, where the foreign object image 215i with relatively low contrast, such as glass or plastic pieces, occupies 4 x 4 pixels, stable detection is possible up to a minimum circle of confusion diameter of 5 pixels (equivalent to a contrast of 0.25). However, when the minimum circle of confusion diameter is 6 pixels, the contrast falls below 0.2, and detection may fail.
[0136] In contrast, with the same foreign object, for example, in a round bottle where the foreign object image 215i occupies 4 x 8 pixels, stable detection was possible up to a minimum confusion circle diameter of 7 pixels (equivalent to a contrast of 0.23). However, at a minimum confusion circle diameter of 8 pixels, the contrast falls below 0.2, and detection may fail.
[0137] The depth of field is an evaluation quantity defined on the object side that specifies the conditions under which the amount of blur on the sensor surface 217 of the foreign object image falls within a predetermined range. In other words, the shooting system should be set to optimal conditions so that the detection range of the foreign object on the object side, i.e., the liquid range inside the bottle container, falls within the depth of field, just as the image falls within the depth of focus. That is, the forward depth of field, which is the distance to the foreign object 216 on the front side of the bottle that is close to the camera, and the backward depth of field, which is the distance to the foreign object 215 on the back side of the bottle that is far from the camera, can be expressed by equations 10 and 11, respectively.
[0138]
number
number
[0139] However, δ is the allowable blur (allowable circle of confusion), Fno is the F-number (aperture F value) of the photographic lens 218, f is the focal length of the photographic lens 218, and D1 is the distance between the photographic lens 218 and the bottle. As can be seen from equations 10 and 11, the rear depth of field is greater than the front depth of field.
[0140] Figure 17 illustrates the relationship between depth of field, depth of focus, and acceptable circle of confusion (blur tolerance) in the case of a round bottle container. Considering the case of a round bottle with a diameter of 2r, the radial direction of the foreign object 215 located at the furthest point from the camera (back side) of the bottle is magnified by the curvature of the bottle's side surface. Therefore, according to Equation 7, the foreign object 215 appears to be located at a position of 2r / (2-n) from the front side of the bottle's side surface. It is also located at a position of nr / (2-n) from the center of the bottle. Near the camera's sensor surface 217, the image 215i of the foreign object 215 is formed on the front side of the sensor surface, and on the sensor surface, it is defocused and blurred to a diameter δ.
[0141] The foreign object 216, located closest to the camera on the bottle (towards the front), is not affected by the curvature of the bottle's side and is therefore located at a position r from the center of the bottle. Near the camera's sensor surface 217, the image 216i of the foreign object 216 is formed on the far side of the sensor surface, and on the sensor surface itself, it is defocused and blurred to a diameter δ.
[0142] The above δ is sufficient if it is less than or equal to the acceptable circle of confusion, which is the blur tolerance. The blur tolerance 2 for the foreign object 215 at the back of the bottle is expanded by n / (2-n) times in the radial direction according to Equation 8, so if n=1.333 it is expanded by 2 times, so as explained in Figure 16 it is sufficient if the minimum circle of confusion diameter is 7 pixels or less. Converting this to depth of field, if it is nr / (2-n) and n=1.333, then if it is within the rear depth of field 2r, the foreign object 215 located near the back of the bottle can be detected with blur tolerance or less.
[0143] On the other hand, the blur tolerance 1 for foreign objects 216 on the near side of the bottle is not affected by the curvature of the bottle, so as explained in Figure 16, it is sufficient if the minimum circle of confusion diameter is 5 pixels or less. Converted to depth of field, if it is within the forward depth of field, foreign objects 216 located near the near side of the bottle can be detected with blur tolerance or less. That is, blur tolerance 2 is approximately 1.4 times blur tolerance 1. Furthermore, if the container has a curved surface with radius r in the direction of the optical axis of the optical system, the lens focal length f, distance D1, aperture F value, etc. as shooting parameters should satisfy equation 12 for the forward depth of field and equation 13 for the rear depth of field.
[0144]
number
number
[0145] For example, a camera with a sensor size of 1.1 inches (x=14.1 mm) and a sensor pixel pitch of p=3.45 μm, along with a lens with a focal length of f=25 mm and a distance D1=440 mm, is used to capture an image with an imaging area of 235 mm in height and 172 mm in width. The bottle was a Bordeaux wine bottle with a diameter of 2r=76 mm.
[0146] The graph shows the front and rear depth of field when the lens Fno is changed. It is sufficient if these values cover the liquid area inside the bottle within a predetermined acceptable circle of confusion, i.e., the acceptable amount of blur. For foreign object 215 located in front of the bottle, it is sufficient if the front depth of field has a radius r (38 mm) or more and the acceptable circle of confusion is 5p (17.25 μm) or less. For foreign object 216 located behind the bottle, it is sufficient if the rear depth of field has a bottle diameter 2r (76 mm) or more and the acceptable circle of confusion is 7p (24.15 μm) or less.
[0147] Figure 18 shows Table 1 in Example 1 of the shooting system, where areas where the front depth of field is greater than or equal to radius r (38 mm) and the rear depth of field is greater than or equal to the bottle diameter 2r (76 mm) are highlighted, and the rest is shown in gray. In Table 1, when the aperture is opened and shooting is performed with Fno=2.8, none of the depth of field values in the table cover the liquid area inside the bottle.
[0148] When shooting with an aperture of Fno=8, the acceptable circle of confusion is 5p (17.25μm), resulting in a front depth of field of 39mm. This allows the amount of blur to be kept below the acceptable circle of confusion even for foreign objects located closest to the front of the bottle. With an acceptable circle of confusion of 7p (24.15μm), the rear depth of field becomes 69mm, and foreign objects located furthest from the back of the bottle cannot be covered.
[0149] Furthermore, when shooting with Fno=9.5, the acceptable circle of confusion is 5p (17.25μm) and the depth of field is 46mm, and the acceptable circle of confusion is 7p (24.15μm) and the depth of field is 85mm. This allows the amount of blur to be kept below the acceptable circle of confusion from the foreign object closest to the viewer to the foreign object furthest from the viewer to the viewer.
[0150] In the above example of the shooting system, by shooting with an aperture value of Fno=9.5 or higher as a shooting parameter, it is possible to capture images with the amount of blur suppressed to below the acceptable circle of confusion, from the foreign object located closest to the front of the bottle to the foreign object located furthest back. When the aperture is stopped down significantly, the depth of field increases, but blur due to diffraction occurs, reducing the ability to detect minute foreign objects. Therefore, it is preferable to set the aperture to around Fno=9.5 to 13.
[0151] The above describes the optimal settings for shooting parameters, but now I will explain why these are difficult to achieve with a typical autofocus (AF) camera. Conventional autofocus (AF) cameras focus on the foreground subject, such as the surface of the bottle (for example, a label on a wine label, hereinafter referred to as the label). Therefore, it is impossible to selectively focus on extremely minute foreign objects that move and whose location inside the bottle is unpredictable.
[0152] Furthermore, the inside of the bottle, where there is no target object to focus on, cannot be designated as the focus area. When using aperture-priority autofocus, the focus will lock onto the bottle surface (seal), and in this case, if the aperture is stopped down, it is possible to set the bottle's depth of field to the rear depth of field as described above. If the plane of focus becomes the bottle surface, the front depth of field will be outside the bottle in front of it and will not be used at all for foreign object detection.
[0153] If foreign objects inside a bottle are to be detected using only the rear depth of field, the focus must be narrowed to the inside of the bottle. In practice, considering the reduction in resolution due to diffraction blur and the decrease in sensitivity when using an inexpensive CMOS sensor, detecting minute foreign objects is difficult. According to the embodiment of the present invention, as described in this embodiment, the focus is adjusted manually, for example (focus adjustment is performed using manual focus / MF). Therefore, the focus position is not taken up by the bottle surface, and the inside of the bottle can be covered within the front depth of field and the rear depth of field, making it possible to effectively image and detect minute foreign objects with unpredictable movements.
[0154] The frame rate for video recording was set to 30fps. Each shooting system may use not only video recording but also interval shooting (time-lapse shooting) to capture still images in succession. The reason for using video or a series of still images is to prevent misidentification of scratches or dirt on the bottle container as foreign objects. Scratches and dirt on the bottle container remain in the same position, but foreign objects move around inside the container, making them distinguishable.
[0155] When detecting foreign objects using still images (still image files) obtained from interval shooting, which involves continuously shooting at predetermined intervals, foreign objects moving faster than a predetermined value are detected using interval images with shorter shooting intervals than foreign objects moving slower. This makes it possible to appropriately capture foreign objects according to their movement. It is desirable that the shooting interval for interval shooting be variable and set by the user.
[0156] Foreign objects with a specific gravity significantly greater than that of the liquid in the container, or foreign objects with a specific gravity less than that of the liquid but relatively large in size, will quickly sink or float inside the container when the container is tilted. In such cases, if the frame rate during video recording is set to at least 20fps, preferably 30fps or higher, the foreign object can be reliably detected by comparing the difference in frames in the time direction between a frame at a certain moment and the next frame, etc. The frame comparison does not have to be between adjacent frames; it may be every few frames.
[0157] On the other hand, foreign objects with a specific gravity almost the same as the liquid, or relatively minute foreign objects, are affected by the resistance of the liquid and will slowly sink or float in the liquid. To detect such foreign objects, setting the frame rate to 15fps or less, preferably 10fps or less, allows for reliable detection of foreign objects by comparing the difference between a given frame and the next frame, as foreign objects will have moved a certain distance or more. Furthermore, since the file size of the video can be kept small, the burden on subsequent image processing and AI processing can be reduced.
[0158] Furthermore, in order to simultaneously detect both fast-moving and slow-moving foreign objects within a container, it is also possible to extract and use a portion of images captured at a constant rate. For example, in images captured at 30fps, the difference between a frame (N) at a certain point in time and the next frame (N+1) can be compared. This allows for the detection of fast-moving foreign objects, and by comparing the difference between frame N and frame (N+3) at a certain point in time, slow-moving foreign objects can be detected. Generally, for fast-moving foreign objects, the Nth frame should be compared with the N+M frame (M=1,2,…), and for slow-moving foreign objects, the Nth frame should be compared with the N+L frame (L>M, L is an integer) to perform foreign object detection.
[0159] Furthermore, the video is output from camera 202 to client terminal 201 as RAW data. Client terminal 201 outputs the compressed image in spatial and temporal directions, for example, in an AVI image file format, to workstations that handle well-known image processing, as well as AI processing such as rule-based machine learning and deep learning.
[0160] <Example of a shooting system 3> Next, we will explain an example using a camera with a sensor size of, for example, APS-C size (x=22.5mm) and a sensor pixel pitch p=4.11μm, and a lens with a focal length f=35mm. The distance D1=440mm, and the imaging area will be set to a height (longitudinal) dimension of 260mm × width dimension of 174mm. The bottle was a Bordeaux-type wine bottle with a 2r=76mm diameter.
[0161] Figure 19 shows Table 2, which represents the front and rear depth of field when the lens Fno is changed in the example shooting system 3. It is sufficient if the front and rear depth of field can cover the liquid area inside the bottle with a value below a predetermined acceptable circle of confusion, i.e., the acceptable amount of blur. In other words, for foreign objects 215 located in front of the bottle, a forward depth of field of radius r (38 mm) or more and an acceptable circle of confusion of 5p (20.56 μm) or less is acceptable. For foreign objects 216 located behind the bottle, a rear depth of field of bottle diameter 2r (76 mm) or more and an acceptable circle of confusion of 7p (28.78 μm) or less is acceptable.
[0162] In shooting system example 3, the sensor size is larger than in shooting system example 1, resulting in a narrower depth of field. When shooting with an aperture of Fno=13, the acceptable circle of confusion is 5p (20.56μm) and the front depth of field is 39mm, which means that the amount of blur can be kept below the acceptable circle of confusion even for foreign objects located at the very front of the bottle. However, with an acceptable circle of confusion of 7p (28.78μm), the rear depth of field becomes 68mm, which does not cover foreign objects located at the very back of the bottle.
[0163] Furthermore, when shooting with Fno=19, the acceptable circle of confusion is 5p (20.56μm) and the depth of field is 54mm, and the acceptable circle of confusion is 7p (24.15μm) and the depth of field is 89mm. Therefore, the amount of blur can be kept below the acceptable circle of confusion from the foreign object located closest to the viewer to the foreign object located furthest from the viewer to the viewer.
[0164] As described above, in the example shooting system 3, by shooting with an aperture value of Fno=19 or higher as a shooting parameter, it is possible to capture images with the amount of blur suppressed to below the acceptable circle of confusion, from the foreign object located closest to the front of the bottle to the foreign object located furthest back. However, if the aperture is stopped down any further, the depth of field will increase, but blur due to diffraction will occur, reducing the ability to detect minute foreign objects. Compared with the example shooting system 1, it can be seen that the optimal imaging conditions are limited.
[0165] On the other hand, reducing the sensor size increases the depth of field, but to ensure sufficient resolution for foreign object detection, it is necessary to narrow the sensor pixel pitch p and increase the number of sensor pixels. Depending on the color of the bottle or liquid, it may be necessary to use near-infrared light (800-2500 nm), so if the sensor pitch is less than 3 μm, the sensor sensitivity to near-infrared light (for example, around 850 nm) will be insufficient.
[0166] The sensor pixel count should be between 10 million and 20 million pixels, and the sensor pixel pitch should be 3 μm or more. Then, the suitable sensor size will be from 1 / 1.7 inch (x=7.5 mm) to APS-C size (x=22.5 mm), preferably from 2 / 3 inch (x=8.8 mm) to 4 / 3 inch (x=17.3 mm).
[0167] Furthermore, considering the case where the bottle is a square bottle with a depth dimension of 2t, as in the example of the shooting system 2, as shown in Figure 14, the foreign object 215 located at the furthest point from the camera (back side) of the bottle appears to be at a position 2t / n from the front side of the bottle according to Equation 9. From the center of the bottle, it is at a position t(2-n) / n. The foreign object 216 located at the closest point from the camera (front side) of the bottle is at a position t from the center of the bottle. If the camera's focus position is shifted forward by t(n-1) / n from the center of the bottle, the foreign objects 215 and 216 will be located at t / n with respect to the focus position.
[0168] Similar to the case of the round bottle, near the camera's sensor surface 217, the image 215i of the foreign object 215 is formed on the near side of the sensor surface, and on the sensor surface, it is defocused and blurred to a diameter δ. Near the camera's sensor surface 217, the image 216i of the foreign object 216 is formed on the far side of the sensor surface, and on the sensor surface, it is defocused and blurred to a diameter δ.
[0169] The above δ is acceptable as long as it is less than or equal to the acceptable circle of confusion, which is the blur tolerance. As explained in Figure 16, the foreign objects 215 and 216 on the far and near sides of the bottle are not affected by the curvature of the bottle, so it is acceptable if the minimum circle of confusion diameter is 5 pixels or less. Converting this to depth of field, if t / n, n=1.333, then if they are within the rear depth of field of 3 / 4t, the foreign objects 215 and 216 can be detected below the blur tolerance.
[0170] On the other hand, the foreign object 216 on the front side of the bottle only needs to have a minimum circle of confusion diameter of 5 pixels or less, as explained in Figure 16. That is, if the container is a square bottle with a depth dimension of 2t in the direction of the optical axis of the optical system, the lens focal length f, distance D1, aperture F value, etc. as shooting parameters only need to satisfy equation 14 for the front depth of field and rear depth of field.
[0171]
number
[0172] Figure 20 is a diagram showing Table 3, which illustrates the relationship between the total depth of field (front and rear depth of field combined) and the allowable circle of confusion when the lens aperture is changed in the example imaging system 2. Here, a camera with a sensor size of 1.1 inches (x=14.1 mm) and a sensor pixel pitch of p=3.45 μm, and a lens with a focal length of f=25 mm and a distance D1=430 mm were used to capture an image with an imaging range of 230 mm in height and 168 mm in width. The bottle was a square bottle of olive oil with a thickness of 2t=66 mm.
[0173] Figure 20 shows the front and rear depth of field when the lens Fno is changed. It is sufficient if the front and rear depth of field are below a predetermined acceptable circle of confusion, i.e., the acceptable amount of blur, and cover the liquid area inside the bottle. In the case of foreign objects located inside the bottle, it is sufficient if the front and rear depth of field is greater than or equal to the bottle depth t / n (25 mm) and the acceptable circle of confusion is 5p (17.25 μm) or less.
[0174] In Table 3, areas where the front and rear depth of field is greater than or equal to the bottle depth t / n (25mm) are shown in bright colors, while other areas are shown in gray. For example, if you shoot with the aperture open at Fno=2.8, none of the depth-of-field values in the table will cover the liquid inside the bottle. If you shoot with the aperture at Fno=5.6, the acceptable circle of confusion is 5p (17.25μm), the front depth of field is 27mm, and the rear depth of field is 31mm. Therefore, the amount of blur can be kept below the acceptable circle of confusion from the foreign object at the very front of the bottle to the foreign object at the very back.
[0175] As described above, in the example shooting system 3, by setting the aperture value as an shooting parameter to Fno=5.6 or higher, it is possible to capture images with the amount of blur suppressed to below the acceptable circle of confusion, from the foreign object located closest to the front of the bottle to the foreign object located furthest away. When the aperture is stopped down significantly, the depth of field increases, but blur due to diffraction occurs, reducing the ability to detect minute foreign objects. Therefore, it is desirable to set the aperture to around Fno=5.6 to 13.
[0176] Furthermore, it is desirable that the characteristic data of the container and the liquid inside the container be stored in advance in the non-volatile memory 303 of the client terminal 201 or in external memory such as the data acquisition server 211. The characteristic data of the container and the liquid inside the container includes at least one of the following: the length L of the shooting range, data related to the curvature of the container (radius r or curvature 1 / radius r, etc.), whether it is a round or square bottle, the depth dimension t, the refractive index of the container, and the refractive index n of the liquid.
[0177] Furthermore, it is desirable to pre-store characteristic data of the optical system and image sensor in external memory, such as the non-volatile memory 303 in the camera 202 or the data acquisition server 211 in the client terminal 201. The characteristic data of the image sensor includes at least one of the following: lens Fno, allowable circle of confusion diameter, sensor size x, sensor pixel pitch p, focal length f, etc.
[0178] The client terminal 201 acquires the characteristic data of the container and the liquid inside the container, as well as the characteristic data of the image sensor, by reading them. Alternatively, it acquires the characteristic data entered by the user. Then, it acquires and displays the shooting parameters by performing calculations using tables or formulas, for example, as shown in Figures 18 to 20.
[0179] In this process, regarding the optical axis direction of the optical system, shooting parameters are acquired according to characteristic data acquired by the acquisition means, such that the shooting range from the front to the back of the container is within the depth of field of the optical system. The shooting parameters include at least one of the distance between the container and the imaging means, the focal length of the optical system, and the aperture value of the optical system.
[0180] In the above example shooting systems, the shooting parameters (distance D1 between the camera and the bottle, focal length f, aperture value Fno, etc.) are displayed on the screen according to the shooting range L in the height direction of the bottle, radius r, or depth dimension t, and the user adjusts and sets them. However, these shooting parameters may also be changed automatically to meet the shooting conditions. The optical system may be a fixed-focus lens or a zoom lens with adjustable focal length.
[0181] Furthermore, the focal length f and aperture value Fno of the zoom lens may be automatically changed according to characteristic data such as the shooting range L in the height direction of the bottle, radius r, or depth dimension t, in order to satisfy the shooting conditions. In other words, at least one of the distance D1, focal length f, and aperture value Fno may be automatically adjusted according to the shooting parameters. In addition, the shooting range L may be divided into multiple shooting ranges, and the system may be configured to shoot each divided shooting range with multiple cameras.
[0182] Next, I will explain the LED lighting 205 used for filming. Figures 21(A) and (B) illustrate the range within which foreign matter can be observed from outside the container when a round bottle container is illuminated with parallel illumination light 221. Figure 21(A) shows the case where the bottle diameter 2r = 80 mm and the foreign matter moves 10 mm, 20 mm, etc. from the center of the bottle outwards.
[0183] In Figure 21(A), foreign objects that reflect and refract light can be observed from outside the bottle as reflected and refracted light from the foreign object. Foreign objects that do not transmit light can be observed from outside the bottle as silhouettes against the background of illumination light. In Figure 21(A), of the reflected and refracted light from the foreign object, the light beam directed towards camera 202 does not reach the outside of the bottle due to total internal reflection caused by the curvature of the bottle when the radius position of the foreign object exceeds 26.7 mm from the center of the round bottle.
[0184] Furthermore, Figure 21(B) illustrates that when a round bottle container is illuminated with parallel illumination light 221, a portion of the light beam is refracted by the curved surface of the bottle, making it impossible to uniformly illuminate the inside of the bottle. The silhouette of an opaque foreign object (background illumination) is significantly refracted by the curved surface of the bottle when the diameter of the foreign object exceeds 26.7 mm, making it difficult to observe from the direction of the camera 202.
[0185] In Figure 21(B), areas where foreign objects illuminated by parallel light are difficult to observe from outside the bottle are shown with diagonal lines, and areas where foreign objects are not illuminated are shown with gray. The refractive index of the bottle glass used in the calculation was 1.5, which is slightly higher than the refractive index of the liquid (1.333), but since the difference in refractive index between the two is small, refraction at the interface between them and the thickness of the bottle were considered negligible.
[0186] Figures 21(A) and (B) show the case where the foreign object is illuminated with parallel light, but we will now explain the case where it is illuminated with light beams at various angles. Figures 22(A) and (B) show the range in which foreign objects can be photographed when a round bottle container is illuminated with diffuse light.
[0187] If the foreign object is located 30 mm from the center of the bottle, the light rays that emerge from the foreign object and point downwards in the diagram will enter the side surface of the bottle at an angle exceeding the critical angle, and therefore cannot exit to the outside. Light rays emitted from the foreign object at an angle of ±27.3° will be refracted outwards from the bottle at an exit angle of 89.9°.
[0188] In the diagram, there is no light beam emitted from the foreign object in region B, but light beams are present in regions A and C. Of these, the light beam from region C does not reach camera 202, but the light beam from region A can be captured by the camera. In other words, to narrow the blind spot where the foreign object cannot be observed, the inside of the container should be illuminated with light beams at various angles. Figure 22(B) shows an example of a configuration for this purpose, in which diffuse illumination light 222 is arranged to illuminate the foreign object inside the bottle with light beams at various angles.
[0189] In the diagram, plots marked with ○ indicate locations where the foreign object could be photographed by camera 202, and plots marked with × indicate locations where it could not be photographed. When the foreign object is located on the side of the bottle that is far from camera 202 (diagonally upward in the diagram), the light rays from the foreign object are greatly refracted by the curved surface of the bottle, making it difficult for them to reach camera 202, which is positioned on the opposite side of the bottle from the light source. Therefore, the area shown by the diagonal lines becomes a blind spot, but compared to Figures 21(A) and (B) in the case of parallel illumination, the area of the blind spot can be significantly reduced.
[0190] Figures 23(A) and (B) illustrate diffuse light illumination for illuminating a round bottle container suitable for the imaging system according to imaging system example 1. Figure 23(A) shows an example of an LED panel lighting system 231 that emits diffused light to illuminate foreign objects inside a round bottle container with light beams at various angles. The LED panel lighting system 231 has multiple LEDs arranged in a two-dimensional manner, and a diffuser plate is placed on the light beam emission surface of the multiple LEDs.
[0191] Furthermore, since bottle containers are often colored to protect the contents from ultraviolet and visible light, it is preferable to use near-infrared light in the 780-2500nm range for the LED emission wavelength. In addition, if the bottle contents are red wine, the transmittance of short-wavelength visible light is low, so it is preferable to use red to near-infrared light in the 600-2500nm range for the light source emission wavelength. If the bottle contents are white wine, both the bottle color and the contents often transmit visible light, so ordinary white lighting or inexpensive red LED lighting can be used.
[0192] By using tunable LEDs or switching between multiple wavelengths of illumination, the wavelength of the illumination can be appropriately changed according to the type of liquid or bottle color of the contents of the bottle. To more effectively illuminate foreign objects with light beams at various angles, it is desirable to make the width of the diffuser plate of the LED panel illumination 231 in the bottle diameter direction (left-right direction in Figure 23) larger than the bottle diameter 2r or the depth dimension 2t, as shown in Figure 23(A).
[0193] In other words, it is desirable that the radial width of the diffuser container be greater than the diameter or depth of the container, and that the vertical length of the diffuser container be longer than the length of the container's imaging range. Furthermore, as shown in Figure 23(B), multiple LED panels 231-1, 231-2, and 231-3 may be arranged on the opposite side of the camera 202, surrounding the bottle 203. It goes without saying that, although LEDs are used as the lighting source 205 in this example shooting system, other light sources may also be used.
[0194] <Image Capture Process> Next, Figure 24 is a flowchart showing an imaging sequence using one of the above-described imaging system examples. This process is a detailed explanation of the process S601 in Figure 6. Note that each process in Figure 24 is realized by the CPU 301 in the client terminal 201 loading the program stored in the non-volatile memory 303 into memory 302 and executing it.
[0195] In step S2401, the CPU 301 of the client terminal 201 detects that the bottle has been set in an upright position (correct position) facing the camera, as shown in Figure 7 of the jig housing 204. The user also inputs the type of container and the name of the liquid that has been set. Then, in step S2402, the CPU 301 of the client terminal 201 turns on the power of the camera 202 and starts displaying the image in live view mode. The type of container includes the shape, material, color, etc.
[0196] Furthermore, an acquisition step is performed to retrieve characteristic data from a database containing pre-stored characteristic data for multiple containers and the liquids within them, according to the names of the containers and liquids entered by the user. At this time, an acquisition step is also performed to retrieve characteristic data for the optical system and image sensor from memory, and a control step is performed to acquire shooting parameters using calculations with tables or formulas, for example, as shown in Figures 18 to 20. The shooting parameters are then displayed on the screen.
[0197] Then, for example, the user adjusts the above shooting parameters, or the shooting parameters are automatically adjusted before shooting. This ensures that the shooting range, at least from the front to the back of the container's interior, is reliably captured within the depth of field of the optical system, along the optical axis.
[0198] In step S2403, the CPU 301 of the client terminal 201 waits until it can determine whether the jig housing 204 has started rotating and the bottle has started rotating. Once it is determined that rotation has started, the bottle is gradually tilted in the vertical direction.
[0199] Triggered by this tilting motion, in step S2404 the CPU 301 of the client terminal 201 automatically starts video mode, begins an illumination step to illuminate the container, and begins an imaging step to image the illuminated container through the optical system. The video captured by the camera is then recorded on a recording medium (not shown) in the camera 202, for example. In step S2405 the CPU 301 of the client terminal 201 determines whether the tilt of the jig housing 204 has reached a predetermined angle (for example, approximately 180 degrees, where the bottle is upside down) and continues operation until it reaches that angle.
[0200] If it is determined in step S2405 that the tilt has reached a predetermined angle, the CPU 301 of the client terminal 201 stops the rotational operation of the jig housing 204 in step S2406. In step S2407, the CPU 301 of the client terminal 201 waits until it is determined that a predetermined time has elapsed with the bottle almost upside down. If it is determined in step S2407 that the predetermined period has elapsed, the CPU 301 of the client terminal 201 stops the video recording operation of the camera 202 in step S2408.
[0201] Next, in step S2409, the CPU 301 of the client terminal 201 generates and writes a video file to the recording medium. In step S2410, the CPU 301 of the client terminal 201 starts the operation to return the jig housing 204 to its correct position and determines whether or not it has returned to the correct position. If it is determined in step S2410 that it has returned to the correct position, the process proceeds to step S2411, and the CPU 301 of the client terminal 201 turns off the camera's power.
[0202] Then, in step S2412, the CPU 301 of the client terminal 201 receives confirmation that the bottle to be inspected has been removed and terminates the flow for the shooting operation. In addition, the video file on the recording medium in the camera 202, which was generated in step S2409, is automatically uploaded to the data acquisition server 211 via wireless communication or wired cable through the client terminal 201.
[0203] According to the embodiments described above, foreign matter inside the liquid container (bottle) can be detected with higher accuracy. Furthermore, in the above-described embodiment, the bottle is changed from a first position where the opening is facing upward relative to the direction of gravity to a second position where the bottle is inverted or tilted at an angle of less than 90 degrees from the inverted position (i.e., a position where the opening is facing downward relative to the direction of gravity). The image of the bottle taken at that time is used as training data (image set 101). This makes it easier to distinguish between the above-described types (1) and (2) foreign objects and improves the accuracy of foreign object detection.
[0204] Then, a trained model is generated using machine learning (training process) with this training data. In this way, since the presence or absence of foreign objects is determined using machine learning rather than a rule-based approach, the accuracy can be improved.
[0205] Furthermore, the number of type (2) images included in the training data is adjusted through a processing step to ensure a sufficient number. This further improves the accuracy of foreign object detection. Additionally, by adjusting the number of type (1) images to prevent them from becoming too numerous, the time required for training is reduced, enabling efficient training while minimizing accuracy degradation.
[0206] In the above embodiment, we described an example in which, by inputting a captured image into a trained model, the result of the estimation process in step S612 is either "foreign object present" or "foreign object absent." However, it is also possible to output information indicating the location of the foreign object in the image. Conversely, it is sufficient if the result of the estimation process in step S612 is either "foreign object present" or "foreign object absent." Therefore, in the step of inputting training data into the trained model in step S604 of the training phase, the foreign object location may not be input as a training label.
[0207] Furthermore, the above-described embodiment described an example in which the bottle's opening is changed from a first position where the opening is facing upward relative to the direction of gravity to a second position where the bottle is inverted or tilted at an angle of less than 90 degrees from the inverted position (i.e., a position where the opening is facing downward relative to the direction of gravity). However, it is not limited to this. As long as it is possible to take a photograph in a state where either type (1) or type (2) movement can be observed, foreign object detection may be performed using an image taken when the bottle is changed from the second position to the first position.
[0208] In other words, it is sufficient to take a picture of the container when its orientation is changed between a first orientation in which a specific part of the container containing the liquid is facing upwards relative to the direction of gravity, and a second orientation in which the specific part is facing downwards relative to the direction of gravity.
[0209] Furthermore, the first posture is not limited to a state where the opening is facing upwards; in short, any change in posture that allows observation of the movement of a foreign object of either type (1) or type (2) is acceptable. That is, by changing the posture of a part of the container other than the opening between a posture where it is facing upwards relative to the direction of gravity and a posture where it is facing downwards relative to the direction of gravity, it is possible to photograph the movement of a foreign object, and then use that image to detect the foreign object.
[0210] Furthermore, the above-described embodiment is applicable to the inspection of liquid-containing bottles such as alcoholic beverages like wine and whiskey, soft drinks like juice and natural water, beverages and foods like seasonings, chemicals like pesticides and alcohol disinfectants, and hygiene and beauty products like shampoo, liquid soap, and lotions. In addition, by using the imaging system described above, it is possible to use images captured under suitable imaging conditions according to the container shape and / or liquid type, thereby improving inspection accuracy.
[0211] Furthermore, when applied to wine, the characteristic data of the container related to determining the imaging conditions may include at least one of the following: the bottle shape type such as Bordeaux type, Burgundy type, or Champagne type; bottle height; bottle diameter or depth; and bottle glass thickness. Additionally, the type of wine, such as red wine, white wine, rosé, or sparkling wine, may be included as characteristic data. Moreover, the training data may be separated for each type of wine, and the accuracy can be further improved by performing estimation processing using the trained models generated for each type of wine.
[0212] Thus, according to the above-described embodiment, foreign matter in a liquid container can be detected with greater accuracy, regardless of the container shape or other characteristics. Furthermore, the various controls described above may be performed by a single piece of hardware, or multiple pieces of hardware (for example, multiple processors or circuits) may share the processing to control the entire device.
[0213] Furthermore, although the present invention has been described in detail based on its preferred embodiments, the present invention is not limited to these specific embodiments, and various forms that do not depart from the spirit of the invention are also included in the present invention. Moreover, each of the embodiments described above is merely one embodiment of the present invention, and it is possible to combine each embodiment as appropriate.
[0214] (Other embodiments) The present invention can also be realized by performing the following process: supplying software (programs) that realize the functions of the embodiments described above to a system or device via a network or various storage media; and having the computer (or CPU, MPU, etc.) of that system or device read and execute the program code.
[0215] In this case, the program and the storage medium storing the program constitute the present invention. A computer may have one or more processors or circuits and may include a plurality of separate computers or a network of a plurality of separate processors or circuits for reading and executing computer executable instructions.
[0216] A processor or circuit may include a central processing unit (CPU), a microprocessing unit (MPU), a graphics processing unit (GPU), an application-specific integrated circuit (ASIC), or a field-programmable gateway (FPGA). Alternatively, a processor or circuit may include a digital signal processor (DSP), a dataflow processor (DFP), or a neural processing unit (NPU). [Explanation of Symbols]
[0217] 202: Camera 203: Bottle 204: Jig housing 205: LED lighting 213: Estimated Server 214: Center of rotation 215: Foreign matter (serious) 215i: Image of foreign object 215 216: Foreign matter (low specific gravity) 216i: Image of foreign object 216 217: Sensor surface 218: Shooting lens 219: Sensor Pixel 221: Illumination light 222: Diffuse illumination light
Claims
1. The acquisition means captures images of the container when its posture is changed between a first posture in which a specific part of the container containing liquid is upward relative to the direction of gravity, and a second posture in which the specific part is downward relative to the direction of gravity, as the image to be detected. A control means controls the input of the image acquired by the acquisition means to a trained model, and controls the notification of the foreign object detection result based on the estimation result information obtained as a result of the input. The system includes a generation means for generating a trained model for estimating whether or not there is a foreign object in a container by performing machine learning using a set of images consisting of multiple images obtained by photographing the containers containing the liquid when they are changed between a first and a second posture, as training data. The generating means further includes an adjusting means that adjusts the number of images in a first image group consisting of multiple images obtained by photographing the containers containing the liquid when they are changed between a first posture and a second posture, and generates a second image group in which the number of second type images in which the foreign object is floating or rising in the liquid is greater than the number of first type images in which the foreign object is falling from top to bottom. The generation means performs the machine learning using the second group of images as training data. A foreign object detection system characterized by the following features.
2. The acquisition means comprises an illumination means for illuminating the container and an imaging means for imaging the container illuminated by the illumination means. The foreign object detection system according to claim 1, characterized in that the acquisition means illuminates the container with the illumination means when the orientation of the container is changed from the first orientation to the second orientation, and acquires the image captured by the imaging means.
3. The foreign object detection system according to claim 1, further characterized in that the adjusting means generates the second image group by adjusting the number of images such that the number of images without foreign objects is greater than the number of images of the first type.
4. The foreign object detection system according to claim 3, further characterized in that the adjusting means generates the second image group by adjusting the number of images such that the number of images without foreign objects is greater than the number of images of the first type, and the number of images of the second type is greater than the number of images without foreign objects.
5. The foreign object detection system according to any one of claims 1 to 4, characterized in that the generation means performs machine learning using a label that indicates whether or not a foreign object is present, which is associated with each image included in the training data.
6. The foreign object detection system according to any one of claims 1 to 5, characterized in that the aforementioned specific portion is a sealing opening.
7. The foreign object detection system according to any one of claims 1 to 6, characterized in that the container is light-transmitting.
8. The acquisition means captures images of the container when its posture is changed between a first posture in which a specific part of the container containing liquid is upward relative to the direction of gravity, and a second posture in which the specific part is downward relative to the direction of gravity, as the image to be detected. A control means controls the input of the image acquired by the acquisition means to a trained model, and controls the notification of the foreign object detection result based on the estimation result information obtained as a result of the input. The system includes a generation means for generating a trained model for estimating whether or not there is a foreign object in a container by performing machine learning using a set of images consisting of multiple images obtained by photographing the containers containing the liquid when they are changed between a first and a second posture, as training data. The generating means further includes an adjusting means that adjusts the number of images in a first image group consisting of multiple images obtained by photographing the containers containing the liquid when they are changed between a first posture and a second posture, and generates a second image group in which the number of second type images in which the foreign object is floating or rising in the liquid is greater than the number of first type images in which the foreign object is falling from top to bottom. The adjustment means generates the second image group by adjusting the number of images so that the number of images without foreign objects is greater than the number of images of the first type. A foreign object detection system characterized by the following features.
9. The foreign object detection system according to claim 8, characterized in that the generation means performs machine learning using the second group of images as training data.
10. The foreign object detection system according to claim 8 or 9, further characterized in that the adjusting means generates the second image group by adjusting the number of images such that the number of images without foreign objects is greater than the number of images of the first type, and the number of images of the second type is greater than the number of images without foreign objects.
11. The acquisition means captures images of the container when its posture is changed between a first posture in which a specific part of the container containing liquid is upward relative to the direction of gravity, and a second posture in which the specific part is downward relative to the direction of gravity, as the image to be detected. A control means controls the input of the image acquired by the acquisition means to a trained model, and controls the notification of the foreign object detection result based on the estimation result information obtained as a result of the input. The system includes a generation means for generating a trained model for estimating whether or not there is a foreign object in a container by performing machine learning using a set of images consisting of multiple images obtained by photographing the containers containing the liquid when they are changed between a first and a second posture, as training data. The generating means further includes an adjusting means that adjusts the number of images in a first image group consisting of multiple images obtained by photographing the containers containing the liquid when they are changed between a first posture and a second posture, and generates a second image group in which the number of second type images in which the foreign object is floating or rising in the liquid is greater than the number of first type images in which the foreign object is falling from top to bottom. The adjustment means generates the second image group by adjusting the number of images such that the number of images without foreign matter is greater than the number of images of the first type, and the number of images of the second type is greater than the number of images without foreign matter. A foreign object detection system characterized by the following features.
12. The foreign object detection system according to claim 11, characterized in that the generation means performs machine learning using the second group of images as training data.
13. The adjustment means further adjusts the number of images so that the number of images without foreign objects is greater than the number of images of the first type, thereby generating the second image group. The foreign object detection system according to claim 11 or 12.
14. A control method for a foreign object detection system, The acquisition step involves capturing an image of the container when its posture is changed between a first posture in which a specific part of the container containing liquid is facing upward relative to the direction of gravity, and a second posture in which the specific part is facing downward relative to the direction of gravity, as the image to be detected. A control step which controls the input of the image acquired in the acquisition step into the trained model, and controls the notification of the foreign object detection result based on the estimation result information obtained as a result of the input, The method includes a generation step of generating a trained model for estimating whether or not there is a foreign object in a container by performing machine learning using a set of images consisting of multiple images obtained by photographing the containers containing the liquid when they are changed to a first position and a second position, as training data. The generation step further includes an adjustment step of adjusting the number of images in a first image group, which consists of multiple images obtained by photographing the containers containing the liquid when they are changed between a first posture and a second posture, and generating a second image group in which the number of second type images in which the foreign object is floating or rising in the liquid is greater than the number of first type images in which the foreign object is falling from top to bottom. The generation step involves performing machine learning using the second set of images as training data. A control method for a foreign object detection system characterized by the following.
15. A control method for a foreign object detection system, The acquisition step involves capturing an image of the container when its posture is changed between a first posture in which a specific part of the container containing liquid is facing upward relative to the direction of gravity, and a second posture in which the specific part is facing downward relative to the direction of gravity, as the image to be detected. A control step which controls the input of the image acquired in the acquisition step into the trained model, and controls the notification of the foreign object detection result based on the estimation result information obtained as a result of the input, The method includes a generation step of generating a trained model for estimating whether or not there is a foreign object in a container by performing machine learning using a set of images consisting of multiple images obtained by photographing the containers containing the liquid when they are changed to a first position and a second position, as training data. The generation step further includes an adjustment step of adjusting the number of images in a first image group, which consists of multiple images obtained by photographing the containers containing the liquid when they are changed between a first posture and a second posture, and generating a second image group in which the number of second type images in which the foreign object is floating or rising in the liquid is greater than the number of first type images in which the foreign object is falling from top to bottom. The adjustment step generates the second image group by adjusting the number of images so that the number of images without foreign objects is greater than the number of images of the first type. A control method for a foreign object detection system characterized by the following.
16. A control method for a foreign object detection system, The acquisition step involves capturing an image of the container when its posture is changed between a first posture in which a specific part of the container containing liquid is facing upward relative to the direction of gravity, and a second posture in which the specific part is facing downward relative to the direction of gravity, as the image to be detected. A control step which controls the input of the image acquired in the acquisition step into the trained model, and controls the notification of the foreign object detection result based on the estimation result information obtained as a result of the input, The method includes a generation step of generating a trained model for estimating whether or not there is a foreign object in a container by performing machine learning using a set of images consisting of multiple images obtained by photographing the containers containing the liquid when they are changed to a first position and a second position, as training data. The generation step further includes an adjustment step of adjusting the number of images in a first image group, which consists of multiple images obtained by photographing the containers containing the liquid when they are changed between a first posture and a second posture, and generating a second image group in which the number of second type images in which the foreign object is floating or rising in the liquid is greater than the number of first type images in which the foreign object is falling from top to bottom. The adjustment step generates the second image group by adjusting the number of images such that the number of images without foreign matter is greater than the number of images of the first type, and the number of images of the second type is greater than the number of images without foreign matter. A control method for a foreign object detection system characterized by the following.
17. An electronic device having each of the means for a foreign object detection system as described in any one of claims 1 to 13.
18. A computer program for causing a computer to function as one of the means of a foreign object detection system described in any one of claims 1 to 13.
19. A computer-readable storage medium storing a computer program for causing the computer to function as one of the means of a foreign object detection system described in any one of claims 1 to 13.