Image acquisition device, learning system, production method for trained model, and inference system
The image acquisition device with non-contact sensors and machine learning enhances the accuracy and speed of reading codes by identifying and processing images of codes, addressing the challenges faced by visually impaired users.
Patent Information
- Application Number
- JP2024032585
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-03-04
- Publication Date
- 2025-09-17
AI Technical Summary
Visually impaired users face difficulties in accurately and quickly reading one-dimensional or two-dimensional codes due to challenges such as code distortion, finger overlap, incomplete capture, or shadowing, leading to incorrect or delayed information retrieval.
An image acquisition device equipped with an imaging unit, non-contact sensors (distance, temperature, luminance, and color sensors) to identify the area of interest in the image, and a communication unit to transmit data for machine learning, enabling accurate code reading by generating a trained model.
Enables accurate and rapid reading of codes by distinguishing between the code and surrounding objects, improving the accuracy and speed of information retrieval for visually impaired users.
Smart Images

Figure 2025134585000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an image acquisition device for collecting images for machine learning, a learning system using the device, and a method for producing a trained model using the device.The present invention also relates to an image acquisition device and an inference system that perform inference using a trained model.In particular, the present invention relates to a technology for collecting images that include readable codes such as one-dimensional codes (barcodes) and two-dimensional codes (QR Codes (registered trademark)). [Background technology]
[0002] A system has been proposed for some time now in which, for example, a matrix-type two-dimensional code is attached to product packaging for the visually impaired, and when the code is read by a camera-equipped information terminal such as a smartphone when shopping, the name and detailed information of the product are read out loud (Patent Document 1). [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2020-135794 Summary of the Invention [Problem to be solved by the invention]
[0004] However, it is often difficult for visually impaired users to operate an information terminal with one hand while holding a product with the other, and ensure that the code on the product's packaging is within the camera's field of view. In particular, because they cannot visually confirm the orientation of the camera or the position of the code on the product, they may not be able to accurately and quickly read the information contained in the code from the captured image if the code is distorted, if a finger is overlapping part of the code, if part of the code is not captured, or if the code is in shadow or bright light.
[0005] If the information terminal cannot accurately read the information contained in the code, it will provide the user with incorrect information, which will confuse the user. Also, if the information terminal is slow to read the code, no information will be read out even if the user has the code within the camera's field of view, which will make the user feel uneasy.
[0006] Therefore, a main object of the present invention is to provide a technique that enables one-dimensional or two-dimensional read codes attached to products and the like to be read accurately and quickly. [Means for solving the problem]
[0007] The inventors of the present invention have thoroughly investigated means for solving the problems of the above-mentioned conventional inventions, and have come up with the idea of improving the accuracy and speed of code reading by collecting a large number of captured images of readable codes attached to products, etc., performing machine learning, and then inferring the information contained in the codes using the resulting trained model. In particular, the inventors have discovered that effective machine learning can be performed by using captured images of the codes along with additional information, such as detection information acquired by a non-contact sensor, as training data. Based on this discovery, the inventors have realized that the problems of the conventional inventions can be solved, and have completed the present invention. Specifically, the present invention has the following configurations or steps.
[0008] A first aspect of the present invention relates to an image acquisition device. The image acquisition device according to the present invention is primarily used to acquire images and information to be used as training data for machine learning. Specifically, the image acquisition device is configured to acquire images of readable codes, including one-dimensional and two-dimensional codes, attached to objects. One-dimensional codes include barcodes established by international standards such as JAN, EAN, and UPC, while two-dimensional codes include matrix codes such as QR Code (registered trademark). The image acquisition device includes an imaging unit, a sensor unit, and an image processing unit. The imaging unit acquires an image. The sensor unit is a non-contact sensor that detects an object within a detection range that at least partially overlaps with the imaging range of the imaging unit. Examples of the object here include a person's hand holding the object (e.g., a product) or a container (e.g., a shopping cart), which are different from the object itself. When the image acquired by the imaging unit includes a readable code, the image processing unit identifies an area in the image where the object is detected based on detection information from the sensor unit. If the area (coordinate information, etc.) of an object such as a human hand reflected in a captured image can be identified, it will be possible to more accurately extract the object and the code attached to it from the captured image, and this can be effectively used as training data for machine learning. Furthermore, since such detection information can also be used to infer the way the user holds the object (product, etc.), there is a high possibility that the size and type of the object can also be determined by using the captured image and detection information as training data.
[0009] In the image acquisition device according to the present invention, the sensor unit may include a distance measuring sensor that measures the distance to an object. In this way, the area of an object, such as a human hand, reflected in the captured image may be identified based on the distance from the sensor unit to the object. Furthermore, it is preferable that the distance measuring sensor detects objects within a distance range of 10 to 100 cm. The range of 10 to 100 cm is the range within which a user can typically hold an object (such as a product) in their hand. Therefore, it can be inferred that an object reflected in the captured image within this range is likely to be a human hand.
[0010] In the image acquisition device according to the present invention, the sensor unit may include a temperature sensor that measures the temperature of an object. In particular, it is preferable that the temperature sensor detects an object in a temperature range of 20 to 40°C (particularly 30 to 40°C). Since the temperature of a human finger is generally 30 to 40°C, the use of a temperature sensor makes it possible to effectively identify the area of a human hand that appears in the captured image.
[0011] In the image acquisition device according to the present invention, the sensor unit may include a luminance sensor that measures the luminance of an object. The sensor unit may also include a color sensor that measures the color of the object. Using the luminance sensor or color sensor in this way makes it possible to distinguish between an object (e.g., a product) and an object (e.g., a human hand) contained in a captured image. Furthermore, by having a user wear gloves that become brighter when captured in an image or gloves with a distinctive color and hold the object, it becomes even easier to distinguish between the object and the human hand (glove).
[0012] The image acquisition device according to the present invention may further include a communication unit. The communication unit transmits a dataset including a captured image and information (such as coordinates) relating to an area in the captured image where an object is detected to an external server device via a communication line such as the Internet. Machine learning is performed in the external server device that receives the dataset. If the image acquisition device does not include a communication unit, the dataset may be temporarily stored in the image acquisition device and transferred to a machine learning device via an exchangeable recording medium such as a memory card or a communication cable.
[0013] In the image acquisition device according to the present invention, the image processing unit may further perform masking on a region of the captured image where an object is detected. Specifically, masking is a process of clipping or hiding a specific region from the captured image, thereby eliminating information contained in the specific region itself. The image processing unit may generate a processed image by processing the original image to mask the object detection region, or may generate a mask image having mask information that identifies the object detection region without processing the original image itself. The image processing device may also transmit the processed image that has undergone masking processing to an external server device via a communication line using the communication unit. Alternatively, the image processing device may transmit the captured image (original image) and the mask image obtained by the masking processing to an external server device via a communication line using the communication unit. Performing such masking processing in the image processing device makes the captured image containing the read code even more useful as data for machine learning. Performing such masking processing in the image processing device also reduces the processing load on the machine learning server device.
[0014] A second aspect of the present invention relates to a learning system including an image acquisition device and a server device. The image acquisition device and the server device are connected to each other via a communication line. The image acquisition device here basically has the same configuration as that of the first aspect. That is, the image acquisition device is for acquiring an image of a readable code, including a one-dimensional code and a two-dimensional code, attached to an object, and includes an imaging unit for acquiring the image, a non-contact sensor unit for detecting the object in a detection range that at least partially overlaps with the imaging range of the imaging unit, an image processing unit for identifying an area in the image where the object is detected based on detection information from the sensor unit when the readable code is included in the captured image, and a communication unit for transmitting a dataset including the captured image and information about the area in the image where the object is detected to the server device via the communication line. The server device includes a learning unit that performs machine learning using the dataset received from the image acquisition device and generates or updates a trained model. The trained model generated or updated here is suitable for use in inferring products to which readable codes are assigned.
[0015] A third aspect of the present invention is a method for producing a trained model executed by an image acquisition device and a server device. A trained model is model data in which parameters (so-called "weights") have been adjusted by performing machine learning. In the trained model production method according to the present invention, first, the image acquisition device acquires an image (image acquisition step), detects an object using a non-contact sensor in a detection range that at least partially overlaps with the imaging range of the imaging unit (detection step), and, if the captured image contains a read code, identifies the area in the captured image in which the object was detected based on the detection information from the sensor unit (area identification step), and transmits a dataset including the captured image and information about the area in the image in which the object was detected to the server device via a communication line (transmission step). The server device then performs machine learning using the dataset received from the image acquisition device to generate or update a trained model (model creation step).
[0016] A fourth aspect of the present invention relates to an inference system for inferring products and the like using the trained model obtained as described above. The inference system includes an image acquisition device for acquiring an image of a read code, including a one-dimensional code and a two-dimensional code, attached to an object, and a server device connected to the image acquisition device via a communication line. The image acquisition device includes an imaging unit, a sensor unit, an image processing unit, and a communication unit. The imaging unit acquires an image. The sensor unit is a non-contact sensor that detects an object in a detection range that at least partially overlaps with the imaging range of the imaging unit. When the read code is included in the captured image, the image processing unit identifies an area in the image where the object is detected based on detection information from the sensor unit. The communication unit transmits a dataset including the captured image and information regarding the area in the image where the object is detected to the server device via the communication line. The server device includes a trained model, a determination unit, and a communication unit. The trained model is a trained model that has been generated or updated by performing machine learning using a dataset previously received from the image acquisition device. When a new data set is received from the image acquisition device, the determination unit uses the trained model to determine the information contained in the read code or the product to which the read code is assigned. The information contained in the read code is product identification information expressed in numbers or letters. The communication unit transmits the determination result by the determination unit to the image acquisition device. With this system, even if the image processing device itself does not have the trained model, the inference result based on the trained model can be obtained from the server device.
[0017] A fifth aspect of the present invention relates to an image acquisition device having a trained model. Similar to the inference system described above, the image acquisition device has an imaging unit, a sensor unit, and an image processing unit. In addition to these elements, the image acquisition device further has a trained model and a determination unit. The trained model is a trained model that is generated or updated by performing machine learning using a dataset containing previously captured images and information about areas in the images where objects are detected. When a new dataset is acquired, the determination unit uses the trained model to determine the information contained in the readable code or the product to which the readable code is assigned. In this way, when the image processing device itself has a trained model, the image processing device can obtain inference results using the trained model without communicating with a server device. [Effects of the Invention]
[0018] According to the present invention, it becomes possible to accurately and quickly read one-dimensional or two-dimensional read codes attached to products and the like. [Brief explanation of the drawings]
[0019] [Figure 1] Figure 1 shows an example of an overall learning and inference system. [Figure 2] FIG. 2 shows an example of a terminal device (image processing device). [Figure 3] Figure 3 is a block diagram showing the configuration of the learning and inference system. [Figure 4] FIG. 4 is a flow diagram showing an example of the learning phase performed by the system. [Figure 5] FIG. 5 shows an example of a method for acquiring a product image by a terminal device. [Figure 6] FIG. 6 shows an example of mask processing by a terminal device. [Figure 7] FIG. 7 is a flow diagram showing an example of the inference phase performed by the system. DETAILED DESCRIPTION OF THE INVENTION
[0020] Hereinafter, embodiments of the present invention will be described with reference to the drawings. The present invention is not limited to the embodiments described below, but also includes appropriate modifications of the embodiments below within the scope obvious to those skilled in the art.
[0021] FIG. 1 schematically illustrates the overall configuration of a system 100 according to an embodiment of the present invention. The system 100 according to this embodiment is intended to use a terminal device 10 to collect captured images of barcodes C attached to objects T, such as products, and to use the captured images as training data for machine learning. While the example shown in FIG. 1 depicts only one terminal device 10, in reality, multiple terminal devices 10 collect captured images of various barcodes. As barcodes, well-known one-dimensional codes defined by international standards, such as JAN, EAN, and UPC, may be used. Furthermore, the system 10 according to this embodiment uses a trained model obtained through the above-described machine learning to quickly and accurately read barcodes included in captured images by the terminal device 10. Therefore, the system 100 according to the present invention is intended for use in a case where, when the terminal device 10 reads the barcode C of a product (object T) displayed in a store or the like, the name and detailed information of the product bearing the barcode C are output as audio from the terminal device 10. The system 100 according to the present invention is particularly suited to assisting visually impaired users in their shopping.
[0022] Currently, barcodes are widely used as identification codes attached to products, etc., and therefore, in this embodiment, captured images of these barcodes are collected. However, the captured images collected in this invention are not limited to those of barcodes, and may also be those of matrix-type codes such as QR Code (registered trademark) or other two-dimensional codes.
[0023] As shown in FIG. 1, the system 100 includes a terminal device 10, a server device 20, and a product management device 30. An example of the terminal device 10 is a wearable device. By using the terminal device 10 as a wearable device, a user's hands are not obstructed when using the terminal device 10, making it easier for the user to read a barcode C attached to an object T while using the terminal device 10. Examples of wearable devices include neck-worn devices, eyeglass-type devices, head-worn devices, and wristwatch-type devices. In this embodiment, a neck-worn wearable device is used as the terminal device 10. The terminal device 10 is equipped with a digital camera for capturing images of barcodes. In the system 10 according to this embodiment, captured images captured by the terminal device 10 are transmitted to the server device 20 via a communication line such as the Internet. The server device 20 performs machine learning using captured images of a wide variety of barcodes collected from multiple terminal devices 10 to generate a trained model. The server device 20 may be implemented as a single web server or multiple web servers interconnected via a network. The product management device 30 is connected to the terminal device 10 and the server device 20 via the Internet, and manages identification information assigned to a barcode in association with information about the object T (product) to which the barcode is attached. Information about the object T is provided to the terminal device 10 and the server device from the product management device 30 in a timely manner.
[0024] Next, the configuration of the system 100 according to one embodiment of the present invention will be described in more detail. Fig. 2 is an external perspective view showing an example of the terminal device 10. Fig. 3 shows an example of the hardware and software configuration of the terminal device 10, server device 20, and product management device 30 that make up the system 100.
[0025] As shown in FIGS. 2 and 3 , the terminal device 10 in this embodiment is a neck-hanging wearable device. The terminal device 10 includes a left arm portion, a right arm portion, and a main body portion that connects them to the back of the wearer's neck. When wearing the terminal device 10, the main body portion is placed in contact with the back of the wearer's neck, and the left arm portion and the right arm portion are hung down from the sides of the wearer's neck toward the chest, so that the entire device is hung around the neck. Various electronic components are stored within the housing of the terminal device 10. Configuring such a terminal device 10 as a neck-hanging wearable device frees both hands of the user wearing the terminal device 10. This makes it easier to have the wearable device read barcodes on product packaging or the like while holding the product in one's hand.
[0026] The left arm and the right arm are each provided with a plurality of sound collection units 14 (microphones). The sound collection units 14 are arranged mainly for the purpose of capturing the wearer's voice, surrounding sounds, and even conversations between the wearer and a person speaking. It is preferable to use an omnidirectional (non-directional) microphone as the sound collection unit 14 so that it can widely collect sounds generated around the wearer. Any known microphone, such as a dynamic microphone, a condenser microphone, or a MEMS (Micro Electrical Mechanical Systems) microphone, may be used as the sound collection unit 14. The sound collection unit 14 converts sound into an electrical signal, amplifies the electrical signal using an amplifier circuit, converts the signal into digital information using an A / D conversion circuit, and outputs the digital information to the control unit 11. The sound signal acquired by the sound collection unit 14 is transmitted to the control unit 11 provided within the housing. The sound signal acquired by the sound collection unit 14 can also be transmitted to the server device 20 via the Internet via the communication unit 13.
[0027] An imaging unit 15 is further provided on the left arm. Specifically, the imaging unit 15 is provided on the distal end surface of the right arm, and can capture still images and moving images of the wearer's front side. The images captured by the imaging unit 15 are transmitted to the control unit 11 in the housing and stored as image data. A typical digital camera may be used as the imaging unit 15. The imaging unit 15 is composed of, for example, a photographing lens, a mechanical shutter, a shutter driver, a photoelectric conversion element such as a CCD image sensor unit, a digital signal processor (DSP) that reads the charge amount from the photoelectric conversion element and generates image data, and an IC memory. The image data captured by the imaging unit 15 is supplied to the control unit 11 and stored in the memory unit 12. The image data may also be subjected to a predetermined image analysis process. The still images and moving images captured by the imaging unit 15 can be transmitted to the control device 20 via the communication unit 13 and the Internet.
[0028] A non-contact object sensor 16 is provided on the right arm. The object sensor 16 is provided on the distal end surface of the left arm, and can detect information about objects mainly included in the imaging range of the imaging unit 15. That is, the imaging range of the imaging unit 15 and the detection range of the object sensor 16 largely overlap, and when an object enters the imaging range of the imaging unit 15, information about the object can be detected by the object sensor 16. Examples of the object sensor 16 are a distance sensor 16a, a temperature sensor 16b, a brightness sensor 16c, a color sensor 16d, and a gesture sensor 16e, as shown in the block diagram of FIG. 3. Information detected by the object sensor 16 is transmitted to the control unit 11 and used to control each element and process the captured image. In addition, the information detected by the object sensor 16, together with the captured image, can also be used as training data for machine learning. Note that the object sensor 16 does not need to include all of the sensors 16a to 16e; it is sufficient to include one or more of these sensors.
[0029] The distance measurement sensor 16a measures the distance to a target with a barcode or other object. The distance measurement sensor 16a can be, for example, an optical or sonic ToF (Time-of-Flight) sensor. The ToF sensor measures the distance to the target by emitting pulsed infrared rays or sound waves toward the target and detecting the reflected light or sound. The distance measurement sensor 16a may also be a camera mounted on the left arm in addition to the imaging unit 15 on the right arm, and measure the distance to the target or object using a stereo camera system. The stereo camera system distance measurement sensor 16a simultaneously captures the target with two cameras and calculates distance information from the parallax between the two images. The closer the distance to the target, the greater the parallax.
[0030] The temperature sensor 16b measures the temperature of an object bearing a barcode or other object in a non-contact manner. Examples of the temperature sensor 16b include an infrared thermistor, a thermographic camera, and a laser temperature sensor. An infrared thermistor measures the temperature of an object in a non-contact manner by detecting infrared radiation from the object using an infrared sensor. A thermographic camera uses a camera sensor capable of detecting infrared rays to measure the surface temperature distribution of an object from a thermal image. A laser temperature sensor shines laser light on an object and measures the temperature from the reflected light. Among these, an infrared thermistor is preferably used as the temperature sensor 16b. In this embodiment, the temperature sensor 16b is used primarily for the purpose of detecting a human hand holding an object (such as a product) bearing a barcode. Generally available, inexpensive infrared thermistors can accurately detect temperatures close to human body temperature (20 to 40°C, especially 30 to 40°C), but have difficulty detecting temperatures outside of this range. Since products held by users are often below 20°C, using an infrared thermistor makes it easy to detect the hand (object) holding the product without detecting the product (object) with a barcode attached.
[0031] The brightness sensor 16c measures the brightness of an object bearing a barcode or other object. An example of the brightness sensor 16c is a photodiode. A photodiode measures the intensity (brightness) of light by measuring the current generated by incident light. The color sensor 16d measures the color value of an object bearing a barcode or other object. An example of the color sensor 16d is an image sensor with a dye filter. An image sensor with a dye filter can extract color information by detecting light through filters of each of the RGB colors. The brightness sensor 16c and the color sensor 16d are used to identify a product bearing a barcode and the user's hand holding it by its brightness or color. For this purpose, for example, having the user wear gloves with a high brightness or a specific color while holding the product makes it easier to identify the product and the user's hand.
[0032] The gesture sensor 16d mainly detects the movement of the wearer's hand on the front side of the terminal device 10. The gesture sensor 16d can detect, for example, the movement and shape of the wearer's fingers. An example of the gesture sensor 16d is an optical sensor that detects the movement and shape of the target by irradiating light from an infrared-emitting LED toward the target and capturing changes in the reflected light with a light-receiving element. Note that detection information by the gesture sensor 16d is transmitted to the control unit 11 and can be used mainly to control the image capture unit 15 and the sound emission unit 18. Specifically, the detection information from the gesture sensor 16d is used to control the activation and deactivation of the image capture unit 15 and the sound emission unit 18. For example, the gesture sensor 16d may detect that an object such as the wearer's hand has approached the gesture sensor 16d and control the image capture unit 15, or may detect that the wearer has made a predetermined gesture within the detection range of the gesture sensor 16d and control the image capture unit 15. The gesture sensor 16d may be replaced with a proximity sensor. The proximity sensor detects, for example, when the wearer's finger approaches within a predetermined range. Known proximity sensors, such as optical, ultrasonic, magnetic, capacitive, and thermal sensors, can be used.
[0033] As shown in FIG. 2, a neck-worn wearable device has a sound emitting unit 18 (speaker) provided on the outer side (opposite side of the wearer) of a main body unit located behind the wearer's neck. In this embodiment, the sound emitting unit 18 is arranged to output sound toward the outer side of the main body unit. By emitting sound from the back of the wearer's neck directly behind the wearer in this manner, the sound output from the sound emitting unit 18 is less likely to reach a conversation partner located directly in front of the wearer. This makes it easier for the conversation partner to distinguish between the voice emitted by the wearer himself / herself and the sound emitted from the sound emitting unit 18 of the terminal device 10. The sound emitting unit 18 is an acoustic device that converts an electrical signal into physical vibrations (i.e., sound). An example of the sound emitting unit 18 is a general speaker that transmits sound to the wearer by air vibrations. Alternatively, the sound emitting unit 18 may be a bone conduction speaker that transmits sound to the wearer by vibrating the wearer's bones. In this case, the sound emitting unit 18 may be provided on the inside (wearer side) of the main body, and the bone conduction speaker may be configured to come into contact with the bones (cervical vertebrae) at the back of the neck of the wearer.
[0034] As shown in FIG. 3, the control unit 11 of the terminal device 10 performs arithmetic processing to control other elements of the terminal device 10. A processor such as a CPU (Central Processing Unit) or a GPU (Graphics Processing Unit) can be used as the control unit 11. The control unit 11 can also process and edit images captured by the imaging unit 15. The control unit 11 basically reads out a program stored in the storage unit 12, loads it into the main memory, and executes predetermined arithmetic processing in accordance with the program. The control unit 11 can also write and read results of calculations performed in accordance with the program to and from the storage unit 12 as appropriate. In this embodiment, the control unit 11 of the terminal device 10 functions as an image processing unit 11a. The function of the image processing unit 11a will be described in detail later.
[0035] The memory unit 12 of the terminal device 10 is an element for storing information used in arithmetic processing and the like in the control unit 11 and the results of the arithmetic processing. The storage function of the memory unit 12 can be realized by a non-volatile memory such as an HDD or an SSD. The memory unit 12 may also function as a main memory for writing or reading intermediate data of arithmetic processing by the control unit 11. The memory function of the memory unit 12 can be realized by a volatile memory such as a RAM or a DRAM. The memory unit 12 may also store ID information specific to the user who owns it. The memory unit 12 may also store an IP address, which is identification information for the terminal device 10 on the network.
[0036] The communication unit 13 of the terminal device 10 is an element for wireless communication with the server device 20 and the product management device 30. To communicate with the server device 20 and the product management device 30 via the Internet, the communication unit 13 may employ a communication module for wireless communication using a known mobile communication standard such as 3G (W-CDMA), 4G (LTE / LTE-Advanced), or 5G, or a wireless LAN system such as Wi-Fi (registered trademark). The communication unit 13 may also employ a communication module for close proximity wireless communication using a system such as Bluetooth (registered trademark) or NFC to directly communicate with another terminal device 10.
[0037] The sensors 17 of the terminal device 10 include, for example, sensor devices for detecting the operation or usage status of the terminal device 10 or biometric information of the wearer. Sensor modules installed in general mobile information terminals or wearable devices may be used as the sensors 17. For example, the sensors 17 include a gyro sensor, an acceleration sensor, a geomagnetic sensor, and a battery sensor. The sensors 17 may also include biosensors for detecting biometric information of the wearer, such as a body temperature sensor, a heart rate sensor, a blood oxygen concentration sensor, a blood pressure sensor, and an electrocardiogram sensor.
[0038] The location information acquisition unit 19 of the terminal device 10 is an element for acquiring current location information of the terminal device 10. Specifically, the location information acquisition unit 19 has a function of performing positioning using a global positioning system (GPS). Based on radio wave transmission time information included in radio waves transmitted from multiple GPS satellites, the location information acquisition unit 19 measures the time required to receive each radio wave and transmits time information indicating the time to the control unit 11. Based on the acquired time information, the control unit 11 can calculate information regarding the latitude and longitude of the location of the terminal device 10. Alternatively, the location information acquisition unit 19 may acquire current location information by scanning radio waves or beacon signals transmitted from wireless base stations such as Wi-Fi (registered trademark) access points.
[0039] 2 and 3, in this embodiment, the terminal device 10 does not have a display device such as a monitor or a display. Therefore, although a worker can perform relatively simple operations such as turning on / off each hardware element using the gesture sensor 16e, it is difficult for the worker to perform complex operations such as operating an application program. In this system, the terminal device 10 without such a display device can be used specifically for collecting images of barcodes attached to products, etc.
[0040] In this embodiment, the server device 20 may utilize a known cloud system configured with one or more web servers. The server device 20 is basically configured to include a central processing unit 21, a storage unit 22, and a communication unit 23. The central processing unit 21 may be a processor such as a central processing unit (CPU) or a graphics processing unit (GPU). The central processing unit 21 reads a predetermined program, loads it into main memory, and executes predetermined arithmetic processing in accordance with the program. The storage function of the storage unit 22 may be realized by a non-volatile memory such as an HDD or SSD. A known communication module may be employed for the communication unit 23. The central processing unit 21 may appropriately write and read arithmetic results in accordance with the program to and from the storage unit 22. In this embodiment, the central processing unit 21 of the server device 20 functions as a reading unit 21a, a learning unit 21b, and a determining unit 21c. The functions of these elements 21a to 21c will be described in detail below.
[0041] In this embodiment, the product management device 30 is a device having a database 33 for managing information about products and the like corresponding to identification information embedded in barcodes. A product management database used in a general POS system can be used as the database 33. The product management device 30 is connected to the terminal device 10 and the server device 20 via a communication line such as the Internet. When the terminal device 10 and the server device 20 transmit identification information obtained by reading the barcode to the product management device 30, they can obtain the name of the product corresponding to the identification information and other detailed information from the product management device 30. The detailed product information includes, for example, the price, nutritional information, the name of the manufacturer, allergy information, usage precautions, and other information that should be communicated to the user.
[0042] Specifically, the product management device 30 includes a processing unit 31, a communication unit 32, and a database 33. The processing unit 31 can be a processor such as a CPU (Central Processing Unit) or a GPU (Graphics Processing Unit). The processing unit 31 reads a predetermined program, loads it into main memory, and executes predetermined arithmetic processing in accordance with the program. The storage function of the database 33 can be realized by a non-volatile memory such as an HDD or SSD. As described above, the database 33 stores barcode identification information and information about the product to which the barcode is attached, in association with each other. Furthermore, a known communication module may be used for the communication unit 32.
[0043] Next, the flow of the learning process by the system according to the present invention will be described with reference to Fig. 4 to Fig. 6. Fig. 4 shows the flow of information processing mainly by the terminal device 10 and the server device 20.
[0044] As shown in FIG. 4, first, the terminal device 10 activates the imaging unit 15 to acquire an image of a barcode C attached to an object T such as a commodity (step S1-1). For example, when the gesture sensor 16d detects that a user has performed a predetermined gesture, the control unit 11 activates the imaging unit 15. After the imaging unit 15 is activated, the control unit 11 of the terminal device 10 may analyze consecutive images captured by the imaging unit 15 and control the imaging unit 15 to release the shutter when an object presumed to be a barcode appears in the image. Furthermore, the terminal device 10 may also control the imaging unit 15 to focus on the barcode when the barcode is captured by the imaging unit 15.
[0045] FIG. 5 schematically illustrates how a barcode C affixed to an object T, such as a product, is captured by a terminal device 10, which is a neck-worn wearable device. FIG. 5 also illustrates an example of a captured image of the barcode C acquired by the imaging unit 15 of the terminal device 10. As illustrated in FIG. 5, this embodiment assumes that a user wears the terminal device 10 around their neck, holds the object T with both hands or one hand, and brings the barcode C into the capturing range of the imaging unit 15 of the terminal device 10. At this time, an object O, such as the user's fingers holding the object T, may be captured in the captured image of the barcode C acquired by the imaging unit 15. Information such as the approximate size of the product and how the product is being held can be obtained from the user's fingers captured in the captured image. Furthermore, the captured image may also include not only the barcode C but also the product packaging to which the barcode C is affixed, the product's contents, and the background at the time of capture. Because the captured image of the barcode C contains a variety of information, it can be said to be useful as training data for machine learning. Furthermore, not only are barcodes C captured completely in the captured image, but depending on how the barcode C is captured, the barcode C may appear incompletely in the captured image due to factors such as the user's fingers overlapping part of the barcode C, the shape of the barcode C being distorted, or a shadow or strong light overlapping the barcode C. Collecting such incomplete captured images as training data can improve the accuracy of machine learning.
[0046] Next, the control unit 11 of the terminal device 10 activates the object sensor 16, particularly the distance measurement sensor 16a, to measure the distance to the object T, such as a commodity with a barcode C, or to another object O, such as the fingers holding the object T (step S1-2). The control unit 11 may keep the distance measurement sensor 16a activated at all times, or may activate the distance measurement sensor 16a temporarily when a barcode is included in the image captured by the imaging unit 15 in order to reduce power consumption. Specifically, the distance measurement sensor 16a may measure the object T or object O within a distance range of 10 to 100 cm. The range of 10 to 100 cm is within the reach of the user's hand, and therefore is sufficient for measuring a commodity being held by the user. The distance (d) from the sensor to the object T is calculated from the detection information of the distance measurement sensor 16a. T ) and the distance from the sensor to the object O (d O ) can be obtained. On the other hand, because objects present in the background of a captured image, such as the object T, are generally outside the measurement range of 10 to 100 cm, the distance from the sensor to these background objects is not measured. In this way, by using the distance measuring sensor 16a to measure the distance to the object T or object O within a predetermined range, it is possible to distinguish between a product and the user's fingers holding the product and other background objects. In addition, information about the distance measured by the distance measuring sensor 16a is stored in the storage unit 12 in approximate correspondence with coordinate values in the captured image. As a result, distance information is added to areas of the captured image where the object T or object O is captured, and distance information is not added to areas of the captured image where the background is captured, making it possible to distinguish between the object T or object O and the rest of the background in the captured image. In addition, the distance information to the object T or object O obtained by the distance measuring sensor 16 can also be used as training data for machine learning as additional information of the captured image.
[0047] Next, the control unit 11 of the terminal device 10 activates the object sensor 16, particularly one or more of the temperature sensor 16b, brightness sensor 16c, and color sensor 16d, to acquire detection information of an object O, such as a user's fingers, that is included in the imaging range of the imaging unit 15 (step S1-3). For example, the temperature sensor 16b, such as an infrared thermistor, can detect the body temperature (20 to 40°C) of the user's fingers. Except in exceptional cases, such as when a product is warmed in a heater, the surface temperature of a product held by the user's fingers is usually below 20°C. Therefore, the temperature sensor 16b can detect an object O in a predetermined temperature range near the user's body temperature and not detect objects in other temperature ranges, thereby distinguishing the user's fingers from the product being held by the user. Furthermore, for example, the brightness sensor 16c and / or the color sensor 16d can detect the brightness or color of the user's fingers. Detection of the user's fingers is even easier if the user is wearing gloves with a distinctive brightness or color. In this way, by using the brightness sensor 16c and / or the color sensor 16d, it is possible to distinguish between the user's fingers and the product being held by the user.
[0048] Next, the control unit 11 of the terminal device 10 identifies the area in the captured image where the object O is detected, mainly based on the detection information acquired in step S1-3, and performs mask processing on the identified area. Note that such mask processing is executed by the function of the image processing unit 11a of the control unit 11. FIG. 6 shows an example of the mask processing. In this embodiment, examples of detection information obtained using the object sensor 16 include the distance (d T ), distance to object O (d O ), the temperature of object O, the brightness of object O, and the color of the object. T ) and the distance to object O (d O) information is primarily used to distinguish between the target object T and object O and the rest of the background in the captured image. Furthermore, information on the temperature, brightness, and color of object O is used to identify the area in the captured image where object O, such as a user's fingers, appears. Here, the image processing unit 11a of the control unit 11 simply performs masking on the area of object O identified in the captured image. Masking is a process of clipping or hiding a specific area from the captured image to remove information contained in the specific area itself. When generating a trained model for barcode inference, the image area of the barcode, the image area of the package or contents of the product to which the barcode is attached, and even the image area of the background provide useful information for inferring the barcode and the product to which the barcode is attached. On the other hand, the image area of the user's fingers holding the product does not contain important information for inferring the barcode or the product. Therefore, by performing masking on the image area of the user's fingers in advance in the captured image, the accuracy of the trained model obtained by machine learning can be improved.
[0049] 6, the image processing unit 11a of the control unit 11 may directly perform mask processing on the original image acquired by the imaging unit 15, and convert this original image into a processed image. In this case, the original image does not remain in the terminal device 10, and therefore the processed image after this mask processing is transmitted from the terminal device 10 to the server device 20. Alternatively, the image processing unit 11a may separately create a mask image indicating a specific area to be masked, based on the original image acquired by the imaging unit 15. In this case, the original image remains in the storage unit 12 of the terminal device 10, and therefore the mask image may be transmitted from the terminal device 10 to the server device 20 together with the original image.
[0050] Furthermore, the image processing unit 11a of the control unit 11 may perform image correction on the captured image in addition to the above-mentioned masking process. The image correction is not particularly limited, but the image processing unit 11a can perform known correction processes on the captured image, such as keystone correction, tilt correction, resizing, brightness correction, noise removal, distortion correction, and sharpness correction.
[0051] Next, the control unit 11 of the terminal device 10 transmits the captured image of the barcode together with the accompanying information to the server device 20 via the communication unit 13 (step S1-5). As described above, if the image processing unit 11a has directly performed mask processing on the original image acquired by the imaging unit 15, the processed image is transmitted to the server device 20. Also, if the image processing unit 11a has generated a mask image separately from the original image acquired by the imaging unit 15, the original image and the mask image are transmitted to the server device 20. Also, the accompanying information transmitted to the server device 20 is mainly detection information acquired by the object sensor 16 when capturing an image of the barcode. The accompanying information transmitted to the server device 20 may include, for example, the distance (d T ) and the distance to object O (d O ) The additional information transmitted to the server device 20 may also include information about the temperature, brightness, color, etc. of the object O.
[0052] Although not shown, when capturing an image of a barcode using the terminal device 10, the user may read out loud the type of product to which the barcode is attached, for example. For example, if the product the user is holding and about to capture the image of the barcode is a sandwich, the user may say "sandwich." In this case, the sound collection unit 14 of the terminal device 10 acquires the user's voice, converts it into an electrical signal, and transmits it to the control unit 11. The control unit 11 performs known voice recognition processing based on this voice signal and converts the voice signal into text information. In the above example, the text data "sandwich" is obtained. The text data obtained in this manner may be included in the supplementary information transmitted to the server device 20 together with the captured image in step S1-5. This text data can be used as training data for machine learning.
[0053] Next, the central processing unit 21 of the server device 20 receives the captured image of the barcode together with the accompanying information from the terminal device 10 via the communication unit 23 (step S1-6). The central processing unit 21 stores the information received from the terminal device 10 in the storage unit 22 as teacher data 22a (step S1-9). In the system of the present invention, it is assumed that a variety of captured images and accompanying information will be transmitted from a plurality of terminal devices 10 to the server device 20, and the central processing unit 21 will sequentially accumulate the information obtained from the plurality of terminal devices 10 as teacher data 22a.
[0054] Next, the central processing unit 21 of the server device 20 reads the barcode included in the captured image received from the terminal device 10 and acquires the identification information attached to the barcode (step S1-7). This reading process is executed by the function of the reading unit 21a of the central processing unit 21. Barcode rules are defined by international standards such as JAN, EAN, and UPC. The reading unit 21a acquires the identification information from the barcode included in the captured image in accordance with known rules. The acquired identification information is associated with the captured image and stored in the storage unit 22 as training data 22a (step S1-9). This identification information is used as a correct answer label in machine learning. In this embodiment, the server device 20 performs the barcode reading process. However, the terminal device 10 may instead perform the barcode reading process. In this case, the terminal device 10 transmits the identification information read from the barcode to the server device 20 as additional information of the captured image.
[0055] Next, the reading unit 21a of the server device 20 transmits the identification information read from the barcode to the product management device 30 and acquires information about the product corresponding to the identification information from the product management device 30 (step S1-8). As described above, the product management device 30 has a database 33 that stores product names and detailed information associated with the identification information, and therefore provides the product names and detailed information in response to an inquiry from the server device 20. This allows the reading unit 21a of the server device 20 to acquire information about the product corresponding to the identification information. The acquired product information is associated with the captured image and stored in the storage unit 22 as training data 22a (step S1-9). This identification information is used as a correct answer label in machine learning. Note that in this embodiment, the server device 20 performs the product information acquisition process. However, the terminal device 10 may instead perform the product information acquisition process. In this case, the terminal device 10 may transmit the product information acquired from the product management device 30 to the server device 20 as supplementary information for the captured image.
[0056] Next, the central processing unit 21 of the server device 20 performs machine learning for image recognition using the training data stored in the storage unit 22 (step S1-10). This machine learning is performed by the function of the learning unit 21b of the central processing unit 21. As a machine learning method, a well-known supervised learning method used for image recognition, such as deep learning, can be used. Examples of machine learning methods include convolutional neural networks (CNNs), recurrent neural networks (RNNs), autoencoders, transfer learning, and generative adversarial networks (GANs). These methods may also be used in combination.
[0057] By performing a learning process by the learning unit 21b of the server device 20, a trained model 22b with adjusted parameters is obtained (step S1-11). Note that if a trained model has already been generated, and a predetermined amount of new training data is accumulated, the trained model 22b may be subjected to the learning process again to update the trained model. The trained model thus obtained basically outputs, as an inference result, identification information of the barcode and / or information on the product to which the barcode is attached, when a captured image containing a barcode is input. In particular, if the detection information acquired by the objective sensor 16 when photographing the barcode is input to this trained model together with the captured image, the accuracy of the inference result obtained from the trained model is improved.
[0058] Next, the flow of inference processing by the system according to the present invention will be described with reference to Fig. 7. Fig. 7 shows the flow of information processing mainly by the terminal device 10 and the server device 20.
[0059] 5, in this inference process, as in the learning process described above, the terminal device 10 first acquires a captured image of a barcode (step S2-1), measures the distance to the target object T or object O (step S2-2), acquires detected information such as the temperature of the object O (step S2-3), masks the captured image (step S2-4), and transmits the captured image and accompanying information (step S2-5). These steps S2-1 to S2-5 are basically the same as the above-described steps S1-1 to S1-5, and therefore detailed description thereof will be omitted.
[0060] Next, the central processing unit 21 of the server device 20 receives the captured image of the barcode together with the accompanying information from the terminal device 10 via the communication unit 23 (step S2-6). The central processing unit 21 temporarily stores the information received from the terminal device 10 in the storage unit 22.
[0061] Next, the central processing unit 21 of the server device 20 inputs the captured image and its associated information received from the terminal device 10 to the trained model 22b (step S2-7). As a result, a determination result is obtained from the trained model 22b (step S2-8). The input process for the trained model 22b and the process for obtaining the determination result are performed by the function of the determination unit 21c of the central processing unit 21. As described above, the trained model 22b has been pre-trained with a correct label using a large number of captured barcode images (including those after masking) and their associated information (such as the distance to the target object or object, the temperature, brightness, and color of the object) as training data. Therefore, by similarly inputting a captured barcode image and its associated information to this trained model 22b, the identification information and / or product information of the barcode corresponding to the correct label is output as a determination result. For example, even if the barcode in the captured image is distorted or overlapped by a user's fingers or a shadow, making it impossible to accurately read the barcode's identification information using a normal method, performing inference using this trained model 22b makes it possible to quickly and accurately read the barcode's identification information. In particular, if the user is visually impaired, it may be difficult to have the terminal device 10 correctly read the barcode attached to the product. In this case, since the barcode in the captured image is significantly distorted or missing as described above, performing inference using this trained model 22b is particularly effective.
[0062] Next, in an embodiment in which the information output from trained model 22b is only the barcode identification information, determination unit 21c of server device 20 transmits the identification information obtained from trained model 22b to product management device 30, and acquires information about the product corresponding to the identification information from this product management device 30 (step S2-9). Since product management device 30 has database 33 that stores product names and detailed information in association with the identification information, it provides the product names and detailed information in response to an inquiry from server device 20. This allows determination unit 21c of server device 20 to acquire information about the product corresponding to the identification information.
[0063] Next, the central processing unit 21 of the server device 20 transmits product information corresponding to the barcode to the terminal device 10 that has received the captured image of the barcode via the communication unit 23 (step S2-10). In this embodiment, since the terminal device 10 is a neck-hanging wearable device that does not have a display, it is preferable that the product information transmitted from the server device 20 to the terminal device 10 be text information so that the product information can be output by voice.
[0064] Next, the control unit 11 of the terminal device 10 receives the product information from the server device 20 via the communication unit 13 (step S2-11). Then, the control unit 11 of the terminal device 10 generates a voice signal from the product information, which is text information, and outputs the voice signal from the sound output unit 18 (speaker) (step S2-12). Methods for reading out text information aloud are well known. For example, the text information may be input to a text analysis engine within the terminal device 10 to analyze the pronunciation of characters, accents, punctuation, etc., and the analyzed text information may be input to a voice generation engine within the terminal device 10 to generate natural voice signals corresponding to symbols and words using various algorithms that mimic human speech, and the generated voice signals may be output from the sound output unit 18. As a result, the user's experience becomes such that by capturing an image of a barcode attached to an object such as a product using the terminal device 10, the name and detailed information of the product can be heard by voice.
[0065] In the above embodiment, the terminal device 10 basically communicates with the server device 20 to acquire product information corresponding to an imaged barcode. However, as indicated by the dotted lines in FIG. 3 , the terminal device 10 itself may store a trained model 24′ corresponding to the trained model 24 generated by the server device 20 and a database 33′ corresponding to the database 33 held by the product management device 30, thereby enabling the terminal device 10 to acquire product information corresponding to the barcode without accessing the server device 20. In this case, the control unit 11 of the terminal device 10 may be provided with a function equivalent to the determination unit 21c of the central processing unit 21 of the server device 20 described above. That is, the control unit 11 of the terminal device 10 may also be provided with the function of the determination unit 11b, and the terminal device 10 may execute processes corresponding to steps S2-7 to S2-9 shown in FIG. 7 within the terminal device 10. This enables the terminal device 10 to read out product information in a standalone manner.
[0066] However, since the storage capacity of the terminal device 10 is generally limited and smaller than that of the server device 20 and the product management device 30, the trained model 22b' and database 33' provided in the terminal device 10 may have a data capacity and functions that are partially reduced compared to those provided in the server device 20 and the product management device 30. For example, if the terminal device 10 is used only in a specific store, it is sufficient that the trained model 22b' and database 33' cover all the products sold in that store. In this way, the server device 20 and the product management device 30 may have trained models 22b' and databases 33' with large data capacities to handle a wide variety of products, while the trained model 22b' and database 33' implemented in the terminal device 10 may have data capacities optimized to suit the intended use of the terminal device 10.
[0067] In addition, if the terminal device 10 is unable to recognize the barcode or acquire product information by referring to the trained model 22b' or database 33' provided in the terminal device 10, the terminal device 10 may make an inquiry to the server device 20 according to the flow shown in Figure 7, and have the server device 20 recognize the barcode or acquire product information.
[0068] In the above description of the present invention, the embodiments of the present invention have been described with reference to the drawings in order to express the contents of the present invention. However, the present invention is not limited to the above embodiments, and includes modifications and improvements that are obvious to those skilled in the art based on the matters described in the present specification. [Explanation of symbols]
[0069] 10... Terminal device 11... Control unit 11a...image processing unit 11b...determination unit 12...Memory unit 13...Communication unit 14...sound collection unit 15...imaging unit 16...object sensor 16a...distance measurement sensor 16b...Temperature sensor 16c...Brightness sensor 16d...Color sensor 16e...Gesture sensor 17...Sensors 18...Sound emission unit 19... location information acquisition unit 20... server device 21...Central processing unit 21a...Reading unit 21b...Learning section 21c...Judgment section 22...Memory unit 22a...Teacher data 22b...Trained model 23...Communication section 23...Training data 24...Trained model 30...Product management device 31...Processing section 32...Communication Department 33...Database 100...System
Claims
1. An image acquisition device for acquiring an image of a read code including a one-dimensional code and a two-dimensional code attached to an object, an imaging unit that acquires an image; a non-contact sensor unit that detects an object in a detection range that at least partially overlaps with the imaging range of the imaging unit; an image processing unit that, when the image includes the read code, identifies an area in the image where the object is detected based on detection information from the sensor unit; Image acquisition device.
2. The sensor unit includes a distance measuring sensor that measures the distance to the object. The image acquisition device of claim 1 .
3. The distance sensor detects the object within a distance range of 10 to 100 cm. The image acquisition device of claim 2 .
4. The sensor unit includes a temperature sensor that measures the temperature of the object. The image acquisition device of claim 1 .
5. The temperature sensor detects the object in a temperature range of 20 to 40°C. The image acquisition device of claim 4 .
6. The sensor unit includes a brightness sensor that measures the brightness of the object. The image acquisition device of claim 1 .
7. The sensor unit includes a color sensor that measures the color of the object. The image acquisition device of claim 1 .
8. The image processing device further includes a communication unit that transmits a data set including information about the image and the area in which the object is detected to an external server device via a communication line. The image acquisition device of claim 1 .
9. The image processing unit further performs a mask process on the region in the image where the object is detected. The image acquisition device of claim 1 .
10. The image processing device further includes a communication unit that transmits the masked image to an external server device via a communication line. The image acquisition device of claim 9.
11. The image processing device further includes a communication unit that transmits the image and the mask image obtained by the mask processing to an external server device via a communication line. The image acquisition device of claim 9.
12. A learning system comprising an image capture device for capturing images of read codes including one-dimensional codes and two-dimensional codes attached to an object, and a server device connected to the image capture device via a communication line, The image acquisition device an imaging unit that acquires an image; a non-contact sensor unit that detects an object in a detection range that at least partially overlaps with the imaging range of the imaging unit; an image processing unit that, when the image includes the read code, identifies an area in the image where the object is detected based on detection information from the sensor unit; a communication unit that transmits a data set including information about the image and a region in the image where the object is detected to the server device via a communication line; The server device a learning unit that performs machine learning using the data set received from the image acquisition device and generates or updates a trained model; Learning system.
13. The trained model is used to make inferences about products to which the reading code is assigned. The system of claim 12.
14. A method for producing a trained model, which is executed by an image acquisition device for acquiring an image of a read code including a one-dimensional code and a two-dimensional code attached to an object, and a server device connected to the image acquisition device via a communication line, The image acquisition device acquiring an image; detecting an object with a non-contact sensor in a detection range that at least partially overlaps with an imaging range of the imaging unit; a step of identifying an area in the image where the object is detected based on detection information from the sensor unit when the image includes the read code; transmitting a data set including the image and information about the area in the image where the object is detected to the server device via a communication line; The server device: performing machine learning using the dataset received from the image acquisition device to generate or update a trained model; How to produce a trained model.
15. An inference system comprising an image acquisition device for acquiring an image of a read code, including a one-dimensional code and a two-dimensional code, attached to an object, and a server device connected to the image acquisition device via a communication line, The image acquisition device an imaging unit that acquires an image; a non-contact sensor unit that detects an object in a detection range that at least partially overlaps with the imaging range of the imaging unit; an image processing unit that, when the image includes the read code, identifies an area in the image where the object is detected based on detection information from the sensor unit; a communication unit that transmits a data set including information about the image and a region in the image where the object is detected to the server device via a communication line; The server device A trained model generated or updated by performing machine learning using the dataset received in advance from the image acquisition device; a determination unit that, when the data set is newly received from the image acquisition device, determines information contained in the read code or a product to which the read code is assigned using the trained model; and a communication unit that transmits the determination result by the determination unit to the image acquisition device; Inference system.
16. An image acquisition device capable of acquiring an image of a read code, including a one-dimensional code and a two-dimensional code, attached to an object and determining a product to which the read code is assigned, an imaging unit that acquires an image; a non-contact sensor unit that detects an object in a detection range that at least partially overlaps with the imaging range of the imaging unit; an image processing unit that, when the image includes the read code, identifies an area in the image where the object is detected based on detection information from the sensor unit; a trained model that has been generated or updated by performing machine learning using a dataset that includes information on the image and a region in the image where the object is detected; and a determination unit that, when a new data set is acquired, determines a product to which the read code is assigned using the trained model. Image acquisition device.
Citation Information
Patent Citations
Commodity package provided with identification code for visually handicapped person
JP2020135794A