Object recognition device, object recognition method, program, and recording medium
The object identification device and method address the challenge of identifying drug types in mixed image scenarios by using a detector to estimate object types and groups, and applying group-specific processing to achieve accurate drug type identification.
Patent Information
- Application Number
- JP2024511824
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2022-03-28
- Filing Date
- 2023-03-17
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2043-03-17
AI Technical Summary
Existing image recognition technologies struggle to identify the type of drugs, such as capsule drugs and half tablets, from images due to incomplete character symbol information and diverse patterns, making it difficult to specify the drug type through template matching or machine learning-based methods.
An object identification device and method that detect objects from images containing a mix of type-identifiable and type-difficult-to-identify objects. The device uses a detector to estimate the type of identifiable objects and the group of difficult-to-identify objects, followed by group-specific processing to identify the type of objects.
Enables the identification of drug types in images where both type-identifiable and type-difficult-to-identify objects are present, improving the efficiency and accuracy of drug type specification in mixed image scenarios.
Smart Images

Figure 0007690684000001 
Figure 0007690684000002 
Figure 0007690684000003
Abstract
Description
Technical Field
[0001] The present disclosure relates to an object recognition device, an object recognition method, and a program, and particularly relates to an image recognition technology for recognizing an object from an image.
Background Art
[0002] Patent Document 1 describes drug identification software that sets a drug to be identified in a drug imaging device, performs imaging of the drug, and searches for the drug with reference to a database based on the data of the captured drug image.
[0003] Patent Document 2 describes a tablet detection method including an imaging step of imaging a reflection light image and a transmission light image of a wrapping paper in which one or more tablets for one dose are wrapped, a cutting step of cutting out a tablet region which is a region corresponding to the tablet in the reflection light image based on the reflection light image and the transmission light image, and a first identification step of identifying the type of each tablet by collating the dimensions and colors of each of the tablet regions cut out in the cutting step with model information regarding the shape and color of the tablet, and a second identification step of identifying the type of at least similar tablets having different types but similar feature amounts based on a learning model generated by executing machine learning using learning data including the image of the tablet.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Patent Document 2
Summary of the Invention
Problems to be Solved by the Invention
[0005] When identifying the drug type of tablets, in many cases, the drug type can be specified by using the information of the engraved or printed characters and / or symbols on the tablets as clues. On the other hand, for drugs such as capsule drugs and half tablets, due to reasons such as only a part of the character symbol information for specifying the drug type being shown in the photographed image, or there being countless patterns of capsule closures and tablet divisions, it is difficult or impossible to specify the drug type by template matching with a master image or a machine learning-based method for the photographed image. That is, drugs can be classified into those for which the drug type can be specified up to the drug type by image recognition technology from the photographed image, such as tablets with engraving or printing, and those for which it is difficult to specify the drug type up to the drug type from the photographed image, such as capsule drugs and half tablets. In this specification, the former is referred to as "drug type specifiable drug", and the latter is referred to as "drug type difficult to specify drug". Also, the characters and / or symbols attached to the drug by engraving or printing are denoted as "character symbols".
[0006] For drugs with specifiable drug types, it is expected that the drug type can be specified up to the drug type by image recognition based on the photographed image. On the other hand, to specify the drug type of drugs with difficult-to-specify drug types, for example, by enlarging the photographed image, presenting the information of the character symbols attached to the drug to the user, and having the user visually identify the partially visible character symbols and perform text input or voice input, the drug type can be specified by a method such as performing a search on a database such as an engraved text master. Thus, for drugs with difficult-to-specify drug types, a method of presenting information useful for specifying the drug type to the user to assist in the drug type specification work is considered to be practically effective.
[0007] It is also conceivable to construct a system that separates drugs with specifiable drug types from drugs with difficult-to-specify drug types at the stage of photographing the drug to be identified, and performs photographing only for drugs with specifiable drug types to identify the types of drugs with specifiable drug types.
[0008] However, from the pharmaceutical perspective and the usability perspective, it is preferable to be able to photograph a plurality of drugs at once in units such as the same dosing time and the same dispensing, and perform the drug type specification work collectively.
[0009] However, when taking pictures in units such as the same dosing time in this way, inevitably, a single captured image will contain drugs with difficult drug type identification and drugs with identifiable drug types, and it is necessary to select whether each drug is a drug with difficult drug type identification or an identifiable drug type. From the perspective of usability, it is preferable that this selection be automatically executed on the drug type identification device as much as possible. Also, this selection is preferably made during a series of drug type identification work flows on the drug type identification device.
[0010] Such a technical problem is grasped as a common problem not only in the application of identifying drugs, but also when identifying the type of an object from an image for various objects. In this specification, an object that can be identified from an image down to the type of the object is called a "type-identifiable object", and an object that is difficult to identify from an image down to the type of the object but can identify the group to which the object belongs is called a "type-difficult-to-identify object".
[0011] In view of such circumstances, the present disclosure has been made, and an object of the present disclosure is to provide an object identification device, an object identification method, and a program that enable identification of the type of each object using an image in which a plurality of objects including type-identifiable objects and type-difficult-to-identify objects may be mixed.
Means for Solving the Problem
[0012] An object identification device according to an aspect of the present disclosure includes a detector that detects objects from an image in which a plurality of objects are captured in object units, and among the objects detected by the detector, for type-identifiable objects that can be identified from the image down to the type of the object, the type of the object is estimated from the image, and for type-difficult-to-identify objects that are difficult to identify from the image down to the type of the object but can identify the group to which the object belongs, a group is estimated from the image, and a processing unit that performs group-specific processing leading to identification of the type for the objects estimated as a group by the identifier.
[0013] According to this aspect, objects are detected from an image in which a plurality of objects are photographed, and an estimator estimates for each individual object. When the object detected from the image is an object whose type can be specified, the type of the object is estimated by the estimator. On the other hand, when the object detected from the image is an object whose type is difficult to specify, the group to which the object belongs is estimated by the estimator, and for the objects for which the group has been estimated up to, the process proceeds to the group-specific process that contributes to the specification of the object type, leading to the specification of the type.
[0014] The image may be an image photographed in a state where objects whose type can be specified and objects whose type is difficult to specify are mixed.
[0015] The objects whose type is difficult to specify may be classified into a plurality of groups, and a group-specific process may be defined for each group.
[0016] The detector may include a first trained model trained by machine learning using first training data labeled in units of objects without distinguishing between objects whose type can be specified and objects whose type is difficult to specify.
[0017] The identifier may include a second trained model trained by machine learning using second training data labeled in units of the types of objects for objects whose type can be specified and labeled in units of the groups to which the objects belong for objects whose type is difficult to specify.
[0018] The label for identifying the group may be defined in a hierarchical structure.
[0019] As an input to the identifier, a configuration may be used that includes at least one of an object image obtained by cutting out the region of the object detected by the detector from the image in units of objects, a character and symbol extraction image including at least one of characters and symbols extracted from the object image, an outer shape image of the object, and size information of the object.
[0020] As an input to the identifier, a configuration may be used in which magnification information indicating the magnification or reduction rate of the object image is further used.
[0021] The group-specific processing may be configured to include a process of displaying a screen for receiving an input of search conditions for searching for the type of object within the estimated group.
[0022] In another aspect of the present disclosure, the object is a drug, the object with difficult type identification includes at least one of capsule drugs, plain drugs, and divided tablets, and the object with identifiable type may include tablets with markings or printing.
[0023] As an input to the identifier, a configuration may be used in which at least one of a drug image obtained by cutting out the area of the drug detected by the detector from the image in drug units, a character and symbol extraction image including at least one of the characters and symbols extracted from the drug image, an external shape image of the drug, and size information of the drug is used.
[0024] An object identification device according to another aspect of the present disclosure is an object identification device including one or more processors and one or more memories in which programs executed by the one or more processors are stored. The one or more processors perform a detection process of detecting an object in units of objects from an image in which a plurality of objects are photographed, and for the identifiable objects among the objects detected by the detection process, the type of the object is estimated from the image, and for the difficult-to-identify objects for which the group to which the object belongs can be estimated although the type of the object cannot be identified from the image, the group is estimated from the image. And a process of transitioning to group-specific processing that leads to the identification of the type of the object estimated as a group by the identification process.
[0025] The one or more processors may be configured to execute the detection process using a detector including a first trained model trained by machine learning using first training data in which objects are labeled in units of objects without distinguishing between identifiable objects and difficult-to-identify objects.
[0026] One or more processors may be configured to perform identification processing using an identifier that includes a second trained model trained by machine learning using second training data in which identifiable objects are labeled in units of object types and difficult-to-identify objects are labeled in units of groups to which the objects belong.
[0027] One or more processors may be configured to cut out a region of an object detected by detection processing from an image, perform processing to generate an object image in units of objects, and perform identification processing based on the object image.
[0028] In another aspect of the present disclosure, the object is a drug, the difficult-to-identify objects include at least one of capsule drugs, plain drugs, and divided tablets, the identifiable objects include tablets having markings or printing, and the group-specific processing may include processing to display a screen for receiving input of search conditions for searching for the type of drug within the estimated group.
[0029] In another aspect of the present disclosure, there is provided a first database in which character symbol information including at least one of characters and symbols indicated by markings or printing attached to a drug is associated with the type of the drug, and a second database in which master images of drugs are stored. One or more processors may be configured to search at least one of the first database and the second database based on the received search conditions and output candidates for drugs that meet the search conditions.
[0030] One or more processors may be configured to perform processing to display a screen including a captured image display unit that displays an image and a candidate display unit that displays information on candidate objects based on an estimation result of identification processing.
[0031] One or more processors may be configured to perform processing to display information on a group to which candidate objects to be displayed on the candidate display unit belong.
[0032] One or more processors may be configured to receive an instruction specifying a group to which a candidate object to be displayed on the candidate display unit belongs, and to control the display on the candidate display unit in accordance with the received instruction.
[0033] In another aspect of the present disclosure, a configuration including a camera and a display that displays information regarding an image captured by the camera and an object estimated from the image may be provided.
[0034] An object identification method according to another aspect of the present disclosure is an object identification method executed by one or more processors, the method including: detecting objects from an image in which a plurality of objects are captured, on a per-object basis; for a type-identifiable object among the detected objects, estimating the type of the object from the image, and for a type-difficult object for which it is difficult to identify the type of the object from the image but for which the group to which the object belongs can be identified, estimating the group of the object from the image; and performing group-specific processing that leads to the identification of the type of the object for the object estimated as a group.
[0035] A program according to another aspect of the present disclosure causes a computer to realize functions of: detecting objects from an image in which a plurality of objects are captured, on a per-object basis; for a type-identifiable object among the detected objects, estimating the type of the object from the image, and for a type-difficult object for which it is difficult to identify the type of the object from the image but for which the group to which the object belongs can be identified, estimating the group of the object from the image; and performing group-specific processing that leads to the identification of the type of the object for the object estimated as a group.
Advantages of the Invention
[0036] According to the present disclosure, it is possible to identify the type of each object in an image even when the image is captured in a state where type-identifiable objects and type-difficult objects are mixed.
Brief Description of the Drawings
[0037]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
Figure 22
Figure 23
Figure 24
Figure 25
Figure 26
Figure 27
Figure 28
Figure 29
Figure 30
Figure 31
Figure 32
Figure 33
Figure 34
Figure 35
Figure 36
Figure 37
Figure 38
Figure 39
Figure 40
Figure 41
Figure 42
Figure 43
Figure 44
Figure 45
Figure 46
Figure 47
BEST MODE FOR CARRYING OUT THE INVENTION
[0038] Hereinafter, preferred embodiments of the present invention will be described in detail with reference to the accompanying drawings.
[0039] <Overview of the Drug Type Identification Device According to the Embodiment> The drug type identification device according to the embodiment is a device that identifies the type (drug type) of a drug from an image of the drug. In this embodiment, it is assumed that a plurality of drugs are photographed in a state where drugs with identifiable drug types and drugs with difficult-to-identify drug types are mixed. However, the drug type identification device can function effectively even when both are not mixed or only one drug is photographed.
[0040] The drug type identification device according to the embodiment detects individual drugs from an image of a plurality of drugs, automatically discriminates whether each detected individual drug is a drug with difficult-to-identify drug type or a drug with identifiable drug type, and automatically branches to a drug type identification flow suitable for each, enabling efficient identification of the drug type. As an example, the drug type identification device is mounted on a portable terminal device. The portable terminal device includes at least one of a smartphone, a mobile phone, a PHS (Personal Handy-phone System), a PDA (Personal Digital Assistant), a tablet computer terminal, a notebook personal computer terminal, and a portable game machine. Hereinafter, taking the drug type identification device realized by the hardware and software of a smartphone as an example, it will be described in detail with reference to the drawings.
[0041] 〔Appearance of Smartphone〕 FIG. 1 is a front perspective view of a smartphone 10, which is a portable terminal device with a camera that functions as a drug type identification device according to the embodiment. As shown in FIG. 1, the smartphone 10 has a flat housing 12. The smartphone 10 includes a touch panel display 14, a speaker 16, a microphone 18, and a front camera 20 on the front surface of the housing 12.
[0042] The touch panel display 14 includes a display unit that displays images and the like, and a touch panel unit that is disposed on the front surface of the display unit and receives touch inputs. The display unit is, for example, a color LCD (Liquid Crystal Display) panel.
[0043] The touch panel unit is, for example, a capacitive touch panel that is provided planar on a substrate body having light transmissivity and has a position detection electrode having light transmissivity and an insulating layer provided on the position detection electrode. The touch panel unit generates and outputs two-dimensional position coordinate information corresponding to a touch operation by a user. The touch operation includes a tap operation, a double-tap operation, a flick operation, a swipe operation, a drag operation, a pinch-in operation, and a pinch-out operation.
[0044] The speaker 16 is an audio output unit that outputs sound during a call and during video playback. The microphone 18 is an audio input unit into which sound is input during a call and during video shooting. The in-camera 20 is an imaging device that shoots videos and still images.
[0045] FIG. 2 is a rear perspective view of the smartphone 10. As shown in FIG. 2, the smartphone 10 includes an out-camera 22 and a light 24 on the back surface of the housing 12. The out-camera 22 is an imaging device that shoots videos and still images. The light 24 is a light source that irradiates illumination light when shooting with the out-camera 22, and is configured by, for example, an LED (Light Emitting Diode).
[0046] Furthermore, as shown in FIGS. 1 and 2, the smartphone 10 includes switches 26 on the front and side surfaces of the housing 12, respectively. The switch 26 is an input member that receives an instruction from the user. The switch 26 is a push-button type switch that turns on when pressed with a finger or the like and turns off by a restoring force of a spring or the like when the finger is released.
[0047] Note that the configuration of the housing 12 is not limited to this, and a configuration having a foldable structure or a slide mechanism may be adopted.
[0048] 〔Electrical Configuration of Smartphone〕 As a main function of the smartphone 10, it has a wireless communication function for performing mobile wireless communication via a base station device and a mobile communication network.
[0049] Figure 3 is a block diagram showing the electrical configuration of the smartphone 10. As shown in Figure 3, in addition to the aforementioned touch panel display 14, speaker 16, microphone 18, in-camera 20, out-camera 22, light 24, and switch 26, the smartphone 10 includes a CPU (Central Processing Unit) 28, a wireless communication unit 30, a call unit 32, a memory 34, an external input / output unit 40, a GPS receiver unit 42, and a power supply unit 44.
[0050] The CPU 28 is an example of a processor that executes instructions stored in the memory 34. The CPU 28 operates according to the control program and control data stored in the memory 34, and comprehensively controls each part of the smartphone 10. The CPU 28 has a mobile communication control function for controlling each part of the communication system and an application processing function in order to perform voice communication and data communication through the wireless communication unit 30.
[0051] In addition, the CPU 28 has an image processing function for displaying videos, still images, characters, etc. on the touch panel display 14. Through this image processing function, information such as still images, videos, and characters is visually transmitted to the user. Also, the CPU 28 acquires two-dimensional position coordinate information corresponding to the user's touch operation from the touch panel part of the touch panel display 14. Furthermore, the CPU 28 acquires an input signal from the switch 26.
[0052] The in-camera 20 and the out-camera 22 capture videos and still images according to the instructions of the CPU 28. Figure 4 is a block diagram showing the internal configuration of the in-camera 20. Note that the internal configuration of the out-camera 22 is common to that of the in-camera 20. As shown in Figure 4, the in-camera 20 includes a photographing lens 50, a diaphragm 52, an imaging device 54, an AFE (Analog Front End) 56, an A / D (Analog to Digital) converter 58, and a lens drive unit 60.
[0053] The photographing lens 50 is composed of a zoom lens 50Z and a focus lens 50F. The lens driving unit 60 drives the zoom lens 50Z and the focus lens 50F forward and backward in response to commands from the CPU 28 to perform optical zoom adjustment and focus adjustment. Further, the lens driving unit 60 controls the aperture 52 in response to commands from the CPU 28 to adjust the exposure. The lens driving unit 60 corresponds to an exposure correction unit that performs camera exposure correction based on the gray color described later. Information such as the positions of the zoom lens 50Z and the focus lens 50F and the aperture degree of the aperture 52 is input to the CPU 28.
[0054] The imaging device 54 has a light-receiving surface on which a large number of light-receiving elements are arranged in a matrix. The subject light that has passed through the zoom lens 50Z, the focus lens 50F, and the aperture 52 is imaged on the light-receiving surface of the imaging device 54. Color filters for R (red), G (green), and B (blue) are provided on the light-receiving surface of the imaging device 54. Each light-receiving element of the imaging device 54 converts the subject light imaged on the light-receiving surface into an electrical signal based on signals of the respective colors of R, G, and B. Thereby, the imaging device 54 acquires a color image of the subject. As the imaging device 54, a photoelectric conversion element such as a CMOS (Complementary Metal-Oxide Semiconductor) or a CCD (Charge-Coupled Device) can be used.
[0055] The AFE 56 performs noise removal, amplification, etc. of the analog image signal output from the imaging device 54. The A / D converter 58 converts the analog image signal input from the AFE 56 into a digital image signal with a gradation width. Note that an electronic shutter is used as the shutter that controls the exposure time of the incident light to the imaging device 54. In the case of an electronic shutter, the exposure time (shutter speed) can be adjusted by controlling the charge accumulation period of the imaging device 54 by the CPU 28.
[0056] The in-camera 20 may convert the captured video and still image data into compressed image data such as MPEG (Moving Picture Experts Group) or JPEG (Joint Photographic Experts Group).
[0057] Returning to the description of FIG. 3, the CPU 28 causes the memory 34 to store the video and still images captured by the in-camera 20 and the out-camera 22. Further, the CPU 28 may output the video and still images captured by the in-camera 20 and the out-camera 22 to the outside of the smartphone 10 through the wireless communication unit 30 or the external input / output unit 40.
[0058] Furthermore, the CPU 28 displays the video and still images captured by the in-camera 20 and the out-camera 22 on the touch panel display 14. The CPU 28 may utilize the video and still images captured by the in-camera 20 and the out-camera 22 within the application software.
[0059] Note that when the out-camera 22 performs shooting, the CPU 28 may irradiate the subject with auxiliary shooting light by lighting the light 24. The lighting and extinguishing of the light 24 may be controlled by a touch operation on the touch panel display 14 by the user or an operation of the switch 26.
[0060] The wireless communication unit 30 performs wireless communication with a base station device accommodated in a mobile communication network according to an instruction from the CPU 28. The smartphone 10 uses this wireless communication to transmit and receive various file data such as voice data and image data, e-mail data, and receive Web (abbreviation for World Wide Web) data and streaming data.
[0061] The call unit 32 is connected to the speaker 16 and the microphone 18. The call unit 32 decodes the voice data received by the wireless communication unit 30 and outputs it from the speaker 16. The call unit 32 converts the voice of the user input through the microphone 18 into voice data that can be processed by the CPU 28 and outputs it to the CPU 28.
[0062] The memory 34 stores instructions for the CPU 28 to execute. The memory 34 is composed of an internal memory unit 36 built into the smartphone 10 and an external memory unit 38 that is detachable from the smartphone 10. The internal memory unit 36 and the external memory unit 38 are realized using known storage media.
[0063] The memory 34 stores the control program of the CPU 28, control data, application software, address data associated with the name and phone number of a communication partner, etc., data of sent and received e-mails, Web data downloaded by Web browsing, and downloaded content data, etc. Also, the memory 34 may temporarily store streaming data, etc.
[0064] The external input / output unit 40 serves as an interface with external devices connected to the smartphone 10. The smartphone 10 is connected directly or indirectly to other external devices through communication or the like via the external input / output unit 40. The external input / output unit 40 transmits data received from an external device to each component inside the smartphone 10, and transmits data inside the smartphone 10 to the external device.
[0065] Means for communication or the like are, for example, Universal Serial Bus (USB), IEEE (Institute of Electrical and Electronics Engineers) 1394, the Internet, Wireless LAN (Local Area Network), Bluetooth (registered trademark), RFID (Radio Frequency Identification), and infrared communication. Also, external devices are, for example, headsets, external chargers, data ports, audio devices, video devices, smartphones, PDAs, personal computers, and earphones.
[0066] The GPS receiver unit 42 detects the position of the smartphone 10 based on the positioning information from the GPS satellites ST1, ST2, …, STn.
[0067] The power supply unit 44 is a power supply source that supplies power to each part of the smartphone 10 via a power supply circuit (not shown). The power supply unit 44 includes a lithium-ion secondary battery. The power supply unit 44 may include an A / D conversion unit that generates a DC voltage from an external AC power supply.
[0068] The smartphone 10 configured as described above is set to the shooting mode by an instruction input from the user using the touch panel display 14 or the like, and can shoot moving images and still images with the in-camera 20 and the out-camera 22.
[0069] When the smartphone 10 is set to the shooting mode, it enters the shooting standby state, a moving image is shot by the in-camera 20 or the out-camera 22, and the shot moving image is displayed on the touch panel display 14 as a live view image.
[0070] The user can view the live view image displayed on the touch panel display 14 to determine the composition, confirm the subject to be shot, or set the shooting conditions.
[0071] When shooting is instructed by an instruction input from the user using the touch panel display 14 or the like in the shooting standby state, the smartphone 10 performs AF (Autofocus) and AE (Auto Exposure) control, and shoots and stores moving images and still images.
[0072] 〔Functional Configuration of Medicine Type Identification Device〕 FIG. 5 is a block diagram showing the functional configuration of a drug type identification device 100 realized by a smartphone 10. The drug type identification device 100 includes a processor 102 and a storage device 104. The processor 102 includes a CPU 28. The processor 102 may include a GPU (Graphics Processing Unit). The storage device 104 is a computer-readable medium which is a non-transitory tangible object and includes a memory 34. The processor 102 is connected to a touch panel display 14. The touch panel display 14 includes a display unit 14A that functions as a display device (display) and an input unit 14B that functions as an input device that receives input by a touch operation.
[0073] The drug type identification device 100 includes an image acquisition unit 112, a drug detector 114, a region correction unit 116, a drug region cutting-out unit 118, a drug identifier 120, a text search unit 122, a display control unit 124, and an input processing unit 126.
[0074] The image acquisition unit 112 acquires a captured image of a drug to be identified. The captured image is, for example, an image captured by an in-camera 20 or an out-camera 22. The captured image may be an image acquired from another device via a wireless communication unit 30, an external storage unit 38, or an external input / output unit 40. The captured image may include a plurality of drugs to be identified. The plurality of drugs to be identified is not limited to drugs to be identified of the same drug type, and may be drugs to be identified of different drug types respectively. In the present embodiment, an aspect of processing an image in which a plurality of drugs are collectively captured (as one image) in a state where different types of drugs are mixed will be described as an example.
[0075] The captured image may be an image in which a drug to be identified and a marker are captured. The marker may be, for example, an ArUco marker, a circular marker, or a square marker. It is preferable that a plurality of markers are included in the captured image. The plurality of markers are arranged, for example, at the four corners of a rectangular area of the drug placement range. The captured image may be an image in which a drug to be identified and a reference gray color are captured.
[0076] The captured image may be an image captured at a standard shooting distance and shooting viewpoint. The shooting distance can be represented by the distance between the drug to be identified and the imaging lens 50 and the focal length of the imaging lens 50. Also, the shooting viewpoint can be represented by the angle formed by the marker printing surface and the optical axis of the imaging lens 50.
[0077] The image acquisition unit 112 includes an image correction unit (not shown). When the captured image contains a marker, the image correction unit standardizes the shooting distance and shooting viewpoint of the captured image based on the marker to obtain a standardized image. The standardized image may be an image in which the area inside a rectangle with the markers at the four corners as vertices is cut out after the captured image has been subjected to standardization processing. For example, the image correction unit designates the coordinates to which the four vertices of the quadrilateral whose coordinates are specified by the marker will go after standardization of the shooting distance and shooting viewpoint. The image correction unit obtains a perspective transformation matrix such that these four vertices are transformed to the positions of the designated coordinates. Such a perspective transformation matrix is uniquely determined if there are four points. For example, if there is a correspondence relationship between four points, the transformation matrix can be obtained by using the getPerspectiveTransform function of OpenCV (Open Source Computer Vision Library).
[0078] The image correction unit uses the obtained perspective transformation matrix to perform perspective transformation on the entire original captured image to obtain a transformed image. Such perspective transformation can be executed by using the warpPerspective function of OpenCV. This transformed image may be the standardized image in which the shooting distance and shooting viewpoint are standardized.
[0079] Also, when the captured image contains a reference gray-colored area, the image correction unit may perform color tone correction of the captured image based on the reference gray color.
[0080] The drug detector 114 includes a first trained model TM1 which is a learning model trained by machine learning. The first trained model TM1 is a model trained to perform a so-called object detection task. When given an image (pre-normalization image or normalized image) as input, the first trained model TM1 outputs position information corresponding to the region of the detected object, the class of the object, and a score indicating the probability of the object. The first trained model TM1 is an example of the "first trained model" in the present disclosure.
[0081] The class of the object in drug detection includes at least "drug", and may further include "marker". The drug detector 114 detects drugs from the captured image acquired by the image acquisition unit 112 and outputs information indicating the region of the detected drug. When the normalized image is acquired by the image correction unit, the drug detector 114 detects the region of the drug from the normalized image. When the captured image contains a plurality of drugs, the drug detector 114 detects the regions of each of the plurality of drugs.
[0082] The output from the first trained model TM1 may be the position information of the bounding box indicating the region of each individual drug detected from within the captured image, or may be a segmentation mask image in which the regions of the individual drugs are filled in pixel units. Details of the learning method for creating the first trained model TM1 and the content of the detection process by the drug detector 114 will be described later. The drug detector 114 is an example of the "detector" in the present disclosure.
[0083] The result of the detection process by the drug detector 114 is displayed on the display unit 14A via the display control unit 124. The display control unit 124 generates a display signal for display on the touch panel display 14 and performs display control. The display control unit 124 includes a magnification change unit 125 that changes the display magnification of the image to be displayed on the display unit 14A. When an instruction to change the display magnification is input, such as when a pinch-out operation or a pinch-in operation is performed on the touch panel display 14, the magnification change unit 125 performs an enlargement or reduction process according to the instruction. The input processing unit 126 receives an input from the input unit 14B or the microphone 18 and sends the received input information to the corresponding processing unit.
[0084] The area correction unit 116 receives an instruction for correction from the user regarding the area of the drug detected by the drug detector 114, and performs a process of correcting the drug area according to the received instruction. That is, the area correction unit 116 makes a correction to the detection result of the drug detector 114 according to the area correction instruction received from the input unit 14B. When an incorrect detection or a detection omission occurs by the drug detector 114, the user can input an instruction to correct the detection result from the input unit 14B and specify the correct area of the drug to be identified.
[0085] The drug area cutting-out unit 118 performs a process of cutting out the drug area for each drug from the captured image based on the detection result of the drug detector 114. When the drug area is corrected by the area correction unit 116, the drug area cutting-out unit 118 performs a process of cutting out the corrected drug area.
[0086] The drug identifier 120 includes a second trained model TM2 which is a learning model trained by machine learning. The second trained model TM2 is a model trained to perform a so-called object recognition task. The drug identifier 120 acquires the region image of each drug (hereinafter referred to as drug image) cut out by the drug region cutting unit 118, and estimates the type of the corresponding drug or the group to which the drug belongs for labeling (i.e., multi-class classification). The class classification performed by the drug identifier 120 will have different classification fineness (granularity) depending on whether the drug shown in the input image is a drug with difficult drug type identification or a drug with possible drug type identification. The drug image is an example of the "object image" in the present disclosure. The second trained model TM2 is an example of the "second trained model" in the present disclosure.
[0087] When the input drug image is an image of a drug with difficult drug type identification, the drug identifier 120 estimates the group to which the drug belongs and outputs the estimation result. The groups to which drugs with difficult drug type identification belong can include types such as "capsule drugs", "divided tablets", or "plain drugs". Also, the definition of the group may be further classified in more detail, or sub-groups may be defined by hierarchical classification. For example, the group may be defined as "white single-color capsule drugs", "red and white two-color capsule drugs", "half tablets", "quarter tablets", "white plain drugs, or "transparent plain drugs".
[0088] When the input drug image is an image of a drug with possible drug type identification, the drug identifier 120 identifies the type (drug type) of the drug and outputs the estimation result of the drug type. The identification information of the drug type output by the drug identifier 120 may be, for example, a unique identification code defined for each drug type. Details of the learning method for creating the second trained model TM2 and the content of the identification process by the drug identifier 120 will be described later. The drug identifier 120 searches a drug master database (not shown) from the identification code unique to the estimated drug type, and acquires drug information about the corresponding drug or similar drugs. The drug identifier 120 is an example of the "identifier" in the present disclosure.
[0089] The result of the identification process by the drug identifier 120 is displayed on the display unit 14A via the display control unit 124. After the user checks the identification result displayed on the display unit 14A, the user can confirm the identification result, correct the identification result, or perform other operations such as text search separately.
[0090] The text search unit 122 receives an input of a search key from the input unit 14B or the microphone 18, accesses the database 130 of the printed text master to search for corresponding information, and outputs the search result. The printed text master includes data in which the text information of the characters and symbols printed or marked on various drugs is associated with the types of drugs. The text information of the characters and symbols stored in the database 130 is an example of the "character symbol information" in the present disclosure.
[0091] The storage device 104 stores the database 130 of the printed text master, the master image database 131 including master images of various drugs, and a drug master database (not shown). The master image database 131 may be included in the drug master database. The data such as the printed text master, the master image, and the drug master may be stored on a network such as a ground server (not shown).
[0092] The search result by the text search unit 122 is displayed on the display unit 14A via the display control unit 124. Also, the master image read from the master image database 131 is displayed on the display unit 14A via the display control unit 124.
[0093] After the user checks the search result and the master image displayed on the display unit 14A, the user can confirm the drug type identification result, perform further narrowing-down searches, or perform re-searches. The database 130 is an example of the "first database" in the present disclosure. The master image database 131 is an example of the "second database" in the present disclosure.
[0094] The memory device 104 includes an identification result storage unit 132. The identification result storage unit 132 is a storage area in which identification results of drug types, such as the identification result of the drug type by the drug identifier 120 and the identification result of the drug type specified based on the search result by the text search unit 122, are stored.
[0095] 〔Explanation of the learning phase〕 In the drug type identification device 100 according to the present embodiment, each of the drug detector 114 and the drug identifier 120 is machine - learning - based trained as follows, and these are combined to configure the drug type identification device 100. FIG. 6 is a flowchart showing an outline of a learning phase by machine learning for realizing the drug type identification device 100 including the drug detector 114 and the drug identifier 120.
[0096] The processing of each step shown in FIG. 6 can be implemented, for example, by a computer executing a program. The machine - learning method for obtaining the drug type identification device 100 includes a step of creating training data for training the drug detector 114 (step S1: first training data creation step), a step of training the drug detector 114 by performing machine learning using the training data (step S2: first training step), a step of creating training data for training the drug identifier 120 (step S3: second training data creation step), a step of training the drug identifier 120 by performing machine learning using the training data (step S4), and a step of configuring the drug type identification device 100 using the drug detector 114 trained in step S2 and the drug identifier 120 trained in step S4 (step S5).
[0097] The training data used for supervised learning includes input data and correct answer data (teacher data). Creating the training data includes creating correct answer data (teacher data) corresponding to the input data.
[0098] The process of creating the drug detector 114 (steps S1 and S2) and the process of creating the drug identifier 120 (steps S3 and S4) may be executed in parallel or sequentially. The execution timing of each of steps 1 to 5 is not particularly limited, and they may be executed continuously, or the processing of each individual step may be executed at different times and using different computers. For example, a first computer may execute the process of creating the drug detector 114 (steps S1 and S2), and a second computer different from the first computer may execute the process of creating the drug identifier 120 (steps S3 and S4). Also, the first computer may execute the process of creating training data (steps S1 and S3), and the second computer may execute the learning process (steps S2 and S4).
[0099] In step S1, for the training images, teacher data (correct data) is assigned with respect to the position information and class of each drug contained in the image. The position information of the drug given as the teacher data may be information specifying the position of the bounding box surrounding the drug, or may be a mask filled with the shape of the drug itself. Also, the class label assigned as the teacher data may be uniformly "drug" regardless of the type of drug (drug species). That is, for both drugs whose species can be specified and drugs for which species identification is difficult, the classification labels are all assigned as the same "drug".
[0100] [Learning of the drug detector 114] The drug detector 114 aims to identify the positions of individual drugs from the input captured image, separate the drug regions from the background, and perform cropping, and estimates object position information, class information, the likelihood of the detected object (here, the drug), etc. The object position information may be, for example, "the vertex coordinates of the four points of the non-rotated bounding box surrounding the drug", "the center coordinates, height, and width of the non-rotated bounding box surrounding the drug", "the center coordinates, height, width, and rotation angle of the rotated bounding box surrounding the drug", or "a mask filled with the shape of the drug itself", etc.
[0101] Many drugs exist in the shape of ellipsoids. When ellipsoidal drugs are arranged vertically in an oblique direction, there may be a situation where multiple drugs are included in a single "bounding box without rotation surrounding the drug", so the object position information is preferably the "center coordinates, height, width, and rotation angle of the bounding box with rotation surrounding the drug" or the "mask filled with the shape of the drug itself".
[0102] In step S2, machine learning is performed to train a learning model (hereinafter referred to as the first learning model) applied to the drug detector 114 using the dataset of the training data created in step S1. The first learning model is configured using, for example, a neural network. As a network model suitable for object detection, for example, a convolutional neural network (CNN) can be used. The drug detector 114 is trained to receive an input of an image and output object position information for each individual drug in the image.
[0103] When the drug detector 114 also detects markers, a dataset including training data with the marker region as the correct data (teacher data) is required for training the first learning model. The training data used for the learning of the drug detector 114 is an example of the "first training data" in the present disclosure.
[0104] [Learning of Drug Identifier 120] For the drug identifier 120, images are input in units of images of individual drugs, and it is trained to estimate the class information of the type of drug or the group to which the drug belongs, the probability of the identified class, and the like.
[0105] When creating the training data for the drug identifier 120, for drugs whose drug types can be specified, a unique identification code defined for each drug is assigned as the correct label. On the other hand, for drugs whose drug types are difficult to specify, an identification code is assigned to the group to which the drug belongs. For example, for drugs whose drug types are difficult to specify, the correct label is assigned in units of groups where the processing flow is desired to be branched in the drug type specification phase, such as "capsule drugs", "plain drugs", "half tablets", or "quarter tablets". For "capsule drugs", the correct label may be further assigned by dividing them into finer groups such as "hard capsule drugs" and "soft capsule drugs".
[0106] In order to improve the discrimination performance for the group, it is preferable to use at least one, preferably a combination of a plurality of, the engraved extraction image, the outer shape image of the drug, and the size information of the drug as the input to the drug identifier 120. The engraved extraction image is an image in which only the engraved part or the printed part of the drug is extracted, and the engraved part or the printed part is mainly represented in white on a black background. The training data used for the learning of the drug identifier 120 is an example of the "second training data" in the present disclosure.
[0107] FIG. 7 is a diagram showing an example of the identification code assigned as teacher data for drugs whose drug types can be specified. For drugs whose drug types can be specified, different identification codes are assigned for each drug type. "P000001", "P000002", ··· "P009999" in FIG. 7 represent examples of identification codes uniquely defined corresponding to different drug types. The identification code shown in FIG. 7 is an example of the label of the "type unit" in the present disclosure.
[0108] FIG. 8 is a diagram showing an example of an identification code assigned as teacher data for drugs with difficult drug type identification. For drugs with difficult drug type identification, an identification code is assigned to the group to which the drug belongs. For example, for an image of a capsule drug, an identification code of "G000001" representing the group of "capsule drugs" is assigned regardless of the drug type. Also, for an image of a plain drug, an identification code of "G000002" representing the group of "plain drugs" is assigned regardless of the drug type. For an image of a half tablet (1 / 2 tablet), an identification code of "G000003" representing the group of "half tablets" is assigned regardless of the drug type. For an image of a quarter tablet, an identification code of "G000004" representing the group of "quarter tablets" is assigned regardless of the drug type. Note that when it is desired to handle half tablets and quarter tablets in the same way, the same identification code may be assigned to them. Alternatively, separate identification codes may be assigned to each of the half tablets and quarter tablets, and the same processing may be performed on the identification codes of "G000003" and "G000004" on the program side.
[0109] Also, regarding the identification code assigned to the group, an identification code for a subgroup obtained by classifying the group in more detail according to a hierarchical structure may be defined.
[0110] FIG. 9 is a diagram showing another example of an identification code assigned as teacher data for capsule drugs. For example, for capsule drugs, by more finely grouping them according to capsule color or printed character color, an improvement in the inference accuracy of the drug identifier 120 is expected. Also, by performing such fine grouping, it may be utilized in the discrimination process, such as improving user-friendliness as a search attribute during capsule search. In FIG. 9, in the group of "capsule drugs", an example is shown in which classification (sub-grouping) is defined between the case where the capsule color is a single color and the case where it is a combination of two colors. Also, regarding the group of single-color capsules, further classification may be performed according to the color of the capsule. For an image of a single-color and white capsule drug, for example, an identification code of "G000010" representing the group of single-color white capsule drugs is assigned. For an image of a single-color and blue capsule drug, an identification code of "G000011" representing the group of single-color blue capsule drugs is assigned. Although not shown in FIG. 9, similarly, individual identification codes representing their respective groups may be assigned to images of capsule drugs of other single colors.
[0111] Also, regarding the group of two-color capsules, further classification may be performed according to the combination of capsule colors. For example, for an image of a two-color capsule drug of red and white, an identification code of "G000100" representing the group of red-and-white two-color capsule drugs is assigned. For an image of a two-color capsule drug of blue and white, an identification code of "G000101" representing the group of blue-and-white two-color capsule drugs is assigned. Although not shown in FIG. 9, similarly, individual identification codes representing their respective groups may be assigned to images of capsule drugs of other two-color combinations.
[0112] In FIG. 9, an example of capsule drugs has been described, but similarly for plain tablets, they may be finely grouped according to color or shape and an identification code may be assigned.
[0113] FIG. 10 is a diagram showing another example of an identification code assigned as teacher data for plain drugs. FIG. 10 shows an example of further finely grouping plain drugs according to the combination of the shape and color of the plain drugs. In FIG. 10, examples are shown where the shape of the plain drug is circular and where it is an ellipsoid.
[0114] As shown in FIG. 10, for an image of a plain drug whose shape is circular and whose color is white, for example, an identification code of "G001001" representing a group of circular white plain drugs is assigned. For an image of a plain drug whose shape is circular and whose color is yellow, an identification code of "G001002" representing a group of circular yellow plain drugs is assigned. Although not shown in FIG. 10, for images of plain drugs whose shape is circular and with other colors, individual identification codes representing their respective groups may be assigned in the same way.
[0115] Also, for an image of a plain drug whose shape is an ellipsoid and whose color is orange, for example, an identification code of "G002001" representing a group of ellipsoid and orange plain drugs is assigned. For an image of a plain drug whose shape is an ellipsoid and whose color is transparent, an identification code of "G002002" representing a group of ellipsoid and transparent plain drugs is assigned. Although not shown in FIG. 10, for images of plain drugs whose shape is an ellipsoid and with other colors, individual identification codes representing their respective groups may be assigned in the same way. Also, for images of plain drugs with combinations of shapes and colors not shown, individual identification codes representing their respective groups may be assigned in the same way.
[0116] [Regarding the input information to the drug identifier 120] When taking pictures with the smartphone 10 in various environments, due to the influence of the shooting environment, color information may not always be a reliable information source. Therefore, when identifying a drug that can be identified by drug type on a per-drug basis, a stamped extraction image with color information excluded is often important. In the drug identifier 120, it is preferable to perform identification while attaching more importance to the stamped extraction image. A learned machine learning model that inputs the original image and outputs the stamped extraction image may be used, and the stamped extraction image may be obtained by inputting the original image and performing inference.
[0117] On the other hand, even if greatly affected by the shooting environment, color information may be useful in some cases, such as for drugs with very rare colors. Also, the shape information of the drug can be a robust information source that is less affected by the shooting environment. The external shape image of the drug can be useful shape information.
[0118] Also, information on the major and minor axis sizes (size information) of the drug obtained by non-machine learning methods can be a robust information source that is less affected by the shooting environment. For example, by using OpenCV or the like to extract the major and minor axes of the drug and using the numerical information, it becomes a robust information source that is less affected by the shooting environment. Thus, the size information measured for each drug from the captured image is an important information source for identifying a huge number of drug types even for drugs difficult to identify by drug type. In particular, for drugs difficult to identify by drug type, such as capsule drugs or plain tablets without an identification symbol printed, the size information can be a very important information source for group identification.
[0119] Considering these facts, as the input to the drug identifier 120, it is preferable to use one or more, preferably a combination of multiple, of the original image which is the region image (drug image) cut out in drug units from the captured image, the engraved extraction image extracted from the original image, the outer shape image, and the size information. Note that the term "engraved extraction image" includes not only an image obtained by extracting an engraving but also an image obtained by extracting printed characters attached to tablets or capsules. The engraved extraction image may be referred to as a character symbol extraction image. The term "engraving" may be understood as a term including concepts such as "printed characters", "printed symbols", "identification symbols", or "character symbols" for tablet or capsule drugs as necessary depending on the context.
[0120] It may be configured to input all the information of the original image including color information, the engraved extraction image, the outer shape image, and the size information into the drug identifier 120 in combination, or it may be configured to input a combination of some of this information into the drug identifier 120.
[0121] FIG. 11 is a conceptual diagram showing an example of information input to the drug identifier 120. The original image Org is a drug image in drug units and corresponds to individual drug images cut out from the captured image for each drug. The engraved extraction image Egm is an image obtained by extracting an engraving from the original image Org. The engraved extraction image Egm may be an image subjected to an enhancement process to improve the visibility of the extracted engraving. The outer shape image Otw is an image showing the outer shape of the drug extracted from the original image Org. Note that the outer shape image Otw is not limited to showing the exact contour of the drug and may show a general shape according to the contour. The size information is, for example, numerical information obtained by measuring the major axis and minor axis of the drug based on the original image Org or the outer shape image Otw.
[0122] In order to construct a system robust against the shooting environment, it is preferable to use a combination of these multiple pieces of information as the input information to the drug identifier 120.
[0123] The learning model applied to the drug identifier 120 is configured using, for example, a neural network. As a network model suitable for image recognition, for example, a CNN can be used. When combining two or more of the original image, the engraved image, and the outer shape image as inputs to the neural network, there can be two methods: combining these multiple images in the channel direction for input, and combining them horizontally or vertically within the image plane to form a composite image for input.
[0124] In a neural network, since the input image size may be subject to the constraint of being a fixed value, there can be the following two methods as input image methods.
[0125] [Method 1] The first method is to enlarge or reduce the image according to the input image size that the input layer of the neural network can accept. Fig. 12 shows an example of input information by the first method.
[0126] In the first method, each image is enlarged or reduced to the upper limit of the input image size that the neural network can accept and then input. Information on the enlargement ratio or reduction ratio of the image used in this image enlargement or reduction process (hereinafter referred to as "resizing process") is stored for future use if necessary. In the case of the first method, the magnification information applied to the resizing process is stored. As input information to the neural network, in addition to the combination of one or more of the original image Org1, the engraved image Egm1, and the outer shape image Otw1, and the size information, the magnification information of the resizing process is used as needed. The original image Org1 shown here is an image of a drug unit cut out from the standardized image of the photographed image, resized to the input image size of the neural network. Also, the engraved image Egm1 is obtained by performing a process of extracting the engraving from the original image Org1. The outer shape image Otw1 is obtained by performing a process of extracting the outer shape of the drug from the original image Org1.
[0127] According to the first method, for example, when the printed characters are small or fine on a small tablet, the image is enlarged and input into the neural network, which has the advantage of improving the inference (identification) accuracy. Also, according to the first method, the input image size of the neural network can be designed to be smaller, and a reduction in execution time can be expected. That is, it is not restricted by the maximum size of the drugs existing in the world, and since the input image size of the neural network can be designed to an appropriate size for the recognition performance, the processing of unnecessary data is suppressed.
[0128] On the other hand, in the case of the first method, the processing becomes somewhat complicated in that image resizing processing and the like are performed when the neural network inputs and outputs. Also, in case it is necessary in subsequent processing (such as when presenting the drug identification result to the user with an image proportional to the size of the tablet), it is necessary to store the information on the magnification ratio (magnification information).
[0129] [Method 2] The second method is a method in which an image is pasted and input at the image center position of the input image size received by the input layer of the neural network in the original image size (without performing enlargement / reduction processing). In this case, the input image size for the neural network is determined in advance, and the same as the original image cut out from the standardized image of the photographed image is input to the input layer with an acceptable input image size. The standardized original image is an image in which the correspondence with the actual size of the drug is grasped.
[0130] Fig. 13 shows an example of input information according to the second method. According to the second method, image resizing processing is unnecessary. Therefore, in the case of the second method, input of magnification information indicating the enlargement / reduction ratio is also unnecessary. In the second method, as the input information of the neural network, a combination of one or more images among the original image Org2, the engraved extraction image Egm2, and the outer shape image Otw2 and the size information can be used. Also, in the second method, since the outer shape image Otw2 itself extracted from the original image Org2 substantially includes information indicating the size of the drug, when using the outer shape image Otw2 as the input, there is an advantage that the numerical information on the major axis and the minor axis is not necessarily required.
[0131] On the other hand, in the case of the second method, since the input image size that the neural network accepts is determined by the largest drug size in the world, in the case of particularly small tablets or the like, extra space is created around the tablet, and waste may occur in the processing. In addition, since the input image size that the neural network accepts is larger than that of the first method, the execution time may be longer than that of the first method. Furthermore, in the case of the second method, since small tablets and tablets with fine engraved characters are processed in their original sizes, the identification accuracy of these tablets may be lower than that of the first method.
[0132] Comparing the first method and the second method, for drug type identification using the smartphone 10, the first method is a more preferable method.
[0133] 〔Examples of input information for drugs with difficult drug type identification〕 Figs. 14 to 16 show examples of input information for drugs with difficult drug type identification. Here, an example of inputting information on drugs with difficult drug type identification into the neural network by the first method is shown.
[0134] Fig. 14 is an example of input information in the case of a capsule drug (with identification mark) having an identification mark printed on the capsule. In the case of a capsule drug with an identification mark, similar to Fig. 12, as input to the neural network, a combination of one or more images among the original image Org3, the engraved character extraction image Egm3, and the outer shape image Otw3 is input. In addition to the images, numerical information on the size (major axis and minor axis) of the capsule drug measured from the original image Org3 may be input. Furthermore, in addition to these information, magnification information for the resizing process may be input. The engraved character extraction image Egm3 is an image of the identification mark extracted from the original image Org3.
[0135] Figure 15 shows an example of input information in the case of a capsule drug without an identification symbol. A capsule drug without an identification symbol where the identification symbol does not appear in the original image Org4 may be a capsule drug originally without an identification symbol printed on the capsule, or may be a case where, even though there is an identification symbol on the capsule, the printing was hidden during shooting and thus the identification symbol does not appear in the original image Org4.
[0136] For a capsule drug without an identification symbol, a combination of one or more of the original image Org4, the engraved extraction image Egm4, and the outer shape image Otw4 is input. However, for a capsule drug without an identification symbol, the engraved extraction image Egm4 is an image that does not contain any information about the identification symbol. The outer shape image Otw4 is an image obtained by extracting the shape of the capsule drug shown in the original image Org4. Also, similar to FIG. 14, in addition to these images, numerical information on the size (major axis and minor axis) of the capsule drug measured from the original image Org4 may be input. Furthermore, in addition to these information, magnification information for the resizing process may be input.
[0137] Figure 16 shows an example of input information in the case of a half tablet. Similarly, in the case of a half tablet, a combination of one or more of the original image Org5, the engraved extraction image Egm5 extracted from the original image Org5, and the outer shape image Otw5 is input. Also, in addition to the image, a combination of numerical information on the size (major axis and minor axis) of the half tablet measured from the original image Org5 and magnification information for the resizing process may be input.
[0138] Figure 17 is a block diagram showing the functional configuration of the machine learning system 150. Here, an example of the case where the combination of input information described in FIG. 12 is input to the learning model 151 is shown.
[0139] The machine learning system 150 is a device that generates a second trained model TM2 to be applied to the drug identifier 120, and is realized using a computer system including one or more computers. The machine learning system 150 includes a learning model 151, a loss calculation unit 152, and an optimizer 154. A neural network such as a CNN is used for the learning model 151.
[0140] The learning model 151 receives, as input information, a combination of the drug image IMj, which is the area image of the drug DRj, the engraved character extraction image IM1j, the outer shape image IM2j, the size information SZj, which is numerical information indicating the size of the drug, and the magnification information MGj of the resizing process, and outputs an inference result PRj of the type of the drug DRj or the group to which the drug DRj belongs. The subscript j represents the index number of the training data.
[0141] The machine learning system 150 shown in FIG. 17 includes an engraved character extraction unit 140, an outer shape extraction unit 142, and a size measurement unit 144 in front of the learning model 151. The engraved character extraction unit 140 extracts engraved or printed characters and symbols from the drug image IMj, and generates an engraved character extraction image IM1j, which is an image of the extracted characters and symbols. The engraved character extraction unit 140 processes the drug area of the input image to exclude the outer shape edge information of the drug DRj and extracts the characters and symbols. The engraved character extraction image IM1j is an image in which the characters and symbols are emphasized by expressing the brightness of the engraved or printed part relatively higher than the brightness of the part other than the engraved or printed part.
[0142] The outer shape extraction unit 142 extracts the outer shape of the drug DRj from the drug image IMj and generates an outer shape image IM2j, which is an image showing the outer shape of the drug DRj. The size measurement unit 144 measures the size of the drug DRj from the drug image IMj and / or the outer shape image IM2j, and generates size information SZj indicating the dimensions of the major axis and the minor axis respectively.
[0143] In the machine learning system 150, a configuration is adopted in which these pieces of information are generated from the drug image IMj and input into the learning model 151. However, at the stage of preparing the training data in advance, a part or all of the engraved extraction image IM1j, the outer shape image IM2j, and the size information SZj may be created and included in the dataset of the training data. In that case, a part or all of the engraved extraction unit 140, the outer shape extraction unit 142, and the size measurement unit 144 are unnecessary in the machine learning system 150.
[0144] The loss calculation unit 152 calculates the loss value (loss) between the two based on the inference result output from the learning model 151 and the correct data (teacher data) GTj associated with the input information.
[0145] The optimizer 154 determines the update amount of the parameters of the learning model 151 based on the calculation result of the loss value indicating the error between the output of the learning model 151 and the correct teacher signal so that the inference result PRj output by the learning model 151 approaches the correct data GTj, and performs the update process of the parameters of the learning model 151. The optimizer 154 updates the parameters based on an algorithm such as the gradient descent method. The parameters of the learning model 151 include the filter coefficients (weights of the connections between nodes) of the filters used in the processing of each layer of the neural network and the biases of the nodes. The machine learning system 150 may perform the acquisition of the training data and the update of the parameters in units of mini-batches that combine a plurality of training data.
[0146] In this way, by performing machine learning using a large number of training data, the parameters of the learning model 151 are optimized, and a learning model 151 with the target inference performance is generated. The learned (trained) learning model 151 for which the allowable inference accuracy has been confirmed is used as the second learned model TM2 of the drug identifier 120.
[0147] In this case, in order for the drug type identification device 100 to obtain a printed matter extraction image, an outer shape image, and size information from the captured image, for example, between the drug region cutting unit 118 and the drug identifier 120 described in FIG. 5, a processing unit similar to the printed matter extraction unit 140, the outer shape extraction unit 142, and the size measurement unit 144 is provided. The drug type identification device 100 is an example of the "object identification device" in the present disclosure.
[0148] 〔Utilization Phase of Drug Type Identification Device 100〕 FIG. 18 is a flowchart showing the operation of the drug type identification device 100 according to the present embodiment. Here, as an example of the utilization mode of the drug type identification device 100, the case of identifying the drugs brought by a certain patient will be described as an example.
[0149] In step S11, the user simultaneously captures a plurality of drugs to be identified. The user can capture a plurality of drugs in a pharmaceutically meaningful unit such as the same administration time. For example, the user takes out a bagged drug included in the drugs brought by the patient and simultaneously (collectively) captures these plurality of drugs using the camera function of the smartphone 10. The processor 102 acquires the captured image obtained by the capture. At this time, the plurality of drugs captured may contain drugs whose types can be specified and drugs whose types are difficult to specify. The drugs whose types can be specified are an example of the "objects whose types can be specified" in the present disclosure, and the drugs whose types are difficult to specify are an example of the "objects whose types are difficult to specify" in the present disclosure.
[0150] Next, in step S12, the processor 102 detects individual drugs from the captured image by the drug detector 114. The drug detector 114 detects drugs whose drug types can be specified and drugs whose drug types are difficult to specify in individual drug units from the captured image. The drug detector 114 estimates, for example, "the center coordinates, height, width, and rotation angle of a bounding box with rotation that surrounds the drug", or "a segmentation mask that fills the drug region with the shape of the drug itself", and outputs the estimation result. The estimation (detection) result by the drug detector 114 is displayed on the touch panel display 14 and provided for confirmation by the user. The processor 102 receives an instruction to correct the detection result or an instruction to approve (confirm) the detection result from the input unit 14B of the touch panel display 14 or the like. If there is over-detection or detection failure (detection omission) in the detection result of the drug detector 114, the user can input an instruction to correct the detection result from the input unit 14B of the touch panel display 14 or the like and specify the correct region of each drug. Note that the correction of the detection result can be performed in units of the detected regions or in units of drugs. The drug unit is an example of the "object unit" in the present disclosure.
[0151] In step S13, the processor 102 determines whether there is over-detection or detection failure in the detection result by the drug detector 114. If there is over-detection or detection failure in the detection result by the drug detector 114 and the determination result in step S13 is a Yes determination, the processor 102 proceeds to step S14 and performs a process of correcting the region of the drug according to the instruction received from the user. For example, if an over-detected region is included, a process of deleting that region is performed. Also, if there is a detection failure where a region is not detected although there is a drug, a process of adding a new region for the drug is performed, etc. After step S14, the processor 102 proceeds to step S15.
[0152] Also, when the determination result in step S13 is a No determination, that is, when there is no over-detection or detection failure in the detection result by the drug detector 114, the processor 102 proceeds to step S15. In step S15, the processor 102 performs identification for each drug in the image using the drug identifier 120. The processor 102 cuts out the drug images in units of the detected drug units, inputs the individual drug images into the drug identifier 120, and performs drug identification.
[0153] The details of the processing in step S15 will be described later with reference to FIG. 19, but the outline is as follows. That is, for drugs for which the drug type can be specified, it is expected that the drug identifier 120 makes an inference in units of drug types (drug species). When it is possible to visually confirm that the drug is correct with respect to the drug species identification result obtained by the drug identifier 120, the inferred result is determined. On the other hand, when the drug identifier 120 erroneously makes an inference in units of a group of drugs for which it is difficult to specify the drug species, select from the drugs presented as other higher-level inference candidates, or specify the drug species by voice search, text search, etc.
[0154] Also, for drugs for which the drug type can be specified, it is expected that the drug identifier 120 makes an inference in units of the group to which the drug belongs. When an inference is made in units of a group as a drug for which it is difficult to specify the drug species, if it is possible to visually confirm that the inference is correct in the photographed image of the drug, proceed to the drug species specification flow defined corresponding to that group. For example, in the case of the group of "capsule drugs", visually observe the alphanumeric symbols attached to the capsule in the photographed image, input the text of the characters and / or symbols by text input or voice input, search the database of the printed text master, and collate the drug to be identified with the master data. Then, shift to a drug species specification flow such as specifying the drug species.
[0155] On the other hand, when the inference result of the drug identifier 120 for drugs for which it is difficult to specify the drug species is incorrect, the user appropriately corrects the inference result and proceeds to an appropriate drug species specification flow (see FIG. 19).
[0156] In step S16, the processor 102 determines whether the drug types of all the drugs in the image have been determined. If the determination result in step S16 is a No determination, the processor 102 returns to step S15, changes the drug to be identified, and continues the process. If the determination result in step S16 is a Yes determination, the processor 102 ends the flowchart of FIG. 18.
[0157] FIG. 19 is a flowchart showing an example of loop processing applied to step S15 and step S16 of FIG. 18.
[0158] When the process of step S15 is started, in step S21, the processor 102 uses the drug identifier 120 to identify the drug, and determines whether the identification result belongs to a drug for which the drug type can be specified. The identification result by the drug identifier 120 is provided to the user through the touch panel display 14. The processor 102 receives inputs of various instructions such as an instruction to determine the drug type, an instruction to correct the identification result, or an instruction to shift to other processes such as text search from the input unit 14B etc. of the touch panel display 14. The user can check the presented identification result and input an instruction to determine the drug type from the input unit etc. of the touch panel display 14, or input an instruction such as correction of the identification result or shift to engraved text search.
[0159] If the determination result in step S21 is a Yes determination, the processor 102 proceeds to step S22. In step S22, the processor 102 determines whether the result that the drug is a drug for which the drug type can be specified by the drug identifier 120 is correct. If the determination result in step S22 is a Yes determination, the processor 102 proceeds to step S23 and determines whether the drug type specified (inferred) by the drug identifier 120 is correct. When the user can visually confirm that it is the correct drug type, the user can input an instruction to determine the drug type.
[0160] If the determination result in step S23 is a Yes determination, the processor 102 proceeds to step S29 and determines the drug type.
[0161] If the determination result in step S23 is a No determination, the processor 102 proceeds to step S28. In step S28, the processor 102 executes a non-machine learning-based drug type identification flow using engraved text or the like. After step S28, the processor 102 proceeds to step S29.
[0162] If the determination result in step S21 is a No determination, the processor 102 proceeds to step S24. In step S24, the processor 102 determines whether the result of the drug type identification difficult drug identified by the drug identifier 120 is correct. If the determination result in step S24 is a No determination, the processor 102 proceeds to step S28.
[0163] If the determination result in step S24 is a Yes determination, the processor 102 proceeds to step S25 and determines whether the type of the group to which the drug type identification difficult drug identified (inferred) by the drug identifier 120 belongs is correct. If the determination result in step S25 is a No determination, or if the determination result in step S22 is a No determination, the processor 102 proceeds to step S26. In step S26, the processor 102 identifies the group of drug type identification difficult drugs to which the drug belongs. In step S26, the processor 102 receives an input of an instruction to specify the type of the group to which the drug belongs from the input unit 14B of the touch panel display 14 or the like, and identifies the group according to the received instruction.
[0164] After step S26, the processor 102 proceeds to step S27. Also, if the determination result in step S25 is a Yes determination, the processor 102 proceeds to step S27. In step S27, the processor 102 executes a drug type identification flow defined by the group of the drug type identification difficult drugs. After step S27, the processor 102 proceeds to step S29 to determine the drug type.
[0165] When the drug type is determined for all the drugs included in the image, the loop process in FIG. 19 is terminated. The drug identification method described using the flowcharts of FIGS. 18 and 19 is an example of the "object identification method" in the present disclosure.
[0166] [Example of GUI (Graphical User Interface) of Drug Type Identification Device 100] [Detection Result Display GUI] FIG. 20 is a diagram showing an example of a screen displayed on the touch panel display 14. FIG. 20 shows an example of a detection result display GUI that displays the detection results of the drug regions detected by the drug detector 114. The screen SC1 provided by the drug type identification device 100 is roughly composed of three areas: an overall image display area EDA, a candidate display area CDA, and a button display area BDA. A standardized image generated from the acquired captured image is displayed in the overall image display area EDA. Note that the entire captured image including the marker may be displayed in the overall image display area EDA. The overall image display area EDA is an example of the "captured image display area" in the present disclosure.
[0167] The overall image display area EDA can display a part of the image enlarged or reduced by a pinch-out or pinch-in operation. Therefore, when the character symbols etc. attached to the drug are small and difficult to see, each drug can be enlarged and displayed.
[0168] The detection results by the drug detector 114 are displayed in the overall image display area EDA. After shooting, the areas of each drug detected by the drug detector 114 are displayed as rectangular frames (bounding boxes). FIG. 20 shows an example of a captured image including three drugs DR1, DR2, and DR3. For the drugs DR1 and DR2 detected by the drug detector 114, frames BX1 and BX2 are respectively displayed. In the example shown in FIG. 20, the detection of the drug DR3 has failed, and no frame for the drug DR3 is displayed. Also, in the example shown in FIG. 20, there is an over-detection for the area Abs where no drug exists, and a frame BX3 based on a false detection is displayed.
[0169] Each area surrounded by frames BX1, BX2, and BX3 can be selected by the user. For example, the selected area is displayed with a red frame, and the unselected area is displayed with a different frame color such as a blue frame. Of course, other colors may be used. It is preferable that the frames BX1, BX2, and BX3 indicating the areas are displayed with different frame colors according to the selected / unselected state of the areas.
[0170] Each of the frames BX1, BX2, and BX3 is provided with a check box CB. The check box CB is marked or left blank according to the status such as "drug type determined" or "undetermined state".
[0171] At the bottom of the entire image display section EDA, an area editing switch SW1, an area addition button BT2, and an area deletion button BT3 are displayed. The area editing switch SW1 is a switch operated when it is desired to correct the detection result of the drug area. When the area editing switch SW1 is slid to the right on the screen, the drug area can be edited on the screen of the entire image display section EDA. When the area editing switch SW1 is turned on, the area addition button BT2 and the area deletion button BT3 can be pressed, and the process moves to the area correction GUI. When the area editing switch SW1 is slid to the left on the screen, the area cannot be edited. In the state where editing is not possible, the area addition button BT2 and the area deletion button BT3 are grayed out.
[0172] Note that the on / off of area editing ability is not limited to the area editing switch SW1 illustrated in FIG. 20. For example, it may be implemented by long-pressing or double-tapping the entire image display section EDA. When the area addition button BT2 is pressed, an area can be specified in the entire image display section EDA to add a drug area. Also, the area deletion button BT3 is used when deleting an over-detected area or the like.
[0173] When each area is tapped and selected with the area editing turned off, identification processing using the drug identifier 120 is executed on the drug image cut out in that area, and drug information of candidate drugs is displayed in the candidate display area CDA below. Also, for each area, the dimensions of the major axis and minor axis are displayed. The on / off of this dimension display may be configured to be switchable with a separate switch or the like. The candidate drugs are an example of the "candidate objects" in the present disclosure.
[0174] The candidate display area CDA is an area mainly for displaying drug information of candidate drugs and the like. In the candidate display area CDA, drug information of candidate drugs is displayed based on the identification result estimated by the drug identifier 120. The drug information of the candidate drugs includes, for example, a master image of the drug, an identification code of the drug, a drug surface, a pharmacological classification, and information on the group to which the drug belongs. In FIG. 20, an example of drug information of candidate drugs for the drug DR1 is shown. Note that in the candidate display area CDA, the image can also be enlarged / reduced by a pinch-out or pinch-in operation.
[0175] The display column for displaying the group information is configured as a selection box SLB1. Also, in the candidate display area CDA, a printed text search button BT4 is displayed.
[0176] In the button display area BDA, various buttons including, for example, a drug type confirmation button BT5, a drug type confirmation hold button BT6, a re-shooting button BT7, and a completion button BT8 are arranged. These buttons can change to a pressable state or a non-pressable state according to the situation. The re-shooting button BT7 is used when there is a problem with the captured image, such as when the captured image is blurred, and is pressed to perform re-shooting.
[0177] [Area Editing GUI] FIG. 21 is an example of a screen display when performing area editing of a drug. As illustrated in FIG. 20, when over-detection and / or detection failure are confirmed in the detection result of the drug detector 114, the user slides the area editing switch SW1 to the right direction of the screen to make the area editable.
[0178] In this state, for example, when the area of drug DR1 is selected, the selected area can have its position, shape, rotation angle, etc. adjusted by operations such as dragging, tapping, etc. For example, as shown in FIG. 22, the area can be scaled by dragging the side of the frame BX, translated by dragging the area, or rotated by pinching the area with two fingers in the rotation direction. Also, by dragging the corner part of the frame BX, the area can be scaled while keeping the aspect ratio constant.
[0179] In the case of the screen SC2 shown in FIG. 21, since the area Abs where there is no drug has been detected as a drug due to over-detection, the user can select the area Abs indicated by the frame BX3 and delete the area Abs by pressing the area deletion button BT3. Note that the "area deletion" function is not limited to the area deletion button BT3. The GUI may be implemented such that the corresponding area is long-pressed to display a pull-down menu (not shown), and deletion is performed by selecting "area deletion" from the pull-down menu.
[0180] Also, in the case of the example shown in FIG. 21, for drug DR3, since the drug DR3 exists in the captured image but the area has not been detected, the user can press the area addition button BT2, select an area, and add the area. The operation of adding an area is, for example, after tapping an arbitrary point, dragging it diagonally down to create a rectangular area, and adjusting the width, height, and rotation angle of the rectangle by the method described in FIG. 22.
[0181] Alternatively, it may be implemented such that a point on the screen of the entire image display area EDA is long-pressed to display a default rectangular area with uncertain area for adding an area, and the rectangle is translated to an appropriate position and the width, height, and rotation angle of the rectangle are adjusted for correction.
[0182] [GUI for Displaying Detection Results with Confirmed Areas] Figure 23 is an example of a screen display in a state where the identification of drugs has been determined. As shown in the screen SC3 of Figure 23, check marks indicating the status of "drug type determined" are entered in the check boxes CB for the respective frames BX1 and BX4 of the drugs DR1 and DR3 for which the drug type has been determined. By long-pressing the check box CB or the like, a pull-down menu can be displayed, and the status can be changed from the pull-down menu. For example, it can be changed from the status of "drug type determined" to "undetermined state" or the like.
[0183] Also, for drugs with difficult-to-identify drug types, etc., the drug type cannot necessarily be identified, so there may be a status of "drug type determination on hold". For example, an "X" mark indicating the status of "drug type determination on hold" may be displayed in the check box CB of the drug DR2 for which the determination of the drug type is on hold. In the case of the "undetermined state" where neither the drug type has been determined nor is it on hold, nothing is displayed in the check box CB (see Figure 22). The default of the check box CB is a blank indicating the "undetermined state".
[0184] After determining the identification status of all drugs, the user presses the completion button BT8 to complete the identification work for this photographed image.
[0185] [Identification result confirmation GUI (in the case of drugs with identifiable drug types)] Figure 24 is an example of a screen display of the identification result confirmation GUI. Figure 24 shows an example of a GUI for displaying drug information about drugs belonging to the group of "drugs with identifiable drug types". Although the illustration of the overall image display section EDA is omitted in Figure 24, the overall image display section EDA exists above the candidate display section CDA of the screen SC4 shown in Figure 24 (see Figure 23).
[0186] In the candidate display unit CDA, the front and back of the master image of the drug estimated by the drug identifier 120 for the drug cut-out image of the area selected by the overall image display unit EDA, the drug code, the drug name, and drug information such as the size (dimensions of the major axis and minor axis) are displayed. Also, at the lower part of the drug information display unit where these drug information are displayed, the type of the group to which the selected drug belongs is displayed. The type of the group is in the form displayed in the selection box SLB1.
[0187] When the right arrow button BT9 at the right end of the screen SC4 is pressed or a left swipe is performed, candidate drugs with lower scores are sequentially displayed in the drug information display unit of the candidate display unit CDA. Also, when the left arrow button BT10 at the left end of the screen is pressed or a right swipe is performed, candidate drugs with higher scores are sequentially displayed in the drug information display unit of the candidate display unit CDA.
[0188] Also, when the inferred group is incorrect, by pressing the pull-down button of the selection box SLB1, the pull-down menu PDM is displayed, and the correct group can be selected from the pull-down menu PDM to transition to the GUI specific to each group.
[0189] When the correct drug cannot be found among the candidate drugs output by the drug identifier 120, when performing a search by a text-based method independent of the drug type estimation of the drug identifier 120, by pressing the engraved text search button BT4, the transition to the engraved text search GUI is made.
[0190] The drug type determination button BT5 is pressed when determining the drug type for the drug displayed or selected in the candidate display unit CDA. When the drug type determination button BT5 is pressed, the drug type is determined, and a check mark is displayed in the check box CB of the corresponding drug in the overall image display unit EDA. The drug type determination hold button BT6 is pressed when completing the discrimination without determining the drug type for the selected drug. When the drug type determination hold button BT6 is pressed, an "×" mark is displayed in the check box CB of the corresponding drug in the overall image display unit EDA.
[0191] When the arrow button BT11 in the upper right of the candidate display section CDA is pressed, the system transitions to the candidate list display GUI (see Figure 25).
[0192] [Candidate List Display GUI] Figure 25 is an example of the screen display of the candidate list display GUI. The candidate list display GUI is a GUI that lists the drugs with the highest scores in the inference by the drug identifier 120 when the drug related to the selection belongs to the group of "drugs whose drug type can be specified". As shown in Figure 25, since a plurality of candidate drugs are listed, comparison and consideration become possible. Although the illustration of the overall image display section EDA and the button display section BDA is omitted in Figure 25, the overall image display section EDA exists above the screen SC5 of the candidate display section CDA shown in Figure 25, and the button display section BDA exists below the screen of the candidate display section CDA (see Figure 23). The same applies to each of Figures 26 to 29.
[0193] Figure 25 shows an example in which drug information about the top 4 drugs in terms of score is listed. When the knob of the scroll bar SRB at the right end of the screen SC5 is operated downward or the screen is swiped, lower-ranked candidates can also be displayed. On the screen SC5 of this candidate image list, any drug is selected and the "drug type confirmation button BT5" in the button display section BDA is pressed, whereby the drug type can be confirmed.
[0194] In addition to the score order of the inference by the drug identifier 120, the display order in the list display may be, for example, to search for drugs similar to the engraved text of the drug that ranked first in the score by the drug identifier 120 from the engraved text database and display them in the order of the similarity scores.
[0195] Also, the display form of the list display is not limited to the form of arranging them in a vertical column as shown in Figure 25, and it may be a multi-line and multi-column display using more of the screen.
[0196] In the candidate image list, as the drug information of each drug, in addition to the master image (front and back), character information indicating the drug name, and numerical information on the size, it is preferable to display the engraved text information (master character information) registered in the engraved text database. In FIG. 25, for the sake of illustration, the notation "master character information" is used, but in the actual screen display, the engraved text information registered for each drug will be displayed.
[0197] When the arrow button BT11 in the upper right of the candidate display section CDA is pressed, it returns to the identification result confirmation GUI (see FIG. 24).
[0198] [Engraved Text Search GUI] FIG. 26 is an example of the screen display of the engraved text GUI. The engraved text GUI is a GUI that displays a list of drugs with text similar to the text entered in the search box SBX1 when the drug related to the selection belongs to the group of "drugs whose drug type can be specified". When the engraved text search button BT4 is pressed on the screen of FIG. 24 or FIG. 25, as shown in FIG. 26, the screen SC6 of the engraved text GUI including the search box SBX1 is displayed.
[0199] The user can search by engraved characters by entering text in the search box SBX1 by keyboard input or voice input and pressing the search button SBT1. The search results are displayed in a list in the same way as in FIG. 25. The display format of the list display of candidate drugs is also the same as in FIG. 25. Similarity score determination may be performed every time one character is entered in the search box SBX1, and the candidate list display may be updated immediately.
[0200] [Capsule Search GUI] FIG. 27 is an example of the screen display of the capsule search GUI. The capsule search GUI is a GUI displayed in the candidate display section CDA when it is identified as a "capsule drug" by the drug identifier 120, or when "capsule" is selected in the pull-down menu PDM of the group selection box SLB1.
[0201] The capsule search screen SC7 includes a text search box SBX2, a size search box SZB, capsule color selection boxes SLB2 and SLB3, a character color selection box SLB4, a capsule image display section SRD, a group selection box SLB1, and a capsule list button BT13.
[0202] By entering text in the text search box SBX2 either through keyboard input or voice input and pressing the search button SBT2, it is possible to search by the printed characters on the capsule. It is also possible to search by the text of the drug name.
[0203] The size search box SZB includes an input box IB1 for entering the major axis value, an input box IB2 for entering the minor axis value, and a search button BT12. By entering numerical values in the input boxes IB1 and IB2 either through keyboard input or voice input and pressing the search button BT12, it is possible to search by the size of the capsule drug.
[0204] The search results are displayed in the capsule image display section SRD. In the case of text search, the front and back of the master image are displayed from the top in descending order of the text similarity score. Similarity score determination can be performed every time a single character is entered in the text search box SBX2, and the candidate list display can be updated immediately. In capsule search, text search is limited to capsules. On the other hand, when searching by size, the front and back of the master image are displayed from the top in descending order of the degree of match with the entered major axis and minor axis values. In this case as well, the search range is limited to capsules for the search.
[0205] Regarding the display order of the master images of the capsule drugs to be displayed in the capsule image display section SRD, past discrimination history data can be referred to, and capsules with a discrimination history can be preferentially displayed. When the knob of the scroll bar SRB at the right end of the screen SC7 is operated downward or the screen SC7 is swiped, lower candidates can also be displayed.
[0206] Regarding the form of candidate display, it is not limited to the form of arranging and displaying in a vertical column as shown in FIG. 27, and a multi-line and multi-column display may also be used.
[0207] When the capsule list button BT13 at the lower right of the screen SC7 is pressed, the capsule image display unit SRD expands to the right, and a list of more capsule drug candidates is displayed at once. In this case, the capsule list button BT13 is replaced with a "return button" (not shown), and when the return button is pressed, it returns to the original capsule candidate screen (FIG. 27).
[0208] By selecting any one of the capsule images displayed in the capsule image display unit SRD and pressing the "drug type confirmation button BT5" of the button display unit BDA, the drug type can be confirmed.
[0209] The capsule color selection boxes SLB2, SLB3, and the character color selection box SLB4 may be used to specify the color of the capsule and the color of the printed characters to narrow down the candidates. Also, the capsule color may be automatically recognized from the photographed image to narrow down the candidates in advance, and the selection boxes SLB2, SLB3 may be displayed by default. In the list of candidate images of the capsule, as the drug information of each drug, in addition to the master image (front and back), the character information indicating the drug name, and the numerical information of the size, it is preferable to display the engraved text information (master character information) registered in the engraved text database.
[0210] [Plain Drug Search GUI] FIG. 28 is an example of the screen display of the plain drug search GUI. The plain drug search GUI is a GUI displayed in the candidate display unit CDA when it is identified as a "plain drug" by the drug identifier 120, or when "plain drug" is selected in the pull-down menu PDM of the group selection box SLB1.
[0211] The screen SC8 for plain drug search includes a color selection box SLB5, a plain drug image display unit SRD2, a group selection box SLB1, and a plain drug list button BT14.
[0212] The plain drug image display section SRD2 displays the front and back of the master image from the top in descending order of the degree of match with the values of the major and minor diameters of the drug selected by the overall image display section EDA. In this case, the search range is limited to plain tablets for the search.
[0213] Regarding the display order of the master images of plain drugs to be displayed on the plain drug image display section SRD2, past discrimination history data may be referred to, and plain drugs with a discrimination history may be preferentially displayed. When the knob of the slider bar at the right end of the screen is operated downward or the screen is swiped, lower candidates can also be displayed.
[0214] Regarding the form of candidate display, it is not limited to the form of arranging them in a vertical column as shown in FIG. 27, and a multi-line and multi-column display may be used.
[0215] When the plain drug list button BT14 at the lower right of the screen SC8 is pressed, the plain drug image display section SRD2 expands to the right, and a list of more plain drug candidates is displayed at once. In this case, the plain drug list button BT14 is replaced by a "return button" (not shown), and when the return button is pressed, it returns to the original plain drug candidate screen (FIG. 28).
[0216] The color of the plain drug may be specified using the color selection box SLB5 to narrow down the candidates. Also, the color of the plain drug may be automatically recognized from the captured image to narrow down the candidates in advance, and the selection box SLB5 may be displayed by default.
[0217] The user can select any capsule from the capsule images displayed on the capsule image display section SRD and press the drug type confirmation button BT5 of the button display section BDA to confirm the drug type.
[0218] [Split Tablet Search GUI] FIG. 29 is an example of a screen display of a split tablet search GUI. The split tablet search GUI is a GUI that is displayed in the candidate display section CDA when it is identified as a "split tablet" by the drug identifier 120 or when "split tablet" is selected from the pull-down menu of the group selection box SLB1.
[0219] The split tablet search screen SC9 includes a split tablet image display section LS1, a split tablet call button BT15, a search box SBX3, a candidate drug image display section LS2, and a group selection box SLB1.
[0220] In the split tablet image display section LS1, a list of the split tablet images selected in the overall image display section EDA and the split tablet images called by pressing the split tablet call button BT15 is displayed. For example, the split tablet image selected in the overall image display section EDA is displayed at the top of the list display in the split tablet image display section LS1. The split tablet image display section LS1 may be configured to clearly distinguish and display the display section for displaying the split tablet images called by pressing the split tablet call button BT15 and the display section for displaying the split tablet images selected in the overall image display section EDA.
[0221] When the user presses the split tablet call button BT15, one or more split tablet images can be selected from the list of previously registered split tablet captured images and added to the list display of the split tablet image display section LS1. As a method for registering split tablet images, for example, the following implementation can be considered. That is, when the drug group selected on the overall image display section EDA is in the state of "split tablet", if the selected area is long-pressed on the overall image display section EDA, a pull-down menu will be displayed, and the item "register split tablet image" can be selected from this pull-down menu. When "register split tablet image" is selected, the corresponding split tablet image is stored in the shared section, and by pressing the split tablet call button BT15 on this screen, the image can be called up. The shared section is a part of the storage area of the storage device 104 and is a storage area in which data etc. shared in the processing in the drug type identification device 100 beyond the work unit of identification is stored. The method for registering split tablet images is not limited to this example, and other methods may be used.
[0222] Search can be performed by the engraved characters by entering text in the search box SBX3 by keyboard input or voice input and pressing the search button SBT3 at the right end.
[0223] The candidate drug image display section LS2 displays the front and back of the master image in descending order of the text similarity score from the top. Similarity score determination may be performed every time one character is entered in the search box SBX3, and the candidate list display may be updated immediately. Regarding the display order of the master images of the drugs to be displayed in the candidate drug image display section LS2, past identification history data may be referred to, and drugs corresponding to split tablets with an identification history may be preferentially displayed. When the knob of the slider bar at the right end of the screen is operated downward or the screen is swiped, lower candidates can also be displayed.
[0224] Regarding the form of candidate display, it is not limited to the form of arranging them in a vertical column as shown in FIG. 27, and a multi-line and multi-column display may be used.
[0225] The user can confirm the drug type by selecting any drug from the drugs displayed on the candidate drug image display unit LS2 and pressing the "Drug Type Confirmation Button BT5" on the button display unit BDA.
[0226] Also in the candidate drug image list of FIG. 29, as the drug information of each drug, in addition to the master image (front and back), the character information indicating the drug name, and the numerical information of the size, it is preferable to display the engraved text information (master character information) registered in the engraved text database.
[0227] In addition, although not shown in FIG. 29, a color selection box may be used to specify the color of the drug so as to narrow down the candidates. In this case, the color of the divided tablets may be automatically recognized from the photographed image to narrow down the candidates in advance, and the color selection box may be displayed by default.
[0228] [Measures for improving user usability in each search GUI] In the search screens illustrated using FIGS. 20 to 29, a plurality of icons representing the shapes of typical drugs may be arranged, and when the user selects one or more icons, only the drugs of the corresponding shape are displayed on the candidate screen.
[0229] FIG. 30 shows an example of an icon representing the shape of a typical drug. The typical shapes of drugs can be circular, oval, elliptical, pentagonal, and hexagonal. As shown in FIG. 30, a configuration may be adopted in which the graphic icons corresponding to these typical shapes and the "Other" button are arranged on the search screen, and the shape is specified by selecting an icon.
[0230] Regarding the selection of color, it is not limited to the configuration of selecting from a pull-down menu, and color options may be provided by icons.
[0231] FIG. 31 shows an example of icons used for color selection. In FIG. 31, an example is shown in which icons of various colors, namely white, yellow, orange, brown, red, blue, green, and transparent, and a "Other" button are arranged from left to right. Of course, the order of color arrangement, the types of colors to be displayed as icons, and the number of colors can be appropriately designed.
[0232] A configuration may be adopted in which icons corresponding to each color and the "Other" button are arranged on the search screen, and color specification is accepted by selecting an icon. When the user selects one or more icons, only the drugs of the corresponding color may be displayed on the candidate screen.
[0233] 〔Example of Configuration for Performing Drug Type Identification Processing Specific to a Group〕 FIG. 32 is a block diagram showing an example of a functional configuration for identifying the types of drugs with difficult drug type identification in the drug type identification device 100 according to the embodiment. The drug type identification device 100 includes a drug type identification processing control unit 160 that controls the processing content based on the information estimated by the drug identifier 120, a drug type estimation result presentation processing unit 170, a capsule drug identification processing unit 172, a plain drug identification processing unit 174, and a divided tablet identification processing unit 176.
[0234] When the drug to be identified is a drug with identifiable drug type, the drug identifier 120 outputs drug type estimation information estimating the type of the drug, and when the drug to be identified is a drug with difficult drug type identification, the drug identifier 120 outputs group estimation information estimating the group to which the drug belongs. The drug type identification processing control unit 160 acquires the estimation information output from the drug identifier 120 and distributes the subsequent processing according to the estimation information.
[0235] When the drug type identification control unit 160 acquires drug type estimation information from the drug identifier 120, it causes the drug type estimation result presentation processing unit 170 to execute its processing. Based on the drug estimation information estimated by the drug identifier 120, the drug type estimation result presentation processing unit 170 performs processing to display the drug information of the candidate drugs on the candidate display unit CDA. Through the processing of the drug type estimation result presentation processing unit 170, the screen displays described in FIGS. 24 and 25 are realized. The drug type estimation result presentation processing unit 170 can cooperate with the text search unit 122 and transition to the processing of the engraved text search when the engraved text search button BT4 is pressed (see FIG. 26).
[0236] The drug type identification control unit 160 includes a group discrimination unit 162. When the estimation information output from the drug identifier 120 is group estimation information, the group discrimination unit 162 discriminates the label of the estimated group. Here, an example of discriminating which group among "capsule drugs", "plain drugs", and "divided tablets" is shown.
[0237] When the group estimation information acquired by the drug type identification control unit 160 from the drug identifier 120 indicates the group of "capsule drugs", it causes the capsule drug identification processing unit 172 to execute its processing. The capsule drug identification processing unit 172 performs processing to provide a capsule search GUI (see FIG. 27) for assisting in identifying the drug type of capsule drugs. The capsule drug identification processing unit 172 includes a capsule drug search unit 173. The capsule drug search unit 173 accepts input of search conditions, executes a search process based on the accepted search conditions, and outputs the search results.
[0238] When the group estimation information acquired by the drug type identification control unit 160 from the drug identifier 120 indicates the group of "plain drugs", it causes the plain drug identification processing unit 174 to execute its processing. The plain drug identification processing unit 174 performs processing to provide a plain drug search GUI (see FIG. 28) for assisting in identifying the drug type of plain drugs. The plain drug identification processing unit 174 includes a plain drug search unit 175. The plain drug search unit 175 accepts input of search conditions, executes a search process based on the accepted search conditions, and outputs the search results.
[0239] When the group estimation information obtained from the drug identifier 120 indicates the "divided drug" group, the drug type specific processing control unit 160 causes the divided tablet specific processing unit 176 to execute the processing. The divided tablet specific processing unit 176 performs a process of providing a divided tablet search GUI (see FIG. 29) for assisting in specifying the drug type of the divided tablet. The divided tablet specific processing unit 176 includes a divided tablet search unit 177 and a divided tablet image registration unit 178. The divided tablet search unit 177 receives an input of search conditions, executes a search process based on the received search conditions, and outputs the search results. The divided tablet image registration unit 178 is a processing unit that performs the registration process and the reading process of the divided tablet images described in FIG. 29.
[0240] When the group label is defined in a hierarchical structure, each of the capsule drug search unit 173, the plain drug search unit 175, and the divided tablet search unit 177 can narrow down the search conditions based on the estimated group label.
[0241] The process including the provision of the group-specific search GUI executed by each of the capsule drug specific processing unit 172, the plain drug specific processing unit 174, and the divided tablet specific processing unit 176 is an example of a process (group-specific process) that leads to the specification of the drug type in each group.
[0242] <Shooting assistance device> FIG. 33 is a top view of a shooting assistance device 70 for shooting a shooting image input to the drug type identification device 100. FIG. 34 is a cross-sectional view taken along line 34-34 of FIG. 33. FIG. 34 also shows a smartphone 10 for shooting an image of a drug using the shooting assistance device 70.
[0243] As shown in FIGS. 33 and 34, the shooting assistance device 70 includes a housing 72, a drug placement table 74, a main light source 75, and an auxiliary light source 78. Although the shape is described based on a square in FIG. 33, the housing 72, the drug placement table 74, and the auxiliary light source 78 may be rectangular.
[0244] The housing 72 is composed of a horizontally supported square bottom plate 72A and four rectangular side plates 72B, 72C, 72D, and 72E that are vertically fixed to the ends of each side of the bottom plate 72A respectively.
[0245] The drug placement table 74 is fixed to the upper surface of the bottom plate 72A of the housing 72. The drug placement table 74 is a member having a surface for placing drugs. Here, it is a thin plate-shaped member made of plastic or paper with a square top view, and the placement surface on which the drug to be identified is placed has a reference gray color. The reference gray color, when expressed in 256 gradation values from 0 (black) to 255 (white), is, for example, in the gradation value range of 130 to 220, and more preferably in the gradation value range of 150 to 190.
[0246] Generally, when a drug is photographed with the smartphone 10 against a white or black background, the color may be distorted due to the automatic exposure adjustment function, and sufficient engraved information may not be obtained. According to the drug placement table 74, since the placement surface is gray, the color is not distorted and the engraving details can also be captured. Also, by acquiring the gray pixel values in the photographed image and correcting them with respect to the true gray gradation values, color tone correction or exposure correction of the photographed image can be realized.
[0247] At the four corners of the placement surface of the drug placement table 74, reference markers 74A, 74B, 74C, and 74D made of black and white respectively are pasted or arranged by printing. The reference markers 74A, 74B, 74C, and 74D can be made of anything, but here, circular markers with high robustness for detection are used.
[0248] The reference markers 74A, 74B, 74C, and 74D preferably have a size of 3 to 30 mm in the vertical and horizontal directions, and more preferably a size of 5 to 15 mm.
[0249] In addition, the distance between the reference marker 74A and the reference marker 74B, and the distance between the reference marker 74A and the reference marker 74D are preferably 20 to 100 mm respectively, and more preferably 20 to 60 mm.
[0250] Furthermore, as shown in FIG. 33, a rectangular drug placement range may be defined by connecting four reference markers 74A, 74B, 74C, and 74D in a straight line. FIG. 33 shows an example where the four reference markers 74A, 74B, 74C, and 74D are arranged at the vertices of a square, but the arrangement form of the reference markers 74A, 74B, 74C, and 74D is not limited to the example in FIG. 33. For example, the four reference markers 74A, 74B, 74C, and 74D may be arranged at the vertices of a rectangle.
[0251] The drug to be photographed is placed inside a rectangular area (drug placement range) with the four reference markers 74A, 74B, 74C, and 74D as vertices. FIG. 32 shows an example where five drugs T1, T2, T3, T4, and T5 are photographed.
[0252] Note that the drug placement table 74 may be formed by the bottom plate 72A of the housing 72.
[0253] The main light source 75 and the auxiliary light source 78 constitute an illumination device used for photographing a photographed image of the drug to be identified. The main light source 75 is used to extract the engraving of the drug to be identified. The auxiliary light source 78 is used to accurately identify the color and shape of the drug to be identified. The photographing auxiliary device 70 may not be provided with the auxiliary light source 78.
[0254] FIG. 35 is a top view of the photographing auxiliary device 70 with the auxiliary light source 78 removed.
[0255] The main light source 75 is composed of a plurality of LEDs 76. Each LED 76 has a light emitting part that is a white light source with a diameter within 10 mm. Here, six LEDs 76 are arranged horizontally at a certain height on each of the four rectangular side plates 72B, 72C, 72D, and 72E. As a result, the main light source 75 irradiates illumination light on the drug to be identified from at least four directions. Note that the main light source 75 only needs to be able to irradiate illumination light on the drug to be identified from at least two directions.
[0256] The angle θ formed by the irradiation light irradiated by the LED 76 and the upper surface (horizontal plane) of the drug to be identified is preferably within the range of 10° to 20° for extracting the imprint. Note that the main light source 75 may be composed of rod-shaped light sources with a width of 10 mm or less arranged horizontally on each of the four rectangular side plates 72B, 72C, 72D, and 72E.
[0257] The main light source 75 may be constantly lit. As a result, the photographing assistance device 70 can irradiate illumination light on the drug to be identified from all directions. An image taken with all the LEDs 76 lit is called a full illumination image. According to the full illumination image, it becomes easier to extract the imprint of the drug to be identified with the imprint added.
[0258] The main light source 75 may have the LEDs 76 that are lit and extinguished according to the timing switched, or the LEDs 76 that are lit and extinguished by a switch (not shown) switched. As a result, the photographing assistance device 70 can irradiate illumination light on the drug to be identified from a plurality of different directions by a plurality of main light sources 75.
[0259] For example, an image captured with only six LEDs 76 provided on the side panel 72B lit is called a partial illumination image. Similarly, by capturing partial illumination images captured with only six LEDs 76 provided on the side panel 72C lit, partial illumination images captured with only six LEDs 76 provided on the side panel 72D lit, and partial illumination images captured with only six LEDs 76 provided on the side panel 72E lit, four partial illumination images irradiated with illumination light from different directions can be obtained. According to a plurality of partial illumination images irradiated with irradiation light from a plurality of different directions, it becomes easier to extract the imprint of the identification target drug to which the imprint is added.
[0260] The auxiliary light source 78 is a flat plate-shaped planar white light source with a square outer shape and a square opening at its center. The auxiliary light source 78 may be an achromatic reflector that diffusely reflects the irradiation light of the main light source 75. The auxiliary light source 78 is disposed between the smartphone 10 and the drug placement table 74 so that the identification target drug is uniformly irradiated with irradiation light from the imaging direction (the optical axis direction of the camera). The illuminance of the irradiation light of the auxiliary light source 78 irradiated on the identification target drug is relatively lower than the illuminance of the irradiation light of the main light source 75 irradiated on the identification target drug.
[0261] FIG. 36 is a top view of an imaging assistance device 80 according to another embodiment. FIG. 37 is a cross-sectional view taken along line 37-37 of FIG. 36. FIG. 37 also shows a smartphone 10 that captures an image of a drug using the imaging assistance device 80. For FIGS. 36 and 37, the same reference numerals are given to the parts common to FIGS. 33 and 34, and the detailed description thereof is omitted. As shown in FIGS. 36 and 37, the imaging assistance device 80 includes a housing 82, a main light source 84, and an auxiliary light source 86. The imaging assistance device 80 may not include the auxiliary light source 86.
[0262] The housing 82 has a cylindrical shape and is composed of a circular bottom plate 82A supported horizontally and a side plate 82B fixed perpendicularly to the bottom plate 82A. A drug placement table 74 is fixed to the upper surface of the bottom plate 82A.
[0263] The main light source 84 and the auxiliary light source 86 constitute an illumination device used for taking a photographed image of the drug to be identified. FIG. 38 is a top view of the photographing assistance device 80 with the auxiliary light source 86 removed.
[0264] The main light source 84 is configured by arranging 24 LEDs 85 in a ring shape at a constant height and at regular intervals in the horizontal direction on the side plate 82B. The main light source 84 may be constantly lit, or the LEDs 85 that are lit and extinguished may be switched.
[0265] The auxiliary light source 86 is a flat plate-shaped planar white light source having a circular outer shape and a circular opening at its center. The auxiliary light source 86 may be an achromatic reflector that diffusely reflects the irradiation light of the main light source 84. The illuminance of the irradiation light of the auxiliary light source 86 irradiated on the drug to be identified is relatively lower than the illuminance of the irradiation light of the main light source 84 irradiated on the drug to be identified.
[0266] The photographing assistance device 70 and the photographing assistance device 80 may be provided with a fixing mechanism (not shown) for fixing the smartphone 10 that photographs the drug to be identified at the positions of the standard photographing distance and the photographing viewpoint. The fixing mechanism may be configured to be able to change the distance between the drug to be identified and the camera according to the focal length of the photographing lens 50 of the smartphone 10.
[0267] 〔Illumination device〕 FIG. 39 is a cross-sectional view showing the configuration of an illumination device 81 as a photographing assistance device according to another embodiment. The illumination device 81 shown in FIG. 39 has a configuration in which the bottom plate 82A and the drug placement table 74 are removed from the configuration of the photographing assistance device 80 described in FIGS. 37 and 38. Other configurations may be the same as those of the photographing assistance device 80. In addition, a diffusion plate cover (not shown) that covers the side LEDs 85 may be arranged.
[0268] <Example of reference marker> Figure 40 is a top view of the medicine placement table 74 using the circular circular marker MC1. As shown in Figure 40, circular markers MC1 are arranged at the four corners of the medicine placement table 74 as reference markers 74A, 74B, 74C, and 74D. The left figure F40A in Figure 40 shows an example where the centers of the reference markers 74A, 74B, 74C, and 74D form the four vertices of a square, and the right figure F40B in Figure 40 shows an example where the centers of the reference markers 74A, 74B, 74C, and 74D form the four vertices of a rectangle. The reason for arranging four reference markers is that four-point coordinates are required to determine the perspective transformation matrix for standardization.
[0269] Here, the four reference markers 74A, 74B, 74C, and 74D are of the same size and the same color, but they may be of different sizes and colors. When the sizes are different, it is preferable that the centers of the adjacent reference markers 74A, 74B, 74C, and 74D are arranged so as to form the four vertices of a square or a rectangle. By making the sizes or colors of the reference markers different, it becomes easier to specify the shooting direction.
[0270] Also, at least four circular markers MC1 may be arranged, and five or more may be arranged. When five or more are arranged, the centers of the four circular markers MC1 form the four vertices of a square or a rectangle, and it is preferable that the centers of the additional circular markers MC1 are arranged on the sides of the square or the rectangle. By arranging five or more reference markers, it becomes easier to specify the shooting direction. Also, by arranging five or more reference markers, even if the detection of any one of the reference markers fails, the probability of simultaneously detecting the minimum four points required to obtain the perspective transformation matrix for standardization can be increased, and there are merits such as reducing the trouble of re-shooting.
[0271] The circular marker MC1 includes an outer first perfect circle with a relatively large diameter and an inner second perfect circle that is concentric with the first perfect circle and has a relatively smaller diameter than the first perfect circle. That is, the first perfect circle and the second perfect circle are circles with different radii arranged at the same center. Further, in the circular marker MC1, the inside of the second perfect circle is white, and the area inside the first perfect circle and outside the second perfect circle is filled with black.
[0272] The diameter of the first perfect circle is preferably 3 millimeters to 20 millimeters. Also, the diameter of the second perfect circle is preferably 0.5 millimeters to 5 millimeters. Further, the ratio of the diameters of the first perfect circle and the second perfect circle (diameter of the first perfect circle / diameter of the second perfect circle) is preferably 2 to 10.
[0273] The center point of the true object in the pre-normalized image and the coordinates of the center point estimated in machine learning may deviate. By providing the coordinates of the center point of the inner second perfect circle as teacher data, such as in the circular marker MC1, the estimation of the center coordinates of the true object in machine learning can be made accurate and easy. Also, due to the existence effect of the relatively large outer first perfect circle, the possibility of false detection due to dust or the like adhering to the circular marker MC1 can be significantly reduced.
[0274] FIG. 41 is a top view of the drug placement table 74 using the circular marker according to the modified example. The left figure F41A in FIG. 41 shows an example using the circular marker MC2 as the reference markers 74A, 74B, 74C, and 74D. In the circular marker MC2, inside the perfectly blackened circle, a cross-shaped figure consisting of two white lines perpendicular to each other is arranged such that the intersection of the two lines coincides with the center of the perfect circle. The reference markers 74A, 74B, 74C, and 74D of the circular marker MC2 are arranged such that their centers form the four vertices of a square, and the lines of the cross-shaped figure of the circular marker MC2 are arranged parallel to the sides of this square. The reference markers 74A, 74B, 74C, and 74D of the circular marker MC2 may be arranged such that their centers form the four vertices of a rectangle. The thickness of the lines of the cross-shaped figure of the circular marker MC2 can be determined as appropriate.
[0275] On the other hand, the right diagram F40B in FIG. 41 shows an example in which the circular marker MC3 is used as the reference markers 74A, 74B, 74C, and 74D. The circular marker MC3 includes two perfect circles with different radii arranged at the same center. The inside of the inner perfect circle is white, and the area inside the outer perfect circle and outside the inner perfect circle is filled with black. Further, inside the inner perfect circle of the circular marker MC3, a cross-shaped figure composed of two black straight lines perpendicular to each other is arranged such that the intersection of the two straight lines coincides with the center of the perfect circle. The reference markers 74A, 74B, 74C, and 74D of the circular marker MC3 are arranged such that their centers form the four vertices of a square, and the straight lines of the cross-shaped figure of the circular marker MC3 are arranged parallel to the sides of this square. The reference markers 74A, 74B, 74C, and 74D of the circular marker MC3 may be arranged such that their centers form the four vertices of a rectangle. The thickness of the lines of the cross-shaped figure of the circular marker MC3 can be determined as appropriate.
[0276] According to the circular markers MC2 and MC3, the estimation accuracy of the center point coordinates can be improved. Also, according to the circular markers MC2 and MC3, since they look different from the drug, it becomes easier to recognize the markers.
[0277] FIG. 42 is a diagram showing a specific example of a reference marker having a quadrangular outer shape. The left diagram F42A in FIG. 41 shows the quadrangular marker MS1. The quadrangular marker MS1 includes an outer square SQ1 with a relatively large side length and an inner square SQ2 arranged concentrically with the square SQ1 and having a relatively smaller side length than the square SQ1. That is, the squares SQ1 and SQ2 are quadrilaterals with different side lengths arranged at the same center (center of gravity). Also, the inside of the square SQ2 of the quadrangular marker MS1 is white, and the area inside the square SQ1 and outside the square SQ2 is filled with black.
[0278] The length of one side of the square SQ1 is preferably from 3 millimeters to 20 millimeters. Also, the length of one side of the square SQ2 is preferably from 0.5 millimeters to 5 millimeters. Further, the ratio of the length of one side of the square SQ1 to the length of one side of the square SQ2 (length of one side of the square SQ1 / length of one side of the square SQ2) is preferably from 2 to 10. For the purpose of further enhancing the estimation accuracy of the center coordinates, a black rectangle (for example, a square) not shown, which is relatively smaller in side length than the square SQ2, may be concentrically arranged inside the square SQ2.
[0279] The right diagram F42B of FIG. 42 shows a top view of the medicine placement table 74 using the quadrilateral marker MS1 in the shape of a quadrilateral. As shown in the right diagram F41B, on the medicine placement table 74, quadrilateral markers MS1 in the shape of a quadrilateral are arranged at the four corners as reference markers 74A, 74B, 74C, and 74D. Here, an example is shown where the lines connecting the centers of adjacent reference markers 74A, 74B, 74C, and 74D form a square, but the lines connecting the centers of adjacent reference markers 74A, 74B, 74C, and 74D may form a rectangle.
[0280] FIG. 43 is a top view of the medicine placement table 74 using a circular marker according to another modified example. The left diagram F43A of FIG. 42 is a diagram showing an example using circular markers MC4 as reference markers 74A, 74B, 74C, and 74D. The circular marker MC4 includes an outer first perfect circle with a relatively large diameter, a second perfect circle arranged concentrically with the first perfect circle and having a relatively smaller diameter than the first perfect circle, and a black circle (third perfect circle) arranged concentrically with the second perfect circle inside the second perfect circle and having a relatively smaller diameter than the second perfect circle. The circular marker MC4 has a white inner side of the second perfect circle, and the region inside the first perfect circle and outside the second perfect circle is filled with black.
[0281] The diameter of the first true circle is preferably 3 millimeters to 20 millimeters. Also, the diameter of the second true circle is preferably 5 millimeters to 18 millimeters. The diameter of the third true circle is preferably 0.5 millimeters to 5 millimeters. Furthermore, the ratio of the diameter of the first true circle to the diameter of the second true circle (diameter of the first true circle / diameter of the second true circle) is preferably 1.1 to 3. Furthermore, the ratio of the diameter of the second true circle to the diameter of the third true circle (diameter of the second true circle / diameter of the third true circle) is preferably 2 to 10.
[0282] The left figure F43A of FIG. 43 shows an example in which the centers of the reference markers 74A, 74B, 74C, and 74D form the four vertices of a square, and the right figure F43B of FIG. 43 shows an example in which the centers of the reference markers 74A, 74B, 74C, and 74D form the four vertices of a rectangle. A straight line connecting the four circular markers MC4 in the vertical and horizontal directions is displayed. According to the circular marker MC4, the estimation accuracy of the center point coordinates can be improved. Also, by connecting the circular markers MC4 with a straight line, it becomes easier to recognize the marker and the drug placement range, and it is also possible to correct the distortion of the photographed image with the straight line appearing in the photographed image, specify the cutout range of the standardized image, etc. Note that these straight lines may not be present.
[0283] FIG. 44 is a diagram showing another specific example of a quadrilateral reference marker. The left diagram F44A in FIG. 44 shows a quadrilateral marker MS2. The quadrilateral marker MS2 includes an outer square SQ1 with a relatively large side length, an inner square SQ2 arranged concentrically with the square SQ1 and having a relatively smaller side length than the square SQ1, and a black square SQ3 arranged concentrically with the square SQ2 inside the square SQ2 and having a relatively smaller side length than the square SQ2. The side length of the square SQ1 is preferably 3 millimeters to 20 millimeters. Also, the side length of the square SQ2 is preferably 5 millimeters to 18 millimeters. The side length of the square SQ3 is preferably 0.5 millimeters to 5 millimeters. Further, the ratio of the side lengths of the square SQ1 and the square SQ2 (side length of the square SQ1 / side length of the square SQ2) is preferably 1.1 to 3. The ratio of the side lengths of the square SQ2 and the square SQ3 (side length of the square SQ2 / side length of the square SQ3) is preferably 2 to 10.
[0284] The right diagram F44B in FIG. 44 shows a top view of a drug placement table 74 using the quadrilateral marker MS2 of the quadrilateral. As shown in the right diagram F44B, on the drug placement table 74, quadrilateral markers MS2 of the quadrilateral are arranged at the four corners as reference markers 74A, 74B, 74C, and 74D. Here, an example is shown in which the lines connecting the centers of adjacent reference markers 74A, 74B, 74C, and 74D form a square, but the lines connecting the centers of adjacent reference markers 74A, 74B, 74C, and 74D may form a rectangle. Also, similar to FIG. 43, the straight lines connecting the four quadrilateral markers MS2 in the vertical and horizontal directions are shown.
[0285] According to the quadrilateral marker MS2, the estimation accuracy of the center point coordinates can be improved. Also, by connecting the quadrilateral markers MS2 with straight lines, it becomes easier to recognize the marker and the drug placement range. Note that these straight lines may not be present.
[0286] The drug placement table 74 may have a mixture of circular markers MC and rectangular markers MS. By mixing the circular markers MC and the rectangular markers MS, effects such as facilitating the identification of the shooting direction can be expected.
[0287] Furthermore, it is preferable to adopt circular markers rather than rectangular markers. This is because the following constraints (a) to (c) exist when detecting reference markers with a mobile terminal device such as a smartphone. (a) In some cases, there is a capacity limit for the app on the mobile terminal device, so it is desirable to perform marker detection and drug detection using the same pre-trained model. (b) In drug detection, it is desirable to use a bounding box with rotation to prevent a part of another drug from entering the cut-out image of the elliptical tablet. In this case, due to requirement (a), it is necessary to prepare reasonable teacher data regarding the rotation angle for marker detection as well. (c) In the case of rectangular markers, there is an arbitrariness in determining the rotation angle of the bounding box due to four-fold symmetry, making it difficult to create reasonable teacher data for the rotation angle of the rectangular marker in the pre-normalization image, which is the input image during marker detection. On the other hand, for circular markers, it is possible to create reasonable teacher data with the rotation angle always set to 0 degrees.
[0288] Also, by making the circular markers concentric, the estimation accuracy of the center coordinates of the markers can be improved. In the pre-normalization image, a simple circular marker is distorted in shape and prone to errors in center coordinate estimation. However, since the inner circle of the concentric circles has a smaller range, even in a distorted pre-normalization image, the pre-trained model can easily identify the center coordinates. Furthermore, the outer circle of the concentric circles has advantages such as a larger structure that is easy for the pre-trained model to find and being robust against noise and dirt. Note that for rectangular markers as well, making them concentric can improve the estimation accuracy of the center point coordinates of the markers.
[0289] <Other aspects of the drug placement table> The placement surface of the drug placement table of the shooting assistance device may be provided with a recessed structure for placing drugs. The recessed structure includes recesses, grooves, indentations, and holes.
[0290] FIG. 45 is a view showing a medicine placement table 410 that is used in place of or in addition to the medicine placement table 74 (see FIG. 33) and has a recessed structure. The medicine placement table 410 is made of paper, synthetic resin, fiber, rubber, or glass. The left view F45A of FIG. 45 is a top view of the medicine placement table 410. As shown in the left view F45A, the placement surface of the medicine placement table 410 has a gray color, and reference markers 74A, 74B, 74C, and 74D are arranged at the four corners. In FIG. 45, the shape of the marker is shown using a circular marker MC1, but it is not limited to the circular marker MC1 and may be a circular marker MC2, MC3, MC4, a square marker MS1, or MS2.
[0291] Furthermore, a total of nine recesses 410A, 410B, 410C, 410D, 410E, 410F, 410G, 410H, and 410I in a 3-row × 3-column arrangement are provided as the recessed structure on the placement surface of the medicine placement table 410. The recesses 410A to 410I are each circular and of the same size in a top view.
[0292] The right view F45B of FIG. 45 is a cross-sectional view taken along line 44-44 of the medicine placement table 410. As shown in the right view F45B, the recesses 410A, 410B, and 410C have hemispherical bottoms and are of the same depth. The same applies to the recesses 410D to 410I. The bottom does not have to be a complete hemisphere and may be a concave curved surface with a certain radius of curvature or a concave curved surface with a gradually changing radius of curvature.
[0293] In addition, the right figure F45B shows tablets T51, T52, and T53 placed in depressions 410A, 410B, and 410C, respectively. Tablets T51 and T52 are circular in top view and rectangular in side view. Also, tablet T53 is circular in top view and elliptical in side view. In top view, tablets T51 and T53 are the same size, and tablet T52 is relatively smaller than tablets T51 and T53. As shown in the right figure F45B, tablets T51 - T53 are statically placed by fitting into depressions 410A - 410C. Tablets T51 - T53 only need to be circular in top view, and in side view, the left and right sides may be linear and the top and bottom may be arc-shaped.
[0294] Thus, according to the drug placement table 410, by having a hemispherical depression structure on the placement surface, circular drugs in top view can be prevented from moving and can be statically placed. Also, since the position of the drug during photography can be determined to be the position of the depression structure, detection of the drug area becomes easier.
[0295] The drug placement table may have a depression structure for rolling-type drugs such as capsule drugs. FIG. 46 is a diagram showing a drug placement table 412 provided with a depression structure for capsule drugs. Parts common to FIG. 45 are denoted by the same reference numerals, and detailed descriptions thereof are omitted.
[0296] The left figure F46A in FIG. 46 is a top view of the drug placement table 412. As shown in the left figure F46A, six depressions 412A, 412B, 412C, 412D, 412E, and 412F in a 3 - row × 2 - column arrangement are provided as the depression structure on the placement surface of the drug placement table 412. Depressions 412A - 412F are rectangles of the same size in top view. The depressions do not necessarily have to be completely hemispherical and may be concave curved surfaces with a certain radius of curvature or concave curved surfaces with a gradually changing radius of curvature.
[0297] The right figure F46B of FIG. 46 is a cross-sectional view taken along line 46-46 of the drug placement table 412. As shown in the right figure F46B, the depressions 412A, 412B, and 412C have semi-cylindrical bottoms and are of the same depth. The same applies to the depressions 412D to 412F. Also, F63B shows the capsule drugs CP1, CP2, and CP3 placed in the depressions 412A, 412B, and 412C respectively. The capsule drugs CP1 to CP3 are cylindrical with hemispherical ends (both bottom surfaces) and have different diameters. As shown in the right figure F46B, the capsule drugs CP1 to CP3 are statically placed by fitting into the depressions 412A to 412C.
[0298] Thus, according to the drug placement table 412, by providing a semi-cylindrical depression structure on the placement surface, it is possible to prevent the cylindrical capsule drugs from moving or rolling and to keep them static. Also, since the position of the drug during imaging can be determined to be the position of the depression structure, it becomes easier to detect the area of the drug.
[0299] Also, the drug placement table may have a depression structure for oval tablets. FIG. 47 is a view showing the drug placement table 414 provided with a depression structure for oval tablets. Note that the same reference numerals are given to the parts common to FIG. 46, and detailed descriptions thereof are omitted.
[0300] The left figure F47A of FIG. 47 is a top view of the drug placement table 414. As shown in the left figure F47A, six depressions 414A, 414B, 414C, 414D, 414E, and 414F in a 3-row × 2-column arrangement are provided as a depression structure on the placement surface of the drug placement table 414. The depressions 414A to 414F are rectangular in top view.
[0301] The depressions 414A and 414B are of the same size. The depressions 414C and 414D are of the same size and are relatively smaller than the depressions 414A and 414B. The depressions 414E and 414F are of the same size and are relatively smaller than the depressions 414C and 414D.
[0302] Also, the left diagram F47A shows tablets T61, T62, and T63 placed in depressions 414B, 414D, and 414F respectively. As shown in the left diagram F47A, the depressions 414B, 414D, and 414F are sized to correspond to tablets T61, T62, and T63 respectively.
[0303] The right diagram F47B in Figure 47 is a cross-sectional view taken along line 47-47 of the drug placement table 414. As shown in the right diagram F47B, the depressions 414A and 414B have flat bottoms. As shown in the right diagram F47B, the tablet T61 is placed stationary by fitting into the depression 414B. The same applies to tablets T62 and T63.
[0304] Thus, according to the drug placement table 414, by providing a rectangular parallelepiped depression structure on the placement surface, the elliptical tablets can be prevented from moving and can be placed stationary. Also, since the position of the drug during imaging can be determined to be the position of the depression structure, the detection of the drug area becomes easier.
[0305] Note that the shape, number, and arrangement of the depression structure are not limited to the embodiments shown in Figures 45 to 47, and may be combined as appropriate, or enlarged or reduced.
[0306] <Regarding the hardware configuration of each processing unit and control unit> The hardware structure of the processing unit that executes various processes such as the image acquisition unit 112, drug detector 114, region correction unit 116, drug region extraction unit 118, drug identifier 120, text search unit 122, display control unit 124, magnification change unit 125, input processing unit 126, stamp extraction unit 140, outer shape extraction unit 142, size measurement unit 144, loss calculation unit 152, optimizer 154, drug type specific processing control unit 160, group discrimination unit 162, drug type estimation result presentation processing unit 170, capsule drug specific processing unit 172, capsule drug search unit 173, plain drug specific processing unit 174, plain drug search unit 175, divided tablet specific processing unit 176, divided tablet search unit 177, and divided tablet image registration unit 178 described with reference to FIG. 5 and FIG. 17 and FIG. 32 is various processors as shown below.
[0307] The various processors include a CPU (Central Processing Unit), which is a general-purpose processor that executes programs and functions as various processing units, a GPU (Graphics Processing Unit), which is a processor specialized for image processing, a programmable logic device (PLD), such as an FPGA (Field Programmable Gate Array), which is a processor whose circuit configuration can be changed after manufacturing, and an application-specific electric circuit, which is a processor having a circuit configuration specifically designed to execute specific processes, such as an ASIC (Application Specific Integrated Circuit).
[0308] One processing unit may be composed of one of these various processors, or may be composed of two or more processors of the same or different types. For example, one processing unit may be composed of a plurality of FPGAs, or a combination of a CPU and an FPGA, or a combination of a CPU and a GPU, etc. Also, a plurality of processing units may be composed of one processor. As an example of composing a plurality of processing units with one processor, first, as represented by a computer such as a client or a server, one processor is composed of a combination of one or more CPUs and software, and this processor functions as a plurality of processing units. Second, as represented by a System On Chip (SoC), etc., there is a form in which a processor that realizes the functions of the entire system including a plurality of processing units is used in one IC (Integrated Circuit) chip. Thus, various processing units are configured as a hardware structure using one or more of the above various processors.
[0309] Furthermore, the hardware structure of these various processors is, more specifically, an electrical circuit (circuitry) that combines circuit elements such as semiconductor elements.
[0310] <Regarding the program for realizing the functions of the drug type identification device 100> The processing functions of the drug type identification device 100 are not limited to the smartphone 10, and can be realized using various forms of information processing devices such as tablet computers, personal computers, or workstations. A program for causing a computer to realize some or all of the processing functions of the drug type identification device 100 described in the above embodiment can be recorded on a computer-readable medium, which is a non-transitory tangible information storage medium such as an optical disk, a magnetic disk, or a semiconductor memory, and the program can be provided through this information storage medium. Also, instead of storing and providing the program in such a non-transitory tangible information storage medium, it is also possible to provide the program signal as a download service using an electric communication line such as the Internet.
[0311] In addition, it is also possible to provide part of the processing functions of the drug type identification device 100 described in the above embodiment as an application server and perform a service that provides the processing functions through a telecommunication line.
[0312] <Effect of the embodiment> According to the drug type identification device 100 according to the embodiment, even for an image taken in a state where drug type identifiable drugs and drugs with difficult drug type identification are mixed, the type of each drug in the image or the group to which the drug belongs can be automatically identified, and for drugs with difficult drug type identification, it is possible to shift to processing specific to each group and provide a GUI that supports identification of the type. For this reason, it becomes possible to take a picture of a plurality of drugs in a meaningful unit from a pharmaceutical point of view and identify individual drugs.
[0313] [Modification 1] The drug identifier 120 may be configured to include an image recognizer using a template matching method.
[0314] [Modification 2] In the above embodiment, an example of the case of identifying drugs has been described, but the technology of the present disclosure can also be applied to the case of performing a drug audit. Further, in the above embodiment, an example of the case where the object to be identified is a drug has been described, but the object to be identified is not limited to drugs. The technology of the present disclosure can be applied as a technology for identifying an object from an image for various objects regardless of the type and use of the object.
[0315] <Explanation of terms> In the present disclosure, the “object” refers to an object that can be conceived of as being classified by a hierarchical structure. The specific depth of the hierarchical structure of the classification is referred to as the granularity, and at this granularity, the object belongs to one of the “types”. “Identification” means determining to which type the object belongs at a certain specific granularity.
[0316] The hierarchical structure may be defined according to the purpose of identification.
[0317] The above-mentioned "certain particle size" shall be determined according to the purpose of identification.
[0318] A set of objects having common characteristics is called a group. The group may or may not have a type and a relationship. The group may refer to the upper hierarchy of the type in the above hierarchical structure, or may be a newly defined (independent of the above hierarchical structure) set for convenience. The group may be defined according to the purpose of identification.
[0319] 〔Specific Example〕 Drugs can be conceived of as being classified in a hierarchical structure based on pharmaceutical efficacy classification or a hierarchical structure based on external characteristics. Drug identification can be defined as an act of determining which pharmaceutical product of a certain YJ code the target drug is (this is an example of the definition of drug identification, and for example, the identification code may be defined using a type other than the YJ code). Each group of "capsule drugs" and "plain tablets" described in the above embodiments is a set of newly defined types for convenience. Also, the group of "half tablets" is a set of newly defined drugs for convenience. Although it is a set independent of the type of YJ code, the purpose of drug identification is ultimately to specify the YJ code of the half tablets.
[0320] 《Others》 The technical scope of the present invention is not limited to the scope described in the above embodiments and modifications. The configurations and the like in the embodiments and modifications can be changed without departing from the gist of the present invention, and can be appropriately combined between the embodiments and modifications.
Explanation of Reference Numerals
[0321] 10 Smartphone 12 Housing 14 Touch Panel Display 14A Display Unit 14B Input Unit 16 Speaker 18 Microphone 20 In-Camera 22 Out-Camera 24 Light 26 switches 28 CPUs 30 wireless communication unit 32 call unit 34 memory 36 internal memory unit 38 external memory unit 40 external input / output unit 42 GPS receiver 44 power supply unit 50 photographing lens 50F focus lens 50Z zoom lens 54 imaging device 58 A / D converter 60 lens drive unit 70 photographing assist device 72 housing 72A bottom panel 72B side panel 72C side panel 72D side panel 72E side panel 74 chemical placement table 74A, 74B, 74C, 74D reference markers 75 main light source 78 auxiliary light source 80 photographing assist device 81 lighting device 82 housing 82A bottom panel 82B side panel 84 main light source 86 auxiliary light source 100 chemical type identification device 102 processor 104 storage device 112 image acquisition unit 114 chemical detector 116 area correction unit 118 chemical area cutting unit 120 chemical identifier 122 text search unit 124 display control unit 125 magnification change unit 126 input processing unit 130 database 131 Master image database 132 Identification result memory unit 140 Engraving extraction unit 142 Outer shape extraction unit 144 Size measurement unit 150 Machine learning system 151 Learning model 152 Loss calculation unit 154 Optimizer 160 Drug type identification processing control unit 162 Group discrimination unit 170 Drug type estimation result presentation processing unit 172 Capsule drug identification processing unit 173 Capsule drug search unit 174 Plain drug identification processing unit 175 Plain drug search unit 176 Divided tablet identification processing unit 177 Divided tablet search unit 178 Divided tablet image registration unit 410 Drug placement table 410A, 410B, 410C, 410D depressions 410E, 410F, 410G, 410H, 410I depressions 412 Drug placement table 412A, 412B, 412C, 412D depressions 412E, 412F depressions 414 Drug placement table 414A, 414B, 414C, 414D depressions 414E, 414F depressions Abs area BDA button display section BT2 Area addition button BT3 Area deletion button BT4 Engraving text search button BT5 Drug type confirmation button BT6 Drug type confirmation hold button BT7 Reshooting button BT8 Completion button BT9 Right arrow button BT10 Left arrow button Arrow button BT11 Search button BT12 Capsule list button BT13 Plain drug list button BT14 Divided tablet call button BT15 Frames BX, BX1, BX2, BX3, BX4 CheckBox CB Candidate display section CDA Capsule drugs CP1, CP2, CP3 Drugs DR1, DR2, DR3, DRj Overall image display section EDA Engraving extraction images Egm, Egm1, Egm2, Egm3, Egm4, Egm5 Left figure F40A Right figure F40B Left figure F41A Right figure F41B Left figure F42A Right figure F42B Left figure F43A Right figure F43B Left figure F44A Right figure F44B Left figure F45A Right figure F45B Left figure F46A Right figure F46B Left figure F47A Right figure F47B Correct data GTj Input box IB1 Input box IB2 Engraving extraction images IM1j Outer shape images IM2j Drug images IMj Divided tablet image display section LS1 Candidate drug image display section LS2 Circular markers MC, MC1, MC2, MC3, MC4 Magnification information MGj Rectangular markers MS, MS1, MS2 Original images Org, Org1, Org2, Org3, Org4, Org5 External views of Otw, Otw1, Otw2, Otw3, Otw4, and Otw5 PDM pull-down menu PRj inference result Search buttons SBT1, SBT2, and SBT3 Search box SBX1 Text search box SBX2 Search box SBX3 Screens SC1, SC2, SC3, SC4, SC5, SC6, SC7, SC8, and SC9 Selection boxes SLB1, SLB2, SLB3, SLB4, and SLB5 Squares SQ1, SQ2, and SQ3 Scroll bar SRB Capsule image display section SRD Plain drug image display section SRD2 GPS satellites ST1 GPS satellites ST2 Region editing switch SW1 Size search box SZB Size information SZj Drugs T1, T2, T3, T4, and T5 Tablets T51, T52, and T53 Tablets T61, T62, and T63 First learned model TM1 Second learned model TM2 Steps S1 to S5 of the learning method for constructing the drug type identification device Steps S11 to S16 of the drug type identification method Steps S21 to S29 of the drug type identification method
Claims
1. A detector that detects the objects from an image in which a plurality of objects are photographed, in units of the objects; Among the objects detected by the detector, for the type-identifiable objects for which the type of the object can be specified up to the type of the object from the image, the type of the object is estimated from the image, and for the type-difficult-to-identify objects for which it is difficult to specify up to the type of the object from the image but the group to which the object belongs can be specified, an identifier that estimates the group from the image; A processing unit that performs group-specific processing leading to the specification of the type for the object estimated as the group by the identifier; An object identification device comprising the above.
2. The image is an image photographed in a state where the type-identifiable objects and the type-difficult-to-identify objects are mixed; The object identification device according to claim 1.
3. The type-difficult-to-identify objects are classified into a plurality of the groups; Group-specific processing is defined for each of the groups; The object identification device according to claim 1 or 2.
4. The detector includes a first trained model trained by machine learning using first training data in which the objects are labeled in units of the objects without distinguishing between the type-identifiable objects and the type-difficult-to-identify objects; The object identification device according to claim 1 or 2.
5. The identifier includes a second trained model trained by machine learning using second training data in which the type-identifiable objects are labeled in units of the type of the object and the type-difficult-to-identify objects are labeled in units of the group to which the object belongs; The object identification device according to claim 1 or 2.
6. The label for identifying the group is defined in a hierarchical structure; The object identification device according to claim 1 or 2.
7. As an input to the identifier, at least one of an object image obtained by cutting out the region of the object detected by the detector from the image in units of the object, a character / symbol extraction image including at least one of characters and symbols extracted from the object image, an outer shape image of the object, and size information of the object is used; The object identification device according to claim 1 or 2.
8. As an input to the identifier, further, magnification information indicating the magnification or reduction ratio of the object image is used; The object identification device according to claim 7.
9. The processing specific to the group includes processing for displaying a screen that receives an input of a search condition for searching for the type of the object within the estimated group. The object identification device according to claim 1 or 2.
10. The object is a drug, The type-specific difficult object includes at least one of a capsule drug, a plain drug, and a divided tablet, The type-specific possible object includes a tablet having an engraving or printing, The object identification device according to claim 1 or 2.
11. As an input to the identifier, at least one of a drug image obtained by cutting out the area of the drug detected by the detector from the image in drug units, a character and symbol extraction image including at least one of characters and symbols extracted from the drug image, an outer shape image of the drug, and size information of the drug is used. The object identification device according to claim 10.
12. An object identification device comprising one or more processors and one or more memories in which a program executed by the one or more processors is stored, The one or more processors, A detection process for detecting the object in object units from an image in which a plurality of objects are photographed, Among the objects detected by the detection process, for the type-specific possible objects for which the type of the object can be specified from the image, the type of the object is estimated from the image, and for the type-specific difficult objects for which it is difficult to specify the type of the object from the image but the group to which the object belongs can be estimated, the group is estimated from the image. An identification process, A process of shifting to a group-specific process that leads to the specification of the type for the object estimated as the group by the identification process. An object identification device that executes.
13. The one or more processors, The detection process is executed using a detector including a first trained model trained by machine learning using first training data in which the type-specific possible objects and the type-specific difficult objects are labeled in object units without distinction. The object identification device according to claim 12.
14. The one or more processors, The identification process is executed using an identifier including a second trained model trained by machine learning using second training data in which the type-specific possible objects are labeled in object type units and the type-specific difficult objects are labeled in group units to which the objects belong. The object identification device according to claim 12 or 13.
15. The one or more processors: execute a process of cutting out a region of the object detected by the detection process from the image and generating an object image in units of the object; perform the identification process based on the object image; The object identification device according to claim 12 or 13.
16. The object is a drug, the difficult-to-identify objects include at least one of capsule drugs, plain drugs, and divided tablets, the identifiable objects include tablets having markings or printing, the group-specific process includes a process of displaying a screen for receiving an input of search conditions for searching for the type of the drug within the estimated group; The object identification device according to claim 12 or 13.
17. a first database in which character symbol information including at least one of characters and symbols indicated by markings or printing attached to the drug is associated with the type of the drug; a second database in which master images of the drug are stored; comprising The one or more processors: search at least one of the first database and the second database based on the received search conditions, and output candidates for the drug that meet the search conditions; The object identification device according to claim 16.
18. The one or more processors: perform a process of displaying a screen including a captured image display unit that displays the image and a candidate display unit that displays information on candidate objects based on an estimation result of the identification process; The object identification device according to claim 12 or 13.
19. The one or more processors: perform a process of displaying information on the group to which the candidate object to be displayed on the candidate display unit belongs; The object identification device according to claim 18.
20. The one or more processors: receive an instruction for designating the group to which the candidate object to be displayed on the candidate display unit belongs, and control the display on the candidate display unit according to the received instruction; The object identification device according to claim 19.
21. A camera and a display that displays the image captured by the camera and information on the object estimated from the image; The object identification device according to claim 1 or 12, comprising.
22. An object identification method executed by one or more processors, wherein the one or more processors: The one or more processors: detecting the objects from an image in which a plurality of objects are photographed, in units of the objects; for the type-identifiable objects among the detected objects, for which the type of the object can be estimated from the image, estimating the type of the object from the image, and for the type-difficult objects for which it is difficult to identify the type of the object from the image but for which the group to which the object belongs can be identified, estimating the group from the image; performing group-specific processing leading to identification of the type, on the objects estimated as the group; An object identification method, comprising the above.
23. A program causing a computer to have a function of detecting the objects from an image in which a plurality of objects are photographed, in units of the objects; have a function of estimating the type of the object from the image, for the type-identifiable objects among the detected objects, for which the type of the object can be identified from the image, and estimating the group from the image, for the type-difficult objects for which it is difficult to identify the type of the object from the image but for which the group to which the object belongs can be identified; have a function of performing group-specific processing leading to identification of the type, on the objects estimated as the group; A program for realizing the above.
24. A non-transitory and computer-readable recording medium, on which the program according to Claim 23 is recorded.
Citation Information
Patent Citations
Tablet detection method, tablet detection device, and table detection program
JP2018027242A
Computer program, information processing device, information processing method and generation method of learned model
JP2021026450A
Food preparation method and system based on ingredient identification
JP2021508811A
Drug discrimination software
JP2022010060A
Method and Device for Identification and / or Sorting of Medicines
US20150154750A1