Target identification device and target identification method
The target identification device addresses the challenge of identifying moving objects with unclear labels by combining visual data processing with machine learning, ensuring accurate and efficient identification through automated training and classification.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- IBARAKI PREFECTURE
- Filing Date
- 2024-10-03
- Publication Date
- 2026-04-15
AI Technical Summary
Existing methods for identifying individuals using visually recognizable identification information on labels face challenges when the target is a moving object, as it is difficult to capture clear images for accurate identification.
A target identification device that utilizes an identification information acquisition unit, a data generation unit, an appearance analysis unit, and a comprehensive determination unit to process visual data, enabling accurate identification even when the label is not clearly visible, by combining recognition results with machine learning-based classification.
The device achieves high-precision identification of moving objects by automatically generating training datasets and using machine learning, allowing for both speed and accuracy in identifying individuals, even when labels are unclear or missing.
Smart Images

Figure 2026065292000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an object identification device and an object identification method.
Background Art
[0002] A method is known in which an identification label with identification information recorded thereon is attached to a target individual, and the identification information is recognized by a sensor or the like to discriminate the target individual.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] Conventionally, a method has been known in which identification information is described on a label in a visually recognizable state, and the label is attached to a target individual for use in individual identification. The label used in this method has a low manufacturing cost and is easy to attach to a target individual. Also, with regard to identification, it can be read visually or by a general-purpose device such as a barcode reader, and has been widely spread in various fields.
[0005] When automating the identification of a target individual using such a label, for example, a method of acquiring and image-analyzing an image in which the label attached to the target individual appears to read the identification information can be considered. In this case, the "label in a state where the identification information can be read" must appear in the input image. However, when the target individual is a moving object (for example, a human, livestock, and a vehicle, etc.), it has sometimes been difficult to prepare an image that satisfies the requirements.
[0006] The present invention solves at least one of the above-described problems of the object identification method using a label on which visually recognizable identification information is described. [Means for solving the problem]
[0007] The first target identification device of the present invention is a target identification device that identifies target individuals based on visual data in which target individuals equipped with a sign bearing identification information are recorded, and comprises: an identification information acquisition unit that recognizes the identification information from the sign recorded in the visual data; a data generation unit that associates and stores the recognition results of the corresponding identification information by the identification information acquisition unit as labels for the target individuals recorded in the visual data and generates a training dataset; an appearance analysis unit that includes a classifier obtained by machine learning using the training dataset; and an overall determination unit that combines the recognition results of the identification information by the identification information acquisition unit and the determination results by the appearance analysis unit to generate a final result for identifying the target individuals, wherein when new visual data in which target individuals are recorded is provided, the overall determination unit generates the final result based on the determination results of the appearance analysis unit if the identification information acquisition unit cannot recognize the identification information based on the new visual data. [Effects of the Invention]
[0008] According to the present invention, at least one of the above-mentioned problems associated with methods for identifying target individuals using markers containing visually recognizable identification information is solved. [Brief explanation of the drawing]
[0009] [Figure 1] This is a hardware configuration diagram of a first embodiment of the object identification device of the present invention. [Figure 2] This is a functional block diagram of the target identification device. [Figure 3] This diagram illustrates the data generation flow by the visual data acquisition unit, sign detection unit, sign information extraction unit, and data generation unit. [Figure 4] This is the structure of the training dataset. [Figure 5] This is a flowchart (learning flow) illustrating the procedure for generating a classifier using a target identification device. [Figure 6] This is a perspective view of the cowshed. [Figure 7] This diagram illustrates the identification flow when a new image is provided to an object identification device that includes an appearance analysis unit containing a trained classifier. [Figure 8] This diagram illustrates the identification flow when a new image is provided to an object identification device that includes an appearance analysis unit containing a trained classifier. [Figure 9] This is an explanatory diagram illustrating a specific example of a method for summarizing recognition results and classification results. [Figure 10] Figure 10(A) is a diagram showing the entrance gate to a parking facility. Figure 10(B) is an example of a planar image (cropped) generated by the sign detection unit based on an image of a vehicle captured by the imaging device. Figure 10(C) is an example of a cropped image generated by the sign detection unit based on an image of a vehicle's license plate captured by the imaging device. [Figure 11] This is a floor plan of the parking space in the parking facility. [Figure 12] This is a flowchart for list generation. [Figure 13] Figure 13(A) is a diagram showing the entrance gate to a parking facility. Figure 13(B) is an example of a planar image (cropped) generated by the sign detection unit based on an image of a vehicle captured by the imaging device. Figure 13(C) is an example of a cropped image generated by the sign detection unit based on an image of a vehicle's license plate captured by the imaging device. [Figure 14] Figure 14(A) is a diagram showing the entrance gate to a parking facility. Figure 14(B) is an example of a planar image (cropped) generated by the sign detection unit based on an image of a vehicle captured by the imaging device. Figure 14(C) is an example of a cropped image generated by the small detection unit based on an image of a vehicle's license plate captured by the imaging device. [Figure 15] This diagram shows the movement trajectory of a specific player on the soccer field during a match. [Figure 16] This is a diagram showing the playing field at kickoff. [Figure 17] It is a diagram showing a modified example of the learning flow.
Embodiments for Carrying Out the Invention
[0010] The first object identification device of the present invention is an object identification device that identifies an object individual provided with a label on which identification information is described based on visual data in which the object individual is recorded, and includes an identification information acquisition unit that recognizes the identification information from the label recorded in the visual data, a data generation unit that associates and stores the recognition result of the corresponding identification information by the identification information acquisition unit with respect to the object individual recorded in the visual data as a label to generate a training data set, an appearance analysis unit including a classifier obtained by machine learning using the training data set, and a comprehensive determination unit that includes the recognition result of the identification information by the identification information acquisition unit. When new visual data in which the object individual is recorded is given, the comprehensive determination unit is an object identification device that summarizes the recognition result of the identification information by the identification information acquisition unit and the discrimination result by the appearance analysis unit to generate a final result of identifying the object individual.
[0011] According to the first object identification device, since the collection, arrangement, generation of the training data set, and machine learning of the classifier of the visual data for the training data set can be automatically performed, no specialized knowledge is required for the operator, so application to various fields is easy. Although machine learning can also be performed under the operation of the operator, even in that case, since the collection, arrangement, and generation of the training data set of the visual data for the training data set can be automatically performed, application to various fields is easy. In addition, since the recognition of the object individual based on the visual data is performed by summarizing the recognition result and the discrimination result, more accurate recognition is possible. For example, the recognition result can be generated even when either the recognition result or the discrimination result cannot be obtained.
[0012] The second object identification device of the present invention is in the first object identification device when the identification information acquisition unit cannot recognize the identification information based on the new visual data. The comprehensive determination unit is an object identification device that generates the final result based on the discrimination result of the appearance analysis unit.
[0013] According to the second object identification device, even from visual data in which no label is recorded or in which the identification information cannot be recognized even if it is recorded, high-precision identification of the target individual is realized by the functions of the appearance analysis unit and the comprehensive determination unit.
[0014] In the third object identification device of the present invention, in the first or second object identification device, when the identification information acquisition unit can recognize the identification information based on the new visual data, the comprehensive determination unit determines whether to generate the final result without using the discrimination result based on the reliability of the recognition result of the identification information acquisition unit. It is an object identification device.
[0015] According to the third object identification device, if the reliability of the recognition result meets a predetermined standard, in order to generate the final result mainly based on the recognition result without using the discrimination result, faster identification of the target individual is realized. On the other hand, if the reliability does not meet the standard, more accurate identification of the target individual is realized in order to generate the final result using the discrimination result (or the discrimination result as well). By properly using or combining the identification information acquisition unit and the appearance analysis unit, both speed and accuracy can be achieved.
[0016] In the fourth identification device of the present invention, in the first or second object identification device, the object identification device is an object identification device that identifies the target individual existing in a predetermined space, and when the comprehensive determination unit acquires a list in which the identification information of the target individual existing in the space is recorded, the recognition result is filtered based on the list, and based on the recognition result after the filtering, it is determined whether to generate the final result without using the discrimination result. It is an object identification device.
[0017] The fourth object identification device is an object identification device that identifies target individuals present in a predetermined space. In this case, the target individuals are present in this space, and the list contains identification information for the target individuals present in the space. By filtering the recognition results based on this, the impact of misrecognition on the final result can be reduced. When the impact of misrecognition is small, the final result is generated without using the discrimination results, making identification faster and more accurate.
[0018] The fifth target identification device of the present invention is a target identification device in which, in the fourth target identification device, the filtering includes deleting the recognition result corresponding to the target individual that does not exist in the space, based on the list.
[0019] The fifth target identification device removes recognition results (misrecognitions) that correspond to target individuals that do not exist in space, based on the list, thus making the filtered recognition results more accurate.
[0020] The sixth target identification device of the present invention is a target identification device in which, in the third target identification device, if the reliability of the recognition result does not meet a predetermined standard, the comprehensive determination unit integrates the recognition result and the determination result to generate the final result.
[0021] The sixth target identification device integrates the recognition result and the discrimination result to generate a final result when the reliability of the recognition result is low, thus obtaining a more accurate recognition result than simply recognizing the target individual from the marker.
[0022] The seventh object identification device of the present invention is an object identification device in which, in the sixth object identification device, the comprehensive determination unit performs weighting when integrating the recognition result and the discrimination result.
[0023] The seventh target identification device weights the results to prioritize one of the outcomes during integration. Since the weighting can be adjusted as needed depending on the target individual, space, and other applicable environmental factors, more flexible operation is possible.
[0024] The eighth target identification device of the present invention is a target identification device in which, in the first or second target identification device, the comprehensive determination unit generates an alert when the recognition result and the determination result do not match.
[0025] The reasons for discrepancies between recognition and discrimination results include not only low recognition / discrimination accuracy in one or both cases, but also human error such as tampering with signs. The seventh target identification device generates an alert when a discrepancy occurs, allowing the operator to recognize the possibility of tampering with signs and to take caution.
[0026] The ninth target identification device of the present invention is a target identification device in which, in the fourth target identification device, the data generation unit generates and / or updates the list based on the recognition result of the identification information acquisition unit based on visual data when the target individual enters and exits the space.
[0027] The list records identification information of target individuals present in the space. The eighth target identification device can generate and manage this list itself in order to identify target individuals upon entry into and exit from the space. This list improves the accuracy of the final result and speeds up processing.
[0028] A tenth object identification device of the present invention is an object identification device that further comprises a learning unit in the first or second object identification device that performs machine learning of the classifier using the training dataset, wherein the learning unit performs machine learning based on the training dataset generated based on the visual data collected and stored over a predetermined period of time, and the overall judgment unit, after the machine learning, considers the visual data, including the stored visual data, as new visual data and generates the final result.
[0029] According to the 10th target identification device, since machine learning is performed using the visual data that has been stored, even when the visual data includes images, contains a large number of target individuals, or needs to be classified into many classes, the recognition of identification information, machine learning, and generation of classification results can be performed after the acquisition and storage of the visual data, thus enabling more efficient use of computing resources. In addition, by retrospectively using the stored visual data as new visual data, tracking of target individuals can be performed more easily.
[0030] A first object identification method of the present invention is an object identification method that identifies an object based on visual data in which an object has a sign on which identification information is recorded, and includes: recognizing the identification information from the sign recorded in the visual data; associating the corresponding identification information as a label with the object object recorded in the visual data and saving it, and generating a training dataset; generating a classifier obtained by machine learning using the training dataset; and generating a final result of object identification by summarizing the recognition result of the identification information and the discrimination result by the classifier for new visual data in which the object object is recorded.
[0031] The target identification device and target identification method will be described below with reference to the drawings. Figure 1 is a hardware configuration diagram of a first embodiment of the target identification device of the present invention. The target identification device 100 is a computer equipped with a processor 1, memory 2, communication interface 3, imaging device interface 4, and input / output interface 5.
[0032] Processor 1 is the central component of the target identification device 100 and has the function of executing the program stored in memory 2 and processing data. Specifically, it has the function of executing instructions included in the program, arithmetic functions, the function of reading data from memory 2 etc., processing it and writing it back to memory 2 etc. as needed, and the function of controlling other things (clock management, power management, etc.). Examples of processor 1 include CPU (Central Processing Unit), GPU (Graphics Processing Unit), DSP (Digital Signal Processor), and FPGA (Field Programmable Gate Array).
[0033] Memory 2 consists of volatile memory and non-volatile memory. Volatile memory is mainly used for storing temporary data or data being processed. Hardware that provides this functionality includes RAM (Random Access Memory), cache memory, and SRAM (Static RAM). On the other hand, non-volatile memory is mainly used for long-term data storage. Hardware that provides this functionality includes flash memory, SSD (Solid State Drive), ROM (Read Only Memory), MRAM (Magnetoresistive RAM), and PRAM (Phase-change RAM).
[0034] The communication interface 3 provides functions such as data transmission and reception between the target identification device 100 and other devices or networks, integrity verification and error checking of transmission and reception, network connection management, and security assurance. Hardware that provides such functions includes wired communication interfaces such as optical fiber interfaces and Ethernet® interfaces, and wireless communication interfaces such as Wi-Fi® modules and Bluetooth® modules.
[0035] The imaging device interface (I / F4) connects the target identification device 100 to an external imaging device and performs transmission, reception, and control of visual data (images and / or video). Hardware with such functions includes "CSI" (Camera Serial Interface), USB camera interface, "HDMI" (registered trademark, High-Definition Multimedia Interface), and "GigE Vision" (registered trademark). The target identification device 100 may be connected to one imaging device or multiple imaging devices. The target identification device 100 may also be equipped with multiple types of imaging device interfaces (I / F4). On the other hand, the target identification device 100 does not need to be equipped with an imaging device interface (I / F4), in which case the imaging device may be connected to a network via the communication interface (I / F3).
[0036] The imaging device connected to the target identification device 100 is a device for acquiring visual data such as still images and / or videos. The imaging device uses visible light, infrared light, and lasers to image the target individual. The imaging device may be fixed or movable. An example of a movable imaging device is an imaging device consisting of a video camera and a drone on which the video camera is mounted. The number and type of imaging devices connected to the target identification device 100 are not particularly limited and can be changed as appropriate depending on the application.
[0037] The I / O interface (I / F5) transmits and receives data to and from external input / output devices (keyboards, mice, printers, and monitors, etc.). It may also supply power to connected external hardware. Hardware with such functionality includes USB (Universal Serial Bus), HDMI, DisplayPort, Bluetooth, and serial ports.
[0038] Figure 2 is a functional block diagram of the target identification device 100. The target identification device 100 comprises a control unit 102, a storage unit 104, a visual data acquisition unit 106, a sign detection unit 110, an identification information acquisition unit 112, a data generation unit 114, a learning unit 116, an appearance analysis unit 118, and an overall determination unit 120.
[0039] The control unit 102 controls the entire target identification device 100. The functions of the control unit 102 are realized by the processor 1, etc. The storage unit 104 stores programs and necessary data. The functions of the storage unit 104 are realized by the memory 2, etc.
[0040] The visual data acquisition unit 106 transmits commands for controlling the imaging device 90 connected to the target identification device 100, and receives visual data in which the target individual is recorded (or not recorded). The functions of the visual data acquisition unit 106 are realized by the execution of a program stored in the imaging device I / F 4 (or communication I / F 3) and memory 2 by the processor 1. In this example, the visual data acquisition unit 106 is described as receiving images from the imaging device 90, but the visual data acquired by the visual data acquisition unit 106 is not limited to images (still images), but may also be video, or a combination of the two.
[0041] In this example, the images acquired by the visual data acquisition unit 106 include the target cows 10A, 10B, and 10C (images without the cows may also be acquired). The target cows 10A, 10B, and 10C, housed in the barn 20, are wearing ear tags as identification markers. The imaging device 90 is installed in the barn 20. The acquired images may include those without the ear tags.
[0042] In this example, the target individual identified by the target identification device 100 is a cow, but the target (target individual) identified by the target identification device of the present invention is not limited to this. The target individual may be any individual equipped with a marker bearing visually recognizable identification information, and may be a human, livestock, agricultural or marine product, industrial product, merchandise, or vehicle. Furthermore, the target individual is an object that can be individually identified by its appearance (characteristic features). The target individual may include both moving and stationary objects. If the target individual is moving, the movement may be due to its own will or due to external instructions. In particular, when the target individual is livestock such as a cow, tracking can be easily performed by identification using the target identification device 100, making it preferably applicable to understanding and managing the health status and estrus status based on the behavior of the target individual.
[0043] Furthermore, the aforementioned "objects that can be individually identified by their appearance" may include not only objects that can be individually identified by their appearance, but also objects that can be separated into a certain group by their appearance. The latter includes, for example, identical products (the same product, but with different manufacturing numbers, etc.). Liquids and gases cannot be considered target individuals in themselves, but they can be considered target individuals when contained in a container, etc.
[0044] The target individual is preferably an individual equipped with a marker bearing visually recognizable identification information, and preferably an object that can be individually identified by its appearance (feature points). Furthermore, if it moves, it is more preferably an object that moves within a predetermined space and / or enters and exits that predetermined space. The predetermined space is preferably a space in which the entry and exit of the target individual can be monitored by visual data.
[0045] The mark detection unit 110 detects target individuals and marks in images acquired by the visual data acquisition unit 106. The function of the mark detection unit 110 is realized by the execution of a program stored in memory 2 by processor 1. An example of a specific procedure is a program flow in which (1) a bounding box containing the target individual (cow) in the image is generated, at which point if the image does not contain a cow it is identified, and if multiple cows are included it a bounding box is generated for each cow [target detection] (2) the generated bounding box is further analyzed to extract the ear tag portion and generate a cropped image, and if the ear tag is not included it it is identified [mark detection].
[0046] General object detection algorithms can be used to detect cows and ear tags in images. For example, "YOLO (You Only Look Once)" and "Faster R-CNN" can be used. For extracting the ear tag portion for each cow in (2), segmentation models such as "U-Net" and "Mask R-CNN" can also be used. Models for extracting target individuals (cows) and markers (ear tags) are generated by first training a model to detect cows using a dataset of cows, and then training a model to detect ear tags from cows using a dataset of ear tag portions. When using video (moving images) as visual data, the video can be divided into frames to generate images, and the above processing can be performed based on the generated images (still images). In this case, by performing object tracking, the same target individual (cow) and marker (ear tag) can be tracked between frames, which can increase the reliability of the recognition result. In this case, known tracking algorithms such as "CSRT (Discriminative Correlation Filter with Channel and Spatial Reliability)", "KCF (Kernelized Correlation Filter)", and "DeepSORT" can be used.
[0047] The identification information acquisition unit 112 recognizes the identification information of the target individual from the trimmed image including the tag. Recognition of identification information is performed using both the learning flow described later and the identification flow based on new visual data. The function of the identification information acquisition unit 112 is realized by the execution of a program stored in memory 2 by processor 1. In the case of cattle, ear tags may have 1 to 5 digits, and often 4 to 5 digits, as identification information. When the identification information is in the form of characters, the identification information acquisition unit 112 recognizes the identification information using optical character recognition (OCR). An example of a specific procedure is a program flow that includes: (1) converting the ear tag image (trimmed image) to grayscale; (2) smoothing the converted image using a Gaussian filter or the like to remove noise (background, etc.); (3) performing binarization; (4) extracting character regions using contour detection or the like; and (5) recognizing characters from the processed image using an OCR engine such as "Tesseract". In addition, deep learning models and other technologies may be used in conjunction to perform model tuning specifically for the font and shape of the identification information.
[0048] The identification information acquisition unit 112 outputs the confidence level of the recognition result of the identification information. The confidence level is the confidence level information for the recognized character, and can be defined as a numerical value in the range of 0 to 100, where a higher value indicates more accurate recognition. Alternatively, "-1" can be defined as a numerical value indicating that recognition was not possible.
[0049] In this example, the identification information of the ear tag has a one-to-one relationship with the target individual (cattle). That is, a certain piece of identification information (number) corresponds to a certain target individual (cattle). In this case, the identification information can be called individual identification information. However, the term "identification information" in this specification includes not only the individual identification information described above, but also product identification information such as barcodes attached to similar products. In this case, there may be many target individuals for one piece of identification information, resulting in a one-to-many relationship. In other words, the target individual corresponding to a certain piece of identification information is not a specific single individual, but rather an individual included in a group (e.g., a product) corresponding to the identification information. Whether the relationship between the identification information and the target individual is one-to-one or one-to-many does not affect the operation of the target identification device 100. That is, the target identification device 100 can generate the final result of identifying the target individual whether the relationship between the identification information and the target individual is one-to-one or one-to-many. However, in the case of one-to-one, the final result is also one-to-one, meaning the target identification device 100 functions as an individual identification device. On the other hand, if the identification information is one-to-many, the final result is also one-to-many; that is, the target identification device 100 functions as a product identification device. Here, "product identification device" means a device that identifies a predefined group to which homogeneous target individuals belong. Specifically, it means identifying two target individuals, A company's chocolate (manufacturing number A001) and A company's chocolate (manufacturing number A002), as "A company's chocolate." Although they are separate individuals with different manufacturing numbers, they are homogeneous and of the same type as products (belonging to the defined group "A company's chocolate").
[0050] The identification information does not have to be text. If the identification information is, for example, a graphical element such as a mark or logo, the identification information acquisition unit 112 includes a model that can classify predetermined types of graphical elements that have been pre-trained. Known methods can be used for training. For example, one method is to collect images of the graphical elements to be recognized, create a dataset by labeling them, and train using an image classification model such as a CNN (Convolutional Neural Network). In this case, a method of fine-tuning a pre-trained model ("ResNet", "MobileNet", etc.) may also be used. In addition to the above, a method of comparing graphical elements with pre-prepared template images and matching them using a feature detection algorithm ("SIFT", "SURF", "ORB", etc.) can also be adopted.
[0051] Furthermore, if the identification information is a barcode or a two-dimensional code, the identification information acquisition unit 112 can use a dedicated barcode reader library or two-dimensional code reader library. If the identification information is a barcode ("EAN", "UPC", etc.), "Zbar", "pyzbar", etc. can be used. If the identification information is a two-dimensional code, "Zbar", "pyzbar", "opencv", etc. can be used. Even if the identification information is not text, image preprocessing (grayscale conversion, noise reduction, binarization, etc.) may be performed to improve reliability (improve recognition accuracy). In any case, the identification information acquisition unit 112 outputs the reliability of the recognition result of the identification information. Depending on the library used, it may not be possible to directly output the reliability, but in this case, the reliability of the recognition result is calculated indirectly (for example, by the identity of the results obtained by repeated reading, and / or by the identity of the results when multiple libraries are used).
[0052] The data generation unit 114 associates the identification results as labels with the target individuals (cows) recorded in the images (visual data) in which the identification information has been recognized, and saves them, generating a training dataset to be used for training the classifier described later. The functions of the data generation unit 114 are realized by the execution of a program stored in memory 2 by processor 1. The generated training dataset is also stored in memory 2.
[0053] One example of the specific procedure is as follows: (1) Using the bounding box generated by the mark detection unit 110, one or more target individuals (cows) are extracted from the image data (generation of trimmed images). (2) The recognition result recognized by the identification information acquisition unit 112 is associated with each trimmed image (specifically, label information is assigned). (3) The trimmed images are saved in separate folders according to the recognition result (class). In this example, (1) is processed by the mark detection unit 110.
[0054] Furthermore, augmentation may be performed on the cropped images by rotating, flipping, translating, zooming, etc., as well as by masking parts of the image to include ear tags, etc. For example, 80% of the data obtained in this way can be used as training data and the remainder as test data to form the training dataset.
[0055] Figure 3 is an explanatory diagram of the data generation flow by the visual data acquisition unit 106, the sign detection unit 110, the identification information acquisition unit 112, and the data generation unit 114. Figure 4 shows the configuration of the training dataset.
[0056] Image 30, captured by the imaging device 90 and acquired by the visual data acquisition unit 106, includes the target individuals, cows 10A and 10B. The mark detection unit 110 first detects the target individuals from image 30. Since image 30 contains two target individuals, cows 10A and 10B, the mark detection unit 110 detects each cow. Figure 3 schematically shows the state in which a bounding box 30A has been generated for cow 10A. Next, the mark detection unit 110 detects a mark for each target individual. Figure 3 shows the generated trimmed image 30B after detecting the ear tag attached to cow 10A. The identification information acquisition unit 112 recognizes the identification information (in this case, the number "2") from the trimmed image 30B. In reality, the identification information acquisition unit 112 outputs the identification result (the number "2") and its confidence level, but here, for the sake of explanation, only the recognition result "2" is shown.
[0057] The data generation unit 114 saves the cropped image 30C of cow 10A, generated based on the bounding box 30A, by associating it with the recognition result ("2") as a label. Specifically, one method is to save the cropped image 30C of cow 10A in folder 16 and name the folder 14 "Recognition Result "2". In this way, cropped images of the target individuals (cow 10A, cow 10B, cow 10C) are generated from multiple images, classified according to the recognition result and saved in the generated folder, which serves as the training dataset. Note that associating the recognition result with a label may include not only using the recognition result itself as the label, but also using other information associated with the recognition result as the label. Other information associated with the recognition result may include, for example, the name of the target individual. Instead of recognition result "2", the label "Taro", which is the name of the cow with the ear tag "2", may be assigned based on pre-stored information.
[0058] Figure 3 illustrates that a single image contains (records) multiple cows, and each of these is detected and subsequently processed. However, when the mark detection unit 110 recognizes that a single image contains multiple target individuals, it may choose not to perform the subsequent processing on that image (i.e., not to create data to be incorporated into the training dataset). In this case, the training dataset will be generated mainly using images containing only one target individual. This can reduce the processing burden of target individual detection.
[0059] The above-described method for detecting the target individual and the marker by the marker detection unit 110 is just one example, and other methods can also be used. For example, pixel-level object classification (instance segmentation) using pre-trained deep learning-based models such as "Mask R-CNN," "Panoptic Segmentation," and "Instance-aware FCN" can also be used.
[0060] Furthermore, in this example, the image targeted for recognition of identification information in the identification information acquisition unit 112 and the original image of the cropped image labeled by the recognition result of the identification information in the data generation unit 114 were described as being the same. In other words, an example was described in which identification information is acquired from a single image and an image for inclusion in the training dataset is generated. However, the above do not have to be the same. For example, if the visual data is video, suppose it is divided into 24 images per second, and identification information is acquired based on the image of the first frame. If the mark is clearly visible in this image, but the target individual is somewhat unclear, the image saved in association with the identification information may be a different image, such as the next or the frame after that, as long as the identity of the target individual can be guaranteed by tracking or the like. On the other hand, in terms of ease of processing, it is preferable that the acquisition of identification information and the generation of visual data to be saved in association with this identification information are based on the same visual data.
[0061] The learning unit 116 is a classifier that constitutes part of the appearance analysis unit 118, which will be described later. It generates a classifier that classifies images of target individuals into classes using machine learning with the training dataset described above. The function of the learning unit 116 is realized when a program stored in memory 2 is executed by processor 1. The generated classifier is then incorporated into the appearance analysis unit 118.
[0062] While there are no particular limitations on the classifier used for image classification, a model based on a Convolutional Neural Network (CNN) can be used, for example. In this example, the classes to be classified by this classifier are each of the cows managed in the barn, which is often more numerous than general image classification (for example, classifying whether an image is a cow or something other than a cow). When such more complex classification is required, it may be more appropriate to fine-tune a pre-trained model such as "ResNet," "MobileNet," or "EfficientNet," that is, to adjust some layers using the training dataset, rather than building a CNN from scratch using a training dataset, and the appropriate method can be used depending on the situation. In this case, semi-supervised learning using images that could not be labeled (images for which recognition results were not assigned) can also be used.
[0063] When the appearance analysis unit 118 is given new visual data (an image in this example), it uses a classifier to classify the target individuals contained in the image into classes derived from identification information. If the identification information is individual identification information, the class becomes the individual identification information. The function of the appearance analysis unit 118 is realized by the execution of a program stored in memory 2 by processor 1. The built-in classifier is generated by the learning unit 116.
[0064] The discrimination performed by the appearance analysis unit 118 is carried out, for example, by the following two-step process: (1) To identify the location of 0, 1, or multiple cows in the new image, an object detection algorithm is used to generate bounding boxes surrounding the cows and extract the region corresponding to each target individual (cow) (generation of a cropped image). (2) The cropped image is input to the classifier and a discrimination result is generated. In this example, process (1) is performed by the mark detection unit 110. The discrimination result in (2) is generated as a probability distribution (confidence level) for each class.
[0065] The comprehensive determination unit 120 combines the recognition results of the identification information acquired by the identification information acquisition unit 112 and the discrimination results of the appearance analysis unit 118, based on the newly recorded visual data of the target individual, to generate the final identification result of the target individual. The function of the comprehensive determination unit 120 is realized by the execution of a program stored in memory 2 by processor 1.
[0066] Methods for summarizing recognition and classification results include, for example, adopting one of the results to form the final result, or integrating both to form the final result. Adopting one of the results might occur, for example, if only one result was obtained, or if the confidence level of one of the results did not meet a certain standard. A certain standard could be, for example, setting a threshold number (e.g., 50) when the confidence level is expressed on a scale of 0 to 100. That is, if the confidence level is below this threshold, the result is rejected, and the final result is generated using the other result (recognition or classification). Methods for integrating recognition and classification results include simply summing the confidence levels of the corresponding classes and selecting the one with the highest total confidence level for each class as the final result. Other methods include weighting one or both results by applying predetermined coefficients. Details of the summarization methods will be described later.
[0067] The comprehensive determination unit 120 may filter the recognition results and / or discrimination results using a list of target individuals housed in the monitored space, and determine the summarization method based on the confidence level after filtering. In this example, the list is a list in which the identification information of the cows (target individuals) housed in the barn 20 is recorded. Currently, cows 10A, 10B, and 10C are housed in the barn 20, so the list contains the identification information (numbers written on the ear tags) of cows 10A, 10B, and 10C.
[0068] Using this list, if the recognition or classification result includes an identification result for a cow that "should not exist," it can be deleted regardless of its confidence level. For example, suppose barn 20 houses cows with ear tag numbers "2," "3," and "8," and this is recorded in the list. In this case, if the recognition result identifies cow "5," this result can be deleted regardless of its confidence level. This filtering method is particularly effective in improving the accuracy of recognition results. In this way, by determining the summarization method—that is, whether to use one of the results or to combine both—based on the filtered recognition and / or classification results, a more accurate final result can be generated. For example, if the filtering results narrow down the recognition results to one, the final result may be generated based on that result (without using the classification result), even if its confidence level is 50. Even if the results are not narrowed down to one, if the confidence level of either of the remaining recognition results meets a predetermined standard when compared, the final result can be generated using that result. This list may be provided externally, or, in a designated space where the entry and exit of target individuals are monitored, such as a cowshed, the target identification device 100 may generate and update it itself using the method described later.
[0069] Next, the operation of the target identification device 100 will be explained using a flowchart. Figure 5 is a flowchart (learning flowchart) showing the procedure for generating a classifier by the target identification device 100. First, as step S101, the visual data acquisition unit 106 acquires the image captured by the imaging device 90. Figure 6 is a perspective view of the barn 20. The barn 20 is a space separated by a fence, where the entry and exit of cows is controlled. The barn 20 houses three cows: cow 10A, cow 10B, and cow 10C. Each cow is given an ear tag, and the individual identification number of each cow is written on the ear tag as identification information.
[0070] Cameras 90A and 90B are positioned in the cowshed 20 at a location that allows for an overview of the entire cowshed 20, and in front of the watering hole 40, respectively. Both cameras 90A and 90B are imaging devices 90. The visual data acquisition unit 106 causes cameras 90A and 90B to acquire images at any desired timing. For example, camera 90A may acquire images at predetermined time intervals to create data that can be tracked, while camera 90B may acquire images each time a cow comes to drink water. In addition to images, video may also be acquired; for example, video may be acquired from camera 90A and images from camera 90B. There is no particular limit to the number of imaging devices 90. In this example, two are used, but one may be used. Marking and detection of target individuals may be performed based on visual data acquired from one imaging device 90.
[0071] Next, in step S102, the mark detection unit 110 detects the target individual and the mark from the image. The image may contain multiple cows, in which case each cow is detected and the mark (ear tag) is detected for each detected cow. The image used at this time may be captured by either camera 90A or 90B. For example, since ear tags are easily captured in images from camera 90B, in the learning flow, images captured by camera 90B are mainly used, and in the identification flow (Figures 7 and 8 described later), for example, if the final result is used for tracking cows, images captured by camera 90A, which makes it easier to track the movement of cows, may be used as "new images".
[0072] Next, in step S103, the mark detection unit 110 determines whether a mark (ear tag) has been detected. If a mark is detected (step S103: YES), in the next step S104, the identification information acquisition unit 112 recognizes the identification information from the mark (ear tag). In this example, the identification information is a number that is individual identification information. The recognition result of the identification information acquisition unit 112 is output along with the confidence level. On the other hand, if a mark is not detected from the image (step S103: NO), the recognition result from the identification information acquisition unit 112 cannot be obtained, so the data generation unit 114 saves the image without associating it with the recognition result (step S106). Cases where a mark is not detected include cases where the mark recorded in the visual data cannot be detected, and cases where the mark is not recorded in the visual data. The latter includes cases where the mark was not recorded due to the influence of the target individual's posture, etc., as well as cases where the mark was removed from the target individual for some reason after the learning flow. In step S106, if the mark detection unit 110 detects one or more cows, the images of each cow (trimmed images of individual cows) trimmed based on the bounding box are saved without being linked to the recognition result. In this example, images are saved in step S106 without being linked to the recognition result, but images that could not be linked do not need to be saved. In that case, images that could not be linked are discarded in step S106.
[0073] Returning to step S104, after the identification information has been recognized, step S105 is performed, comparing the confidence level of the recognition result with a threshold. At this time, at least the recognition result with the highest confidence level is compared with the threshold. For example, if the recognition results are "2" with a confidence level of 50 and "3" with a confidence level of 30, at least the confidence level of "2" with the highest confidence level of 50 is compared with the threshold. Furthermore, the confidence level of "3" (and subsequent results) may also be compared with the threshold. Note that the confidence level does not need to be a number between 1 and 100; an arbitrary numerical range can be defined. The threshold can also be determined appropriately accordingly. If, as a result of the comparison with the threshold, the confidence level of the recognition result is below the threshold (step S105: YES), the trimmed image of the cow is saved without being linked to the recognition result (step S106). If the confidence level is below the threshold, it is possible that the accuracy of the recognition result is insufficient, etc., so linking with the recognition result is postponed. In this case as well, the image may be discarded without being saved.
[0074] On the other hand, if the confidence level of the recognition result exceeds the threshold (step S105: NO), in step S107, the data generation unit 114 saves the image associated with the recognition result. That is, it saves the cropped image labeled according to the recognition result. In this case, if the mark detection unit 110 has detected one or more cows, the recognition result for each cow is associated with the image of each cow that has been cropped based on the bounding box. In other words, the comparison with the threshold in step S105 is performed on a per-cow (target individual) basis in the image, not on an image basis. The saved image is the corresponding cropped image. A large number of these labeled images are accumulated and organized into a group of folders with a configuration that reflects the recognition results, becoming the training dataset. The training dataset may also include unlabeled images saved in step S106 as additional data. It is not necessary to use all of the saved images; some may be removed by operator operations, etc., and used for machine learning. For example, images of cattle that have been shipped may be removed, or images with low learning effectiveness (similar images) may be excluded to reduce computational cost. Annotated images may also be added externally and used in machine learning. In this example, cropped images are described as being used for machine learning, but it is not necessary to use cropped images; images according to the object detection method in the mark detection unit may be used. Images including the background may also be used.
[0075] Next, in step S108, the learning unit 116 performs machine learning using the training dataset and generates a classifier to be incorporated into the appearance analysis unit 118.
[0076] Figures 7 and 8 illustrate the identification flow when a new image is provided to an object identification device 100, which includes an appearance analysis unit 118 containing a trained classifier. First, the initial steps S101 to S103 are the same as the flow in Figure 5. If the mark detection unit 110 does not detect a mark from the image (S103: NO, connector A), this includes cases where, for example, a cow is included in the image, but the mark (ear tag) is hidden and not visible, and identification information cannot be obtained. In this case, the next step is S201, when an image of the target individual (a cropped image of the target individual) is input to the appearance analysis unit 118. Then, in step S202, the appearance analysis unit 118 generates a discrimination result based on the image of the target individual. Based on this discrimination result, the comprehensive judgment unit 120 generates the final identification result (step S203). In this case, the recognition result is not referenced in generating the final result, and the discrimination result is mainly used, but the information in the list may also be used. In other words, the discrimination results may be filtered based on the list, and the final result may be generated from the filtered discrimination results. However, the effectiveness of filtering is limited compared to filtering the recognition results.
[0077] On the other hand, if the sign detection unit 110 detects a sign from the image (step S103: YES), in step S104, the identification information acquisition unit 112 recognizes identification information from the sign. Next, the confidence level of the recognition result is compared with a threshold (step S105). If the confidence level of the recognition result exceeds the threshold, that is, "not below the threshold" (step S105: NO), the comprehensive judgment unit 120 generates the final result from the recognition result (step S204). In this case, the classification result by the classifier is not used to generate the final result, but filtering by list may be performed as in step S203. Note that in this flow, if the confidence level of the recognition result exceeds the threshold (step S105: NO), the comprehensive judgment unit 120 generates the final result without using the classification result, but it may be configured to integrate the classification results to generate the final result. In this case, even when (step S105: NO), the same flow as in (step S105: YES) is performed as in step S201 and onward in Figure 8.
[0078] On the other hand, if the confidence level of the recognition result is below a threshold (step S105: YES, connector B), in step S201, the image of the target individual (the cropped image of the target individual) is input to the appearance analysis unit 118. Then, in step S202, the appearance analysis unit 118 generates a discrimination result based on the image of the target individual. Next, in step S205, the comprehensive judgment unit 120 combines the recognition result and the judgment result to generate a final result.
[0079] Figure 9 is an explanatory diagram illustrating a specific example of the method for summarizing the recognition results and discrimination results in step S205. Although presented in a table format, this is for illustrative purposes only and does not illustrate the data structure, etc. Figure 9 shows examples of the generation of final results for three cropped images 300A, 300B, and 300C. The "Input Image" column shows the cropped images 300A, 300B, and 300C generated by the mark detection unit 110. These were all cropped based on the bounding box of the target individual (cow) detected from the image acquired by the visual data acquisition unit 106. While all are images of the cow's head, this is for illustrative purposes only; in reality, images of other parts, such as the whole body, may also be used.
[0080] The "Tag Extraction" column shows the trimmed images of ear tags extracted from the trimmed images 300A, 300B, and 300C on the left, respectively. Specifically, trimmed image 301A is extracted from trimmed image 300A, trimmed image 301B is extracted from trimmed image 300B, and trimmed image 301C is extracted from trimmed image 300C. In the following, trimmed images 300A, 300B, and 300C will be collectively referred to as "Cattle Images," and trimmed images 301A, 301B, and 301C will be collectively referred to as "Ear Tag Images."
[0081] Next, the "Identification Information Recognition Result" column shows the recognition result by the identification information acquisition unit 112 based on the ear tag image, along with the confidence level. In this example, the confidence level is represented by a numerical value from 0 to 100. From the cropped image 301A, it is shown that the confidence level of the recognition result for "Cow (2)," i.e., identification information "2," is 50%, and the confidence level of the recognition result for "Cow (5)," i.e., identification information "5," is also 50% (code 62A). Similarly, the recognition results and confidence levels are shown for the cropped images 12B and 12C, respectively (codes 62B and 62C).
[0082] The "Image Recognition Results" column shows the recognition results from the appearance analysis unit 118 based on the cow images, along with their confidence levels. From the cropped image 300, it is shown that "Cow (2)" and "Cow (8)" were recognized with a confidence level of 50% (code 64A). Similarly, the recognition results and confidence levels are shown for the cropped images 300B and 300C, respectively (codes 64B and 64C).
[0083] Figure 9 illustrates a specific example of the method for summarizing the recognition result and discrimination result in step S205. Therefore, in step S105, the confidence level of the recognition result is determined to be below the threshold, meaning that the final result is not generated solely from the recognition result.
[0084] The "Final Result" column shows the recognition result by the comprehensive judgment unit 120, an example of a method for summarizing the discrimination result, and the final result generated as a result. In this example, the summarization method includes not using one of the results and integrating both results.
[0085] For the trimmed image 300A, a filtering method using a list is shown as an example of "not used". The comprehensive determination unit 120 can confirm that cows (2), (3), and (8) are housed in the barn using the given (or created / updated) list, and can delete the recognition result corresponding to cow (5), which should not be in the barn. In this case, filtering leaves only cow (2) as the recognition result. Although the confidence level is 50%, since it does not include other results, it is possible to generate the final result "cow (2)" without using the discrimination results (code 66A). While it is not essential to use a list to generate the final result, it is preferable because it allows for more accurate identification.
[0086] Furthermore, the trimmed image 300A can also be "integrated". An example of integration is to compare the sum of the confidence scores for each class of the recognition result and the classification result. In this case, since cow (2) is 100, cow (5) is 50, and cow (8) is 50, the final result can be determined to be cow (2). Weighting may also be applied when summing. Specifically, the recognition result and / or the classification result may be multiplied by a predetermined coefficient before summing.
[0087] For the cropped images 300B and 300C, the "Final Result" column shows the method for "integrating" the recognition and classification results. In both cases, the confidence levels for each class of the recognition and classification results are multiplied by a coefficient of 0.5 and summed, and the final result is generated from the comparison of these sums (codes 66B and 66C). By using different coefficients for the recognition and classification results, weighting can be applied. By applying weighting, it is possible to arbitrarily choose which result to give more weight to, the recognition result or the classification result, depending on the environment. For example, in an environment where misrecognition of identification information is likely to occur, the coefficient for the classification result may be increased.
[0088] Returning to Figure 8, in step S205, the recognition result and the discrimination result are combined as described above to generate the final result. The identification flow may end here, but it may be determined before or after generating the final result whether the recognition result and the discrimination result differ. In step S206, it is determined whether the recognition result and the discrimination result differ. A discrepancy in results means that the recognition result and the discrimination result are different. In this case, both may be after filtering by the list. If the results differ (step S206: YES), the target identification device 100 generates an alert (step S207). On the other hand, if the results do not differ, the identification flow ends. Note that "no discrepancy in results" includes not only cases where the recognition result and the discrimination result are the same, but also cases where one of the results cannot be obtained. In other words, "no discrepancy" is defined as "cases other than those where there is a discrepancy." Note that in this example, the processing in step S206 is performed after the processing in step S205, but step S206 may be performed before step S205, or simultaneously.
[0089] Furthermore, if the classifier's machine learning is insufficient, or if machine learning has not yet been performed, such as immediately after the cows are brought into barn 20, a classification result may not be obtained, or the reliability of the classification result may be insufficient. In this case, the overall judgment unit may generate the final result based on the recognition result, regardless of the result of step S105 (reliability of the recognition result) in the flow chart of Figure 7.
[0090] In this example, using the target identification device 100 equipped with a machine learning-trained classifier, data for tracking cattle can be generated. For example, by periodically capturing images or videos with the camera 90A and identifying the cattle recorded, it becomes possible to track cattle within the barn 20. From the cattle tracking data, the behavioral history of each cattle, such as the amount of movement, can be obtained. Traditionally, there has been a demand to know the estrus period of cattle for artificial insemination, and it is known that there is a correlation between the estrus period and behavioral history such as the amount of movement. Using the target identification device 100, the estrus period can be determined in a stress-free and effortless manner without attaching special equipment to the cattle. Furthermore, because identification is performed by combining identification information and image discrimination based on a single new visual data, more accurate identification is achieved in various environments. In addition, regarding image discrimination, the generation of images for classifier training, labeling, training dataset generation, and machine learning are all performed automatically, which is an advantage as it does not require excessive expertise from the operator.
[0091] When tracking is performed, the generation of a final result based on visual data is repeated. In this case, for each visual data, when summarizing the recognition result and judgment result in step S205, the final result based on visual data acquired earlier may be used as a reference. For example, to avoid the contradiction of generating the same final result for target individuals that have been identified as different cows through tracking, a combination can be adopted that minimizes such contradictions (errors), regardless of the reliability of the judgment result. Also, when referring to previous results, weighting may be adjusted so that the most recent result has a greater influence. This is because the most recent result is more likely to reflect the current appearance of the target individual.
[0092] In the above example, in step S102 of Figure 7, the target individual and the marker are detected by the marker detection unit, and in the subsequent step S104, the identification information is recognized by the identification information acquisition unit and the recognition result is generated. However, with the target identification device 100 equipped with a classifier that has undergone machine learning, it is possible to identify the target individual without using the identification information. In this case, the marker detection unit may detect the target individual, and steps S104 and S105 may be omitted, and steps S201 onwards may be performed.
[0093] Next, based on a second embodiment in which the object identification device of the present invention is applied to a vehicle, steps S206 and S207 described above will be explained further, along with a method for generating and updating the list.
[0094] Figure 10(A) is a diagram showing the entrance gate to a parking facility. The parking facility 80A comprises a parking space capable of accommodating multiple vehicles, an entrance gate 80B to the parking space, and an exit gate 80D from the parking space (see Figures 13 and 14 below). The parking facility 80A is, for example, a time-based paid parking facility where vehicle entry and exit are monitored by visual data.
[0095] Near the entrance gate 80B of the parking facility 80A, imaging devices 90C and 90D are provided to acquire visual data of a vehicle 82A (target individual) attempting to enter the parking space. These imaging devices 90C and 90D are connected to the target identification device 100 via a visual data acquisition unit 106. Of these, imaging device 90C is configured to image the vehicle 82A from above and capture a planar image of it. Figure 10(B) is an example of a planar image (cropped) generated by the sign detection unit 110 based on the image of the vehicle 82A captured by imaging device 90C, and is a color image so that information such as body color can be obtained in addition to the shape. The detection of the vehicle 82A and cropping based on the bounding box are the same as in the first embodiment.
[0096] The imaging device 90D is configured to capture images of the area around the license plate 84A (sign) of vehicle 82A from the front. Figure 10(C) is an example of a cropped image of license plate 84A generated by the sign detection unit 110 based on the image of vehicle 82A captured by the imaging device 90D. The license plate 84A has the identification information (in this case, individual identification information), which is the number 86. The image in Figure 10(B) is saved in association with the recognition result extracted from Figure 10(C), becomes part of the training dataset, and is used for machine learning of the classifier. In this example, only a planar image of vehicle 82A is shown to be saved, but the images included in the training dataset are not limited to the above. Images captured by the imaging device 90D (the portion of vehicle 82A detected from it) may be included, or images acquired by other imaging devices not shown may be included. As described above, by collecting visual data of the target individual from various angles using multiple imaging devices, the accuracy of learning and identification can be improved. For example, an imaging device that captures a target object from above or from a bird's-eye view can collect visual data that can be used not only for learning and identification but also for tracking. Furthermore, if the target object is a vehicle, visual data including feature points that are difficult to see in anything other than a planar image (e.g., skylights, sunroofs, etc.) can be collected. On the other hand, imaging devices that capture images from above or from a bird's-eye view may have difficulty capturing the vehicle's license plate, and in such cases, it is preferable to use multiple imaging devices.
[0097] In this example, the visual data of vehicles acquired at the entry gate 80B is used for acquiring identification information, creating lists, and machine learning of the classifier. In terms of the flowchart explained in the first embodiment, it is used in the learning flow of Figure 5, and further used for generating lists, which will be described later. The target identification device 100, configured in this way to identify vehicles present in the parking space after the classifier has been adjusted, is used for tracking vehicles based on visual data captured by an imaging device (not shown) installed in the parking space, and for identifying exiting vehicles based on visual data captured by an imaging device installed at the exit gate 80D.
[0098] Figure 11 is a plan view of the parking space of parking facility 80A. Parking facility 80A is provided with multiple parking spaces 80C, where vehicles 82A, 82B, and 82C are parked. All of these vehicles entered the parking space by passing through the entrance gate 80B. The vehicle identification device 100 recognizes the identification information (number) written on the license plate of each vehicle from the visual data (images and / or video) acquired at the entrance gate 80B, and identifies each vehicle from the visual data.
[0099] Next, a method for generating a list of vehicles parked in the parking space will be described. Figure 12 is a flowchart of the list generation process. When generating the list, first, in step S401, the visual data acquisition unit 106 acquires an image captured by the imaging device 90D of the parking space entrance gate 80B. Next, in step S402, the sign detection unit 110 detects signs from the image (generating a cropped image). Next, in step S403, the identification information acquisition unit 112 recognizes identification information from the cropped image and generates a recognition result. Next, in step S404, the data generation unit 114 generates a list based on the recognition result. If a list has already been generated, the data generation unit 114 updates that list. The list should at least include the identification information of vehicles currently parked in the parking space, and may also include other information.
[0100] Each vehicle that enters the parking space exits through the exit gate 80D (see Figures 13 and 14). Therefore, the list can be updated by using the recognition result of the identification information acquisition unit at the time of exit. In step S405, the visual data acquisition unit 106 acquires the image captured by the imaging device 90F of the exit gate 80D from the parking space. Steps S402 and S403 are the same as above, and finally, in step S406, the data generation unit 114 updates the list based on the recognition result (removing the identification information of the vehicle in question from the list). In this way, by acquiring visual data including the markings of target individuals when they enter and exit a predetermined space and recognizing the identification information, a list of target individuals contained in that space can be generated. Examples of such spaces include sports playfields and courts where the entry and exit of players can be monitored, childcare and educational facilities where the entry and exit of children can be monitored, medical facilities where the entry and exit of patients can be monitored, and workplaces where the entry and exit of employees can be monitored. In addition to the above, highways and expressways where the entry and exit of vehicles can be monitored by visual data at each interchange can also be mentioned.
[0101] In this example, the parking facility 80A is equipped with an entry gate 80B and an exit gate 80D, but these may be shared, or there may be multiple of each. Furthermore, the target "space" does not need to be indoors, but any designated space (it does not need to be indoors) where visual data can be acquired when the target individual enters or exits, in other words, where entry and exit can be monitored by visual data.
[0102] Figure 13(A) shows the exit gate of a parking facility. Near the exit gate 80D of the parking facility 80A, similar to the entry gate 80B, imaging devices 90E and 90F are provided to acquire visual data of a vehicle 82A (target individual) attempting to exit the parking space. These imaging devices 90E and 90F are connected to the target identification device 100 via a visual data acquisition unit. Based on the images from imaging devices 90E and 90F, a plan view (cropped) of vehicle 82A in Figure 13(B) and an image of the license plate 84A of vehicle 82A in Figure 13(C) are generated. These images are treated as input images (new images) to the target identification device 100 for identifying the exiting vehicle 82A.
[0103] The imaging device 90F is configured to capture the license plate 84A (sign) of vehicle 82A from the front. Figure 13(C) is an example of an image (cropped) of the license plate 84A of vehicle 82A generated from an image captured by the imaging device 90F. The license plate 84A has the identification information (in this case, individual identification information), number 86, written on it. This image is also treated as an input image (new image) to the target identification device 100. Although the images in Figure 13(B) and Figure 13(C) were obtained from different imaging devices, they are images of the same vehicle acquired at almost the same time, so they can be treated as a set of input images. Hereafter, these will be collectively described as "exit input images".
[0104] The identification of vehicles based on the exit input image will be briefly explained following the flow in Figures 7 and 8. Since the details of each process are as described in the first embodiment, this section will mainly explain how the exit input image is used. The exit input image is generated through steps S101, S102, S103 (YES), and S104. Next, the confidence level of the recognition result is compared with a threshold (step S105). As shown in Figure 13(C), when vehicle 82A exits the parking space, dirt 88 is attached to the license plate 84A, and as a result, the confidence level of the recognition result is below the threshold (step S105: YES). In this case, the comprehensive determination unit 120 combines the recognition result and the discrimination result to generate the final result (steps S201, S202, and S205).
[0105] Next, we will explain an example where the recognition result and the discrimination result differ. Figure 14(A) is a diagram showing the exit gate of a parking facility. The arrangement of the imaging devices 90E and 90F at the exit gate 80D is the same as in Figure 14(A). Also, Figure 14(B) is a cropped image based on the image captured by imaging device 90E, and Figure 14(C) is a cropped image based on the image captured by imaging device 90F.
[0106] Based on these exit input images, the identification information acquisition unit recognizes the number 86 as identification information. At this time, the recognized identification information is the one that was attached to vehicle 82A upon entry. On the other hand, the appearance analysis unit 118 determines that it is vehicle 82C. From this result, it is recognized that the recognition result and the determination result do not match (Figure 8, step S206: YES), and an alert is generated. In this example, the recognized number belonged to vehicle 82A, but it is not necessary for the identification information to be one that exists within the parking space (recorded in the list). If either the identification information recorded upon entry or the appearance is inconsistent, an alert can be generated. Generating an alert can deter the avoidance of payment for paid parking facilities through impersonation, etc.
[0107] Next, a modified example of the learning flow shown in Figure 5 will be described based on a third embodiment in which the target identification device of the present invention is applied to athletes.
[0108] Figure 15 shows the movement trajectory 76 of a specific player 72 during a soccer match on a playing field 70. One known method for generating such a movement trajectory 76, or a so-called "heatmap" that visually represents which areas the player 72 moved to and how frequently, is to attach a GPS transmitter to the player 72. In contrast, using the target identification device 100, tracking data can be generated based on visual data captured by the imaging device, thus reducing the burden of attaching GPS devices to the player.
[0109] An example of a specific procedure is described below. First, one or more imaging devices are pre-arranged to capture images of various parts of the playfield 70 and connected to the target identification device 100. The visual data acquisition unit 106 uses these imaging devices to capture images of each player (target individual) during the match. In this example, the markings are predetermined locations on the uniform (for example, a certain area on the back or front), and the identification information is, for example, the "back number." A soccer playfield usually accommodates 22 players, and the back numbers may be the same for players on both teams. However, in many cases, the font, color, and position of the back numbers differ from team to team, so by pre-adjusting the marking detection unit 110 and the identification information acquisition unit 112 based on this image data, the back numbers can also function as individual identification information. It is also possible that the above adjustment is not performed, and the possibility remains that the same identification information (same back number) may be assigned to multiple players (target individuals) (it does not have to function as individual identification information). In that case, for example, even if the reliability is high, the final result should not be generated using only identification information. Instead, the flow should be adjusted so that the final result is generated by combining it with discrimination information, member information (list), etc. Furthermore, the identification information does not need to be of only one type; it may also include the jersey number, player name, and team logo. By combining these, the identification information can function as individual player identification information.
[0110] Figure 16 shows the playfield 70 at kickoff. Including player 72 whose movement trajectory should be recorded, the playfield 70 accommodates 22 players. Including substitutions, more than 22 players will be present on the playfield 70 throughout a match. At kickoff, each player is in their designated position, so depending on the placement of the imaging device (not shown), it may be easy to capture each player's jersey number 74. On the other hand, during the match, players are often mixed together on the playfield 70, making it difficult to obtain individual images of player 72 while recording their jersey number 74.
[0111] In such cases, the learning flow may be adjusted so that visual data is acquired and stored first, and then machine learning of the classifier is performed after that is complete. Figure 17 shows a modified example of the learning flow. First, in this example, it is preferable to use video as the visual data. The images obtained by dividing the video into frames are processed according to this flow. The difference from the flow in Figure 5 is that after the data generation unit saves the image either linked to the recognition result (step S107) or not linked (step S106), before machine learning (step S108), a decision is made in step S501 as to whether to continue acquiring visual data. The criteria for deciding whether to continue acquiring visual data are not particularly limited, but one example is to decide based on the elapsed time since the start of acquiring visual data. As for the elapsed time, in the case of soccer, a certain period of time can be considered, for example, 45 minutes in the first half, 45 minutes in the second half, or 90 minutes in a match. For example, visual data acquisition and image saving continue during the match (step S301: YES). On the other hand, when the match ends, visual data acquisition ends (step S301: NO). Subsequently, machine learning of the classifier is performed using the training dataset containing the obtained labeled (and unlabeled) images (step S108). In this way, even when the visual data is enormous due to reasons such as including video, data generation and machine learning can be performed sequentially, resulting in improved recognition accuracy.
[0112] The learning flow ends in step S108 above. The target identification device 100, equipped with this trained classifier, is capable of identifying each player (target individual) during a match. In this case, by using the target identification device 100, and by using visual data that was used (stored) or not used for learning as new input data, it becomes possible to identify each player during a match.
[0113] The flow in Figure 17 shows step S502, in which the visual data (video) used for training is used as new input visual data, in other words, the target players 72 are identified by working backward through the game. Tracking data can be generated by concatenating the identification results of the target players recorded in the video. The new data may consist only of visual data from an overhead viewpoint that is more suitable for creating tracking data, or it may include other data. According to this modified example, even if there are many classes to be classified by the classifier, or if there is a large amount of input visual data, training and identification can be performed sequentially, making it possible to apply it to various fields with high accuracy. In this example, the timing of machine learning is set to the end of one game, but this timing is arbitrary. It may also be performed periodically. For example, in the case of a target identification device that monitors the entry and exit of employees, by periodically performing the training flow, the classifier can be adjusted, and more accurate identification can be maintained in response to changes in the appearance of employees over time.
[0114] When the comprehensive judgment unit generates the final result, it may be possible to achieve more accurate identification by using additional information other than the identification result and discrimination result. In this example, one method is to add location information such as the player's position and time information to perform identification. For example, when identifying an unknown outfielder on a baseball playfield, the identification result and discrimination result indicate a 50% probability that it is player A and a 50% probability that it is player B. Considering the position and members, if the target player (target individual) is in the outfield and player B is the pitcher, the final result may be that it is player A.
[0115] The object identification device of the present invention has been described above with reference to the first to third embodiments. However, these are merely embodiments and examples that embody the technical concept of the present invention. The object identification device of the present invention can be modified and applied in various ways within the range that achieves the desired effect.
[0116] For example, regarding markings, the first embodiment describes an ear tag, and the second embodiment describes a license plate, both of which are clearly distinguishable from the target individual (cattle or vehicle). However, the markings are not limited to the above; they may also be applied to a part of the target individual, as in the third embodiment. Other examples of such markings include branding and tattoos.
[0117] Furthermore, while we have described a configuration in which the label detection unit detects each target individual in an image containing multiple target individuals and then detects a label for each individual, it is also possible to choose not to perform label detection in images containing multiple target individuals. In this case, images containing multiple target individuals do not need to be saved; that is, they do not need to be used for machine learning.
[0118] Furthermore, images in which the label detection unit failed to detect the target individual do not need to be saved, but they may be saved as "images that do not contain the target individual." These may also be used for machine learning.
[0119] Alternatively, the system may include an object identification device and one or more imaging devices.
[0120] Furthermore, when collecting images to obtain identification information, electronic tags such as "RFID" may be used in conjunction. [Explanation of Symbols]
[0121] 90 Imaging device 100 Target identification device 102 Control Unit 104 Storage section 106 Visual data acquisition unit 110 Label detection unit 112 Identification Information Acquisition Unit 114 Data Generation Unit 116 Learning Department 118 Appearance Analysis Section 120 Overall Judging Section
Claims
1. An object identification device that identifies an object based on visual data recorded of an object equipped with a marker bearing identification information, An identification information acquisition unit that recognizes the identification information from the sign recorded in the visual data, A data generation unit that associates the recognition results of the corresponding identification information obtained by the identification information acquisition unit with the target individuals recorded in the visual data and stores them as labels, and generates a training dataset, An appearance analysis unit including a classifier obtained by machine learning using the aforementioned training dataset, It comprises a comprehensive judgment unit, A target identification device in which, upon receiving new visual data on which the target individual is recorded, the comprehensive determination unit combines the recognition result of the identification information by the identification information acquisition unit and the discrimination result by the appearance analysis unit to generate a final result for the identification of the target individual.
2. If the identification information acquisition unit cannot recognize the identification information based on the new visual data, The target identification device according to claim 1, wherein the comprehensive determination unit generates the final result based on the determination result of the appearance analysis unit.
3. If the identification information acquisition unit is able to recognize the identification information based on the new visual data, The target identification device according to claim 1 or 2, wherein the comprehensive determination unit determines whether or not to generate the final result without using the determination result, based on the reliability of the recognition result of the identification information acquisition unit.
4. The aforementioned target identification device is a target identification device that identifies the target individual located in a predetermined space, The target identification device according to claim 1 or 2, wherein the comprehensive determination unit, upon obtaining a list in which identification information of the target individuals present in the space is recorded, filters the recognition results based on the list and, based on the filtered recognition results, determines whether or not to generate the final result without using the discrimination result.
5. The object identification device according to claim 4, wherein the filtering includes deleting the recognition results corresponding to the target individual that does not exist in the space, based on the list.
6. The target identification device according to claim 3, wherein the comprehensive determination unit, if the reliability of the recognition result does not meet a predetermined standard, integrates the recognition result and the discrimination result to generate the final result.
7. The target identification device according to claim 6, wherein the comprehensive determination unit performs weighting when integrating the recognition result and the discrimination result.
8. The target identification device according to claim 1 or 2, wherein the comprehensive determination unit generates an alert if the recognition result and the discrimination result do not match.
9. The target identification device according to claim 4, wherein the data generation unit generates and / or updates the list based on the recognition result of the identification information acquisition unit based on the visual data when the target individual enters and exits the space.
10. The system further comprises a learning unit that performs the machine learning of the classifier using the training dataset. The learning unit performs machine learning based on the training dataset generated based on the visual data collected and stored over a predetermined period of time. The target identification device according to claim 1 or 2, wherein the comprehensive determination unit, after machine learning, considers the visual data, including the accumulated visual data, as new visual data to generate the final result.
11. A method for identifying target individuals, which includes a marker bearing identification information, based on recorded visual data, wherein the target individual is identified based on the visual data of the target individual, Recognizing the identification information from the sign recorded in the visual data, The process involves associating the corresponding identification information as a label with the target individual recorded in the visual data and saving it, thereby generating a training dataset. To generate a classifier obtained by machine learning using the aforementioned training dataset, A method for identifying an object, comprising generating a final result for identifying the object by combining the recognition result of the identification information and the discrimination result by the classifier with respect to new visual data on which the object is recorded.
Citation Information
Patent Citations
Individual authentication system and method and system for tracing farm product by using the same
JP2006146570A