Container information recognition device, system, method, and program
The container information recognition device uses machine learning and YOLO deep learning models to address layout and environmental challenges, achieving high-accuracy and robust recognition of alphanumeric strings on containers.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-09-11
- Publication Date
- 2026-03-24
AI Technical Summary
Existing methods struggle to accurately recognize container numbers on containers due to variations in layout, font, color, clarity, and environmental conditions, leading to inconsistent recognition performance.
A container information recognition device utilizing machine learning techniques, including image acquisition and information recognition units, to identify alphanumeric strings on containers with high accuracy and robustness, employing YOLO deep learning models for region and character recognition.
Enables reliable and stable recognition of container displays from images, even under varying conditions, improving accuracy and robustness in identifying alphanumeric sequences.
Smart Images

Figure 2026052270000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an image recognition device or the like, particularly to an image processing device or the like that recognizes information displayed on a container.
Background Art
[0002] Generally, on the surface of a container used for transporting goods or the like, in accordance with the ISO standard (ISO6346), there are displayed characters and numbers such as an owner code consisting of three alphabetic characters, a category identifier consisting of one alphabetic character (for example, "U" for a cargo container), a serial number consisting of six digits, and a check digit (CD) consisting of one digit (hereinafter, these may be referred to as container numbers).
[0003] In recent years, attempts have been made to recognize the container number displayed on a container from an image or moving image in which a container with these container numbers is shown. For example, Patent Document 1 discloses a system that recognizes a container number by reading character information on the door and ceiling parts of a container with an existing character recognition software using a camera arranged on the portal tie beam of a gantry crane.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] Incidentally, the aforementioned ISO standard does not specify in detail how the letters or numbers mentioned above should be arranged on the container. Therefore, in actual containers, the layout of the letters or numbers (e.g., position, arrangement, size, height / width, range, etc.), font, color, etc. vary from container to container. Also, containers are often used outdoors and are exposed to various environments. As a result, some or almost all of the letters may be missing due to dirt, scratches, fading, etc., and the clarity of the letters or numbers also varies.
[0006] When attempting to recognize the container number attached to this type of container from an image using conventional methods, as mentioned above, recognition sometimes fails due to the various forms in which the letters or numbers are displayed, and the lack of clarity in some or all of them. Therefore, there were certain limitations to the accuracy of recognizing the markings on the container.
[0007] This invention has been made in view of the above-mentioned technical background, and its purpose is to reliably recognize the display on a container from an image. [Means for solving the problem]
[0008] The technical challenges described above can be solved by a container information recognition device, system, method, and program having the following configuration.
[0009] In other words, the container information recognition device according to the present invention includes an image acquisition unit that acquires a container image which is an image of a container displaying an alphabet string including an owner code consisting of three letters and a number string including a serial number consisting of six digits, and an information recognition unit that uses machine learning technology to recognize the alphabet string and the number string from the container image.
[0010] With this configuration, alphanumeric strings containing owner codes and numerical strings containing serial numbers, displayed on a container in various forms (e.g., various layouts (position, arrangement, size, height / width, range, etc.), font, color, clarity, etc.), can be recognized with high accuracy and robustness using generalization techniques based on machine learning. This allows for stable recognition of the display on the container from images. Note that images include images contained in moving images. [Effects of the Invention]
[0011] According to the present invention, it is possible to reliably recognize the display on a container from an image. [Brief explanation of the drawing]
[0012] [Figure 1] Figure 1 is an overall diagram of the container information recognition system. [Figure 2] Figure 2 is a hardware configuration diagram of the information processing device. [Figure 3] Figure 3 is a functional block diagram of the information processing device. [Figure 4] Figure 4 is a detailed flowchart of the learning process. [Figure 5] Figure 5 is a detailed flowchart of the information recognition process. [Figure 6] Figure 6 is an explanatory diagram showing an example of the target image data. [Figure 7] Figure 7 is an explanatory diagram showing another example of the target image data. [Figure 8] Figure 8 is an explanatory diagram showing examples of the first and second regions. [Figure 9] Figure 9 is an explanatory diagram showing an example of the output results of an alphabet recognition model. [Figure 10] Figure 10 is an explanatory diagram showing an example of the output results of a digit recognition model. [Figure 11] Figure 11 is a table showing the assignment of letters and numbers as defined in the ISO standard. [Figure 12]FIG. 12 is an explanatory diagram showing an example of check digit calculation. [Figure 13] FIG. 13 is an explanatory diagram showing an example of a temporary check digit database.
Embodiments for Carrying out the Invention
[0013] Hereinafter, preferred embodiments of the present invention will be described in detail with reference to the accompanying drawings.
[0014] (1. First Embodiment) As a first embodiment, an example in which the present invention is applied to a container information recognition system including a camera installed in a container yard and an information processing device will be described. Note that the application scenario of the present invention is not limited to the yard. Therefore, for example, it can be applied to various scenarios such as the scenario of operating a container by a gantry crane or inside a warehouse.
[0015] (1.1 Configuration) FIG. 1 is an overall configuration diagram of a container information recognition system 100. As is clear from the figure, the container information recognition system 100 includes at least a camera 10 and an information processing device 30, and they are connected to each other via a network.
[0016] In the present embodiment, the camera 10 is installed in the container yard. More specifically, the camera 10 is installed at the upper rear diagonal of the stop position of the container transport vehicle 13 that transports the container 15 into the container yard so as to include at least the top surface and the back surface of the container 15 within the angle of view. With this camera 10, an image or a moving image including the container 15 can be captured.
[0017] In this embodiment, the rear end and back of the top surface of the container 15 are marked or painted with at least three alphabetic owner codes, one alphabetic category identifier (for example, "U" for a cargo container), a six-digit serial number, and a one-digit check digit (CD) in accordance with the ISO standard (ISO 6346) (see also Figure 6). For convenience, the four alphabetic characters including the three-alphabetic owner code and the one-alphabetic category identifier may be referred to as the container code below.
[0018] Furthermore, while this embodiment shows an example where the container number is displayed on the top of the container in addition to the back (for example, Figure 6), there may be cases where there is no top panel at all, or where the container number is not displayed on the top panel. In this case, the container number is recognized only from the display on the back of the container.
[0019] Figure 2 is a hardware configuration diagram of the information processing device 30. As is clear from the figure, the information processing device 30 comprises a processing unit 31, a storage unit 32, a communication unit 33, an input unit 35, a display processing unit 36, an audio processing unit 37, and an I / O unit 38, which are connected to each other via a bus. The information processing device 30 is, for example, a personal computer, a tablet terminal, or the like.
[0020] The processing unit 31 is a processor such as a CPU / GPU, and performs various processes to execute programs stored in the memory unit 32, etc. The memory unit 32 is a storage device such as ROM, RAM, flash memory, or hard disk, and stores programs and data for executing the processes described later.
[0021] The communication unit 33 is a communication unit for communicating with external devices wirelessly or via wired connections, and it exchanges information with external devices. The input unit 35 processes information input from input devices 351 such as a mouse, keyboard, or touch panel and provides it to the processing unit 31. The display processing unit 36 processes visual information to be displayed on a display device 361 such as a display, either in cooperation with the processing unit 31 or independently.
[0022] The audio processing unit 37 processes the audio signal to be output to the speaker 371, either in cooperation with the processing unit 31 or independently. The I / O unit 38 performs input / output processing with external devices.
[0023] The configuration of the information processing device 30 is not limited to that of this embodiment, and may also incorporate an input device 351, a display device 361, or a speaker 371.
[0024] Figure 3 is a functional block diagram of the information processing device 30. Each function shown in the diagram is implemented by the processing unit 31.
[0025] The information processing device 30 includes a learning processing unit 311 for performing the learning process described later. The learning processing unit 311 is connected to the storage unit 32.
[0026] Furthermore, the information processing device 30 includes an information recognition unit 320 for information recognition processing, which will be described later. The information recognition unit 320 includes an image acquisition unit 312 that reads and acquires image information from the storage unit 32, an object determination unit 315 that determines objects in the image, and an area determination unit 313 that extracts a specific area on the image based on the object determination result.
[0027] The region identification unit 313 provides the region identification result to the alphabet recognition unit 316, which recognizes alphabets, and the digit recognition unit 317, which recognizes digits. The alphabet recognition unit 316 and the digit recognition unit 317 each provide the recognition result to the output unit 318.
[0028] If recognition fails by the alphabet recognition unit 316 or the digit recognition unit 317, the estimation unit 319 operates. The output of this estimation unit 319 is provided to the output unit 318.
[0029] The output unit 318 processes the provided recognition results and other information to output to the storage unit 32, display device 361, speaker 371, etc.
[0030] The above configuration is merely an example. Therefore, it can be implemented in various modified forms. For example, a network storage 20 or file server for storing camera images may be connected to the network (see Figure 1). Various processes, including learning and information recognition processing, described later, may be performed while referencing this network storage 20. Alternatively, all functions may be configured without using a network.
[0031] (1.2 Operation) (1.2.1 Learning Process) First, we will explain the data learning process that is performed prior to the information recognition process.
[0032] Figure 4 is a detailed flowchart of the learning process according to this embodiment. As is clear from the figure, when the process starts, the learning processing unit 311 reads the data to be learned from the storage unit 32 (S11). Here, the data to be learned is color image data including at least the top and back surfaces of containers observed by the camera 10 in a specific container yard within a predetermined period (see, for example, Figure 6). In addition, the image data is annotated in a predetermined manner and also serves as training data.
[0033] With this configuration, the model is trained based on container images observed in a container yard within a predetermined period, resulting in a relatively larger proportion of data frequently observed in that container yard. This improves the accuracy of identifying and recognizing frequently observed strings and number sequences.
[0034] Furthermore, the image data is in color. With this configuration, the background color of the container, the color of the text, and the contrast between the text and background colors also become features during training, allowing for more accurate information identification. Additionally, the rectangle surrounding the check digit and the fact that the fourth character of the container code is "U" also become features during training.
[0035] After the readout process, the learning processing unit 311 first performs a learning process (or training process) for a region identification model that identifies regions from the image data that satisfy predetermined conditions, specifically, a first region displaying a container code consisting of four alphabet characters, and a second region containing a six-digit serial number and a one-digit check digit (CD) through object detection (S12).
[0036] With this configuration, machine learning techniques are used to identify image regions related to alphabetic strings and image regions related to digit sequences. Then, character or digit recognition processing is performed on each image region. By performing recognition processing only on the identified image regions, recognition can be performed efficiently and accurately.
[0037] In this embodiment, the learning process is performed using YOLO (You Only Look Once), which performs object detection using a deep learning model. By repeatedly inputting image data and training the model based on the difference between the model output and annotation data, a trained region identification model can be generated that identifies a first region containing a container code and a second region containing a sequence of numbers through object detection. The output of the region identification model is tag information indicating, for example, the position and dimensions of an object (or region), the object name, and the probability. The format of the tag information is not limited; for example, it may be raw data or output as text data such as JSON or XML.
[0038] With this configuration, YOLO can achieve high-speed processing.
[0039] Furthermore, with this configuration, object detection technology can reliably identify areas related to strings of characters or numbers.
[0040] After training the region identification model, the training processing unit 311 performs training on the alphabet recognition model (S13). This alphabet recognition model allows for the recognition of each letter of the alphabet contained within a first region that displays a container code consisting of four letters of the alphabet, through object detection.
[0041] With this configuration, machine learning techniques can be used to identify alphabetic strings with high accuracy and robustness.
[0042] In this embodiment, the learning process is performed by YOLO (You Only Look Once), which uses a deep learning model for object detection. By repeatedly inputting image data of a rectangular region corresponding to the first region and training the model based on the difference between the model output and annotation data indicating the alphabet, a trained alphabet recognition model that can identify each of the four letters of the alphabet can be generated. The output of the alphabet recognition model is, for example, tag information indicating the position and dimensions of the object (or alphabet), the object name, and the probability. The format of the tag information is not limited; for example, it may be raw data or output as text data such as JSON or XML.
[0043] With this configuration, YOLO can achieve high-speed processing.
[0044] Furthermore, with this configuration, the region containing characters can be reliably identified using object detection technology.
[0045] After training the alphabet recognition model, the training processing unit 311 performs training on the digit recognition model (S15). This digit recognition model allows the system to recognize each digit from a second area displaying a 6-digit serial number and a 1-digit check digit (CD) by object detection.
[0046] With this configuration, machine learning techniques can be used to identify sequences of numbers with high accuracy and robustness.
[0047] In this embodiment, the learning process is performed by YOLO (You Only Look Once), which uses a deep learning model for object detection. By repeatedly inputting image data of a rectangular region corresponding to the second region and training the model based on the difference between the model output and annotation data indicating the numbers, a trained digit recognition model that can identify each of the seven digits can be generated. The output of the digit recognition model is, for example, tag information indicating the position and dimensions of the object (or number), the object name, and the probability. The format of the tag information is not limited; for example, it may be raw data or output as text data such as JSON or XML.
[0048] With this configuration, YOLO can achieve high-speed processing.
[0049] Furthermore, with this configuration, object detection technology can reliably identify areas related to numbers. In particular, it can reliably identify check digits that are displayed with a frame.
[0050] Furthermore, with the above configuration, the recognition means can be divided into one suitable for recognizing alphabetic strings and another suitable for recognizing number sequences, thereby improving the recognition accuracy for characters that are easily confused between alphabets and numbers (for example, the number "1" and the alphabet "I").
[0051] After the training process for the digit recognition model is complete, the training processing unit 311 stores all the generated trained models in the storage unit 32 (S16), and the training process is completed.
[0052] It should be noted that the learning methods shown in this embodiment are merely examples, and other known machine learning methods can be used.
[0053] (1.2.2 Information Recognition Processing) Next, we will explain the information recognition process performed on the image data acquired by camera 10.
[0054] Figure 5 is a detailed flowchart of the information recognition process. As is clear from the figure, when the process starts, the image acquisition unit 312 reads out the image data that has been captured by the camera 10 and stored in the storage unit 32, which is the subject of the information recognition process (S21). In this embodiment, this image data is in color.
[0055] Figure 6 is an explanatory diagram showing an example of the target image data. This image was obtained by photographing a container 15, which was placed on a container transport vehicle 13 during transport to a container yard, from slightly above and behind the container 15 using a camera 10. Therefore, the image includes the back and top surfaces of the container 15.
[0056] The container number is displayed on the rear edge of the top surface and the upper right corner of the back surface of container 15. As is clear from the figure, the container number consists of a string of four uppercase letters (area 151 in the figure), a six-digit serial number, and a seven-digit number sequence consisting of a one-digit check digit (area 152 in the figure). The check digit is generally enclosed in a rectangle (square) as shown in the figure. The container number on the top surface is arranged in a single horizontal row, while the container number on the back surface is displayed in two rows, one above the other.
[0057] In the example shown in the figure, a four-digit code 153 is located directly below the sequence of numbers on the back (area 152). Of these four digits, the first two represent the size code indicating the dimensions of the container, and the latter two represent the type code indicating information about the goods loaded into the container.
[0058] Furthermore, a weight information display area 155 is provided directly below code 153. This area displays "MAX GROSS," which indicates the maximum load capacity of the container; "TARE," which indicates the weight of the container itself; and "NET," which indicates the actual capacity of the container, along with their corresponding numerical values and units.
[0059] Note that while Figure 6 shows an example where the container number is displayed on the back of the container in a single horizontal line, in reality, container numbers are displayed in various ways. For example, it should be noted that some containers may display the container number in multiple lines. Figure 7 shows an example where the container numbers (151,152) are displayed in two lines, one above the other. In this example, the serial number is divided into the first four digits and the last two digits.
[0060] Returning to Figure 5, once the target image data reading process is complete, the region identification unit 313 reads the learned region identification model from the storage unit 32 and uses it to identify a predetermined region from the target image data (S22). In this embodiment, the predetermined region includes a first region displaying the container code and a second region containing a 7-digit number. Note that both the container number displayed on the back of the container and the container number displayed on the top may be targeted, or only one of them may be targeted.
[0061] Figure 8 is an explanatory diagram showing examples of the first and second regions identified in the target image. Figure (a) is an example of the first region containing a four-letter alphabet container code, and Figure (b) is an example of the second region containing a six-digit serial number and a one-digit check digit (CD).
[0062] As is clear from Figure (a), a rectangular region containing a four-letter alphabetical container code is identified as the first region. In this embodiment, the first region is output as tag information for the image data by inputting the target image data into a trained region identification model. The tag information is text data that includes the location and dimensions of the first region, the region name, and the probability. For example, the tag information identifies a rectangular region with A pixels vertically and B pixels horizontally at a predetermined coordinate position ((X,Y)) in the image as the first region (region name). The tag information also includes probability information, such as the probability that the rectangular region is the first region being Z%.
[0063] As is clear from Figure (b), a rectangular area containing a 6-digit serial number and a 1-digit check digit is identified as the second area. Similarly, in this embodiment, the second area is also output as tag information for the image data by inputting the target image data into a trained area identification model. The tag information is text data that includes the location and dimensions of the second area, the area name, and the probability.
[0064] With this configuration, machine learning techniques are used to identify image regions related to alphabetic strings and image regions related to digit sequences. Then, character or digit recognition processing is performed on each image region. By performing recognition processing only on the identified image regions, recognition can be performed efficiently and accurately.
[0065] Returning to Figure 5, once the region identification process is complete, the object determination unit 315 determines the type of each region to be referenced based on the object name (region name) identified as tag information by the region identification unit 313 (S23). If the region name of the referenced region is determined to be the first region (S23 "alphabet"), the region identification unit 313 provides information about the first region to the alphabet recognition unit 316. On the other hand, if the region name of the referenced region is determined to be the second region (S23 "digit"), the region identification unit 313 provides information about the second region to the digit recognition unit 317.
[0066] With this configuration, the recognition means is switched according to tag information including object names, so that recognition processing can be performed using the most optimal method for each region. For example, for a first image region containing an alphabet string, a recognition means suitable for recognizing alphabets can be used, and for a second image region containing a sequence of numbers, a recognition means suitable for recognizing numbers can be used, thereby optimizing recognition.
[0067] The alphabet recognition unit 316 performs alphabet recognition processing on the provided first area (S25). More specifically, the alphabet recognition unit 316 reads a trained alphabet recognition model generated by the learning process from the storage unit 32, applies it to the first area, and performs the process of recognizing the alphabet one character at a time.
[0068] In this embodiment, the recognition result is output as tag information for the first region by inputting image data corresponding to the first region into a trained alphabet recognition model. The tag information is text data that includes the position and dimensions of each alphabet, the object type (A, B, C, etc.), and the probability.
[0069] Figure 9 is an explanatory diagram showing an example of the output results of an alphabet recognition model. For explanatory purposes, information corresponding to tag information is shown in the upper right corner of the rectangle surrounding each character in the figure. As is clear from the four rectangles in the figure, the alphabet recognition model identifies the four letters contained in the first region as "S", "K", "L", and "U". Furthermore, the recognition probability for each character is "100%" for "S", "100%" for "K", "100%" for "L", and "100%" for "U".
[0070] With this configuration, machine learning techniques can be used to identify alphabetic strings with high accuracy and robustness.
[0071] Returning to Figure 5, the digit recognition unit 317 performs digit recognition processing on the provided second area in parallel with the alphabet recognition processing (S26). More specifically, the digit recognition unit 317 reads the trained digit recognition model generated by the learning process from the memory unit 32, applies it to the second area, and performs the process of recognizing the digits one by one.
[0072] In this embodiment, the recognition result is output as tag information for the second region by inputting image data corresponding to the second region into a trained digit recognition model. The tag information is text data that includes the position and dimensions of each digit, the object type (1, 2, 3...), the probability, etc.
[0073] Figure 10 is an explanatory diagram showing an example of the output of a digit recognition model. For illustrative purposes, in the figure, information corresponding to tag information is shown in the upper right corner of the rectangle surrounding each digit. As is clear from the seven rectangles in the figure, the digit recognition model has identified the seven letters contained in the second region as "1", "6", "0", "7", "2", "3", and "8", respectively. In this example, each digit is recognized with a probability of "100%".
[0074] With this configuration, machine learning techniques can be used to identify sequences of numbers with high accuracy and robustness.
[0075] Returning to Figure 5, once the alphabet recognition process and the digit recognition process are complete, the output unit 318 determines whether the recognition of either target was successful. This success is determined by whether or not predetermined conditions are met. Here, the predetermined conditions are, for example, in this embodiment, that in the first region, all four alphabets are recognized with a predetermined probability or higher, and in the second region, all seven digits are recognized with a predetermined probability or higher.
[0076] If it is determined that all recognition processes have been successful (S27YES), the output unit 318 outputs the recognized four-letter alphabet container code, a six-digit serial number, and a one-digit check digit (S28), and the process ends. The output format is not limited to a specific method. Therefore, it may be output to the display device 361, speaker 371, etc., or it may be output in a format that is stored in the storage unit 32, etc., as some kind of file or data. It may also be output to an external device, such as a terminal management device, via the communication unit 33.
[0077] On the other hand, if the above-mentioned predetermined conditions are not met and the recognition of the target fails (S27NO), the estimation unit 319 performs estimation processing for the target that was not successfully recognized (S31).
[0078] Here, according to ISO 6346, the check digit is calculated as follows. Figure 11 is the alphabet and number assignment table defined in the ISO standard. First, according to the table, a number from 0 to 38 is assigned to each digit, either an alphabet or a number. For each of these assigned numbers, 2 0 ~2 9 Multiply the numbers in order and get the sum of them. The remainder when this sum is divided by 11 is the check digit. 0 ~2 9 Instead, 2 0 ~2 9 You can also multiply each of them by the remainder when divided by 11.
[0079] Figure 12 is an explanatory diagram illustrating an example of check digit calculation. In this example, the container code is "NYKU" and the serial number is "484074". In this case, the four letters "NYKU" are each assigned the numbers shown in Figure 10, namely "25, 37, 21, and 32". For each of these numbers, 2 0 ~2 9Each of these numbers is divided by 11, and the remainders are multiplied in order to calculate their sum (=25×1+37×2+21×4+32×8+4×5+8×10+4×9+0×7+7×3+4×6). Dividing this sum by 11 yields the check digit, 4. This embodiment utilizes this process.
[0080] As an example, let's consider a case where the 6-digit serial number and 1-digit check digit are successfully recognized, but the container code (or owner code) is not. In this case, the estimation unit 319 performs the above calculation on the 6-digit serial number and calculates a provisional check digit. Then, the estimation unit 319 subtracts this provisional check digit from the recognized correct check digit. If the value is negative, 11 is added. The value obtained by subtracting the provisional check digit calculated from the serial number from the correct check digit becomes the provisional check digit of the container code.
[0081] After calculating a provisional check digit for the container code, the estimation unit 319 refers to the provisional check digit database stored in the storage unit 32 and performs a process to identify a single container code based on the provisional check digit.
[0082] Figure 13 is an explanatory diagram illustrating an example of a provisional check digit database. As is clear from the figure, in this database (or table), a pre-calculated provisional check digit for a container code is associated with each four-letter alphabetical container code. In the example shown in the figure, each container code is sorted in descending order of frequency of occurrence. Here, frequency of occurrence refers to the number of times a given yard has been observed within a certain period.
[0083] The estimation unit 319 uses the provisional check digit obtained by calculation as a key to search and identify a container code in the database that matches the provisional check digit. If two or more container codes are candidates and at least some characters of the container code are recognized (especially characters other than the fourth letter "U"), the unit selects the container code with the highest character match rate.
[0084] With this configuration, even if part of the container code is unrecognizable, the check digit can be used to estimate the most likely container code (or owner code). This allows for more stable recognition of container information.
[0085] Furthermore, when two or more container codes are candidates based on the provisional check digit, the method for selecting one of them is not limited to the method described above. For example, if no characters corresponding to any of the container codes can be recognized, the estimation unit 319 may refer to the provisional check digit database and select the one with the highest frequency of occurrence as the container code.
[0086] With this configuration, even if the entire container code cannot be recognized, the container code can be estimated using a check digit. This allows for more stable recognition of container information. Furthermore, by using a method that utilizes the frequency of occurrence, it is possible to estimate container codes that are frequently observed and have a high probability of being correct.
[0087] Alternatively, the recognized container code may be verified using a 6-digit serial number and a 1-digit check digit.
[0088] In this case, the estimation unit 319 performs the above calculation on the 6-digit serial number and calculates a provisional check digit. Then, it subtracts this check digit from the recognized correct check digit. If the value is negative, it adds 11. The value obtained by subtracting the provisional check digit calculated from the serial number from the correct check digit becomes the provisional check digit of the container code.
[0089] After calculating a provisional check digit for the container code, the estimation unit 319 refers to the provisional check digit database stored in the storage unit 32 and performs a process to identify a single container code based on the provisional check digit.
[0090] Here, the recognized container code is compared with the container code identified through the above process. If the two do not match, it is possible that recognition of some characters in the container code has failed.
[0091] In this case, the estimation unit 319 may refer to the character conversion table from the storage unit 32 and correct the characters recognized in the container code. Here, in this embodiment, the character conversion table includes pairs of characters that are prone to misrecognition, for example, pairs of characters that are prone to misrecognition, such as "E" and "F," are stored. After this correction process, a comparison is made with the identified container code again to confirm that the two match, and the final container code may be output.
[0092] With this configuration, the correct alphabetical string can be quickly identified by replacing characters that are easily misrecognized.
[0093] Returning to Figure 5, after the estimation process estimates the characters related to the container code, the digits related to the serial number, and / or the digits related to the check digit, the output unit 318 performs the process of outputting the four characters related to the container code, the six digits related to the serial number, and the one digit check digit that were finally obtained (S32). Note that the output format is not limited to a specific method. Therefore, it may be output to the display device 361, speaker 371, etc., or it may be output in a format that is stored in the storage unit 32, etc., as some kind of file or data. It may also be output to an external device, such as a terminal management device, via the communication unit 33.
[0094] After this output processing, the target image data is stored in the storage unit 32 as data to be used for further training of each model (S33).
[0095] With this configuration, the accuracy of region identification can be further improved by using additional image data from cases where the identification or recognition of an image region fails for further training.
[0096] After the data is stored for further training, the process ends.
[0097] With the above configuration, alphanumeric strings containing owner codes and numerical strings containing serial numbers, displayed on a container in various forms (e.g., various layouts (position, arrangement, size, height / width, range, etc.), font, color, clarity, etc.), can be recognized with high accuracy and robustness using generalization techniques based on machine learning. This enables stable recognition of the display on the container from images or videos.
[0098] (2. Variant) The present invention can be implemented in various modified forms.
[0099] Although the above-described embodiment employs a configuration in which separate models are provided for alphabet recognition and digit recognition, the present invention is not limited to such a configuration. Therefore, instead, a single recognition model capable of recognizing both letters and digits may be provided. Furthermore, although the above-described embodiment employs a configuration in which a character or digit recognition model is provided after the region identification model, the present invention is not limited to such a configuration. Therefore, for example, the preceding region identification model and the subsequent recognition model may be integrated into a single machine learning model that recognizes letters or digits all at once.
[0100] This configuration allows for a simpler structure and faster processing.
[0101] Although the above-described embodiment shows an example where characters or numbers are displayed horizontally (horizontally) as annotated image data (e.g., Figure 6), the present invention is not limited to such configurations. Therefore, the training data may include examples where characters or numbers are displayed vertically (vertically) as annotated image data. It may also include examples where characters or numbers are displayed diagonally, for example, because the door on the back of the container is open. Furthermore, it may include image data with irregular displays, such as an owner code displayed below the serial number. In addition, it may include image data where some or all of the characters or numbers are missing or lack clarity. Moreover, it may include images of various types of containers, such as flat rack containers and tank containers.
[0102] This configuration makes it less susceptible to variations in the orientation, layout, and clarity of displayed text and number sequences, enabling robust recognition.
[0103] Although the above-described embodiment uses the container number as the target of recognition, the present invention is not limited to such a configuration. Therefore, the system may be configured to recognize other characters, numbers, symbols, or figures displayed on the container. For example, supplementary information such as the IMO label may be used as the target of recognition.
[0104] With this configuration, additional information such as container type can also be recognized, allowing even more information to be obtained from the container image.
[0105] Although the above-described embodiment illustrates a configuration in which one camera 10 is provided for the parking area of one container transport vehicle 13 in the yard, the present invention is not limited to such a configuration. Therefore, for example, multiple cameras may be installed and recognition processing may be performed on each image, or one camera 10 may be used to monitor multiple parking areas and recognize multiple container numbers, etc.
[0106] In the embodiment described above, the camera 10 is installed at the upper rear of the container 15 to photograph both the top and back surfaces of the container 15, but the present invention is not limited to such a configuration. For example, the camera 10 may be installed to photograph either one of them.
[0107] Although embodiments of the present invention have been described above, these embodiments represent only a part of the application examples of the present invention, and are not intended to limit the technical scope of the present invention to the specific configurations of the above embodiments. Furthermore, the above embodiments can be combined as appropriate without creating any contradictions. [Industrial applicability]
[0108] This invention can be used in industries that manufacture image recognition devices and the like. [Explanation of Symbols]
[0109] 10 Cameras 13 Container transport vehicles 15 containers 20 storage 30 Information Processing Devices 100 Container Information Recognition System
Claims
1. An image acquisition unit acquires a container image which is an image containing a container that displays an alphanumeric string including an owner code consisting of three letters and a numeric string including a serial number consisting of six numbers. An information recognition unit that uses machine learning technology to recognize the alphabet string and the number string from the container image, A container information recognition device equipped with the following features.
2. The aforementioned information recognition unit further, A region identification unit identifies a first image region containing the alphabet string and a second image region containing the digit sequence within the container image, based on the container image and a first trained model. A region information recognition unit that recognizes the alphabet string from the first image region and the number sequence from the second image region, A container information recognition device according to claim 1, comprising:
3. The container information recognition device according to claim 2, wherein the first trained model has an object detection function, and identifies the first image region and the second image region in the container image based on the object detection function.
4. The container information recognition device according to claim 3, wherein the first trained model identifies the first image region and the second image region by generating tag information including the type, position, and dimensions of an object.
5. The container information recognition device according to claim 4, wherein the region information recognition unit switches the recognition means applied to the first image region or the second image region according to the tag information.
6. The container information recognition device according to claim 5, wherein the recognition means includes a first recognition means suitable for recognizing an alphabet string and a second recognition means suitable for recognizing a sequence of numbers.
7. The first recognition means is a means using a second trained model suitable for recognizing the alphabet string, The container information recognition device according to claim 6, wherein the second recognition means is a means using a third trained model suitable for recognizing the sequence of digits.
8. The second trained model has an object detection function, and based on the object detection function, recognizes each character of the alphabet string in the first image region. The container information recognition device according to claim 7, wherein the third trained model has an object detection function and recognizes each digit in the sequence of digits in the second image region based on the object detection function.
9. The container information recognition device according to claim 8, wherein the first trained model, the second trained model, and the third trained model are generated by YOLO.
10. The container information recognition device according to claim 7, further comprising an additional training data storage unit that stores the container image as additional training data for the first trained model, the second trained model, and / or the third trained model if it fails to identify the first image region and / or the second image region, or if it fails to recognize the alphabet string and / or the number string.
11. The container information recognition device according to claim 1, wherein the alphabet string is a four-letter alphabet, and the number string is a seven-digit number string followed by a one-digit check digit at the end of the serial number.
12. The container information recognition device according to claim 1, wherein the container image is a color image.
13. The container information recognition device according to claim 8, wherein the first trained model, the second trained model, and / or the third trained model are trained based on a set of observed images including containers observed in a container yard within a predetermined period of time.
14. The container information recognition device according to claim 13, wherein the group of observed images includes an image containing a container in which the alphabet string and the number string are each written vertically, and an image containing a container in which the alphabet string and the number string are each written horizontally.
15. The aforementioned container also displays additional information, The container information recognition device according to claim 1, wherein the information recognition unit further recognizes the associated information from the container image.
16. The aforementioned alphabet string is a four-letter alphabet string, with the owner code followed by a single letter indicating a category identifier. The aforementioned sequence of numbers is a seven-digit sequence of numbers, with a single check digit following the end of the serial number. The model used in the aforementioned machine learning technique is trained based on a set of observed images, including containers, observed in a container yard within a predetermined period. The container information recognition device further, A provisional check digit calculation unit calculates a first provisional check digit based on the recognized serial number and the check digit, A first table storage unit stores a table in which the aforementioned alphabet string and a second provisional check digit calculated based on the aforementioned alphabet string are associated, A container information recognition device according to claim 1, comprising: a code estimation unit that refers to the table and identifies an alphabet string in which the first provisional check digit and the second provisional check digit match.
17. The table storage unit further includes observation frequency information of the alphabet string in the observed image group, The code estimation unit, The container information recognition device according to claim 16, which identifies the alphabet string with the highest observation frequency among the alphabet strings in which the first provisional check digit and the second provisional check digit match.
18. The code estimation unit, The container information recognition device according to claim 16, which identifies the alphabet string with the highest degree of match with the recognized alphabet string among the alphabet strings in which the first provisional check digit and the second provisional check digit match.
19. It further includes a misrecognition table storage unit that stores a misrecognition table, which is a table of letters that are easily misrecognized. The code estimation unit, The container information recognition device according to claim 16, further comprising a character replacement unit that, if there is no second provisional check digit that matches the first provisional check digit, refers to the misrecognition table and replaces part or all of the alphabet string.
20. An image acquisition unit acquires a container image which is an image containing a container that displays an alphanumeric string including an owner code consisting of three letters and a numeric string including a serial number consisting of six numbers. An information recognition unit that uses machine learning technology to recognize the alphabet string and the number string from the container image, A container information recognition system equipped with the following features.
21. Image acquisition step: Obtain a container image which is an image containing a container that displays an alphanumeric string containing an owner code consisting of three letters and a numeric string containing a serial number consisting of six numbers. An information recognition step in which machine learning technology is used to recognize the alphabet string and the number string from the container image, A container information recognition method comprising the following features.
22. On the computer, Image acquisition step: Obtain a container image which is an image containing a container that displays an alphanumeric string containing an owner code consisting of three letters and a numeric string containing a serial number consisting of six numbers. An information recognition step in which machine learning technology is used to recognize the alphabet string and the number string from the container image, A container information recognition program that executes the command.
Citation Information
Patent Citations
Container number recognition system
JP2021096864A