Terminal, system, method, and program for counting stacked objects

The terminal and method improve object counting accuracy by using a colored margin to highlight edges and boundaries, addressing miscounting issues with varying shapes and thicknesses in stacked objects.

JP2025135076AActive Publication Date: 2025-09-18AYATAKA SYST DEV CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2024032656
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-03-05
Publication Date
2025-09-18
Estimated Expiration
2044-03-05

Smart Images

  • Figure 2025135076000001_ABST
    Figure 2025135076000001_ABST
Patent Text Reader

Abstract

To count objects stacked one upon another by detecting the objects through image recognition, even when shapes of the stacked objects vary vertically or apparent thicknesses of the objects differ.SOLUTION: A terminal for counting stacked objects acquires an image including the whole image of the stacked objects, extracts a partial image of the stacked objects including the upper edge portion and the lower edge portion of the whole image of the stacked objects from the image, creates a recognition image in which a margin portion of a color different from the colors of the stacked objects is provided around the partial image; detects the upper edge portion and the lower edge portion of the whole image of the stacked objects and an image of the boundaries among the objects in the recognition image, and counts the stacked objects on the basis of the detected images of the upper edge portion, the lower edge portion, and the boundaries.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a terminal, a system, a method and a program for counting piled objects. [Background technology]

[0002] In a wide variety of industries and situations, it is common to count stacked objects such as boards, food containers, and sheets, but when counting by hand, mistakes such as miscounting can occur. In particular, stacked objects often have simple shapes, making counting mistakes more frequent.

[0003] For this reason, the number of such objects is sometimes calculated and managed based on the thickness of each object and the overall height, but soft objects such as sheets will appear to have different thicknesses when stacked.

[0004] For this reason, in recent years, objects have been counted from images of the objects to be counted by image recognition using a learning model based on AI (Artificial Interigence).For example, in an image of stacked objects, each object in the image is detected and counted (Non-Patent Document 1), or in an image of objects contained in a container, the outer frame of the container, the outer edge of the entire object, and even the density of the wet towels are detected to estimate the number of contents (Patent Document 1).

[0005] Incidentally, in order to count objects in an image using an AI learning model, it is necessary to annotate the image data of objects detected by image recognition and use the annotated data as training data for machine learning. There are several annotation methods, but among them, bounding box annotation, which surrounds the object to be annotated in an image with a rectangle, is relatively simple compared to other methods and is therefore used in many situations. [Prior art documents] [Patent documents]

[0006] [Patent Document 1] Patent No. 7004375 specification [Non-patent literature]

[0007] [Non-Patent Document 1] Skylogic Inc., "AI counting app "cazoeTell"", [online], November 21, 2022, Kadokawa ASCII Research Institute, [Retrieved January 21, 2024], Internet<URL:https: / / ascii.jp / elem / 000 / 004 / 114 / 4114025 / > Summary of the Invention [Problem to be solved by the invention]

[0008] However, when applying bounding box annotation to the technologies of Non-Patent Document 1 and Patent Document 1 to count stacked objects with different apparent thicknesses, the shapes of the objects move up and down along the way, making it difficult to extract feature points of individual objects and leading to false detection of objects.In addition, it is difficult to enclose objects with rectangles during annotation, making it difficult to create appropriate training data.

[0009] Therefore, the inventors noticed that by extracting a portion of an image capturing the entire stack of objects and surrounding the extracted image with white space of a different color than the stack of objects, the characteristics of the boundaries between each object become more prominent, making it possible to reliably extract the characteristic points of the boundaries of each object that are necessary when counting the stack of objects.Therefore, by using highly accurate bounding box annotation, it becomes possible to accurately count even soft objects whose shapes vary up and down or whose apparent thicknesses vary.

[0010] In view of these problems, the present invention aims to provide a terminal, system, method and program for counting stacked objects, which can detect objects using image recognition and count objects even if the shapes of the stacked objects are up or down or the apparent thicknesses are different. [Means for solving the problem]

[0011] The present invention provides the following solutions.

[0012] A first aspect of the present invention is a terminal for counting piled objects, comprising: an acquisition unit that acquires an image including an overall image of the piled objects; a creating unit that extracts a partial image of the piled objects, including an upper edge portion and a lower edge portion of the overall image of the piled objects, from the image, and creates an image for recognition in which a marginal portion of a different color from the piled objects is provided around the partial image; a detection unit that detects, in the recognition image, an image of the upper edge portion and the lower edge portion of the overall image of the piled up objects and an image of a boundary between each object; a counting unit that counts the piled objects based on the detected images of the upper edge portion, the lower edge portion, and the boundary; It is a terminal equipped with.

[0013] According to a first feature of the present invention, by extracting a portion of an image capturing the entire image of stacked objects and surrounding the extracted image with a margin of a different color from the stacked objects, the features of the upper and lower edges of the entire image of the stacked objects and the boundaries between each object become more prominent, and it is possible to reliably extract the feature points of the boundaries between each object that are necessary when counting stacked objects, making it possible to accurately count even soft objects whose shapes are up and down or that appear to be of different thicknesses.

[0014] The second feature of the present invention is the invention according to the first feature, the detection unit detects an image of each of the piled objects in the recognition image; The counting unit provides a terminal that counts the piled objects based on an image of each of the detected objects.

[0015] According to the second feature of the present invention, by detecting each object itself, even if the upper or lower edges of the stacked objects or the boundaries between the objects cannot be detected, by counting each object itself, it is possible to accurately count even soft objects whose shapes are up or down when stacked or whose apparent thicknesses vary.

[0016] The third feature of the present invention is the invention according to the first feature, The terminal further includes an output unit that outputs the counting result.

[0017] According to the third aspect of the present invention, the count result of the piled objects is output to the terminal, so that the user can immediately grasp the count number of the piled objects and the like.

[0018] The fourth feature of the present invention is the invention according to the first feature, a learning unit that learns the recognition image and images of the upper edge portion, the lower edge portion, and the boundary to create a learned model; an estimation unit that estimates images of the upper and lower edges of the stacked objects and the boundaries between the objects in the recognition image based on the trained model, The counting unit provides a terminal that counts the piled objects based on images of the estimated upper edge, lower edge, and boundary.

[0019] According to a fourth feature of the present invention, by learning the recognition image and images of the upper and lower edges and boundaries to create a trained model, it is possible to improve the recognition accuracy of the upper and lower edges of stacked objects and the boundaries between each object, thereby making it possible to improve the counting accuracy of soft objects whose stacked shapes vary up and down or whose apparent thicknesses are different.

[0020] A fifth feature of the present invention is the invention according to the fourth feature, the learning unit further learns an image of each of the objects; the estimation unit estimates an image of each of the stacked objects in the recognition image based on the trained model; The counting unit provides a terminal that counts the piled objects based on the estimated image of each object.

[0021] According to the fifth feature of the present invention, the recognition accuracy of stacked objects can be improved by learning the recognition images and images of each object to create a trained model. Therefore, even if the upper or lower edges of the stacked objects or the boundaries between each object cannot be recognized, by counting each object itself, it is possible to improve the counting accuracy of soft objects whose stacked shapes are up or down or whose apparent thicknesses vary.

[0022] The sixth feature of the present invention is the invention according to the fourth feature, The learning unit provides a terminal that learns using bounding box annotation that includes a part of the marginal area.

[0023] According to a sixth feature of the present invention, even if the shapes of stacked objects are upside down or the apparent thicknesses are different, bounding box annotation makes it possible to simply tag the information on the image to the image data to be learned, learn, create a learning model, and maintain the accuracy of automatic object counting.

[0024] Although the present invention is in the category of terminals, it also exerts similar functions and effects according to other categories such as systems. [Effects of the Invention]

[0025] According to the present invention, it is possible to provide a terminal, system, method and program for counting stacked objects, which can detect objects using image recognition and can count soft objects that have different apparent thicknesses even if the stacked objects have different shapes or appear to be upside down. [Brief explanation of the drawings]

[0026] [Figure 1] FIG. 1 is a diagram illustrating an overview of a terminal 1 for counting objects according to a first embodiment of the present invention. [Figure 2] 1 is a configuration diagram of a terminal 1 for counting objects according to the present embodiment. [Figure 3] 10 is a flowchart of an object counting process executed by a terminal 1 for counting objects according to the present embodiment. [Figure 4] 1 is a diagram showing an example of an image 100 obtained by capturing an overall image 110 of piled objects, which is acquired by the terminal 1 in the object counting process of this embodiment. [Figure 5] 10 is a diagram for explaining a recognition image 200 created by the terminal 1 in the object counting process of the present embodiment. FIG. [Figure 6] 10A and 10B are diagrams for explaining a detection process in an image for recognition 200 executed by a terminal 1 for counting objects according to the present embodiment. [Figure 7] FIG. 10 is a diagram illustrating an overview of a terminal 1 for counting objects according to a second embodiment of the present invention. [Figure 8] 10 is a flowchart of an object counting process executed by a terminal 1 for counting objects according to a second embodiment of the present invention. [Figure 9] 10 is a flowchart of a learning process executed by the object counting terminal 1 of the present embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0027] The best mode for carrying out the present invention will be described below with reference to the drawings. However, this is merely an example, and the technical scope of the present invention is not limited to this example.

[0028] [Overview of Terminal 1 for counting objects] An overview of a terminal 1 for counting objects according to a first embodiment of the present invention will be described with reference to Fig. 1. Fig. 1 is a diagram for explaining the overview of a terminal 1 for counting objects according to a first embodiment of the present invention.

[0029] The terminal 1 for counting objects is, for example, a mobile terminal such as a handheld terminal, smartphone, or tablet terminal, or a wearable terminal such as a head-mounted display such as smart glasses or a smart watch, and is equipped with an imaging device such as a camera that captures images such as color video and / or still images.

[0030] The terminal 1 may be realized, for example, by one terminal device, or may be realized by a plurality of terminal devices, or may be realized by a virtual device such as a cloud computer.

[0031] Next, an outline of the process executed by the terminal 1 to count objects will be described. First, the terminal 1 acquires an image 100 including an overall image 110 of the piled objects (step S1). Specifically, the terminal 1 acquires an image 100 of the overall image 110 of the piled objects to be counted, which is taken by the user using the terminal 1.

[0032] Next, the terminal 1 extracts a partial image 210 from the image 100, including the upper edge 120 and the lower edge 130 of the overall image 110 of the stacked objects, and creates a recognition image 200 by surrounding the partial image 210 with a margin 220 of a different color from the stacked objects (step S2). Specifically, the terminal 1 extracts a portion of the image 100 acquired in step S1 as the partial image 210 in a direction that includes the upper edge 120 and the lower edge 130 of the overall image 110 of the stacked objects, and then combines the extracted partial image 210 with the partial image 210 of a different color from the stacked objects, by surrounding the margin 220. The color of the margin 220 may be any color as long as it is different from the color of the stacked objects, but is preferably an opposite color (complementary color) on the color wheel. The terminal 1 may accept a selection input from the user regarding the color of the margin 220. In this case, if the accepted color is the same color or a similar color as the color of the piled up object, the terminal 1 may not accept the selection input, and in this case may output an error message, etc. Also, the size of the margin 220 is not particularly important, but if the margin 220 is too large, the proportion of the piled up object in the recognition image 200 becomes relatively small, so it is preferable that the size of the margin 220 is 20 to 30 percent of the width of the extracted partial image 210.

[0033] Next, terminal 1 detects, in recognition image 200, images of upper edge 120 and lower edge 130 of overall image 110 of the piled objects and images of boundaries 230 between each object 240 (step S3). Specifically, terminal 1 detects image-specific features of upper edge 120 and lower edge 130 of overall image 110 of the piled objects shown in recognition image 200 and boundaries 230 that distinguish each object 240. The method of detecting upper edge 120, lower edge 130, and each boundary 230 is not particularly limited.

[0034] In step S3, the terminal 1 may re-learn, by machine learning, the image-specific feature amounts of the detected upper edge portion 120, lower edge portion 130, and each border 230, thereby updating the trained model 300. Note that this re-learning by machine learning may be performed at any time after step S3, or may be performed by the terminal 1 receiving input from the user. Machine learning will be described later.

[0035] Next, the terminal 1 counts the piled objects based on the detected images of the upper edge 120, the lower edge 130, and the boundary 230 (step S4). Specifically, the terminal 1 counts the number of detected images of the boundary 230 and corrects it by +1. Here, if the terminal 1 detects images of the upper edge 120 and the lower edge 130, it counts the number of images of the boundary 230 and corrects it by -1. Furthermore, if it detects images of either the upper edge 120 or the lower edge 130, it counts the number of images of the boundary 230 and does not correct it. In this way, the number of piled objects is calculated. The terminal 1 may output the count result and notify the user. The output method is not particularly limited, and may be output on its own display unit, may be output as audio through a speaker, or may be distributed to another terminal or device.

[0036] The above is an outline of the process executed by the terminal 1 to count objects.

[0037] [System configuration of Terminal 1 for counting objects] The system configuration of the terminal 1 for counting objects according to this embodiment will be described with reference to FIG.

[0038] The terminal 1 may be realized, for example, by one terminal device or by a plurality of terminal devices.

[0039] Terminal 1 is, for example, a mobile terminal such as a handheld terminal, smartphone, or tablet terminal, or a wearable terminal such as a head-mounted display such as smart glasses or a smart watch, and is equipped with an imaging device such as a camera that captures images such as color video and / or still images.

[0040] Terminal 1 is equipped with a control unit and a processing unit, such as a CPU (Central Processing Unit), GPU (Graphics Processing Unit), RAM (Random Access Memory), and ROM (Read Only Memory). The control unit issues execution commands to the processing unit, imaging unit, communication unit, input unit, output unit, and memory unit, which will be described later. The processing unit calculates data and determines the calculation results.

[0041] The imaging unit includes a device for capturing images such as moving images and / or still images.

[0042] The terminal 1 includes a device as a communication unit for enabling communication with other terminals, devices, etc. The communication method may be wireless or wired.

[0043] The terminal 1 is assumed to have, as an input unit, functions necessary for a user to operate the terminal 1. Examples of devices for realizing input include an LCD display that realizes a touch panel function, a keyboard, a mouse, a pen tablet, hardware buttons on the device, and a microphone for voice recognition. The present invention is not particularly limited in function depending on the input method.

[0044] The terminal 1 is assumed to have, as an output unit, functions necessary for the user to operate the terminal 1. Examples of output methods include display such as projection on a display unit such as an LCD display, a PC display, or a projector, and audio output. The present invention is not particularly limited in function by the output method.

[0045] The terminal 1 includes a storage unit for storing data such as a hard disk, semiconductor memory, recording medium, memory card, etc. The data may be stored in a cloud service, a database, etc.

[0046] The control unit cooperates with the processing unit to implement an acquisition unit 10, a creation unit 11, a detection unit 12, a counting unit 13, a learning unit 14, and an estimation unit 15.

[0047] The above is the system configuration of the terminal 1 for counting objects.

[0048] [Object counting process] The object counting process executed by the terminal 1 for counting objects will be described with reference to Fig. 3. Fig. 3 is a diagram showing a flowchart of the object counting process executed by the terminal 1 for counting objects according to this embodiment. As shown in Fig. 3, the object counting process is made up of steps S11 to S14, which correspond to the above-mentioned steps S1 to S4.

[0049] First, the acquisition unit 10 of the terminal 1 acquires an image 100 including an overall image 110 of the piled objects (step S11). Specifically, the acquisition unit 10 acquires an image 100 of the overall image 110 of the piled objects to be counted, which is taken by the user using the terminal 1. FIG. 4 shows an example of the image 100 of the overall image 110 of the piled objects acquired by the acquisition unit 10. The image 100 may be a monochrome image, but is preferably a color image. The image may be a moving image or a still image. The image 100 may also be acquired from another terminal, computer, or device via data communication or the like.

[0050] Next, the creation unit 11 of the terminal 1 extracts from the image 100 a partial image 210 including the upper edge 120 and the lower edge 130 of the overall image 110 of the stacked objects, and creates a recognition image 200 in which a margin 220 of a different color from the color of the stacked objects is provided around the partial image 210 (step S12). Specifically, as shown in FIG. 5, the creation unit 11 extracts a part of the image 100 acquired in step S11 as the partial image 210 in a direction including the upper edge 120 and the lower edge 130 of the overall image 110 of the stacked objects, and combines the extracted partial image 210 with the partial image 210 in a color different from the color of the stacked objects so that a margin 220 is provided around the partial image 210. The color of the margin 220 may be any color as long as it is different from the color of the stacked objects, but an opposite color (complementary color) on the color wheel is preferable. The input unit of the terminal 1 may receive a selection input from the user regarding the color of the marginal portion 220. In this case, if the received color is the same color or a similar color as the color of the stacked objects, the input unit may not receive the selection input, and in this case, an error message or the like may be output to the output unit. The size of the marginal portion 220 is not particularly important, but if the marginal portion 220 is too large, the proportion of the stacked objects in the recognition image 200 will become relatively small. Therefore, it is preferable that the size of the marginal portion 220 be 20 to 30 percent of the width of the extracted partial image 210.

[0051] Next, the detection unit 12 of the terminal 1 detects, in the recognition image 200, an image of the upper edge 120 and the lower edge 130 of the overall image 110 of the stacked objects and an image of the boundaries 230 between each object 240 (step S13). Specifically, as shown in FIG. 6, the detection unit 12 detects image-specific feature amounts of the upper edge 120 and the lower edge 130 of the overall image 110 of the stacked objects shown in the recognition image 200 and the boundaries 230 that distinguish each object 240, and extracts each boundary 230 by enclosing it in a rectangular shape so that it covers the entire boundary 230, that is, extracts a rectangular image including the upper edge 120, the lower edge 130, and each boundary 230. The method of detecting and extracting the upper edge 120, the lower edge 130, and each boundary 230 is not particularly limited. Furthermore, in this embodiment, the upper edge 120, the lower edge 130 and the borders 230 are surrounded by a rectangular shape, but they may be surrounded by any other shape.

[0052] In step S13, the detection unit 12 of the terminal 1 may further detect images of each object 240 in the recognition image 200. Specifically, similar to the detection of the upper edge 120, the lower edge 130, and each boundary 230 described above, as shown in FIG. 6, the detection unit 12 detects image-specific features of each object 240 appearing in the recognition image 200 and extracts each object 240 by surrounding it in a rectangular shape so that the entire object 240 is encompassed, that is, extracts a rectangular image including each object 240. The method of detecting and extracting each object 240 is not particularly limited. Furthermore, in this embodiment, each object 240 is surrounded by a rectangle, but it may be surrounded by any other shape.

[0053] Furthermore, in step S13, the learning unit 14 of the terminal 1 may re-learn, by machine learning, image-specific feature quantities of the detected upper edge portion 120, lower edge portion 130, and each border 230, thereby updating the trained model 300. Furthermore, when image-specific feature quantities of each object 240 are detected, the learning unit 14 may also re-learn, by machine learning, the image-specific feature quantities of each detected object 240, thereby updating the trained model 300. Note that this re-learning by machine learning may be performed at any timing after step S13, or may be performed by the input unit of the terminal 1 receiving input from the user. Machine learning will be described later.

[0054] Next, the counting unit 13 of the terminal 1 counts the piled objects based on the detected images of the upper edge 120, the lower edge 130, and the boundary 230 (step S14). Specifically, the terminal 1 counts the number of detected images of the boundary 230 and corrects it by +1. Here, if images of the upper edge 120 and the lower edge 130 are further detected, the terminal 1 counts the number of images of the boundary 230 and corrects it by -1. Furthermore, if images of either the upper edge 120 or the lower edge 130 are detected, the terminal 1 counts the number of images of the boundary 230 and does not correct it. In this way, the number of piled objects is calculated. The output unit of the terminal 1 may output the count result and notify the user. The output method is not particularly limited, and may be output to the terminal's own display unit, may be output as audio through a speaker, or may be distributed to another terminal or device.

[0055] In step S14, if an image of each object 240 is detected in step S13 described above, the counting unit 13 of the terminal 1 may count the piled objects based on the detected image of each object 240. Specifically, the counting unit 13 may calculate the number of each object 240 in the piled objects by counting the number of detected images of each object 240.

[0056] Furthermore, when an image of each object 240 is detected in the above-mentioned step S13, it is compared with the count result based on the detected images of the upper edge 120, lower edge 130, and border 230, and if the comparison result does not match, the output unit of the terminal 1 may notify the user by outputting a notification urging the user to re-acquire the image 100 of the overall image 110 of the piled objects. The output method is not particularly limited, and may be output on the display unit of the terminal 1 itself, or may be output as a sound from a speaker or the like. The notification may also be distributed from the communication unit of the terminal 1 to another terminal or device.

[0057] Furthermore, when an image of each object 240 is detected in the aforementioned step S13, the user may input from terminal 1 whether to count the stacked objects based on the image of the detected boundary 230, based on each detected object 240, or based on both.

[0058] This completes the object counting process.

[0059] According to the terminal 1 for counting objects, a portion of an image taken of the entire stack of objects is extracted, and a margin of a different color from the stacked objects is provided around the extracted image, which makes the features of the upper and lower edges of the entire image of the stacked objects and the boundaries between each object more prominent. This makes it possible to reliably extract the feature points of the boundaries of each object that are necessary when counting the stacked objects, and therefore makes it possible to accurately count even soft objects whose shapes are up and down or whose apparent thicknesses vary.

[0060] Furthermore, according to the terminal 1 for counting objects, by detecting each object itself in addition to the boundaries between each object, it is possible to further extract the feature points required when counting stacked objects. Therefore, when it is not possible to accurately count the objects based on the boundaries between each object due to the characteristics of each object or the characteristics of the background of the stacked objects, by counting each object itself, it becomes possible to accurately count even soft objects whose stacked shapes are up and down or whose apparent thicknesses vary.

[0061] Furthermore, according to the terminal 1 for counting objects, the count result of the piled up objects is output to the terminal, so that the user can immediately know the number of the piled up objects and the like.

[0062] [Second embodiment] [Overview of Terminal 1 for counting objects] An overview of a terminal 1 for counting objects according to a second embodiment of the present invention will be described with reference to Fig. 7. Fig. 7 is a diagram for explaining the overview of a terminal 1 for counting objects according to a first embodiment of the present invention. Note that the same functions and configurations as those in the first embodiment are given the same reference numerals, and descriptions thereof will be omitted. This embodiment differs from the first embodiment in that objects are counted using learned data.

[0063] [Overview of Terminal 1 for counting objects] An overview of a terminal 1 for counting objects according to a second embodiment of the present invention will be described with reference to Fig. 7. Fig. 7 is a diagram for explaining the overview of a terminal 1 for counting objects according to a second embodiment of the present invention.

[0064] As in the first embodiment, the terminal 1 for counting objects is a terminal such as a mobile terminal such as a handheld terminal, smartphone, or tablet terminal, or a wearable terminal such as a head-mounted display such as smart glasses or a smart watch, and is equipped with an imaging device such as a camera that captures images such as color video and / or still images.

[0065] The terminal 1 may be realized, for example, by one terminal device, as in the first embodiment, or may be realized by a plurality of terminal devices, or may be realized by a virtual device such as a cloud computer.

[0066] Next, an outline of the process executed by the terminal 1 to count objects will be described. First, the terminal 1 acquires an image 100 including an overall image 110 of the piled objects (step S21). Specifically, this is the same as step S1 in the first embodiment described above.

[0067] Next, the terminal 1 extracts a partial image 210 including the upper edge 120 and the lower edge 130 of the overall image 110 of the piled objects from the image 100, and creates a recognition image 200 in which a margin 220 of a different color from the color of the piled objects is provided around the partial image 210 (step S22). Specifically, this is the same as step S2 in the first embodiment described above.

[0068] Next, the terminal 1 estimates images of the upper edge 120 and lower edge 130 of the overall image 110 of the stacked objects and the boundaries 230 between each object 240 in the recognition image 200 (step S23). Specifically, the terminal 1 estimates and extracts the upper edge 120 and lower edge 130 of the overall image 110 of the stacked objects shown in the recognition image 200 and the boundaries 230 that distinguish each object 240, based on a trained model 300 created in advance through machine learning. There is no particular restriction on the type or method of machine learning. There is no particular limitation on the method for estimating the images of the upper edge 120, lower edge 130, and boundaries 230.

[0069] In step S23, the terminal 1 may re-learn the image-specific features of the estimated upper edge portion 120, lower edge portion 130, and each border 230 by machine learning, thereby updating the trained model 300. Note that this re-learning by machine learning may be performed at any time after step S23, or may be performed by the terminal 1 receiving input from the user. Machine learning will be described later.

[0070] Next, the terminal 1 counts the piled objects based on the images of the estimated upper edge 120, lower edge 130, and boundary 230 (step S24). Specifically, this is the same as step S4 in the first embodiment, except that the counting is based on the images of the estimated upper edge 120, lower edge 130, and boundary 230.

[0071] The above is an outline of the process executed by the terminal 1 to count objects.

[0072] [System configuration of Terminal 1 for counting objects] The system configuration of the terminal 1 for counting objects in this embodiment is the same as that in the first embodiment described above, and therefore a description thereof will be omitted.

[0073] [Object counting process] The object counting process executed by the terminal 1 for counting objects will be described with reference to Fig. 8. Fig. 8 is a diagram showing a flowchart of the object counting process executed by the terminal 1 for counting objects according to this embodiment. As shown in Fig. 8, the object counting process is made up of steps S31 to S34, which correspond to the above-mentioned steps S21 to S24.

[0074] First, the acquisition unit 10 of the terminal 1 acquires an image 100 including an overall image 110 of the piled objects (step S31). Specifically, this is the same as step S11 in the first embodiment described above.

[0075] Next, the creation unit 11 of the terminal 1 extracts a partial image 210 including an upper edge 120 and a lower edge 130 of the overall image 110 of the piled objects from the image 100, and creates a recognition image 200 in which a margin 220 of a different color from the color of the piled objects is provided around the partial image 210 (step S32). Specifically, this is the same as step S12 in the first embodiment described above.

[0076] Next, the estimation unit 15 of the terminal 1 estimates an image of the upper edge 120 and the lower edge 130 of the overall image 110 of the piled objects and the boundaries 230 between each object 240 in the recognition image 200 (step S33). Specifically, as shown in Fig. 6, the estimation unit 15 estimates the upper edge 120 and the lower edge 130 of the overall image 110 of the piled objects shown in the recognition image 200 and the boundaries 230 that distinguish each object 240, and extracts each boundary 230 by enclosing it in a rectangular shape so that it covers the entire area, that is, extracts a rectangular image including the upper edge 120, the lower edge 130, and each boundary 230. The method of estimating and extracting the upper edge 120, the lower edge 130, and each boundary 230 is not particularly limited. Furthermore, in this embodiment, the upper edge 120, the lower edge 130 and the borders 230 are surrounded by a rectangular shape, but they may be surrounded by any other shape.

[0077] In step S33, the detection unit 12 of the terminal 1 may further estimate an image of each object 240 in the recognition image 200. Specifically, similar to the estimation of the upper edge 120, the lower edge 130, and each boundary 230 described above, as shown in FIG. 6, the detection unit 12 estimates image-specific features of each object 240 appearing in the recognition image 200, and extracts each object 240 by enclosing it in a rectangular shape so that the entire object 240 is encompassed, that is, extracts a rectangular image including each object 240. The method of estimating and extracting each object 240 is not particularly limited. Furthermore, in this embodiment, each object 240 is surrounded by a rectangle, but it may be surrounded by any other shape.

[0078] Furthermore, in step S33, the learning unit 14 of the terminal 1 may re-learn the image-specific features of the estimated upper edge portion 120, lower edge portion 130, and each border 230 by machine learning, thereby updating the trained model 300. Furthermore, if the image-specific features of each object 240 have been estimated, the learning unit 14 may also re-learn the image-specific features of the estimated object 240 by machine learning, thereby updating the trained model 300. Note that this re-learning by machine learning may be performed at any timing after step S33, or may be performed by the input unit of the terminal 1 receiving input from the user. Machine learning will be described later.

[0079] Next, the counting unit 13 of the terminal 1 counts the piled objects based on the images of the estimated upper edge 120, lower edge 130, and boundary 230 (step S34). Specifically, this is the same as step S14 in the first embodiment, except that the counting is based on the images of the estimated upper edge 120, lower edge 130, and boundary 230.

[0080] This completes the object counting process.

[0081] [Learning process] The learning process executed by the terminal 1 for counting objects will be described with reference to Fig. 9. Fig. 9 is a diagram showing a flowchart of the learning process executed by the terminal 1 for counting objects.

[0082] The acquisition unit 10 of the terminal 1 acquires a learning image (step S51). Specifically, the recognition image 200 created by carrying out the above steps S21 to S22 is acquired as the learning image.

[0083] The learning unit 14 of the terminal 1 analyzes the created learning image (step S52). Specifically, the learning unit 14 detects and analyzes the images of the upper edge 120 and the lower edge 130 of the overall image 110 of the stacked objects and the boundaries 230 between each object 240, which are shown in the learning image, and extracts feature amounts specific to the upper edge 120, the lower edge 130, and each boundary 230. The learning unit 14 may also execute this step by receiving input from the user via the input unit of the terminal 1.

[0084] In step S52, the learning unit 14 of the terminal 1 may further detect each object 240 in the learning image, analyze the image, and extract a feature amount specific to each object 240.

[0085] The learning unit 14 of the terminal 1 performs machine learning based on the feature amounts specific to the upper edge portion 120, the lower edge portion 130, and each boundary 230 extracted in step S52 (step S53). Specifically, as machine learning, the learning unit 14 creates training data by tagging the feature amounts specific to the upper edge portion 120, the lower edge portion 130, and each boundary 230 extracted in step S52 as the upper edge portion 120, the lower edge portion 130, and each boundary 230 by annotation, and performs supervised learning. Furthermore, if feature amounts specific to each object 240 are extracted in step S52, the feature amounts may be tagged as each object 240 by annotation to create training data and perform supervised learning.

[0086] The annotation may be performed, for example, using a bounding box method in which an annotation tool encloses the regions of the upper edge 120, the lower edge 130, and each boundary 230, or the region of each object 240, in a rectangle. Needless to say, other methods may also be used. Any annotation tool may also be used. When performing annotation, the regions of the upper edge 120, the lower edge 130, and each boundary 230, or the region of each object 240 may be enclosed in a shape other than a rectangle. The created training data may also be learned using deep learning, which automatically defines and learns using a multi-layered neural network.

[0087] The learning unit 14 of the terminal 1 creates a trained model for recognition based on the learning result (step S54). Specifically, the learning unit 14 creates a trained model 300 incorporating the training data created in step S53.

[0088] The learning unit 14 of the terminal 1 stores the created trained model 300 in its own storage unit (step S55).

[0089] This completes the learning process.

[0090] According to the terminal 1 for counting objects, in addition to the effects obtained in the first embodiment, the following effects can be obtained.

[0091] By learning the recognition image and images of the upper and lower edges and boundaries to create a trained model, it is possible to improve the recognition accuracy of the upper and lower edges of stacked objects and the boundaries between each object, thereby improving the counting accuracy of soft objects whose stacked shapes vary up and down or whose apparent thicknesses are different.

[0092] In addition, by learning the recognition images and images of each object and creating a trained model, the recognition accuracy of stacked objects can be improved. Therefore, even if the upper or lower edges of the stacked objects or the boundaries between each object cannot be recognized, by counting each object itself, it is possible to improve the counting accuracy of soft objects whose stacked shapes are up or down or whose apparent thicknesses vary.

[0093] Furthermore, even if the stacked objects are shaped differently, or have different apparent thicknesses, bounding box annotation makes it possible to simply tag the image data to be learned with information on the image, learn from it, create a learning model, and maintain the accuracy of automatic object counting.

[0094] The above-described means and functions are realized by a computer (including a CPU, an information processing device, and various terminals) reading and executing a predetermined program. The program is provided, for example, in the form of a cloud service or SaaS (Software as a Service) provided from one or more terminals via a network. The program is also provided, for example, in the form of a program recorded on a computer-readable recording medium. In this case, the computer reads the program from the recording medium, transfers it to an internal or external recording device, records it, and executes it. The program may also be pre-recorded on a recording device (recording medium) such as a magnetic disk, optical disk, or magneto-optical disk, and provided to the terminal from the recording device via a communication line.

[0095] Although the embodiments of the present invention have been described above, the present invention is not limited to these embodiments. Furthermore, the effects described in the embodiments of the present invention are merely a list of the most preferable effects resulting from the present invention, and the effects of the present invention are not limited to those described in the embodiments of the present invention. [Explanation of symbols]

[0096] 1 terminal, 10 acquisition unit, 11 creation unit, 12 detection unit, 13 counting unit, 14 learning unit, 15 estimation unit, 100 image, 110 overall image of stacked objects, 120 upper edge, 130 lower edge, 200 recognition image, 210 partial image, 230 boundary, 240 each object, 300 trained model

Claims

1. A terminal for counting piled objects, an acquisition unit that acquires an image including an overall image of the piled objects; a creating unit that extracts a partial image of the piled objects, including an upper edge portion and a lower edge portion of the overall image of the piled objects, from the image, and creates an image for recognition in which a marginal portion of a different color from the piled objects is provided around the partial image; a detection unit that detects, in the recognition image, an image of the upper edge portion and the lower edge portion of the overall image of the piled up objects and an image of a boundary between each object; a counting unit that counts the piled objects based on the detected images of the upper edge portion, the lower edge portion, and the boundary; A terminal comprising:

2. the detection unit detects an image of each of the piled objects in the recognition image; The terminal according to claim 1 , wherein the counting unit counts the piled objects based on an image of each of the detected objects.

3. The terminal according to claim 1 , further comprising an output unit that outputs the counting result.

4. a learning unit that learns the recognition image and images of the upper edge portion, the lower edge portion, and the boundary to create a learned model; an estimation unit that estimates images of the upper and lower edges of the stacked objects and the boundaries between the objects in the recognition image based on the trained model, The terminal according to claim 1 , wherein the counting unit counts the piled objects based on images of the estimated upper edge, lower edge, and boundary.

5. the learning unit further learns an image of each of the objects; the estimation unit estimates an image of each of the stacked objects in the recognition image based on the trained model; The terminal according to claim 4 , wherein the counting unit counts the piled objects based on an image of each of the estimated objects.

6. The terminal according to claim 4 , wherein the learning unit learns using bounding box annotation that includes a part of the margin.

7. 1. A system for counting stacked objects, comprising: an acquisition unit that acquires an image including an overall image of the piled objects; a creating unit that extracts a partial image of the piled objects, including an upper edge portion and a lower edge portion of the overall image of the piled objects, from the image, and creates an image for recognition in which a marginal portion of a different color from the piled objects is provided around the partial image; a detection unit that detects, in the recognition image, an image of the upper edge portion and the lower edge portion of the overall image of the piled up objects and an image of a boundary between each object; a counting unit that counts the piled objects based on the detected images of the upper edge portion, the lower edge portion, and the boundary; A system comprising:

8. 1. A computer-implemented method for counting piled objects, comprising: acquiring an image including a full view of the pile of objects; a step of extracting a partial image of the stacked objects, including the upper and lower edge portions of the overall image of the stacked objects, from the image, and creating an image for recognition in which a margin of a different color from the color of the stacked objects is provided around the partial image; detecting, in the recognition image, images of the upper and lower edges of the entire images of the piled objects and the boundaries between the objects; counting the piled objects based on the detected images of the upper edge, the lower edge, and the boundary; A method for providing the above.

9. On the computer, acquiring an image including an overall view of the pile of objects; a step of extracting a partial image of the piled objects, including the upper and lower edge portions of the overall image of the piled objects, from the image, and creating an image for recognition in which a margin of a different color from the color of the piled objects is provided around the partial image; detecting, in the recognition image, an image of the upper edge portion and the lower edge portion of the overall image of the piled up objects and an image of a boundary between each object; counting the piled objects based on the detected images of the upper edge, the lower edge, and the boundary; A computer-readable program that causes a

Citation Information

Patent Citations

  • log scanning system

    JP2017527057A

  • Image analysis program

    JP2019185407A

  • Counting method, counting system, and computer program for counting

    JP2022167756A

  • Mobile terminal and hand towel management system

    JP7004375B1