Information processing device, information processing method, and program

The information processing device sets image areas using tag information from physical tags, addressing computational inefficiencies in AI-based object detection by focusing on relevant regions, thus reducing processing time and improving accuracy.

JP2025172654APending Publication Date: 2025-11-26SATO CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024078292
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-05-13
Publication Date
2025-11-26

AI Technical Summary

Technical Problem

Existing AI-based object detection and recognition processes require significant computational resources and time due to the need for extensive processing of entire images, necessitating high-performance equipment.

Method used

An information processing device that acquires images and sets image areas based on tag information from physical tags attached to objects, reducing computational load and processing time by focusing only on relevant image regions.

Benefits of technology

This approach reduces computational load and shortens processing time while improving the accuracy of annotation and estimation by limiting processing to only the image areas containing the objects, thereby enhancing overall efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025172654000001_ABST
    Figure 2025172654000001_ABST
Patent Text Reader

Abstract

To reduce the computation load of each process and time required for each process when executing AI processing for setting an image region.SOLUTION: An information processing device is provided, comprising an image acquisition unit for acquiring an image including a target object, a tag information acquisition unit for acquiring tag information for identifying the target object, and a setting unit for setting an image region related to the target object on the image on the basis of the tag information.SELECTED DRAWING: Figure 6
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an information processing device, an information processing method, and a program that are capable of setting an image area related to an object. [Background technology]

[0002] Conventionally, there are technologies for detecting objects through image recognition using AI. For example, a technology has been proposed in which multiple learning devices are used to detect multiple specification areas of an analog meter from a captured image of the analog meter (see, for example, Patent Document 1). [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2020-135353 Summary of the Invention [Problem to be solved by the invention]

[0004] Generally, when detecting an object using image recognition with AI, it is necessary to perform object detection processing and object recognition processing. Performing these detection and recognition processing requires a large amount of computational processing, which requires high-performance equipment and increases the time required for each processing step.

[0005] The present invention aims to reduce the computational load associated with each process and shorten the time required for each process when performing AI processing to set an image area. [Means for solving the problem]

[0006] One aspect of the present invention is an information processing device having an image acquisition unit that acquires an image including an object, a tag information acquisition unit that acquires tag information that identifies the object, and a setting unit that sets an image area related to the object in the image based on the tag information. [Effects of the Invention]

[0007] According to an aspect of the present invention, when performing AI processing to set an image area, it is possible to reduce the computational load associated with each process and shorten the time required for each process. [Brief explanation of the drawings]

[0008] [Figure 1] FIG. 1 is a diagram illustrating an example of use of an information processing system. [Figure 2] FIG. 2 is a top view showing an example of the configuration of a physical tag. [Figure 3] FIG. 3 is a top view showing an example of the configuration of a physical tag related to a barcode. [Figure 4] FIG. 4 is a top view showing an example of the configuration of a physical tag related to a graphic code. [Figure 5] FIG. 5 is a top view showing an example of the configuration of a physical tag relating to another graphic code. [Figure 6] FIG. 6 is a block diagram illustrating an example of the system configuration of the information processing system. [Figure 7] FIG. 7 is a diagram showing each piece of information stored in the setting information DB. [Figure 8] FIG. 8 is a diagram showing each piece of information stored in the trained model DB. [Figure 9A] FIG. 9A is a diagram showing an example of rectangular image region information. [Figure 9B] FIG. 9B is a diagram showing an example of circular image region information. [Figure 10A] FIG. 10A is a diagram showing an example of transition when a figure is subjected to affine transformation and translation. [Figure 10B] FIG. 10B is a diagram showing an example of transition when a figure is enlarged or reduced by affine transformation. [Figure 10C] FIG. 10C is a diagram showing an example of a transition when a figure is rotated by affine transformation. [Figure 10D] FIG. 10D is a diagram showing an example of transition when a figure is sheared by affine transformation. [Figure 11] FIG. 11 is a diagram showing an example of setting a bounding box for a box-shaped container. [Figure 12] FIG. 12 is a diagram showing an example of transformation when image region information is subjected to affine transformation. [Figure 13] FIG. 13 is a diagram showing an example in which physical tags are provided on three sides of a container. [Figure 14] FIG. 14 is a diagram showing an example of setting a bounding box for a cylindrical container. [Figure 15] FIG. 15 is a diagram showing an example of a display screen on a user terminal. [Figure 16] FIG. 16 is a diagram illustrating an example of an annotation method. [Figure 17] FIG. 17 is a flowchart illustrating an example of annotation processing. [Figure 18] FIG. 18 is a diagram illustrating an example of selecting a trained model. [Figure 19] FIG. 19 is a diagram showing an example of a display screen on a user terminal. [Figure 20] FIG. 20 is a flowchart illustrating an example of the estimation process. [Figure 21A] FIG. 21A is a perspective view showing an example in which physical tags are provided on two edges of a box-shaped container. [Figure 21B] FIG. 21B is a diagram showing an example of image area information when physical tags are provided on two sides of the edge of a box-shaped container. [Figure 22A] FIG. 22A is a perspective view showing an example in which physical tags are provided at two positions on the edge of a cylindrical container. [Figure 22B] FIG. 22B is a diagram showing an example of image area information when physical tags are provided at two positions on the edge of a cylindrical container. [Figure 23A] FIG. 23A is a perspective view showing an example in which a physical tag is provided on one side of the edge of a container according to a modified example. [Figure 23B] FIG. 23B is a diagram showing an example of image area information when a physical tag is provided on one side of the edge of a container according to a modified example. [Figure 24]FIG. 24 is a diagram showing a modified example in which an information processing system is used. [Figure 25] FIG. 25 is a diagram showing an example of identification information provided on the surface of a physical tag. [Figure 26] FIG. 26 is a block diagram showing an example of a system configuration according to a modified example of the information processing system. DETAILED DESCRIPTION OF THE INVENTION

[0009] The embodiments described below are not limited to the drawings described by the brief description of the drawings.

[0010] (1) One aspect of the present invention is an information processing device having an image acquisition unit that acquires an image including an object, a tag information acquisition unit that acquires tag information that identifies the object, and a setting unit that sets an image area related to the object in the image based on the tag information.

[0011] This makes it possible to set an image area for an object based on tag information that identifies the object, thereby reducing the computational load and time required for each process when performing AI processing to set the image area.

[0012] (2) One aspect of the present invention is an information processing device according to (1), wherein the setting unit extracts and sets the image area including the object from the image based on the tag information.

[0013] This allows only the image area containing the object to be the processing target (for example, annotation processing, object detection processing, object state estimation processing), thereby reducing the amount of calculation required for each process and shortening the time required for each process. For example, since only the image area containing the object can be the target of annotation processing, it is possible to improve the accuracy of annotation. Furthermore, for example, since only the image area containing the object can be the target of object state estimation processing, it is possible to improve the accuracy of the estimation processing.

[0014] (3) One aspect of the present invention is an information processing device as described in (1) or (2), in which a physical tag capable of reading the tag information is provided on or near the object, the tag information acquisition unit acquires the tag information using the physical tag, and the setting unit sets the image area based on the position of the physical tag in the image.

[0015] This makes it possible to acquire tag information based on a physical tag attached to or near an object, and to set an image area based on the position of the physical tag in the image, thereby making it possible to appropriately set an image area for each object.

[0016] (4) One aspect of the present invention is an information processing device described in (3), in which the physical tag is information having a predetermined shape, and the tag information acquisition unit acquires the tag information based on the shape of the physical tag.

[0017] This makes it possible to acquire tag information based on the shape of a physical tag attached to or near an object, and by attaching a physical tag of each shape to or near the waste, it becomes possible to appropriately set an image area related to the object.

[0018] (5) One aspect of the present invention is an information processing device as described in (3) or (4), in which the physical tag is arranged in a predetermined orientation relative to the object, and the setting unit sets the image area based on the orientation of the physical tag in the image.

[0019] This makes it possible to appropriately set an image area relating to an object based on the orientation of a physical tag attached to the object or near the object.

[0020] (6) One aspect of the present invention is an information processing device described in any of (1) to (5), wherein the tag information is associated with graphic selection information, and the setting unit selects the graphic selection information associated with the tag information and sets the image area using the graphic selection information.

[0021] According to this, by selecting the graphic selection information associated with the tag information corresponding to the object, it is possible to appropriately set the image area relating to the object.

[0022] (7) One aspect of the present invention is an information processing device described in (3) to (5), in which the physical tag is provided at a predetermined position and in a predetermined orientation relative to the object, and the setting unit sets the image area based on the position and orientation of the physical tag in the image.

[0023] This makes it possible to appropriately set an image area relating to the waste based on the position and orientation of a physical tag attached to or near the object.

[0024] (8) One aspect of the present invention is an information processing device described in any of (3) to (5) and (7), wherein the physical tag is provided at a predetermined position and in a predetermined orientation relative to the object, the tag information is associated with graphic selection information, and the setting unit selects the graphic selection information associated with the tag information and sets the image area based on the graphic selection information and the position and orientation of the physical tag in the image.

[0025] This makes it possible to appropriately set an image area for an object based on the position and orientation of a physical tag placed on or near the object and the graphic selection information associated with the tag information corresponding to the physical tag.

[0026] (9) One aspect of the present invention is an information processing device described in any of (1) to (8), further comprising an annotation unit that performs annotations regarding the object on the image area based on the tag information.

[0027] This allows annotation of an object based on tag information that identifies the object to be performed on an image region related to the object, thereby reducing the time required for annotation when generating a trained model and improving the accuracy of annotation.

[0028] (10) One aspect of the present invention is an information processing device described in any of (1) to (9), further comprising a selection unit that selects at least one trained model from among a plurality of trained models based on the tag information, and an estimation unit that inputs the image in which the image area is set to the selected trained model and estimates the state of the object based on the output from the trained model.

[0029] According to this, when estimating the state of an object (e.g., weight, filling rate) from an image region related to the object using a trained model, it is possible to select and use a trained model corresponding to the object from among multiple trained models based on tag information that identifies the object. This eliminates the need to simultaneously use multiple trained models to perform estimation processing to estimate the state of the object from an image including the object, making it possible to reduce the amount of calculation involved in the estimation processing and shorten the processing time.

[0030] (11) An aspect of the present invention is the information processing device according to (10), wherein the selection unit selects, from the plurality of trained models, a trained model corresponding to the volume of the object, a trained model corresponding to the shape of the object, and a trained model corresponding to the density of the object based on attributes of the object identified according to the tag information, and the estimation unit estimates the state of the object from the image region using each selected trained model.

[0031] For example, when the target object is a resin such as PET, it is possible that a mixture of standardized objects such as PET bottles and non-standardized objects such as rolls or plates is present. In this case, the volume of the standardized object differs from the volume of the non-standardized object, which may result in a large error in the weight calculated based on the volume. Furthermore, it is also possible that a mixture of standardized objects (e.g., uncrushed PET bottles) and compressed standardized objects (e.g., crushed PET bottles) is present. In this case, the volume of the standardized object differs from the volume of the compressed standardized object, which may result in a large error in the weight calculated based on the volume. Therefore, it is possible to improve the estimation accuracy by estimating the weight of an object using multiple trained models according to the attributes of the object.

[0032] (12) One aspect of the present invention is an information processing device described in any of (3) to (5), (7), and (8), wherein the physical tag is provided on each of a plurality of containers, and the setting unit sets the image area corresponding to each of the plurality of containers in the image for each of the containers.

[0033] This makes it possible to set an image area corresponding to each of multiple containers for each container, thereby improving the accuracy of setting the image area for the object contained in each container.

[0034] (13) One aspect of the present invention is an information processing device described in any one of (3) to (5), (7), (8), and (12), in which the physical tag is at least one of a barcode, a two-dimensional code, a color code capable of expressing multiple pieces of information by an arrangement of multiple colors, a graphic code capable of expressing multiple pieces of information by an arrangement of multiple figures each having one of multiple colors, and a wireless tag capable of transmitting the tag information using wireless communication.

[0035] This makes it possible to provide a physical tag appropriate for the environment of the installation location of a container that contains an object on or near the container, thereby making it possible to acquire tag information from the physical tag in accordance with the environment of the installation location of the container.

[0036] (14) One aspect of the present invention is an information processing device as described in (13), which, when using a wireless tag as the physical tag, provides direction identification information indicating the direction of the physical tag in the image together with the wireless tag.

[0037] This allows the position of the physical tag to be identified by the wireless tag, and the direction of the physical tag to be identified by the direction identification information, thereby improving the accuracy of detecting the position and direction of the container associated with the physical tag.

[0038] (15) One aspect of the present invention is an information processing device according to any one of (1) to (14), wherein the object is at least one of waste, agricultural products, marine products, and industrial products.

[0039] This makes it possible to reduce the computational load associated with each process and shorten the time required for each process when setting an image area related to the target object (at least one of waste, agricultural products, marine products, and industrial products).

[0040] (16) One aspect of the present invention is an information processing method including an image acquisition process for acquiring an image including an object, a tag information acquisition process for acquiring tag information that identifies the object, and a setting process for setting an image area related to the object in the image based on the tag information.

[0041] This makes it possible to set an image area for an object based on tag information that identifies the object, thereby reducing the computational load and time required for each process when performing AI processing to set the image area.

[0042] (17) One aspect of the present invention is a program that causes a computer to execute an image acquisition procedure for acquiring an image including an object, a tag information acquisition procedure for acquiring tag information that identifies the object, and a setting procedure for setting an image area related to the object in the image based on the tag information.

[0043] This makes it possible to set an image area for an object based on tag information that identifies the object, thereby reducing the computational load and time required for each process when performing AI processing to set the image area.

[0044] Hereinafter, embodiments will be described with reference to the accompanying drawings.

[0045] [Example of use of information processing system] Fig. 1 is a diagram showing an example of using the information processing system 1. Fig. 1 shows an example in which a plurality of types of waste are stored in a plurality of containers C1 to C11 at a waste collection site WP1. Fig. 1 also shows an example in which physical tags PT1 to PT11 are provided on the plurality of containers C1 to C11, respectively, for identifying the waste stored in each container.

[0046] For example, containers C1 and C5 contain plastic bags, containers C2 and C6 contain paper waste, containers C3 and C7 contain PP (polypropylene) resin, containers C4 and C8 contain hard plastic, container C9 contains glass waste, container C10 contains iron scrap, and container C11 contains circuit boards. The containers C1 to C11 can be made of a material and have a structure that suits the waste to be stored. For example, they can be made of a metal material with a mesh structure that is visible from the outside, or they can be made of a metal material with a box-like or cylindrical structure. Various known techniques can be used to determine the material and structure of containers C1 to C11 depending on the waste to be stored, and therefore detailed explanations will be omitted here.

[0047] In addition, in this embodiment, an example is shown in which waste stored in containers C1 to C11 is identified using physical tags PT1 to PT11 included in the imaging range IM1 of the imaging device 200. The physical tags PT1 to PT11 will be described in detail with reference to FIG. 2. Also, configuration examples of the information processing device 100 and the imaging device 200 will be described in detail with reference to FIG. 6. Also, the bounding box BB1 and the like will be described in detail with reference to FIGS. 11 to 14, etc.

[0048] [Physical tag configuration example] FIG. 2 is a top view showing an example of the configuration of the physical tag PT1. Here, only the physical tag PT1 is shown as a representative example, but the same applies to the other physical tags PT2 to PT11. Furthermore, this embodiment shows an example in which the physical tags PT1 to PT11 are used to acquire the attributes and positions of the target object (waste) (or the attributes and positions of the corresponding containers C1 to C11). Note that the physical tags PT1 to PT11 can be changed as appropriate depending on the positions at which they are attached to the containers. For example, when the physical tag PT11 is attached to the edge of the opening of the container C11, the physical tag PT11 can be configured along the circumferential edge.

[0049] FIG. 2 shows an example of a color code in which five rectangles of a predetermined color are arranged in a row as the physical tag PT1. This color code can be configured by coloring the rectangle R1 at the end in the arrangement direction black, and coloring the other four rectangles R2 to R5 other than the black rectangle R1 in a color other than black. In this case, it is possible to represent multiple pieces of information by changing the colors of the four rectangles R2 to R5 arranged in a row. For example, an arrangement in which rectangle R2 is blue, rectangle R3 is red, rectangle R4 is yellow, and rectangle R5 is green can be used as a color code representing a "plastic bag."

[0050] Note that, other than the color code shown in FIG. 2, other physical tags readable by the imaging device 200 may be used, or other wireless tags may be used. Examples of other physical tags readable by the imaging device 200 are shown in FIGS. 3 to 5. Ordinary barcodes and two-dimensional codes (e.g., QR Code (registered trademark)) may also be used as physical tags. Examples of wireless tags are shown in FIGS. 24 to 26. These physical tags are associated with unique identification information (tag information). For example, the information processing device 100 stores a tag DB (Data Base) 140 indicating the relationship between physical tags and tag information in the storage unit 130 (see FIG. 6), and can acquire tag information associated with the read physical tag using the tag DB 140. Note that the tag DB 140 may be stored in an external device and acquired from the external device for use.

[0051] For example, the physical tag PT1 can be a sheet-like member (not shown) that is attached to the object, and five rectangles R1 to R5, each colored a specific number, can be arranged in a line on this sheet-like member. The rectangles R1 to R5 have the same length L1 in the arrangement direction. The lengths CR1 to CR4 between the rectangles R1 to R5 in the arrangement direction are also the same. The lengths L1 in the arrangement direction may be different from each other, and the lengths CR1 to CR4 in the arrangement direction may also be different from each other. The lengths CR1 to CR4 in the arrangement direction may also be zero.

[0052] [Barcode configuration example] FIG. 3 is a top view showing an example of the configuration of a physical tag PT20 relating to a barcode.

[0053] FIG. 3 shows an example of a color code in which multiple rectangles are arranged in a row as the physical tag PT20. This color code can be configured by placing standard bars (rectangle R21 of length L11 and rectangle R26 of length L12) at both ends to indicate the direction of the arrangement, and arranging four bars of two different lengths (long bars, short bars) between them. Note that the arrangement may also be configured with three types of information: bars of two different lengths (long bars, short bars) and blank spaces where no bars are arranged. In this case, multiple pieces of information can be represented by changing the bars of two different lengths (long bars, short bars) and blank spaces arranged between the standard bars at both ends (rectangles R21, R26).

[0054] 3 shows an example of arranging a long bar (rectangle R22 of length L13), a short bar (rectangle R23 of length L14), a short bar (rectangle R24 of length L14), and a long bar (rectangle R25 of length L13). For example, this arrangement can be used as a barcode representing a "plastic bag." In this way, the physical tag PT20 can represent multiple pieces of information by using the horizontal length of each rectangle.

[0055] [Example of shape code configuration] FIG. 4 is a top view showing an example of the configuration of a physical tag PT30 relating to a graphic code.

[0056] FIG. 4 shows an example of a graphic code in which four types of graphics (circle (oval), triangle, square, inverted triangle) are arranged in a row as the physical tag PT30. This graphic code can be configured by placing standard bars (rectangles R31, R36) at both ends to indicate the direction of the arrangement, and arranging four of the four types of graphics between them. The four types of graphics may also be colored. In this case, multiple pieces of information can be represented by changing the four types of graphics arranged between the standard bars (rectangles R31, R36) at both ends and the colors assigned to them.

[0057] 4 shows an example of arranging a blue circle (oval) graphic R32, a red triangle graphic R33, a yellow square graphic R34, and a green inverted triangle graphic R35. For example, this arrangement can be used as a graphic code representing a "plastic bag." In this way, the physical tag PT30 can represent multiple pieces of information by combining multiple types of graphics and their colors.

[0058] [Example of shape code for changing vertical length] FIG. 5 is a top view showing an example of the configuration of a physical tag PT40 relating to a graphic code whose length in the up-down direction (vertical direction, height direction) is changed.

[0059] FIG. 5 shows an example of a graphic code as a physical tag PT50, in which four types of graphics (circle (oval), triangle, square, inverted triangle) whose lengths can be changed in two directions (vertical and height directions) are arranged in a row. This graphic code can be configured by placing standard bars (rectangles R41 and R46) at both ends to indicate the direction of the arrangement, and arranging four of the four types of graphics whose lengths can be changed in the vertical direction between them. Furthermore, each of the four types of graphics may be assigned a color. In this case, multiple pieces of information can be represented by changing the four types of graphics arranged between the standard bars (rectangles R41 and R46) at both ends, the colors assigned to them, and the vertical lengths.

[0060] 5 shows an example in which a circle (oval) R42 with a long vertical line of blue, a triangle R43 with a short vertical line of red, a square R44 with a short vertical line of yellow, and an inverted triangle R45 with a long vertical line of green are arranged. For example, this arrangement can be used as a graphic code representing a "plastic bag." In this way, the physical tag PT40 can represent multiple pieces of information by combining multiple types of shapes, their colors, and their vertical lengths.

[0061] [Example of information processing system configuration] FIG. 6 is a block diagram showing an example of the system configuration of the information processing system 1. As shown in FIG.

[0062] The information processing system 1 is composed of multiple devices that can be connected via a network N1. FIG. 6 shows an example of the information processing system 1 including an information processing device 100, an imaging device 200, and a user terminal 300. The information processing device 100, the imaging device 200, and the user terminal 300 are each connected to the network N1 by a communication method using wired communication or wireless communication. The network N1 is a network such as a public line network or the Internet. In this case, the wireless communication may be a mobile communication network (e.g., standards such as 3G (3rd Generation), 4G (4th Generation), 5G (5th Generation), and 6G (6th Generation)). Alternatively, at least one of wireless communication standards such as wireless LAN (e.g., Wi-Fi (Wireless Fidelity)), Bluetooth (registered trademark), and ZigBee (registered trademark) may be used. In addition, multiple frequency bands (e.g., UHF band and 2.4 GHz band) may be used in combination.

[0063] Note that the information processing device 100, the imaging device 200, and the user terminal 300 may be directly connected using wired or wireless communication without going through the network N1. Also, although only the imaging device 200 and the user terminal 300 are shown as representative examples in Fig. 6, other imaging devices and other user terminals installed at various locations may also constitute the information processing system 1.

[0064] [Configuration example of information processing device] The information processing device 100 includes a communication unit 110, a control unit 120, and a storage unit 130. The information processing device 100 can be, for example, a server realized by one or more devices.

[0065] The communication unit 110, under the control of the control unit 120, exchanges various types of information with other devices using wired or wireless communication.

[0066] The control unit 120 controls each unit of the information processing device 100 based on a control program stored in the storage unit 130. The control unit 120 is realized by a processing device such as a CPU (Central Processing Unit) or a GPU (Graphics Processing Unit). Specifically, the control unit 120 includes an image acquisition unit 121, a tag information acquisition unit 122, a setting unit 123, an annotation unit 124, a generation unit 125, a selection unit 126, an estimation unit 127, and an output control unit 128.

[0067] The image acquisition unit 121 acquires the captured image (including image data on which image processing has been performed and its accompanying information) generated by the imaging device 200 via the communication unit 110. Then, the image acquisition unit 121 outputs the acquired captured image to the tag information acquisition unit 122 and the setting unit 123.

[0068] Here, the object shown in this embodiment includes an object in which a regular object and an irregular object are mixed. For example, when the object is PET (Polyethyleneterephthalate) resin, it is assumed that a regular object such as a PET bottle and an irregular object such as a roll or plate are mixed. Furthermore, the measurement data of the object shown in this embodiment also includes data measured in a state in which one or more regular or irregular objects are present. For example, when the object is PET resin, it is assumed that one or more regular objects such as a PET bottle are mixed with one or more irregular objects (e.g., roll or plate). Furthermore, it is assumed that the regular objects include a mixture of normal regular objects (e.g., uncrushed PET bottles) and compressed regular objects (e.g., crushed PET bottles).

[0069] The tag information acquisition unit 122 reads a physical tag from the captured image output from the image acquisition unit 121 and acquires tag information based on the read physical tag. Then, the tag information acquisition unit 122 associates the position of the read physical tag (the position in the captured image) with the acquired tag information and outputs the associated information to the setting unit 123, the annotation unit 124, the generation unit 125, and the selection unit 126. For example, when using the physical tag PT1 shown in FIG. 2, it is possible to read the physical tag PT1 from the captured image using a known image recognition technique (e.g., an object detection technique using pattern matching). Then, the tag information acquisition unit 122 acquires tag information associated with the read physical tag PT1 using the tag DB 140. As described above, the tag DB 140, which indicates the relationship between the physical tag and the tag information, is stored in the storage unit 130.

[0070] The setting unit 123 sets a predetermined image area from the captured image output from the image acquisition unit 121 based on the tag information acquired by the tag information acquisition unit 122, and outputs setting information regarding the set image area to the annotation unit 124, the generation unit 125, the selection unit 126, and the estimation unit 127. This setting information includes the position, size, shape, etc. of the image area in the captured image. For example, assume a case where an image area corresponding to waste contained in a container C1 is set from a captured image corresponding to an imaging range IM1 (see FIG. 1). In this case, the setting unit 123 identifies the edge of the opening of the container C1 in the captured image based on information about the physical tag PT1 output from the tag information acquisition unit 122 (the position (base point), length, and direction (angle with respect to a specific direction) of the physical tag PT1 in the captured image)) and the tag information output from the tag information acquisition unit 122 (tag information corresponding to the physical tag PT1).

[0071] Specifically, the setting unit 123 extracts image area information 154 corresponding to the tag information (tag information corresponding to physical tag PT1) output from the tag information acquisition unit 122 from the setting information DB 150 (see FIG. 7). Then, the setting unit 123 identifies the edge of the opening of the container C1 based on the extracted image area information and information related to the physical tag PT1 (the position, length, and angle described above). Next, the setting unit 123 identifies an image area including the identified edge of the opening of the container C1, and sets the image area in the captured image. These setting methods will be described in detail with reference to FIGS. 11 to 14, etc.

[0072] The annotation unit 124 performs annotation on an object included in the captured image output from the image acquisition unit 121 based on the tag information output from the tag information acquisition unit 122. The annotation unit 124 then outputs the annotated image to the generation unit 125. For example, the annotation unit 124 tags an image area corresponding to the setting information output from the setting unit 123 with the attribute of the object, the state of the object, etc. corresponding to the tag information output from the tag information acquisition unit 122. As shown in FIG. 15 , the state of the object can be manually input by the user. A method of performing annotation will be described in detail with reference to FIG. 16 .

[0073] Here, annotation means adding various information to data. For example, the process of adding attribute information of an object to a captured image (image data) including the object can be called annotation. Note that various information added to data is sometimes called a label. Also, adding various information to data is sometimes called tagging. Also, annotated data (tagged data) is sometimes called training data. As will be described later, it is possible to generate a trained model by machine learning using training data.

[0074] The generation unit 125 generates a trained model using an image annotated by the annotation unit 124, and stores the generated trained model in the trained model DB 160. Note that a known generation method can be used to generate the trained model.

[0075] Here, the trained model is an AI (Artificial Intelligence) model that has been machine-learned using teacher data or training data. Examples of the trained model that can be used include SVM (Support Vector Machine), CNN (Convolutional Neural Network), ViT (Vision Transformer), and YOLO (You Only Look Once).

[0076] Furthermore, the trained model shown in this embodiment is generated by learning using images of waste stored in a container, and is capable of detecting, for example, the state (e.g., weight, filling rate) of the waste stored in the container. For example, a trained model can be generated by learning a large amount of training data in which example problems and corresponding correct answers are associated. Furthermore, when new input data is input to the trained model, it is possible to output output data that is the correct answer based on the learning results of the example problems and the corresponding correct answers.

[0077] For example, when learning the filling rate as the state of waste, it is possible to generate a learned model by learning a large amount of tagged data (teacher data) in which an image including the target waste (the waste contained in the container), the attributes of the waste, and the relationship between the filling rate of the waste in the container are associated. For example, when the filling rate when there is no waste in the container is set to 0, the filling rate when the container is full of waste is set to 1, and the filling rate of the waste in other containers is learned as a value between 0 and 1. In this case, a value between 0 and 1 is output as the output result of the learned model. That is, when there is no waste in the container, a filling rate of 0 is output, and when the container is full of waste, a filling rate of 1 is output. Also, when there is waste in the container and it is not full, a numerical value t (0 < t < 1) corresponding to the filling rate is output. Similarly, when learning the weight as the state of waste, it is possible to generate a learned model by learning a large amount of tagged data (teacher data) in which an image including the target waste (the waste contained in the container) and the relationship between the weight of the waste in this container are associated.

[0078] When generating a learned model, it is possible to cut out a rectangular image area including one container (including the waste) in which the target waste is contained from the captured images of each waste contained in each of a plurality of containers, and use the cut-out image. For example, as shown in FIG. 1, when waste is contained in 11 containers, it is possible to cut out 11 images for each container. Also, for example, as shown in FIG. 1, when 7 types of waste (plastic bags, paper scraps, PP resin, hard plastic, glass scraps, iron scraps, substrates) are contained in 11 containers, it is possible to generate a learned model according to the 7 types of waste. Then, when estimating the state of the waste, it is possible to select a learned model according to the type of the waste to be estimated, and use the selected learned model to output the state of the waste to be estimated.

[0079] The selection unit 126 selects a trained model stored in the trained model DB 160 based on the tag information acquired by the tag information acquisition unit 122, and outputs the selection result to the estimation unit 127. Specifically, the selection unit 126 selects trained model identification information 153 corresponding to the tag information 151 acquired by the tag information acquisition unit 122 from the setting information DB 150 (see FIG. 7 ). Next, the selection unit 126 outputs the selection result to the estimation unit 127. That is, the selection unit 126 outputs the selected trained model identification information to the estimation unit 127.

[0080] The estimation unit 127 estimates the state of waste contained in a container to which a physical tag corresponding to the tag information acquired by the tag information acquisition unit 122 is attached, and stores the estimation result in the estimation result DB 170 of the storage unit 130. Specifically, the estimation unit 127 estimates the state of waste contained in a container included in the captured image output from the image acquisition unit 121 using the trained model selected by the selection unit 126. For example, the estimation unit 127 can identify an image area set in the captured image based on the setting information output from the setting unit 123. Therefore, the estimation unit 127 inputs the captured image in which the image area is set by the setting unit 123 into the trained model selected by the selection unit 126, and estimates the output result from the trained model as the state of the waste contained in that image area. This waste estimation method will be described in detail with reference to Figures 18 to 20, etc.

[0081] The output control unit 128 executes control to output various types of information from the user terminal 300. For example, in response to a request from the user terminal 300, the output control unit 128 provides the estimation results stored in the estimation result DB 170 of the storage unit 130 to the user terminal 300 to display them on the display unit 306, or causes the sound output unit 307 to output audio information corresponding to the estimation results.

[0082] The storage unit 130 is a storage medium that stores various types of information. For example, the storage unit 130 stores various types of information (e.g., a control program, a tag DB 140, a setting information DB 150 (see FIG. 7), a trained model DB 160 (see FIG. 8), and an estimation result DB 170) that are required for the control unit 120 to perform various processes. The storage unit 130 also stores various types of information acquired via the communication unit 110. The storage unit 130 can be, for example, a read-only memory (ROM), a random access memory (RAM), a static random access memory (SRAM), a hard disk drive (HDD), a solid state drive (SSD), or a combination thereof.

[0083] [Configuration example of imaging device] The imaging device 200 includes a communication unit 201, a control unit 202, a storage unit 203, and an imaging unit 204. As the imaging device 200, for example, a remote camera that can be remotely controlled can be used.

[0084] The communication unit 201 exchanges various types of information with the information processing device 100 using wireless communication under the control of the control unit 202 .

[0085] The control unit 202 controls each unit of the imaging device 200 based on a control program stored in the storage unit 203. The control unit 202 is realized by a processing device such as a CPU or a GPU. For example, the control unit 202 executes control to associate the captured image generated by the imaging unit 204 with imaging device identification information and transmit the image to the information processing device 100. Based on the imaging device identification information, the information processing device 100 can identify that the captured image was generated at the waste collection site WP1.

[0086] The storage unit 203 is a storage medium that stores various types of information. For example, the storage unit 203 stores various types of information (for example, a control program, imaging device identification information for identifying the imaging device 200) that is required for the control unit 202 to perform various processes. The storage unit 203 also stores various types of information acquired via the communication unit 201. The storage unit 203 can be, for example, a ROM, a RAM, an SRAM, an HDD, an SSD, or a combination thereof.

[0087] The imaging unit 204 captures an image of a subject under the control of the control unit 202 to generate an image (image data), and outputs the generated image to the control unit 202. The imaging unit 204 is configured, for example, with an imaging element (image sensor) that receives light from the subject collected by lenses (e.g., multiple lenses that collect light from the subject), and an image processing unit that performs predetermined image processing on the image data generated by the imaging element. As the imaging element, for example, a CCD (Charge Coupled Device) type or a CMOS (Complementary Metal Oxide Semiconductor) type imaging element can be used. For example, the imaging unit 204 captures an image of a subject included in an imaging range IM1 (see FIG. 1 ) to generate a captured image, and outputs the captured image to the control unit 202. The imaging unit 204 may also generate still images periodically or irregularly, or may generate moving images.

[0088] [Example of user terminal configuration] The user terminal 300 includes a communication unit 301 , a control unit 302 , a storage unit 303 , an operation unit 304 , a sound acquisition unit 305 , a display unit 306 , and a sound output unit 307 .

[0089] The communication unit 301 exchanges various types of information with other devices using wired or wireless communication under the control of the control unit 202.

[0090] The control unit 302 controls each unit based on various programs stored in the storage unit 303. The control unit 302 is realized by, for example, a processing device such as a CPU or a GPU. For example, the control unit 302 executes control to transmit status information (for example, filling rate, weight) inputted through the operation unit 304 to the information processing device 100. Furthermore, for example, the control unit 302 executes control to display status information relating to the status of the object provided from the information processing device 100 on the display unit 306.

[0091] The storage unit 303 is a storage medium that stores various types of information. For example, the storage unit 303 stores various types of information (for example, control programs and applications) required for the control unit 202 to perform various processes. The storage unit 303 also stores various types of information acquired via the communication unit 301. The storage unit 303 can be, for example, a ROM, a RAM, an SRAM, an HDD, an SSD, or a combination thereof.

[0092] The operation unit 304 accepts various operations from the user U1 and outputs the contents of the accepted operation to the control unit 302. For example, when a user operation to input status information (e.g., filling rate, weight) regarding the status of waste contained in each of the multiple containers C1 to C11 installed at the waste collection site WP1 (see FIG. 1) is accepted, the operation unit 304 outputs the input information to the control unit 302. Furthermore, when a user operation to display status information regarding the status of waste contained in each of the multiple containers C1 to C11 installed at the waste collection site WP1 (see FIG. 1) on the display unit 306 is accepted, the operation unit 304 outputs the operation information regarding the user operation to the control unit 302.

[0093] The sound acquisition unit 305 acquires sounds near or around the user terminal 300 under the control of the control unit 302. As the sound acquisition unit 305, for example, one or more microphones can be used.

[0094] The display unit 306 is a display unit that displays various images under the control of the control unit 302. For example, a display panel such as an organic EL (Electro Luminescence) panel or an LCD (Liquid Crystal Display) panel can be used as the display unit 306. For example, a display screen 310 (see FIG. 15) for inputting the state of an object to be used for annotation when generating a trained model can be displayed on the display unit 306. Furthermore, for example, a display screen 320 (see FIG. 19) for displaying status information regarding the state of waste stored in each of a plurality of containers C1 to C11 installed at the waste collection site WP1 (see FIG. 1) can be displayed on the display unit 306.

[0095] The sound output unit 307 outputs various sounds based on the control of the control unit 302. As the sound output unit 307, for example, one or more speakers can be used.

[0096] [Example of settings information DB content] 7 is a diagram schematically showing each piece of information stored in the setting information DB 150. The setting information DB 150 is a database for managing each piece of information used when performing annotation on the waste contained in the containers C1 to C11 and each piece of information used when estimating the state of the waste contained in the containers C1 to C11.

[0097] The setting information DB 150 stores tag information 151, attribute information 152, trained model identification information 153, and image region information 154 in association with each other. Note that each of these pieces of information is an example, and other information may be stored in the setting information DB 150.

[0098] The tag information 151 is tag information acquired based on the physical tags PT1 to PT11 attached to the containers C1 to C11. In Fig. 7, an example is shown in which the tag information corresponding to the physical tag PT1 attached to the container C1 is "TG001," the tag information corresponding to the physical tag PT2 attached to the container C2 is "TG002," and the tag information corresponding to the physical tag PT11 attached to the container C11 is "TG011." Note that other tag information is omitted from the illustration.

[0099] The attribute information 152 is attribute information that indicates the attributes of the object (waste) contained in a container that is provided with a physical tag that corresponds to the tag information stored in the tag information 151. Note that the attributes shown in FIG. 7 are just an example, and other attributes relating to multiple classifications, other subdivided attributes, etc. may be stored. For example, if the object is a resin such as PET, since there is a mixture of standardized objects such as PET bottles and non-standardized objects such as rolls or plates, these may be classified and set as different attributes. Note that the attribute information and tag information correspond to each other, and therefore the attribute information may be included in the tag information.

[0100] The trained model identification information 153 is identification information that indicates the trained model to be used for the image area corresponding to the tag information stored in the tag information 151 when that tag information is acquired. Note that the same trained model is often used for waste with the same attributes. Therefore, the trained model identification information 153 may also be the same for waste with the same attribute information 152.

[0101] The image area information 154 is information used when tag information stored in the tag information 151 is acquired and an image area corresponding to the tag information is set. This image area information will be described in detail with reference to FIGS. 11 to 14, etc. The image area to be set is sometimes called a bounding box. This bounding box refers to an image area that surrounds an object included in a captured image. When a wireless tag is used as a physical tag (see FIGS. 24 to 26), each setting information such as attribute information 152, trained model identification information 153, and image area information 154 may be stored in the wireless tag. In this case, the information processing device 100 can acquire and use each setting information stored in the wireless tag. The image area information 154 is information for selecting a figure associated with the tag information and is an example of figure selection information.

[0102] [Example of contents of trained model DB] 8 is a diagram schematically illustrating each piece of information stored in the trained model DB 160. The trained model DB 160 is a database for managing a plurality of trained models used to estimate the state of the waste contained in the containers C1 to C11.

[0103] Trained model DB 160 stores trained model identification information 161, trained models 162, and attribute information 163 in association with each other. Note that each of these pieces of information is an example, and other information may also be stored in trained model DB 160.

[0104] Trained model identification information 161 is identification information for identifying each trained model stored in trained model 162. Note that trained model identification information 161 corresponds to trained model identification information 153 (see FIG. 7).

[0105] The trained model 162 is a trained model used to estimate the state of the waste contained in the containers C1 to C11. In this embodiment, an example is shown in which a trained model generated by the generation unit 125 (see FIG. 6) is stored in the trained model DB 160. However, a trained model generated by another device may be stored in the trained model DB 160 and used. Furthermore, this embodiment shows an example in which one or more trained models according to the attributes of the waste are selected and used from among a plurality of trained models. This selection example is shown in FIG. 18. Note that, although FIG. 8 shows an example in which a trained model is stored in the trained model DB 160 and used, a trained model stored in an external device may also be used.

[0106] The attribute information 163 is attribute information indicating the attributes of the object (waste) that uses the trained model stored in the trained model 162. The attribute information 163 corresponds to the attribute information 152 (see FIG. 7).

[0107] [Image area information] 9A and 9B are diagrams schematically showing image region information stored in image region information 154 (see FIG. 7). 9A and 9B show an example in which a figure for a bounding box is set in advance, and the bounding box is set based on this figure and information about the physical tag (for example, the position, length, and angle of the physical tag in the captured image).

[0108] 9A shows an example of constructing image area information based on the relationship between the shape of the edge SE1 of the opening of a container C1 when viewed from above and the physical tag PT1 attached to the edge SE1. In FIG. 9A, the length of the physical tag PT1 is T1, and the length of one side on which the physical tag PT1 is provided is T2. The ends of each side of the rectangle corresponding to the edge SE1 are indicated by E1, E3 to E5. FIG. 12 shows an example using the ratio of the length T1 of these ends E1, E3 to E5 and the physical tag PT1 to the length T2 of one side. The same applies to the relationship between the other containers C2 to C10, whose edges SE1 when viewed from above are rectangular, and the physical tags PT2 to PT10.

[0109] 9B shows an example of constructing image area information based on the shape of the edge SE11 of the opening of the container C11 when viewed from above and the relationship with the physical tag PT11 attached to the edge SE11 of the opening of the container C11. The same applies to the relationship between other containers and physical tags whose edge SE1 has a circular shape when viewed from above.

[0110] 9A, the shape of the edge of the opening of the container C1 to which the physical tag PT1 is attached (the shape of the edge when viewed from above) can be identified based on the physical tag PT1. Here, when using a captured image (a captured image from directly above the container C1) generated by the imaging device 200 provided directly above the container C1, it is assumed that the shape of the edge of the opening of the container C1 included in the captured image matches the shape of the edge SE1 associated with the physical tag PT1. Therefore, the shape of the edge of the opening of the container C1 to which the physical tag PT1 is attached can be identified based on the physical tag PT1.

[0111] However, when using a captured image generated by an imaging device 200 disposed diagonally above the container C1 (a captured image captured from diagonally above the container C1), it is assumed that the shape of the edge of the opening of the container C1 included in the captured image does not match the shape of the edge SE1 associated with the physical tag PT1. Therefore, in this case, the shape of the edge SE1 associated with the physical tag PT1 can be transformed based on the state of the physical tag PT1 included in the captured image (e.g., the position, length, and angle of the physical tag in the captured image), and the transformed shape can be identified as the shape of the edge of the opening of the container C1. For example, an affine transformation can be used as this transformation. Examples of this are shown in FIGS. 10A to 10D.

[0112] Affine Transformation 10A to 10D are diagrams showing an example of transition when a square figure S1 is subjected to affine transformation. In Fig. 10A to 10D, figures S2 to S5 after affine transformation are shown by thick lines.

[0113] Here, when affine transforming each coordinate (x, y) in two-dimensional space to coordinate (x', y'), the following formula is used: Therefore, it is possible to affine transform each coordinate (x, y) in a plane image, which is two-dimensional space, to coordinate (x', y') using the following formula.

[0114]

number

[0115] 10A shows an example of a transition when translating a figure S1 into a square figure S2. When translating the figure S1 in this way, the following affine matrix is ​​used. Note that tx and ty are parameters that specify the distance when translating.

[0116]

number

[0117] Figure 10B shows an example of a transition when scaling the figure S1 to convert it into a square figure S3. When scaling the figure S1 in this way, the following affine matrix is ​​used. Note that sx and sy are parameters that specify the scaling ratio when scaling.

[0118]

number

[0119] 10C shows an example of a transition when rotating the figure S1 to convert it into a square figure S4. When rotating the figure S1 in this way, the following affine matrix is ​​used. Note that θ is a parameter that specifies the angle when rotating.

[0120]

number

[0121] FIG. 10D shows an example of the transition when shearing (skewing) the figure S1 to transform it into a parallelogram figure S5. When shearing and transforming the figure S1 in this way, the following affine matrix is ​​used. Note that sh x , sh y is a parameter that specifies the distance at which shearing occurs.

[0122]

number

[0123] [Bounding box setting example] FIG. 11 is a diagram schematically showing an example of a setting method for setting a bounding box that includes waste (plastic bags) contained in a container C1 included in the imaging range IM1 (see FIG. 1).

[0124] As shown in FIG. 11(A), the tag information acquisition unit 122 (see FIG. 6) acquires tag information included in the captured image (corresponding to the imaging range IM1) acquired by the image acquisition unit 121. Specifically, the tag information acquisition unit 122 extracts physical tags PT1 to PT11 included in the captured image using a known image recognition technique as described above. Next, the tag information acquisition unit 122 acquires tag information corresponding to each of the physical tags PT1 to PT11 extracted from the captured image based on the arrangement order of the colors constituting the physical tags PT1 to PT11 as described above. As described above, the tag information corresponding to each of the physical tags PT1 to PT11 can be acquired using the tag DB 140 (see FIG. 6) that indicates the relationship between physical tags and tag information.

[0125] Next, the setting unit 123 (see FIG. 6) extracts, from the setting information DB 150, image area information 154 (see FIG. 7) corresponding to the tag information (corresponding to the physical tag PT1) acquired by the tag information acquisition unit 122. For example, as shown in FIG. 10A, the shape of a rectangular edge SE1 having the physical tag PT1 provided on one side is extracted.

[0126] 11(B), based on the physical tag PT1 extracted from the captured image (corresponding to the imaging range IM1), the setting unit 123 fits the extracted image region information (the shape of the rectangular edge SE1) to the shape of the edge of the opening of the container C1 included in the captured image. For example, the setting unit 123 identifies both ends E1' and E2' of the physical tag PT1 attached to the container C1 included in the imaging range IM1. Then, the setting unit 123 performs affine transformation on the rectangular edge SE1 so that both ends E1 and E2 of the rectangular edge SE1 coincide with both ends E1' and E2' of the physical tag PT1 included in the captured image. This affine transformation will be described in detail with reference to FIG. 12.

[0127] 11(C) shows an example in which the rectangular edge SE1 is affine transformed so that both ends E1 and SE2 of the shape of the rectangular edge SE1 coincide with both ends E1' and E2' (see FIG. 1) of the physical tag PT1 included in the captured image (corresponding to the imaging range IM1). That is, the rectangular edge SE1 is affine transformed so that the rectangular edge SE1 coincides with the edge of the opening of the container C1 included in the captured image.

[0128] 11(D1), the setting unit 123 fits the affine-transformed rectangular edge SE1 to the edge of the opening of the container C1 included in the imaging range IM1 based on the position (base point) of the physical tag PT1 in the captured image. Then, the setting unit 123 sets an image area of ​​a predetermined size including the fitted rectangular edge SE1 in the captured image (corresponding to the imaging range IM1). This image area can be used as a bounding box BB1.

[0129] The size of the bounding box BB1 can be set appropriately depending on the size of the fitted rectangular edge SE1. For example, it is possible to select and use the smallest size that can contain the fitted rectangular edge SE1 from among a plurality of preset rectangle sizes. Also, for example, the maximum size (e.g., horizontal length) of the fitted rectangular edge SE1 may be used as a reference, and the size of the rectangle that contains that maximum size may be selected. Note that while FIG. 11(D1) shows an example of setting a rectangular bounding box BB1, other shapes may also be set as the bounding box.

[0130] For example, as shown in Fig. 11(D2), the shape of the fitted rectangular edge SE1 may be set as a bounding box BB2. Alternatively, the bounding box may be set based on the position and size of the physical tag PT1. For example, it is possible to set a rectangular bounding box centered on the position of the physical tag PT1 and with a rectangular size corresponding to the length of the physical tag PT1.

[0131] In this way, the bounding boxes BB1 and BB2 can be set based on the position, size, shape (degree of deformation), etc. of the physical tag PT1. In this case, as described above, the image area can be set to any shape, figure, etc.

[0132] [Affine transformation example] FIG. 12 is a diagram showing an example of transformation when affine transformation is performed on image region information (the shape of the rectangular edge SE1) so that it matches the edge of the opening of the container C1.

[0133] For example, assume that the position of container C1 and the position and orientation (imaging direction) of imaging device 200 are fixed at waste collection site WP1. In this case, the distance from imaging device 200 to container C1 can be represented by r, and the position of container C1 relative to imaging device 200 can be represented by depression angle θ and azimuth angle φ. Furthermore, the spherical coordinates (r, θ, φ) can be converted to Cartesian linear coordinates (x, y, z) using the following equations.

[0134]

number

[0135] Furthermore, the opening surface of container C1 (the surface including the edge of the opening of container C1, the surface including physical tag PT1) is assumed to be on plane P (a virtual plane parallel to the floor surface). The height from the floor surface at the installation location of container C1 to the edge of the opening of container C1 is assumed to be h. Furthermore, it is assumed that the setting unit 123 is capable of grasping in advance the relationship between the three-dimensional space (three-dimensional coordinates) at waste collection site WP1 and the two-dimensional coordinates corresponding to the imaging range IM1.

[0136] In this case, the setting unit 123 identifies both ends E1', E2' of the physical tag PT1 acquired by the tag information acquisition unit 122, and calculates the length D1 of both ends E1', E2'. Next, the setting unit 123 calculates the length D2 (the length of ends E1', E3') of one side of the container C1 on which the physical tag PT1 is provided, based on the ratio (T1 / T2) of the length T1 (see FIG. 9A) of the physical tag PT1 in the image region information (the shape of the rectangular edge SE1) to the length T2 (see FIG. 9A) of one side on which the physical tag PT1 is provided. That is, the length D2 (=D1·T2 / T1) of the ends E1' and E3' on the plane P is calculated. This allows the setting unit 123 to identify the end E3' of the physical tag PT1 in three-dimensional space.

[0137] Next, the setting unit 123 changes the size of the image region information (the shape of the rectangular edge SE1) (see FIG. 9A) so that the length D2 of the ends E1' and E3' on the plane P matches the length T2 of the ends E1 and E3 of the image region information (the shape of the rectangular edge SE1). For example, the size of the image region information (the shape of the rectangular edge SE1) can be changed by an affine transformation (a transformation that enlarges or reduces a figure) shown in FIG. 10B. Then, the setting unit 123 draws the image region information (the shape of the rectangular edge SE1) after the size conversion on the plane P. This allows the setting unit 123 to identify both ends E4' and E5' of the physical tag PT1 in three-dimensional space.

[0138] The setting unit 123 can also identify the ends E1', E3', E4', and E5' of the physical tag PT1 included in the captured image corresponding to the imaging range IM1 based on the relationship between the three-dimensional space (three-dimensional coordinates) at the waste collection site WP1 and the two-dimensional coordinates corresponding to the imaging range IM1. Therefore, the setting unit 123 can calculate the crushed (sheared) area of ​​the opening of the container C1 in the captured image based on the relationship between the line segment between the ends E1' and E3' and the line segment between the ends E4' and E5' of the physical tag PT1 included in the captured image. That is, the shearing amount of the rectangle corresponding to the opening of the container C1 in the captured image can be calculated. The scaling size can be calculated based on the relationship between the length D1 of the line segment between the ends E1' and E2' of the physical tag PT1 and the length T1 of the line segment between the ends E1 and E2 of the image area information (the shape of the rectangular edge SE1). The rotation angle can be calculated based on the angle of the line segment between the ends E1' and E2' of the physical tag PT1 relative to a specific direction (e.g., the horizontal direction in the captured image).

[0139] Using the shear amount, scaling size, and rotation angle thus determined, it is possible to perform affine transformation so that the image region information (the shape of the rectangular edge SE1) coincides with the edge of the opening of the container C1, as shown in Figure 11(C). Furthermore, the image region information (the shape of the rectangular edge SE1) thus affine transformed is placed in the captured image so that the reference points (e.g., end E1' of the physical tag PT1 and end E1 of the physical tag PT1) coincide, as shown in Figures 11(D1) and 11(D2). This allows the bounding box BB1 or BB2 to be set.

[0140] In this way, if the position and imaging direction of the imaging device 200 are fixed and the size of the container C1 can be known in advance, the shear amount can be calculated based on the depression angle and azimuth angle.

[0141] The above example shows how affine transformation is performed to fit image area information (rectangular or circular edge shape) to each container based on the relationship between the three-dimensional space (three-dimensional coordinates) at waste collection point WP1 and the two-dimensional coordinates corresponding to imaging range IM1. It is also possible to measure container C1 included in imaging range IM1 in advance and calculate the shear amount of the opening (rectangular) edge of container C1 in advance based on the measurement results. In this case, the calculation results can be used to perform affine transformation.

[0142] Furthermore, for example, affine transformation can be performed based on the length (size) D1 and angle (e.g., angle relative to the horizontal direction) of both ends E1', E2' of the physical tag PT1. That is, affine transformation can be performed so that the length and angle of both ends E1, E2 of the rectangular edge SE1 coincide with the length and angle of both ends E1', E2' of the physical tag PT1. For example, using the difference value (e.g., angle, length) between end E2 and end E2' when end E1 and end E1' are coincident, an affine transformation can be performed by setting an affine matrix so that the difference value between end E2 and end E2' becomes 0. For each of these affine transformations, a known transformation method can be used.

[0143] The edge of the opening of the container C1 can be detected using a known image recognition technology (e.g., edge detection technology). Therefore, the angle between one of the four edges of the container C1, on which the physical tag PT1 is provided, and another adjacent edge can be obtained. The angle between these two adjacent edges can be used to estimate the shape (degree of deformation) of the rectangle corresponding to the edge of the opening of the container C1. For example, when using a captured image (a captured image from directly above the container C1) generated by the imaging device 200 provided directly above the container C1, the rectangle corresponding to the edge of the opening of the container C1 included in the captured image is rectangular. Therefore, a 90-degree angle can be obtained as the angle between the one side on which the physical tag PT1 is provided and the other adjacent side. On the other hand, when using a captured image (a captured image from diagonally above the container C1) generated by the imaging device 200 provided diagonally above the container C1, the rectangle corresponding to the edge of the opening of the container C1 included in the captured image has a sheared rectangular shape. In this case, a value less than 90 degrees is acquired as the angle between one side on which the physical tag PT1 is provided and another side adjacent thereto.

[0144] By providing physical tags on two or more of the four sides of the container C1, it is possible to calculate the shear amount by calculating the angles of the two or more physical tags without detecting the edge of the opening of the container C1. Examples of providing physical tags on two or more of the four sides of the container C1 are shown in Figures 13, 21A and 21B, 22A and 22B.

[0145] Furthermore, if the positions and orientations of the containers C1 to C11 and the position and orientation (imaging direction) of the imaging device 200 are fixed at the waste collection site WP1, the shapes of the edges of the openings of the containers C1 to C11 included in the imaging range IM1 can be acquired in advance by measurement or the like. Therefore, in such a case, the shapes of the edges of the openings of the containers C1 to C11 included in the imaging range IM1 (deformed rectangular shapes) can be measured and acquired in advance and stored in the image region information 154 (see FIG. 7) corresponding to the tag information 151 (see FIG. 7). For example, the graphic information shown in FIG. 11(C) can be stored in the image region information 154 as the image region information corresponding to the container C1. This makes it possible to omit the above-mentioned processes such as affine transformation, and the graphic information stored in the image region information 154 can be used as is.

[0146] The above example illustrates the affine transformation of the rectangular edge SE1 based on the length and angle of both ends E1', E2' of the physical tag PT1. However, the affine transformation may also be performed based on other information. For example, the rectangular edge SE1 after the affine transformation may be compared with the edge of the opening of the container C1, and whether or not to use the affine-transformed rectangular edge SE1 may be determined based on the degree of similarity. For example, if the degree of similarity is equal to or greater than a threshold, the affine-transformed rectangular edge SE1 is used. If the degree of similarity is less than the threshold, the affine-transformed rectangular edge SE1 is again affine-transformed. The edge of the opening of the container C1 can be detected using known image recognition technology (e.g., edge detection technology). Furthermore, known image recognition technology (e.g., a determination technology based on brightness difference values) can be used to compare images.

[0147] [Example of physical tags on three sides] 12 shows an example of affine transformation in which one physical tag PT1 is attached to one side of one container C1. However, one container may be attached with multiple physical tags. This makes it possible to simplify the calculation of the affine transformation.

[0148] 13 is a diagram showing an example in which physical tags PT1A, PT1B, and PT1C are provided on three sides of the edge of a container C1. In this case, image region information (the shape of the rectangular edge SE1A) corresponding to the physical tags PT1A, PT1B, and PT1C provided on the container C1 is stored in the image region information 154 (see FIG. 7).

[0149] 13 shows a perspective view of the container C1 on the left side, and image area information (the shape of the rectangular edge SE1A) on the right side. In this case, similar to the example shown in FIG. 11(B), the setting unit 123 fits the extracted image area information (the shape of the rectangular edge SE1A) to the shape of the edge of the opening of the container C1 included in the captured image based on the physical tags PT1A, PT1B, and PT1C extracted from the captured image (corresponding to the imaging range IM1). In the example shown in FIG. 13, the physical tags PT1A, PT1B, and PT1C are provided on three sides of the container C1, so the setting unit 123 identifies the ends, lengths, and angles (e.g., the angles between the physical tags) of the physical tags PT1A, PT1B, and PT1C provided on the container C1 included in the imaging range IM1. The setting unit 123 then performs affine transformation on the image region information (the shape of the rectangular edge SE1A) so that the physical tags PT1A, PT1B, and PT1C on three sides of the image region information (the shape of the rectangular edge SE1A) match the physical tags PT1A, PT1B, and PT1C included in the captured image. A known affine transformation method can be used for this affine transformation. As shown in FIG. 13, by providing physical tags PT1A, PT1B, and PT1C on three sides of the container C1, it is possible to calculate the angle by which the rectangle of the container C1 is deformed, which makes it easier to calculate the shear amount.

[0150] [Example of setting a bounding box for a cylindrical container] FIG. 14 is a diagram schematically showing an example of a setting method for setting a bounding box that includes the object (substrate) contained in the container C11 included in the imaging range IM1 (see FIG. 1).

[0151] 14(A), the tag information acquisition unit 122 (see FIG. 6) acquires tag information included in the captured image (corresponding to the imaging range IM1) acquired by the image acquisition unit 121. The method for acquiring this tag information is the same as the example shown in FIG.

[0152] Next, the setting unit 123 (see FIG. 6) extracts, from the setting information DB 150, image region information 154 (see FIG. 7) corresponding to the tag information (corresponding to the physical tag PT11) acquired by the tag information acquisition unit 122. For example, as shown in FIG. 10B, the shape of a circular edge SE11 on which the physical tag PT11 is provided is extracted.

[0153] As shown in FIG. 14(B), based on the physical tag PT11 extracted from the captured image (corresponding to the imaging range IM1), the setting unit 123 fits the extracted image region information (the shape of the rectangular edge SE11) to the shape of the edge of the opening of the container C11 included in the captured image. For example, the setting unit 123 identifies both end portions E11′, E12′ and a central portion M13′ of the physical tag PT11 attached to the container C11 included in the imaging range IM1. Then, the setting unit 123 performs affine transformation on the circumferential edge SE11 so that both end portions E11, E12, and the central portion M13 of the circumferential edge SE11 coincide with both end portions E11′, E12′, and the central portion M13′ of the physical tag PT11 included in the captured image. A known transformation method can be used for this affine transformation.

[0154] For example, an affine transformation can be performed based on the circumferential shape specified by the lengths (sizes) of both ends E11', E12' of the physical tag PT11 and the central portion M13'. That is, an affine transformation can be performed so that the circumferential shape specified by both ends E11, E12 and the central portion M13 of the circumferential edge SE11 matches the circumferential shape specified by both ends E11', E12' and the central portion M13' of the physical tag PT11. The method for calculating the shear amount of the opening (circular shape) of the container C11 is the same as that shown in FIG. 12.

[0155] 14(C), the circumferential edge SE11 is affine transformed so that both ends E11, E12 and a central portion M13 of the rectangular shape of the edge SE11 coincide with both ends E11', E12' and a central portion M13' of the physical tag PT11 included in the captured image (corresponding to the imaging range IM1). That is, the circumferential edge SE11 is affine transformed so that the circumferential edge SE11 coincides with the edge of the opening of the container C11 included in the captured image.

[0156] 14(D), the setting unit 123 fits the affine-transformed circumferential edge SE11 to the edge of the opening of the container C11 included in the imaging range IM1 based on the position (base point) of the physical tag PT11 in the captured image. Then, the setting unit 123 sets an image area of ​​a predetermined size including the fitted circumferential edge SE11 in the captured image (corresponding to the imaging range IM1). This image area can be used as a bounding box BB11.

[0157] 11, the circumferential edge SE11 after the affine transformation may be compared with the edge of the opening of the container C11, and whether or not to use the circumferential edge SE11 after the affine transformation may be determined based on the degree of similarity. Bounding boxes can also be set for edges of other shapes in a similar manner.

[0158] [Example of input using a user terminal] Fig. 15 is a diagram showing an example of input when the state of waste is input using the display screen 310 displayed on the display unit 306 of the user terminal 300. Annotation is performed using the state of waste input in this way. Fig. 15 shows an example of inputting the filling rate as the state of waste contained in containers C1 to C11 installed at the waste collection site WP1 (see Fig. 1). Fig. 17 shows a learning method using the filling rate of waste input in Fig. 15.

[0159] A container information display area 311, a waste information display area 312, and a filling rate input area 313 are displayed in association with each other on the display screen 310. The display screen 310 can be displayed based on the control of the information processing device 100. The container information display area 311 and the waste information display area 312 can be displayed based on the information stored in the tag DB 140 and the setting information DB 150 of the storage unit 130. FIG. 15 shows an example in which only some (containers C1 to C4) of the containers C1 to C11 installed at the waste collection point WP1 are displayed. The other containers can be displayed using a scroll bar 314.

[0160] Container information for identifying each of the containers C1 to C11 installed at the waste collection point WP1 is displayed in the container information display area 311. Fig. 15 shows an example in which the container information corresponding to container C1 is displayed as "first container," the container information corresponding to container C2 is displayed as "second container," the container information corresponding to container C3 is displayed as "third container," and the container information corresponding to container C4 is displayed as "fourth container."

[0161] The waste information display area 312 displays the names of the waste materials to identify the waste materials contained in the containers C1 to C11 installed at the waste material collection point WP1.

[0162] The filling rate input area 313 displays the filling rate of each waste material input using the operation unit 304. In Fig. 15, in the filling rate input area 313, the column of the waste material to be input is indicated by being surrounded by a thick line 315. Note that information may be input to the filling rate input area 313 based on the user's voice acquired using the sound acquisition unit 305.

[0163] In this way, the inputted filling rate of the waste is transmitted to the information processing device 100 in association with the identification information (tag information 151 (see FIG. 7)) of each waste.

[0164] [Annotation execution example] 16 is a diagram illustrating an example of an annotation method for generating a trained model for waste (plastic bags) contained in a container C1 included in the imaging range IM1 (see FIG. 1). Here, an example is shown in which the attribute of the waste (plastic bags) and the filling rate of the waste are tagged.

[0165] 16(A), the tag information acquisition unit 122 (see FIG. 6) acquires tag information included in the captured image (corresponding to the imaging range IM1) acquired by the image acquisition unit 121. The method of acquiring this tag information is similar to the examples shown in FIGS. 11 to 14, etc.

[0166] Next, the setting unit 123 (see FIG. 6) extracts attribute information 152 and image area information 154 (see FIG. 7) corresponding to the tag information (corresponding to physical tag PT1) acquired by the tag information acquisition unit 122 from the setting information DB 150. For example, "plastic bag" is extracted as the attribute corresponding to physical tag PT1, and the shape of edge SE1 shown in FIG. 9A is extracted as the image area information corresponding to physical tag PT1.

[0167] As shown in Fig. 16(B), the setting unit 123 sets a bounding box BB11 using the extracted image region information. The method for setting this bounding box BB1 is the same as the example shown in Fig. 11. Note that instead of the bounding box BB1, another shape (for example, a bounding box BB2) may be set.

[0168] Also, in FIG. 16(B), the attribute "plastic bag" extracted as the attribute corresponding to the physical tag PT1 and the filling rate "89%" acquired in response to a user operation (see FIG. 15) are enclosed in a rectangle TG1. The attribute "plastic bag" and the filling rate "89%" shown in the rectangle TG1 are tagged to the image included in the bounding box BB1. That is, FIG. 16 shows an example in which the attribute "plastic bag" and the filling rate "89%" shown in the rectangle TG1 are used as labels. Note that other labels may be added as necessary.

[0169] Next, the annotation unit 124 (see FIG. 6) tags the image included in the bounding box BB1 with the attribute "plastic bag" and the filling rate "89%" shown in the rectangle TG1.

[0170] Next, the generation unit 125 (see FIG. 6) generates a trained model using the image included in the bounding box BB1 tagged with "plastic bag" and a filling rate of "89%" by the annotation unit 124. A known learning method can be used to generate this trained model.

[0171] Here, if the physical tag PT1 falls within the bounding box BB1, there is a risk that the accuracy of the image data to be learned will decrease. Therefore, for example, taking into consideration the image area corresponding to the bounding box BB1, the physical tag PT1 may be attached to the container C1 (or the periphery of the contained waste) so that the physical tag PT1 does not fall within the bounding box BB1. Also, for example, a masking process may be performed to mask the physical tag PT1 included in the bounding box BB1. Note that known image processing can be used for this masking process.

[0172] [Example of operation of information processing device] Fig. 17 is a flowchart showing an example of annotation processing in the information processing device 100. This annotation processing is executed by the control unit 120 (see Fig. 6) based on a program stored in the storage unit 130 (see Fig. 6). Fig. 17 also shows an example of tagging the attributes and filling rate of each waste contained in containers C1 to C11 installed at the waste collection site WP1 (see Fig. 1). This annotation processing will be explained with appropriate reference to Figs. 1 to 16.

[0173] In step S501 , the image acquisition unit 121 acquires a captured image generated by the imaging device 200 .

[0174] In step S502, the tag information acquisition unit 122 reads the physical tags PT1 to PT11 attached to the containers C1 to C11, respectively, based on the captured images acquired in step S501. A known image recognition technique can be used to read these physical tags.

[0175] In step S503, the tag information acquisition unit 122 acquires tag information of the objects (waste) contained in each of the containers C1 to C11 based on each of the physical tags PT1 to PT11 read in step S502.

[0176] In step S504, the setting unit 123 acquires the attributes of the object based on the tag information of the object acquired in step S503. For example, the setting unit 123 extracts and acquires attribute information 152 corresponding to tag information 151 of the object from the setting information DB 150 (see FIG. 7). Furthermore, the annotation unit 124 acquires the state of the object when tagging the state of the object (e.g., filling rate, weight) together with the attribute of the object. For example, as shown in FIG. 15, it is possible to acquire the state of the object (e.g., filling rate, weight) input by a user operation.

[0177] In step S505, setting unit 123 sets an image area including the object based on the tag information of the object acquired in step S503. The method for setting this image area can be the same as the examples shown in Figs. 11 to 14, etc.

[0178] In step S506, the annotation unit 124 performs annotation to tag the image region of the object set in step S505 with the information about the object acquired in step S504 (attributes of the object, state of the object). Then, the generation unit 125 performs a learning process on the annotated image region of the object.

[0179] These processes may be executed sequentially for each object, or may be executed in parallel for each object.

[0180] The generation unit 125 sequentially stores the trained models generated in step S506 in the trained model DB 160 of the storage unit 130.

[0181] [Example of estimating the state of waste] The above describes an example of generating a trained model used to estimate the state of waste. Below, we will describe an example of using the trained model thus generated to estimate the state of waste.

[0182] [Example of selecting a trained model] 18 is a diagram showing an example of a selection method for selecting a trained model to be used when estimating the state of waste (plastic bags) contained in a container C1 included in the imaging range IM1 (see FIG. 1). Here, an example is shown in which the filling rate of the target object is calculated as the state of the waste (plastic bags).

[0183] 18(A), the tag information acquisition unit 122 (see FIG. 6) acquires tag information included in the captured image (corresponding to the imaging range IM1) acquired by the image acquisition unit 121. The method for identifying this tag information is similar to the examples shown in FIGS. 11 to 14, etc.

[0184] Next, the selection unit 126 (see FIG. 6) extracts, from the setting information DB 150, trained model identification information 153 (see FIG. 7) corresponding to the tag information (corresponding to the physical tag PT1) acquired by the tag information acquisition unit 122. For example, as shown in FIG. 18(B), trained model identification information "LL4" corresponding to the tag information (corresponding to the physical tag PT1) is selected. In this case, the fourth trained model LL4 corresponding to the trained model identification information "LL4" is selected from the trained model DB 160 (see FIG. 7) and used.

[0185] Next, the estimation unit 127 uses the trained model selected by the selection unit 126 to estimate the state of the waste contained in the container included in the captured image output from the image acquisition unit 121. For example, as shown in FIG. 18(C), the estimation unit 127 inputs the captured image in which the image area (bounding box BB1) is set by the setting unit 123 into the fourth trained model LL4 selected by the selection unit 126, and estimates the output result from the fourth trained model LL4 as the state (filling rate) of the waste (plastic bags) included in the image area (bounding box BB1). Then, the estimation unit 127 stores the estimation result in the estimation result DB 170 of the storage unit 130.

[0186] [Example of display on user terminal] Fig. 19 is a diagram showing an example of a display screen 320 displayed on the display unit 306 of the user terminal 300. Fig. 19 shows an example of displaying the estimated results when the filling rate of waste stored in the containers C1 to C11 installed at the waste collection site WP1 (see Fig. 1) is estimated.

[0187] A container information display area 321, a waste information display area 322, and a filling rate display area 323 are displayed in association with each other on the display screen 320. Each of these pieces of information can be displayed based on each piece of estimation result information stored in the estimation result DB 170 of the storage unit 130. Note that Fig. 19 shows an example in which only some (containers C1 to C4) of the containers C1 to C11 installed at the waste collection point WP1 are displayed. The other containers can be displayed using a scroll bar 324.

[0188] Container information for identifying each of the containers C1 to C11 installed at the waste collection point WP1 is displayed in the container information display area 321. Fig. 19 shows an example in which the container information corresponding to container C1 is displayed as "first container," the container information corresponding to container C2 is displayed as "second container," the container information corresponding to container C3 is displayed as "third container," and the container information corresponding to container C4 is displayed as "fourth container."

[0189] The waste information display area 322 displays the names of the waste materials to identify the waste materials contained in the containers C1 to C11 installed at the waste collection site WP1.

[0190] The filling rate display area 323 displays the filling rate of each waste estimated by the estimation unit 127. When images generated by the imaging device 200 are acquired in real time and the filling rates of the waste are sequentially estimated, it is possible to change the display content of the filling rate display area 323 each time the estimation process is executed.

[0191] Furthermore, each piece of information that can be displayed on the display screen 320 can be output from the sound output unit 307. For example, each piece of information for each object (for example, audio information SS1) can be output from the sound output unit 307.

[0192] [Example of operation of information processing device] Fig. 20 is a flowchart showing an example of estimation processing in the information processing device 100. This estimation processing is executed by the control unit 120 (see Fig. 6) based on a program stored in the storage unit 130 (see Fig. 6). Fig. 20 also shows an example of estimating the filling rate of each waste contained in containers C1 to C11 installed at the waste collection site WP1 (see Fig. 1). This estimation processing will be explained with appropriate reference to Figs. 1 to 19.

[0193] In step S511 , the image acquisition unit 121 acquires the captured image generated by the imaging device 200 .

[0194] In step S512, the tag information acquisition unit 122 reads the physical tags PT1 to PT11 attached to the containers C1 to C11, respectively, based on the captured images acquired in step S511. A known image recognition technique can be used to read these physical tags.

[0195] In step S513, the tag information acquisition unit 122 acquires tag information of the objects (waste) contained in each of the containers C1 to C11 based on each of the physical tags PT1 to PT11 read in step S512.

[0196] In step S514, the setting unit 123 sets an image area including the object based on the tag information of the object acquired in step S513. The method for setting this image area can be the same as the examples shown in Figs. 11 to 14, etc.

[0197] In step S515, the selection unit 126 selects a trained model to be used when estimating the filling rate of the object based on the tag information of the object acquired in step S513. The method for selecting this trained model can be the same as the example shown in FIG.

[0198] In step S516, the estimation unit 127 uses the trained model selected in step S515 to estimate the filling rate of the object for which tag information was acquired in step S513. Specifically, the estimation unit 127 inputs the captured image for which the image area was set in step S514 to the selected trained model, and estimates the response result as the filling rate of the object.

[0199] These processes may be executed sequentially for each object, or may be executed in parallel for each object.

[0200] The estimation unit 127 sequentially stores the estimation results estimated in step S516 in the estimation result DB 170 of the storage unit 130. Furthermore, the output control unit 128 provides the estimation results stored in the estimation result DB 170 of the storage unit 130 to the user terminal 300 in response to a request from the user terminal 300. Furthermore, the control unit 302 of the user terminal 300 causes the display unit 306 to display the estimation results provided from the information processing device 100. For example, as shown in FIG. 19 , it is possible to cause the display unit 306 to display a display screen 320. Furthermore, the control unit 302 of the user terminal 300 causes the sound output unit 307 to output audio information related to the estimation results provided from the information processing device 100. For example, as shown in FIG. 19 , it is possible to cause the sound output unit 307 to output audio information SS1.

[0201] [Modification of physical tags] In Figure 1 and other figures, one physical tag is provided for one container, and in Figure 13, three physical tags are provided for one container. However, two, four, or more physical tags may be provided for one container. This makes it possible to improve the accuracy of setting the bounding box.

[0202] 21A and 21B are diagrams illustrating an example in which physical tags PT1a and PT1b are provided on two edges of a container C1. FIG. 21A illustrates a perspective view of the container C1, and FIG. 21B illustrates the shape of the edge SE1a corresponding to the container C1, which is stored in the image region information 154 (see FIG. 7). In this manner, image region information corresponding to the physical tags PT1a and PT1b provided on the container C1 is stored in the image region information 154. The setting unit 123 can also set a bounding box using the positions, lengths, angles, etc. of the physical tags PT1a and PT1b provided on the two edges of the container C1. In this manner, providing two physical tags PT1a and PT1b makes it possible to obtain the angles of the two edges of the container C1, thereby facilitating the calculation of the shear amount.

[0203] 22A and 22B are diagrams illustrating an example in which physical tags PT11a and PT11b are provided at two opposing positions on the edge of a container C11. FIG. 22A illustrates a perspective view of the container C11, and FIG. 22B illustrates the shape of the edge SE11a corresponding to the container C11, which is stored in the image region information 154 (see FIG. 7). In this manner, image region information corresponding to the physical tags PT11a and PT11b provided on the container C11 is stored in the image region information 154. The setting unit 123 can also set a bounding box using the positions, lengths, angles, and the like of the physical tags PT11a and PT11b provided at two opposing positions on the container C11. In this manner, providing two physical tags PT11a and PT11b makes it possible to acquire the positional relationship between the two opposing positions on the container C11, thereby facilitating the calculation of the shear amount.

[0204] [Container Modification] In the above, examples have been shown in which physical tags are provided on box-shaped and cylindrical containers. However, one or more physical tags may be provided on containers of other shapes. Figures 23A and 23B show examples of other shapes.

[0205] 23A and 23B are diagrams showing an example in which a physical tag PT13 is provided on one side of the edge of an L-shaped container C13 when viewed from above. FIG. 23A shows a perspective view of the container C13, and FIG. 23B shows the shape of the edge SE13 corresponding to the container C13, which is stored in the image area information 154 (see FIG. 7). In this way, the image area information corresponding to the physical tag PT13 provided on the container C13 is stored in the image area information 154. Note that the method of setting the image area is the same as in the above-described example. Furthermore, as in FIGS. 13, 21A, 21B, 22A, and 22B, multiple physical tags may be provided on the container C13.

[0206] [Modification of information processing system] The above describes an example in which tag information of the containers C1 to C11 is acquired based on captured images generated by the imaging device 200. Here, the tag information of the containers C1 to C11 may be acquired using wireless communication. Therefore, the following describes an example in which tag information of the containers C1 to C11 is acquired using wireless communication.

[0207] [Example of use of information processing system] Fig. 24 is a diagram showing an example of use of the information processing system 1a. The information processing system 1a shown in Fig. 24 shows an example in which a communication device 400 is added to the information processing system 1 shown in Fig. 1, and physical tags PTT1 to PTT11 capable of wireless communication are provided on the containers C1 to C11. Note that, since other parts are common to the information processing system 1, the same reference numerals as in Fig. 1 are used and their description will be omitted. In the following, the physical tag PTT1 will be mainly described as an example, but the same applies to the physical tags PTT2 to PTT11.

[0208] The physical tag PTT1 is an integrated circuit (IC) tag capable of wireless communication. For example, the physical tag PTT1 can be provided inside a sheet that can be attached to the container C1. The physical tag PTT1 is an example of RFID (Radio Frequency Identification) and is a device capable of low-power wireless communication. The physical tag PTT1 can be configured as a wireless communication tag (RFID tag) compatible with RFID technology. For example, the physical tag PTT1 can be configured as an IC tag that employs a communication method such as Bluetooth Low Energy (BLE), Bluetooth, ZigBee, Low Power Wide Area (LPWA), or Ultra Wide Band (UWB), which are low-power communication modes. The maximum communication distance of the physical tag PTT1 is not particularly limited, but can be, for example, in the range of several tens of centimeters to several meters. For example, if the physical tag PTT1 complies with the BLE standard, it broadcasts packets at predetermined intervals (for example, every 1 to 10 seconds). A packet transmitted by this physical tag PTT1 includes a unique ID (Identification) (tag ID) that is identification information of the IC tag. Figures 24 to 26 show an example in which this tag ID is used as identification information (tag information) of the container C1 (waste).

[0209] That is, by reading the tag ID of the physical tag PTT1 attached to the container C1, it is possible to obtain the identification information of the waste (vinyl bag) contained in the container C1.

[0210] Note that identification information that allows tag information to be acquired from the captured image generated by the imaging device 200 may be provided on the surface of the sheet provided with the physical tag PTT1. For example, triangular identification information PTT1a (see FIG. 25) that allows identification of the base point, size, and direction may be provided. Note that color codes or the like shown in FIGS. 2 to 5 may also be provided.

[0211] [Physical tag configuration example] 25 is a diagram showing an example of identification information PTT1a provided on the surface of a sheet having a physical tag PT1a. Note that the identification information PTT1a can be colored to identify the attributes of the waste contained in the container C1.

[0212] The identification information PTT1a is a figure formed by an isosceles triangle. The center position of the base of the isosceles triangle is defined as a base point PP1, the length of a line segment PP2 from the base point PP1 to the apex angle of the isosceles triangle is defined as a size PPL1 of the identification information PTT1a, and the angle of the line segment PP2 with respect to the horizontal direction is defined as a direction θ1 of the identification information PTT1a. The base point PP1, the size PPL1 of the identification information PTT1a, and the direction θ1 of the identification information PTT1a can be determined based on the identification information PTT1a included in the captured image. The base point PP1, the size PPL1 of the identification information PTT1a, and the direction θ1 of the identification information PTT1a can be used to perform affine transformation on the image region information stored in the image region information 154 (see FIG. 7).

[0213] [Examples of using information processing equipment] Fig. 26 is a block diagram showing an example of the system configuration of the information processing system 1a. The information processing system 1a is a partial modification of the information processing system 1 shown in Fig. 6. Specifically, the information processing system 1a is provided with a communication device 400 for reading the physical tag PTT1. Note that apart from the provision of the communication device 400, the information processing system 1a has the same components as the information processing system 1. Therefore, the same reference numerals are used for the components common to the information processing system 1, and some of these components will not be illustrated or described. Furthermore, the components different from the information processing system 1 will be described below as appropriate.

[0214] The information processing system 1a is configured by a plurality of devices that can be connected via a network N1. Fig. 26 shows, as an example, the information processing system 1a including an information processing device 100, an imaging device 200, a user terminal 300, and a communication device 400. Each of the information processing device 100, the imaging device 200, the user terminal 300, and the communication device 400 is connected to the network N1 by a communication method using wired communication or wireless communication.

[0215] Note that the information processing device 100, the imaging device 200, the user terminal 300, and the communication device 400 may be connected directly using wired or wireless communication without going through the network N1. Also, while Fig. 26 shows only the imaging device 200, the user terminal 300, and the communication device 400 as representative examples, the information processing system 1a may also include other imaging devices, other user terminals, and other communication devices installed at each location. Furthermore, as will be described later, if the position of the physical tag PTT1 can be identified by the communication device 400, the installation of the imaging device 200 may be omitted.

[0216] The communication device 400 and the physical tag PTT1 are connected by a direct connection using wireless communication without going through the network N1. Also, in Fig. 26, only the physical tag PTT1 and the container C1 are illustrated as representative examples, but the other physical tags PTT2 to PTT11 and the other containers C2 to C11 also constitute the information processing system 1a. Furthermore, as will be described later, when a plurality of communication devices 400 are used, the information processing system 1a is constituted by the plurality of communication devices 400.

[0217] The communication device 400 includes a communication unit 401 , a reading unit 402 , a control unit 403 , and a storage unit 404 .

[0218] The communication unit 401 exchanges various types of information with the information processing device 100 using wireless communication under the control of the control unit 403 .

[0219] The reading unit 402 is a communication unit that receives radio waves from the physical tag PTT1 provided on the container C1 and exchanges various information with the physical tag PTT1. Note that the wireless communication used by the physical tag PTT1 described above can be adopted as the wireless communication exchanged between the reading unit 402 and the physical tag PTT1. For example, the reading unit 402 reads the tag ID of the physical tag PTT1 and outputs this tag ID to the control unit 403.

[0220] The control unit 403 controls each unit of the communication device 400 based on a control program stored in the storage unit 404. The control unit 403 is realized by a processing device such as a CPU or a GPU. For example, the control unit 403 executes control to transmit the tag ID read by the reading unit 402 to the information processing device 100.

[0221] The storage unit 404 is a storage medium that stores various types of information. For example, the storage unit 404 stores various types of information (for example, a control program, location identification information for identifying the location where the communication device 400 is installed) that is required for the control unit 403 to perform various processes. The storage unit 404 also stores various types of information acquired via the communication unit 401. The storage unit 404 can be, for example, a ROM, a RAM, an SRAM, an HDD, an SSD, or a combination thereof.

[0222] [Example of location measurement using wireless tags] Here, a description will be given of a measurement method for measuring the position of a container using a wireless tag as the physical tag PTT 1. For example, the position of the wireless tag (position of the container) can be measured using a known indoor positioning system or the like.

[0223] For example, it is possible to use a measurement method that estimates the position of a wireless tag based on radio waves (radio waves emitted by the wireless tag) received by multiple communication devices. For example, three or more receivers can be installed at a waste collection site WP1 (see Figure 24), and the radio waves (radio waves emitted by the wireless tag) received by these receivers can be acquired. The position of the wireless tag can then be estimated by triangulation (cross-azimuth method) using the radio wave intensities.

[0224] Furthermore, for example, a measurement method can be used in which angle information between one or more communication devices and a wireless tag is calculated based on communication between the communication devices and the wireless tag, and the position of the wireless tag is estimated based on the angle information. For example, a reception angle detection technique (AoA (Angle of Arrival)) or a radiation angle detection technique (AoD (Angle of Departure)) can be used to calculate the angle information. For example, assume that containers C1 to C11 are located within a one-floor facility (waste collection site WP1) and an AoA receiver is installed on the ceiling of that floor. In this case, it is possible to calculate the angle of incidence of the radio waves (radio waves emitted by the wireless tag) received by the receiver installed on the ceiling. Then, the position of each wireless tag present on the floor of the waste collection site WP1 can be estimated based on the angle of incidence. Alternatively, the position of each wireless tag may be estimated using a receiver that can detect the position of the wireless tag that emitted the radio waves based on the directionality and reception strength of the received radio waves.

[0225] Using these position estimation techniques, it is possible to estimate the position of the physical tag PTT1 in three-dimensional space (waste collection point WP1). The estimation result (tag information and position) is transmitted by the communication device 400 to the information processing device 100. Note that the communication device 400 may transmit measurement data of the received radio waves to the information processing device 100, and the information processing device 100 may estimate the position of the physical tag PTT1 using the above-mentioned position estimation method.

[0226] Furthermore, by fixing the imaging range IM1 of the imaging device 200, it is possible to associate a three-dimensional space (the waste collection point WP1) with a two-dimensional space (the captured image corresponding to the imaging range IM1). Therefore, the tag information acquisition unit 122 can estimate the position of the physical tag PTT1 in the captured image (corresponding to the imaging range IM1) based on the position of the physical tag PTT1 at the waste collection point WP1.

[0227] In this way, the tag information acquisition unit 122 acquires the tag ID (tag information) of the physical tag PTT1 acquired by the communication device 400 and the position of the physical tag PTT1 in the captured image (corresponding to the imaging range IM1) estimated by the above-mentioned estimation method. Then, the tag information acquisition unit 122 associates the position of the physical tag PTT1 (position in the captured image) with the acquired tag information and outputs them to the setting unit 123, the annotation unit 124, the generation unit 125, and the selection unit 126.

[0228] The setting unit 123 sets a predetermined image area from the captured image output from the image acquisition unit 121 based on the position and tag information of the physical tag PTT1 acquired by the tag information acquisition unit 122. For example, it is possible to set an image area of ​​a predetermined range based on the position of the physical tag PTT1 from the captured image corresponding to the imaging range IM1 (see FIG. 1). For example, it is possible to set a rectangular image area with the position of the physical tag PTT1 as the center and a size according to the attribute corresponding to the tag information.

[0229] 25, the identification information PTT1a of the physical tag PTT1 can be read, and based on this identification information PTT1a, the base point PP1, the size PPL1 of the identification information PTT1a, and the direction θ1 of the identification information PTT1a can be identified. Therefore, the setting unit 123 can affine transform the image region information (corresponding to the tag information of the physical tag PTT1) stored in the image region information 154 (see FIG. 7) using the base point PP1, the size PPL1 of the identification information PTT1a, and the direction θ1 of the identification information PTT1a. In this case, the setting unit 123 identifies the edge of the opening of the container C1 in the captured image based on the information about the physical tag PTT1 (the base point PP1, the size PPL1 of the identification information PTT1a, and the direction θ1 of the identification information PTT1a) and the tag information output from the tag information acquisition unit 122 (tag information corresponding to the physical tag PTT1).

[0230] Specifically, the setting unit 123 extracts image area information 154 corresponding to the tag information (tag information corresponding to the physical tag PTT1) output from the tag information acquisition unit 122 from the setting information DB 150 (see FIG. 7). Then, the setting unit 123 identifies the edge of the opening of the container C1 based on the extracted image area information and information related to the physical tag PTT1 (base point PP1, size PPL1 of the identification information PTT1a, and direction θ1 of the identification information PTT1a). This identification method is the same as the identification method shown in FIGS. 11 to 14, etc. Next, the setting unit 123 identifies an image area including the edge of the opening of the identified container C1, and sets the image area in the captured image.

[0231] The annotation unit 124 performs annotation on the object included in the captured image output from the image acquisition unit 121 based on the tag information output from the tag information acquisition unit 122. Here, if it is possible to estimate the position of the physical tag PTT1 in the captured image (corresponding to the imaging range IM1) using wireless communication, it is possible to associate the physical tag PTT1 with the image area including it based on the position. Also, if it is not possible to estimate the position of the physical tag PTT1 using wireless communication, it is possible to associate the identification information PTT1a of the physical tag PTT1 with the image area specified based on the color of the identification information PTT1a.

[0232] The estimation unit 127 uses the trained model selected by the selection unit 126 to estimate the state of waste stored in a container included in the captured image output from the image acquisition unit 121. Here, if it is possible to estimate the position of the physical tag PTT1 in the captured image (corresponding to the imaging range IM1) using wireless communication, it is possible to associate the trained model with the image area input thereto based on the position. Also, if it is not possible to estimate the position of the physical tag PTT1 using wireless communication, it is possible to associate the trained model corresponding to the identification information PTT1a of the physical tag PTT1 with the image area identified based on the color of the identification information PTT1a.

[0233] [Example of estimating waste weight using multiple trained models] In the above example, the state of waste (filling rate, weight) is estimated using one trained model according to the attributes of the waste contained in a container. However, it is also possible to estimate the state of waste (filling rate, weight) using multiple trained models according to the attributes of the waste.

[0234] Therefore, an example of estimating the weight of waste using multiple trained models will be described below. For example, it is possible to use a trained model (volume) that estimates the volume of waste contained in a container, a trained model (shape) that estimates the shape of the waste contained in the container, and a trained model (density) that estimates the density of the waste contained in the container.

[0235] The trained model (capacity) is, for example, a model that learns the volume (filling rate) as a state of waste. Specifically, it is possible to generate the trained model (capacity) by training a large amount of training data that associates images including the target waste (waste contained in a container) with the filling rate and weight of the waste in the container. The trained model (shape) is, for example, a model that learns the shape (volume) as a state of waste. Specifically, it is possible to generate the trained model (shape) by training a large amount of training data that associates images including the target waste (waste contained in a container) with the shape and weight of the waste in the container. The trained model (density) is, for example, a model that learns the density as a state of waste. Specifically, it is possible to generate the trained model (density) by training a large amount of training data that associates images including the target waste (waste contained in a container) with the density and weight of the waste in the container.

[0236] In addition, combinations of these trained models (trained model (volume), trained model (shape), trained model (density)) are generated for each waste attribute.

[0237] Furthermore, when estimating the weight of waste, it is possible to use the output values ​​(respective output values ​​for the volume, shape, and density of the target waste) from three trained models (trained model (volume), trained model (shape), trained model (density)) according to the attributes of the target waste. In other words, it is possible to estimate the weight of waste based on the three output values.

[0238] [Example of using a trained model (capacity)] For example, it is possible to estimate the weight of waste using the output result (filling rate v) of the trained model (capacity) and the following equation 1. w v =f(v)=k1·v …Equation 1

[0239] Here, the filling rate v is output as a value between 0 and 1, where 1 is the value when the container for the target waste is full. Also, k1 is a coefficient set based on the weight at a filling rate of 1 (100%).

[0240] For example, suppose the weight of PET garbage is to be estimated as waste. In this case, the ratio of PET bottles of various sizes and the filling rate are taken into consideration, and it is assumed that the weight when full is 60 kg. In this case, it is possible to set k1 = 60. The coefficient k1 can be set appropriately based on experiments, simulations, etc.

[0241] [Example using a trained model (shape)] For example, it is possible to estimate the weight of atypical waste using the output result (volume s) of the trained model (shape) and the following equation 2. w s =f(s)=k2·s …Equation 2

[0242] Here, k2 is a coefficient set based on specific gravity.

[0243] For example, let us consider the case where the weight of PET waste is to be estimated. In this case, using a known image recognition technique, it is possible to detect PET waste (non-standard waste) other than standard PET bottles (uncompressed and compressed shapes) based on their shape from among the objects contained in the captured image. For example, in the case of a standard PET board (board thickness 5 mm), the volume s when the length h (cm) × width I (cm) is 0.05 × h × I (cm 3 In this case, if k2 = 1.4 (specific gravity), the weight w s is 1.4×s(g).

[0244] In this way, when atypical waste is detected, it is possible to add the weight of the atypical waste estimated using the trained model (shape).

[0245] [Example using a trained model (density)] For example, it is possible to estimate the weight of waste using the output result (density d) of the trained model (density) and the following equation 3. w d =f(d)=k1·d·Δv …Equation 3

[0246] Here, the density d is the reciprocal of the compression rate. For example, if the compression rate is 30%, then d = 1 / 0.3.

[0247] For example, suppose you want to estimate the weight of PET waste. In this case, you can check the captured images in chronological order and measure the volume Δv and density d at the point when the waste density changes. Then, you can subtract the measurement results to estimate the weight of the waste.

[0248] Taking these into consideration, the weight w of the waste can be calculated using the following equation 4. Note that the following equation 4 is an example in which captured images generated in time series are used, and k is a value indicating the time axis.

[0249]

number

[0250] The first term of Equation 4 represents the estimated weight of standard waste. The second term of Equation 4 represents the estimated weight of non-standard waste. The third term of Equation 4 represents the estimated weight of compressed standard waste. That is, when non-standard waste is detected based on its shape from among the objects included in the captured image, the volume at that time is subtracted from the first term. When a change in density of compressed or other waste is detected from among the objects included in the captured image, the compressed volume at that time is subtracted from the first term. The weight based on the subtracted volume of the non-standard waste is then calculated using the second term. The weight based on the volume of compressed or other waste is then calculated using the third term. Note that Equation 4 is an example shown for ease of explanation, and other equations determined based on experiments, simulations, etc. may also be used.

[0251] In this way, the information processing device 100 can estimate the weight of the waste using multiple trained models according to the attributes of the target waste (a trained model corresponding to the volume of the waste, a trained model corresponding to the shape of the waste, and a trained model corresponding to the density of the waste). This makes it possible to improve the accuracy of estimating the weight of the waste.

[0252] [Example of effect of this embodiment] In recent years, image recognition processing using AI has become widespread. For example, when performing object detection using YOLO, detection and classification processes must be performed simultaneously to detect the target object in the image. For this purpose, the image is divided into grid cells, multiple bounding boxes (BBoxes) and confidence scores are calculated, and class (attribute) predictions are calculated at the same time. Finally, these are multiplied to calculate a confidence score for the class (attribute) of the object contained in the BBox.

[0253] In this way, when AI-based image recognition processing (for example, object detection using YOLO) simultaneously performs detection processing and classification processing for object detection, a huge amount of computational processing is required. This requires high-performance computing equipment and consumes a large amount of power. Furthermore, with AI, the confidence score is merely a probability and does not indicate 100% certainty that the object is the one in question.

[0254] Therefore, in this embodiment, containers C1 to C11 that store waste are provided with physical tags PT1 to PT11 and PTT1 to PTT11 from which tag information can be read. An imaging device 200 (or a communication device 400) is also installed at the waste disposal site. The information processing device 100 acquires tag information for identifying each waste based on a captured image (or wireless communication by the communication device 400) that includes the containers C1 to C11 generated by the imaging device 200. Next, the information processing device 100 automatically determines the attributes of each waste using the tag information and automatically sets an image area including each waste. That is, the attributes of each waste and the image area including each waste can be acquired using the tag information. This makes it possible to set the desired image area more reliably and accurately, reduce the computational load related to calculation processing, and shorten the time required for each process. It also makes it possible to reduce power consumption related to each process. For example, it is also possible to facilitate annotation work, such as specifying complex image areas from a target image. In this way, in this embodiment, when performing AI processing to set an image area, it is possible to reduce the calculation load associated with each process and shorten the time required for each process.

[0255] [Examples of application to other objects] In the above, waste has been described as an example of an object contained in a container. That is, an example has been shown in which annotation is performed on waste (object) contained in a container. However, this embodiment can also be applied to objects other than waste as objects contained in a container. For example, this embodiment can also be applied when annotation is performed on agricultural products, marine products, industrial products, etc. as objects contained in a container. This makes it possible to easily perform annotation on objects at the processing site of each object, and to realize efficient generation of trained models.

[0256] In the above, an example of performing annotation on an object contained in a container has been shown. However, this embodiment can also be applied to an object that is not contained in a container. In this case, for example, a physical tag can be attached to the object itself or its surroundings (for example, a sign behind it), and annotation can be performed on the object using this physical tag.

[0257] [Example of executing processing on other devices or systems] Although the above describes an example in which tag information acquisition processing, setting processing, annotation processing, generation processing, etc. are executed in the information processing device 100, all or part of each of these processes may be executed in other devices. In this case, an information processing system is configured by the devices that execute part of each of these processes. For example, at least part of each process can be executed using devices available to the user (e.g., smartphones, tablet terminals, personal computers), various information processing devices such as servers that can be connected via a predetermined network such as the Internet, and various electronic devices.

[0258] Furthermore, a part (or all) of the information processing system capable of executing the functions of the information processing device 100 (or the information processing system 1, 1a) may be provided by an application that can be provided via a predetermined network such as the Internet. This application is, for example, SaaS (Software as a Service).

[0259] [Configuration example and effect example of this embodiment] The main effects of the information processing device 100, the information processing method, and the program according to the embodiment of the present invention will be described below.

[0260] The information processing device 100 has an image acquisition unit 121 that acquires an image including waste (an example of an object), a tag information acquisition unit 122 that acquires tag information that identifies the waste, and a setting unit 123 that sets an image area related to the object in the image based on the tag information.

[0261] This makes it possible to set an image area for waste based on tag information that identifies that waste, which reduces the computational load and time required for each process when performing AI processing to set the image area.

[0262] The setting unit 123 extracts and sets an image area including the object from the image based on the tag information.

[0263] This allows only the image area containing waste to be the target of processing (e.g., annotation processing, object detection processing, object state estimation processing), thereby reducing the amount of calculation required for each process and shortening the time required for each process. For example, since only the image area containing waste can be the target of annotation processing, the accuracy of annotation can be improved. Furthermore, for example, since only the image area containing waste can be the target of object state estimation processing, the accuracy of estimation processing can be improved.

[0264] Physical tags PT1 to PT11 and PTT1 to PTT11 from which tag information can be read are provided on or near waste (an example of a target object). A tag information acquisition unit 122 acquires tag information using the physical tags PT1 to PT11 and PTT1 to PTT11. A setting unit 123 sets an image area based on the positions of the physical tags PT1 to PT11 and PTT1 to PTT11 in the image.

[0265] This makes it possible to acquire tag information based on the physical tags PT1 to PT11 and PTT1 to PTT11 attached to the waste or nearby, and to set an image area based on the positions of the physical tags PT1 to PT11 and PTT1 to PTT11 in the image, thereby making it possible to appropriately set an image area for each waste.

[0266] The physical tags PT1 to PT11 and PTT1 to PTT11 are information having a predetermined shape. For example, the shape is made up of multiple rectangles (see FIGS. 2 and 3), or multiple figures (see FIGS. 4 and 5). The tag information acquisition unit 122 acquires tag information based on the shapes of the physical tags PT1 to PT11 and PTT1 to PTT11. For example, the tag information acquisition unit 122 can acquire tag information based on the shapes of the figures (e.g., rectangle, triangle, circle) that make up the physical tags PT1 to PT11 and PTT1 to PTT11.

[0267] This makes it possible to acquire tag information based on the shapes of the physical tags PT1 to PT11 and PTT1 to PTT11 placed on or near the waste. By placing physical tags of each shape on or near the waste, it becomes possible to appropriately set an image area related to the waste.

[0268] The physical tags PT1 to PT11 and PTT1 to PTT11 are provided in a predetermined orientation with respect to the waste (an example of a target object). The setting unit 123 sets an image area based on the orientations of the physical tags PT1 to PT11 and PTT1 to PTT11 in the image. For example, as shown in FIG. 12, the setting unit 123 can set the image area by affine transforming the image area information (e.g., FIG. 9A) based on the orientation of the physical tag PT1.

[0269] This makes it possible to appropriately set an image area relating to the waste based on the orientation of the physical tags PT1 to PT11 and PTT1 to PTT11 provided on or near the waste.

[0270] In the setting information DB 150, tag information 151 is associated with image area information 154 (an example of graphic selection information). The setting unit 123 selects the image area information 154 associated with the tag information 151, and sets the image area using the selected image area information.

[0271] According to this, by selecting image area information associated with tag information corresponding to waste, it is possible to appropriately set an image area relating to that waste.

[0272] The physical tags PT1 to PT11 and PTT1 to PTT11 are provided at predetermined positions and in predetermined orientations relative to the waste (an example of a target object). The setting unit 123 sets an image area based on the positions and orientations of the physical tags PT1 to PT11 and PTT1 to PTT11 in the image. For example, as shown in FIG. 12, the setting unit 123 can set the image area by affine transforming the image area information (e.g., FIG. 9A) based on the position and orientation at which the physical tag PT1 is provided.

[0273] This makes it possible to appropriately set an image area relating to the waste based on the positions and orientations of the physical tags PT1 to PT11 and PTT1 to PTT11 provided on or near the waste.

[0274] The physical tags PT1 to PT11 and PTT1 to PTT11 are provided at predetermined positions and in predetermined orientations relative to the waste (an example of a target object). In the setting information DB 150, tag information 151 is associated with image area information 154 (an example of graphic selection information). The setting unit 123 selects the image area information 154 associated with the tag information 151, and sets the image area based on the selected image area information and the positions and orientations of the physical tags PT1 to PT11 and PTT1 to PTT11 in the image. For example, as shown in FIG. 12, the setting unit 123 can set the image area by affine transforming the selected image area information (e.g., FIG. 9A) based on the position and orientation of the physical tag PT1.

[0275] This makes it possible to appropriately set an image area related to the waste based on the position and orientation of the physical tags PT1 to PT11, PTT1 to PTT11 placed on or near the waste, and the image area information associated with the tag information corresponding to the physical tag.

[0276] The information processing device 100 further includes an annotation unit 124 that performs annotation on the image area regarding waste (an example of an object) based on the tag information.

[0277] This allows annotation of the waste based on tag information that identifies the waste to be performed on the image area related to the waste, thereby reducing the time required for annotation when generating a trained model and improving the accuracy of the annotation.

[0278] The information processing device 100 further includes a selection unit 126 that selects at least one trained model from among a plurality of trained models based on tag information, and an estimation unit 127 that inputs an image in which an image area is set into the selected trained model and estimates the state (e.g., filling rate, weight) of waste (an example of a target object) based on the output from the trained model.

[0279] According to this, when using a trained model to estimate the state of waste (e.g., weight, filling rate) from an image area related to the waste, it is possible to select and use a trained model corresponding to the waste from among multiple trained models based on tag information that identifies the waste. This eliminates the need to simultaneously use multiple trained models to perform estimation processing to estimate the state of the waste from an image containing the waste, making it possible to reduce the amount of calculation involved in the estimation processing and shorten the processing time.

[0280] The selection unit 126 selects from among the multiple trained models a trained model corresponding to the volume of the waste (an example of an object) identified according to the tag information, a trained model corresponding to the shape of the waste, and a trained model corresponding to the density of the waste, and the estimation unit 127 estimates the state of the object from the image area using each of the selected trained models.

[0281] For example, when waste is made of resins such as PET, it is possible that the waste contains a mixture of standardized objects such as PET bottles and non-standardized objects such as rolls or plates. In this case, the volume of the standardized objects differs from the volume of the non-standardized objects, which may result in a large error in the weight calculated based on the volume. It is also possible that the waste contains a mixture of standardized objects (e.g., uncrushed PET bottles) and compressed standardized objects (e.g., crushed PET bottles). In this case, the volume of the standardized objects differs from the volume of the compressed standardized objects, which may result in a large error in the weight calculated based on the volume. Therefore, it is possible to improve the estimation accuracy by estimating the weight of waste using multiple trained models according to the attributes of the waste.

[0282] The physical tags PT1 to PT11 and PTT1 to PTT11 are provided on the plurality of containers C1 to C11, respectively. The setting unit 123 sets image areas corresponding to the plurality of containers C1 to C11 in the image for each of the containers C1 to C11.

[0283] This makes it possible to set an image area corresponding to each of the multiple containers C1 to C11 for each container C1 to C11, thereby improving the accuracy of setting the image area related to the waste contained in each of the containers C1 to C11.

[0284] The physical tags PT1 to PT11, PTT1 to PTT11 can be at least one of a barcode (see Figures 2 and 3), a two-dimensional code, a color code capable of expressing multiple pieces of information using an arrangement of multiple colors (see Figures 2 and 3), a graphic code capable of expressing multiple pieces of information using an arrangement of multiple figures each colored in one of multiple colors (see Figures 4 and 5), and a wireless tag capable of transmitting tag information using wireless communication (see Figures 18 to 20).

[0285] This makes it possible to provide appropriate physical tags on or near the containers C1 to C11 that accommodate waste, depending on the environment of the installation location of the containers C1 to C11. This makes it possible to acquire tag information from the physical tags PT1 to PT11 and PTT1 to PTT11 depending on the environment of the installation location of the containers C1 to C11.

[0286] When a wireless tag is used as the physical tag PTT1, identification information PTT1a (an example of direction specifying information) indicating the direction of the physical tag PTT1 in the image is provided together with the wireless tag.

[0287] This makes it possible to identify the position of the physical tag PTT1 by the wireless tag, and to identify the direction of the physical tag PTT1 by the identification information PTT1a, thereby improving the accuracy of detecting the position and direction of the container C1 associated with the physical tag PTT1.

[0288] In addition to waste, at least one of agricultural products, marine products, and industrial products may be the target object.

[0289] This makes it possible to reduce the computational load associated with each process and shorten the time required for each process when setting an image area related to the target object (at least one of waste, agricultural products, marine products, and industrial products).

[0290] An information processing method according to an embodiment of the present invention includes an image acquisition process (steps S501, S511) for acquiring an image including waste (an example of a target object), a tag information acquisition process (steps S502, S503, S512, S513) for acquiring tag information for identifying the waste, and a setting process (steps S505, S514) for setting an image area relating to the waste in the image based on the tag information.

[0291] This makes it possible to set an image area for waste based on tag information that identifies that waste, which reduces the computational load and time required for each process when performing AI processing to set the image area.

[0292] A program according to an embodiment of the present invention causes a computer to execute an image acquisition procedure (steps S501, S511) for acquiring an image including waste (an example of a target object), a tag information acquisition procedure (steps S502, S503, S512, S513) for acquiring tag information for identifying the waste, and a setting procedure (steps S505, S514) for setting an image area relating to the waste in the image based on the tag information.

[0293] This makes it possible to set an image area for waste based on tag information that identifies that waste, which reduces the computational load and time required for each process when performing AI processing to set the image area.

[0294] Note that each processing procedure shown in this embodiment is an example for realizing this embodiment, and the order of some of the processing procedures may be changed within the scope that makes it possible to realize this embodiment, and some of the processing procedures may be omitted or other processing procedures may be added.

[0295] Each process described in this embodiment is executed based on a program that causes a computer to execute each processing procedure. Therefore, this embodiment can also be understood as an embodiment of a program that realizes the function of executing each process and a recording medium that stores the program. For example, an update process for adding a new function to an information processing device can store the program in the storage device of the information processing device. This makes it possible to cause the updated information processing device to execute each process described in this embodiment.

[0296] Although the embodiments of the present invention have been described above, the above embodiments merely illustrate some of the application examples of the present invention, and it is not intended that the technical scope of the present invention be limited to the specific configurations of the above embodiments. [Explanation of symbols]

[0297] 1, 1a Information Processing System 100 Information processing device 110 Communications Department 120 control section 121 Image acquisition unit 122 Tag information acquisition unit 123 Settings 124 Annotation Section 125 Generation part 126 Selection Section 127 Estimation Department 128 Output control section 130 Storage section 140 Tag DB 150 Setting information DB 160 trained models DBDB 170 Estimation result DB 200 Imaging device 201 Communications Department 202 Control section 203 Storage section 204 Imaging unit 300 User Terminals 301 Communications Department 302 Control Unit 303 Storage section 304 Operation section 305 Sound acquisition section 306 Display section 307 Sound output unit 400 Communication equipment 401 Communications Department 402 Reading unit 403 Control Unit 404 Storage section C1~C11, C13 containers N1 Network PT1~PT11, PT1a, PT11a, PT13, PT20, PT30, PT40, PTT1~PTT11 Physical tags WP1 Waste collection point

Claims

1. an image acquisition unit that acquires an image including an object; a tag information acquisition unit that acquires tag information that identifies the object; a setting unit that sets an image area related to the object in the image based on the tag information; An information processing device having the above.

2. 2. The information processing device according to claim 1, the setting unit extracts and sets the image area including the object from the image based on the tag information. Information processing device.

3. 2. The information processing device according to claim 1, a physical tag capable of reading the tag information is provided on or near the object; the tag information acquisition unit acquires the tag information using the physical tag; the setting unit sets the image area based on a position of the physical tag in the image. Information processing device.

4. 4. The information processing device according to claim 3, The physical tag is information having a predetermined shape, the tag information acquisition unit acquires the tag information based on the shape of the physical tag; Information processing device.

5. 4. The information processing device according to claim 3, the physical tag is provided in a predetermined orientation with respect to the object; the setting unit sets the image area based on an orientation of the physical tag in the image. Information processing device.

6. 2. The information processing device according to claim 1, the tag information is associated with feature selection information; the setting unit selects the graphic selection information associated with the tag information, and sets the image area using the graphic selection information. Information processing device.

7. 4. The information processing device according to claim 3, the physical tag is provided at a predetermined position and in a predetermined orientation relative to the object; the setting unit sets the image area based on a position and an orientation of the physical tag in the image. Information processing device.

8. 4. The information processing device according to claim 3, the physical tag is provided at a predetermined position and in a predetermined orientation relative to the object; the tag information is associated with feature selection information; the setting unit selects the graphic selection information associated with the tag information, and sets the image area based on the graphic selection information and the position and orientation of the physical tag in the image; Information processing device.

9. 2. The information processing device according to claim 1, The method further includes an annotation unit that performs annotation on the image region regarding the object based on the tag information. Information processing device.

10. 2. The information processing device according to claim 1, a selection unit that selects at least one trained model from among a plurality of trained models based on the tag information; An estimation unit that inputs the image in which the image area is set to the selected trained model and estimates the state of the object based on an output from the trained model, Information processing device.

11. The information processing device according to claim 10, the selection unit selects, from the plurality of trained models, a trained model corresponding to a volume of the object, a trained model corresponding to a shape of the object, and a trained model corresponding to a density of the object based on attributes of the object identified according to the tag information; The estimation unit estimates the state of the object from the image region using each selected trained model. Information processing device.

12. 4. The information processing device according to claim 3, the physical tag is provided on each of the plurality of containers; the setting unit sets the image area corresponding to each of the plurality of containers in the image for each of the containers. Information processing device.

13. 4. The information processing device according to claim 3, The physical tag is at least one of a barcode, a two-dimensional code, a color code capable of expressing a plurality of pieces of information by an arrangement of a plurality of colors, a graphic code capable of expressing a plurality of pieces of information by an arrangement of a plurality of graphics to which any of a plurality of colors is applied, and a wireless tag capable of transmitting the tag information by wireless communication. Information processing device.

14. 14. The information processing device according to claim 13, When the wireless tag is used as the physical tag, direction specifying information indicating the direction of the physical tag in the image is provided together with the wireless tag. Information processing device.

15. 15. An information processing device according to claim 1, The object is at least one of waste, agricultural products, marine products, and industrial products. Information processing device.

16. an image acquisition process for acquiring an image including the object; a tag information acquisition process for acquiring tag information for identifying the object; a setting process of setting an image area relating to the object in the image based on the tag information; An information processing method including:

17. an image acquisition step for acquiring an image including the object; a tag information acquisition step of acquiring tag information that identifies the object; a setting step of setting an image area relating to the object in the image based on the tag information; A program that causes a computer to execute the following.

Citation Information

Patent Citations

  • Analog meter specification recognition device, computer program, and analog meter specification recognition method

    JP2020135353A