Image processing apparatus and image processing method
Patent Information
- Application Number
- JP2022071824
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-04-25
- Publication Date
- 2026-09-04
- Estimated Expiration
- 2042-04-25
AI Technical Summary
【0010】 本発明によれば、ユーザが指定した属性を有する教師画像の収集状況、すなわち、必要な属性の教師画像が偏りなく揃っているか否かを、ユーザが容易に確認することができる。特に、教師画像の抽出枚数の多い位置または教師画像の抽出枚数の少ない位置が可視化画像で提示されるため、教師画像の収集状況に問題のあるエリア画像上の位置を、ユーザが容易に把握することができる。これにより、学習に先だって、教師画像のアノテーション状況をユーザが目視で容易に確認でき、効率よく高精度な学習モデルを作成することができる。
Smart Images

Figure 0007915432000001 
Figure 0007915432000002 
Figure 0007915432000003
Abstract
Description
Technical Field
[0001] The present invention relates to an image processing apparatus and an image processing method for visualizing the collection status of teacher images for constructing an image recognition model corresponding to a monitoring area.
Background Art
[0002] Systems that detect predetermined events such as a person visiting a store from images captured by a camera using an image recognition model (machine learning model) constructed by machine learning such as deep learning are in use. An image recognition model is constructed by machine learning using a large number of collected teacher images (training images). However, if there is a bias in the teacher images, an image recognition model with stable accuracy cannot be constructed.
[0003] In order to avoid a decrease in accuracy of the image recognition model caused by such bias in teacher images, there has been conventionally known a technique for reducing bias in training data (teacher images) by changing the probability distribution of entities of persons and stores (components) existing in a monitoring area (application environment) that is a processing target of the image recognition model (see Patent Document 1).
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Summary of the Invention
Problem to be Solved by the Invention
[0005] Conventional techniques update the training image dataset to reduce bias in the training images, increasing the likelihood of building a highly accurate machine learning model. However, the accuracy of the machine learning model may still be insufficient. Therefore, when evaluating the built machine learning model, if its accuracy is deemed insufficient, the training image dataset is updated by adding missing training images, and the machine learning process is repeated to evaluate the built machine learning model. Thus, conventional techniques require repeatedly updating the training image dataset, performing machine learning, and evaluating the machine learning model, which can be very time-consuming before a sufficiently accurate machine learning model is completed.
[0006] On the other hand, if the status of training image collection (annotation status), that is, whether a sufficient number of training images with the necessary attributes are available and distributed appropriately, is visualized and presented to the user, the user can immediately grasp the status of training image collection and efficiently perform annotation work to add any missing training images.
[0007] Therefore, the main objective of the present invention is to provide an image processing device and an image processing method that allow users to easily visually confirm the status of teacher image collection prior to training, and to efficiently create highly accurate learning models. [Means for solving the problem]
[0008] The present invention is an image processing apparatus that uses a processor to perform a process to visualize the status of collecting training images for constructing an image recognition model corresponding to a monitoring area, wherein the processor generates a training image including the object to be detected and the background from an area image relating to the monitoring area, and relates to the characteristics of the object to be detected included in the training image A region attribute value representing an attribute is assigned to each of the aforementioned training images. Targeting the training image having the attributes specified by the user, A mapping process is performed to assign the region attribute values assigned to each of the aforementioned training images to each position in the area image, the number of training images extracted at each position in the area image is measured, and positions with a large number of extracted training images or positions with a small number of extracted training images are determined. The system generates a visualized image and outputs display information by superimposing the visualized image onto the area image.
[0009] Furthermore, the present invention is an image processing method in which a processor performs a process to visualize the status of collecting training images for constructing an image recognition model corresponding to a monitoring area, and generates a training image including the object to be detected and the background from an area image relating to the monitoring area, and relates to the characteristics of the object to be detected included in the training image A region attribute value representing an attribute is assigned to each of the aforementioned training images. Targeting the training image having the attributes specified by the user, A mapping process is performed to assign the region attribute values assigned to each of the aforementioned training images to each position in the area image, the number of training images extracted at each position in the area image is measured, and positions with a large number of extracted training images or positions with a small number of extracted training images are determined. The system generates a visualized image and outputs display information by superimposing the visualized image onto the area image. [Effects of the Invention]
[0010] According to the present invention, users can easily check the status of training image collection, that is, whether or not training images with the required attributes are collected without bias. In particular, since locations with a large number of extracted training images or locations with a small number of extracted training images are presented in a visualized image, users can easily identify the locations on the image area where there are problems with the training image collection status. As a result, users can easily visually check the annotation status of training images prior to training, and efficiently create highly accurate learning models. [Brief explanation of the drawing]
[0011] [Figure 1] Overall configuration diagram of the image recognition model construction system according to this embodiment. [Figure 2] Block diagram showing the schematic configuration of the image processing device. [Figure 3] This diagram illustrates the detection area and pre-detection area set on the area image in the case of a people counting system. [Figure 4] An explanatory diagram showing an overview of the training image generation process. [Figure 5] An explanatory diagram showing information about area images registered in the database. [Figure 6] Diagram illustrating the rectangle representing a person [Figure 7] Explanatory diagram showing training images [Figure 8]Explanatory diagram showing information about teacher images registered in a database [Figure 9] Explanatory diagram showing deficiencies in the annotation status regarding extraction positions of teacher images on an area image [Figure 10] Explanatory diagram showing extracted number counting processing [Figure 11] Explanatory diagram showing extracted number counting processing when overlapping of persons occurs [Figure 12] Explanatory diagram showing a screen in an annotation work mode [Figure 13] Explanatory diagram showing a screen when a list is selected in an annotation status confirmation mode [Figure 14] Explanatory diagram showing a screen when a graph is selected in an annotation status confirmation mode [Figure 15] Explanatory diagram showing a screen when a graph is selected in an annotation status confirmation mode [Figure 16] Explanatory diagram showing a screen in an annotation status detail confirmation mode [Figure 17] Explanatory diagram showing a screen in an annotation status detail confirmation mode [Figure 18] Explanatory diagram showing a screen in an annotation status detail confirmation mode [Figure 19] Explanatory diagram showing a screen in an annotation status detail confirmation mode [Figure 20] Explanatory diagram showing another example of a screen in an annotation status detail confirmation mode [Mode for Carrying Out the Invention]
[0012] A first invention made to solve the above problem is an image processing apparatus in which a processor executes processing for visualizing a collection status of teacher images for constructing an image recognition model corresponding to a monitoring area, wherein the processor generates a teacher image including a detection target and a background from an area image related to the monitoring area, and processes characteristics of the detection target included in the teacher image A region attribute value representing an attribute is assigned to each of the aforementioned training images. targeting the teacher images having the attribute specified by a user, A mapping process is performed to assign the region attribute values assigned to each of the aforementioned training images to each position in the area image, the number of training images extracted at each position in the area image is measured, and positions with a large number of extracted training images or positions with a small number of extracted training images are determined.The system generates a visualized image and outputs display information by superimposing the visualized image onto the area image.
[0013] According to this, users can easily check the status of the collection of training images with the attributes they specify, that is, whether or not training images with the necessary attributes are available without bias. In particular, Since the visualization image shows locations with a large number of extracted training images or locations with a small number of extracted training images, This allows users to easily identify areas on the image where there are problems with the collection of training images. This enables users to easily visually check the annotation status of training images prior to training, allowing for the efficient creation of highly accurate learning models.
[0014] Furthermore, the second invention is that the processor, as the visualization image, provides the training image at each position of the area image. Extraction status The system is configured to generate a heatmap image representing [the specified area].
[0015] According to this, users can easily understand the status of training image collection at each location in the area image.
[0016] Furthermore, the third invention is that the processor uses the training image as the visualization image. Extraction status The system is configured to generate a mark image representing the area on the aforementioned area image where there is a problem.
[0017] According to this, users can easily identify areas on the area image where there are problems with the collection of training images.
[0018] Furthermore, the fourth invention is configured such that the processor generates the training image from a real area image captured by a camera or a virtual area image created with computer graphics, as the area image.
[0019] According to this method, training images can be generated efficiently.
[0020] Furthermore, the fifth invention is that the processor processes the training images for each color category relating to the person as an attribute. Extraction status The system is configured to generate the visualization image that visualizes the above.
[0021] According to this, the accuracy of image recognition models can vary greatly depending on the color type of the person. Therefore, by showing users the status of training image collection for each color type of person, it is possible to easily create highly accurate image recognition models (machine learning models).
[0022] Furthermore, the sixth invention is that the processor processes the training images for each attribute. Extraction status The system is configured to generate a statistical graph that visualizes the data and to output the aforementioned display information, including this statistical graph.
[0023] This allows users to easily understand the status of training image collection for each attribute. In this case, a 3D statistical graph visualizing the training image collection status for each combination of multiple attributes may be generated.
[0024] Furthermore, the seventh invention is configured such that the processor outputs the display information including a first screen which generates the training image and sets attributes to the training image in response to user operation, displays the visualization image superimposed on the area image, and outputs the display information including a second screen which is provided with an operation unit for returning to the first screen.
[0025] According to this, if there are problems with the collection of training images, the process can quickly proceed to adding missing training images on the first screen, which generates training images and sets attributes for them.
[0026] Furthermore, the eighth invention is an image processing method in which a processor performs a process to visualize the status of collecting training images for constructing an image recognition model corresponding to a monitoring area, wherein a training image including the object to be detected and the background is generated from an area image relating to the monitoring area, and the characteristics of the object to be detected included in the training image A region attribute value representing an attribute is assigned to each of the aforementioned training images. Targeting the training image having the attributes specified by the user, A mapping process is performed to assign the region attribute values assigned to each of the aforementioned training images to each position in the area image, the number of training images extracted at each position in the area image is measured, and positions with a large number of extracted training images or positions with a small number of extracted training images are determined. The system generates a visualized image and outputs display information by superimposing the visualized image onto the area image.
[0027] According to this, similar to the first invention, users can easily visually check the annotation status of training images prior to training, enabling the efficient creation of highly accurate learning models.
[0028] Hereinafter, embodiments of the present invention will be described with reference to the drawings.
[0029] Figure 1 is an overall diagram of the image recognition model construction system according to this embodiment.
[0030] This system comprises an image processing device 1 (information processing device), a camera 2, and a recorder 3.
[0031] Camera 2 captures images of the surveillance area. Recorder 3 stores the images captured by Camera 2. The images stored in Recorder 3 are input to Image Processing Device 1.
[0032] The image processing device 1 is comprised of a PC or similar device. The image processing device 1 is connected to a display 4 and an input device 5 such as a keyboard or mouse. Alternatively, the display 4 and input device 5 may be integrated into a single touch panel display.
[0033] The image processing device 1 constructs an image recognition model (machine learning model) that detects predetermined events from images captured by the camera 2 using machine learning such as deep learning. The image processing device 1 also generates training images used for training the image recognition model. Furthermore, the image processing device 1 evaluates the performance of the machine learning model using evaluation images that are different from the training images.
[0034] Furthermore, prior to training, the image processing device 1 visualizes and presents to the user the status of teacher image collection (annotation status), that is, whether a sufficient number of teacher images with the necessary attributes are available and in an appropriate distribution.
[0035] In this embodiment, the image processing device 1 generates training images and performs a learning process to construct an image recognition model (machine learning model) using those training images. However, the learning process may be performed by a device other than the image processing device 1.
[0036] Next, the schematic configuration of the image processing device 1 will be described. Figure 2 is a block diagram showing the schematic configuration of the image processing device 1.
[0037] The image processing device 1 comprises a communication unit 11, a storage unit 12, and a processor 13.
[0038] The communication unit 11 communicates with the recorder 3.
[0039] The memory unit 12 stores programs executed by the processor 13, etc. It also stores registration information for a database that manages training images generated by the processor 13 and their attributes.
[0040] The processor 13 performs various processes by executing programs stored in the memory unit 12. In this embodiment, the processor 13 performs processes such as training image generation, extraction count measurement, visualization, output, and learning.
[0041] In the training image generation process (annotation process), processor 13 generates training images (training images) that are used for learning to build an image recognition model. In addition, in the training image generation process, processor 13 sets attributes related to the features of the object to be detected and its background contained in the training image for each training image, in response to user input.
[0042] In the extraction count measurement process, the processor 13 measures the number of extracted training images at each position on the area image for each attribute of the training image, such as the attributes of the person included in the training image (for example, the color of the person's clothing).
[0043] In the visualization process, processor 13 visualizes the status of teacher image collection (annotation status). The visualization process targets teacher images with attributes specified by the user and generates visualization images that visualize the status of teacher image collection at each location in the area image. Specifically, it generates heatmap images that represent the status of teacher image collection at each location in the area image, and frame images (mark images) that represent areas in the area image where there are problems with the status of teacher image collection.
[0044] During output processing, the processor 13 outputs the annotation work mode screen (see Figure 12), the annotation status confirmation mode screen (see Figures 13 to 15), and the annotation status confirmation mode screen (see Figures 16 to 19) to the display 4.
[0045] During the learning process, processor 13 constructs an image recognition model (machine learning model) using machine learning to detect predetermined events from images captured by camera 2. The training process uses the training images generated in the training image generation process.
[0046] Next, we will describe an event detection system that utilizes an image recognition model. Figure 3 is an explanatory diagram showing the detection area and pre-detection area set on an area image in the case of a people counting system.
[0047] In this embodiment, an image recognition model (machine learning model) is constructed for use in a people counting system that measures the number of people passing through a monitoring area. Specifically, in order to measure the number of people entering a store (number of customers), the image recognition model is used to detect the arrival of people as a target event from an area image taken by camera 2 at the entrance of the store, which serves as the monitoring area.
[0048] In this case, the object to be detected is a person (customer) entering the store. Furthermore, a detection area and a pre-detection area are pre-defined in the area image. The detection area is set at the store entrance in the area image. The pre-detection area is the area a person passes through before entering the detection area, and is set at the location of the passage adjacent to the detection area.
[0049] The image recognition model determines that a person (detected object) has entered the store when it moves from the pre-detection area to the detection area, and adds one person to the measurement result (number of customers). At this time, it is determined that the person has passed through the pre-detection area and entered the detection area based on the position of the person's representative point. The representative point is either the center point of the person's rectangle or the center point of their feet.
[0050] The detection area and pre-detection area are defined as polygons on the area image. The coordinates of the polygon's vertices are registered as information regarding the positions of the detection area and pre-detection area defined on the area image.
[0051] This embodiment describes an image recognition model used in a people counting system that measures the number of people passing through a monitoring area, but it may also be an image recognition model used in an event detection system that detects various events.
[0052] For example, it could be an image recognition model used in an intrusion detection system that detects when a person (object to be detected) enters a restricted area. In this case, in an area image taken by camera 2 of a monitoring area including the restricted area, a detection area is set at the location of the restricted area, and it is determined that the person has entered the restricted area when they move from the pre-detection area to the detection area.
[0053] Alternatively, it may be an image recognition model used in an abandoned item detection system that detects when a person has left an item behind. In this case, for example, in an area image taken by camera 2 of a surveillance area that includes an area where abandonment is prohibited, such as an emergency entrance (fire department entrance), a detection area is set at the location of the area where abandonment is prohibited. When a person carrying an item moves from the pre-detection area to the detection area, and then leaves the detection area while leaving the item behind, it is determined that the person has abandoned the item.
[0054] Alternatively, it may be an image recognition model used in a dwell time detection system that detects when a person stays in a specific location for an extended period of time. In this case, for example, if it concerns customers waiting in line at a retail store, a detection area is set at the location of the waiting area in an area image captured by camera 2 of a monitoring area including the waiting area. After a person moves from the pre-detection area to the detection area, if the person stays in the detection area for a predetermined time or longer, it is determined that the person has stayed in the waiting area for an extended period of time.
[0055] Next, we will explain the overview of the training image generation process performed by the image processing device 1. Figure 4 is an explanatory diagram showing the overview of the training image generation process.
[0056] Image processing device 1 performs the process of generating training images (training images) used in machine learning to build an image recognition model (machine learning model) (training image generation process).
[0057] Here, as shown in Figure 4(A), if a person (object to be detected) is included in the real-world area image (actual area image) captured by camera 2 of the target monitoring area, a training image is generated by cutting out the region containing the person from the real-world area image.
[0058] Furthermore, as shown in Figure 4(B), a composite area image including a person is created by superimposing a person image (real person image or CG person image) onto an area image (real-life area image or CG area image), and a training image is generated by cutting out the region containing the person from this composite area image. Here, the real-life area image (real area image) is an image of the surveillance area captured by camera 2. The CG area image (virtual area image) is an image that simulates the real-life area image using CG (Computer Graphics). The real-life person image (real person image) is an image generated by cutting out the region of a person from an image captured by camera 2, etc. The CG person image (virtual person image) is an image created with CG.
[0059] Furthermore, the training image is created from the area image frame by frame. Therefore, as shown in Figure 4(A), when the training image is generated from a live-action area image that includes a person, the person moves on the live-action area image frame by frame as the person walks within the monitoring area. In this case, the position from which the training image is extracted should be gradually changed in accordance with the person's movement. Alternatively, as shown in Figure 4(B), when the training image is generated from a composite area image in which a person image is superimposed on the area image, the training image may be extracted from the area image frame by frame while moving the person image on the area image to reproduce the state of a person walking within the monitoring area.
[0060] Furthermore, when a training image is generated from a composite area image in which a person image (real person image, CG person image) is superimposed on an area image (real-life area image, CG area image), the person image is superimposed at positions in the area image where a person is likely to appear.
[0061] Furthermore, even if the target monitoring area is the same, if images are taken from different directions by separate cameras 2, the orientation of the person (detection target) will differ, so separate training images are prepared and separate image recognition models (machine learning models) are constructed.
[0062] Next, we will explain information regarding area images. Figure 5 is an explanatory diagram showing information about area images registered in the database.
[0063] Area images (live-action background images, CG background images) are images that represent the surveillance area. Area images include structures such as pillars, walls, and shutters that exist within the surveillance area. Area images also include people present within the surveillance area.
[0064] Image processing device 1 registers and manages information about area images in a database. The information about area images registered in the database includes structural information and person information. Structural information is information about structures included in the area image. Person information is information about people included as background in the area image.
[0065] Structural information includes information about the type of structure and information about the attributes of the structure. Information about the type of structure indicates whether it is a fixed structure or a movable structure. For example, columns and walls are fixed structures, while shutters installed at the entrance of a store are movable structures. Information about the attributes of a structure includes, for example, the opening and closing times of a shutter as a movable structure.
[0066] Person information includes details about the color scheme of the person's clothing (color of the upper body clothing, color of the lower body clothing) and information about whether or not the person is carrying any belongings. Belongings include not only luggage but also objects that the person moves, such as strollers or carts.
[0067] Next, we will explain the person rectangle set in the image processing device 1. Figure 6 is an explanatory diagram showing the person rectangle. Note that when a training image is extracted from a composite area image in which person images (real person images, CG person images) are superimposed, the person region corresponds to the person image.
[0068] As shown in Figure 6(A), in this embodiment, a rectangle surrounding the person to be detected is set, and the height H and width W of the rectangle are registered in the database as information about the rectangle. The example shown in Figure 6(B) is when the person to be detected is moving a stroller. In this case as well, the rectangle is set to surround only the person.
[0069] Furthermore, as shown in Figure 6(C), in the case of CG human images, the coordinates of the points that make up the contour (contour points) are registered in the database as information about the contour of the person.
[0070] Furthermore, as shown in Figures 6(D) and (E), information regarding the person's position is registered in the database, including the coordinates of the reference point (the top-left point of the person rectangle), the center point (the center point of the person rectangle), and the center point of the feet (the intersection of the perpendicular line passing through the center point of the person rectangle and the base of the person rectangle).
[0071] Next, we will describe the training images generated by the image processing device 1. Figure 7 is an explanatory diagram showing the training images.
[0072] The training image is generated by extracting the region containing the person from an area image that includes the person. The training image includes a region representing the person to be detected and a background region representing the person's background. The background region includes building structures such as pillars and floors.
[0073] As shown in Figures 7(A-1) and (A-2), the training image is extracted from the area image based on a rectangle surrounding the person to be detected. That is, the training image is created by extracting a rectangular area from the area image that is enlarged by a predetermined width around the person rectangle. Specifically, the training image has a size where the area of the person rectangle is enlarged horizontally by a predetermined horizontal enlargement width α, and the area of the person rectangle is enlarged vertically by a predetermined vertical enlargement width β. The training image consists of the person area included in the person rectangle, the background area included in the person rectangle, and the background area around the person rectangle. The enlargement widths α and β may be set appropriately, for example, within the range of 0 to 10 pixels.
[0074] The examples shown in Figures 7(B-1) and (B-2) illustrate the case where a person, as the object to be detected, is moving a stroller. In this case as well, the training image is created by cutting out a rectangular area from the area image that is enlarged by a predetermined width around a rectangle that surrounds only the person.
[0075] Furthermore, in actual operation, situations may occur in the monitoring area where multiple people overlap in the front and back as seen from camera 2 (overlapping people). To ensure the performance of the image recognition model (machine learning model) even in such situations, machine learning is performed to build the image recognition model (machine learning model) using training images where overlapping people occur.
[0076] Furthermore, when overlapping of people occurs, there are two possibilities: as shown in Figures 7(C-1) and (C-2), another person appears behind the person to be detected; and as shown in Figures 7(D-1) and (D-2), another person appears in front of the person to be detected. In this case, if the training image is extracted with another person appearing behind the person to be detected, the training image will include the person's region as the background. If the training image is extracted with another person appearing in front of the person to be detected, the training image will include the person's region as the foreground.
[0077] Furthermore, training images in which people overlap are generated by extracting training images from live-action area images in which people overlap. Training images in which people overlap can also be generated by creating a composite area image in which a person image (live-action or CG) is superimposed on an area image containing people so that it overlaps the people in that area image. Additionally, training images in which people overlap can also be generated by creating a composite area image in which multiple person images (live-action or CG) are superimposed on an area image so that they overlap.
[0078] Furthermore, when a training image is generated from an area image, information regarding the training image's position on the area image (coordinates) and information regarding its size (height, width) are registered in the database.
[0079] Next, we will explain the information regarding the training images managed by the image processing device 1. Figure 8 is an explanatory diagram showing the information regarding the training images registered in the database.
[0080] The image processing device 1 registers and manages information about training images in a database. The information about training images registered in the database includes the image number (image identification information), attribute information, and image information.
[0081] Attribute information includes information about the clothing of the person in the training image (color of the upper body clothing, color of the lower body clothing), information about whether the person is carrying any belongings, and information about concealment. In addition, the person's gender, height, and body type may also be included in the attribute information. Belongings include not only luggage but also objects that the person moves, such as strollers and carts.
[0082] Information regarding concealment includes information about overlapping of people and information about concealment by non-people objects. Information about overlapping people indicates whether other people are present in the background of the person being detected, or in the foreground of the person being detected. Information about concealment by non-people objects indicates whether the person being detected is partially concealed by an object other than a person.
[0083] The image information includes information about the rectangle surrounding the person in the training image (height H, width W), information about the person's outline (coordinates of the points that make up the outline) in the case of a CG person image, and information about the position of the person's region.
[0084] Information regarding the location of the person's area includes the coordinates of the reference point (the top-left point of the person's rectangle), the center point (the center point of the person's rectangle), and the center point at the feet (the intersection of the perpendicular line passing through the center point of the person's rectangle and the base of the person's rectangle). The center point at the feet is used to determine whether or not a person has entered the detection area and the pre-detection area. The coordinates of each point are on the area image.
[0085] Next, we will explain the procedure for improving the deficiencies in the annotation. Figure 9 is an explanatory diagram illustrating the deficiencies in the annotation regarding the extraction position of the training image on the area image.
[0086] In this embodiment, a training image is generated by extracting the region containing the person from an area image containing the person. The training image includes both the person region and the background region, and the accuracy of person recognition in the image recognition model (machine learning model) changes depending on the differences in features between the person region and the background region within the training image. In other words, even for people with similar features, the accuracy of person recognition changes if the position of the training image on the area image from which it was extracted is different. Furthermore, even if the position of the training image on the area image from which it was extracted is the same, the accuracy of person recognition changes if the person's features, such as the color of the person's clothing, are different. For this reason, if there is a bias in the position from which the training image is extracted from the area image for each person's features (e.g., the color of the person's clothing), it becomes impossible to construct an image recognition model (machine learning model) with stable accuracy.
[0087] Therefore, in this embodiment, the image processing device 1 measures the number of extracted training images at each position on the area image for each attribute of a person included in the training image (extraction count measurement process). Next, the number of extracted training images at each position on the area image is compared for each attribute of a person included in the training image, and information representing the annotation status regarding whether training images are extracted evenly in areas of the area image where a person may appear is presented to the user. Specifically, if a position is detected where the number of extracted training images is significantly lower than at other positions in the area image, that position with the significantly lower number of extracted training images is presented to the user.
[0088] In response, users perform annotation work to add training images to locations where the number of extracted training images is significantly low. This replenishes the missing training images, improving the deficiency of annotation due to positional bias in training images and avoiding the problem where the recognition accuracy of the image recognition model (machine learning model) changes drastically depending on the location where the person to be detected appears within the monitoring area.
[0089] In the example shown in Figure 9, the number of training images extracted at each location on the area image is measured, focusing on the color of the person's clothing as an attribute of the person included in the training image. In this example, four color systems are focused on: yellow, light blue, green, and red. Furthermore, Figure 9 shows the analysis results, indicating on the area image regions where there are few training images for the specified attribute (person's clothing color is yellow), regions where there are few training images regardless of the attribute, and regions where there are sufficient training images for all attributes.
[0090] Furthermore, the number of training images extracted may be measured by focusing on other color systems such as black and white, and the color system may be changed as needed. In addition, the number of training images extracted may be measured by dividing the colors of people's clothing into warm colors, cool colors, and black and white.
[0091] Furthermore, in this example, we focus on the color of the person's entire body clothing. In this case, the upper and lower body of the person's clothing are the same color. However, there are cases where the upper and lower body of the person's clothing are different colors. In this case, the number of training images extracted may be measured by focusing on the combination of upper and lower body colors.
[0092] Next, we will explain the extraction count measurement process performed by the image processing device 1. Figure 10 is an explanatory diagram showing the extraction count measurement process.
[0093] In the extraction count measurement process, first, region attribute values are assigned to each position (pixel) on the training image. Then, the region attribute values assigned to each position (pixel) on the training image are assigned to the corresponding positions (pixels) on the area image (mapping process). Specifically, for example, a region attribute value of "1" is assigned to the person region and a region attribute value of "0" is assigned to the background region. The mapping process is performed on all of the target training images.
[0094] Next, a counting process is performed to count the number of times each region attribute value (e.g., 1, 0) is assigned to each position (pixel) on the area image. The resulting count value represents the number of training images extracted for each attribute at each position (pixel) on the area image.
[0095] Here, by focusing on training images with a specific attribute, we can measure the number of training images extracted at each position (pixel) on the area image. Specifically, we target training images where the person's clothing is the color of interest (e.g., yellow), and count the number of times the area attribute value "1" is assigned to each position (pixel) on the area image. This count value represents the number of training images extracted where the person's clothing color is the color of interest (e.g., yellow). Furthermore, by performing this process similarly for each color of the person's clothing, we can obtain the number of training images extracted for each color of the person's clothing. This makes it possible to visualize the positions on the area image where the number of training images extracted for each color of the person's clothing is low.
[0096] Furthermore, the number of times a region attribute value of "0" is assigned to each position (pixel) on the area image is counted. This count value represents the number of training images extracted for each position (pixel) on the area image. This makes it possible to visualize positions on the area image where the number of training images extracted is small.
[0097] Next, we will explain the process for measuring the number of extracted images when there is overlapping of people. Figure 11 is an explanatory diagram showing the process for measuring the number of extracted images when there is overlapping of people.
[0098] When overlapping people occur, the mapping process, which assigns region attribute values to each position (pixel) on the training image, assigns different region attribute values to the background person region and the foreground person region. Specifically, for example, as shown in Figure 11(A), if there is a background person, the background person region is assigned a region attribute value of "2". As shown in Figure 11(B), if there is a foreground person, the foreground person region is assigned a region attribute value of "3". Note that the assignment of a region attribute value of "1" to the person region and a region attribute value of "0" to the background region that does not contain a person is the same as in the example shown in Figure 10.
[0099] Furthermore, in the counting process that counts the number of times a region attribute value has been assigned, the number of times each region attribute value (for example, 1, 0, 2, 3) has been assigned is counted at each position (pixel) on the area image.
[0100] Here, as training images for the attribute of interest, we measure the number of training images extracted at each position in the area image, specifically for training images that include overlapping people. That is, we count the number of times the area attribute value "2" or "3" is assigned to each position (pixel) on the area image. This count value represents the number of training images extracted that include overlapping people. This makes it possible to visualize the positions on the area image where the number of training images extracted that include overlapping people is low.
[0101] When overlapping of people occurs in this way, especially when another person appears in front of the person to be detected, the person to be detected will be partially obscured by the other person. On the other hand, there may be objects other than people in front of the person to be detected. In this case, the person to be detected will be obscured by the object other than people. To ensure the performance of the image recognition model (machine learning model) even in such situations, it is desirable to perform machine learning to build the image recognition model (machine learning model) using training images in which obscuration occurs, that is, training images in which the person to be detected is obscured by an object in the foreground.
[0102] In this case, by using live-action area images where occlusion is occurring, composite area images in which object images (live-action object images, CG object images) are superimposed on area images containing people, or composite area images in which object images are superimposed on person images (live-action person images, CG person images), it is possible to obtain training images in a state where occlusion is occurring.
[0103] Next, we will explain the screens displayed on Display 4. Figure 12 is an explanatory diagram showing the screen in annotation work mode. Figure 13 is an explanatory diagram showing the screen when a list is selected in annotation status confirmation mode. Figures 14 and 15 are explanatory diagrams showing the screen when a graph is selected in annotation status confirmation mode. Figures 16, 17, 18, and 19 are explanatory diagrams showing the screen in annotation status detail confirmation mode.
[0104] As shown in Figures 12 to 19, screens 21, 61, 71, and 81 displayed on the display 4 are provided with tabs 22 (operation sections) for annotation work, annotation status confirmation, and detailed annotation status confirmation. When the user operates the annotation work tab 22, the annotation work mode screen shown in Figure 12 is displayed. When the user operates the annotation status confirmation tab 22, the screen transitions to the annotation status confirmation mode screen shown in Figures 13 to 15. When the user operates the detailed annotation status confirmation tab 22, the screen transitions to the detailed annotation status confirmation mode screen shown in Figures 16 to 19.
[0105] The annotation work mode (training image creation mode) screen 21 (first screen) shown in Figure 12 is provided with an image input unit 31. The image input unit 31 is provided with separate tabs 32 for CG and live-action images. When the user operates the CG tab 32, it enters CG image input mode. When the user operates the live-action tab 32, it enters live-action image input mode. The image input unit 31 is also provided with a person image input unit 33 and an area image input unit 34.
[0106] In the person image input section 33, a list of person images is displayed when the user operates the input button 35, and a person image can be input by selecting one from this list. In this case, a CG person image is input in CG image input mode, and a real person image is input in live-action image input mode.
[0107] In the area image input section 34, a list of area images is displayed when the user operates the input button 36, and an area image can be input by selecting an area image from this list. In this case, a CG area image is input in CG image input mode, and a live-action area image is input in live-action image input mode.
[0108] Furthermore, the annotation work mode screen (21) is equipped with an area image display unit (41). The area image display unit (41) displays the input area image (CG area image, live-action area image). In addition, the area image display unit (41) overlays the input person image (CG person image, live-action person image) onto the area image. The area image display unit (41) also allows the user to adjust the position and size of the person image overlaid on the area image by dragging with the mouse.
[0109] Furthermore, the area image display unit 41 allows the user to specify a target person on the area image, that is, to specify the position on the area image from which the training image will be extracted. Specifically, the user can input a person frame 42 surrounding a person image (CG person image or live-action person image) on the area image (CG area image or live-action area image) by dragging with the mouse or other means. This person frame 42 becomes a candidate for the range of the training image. In other words, a person rectangle is set based on the position of the person frame 42, and the training image is generated based on that person rectangle. At this time, the training image is extracted from the composite area image, which is a composite of the area image and the person image.
[0110] Furthermore, when creating a training image using a person included in a real-life area image displayed on the area image display unit 41, it is sufficient to input a person frame 42 surrounding the person included in the real-life area image. In this case, no user input operation is required in the person image input unit 33.
[0111] Furthermore, by performing the person detection process, the user's input operation for the person frame 42 may be omitted. That is, the person detection process may be performed on a live-action area image containing a person, or on a composite area image obtained by combining an area image and a person image, and a training image may be extracted based on the person detection frame obtained by the person detection process.
[0112] Furthermore, the annotation work mode screen 21 is provided with a frame operation unit 45. The frame operation unit 45 is provided with a button 46 to move the area image back one frame and a button 47 to move the area image back one frame. By operating this frame operation unit 45, a training image can be created for each frame of the area image.
[0113] Furthermore, the annotation work mode screen 21 is provided with a detection area input unit 51. When the user operates the input button 52 in the detection area input unit 51, the system transitions to the detection area input mode, and the user can input the range of the detection area and the pre-detection area on the area image displayed on the area image display unit 41 by dragging with the mouse or other means.
[0114] Furthermore, the annotation work mode screen 21 is provided with an attribute input section 53. In the attribute input section 53, the user can input the colors of the upper body clothing and the lower body clothing as attributes related to the person included in the training image. If a CG person image is selected, the clothing colors are already known, so no user input is required.
[0115] Furthermore, the annotation work mode screen 21 is provided with a title input section 55. In the title input section 55, the user can input the title of the training image, specifically the name assigned to a group of training images. The title of the training image (group name) identifies the set of training images when checking the annotation status of the training images.
[0116] Furthermore, the annotation work mode screen 21 is provided with a save button 56. When the user operates the save button 56, a training image is generated based on the input content of each section, and this training image is saved in the storage unit 12. At the same time, attribute information related to the training image is registered in the database.
[0117] The annotation status confirmation mode screen 61 shown in Figures 13, 14, and 15 is equipped with a training image selection unit 62. The training image selection unit 62 allows the user to select a group of training images to be analyzed. For example, when the user operates the training image selection unit 62, a list of training image groups is displayed, and the user can select a group of training images from here. In this example, a group of training images with entrance / exit A as the monitoring area has been selected.
[0118] Furthermore, the annotation status confirmation mode screen 61 is equipped with a list button 63 and a graph button 64. When the user operates the list button 63, analysis processing is performed on the group of training images selected by the training image selection unit 62, and the screen 61 shown in Figure 13 for list selection is displayed as the analysis result. Similarly, when the user operates the graph button 64, analysis processing is performed on the group of training images selected by the training image selection unit 62, and the screen 71 shown in Figure 14 for graph selection is displayed as the analysis result.
[0119] In the annotation status confirmation mode shown in Figure 13, the screen 61 when a list is selected displays a list table 66 in the visualization result display unit 65. The list table 66 displays a list of attributes for each training image registered in the database. In this example, the attributes of the training image related to entrance / exit A are displayed. In addition, the attributes of the training image include the characteristics of the person included in the training image, particularly the color of the person's clothing, and the time of day of the area image that serves as the background for the training image.
[0120] Table 66 visualizes the annotation status of training images using text information, allowing users to visually check for deficiencies in the annotation status, such as a lack of training images for specific attributes. Specifically, a user can see that in area images for a specific time period (for example, around 8 AM), there are fewer training images containing people wearing a specific color (for example, yellow) than training images for other attributes.
[0121] In the example shown in Figure 13, Table 66 displays the color of the person's clothing and the time of day of the area image as attributes related to the training image, but other attributes may also be displayed. For example, whether or not there is overlap between people, i.e., whether or not the training image includes background or foreground areas of people, may be displayed. Also, the image quality (resolution, presence or absence of blur, etc.) of the person and background areas of the training image may be displayed. Furthermore, the season (spring, summer, autumn, winter) of the area image may be displayed.
[0122] In the annotation status confirmation mode shown in Figures 14 and 15, the screen 71 when a graph is selected displays a statistical graph 72 in the visualization result display unit 65. The statistical graph 72 is a three-dimensional bar graph. In the statistical graph 72, the first horizontal axis represents the time period of the area image that serves as the background of the training image as the first attribute of the training image, the second depth axis represents the characteristics of the person included in the training image, particularly the color of the person's clothing, as the second attribute of the training image, and the third vertical axis represents the number of training images. That is, a bar graph is drawn for each combination of the first attribute (time period of the area image) and the second attribute (color of the person's clothing) of the training image, and the height of the bar graph represents the number of training images that possess both the first and second attributes.
[0123] Statistical graph 72 visualizes the annotation status of training images using a bar graph, allowing users to visually check for deficiencies in the annotation status, such as a lack of training images for specific attributes. Specifically, a user can see that in area images for a specific time period (for example, around 8 AM), there are fewer training images containing people wearing a specific color (for example, yellow) than training images for other attributes.
[0124] In the screen 71 shown in Figure 14, there are fewer training images where the person's clothing is yellow compared to training images where the person's clothing is light blue, green, and red, indicating that additional training images with this attribute are needed. On the other hand, in the screen 71 shown in Figure 15, the number of training images is uniform for all colors of clothing (yellow, light blue, green, and red), indicating an improvement in the annotation status of the training images.
[0125] Furthermore, the statistical graph 72 may be configured with an operation panel on the screen to select a first attribute of the training image represented by the first horizontal axis and a second attribute of the training image represented by the second depth axis, allowing the user to specify the attributes of the training image represented by each axis in the statistical graph 72 on the screen.
[0126] The annotation status details confirmation mode screen 81 (second screen) shown in Figures 16 to 19 is provided with a training image selection unit 82. The training image selection unit 82 allows the user to select a group of training images to be analyzed. For example, when the user operates the training image selection unit 82, a list of training image groups is displayed, and the user can select a group of training images from here. In this example, a group of training images where the monitoring area is entrance / exit A and the time period is 8:00 is selected.
[0127] Furthermore, the annotation status details confirmation mode screen 81 is provided with an analysis button 83 and a visualization results display unit 84. When the user operates the analysis button 83, the analysis process is executed on the group of training images selected by the training image selection unit 82, and as an analysis result, a heatmap image 85 is displayed on the visualization results display unit 84. The heatmap image 85 is displayed in a transparent state, superimposed on the area image.
[0128] Heatmap image 85 visualizes the annotation status of training images at each location on the area image. Specifically, heatmap image 85 is divided into multiple cells in a mesh-like structure, and in each cell, the number of training images in which a representative point is located is represented by a change in grayscale.
[0129] The number of training images in which a representative point is located within a cell may be represented by a change in hue. In this case, as the number of training images increases, the color of the cell may change in the order of, for example, blue, yellow, orange, and red. Alternatively, the number of training images in which a representative point is located within a cell may be represented by a change in pattern (pattern image).
[0130] Furthermore, the annotation status details confirmation mode screen 81 is provided with an attribute selection section 87. In the attribute selection section 87, the user can select attributes of the training image, particularly the clothing color of people included in the training image. Specifically, the attribute selection section 87 is provided with buttons 88 for "All," "Yellow," "Light Blue," "Green," and "Red." The user can select the clothing color of a person by operating the buttons 88. In addition, the attribute selection section 87 allows the user to select multiple clothing colors for a person.
[0131] When the attribute selection unit 87 selects the color of a person's clothing as an attribute of the training image, particularly an attribute of a person included in the training image, the heatmap image 85 displays in each cell the number of training images that have a representative point located within the cell and that correspond to the specified attribute, represented by a change in grayscale.
[0132] In this example, the heatmap image 85 displayed visualizes the annotation status by focusing on attributes related to the characteristics of the people included in the training image, particularly the color of their clothing. However, a heatmap image 85 focusing on other attributes of the training image may also be used. For example, the presence or absence of overlapping people, i.e., whether or not background or foreground areas of people are included in the training image, may be represented in the heatmap image 85. Alternatively, the image quality of the people or background areas of the training image (e.g., resolution, presence or absence of blur) may also be represented in the heatmap image 85.
[0133] Here, as shown in Figure 16, the user first checks the annotation status of the training image without limiting the attributes of the training image. Specifically, the user operates the button 88 in the attribute selection unit 87 to select all colors (yellow, light blue, green, red) and checks the annotation status of the training image without limiting the colors of the clothing of the people included in the training image.
[0134] In this example, in heatmap image 85, the cells on the left within the area image have a small number of training images for the specified attribute, the cells in the center within the area image have a large number of training images for the specified attribute, and the cells on the right within the area image have no training images for the specified attribute at all.
[0135] Next, as shown in Figure 17, the user checks the annotation status limited to training images with specific attributes. In this example, the user operates button 88 to select yellow in the attribute selection unit 87 to check the annotation status for training images with the attribute that makes a person's clothing yellow.
[0136] In this example, similar to the state shown in Figure 16, in the heatmap image 85, the cells on the left within the area image have a small number of training images for the specified attribute, the cells in the center within the area image have a large number of training images for the specified attribute, and the cells on the right within the area image have no training images for the specified attribute at all.
[0137] Therefore, the user operates the annotation tab 22 (operation panel) to return to the annotation work mode screen 21 shown in Figure 12, and performs the annotation work to add the missing training images. In this example, training images are added in which the representative points are located in the left and right regions of the area image and in which the person's clothing is yellow. That is, training images are added that include a person whose clothing is yellow, with the left and right regions of the area image as the background. This improves the shortage of training images with the attribute that the representative points are located in the left and right regions of the area image and the person's clothing is yellow.
[0138] Next, as shown in Figure 18, the user selects the group of teacher images to be analyzed (Entrance / Exit A_8:00_Teacher Image_Additional 1) in the teacher image selection unit 82 and checks the annotation status. The user also operates the button 88 in the attribute selection unit 87 to select a color other than yellow (light blue, green, red) and checks the annotation status of the teacher images for the attributes corresponding to the colors other than yellow (light blue, green, red).
[0139] In this example, in heatmap image 85, the left and center cells within the area image contain a large number of training images with the specified attribute, i.e., training images of people whose clothing is either light blue, green, or red. However, the right cell within the area image contains no training images with the specified attribute at all.
[0140] Therefore, the user operates the annotation tab 22 to return to the annotation work mode screen 21 shown in Figure 12 and performs the annotation work to add the missing training images. In this example, a training image is added in which the representative point is located in the right-hand region of the area image and the color of the person's clothing is one of light blue, green, or red. That is, a training image is added in which the right-hand region of the area image is used as the background and the person's clothing color is one of light blue, green, or red. This improves the shortage of training images in which the representative point is located in the right-hand region of the area image and the person's clothing color is one of light blue, green, or red.
[0141] Next, as shown in Figure 19, the user selects the group of teacher images that have been added again (Entrance / Exit A_8:00_Teacher Image_Additional 2) in the teacher image selection unit 82 as the analysis target and checks the annotation status. The user also operates the button 88 in the attribute selection unit 87 to select all colors (yellow, light blue, green, red) and checks the annotation status of the teacher images for the attributes corresponding to all colors (yellow, light blue, green, red).
[0142] In this example, in heatmap image 85, the number of training images is uniform in all cells within the range where a person could exist, allowing the user to confirm that the annotation situation has improved.
[0143] Furthermore, in the annotation status details confirmation mode screen 81, when the user specifies the measurement target area 91 on the heatmap image 85 by dragging with the mouse, a list table 92 (frequency distribution table) is displayed. The list table 92 shows the distribution of the color of a person's clothing as an attribute of the training images, targeting training images where representative points are located in the specified measurement target area 91. Specifically, the number of training images for each color of a person's clothing (yellow, light blue, green, red) is displayed. The user can check the details of the annotation status by looking at the number of training images for each color of a person's clothing (yellow, light blue, green, red). In this example, the number of training images for all colors is equal, and the user can confirm that the training images for the necessary attributes are available without bias.
[0144] Furthermore, the annotation status details confirmation screen 81 is equipped with a training button 88. Once the user confirms that the required attribute training images are available without bias, they operate the training button 88. This triggers a training process based on the generated training images, creating an image recognition model (machine learning model).
[0145] Next, we will describe another example of the annotation status details confirmation mode screen. Figure 20 is an explanatory diagram showing another example of the annotation status details confirmation mode screen.
[0146] In the screen 81 shown in Figures 16 to 19, a heatmap image 85 (visualization image) in which the number of extracted training images is represented by changes in grayscale is superimposed on the area image in the visualization result display unit 84. On the other hand, in the screen 101 shown in Figure 20, a frame image 102 (mark image) representing the area on the area image where there are problems with the training image collection status is superimposed on the area image in the visualization result display unit 84.
[0147] In this example, frame image 102 is superimposed on the area image to surround the area where the number of extracted training images matching the specified attribute (for example, a person's clothing is yellow) is small. Alternatively, frame image 102 can be superimposed on the area image to surround the area where the number of extracted training images is small, without limiting the attributes.
[0148] Furthermore, the area on the area image where there are no problems with the collection of training images may be displayed in the frame image 102. For example, the frame image 102 may be drawn to surround the area on the area image where a large number of training images with a specific attribute have been extracted. Also, the frame image 102 may be drawn in different colors depending on the number of training images extracted.
[0149] In this example, the color of a person's clothing was specified as an attribute of the training image, and the region where the number of training images in which the person's clothing is of a specific color is small was displayed in frame image 102. However, attributes other than the color of a person's clothing may also be specified. For example, overlapping of people may be specified as an attribute of the training image, and the region where the number of training images in which overlapping people occurs is small may be displayed in frame image 102.
[0150] In this example, a frame image 102 (mark image) representing the area on the area image where there are problems with the training image collection is superimposed on the area image, but the mark image is not limited to frame image 102. For example, a semi-transparent image with a pattern drawn on it may be superimposed as the mark image on the area on the area image where there are problems with the training image collection. In addition to frame image 102 (mark image), comments (not shown) regarding the addition of training images may also be displayed.
[0151] As described above, embodiments have been explained as examples of the technology disclosed in this application. However, the technology in this disclosure is not limited to these embodiments and can be applied to embodiments that have been modified, replaced, added, or omitted. Furthermore, it is possible to create new embodiments by combining the components described in the above embodiments. [Industrial applicability]
[0152] The image processing apparatus and image processing method according to the present invention have the effect of allowing users to easily visually confirm the status of training image collection prior to training, and to efficiently create highly accurate training models. They are useful as an image processing apparatus and image processing method for visualizing the status of training image collection for constructing an image recognition model corresponding to a monitoring area. [Explanation of symbols]
[0153] 1 Image Processing Device 13 processors 21 Annotation work mode screen 22 tabs 61 Annotation status confirmation mode screen 66 List 71 Annotation status confirmation mode screen 72 Statistical Graphs 81 Annotation Status Details Confirmation Mode Screen 85 Heatmap Images 101 Annotation Status Details Confirmation Mode Screen 102 Frame image (mark image)
Claims
1. An image processing device that uses a processor to perform a process to visualize the status of collecting training images for building an image recognition model corresponding to a monitoring area, The aforementioned processor, From the area image relating to the aforementioned monitoring area, a training image including the object to be detected and the background is generated. A region attribute value representing an attribute related to the characteristics of the object to be detected included in the training image is assigned to each of the training images. Targeting the training images having the attributes specified by the user, a mapping process is performed to assign the region attribute values assigned to each training image to each position in the area image, and the number of training images extracted at each position in the area image is measured. A visualization image is generated that visualizes the locations where the number of extracted training images is large or where the number of extracted training images is small. An image processing apparatus characterized by outputting display information obtained by superimposing the aforementioned visualization image onto the aforementioned area image.
2. The aforementioned processor, The image processing apparatus according to claim 1, characterized in that it generates a heatmap image as the visualization image, which represents the extraction status of the training image at each position of the area image.
3. The aforementioned processor, The image processing apparatus according to claim 1, characterized in that it generates a mark image representing the area on the area image where there is a problem in the extraction status of the training image, as the visualization image.
4. The aforementioned processor, The image processing apparatus according to claim 1, characterized in that it generates the training image from a real area image captured by a camera or a virtual area image created with computer graphics as the area image.
5. The aforementioned processor, The image processing apparatus according to claim 1, characterized in that it generates a visualization image that visualizes the extraction status of the training images for each color type related to the person as an attribute.
6. The aforementioned processor, The image processing apparatus according to claim 1, characterized in that it generates a statistical graph that visualizes the extraction status of the training images for each attribute, and outputs the display information including this statistical graph.
7. The aforementioned processor, In response to user operations, the system outputs the display information, including a first screen that generates the training image and sets attributes for the training image. The image processing apparatus according to claim 1, characterized in that it displays the visualization image superimposed on the area image and outputs the display information including a second screen provided with an operation unit for returning to the first screen.
8. An image processing method that uses a processor to perform a process to visualize the status of collecting training images for building an image recognition model corresponding to a monitoring area, From the area image relating to the aforementioned monitoring area, a training image including the object to be detected and the background is generated. A region attribute value representing an attribute related to the characteristics of the object to be detected included in the training image is assigned to each of the training images. Targeting the training images having the attributes specified by the user, a mapping process is performed to assign the region attribute values assigned to each training image to each position in the area image, and the number of training images extracted at each position in the area image is measured. A visualization image is generated that visualizes the locations where the number of extracted training images is large or where the number of extracted training images is small. An image processing method characterized by outputting display information obtained by superimposing the visualization image onto the area image.
Citation Information
Patent Citations
Machine learning system, training dataset generation system, and machine learning program
JP2021111101A
Object detection device, object detection system, and object detection method
JP2021131734A
Supervised domain adaptation
US11170581B1
Human body attribute recognition method and apparatus, electronic device, and storage medium
US20220036059A1
Information processing device, information processng method, and program
WO2020071233A1