Image processing system
The image processing system addresses the limitation of not being able to use task-specific feature amounts across tasks by using a learned model that extracts and combines these features, enhancing the system's ability to perform multiple image processing tasks effectively.
Patent Information
- Application Number
- JP2023555935
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-10-26
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2041-10-26
AI Technical Summary
Existing image processing systems using multi-task learning are limited because they cannot mutually utilize task-specific feature amounts across different tasks, restricting the ability to perform learning and estimation considering feature amounts specific to each task and others.
An image processing system is configured with a learned model that includes components to extract common and task-specific feature amounts from images, combine these feature amounts, and output inference results for each task, allowing for mutual utilization of task-specific feature amounts.
This configuration enables effective learning and estimation across multiple tasks by allowing the system to consider and combine task-specific feature amounts, improving the performance and efficiency of image processing tasks.
Smart Images

Figure 0007683723000001 
Figure 0007683723000002 
Figure 0007683723000003
Abstract
Description
Technical Field
[0001] The present invention relates to an image processing system, an image processing method, and a recording medium.
Background Art
[0002] There is a method of simultaneously learning and estimating a plurality of tasks using a single multi-layer neural network DNN (Deep Neural Network). This method is called multi-task learning. Multi-task learning can reduce the learning and estimation time that increases in proportion to the number of tasks. Thus, multi-task learning has become one of the effective methods in applications such as human image analysis where information obtained from a plurality of tasks is required.
[0003] An example of multi-task learning is described in Patent Document 1. In the technique described in Patent Document 1 (hereinafter referred to as the technique related to the present invention), the DNN extracts a feature amount x L common to a plurality of tasks from an image in which a human face is depicted. Next, the DNN extracts a feature amount specific to the task of identifying the facial expression from the feature amount x L and outputs an estimation result y c . In parallel therewith, the DNN extracts a feature amount specific to the task of estimating the positions of eyes and nose in the face region from the feature amount x L and outputs an estimation result y r .
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] However, in the technology related to the present invention, it is configured to extract common feature amounts for all tasks from an image, extract task-specific feature amounts from the common feature amounts, and estimate the estimation results of each task. Therefore, there is a problem that the feature amounts specific to a certain task cannot be used for the estimation of other tasks.
[0006] An object of the present invention is to provide an image processing system that solves the above-described problem, that is, the problem that task-specific feature amounts cannot be mutually used among a plurality of tasks.
Means for Solving the Problem
[0007] An image processing system according to one embodiment of the present invention includes a learning unit that generates a learned model for performing a plurality of different inference tasks from an image, wherein the learned model a first component that extracts a first feature amount common to the plurality of inference tasks from the image, a second component that is provided corresponding to the inference task and extracts a second feature amount specific to the corresponding inference task from the first feature amount, a third component that combines the second feature amounts extracted for each inference task to generate a third feature amount, a fourth component that is provided corresponding to the inference task and outputs an inference result of the corresponding inference task from the third feature amount, and is configured to include.
[0008] An image processing system according to another embodiment of the present invention includes an inference unit that outputs inference results of a plurality of different inference tasks from an image using a learned model, wherein the learned model a first component that extracts a first feature amount common to the plurality of inference tasks from the image, a second component that is provided corresponding to the inference task and extracts a second feature amount specific to the corresponding inference task from the first feature amount, A third component that combines the second feature amounts extracted for each of the inference tasks to generate a third feature amount; A fourth component that is provided corresponding to the inference task and outputs an inference result of the corresponding inference task from the third feature amount; and is configured to include the above.
[0009] An image processing method according to another aspect of the present invention generates a learned model that performs a plurality of different inference tasks on an image, and in the generation, the learned model extracts a first feature amount common to the plurality of inference tasks from the image, for each of the inference tasks, extracts a second feature amount specific to the corresponding inference task from the first feature amount, combines the second feature amounts extracted for each of the inference tasks to generate a third feature amount, and for each of the inference tasks, is configured to output an inference result of the corresponding inference task from the third feature amount.
[0010] An image processing method according to another aspect of the present invention uses a learned model to estimate and output inference results of a plurality of different inference tasks from an image, and in the estimation, the learned model extracts a first feature amount common to the plurality of inference tasks from the image, for each of the inference tasks, extracts a second feature amount specific to the corresponding inference task from the first feature amount, combines the second feature amounts extracted for each of the inference tasks to generate a third feature amount, and for each of the inference tasks, is configured to output an inference result of the corresponding inference task from the third feature amount.
[0011] A computer-readable recording medium according to another aspect of the present invention A program for causing a computer to perform a process of generating a learned model that performs a plurality of different inference tasks from an image, In the generation, the learned model is caused to extract a first feature amount common to the plurality of inference tasks from the image, extract, for each inference task, a second feature amount specific to the corresponding inference task from the first feature amount, generate a third feature amount by combining the second feature amounts extracted for each inference task, output, for each inference task, an inference result of the corresponding inference task from the third feature amount, and is configured to record the program.
[0012] A computer-readable recording medium according to another aspect of the present invention is a program for causing a computer to perform a process of estimating and outputting inference results of a plurality of different inference tasks from an image using a learned model, In the estimation, the learned model is caused to extract a first feature amount common to the plurality of inference tasks from the image, extract, for each inference task, a second feature amount specific to the corresponding inference task from the first feature amount, generate a third feature amount by combining the second feature amounts extracted for each inference task, output, for each inference task, an inference result of the corresponding inference task from the third feature amount, and is configured to record the program.
Advantages of the Invention
[0013] By having the configuration as described above, the present invention can mutually utilize feature amounts specific to tasks among a plurality of tasks. For this reason, in each of the plurality of tasks, learning and estimation considering the feature amount specific to the task and the feature amounts specific to other tasks become possible.
Brief Description of the Drawings
[0014]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Embodiments for Carrying Out the Invention
[0015] Next, embodiments of the present invention will be described in detail with reference to the drawings. [First Embodiment] FIG. 1 is a block diagram of an image processing apparatus 10 according to a first embodiment of the present invention. This image processing apparatus 10 is configured to perform a plurality of inference tasks that are different from each other on an image. Referring to FIG. 1, the image processing apparatus 10 includes a camera I / F (interface) unit 11, a communication I / F unit 12, an operation input unit 13, a screen display unit 14, a storage unit 15, and an arithmetic processing unit 16.
[0016] The camera I / F unit 11 is connected to an image server 17 by wire or wirelessly, and is configured to transmit and receive data between the image server 17 and the arithmetic processing unit 16. The image server 17 is connected to a camera 18 by wire or wirelessly, and is configured to store a plurality of images taken by the camera 18 at different shooting times for a certain period in the past. The camera 18 may be, for example, a color camera or a monochrome camera equipped with a CCD (Charge-Coupled Device) image sensor or a CMOS (Complementary MOS) image sensor having a pixel capacity of about several million pixels. The camera 18 may be a camera installed on the street, indoors, etc. where many people pass by for the purpose of crime prevention and surveillance. Alternatively, the camera 18 may be a camera mounted on a moving body such as a vehicle and configured to capture the same or different shooting areas while moving. The camera 18 is not limited to one, and may be a plurality of cameras that capture different shooting areas from different locations.
[0017] The communication I / F unit 12 is composed of a data communication circuit and is configured to perform data communication with an external device (not shown) by wire or wirelessly. The operation input unit 13 is composed of an operation input device such as a keyboard or a mouse, and is configured to detect an operator's operation and output it to the arithmetic processing unit 16. The screen display unit 14 is composed of a screen display device such as an LCD (Liquid Crystal Display), and is configured to display various information on the screen according to an instruction from the arithmetic processing unit 16.
[0018] The storage unit 15 is composed of a storage device such as a hard disk or a memory, and is configured to store processing information and programs 151 necessary for various processes in the arithmetic processing unit 16. The program 151 is a program that realizes various processing units when read and executed by the arithmetic processing unit 16, and is pre-read from an external device or a recording medium (not shown) via a data input / output function such as the communication I / F unit 12 and stored in the storage unit 15. The main processing information stored in the storage unit 15 includes image information 152, a model 153, and estimation result information 154.
[0019] The image information 152 is a frame image of the camera 18 acquired from the image server 17 through the camera I / F unit 11.
[0020] The model 153 is a machine learning model that simultaneously learns and estimates a plurality of different inference tasks from the frame image of the camera 18. The model 153 may be configured using, for example, a DCNN (Deep Convolutional Neural Network). In the present embodiment, the parameters of the model 153 are learned so as to perform three inference tasks: object detection, pose estimation, and semantic segmentation estimation. The model with learned parameters is called a learned model and is distinguished from the model before learning.
[0021] Object detection detects the class and object position in the image. The result of object detection includes the class name, the estimated confidence of the class, and a bounding box (hereinafter referred to as a rectangle) representing the object position. The class to be detected may be, for example, a person. However, the class to be detected is not limited to a person and may be an animal or an object.
[0022] Pose estimation estimates the skeletal information of a person in the image. The skeletal information of a person includes information representing the positions of joints that make up the human body. The joints may include not only joints such as the neck and shoulders but also parts of the face such as the eyes and nose. The result of pose estimation includes the joint name (joint ID), the position of the joint, and the confidence of the joint.
[0023] Semantic segmentation estimation estimates the class of each pixel in an image. The result of semantic segmentation estimation includes the class of each pixel. The classes to be estimated are the same as the classes detected by object detection.
[0024] The estimation result information 154 represents information on the result estimated from an image using the learned model 153. The estimation result information 154 includes an object detection result, a pose estimation result, and a semantic segmentation estimation result.
[0025] The arithmetic processing unit 16 has one or more processors such as an MPU and its peripheral circuits, and is configured to cooperate the above hardware and the program 151 to realize various processing units by reading and executing the program 151 from the storage unit 15. The main processing units realized by the arithmetic processing unit 16 include an acquisition unit 161, a learning unit 162, and an estimation unit 163.
[0026] The acquisition unit 161 is configured to acquire, from the image server 17 through the camera I / F unit 11, a frame image constituting a video captured by the camera 18 or a frame image obtained by downsampling the same, and store it in the storage unit 15 as image information 152. The acquired frame image is added with a camera ID and a shooting time. The shooting time of the frame image is different for each frame.
[0027] The learning unit 162 is configured to simultaneously train the model 153 for the above three inference tasks using training data. That is, the learning unit 162 generates a learned model 153 that performs the above three inference tasks from an image. In the above generation, the learning unit 21 causes the model 153 to extract a first feature amount common to the above three inference tasks from an image, and then, for each inference task, extracts a second feature amount specific to the corresponding inference task from the first feature amount, and then combines the second feature amounts extracted for each inference task to generate a third feature amount, and then, for each inference task, outputs an inference result of the corresponding inference task from the third feature amount.
[0028] The estimation unit 163 is configured to estimate and output the inference results of the above three inference tasks from the image using the learned model 153. In the above estimation, the estimation unit 31 first causes the learned model 153 to extract a first feature amount common to the above three inference tasks from the image, and then, for each inference task, extracts a second feature amount specific to the corresponding inference task from the first feature amount. Next, the second feature amounts extracted for each inference task are combined to generate a third feature amount. Then, for each inference task, the inference result of the corresponding inference task is output from the third feature amount.
[0029] Next, the operation of the image processing apparatus 10 will be described. The phases of the image processing apparatus 10 are roughly classified into a learning phase and an estimation phase. The learning phase is a phase of machine learning of the model 153. The estimation phase is a phase of estimating and outputting the inference results of the above three inference tasks from the image using the learned model 153.
[0030] FIG. 2 is a flowchart showing an example of the operation in the learning phase. Referring to FIG. 2, first, the acquisition unit 161 acquires the frame image captured by the camera 18 from the image server 17 through the camera I / F unit 11 and stores it in the storage unit 15 as the image information 152 (step S1). Next, the learning unit 162 creates training data to be used for the machine learning of the model 153 (step S2). Next, the learning unit 162 uses the training data to perform machine learning on the model 153 with the input being the image and the output being the estimation results of the above three inference tasks, and generates the learned model 153 (step S3).
[0031] FIG. 3 is a flowchart showing an example of the operation in the estimation phase. Referring to FIG. 3, first, the acquisition unit 161 acquires the frame image captured by the camera 18 from the image server 17 through the camera I / F unit 110 and stores it in the storage unit 15 as the image information 152 (step S11).
[0032] Next, the estimation unit 163 simultaneously estimates the estimation results of the above three inference tasks from the frame images included in the image information 152 using the learned model 153 (step S12). Next, the estimation unit 163 displays the estimated results of the three inference tasks on the screen display unit 14 and / or transmits them to an external device through the communication I / F unit 12 (step S13).
[0033] Subsequently, the model 153 and the learning unit 162 will be described in detail.
[0034] First, the details of the model 153 will be described.
[0035] FIG. 4 is a configuration diagram showing an example of a multi-task model that can be used as the model 153. The model 153 in this example is composed of eight components CM and is a single multi-layer neural network as a whole.
[0036] The component CM1 is provided on the lower layer side of the multi-layer neural network, configured to input an image and extract a low-order feature amount FM1 common to all tasks. The component CM1 is also called a backbone. The feature amount FM1 extracted by the component CM1 is also referred to as a low-order feature map. The component CM1 may be configured to include one or more convolutional layers. For example, the component CM1 may use VGG-16, which is a component of SSD (Single Shot MultiBox Detector). Alternatively, the component CM1 may use, for example, VGG-19, which is a component of OpenPose. Alternatively, the component CM1 may use, for example, the encoder, which is a component of SegNet. Alternatively, the component CM1 may use the backbone of a model other than SSD, OpenPose, or SegNet, for example.
[0037] Component CM2-1 is configured to input the feature quantity FM1 from component CM1 and extract a higher-order feature quantity FM2-1 specific to the object detection task. Component CM2-1 may be configured to include one or more convolutional layers. For example, component CM2-1 may use a special convolutional layer (Extra Feature Layers) that is a component of SSD. However, component CM2-1 is not limited to the above, and a convolutional layer that extracts a higher-order feature quantity specific to the object detection task in an object detection model other than SSD may be used.
[0038] Component CM2-2 is configured to input the feature quantity FM1 from component CM1 and extract a higher-order feature quantity FM2-2 specific to the pose estimation task. Component CM2-2 may be configured to include one or more convolutional layers. For example, component CM2-2 may use a convolutional layer that generates a Part Confidence Map representing the position of keypoints, which is a component of OpenPose, a convolutional layer that generates Part Affinity Fields representing the relevance between keypoints, and a layer that concatenates the generated Part Confidence Map, Part Affinity Fields, and the source feature quantity FM1 (the feature map obtained by concatenation will be hereinafter referred to as the OpenPose feature map). However, component CM2-2 is not limited to the above, and a convolutional layer that extracts a higher-order feature quantity specific to the pose estimation task in a pose estimation model other than OpenPose may be used.
[0039] Component CM2-3 is configured to input feature amount FM1 from component CM1 and extract higher-order feature amount FM2-3 specific to the semantic segmentation estimation task. Component CM2-3 may be configured to include one or more convolutional layers. For example, component CM2-3 may use a decoder which is a component of SegNet. However, component CM2-3 is not limited to the above, and may use a convolutional layer that extracts higher-order feature amounts specific to the semantic segmentation estimation task in a semantic segmentation estimation model other than SegNet.
[0040] Component CM3 is configured to input feature amounts FM2-1, FM2-2, FM2-3 from components CM2-1, CM2-2, CM-2-3, and generate feature amounts FM3-1, FM3-2, FM3-3 obtained by concatenating these three feature amounts FM2-1, FM2-2, FM2-3.
[0041] Figure 5 is a configuration diagram showing an example of component CM3. The component CM3 in this example is configured to include a resizing unit CM3-1, a combining unit CM3-2, and a resizing unit CM3-3.
[0042] The resizing unit CM3-1 is configured to adjust the sizes of the feature quantities FM2-1, FM2-2, and FM2-3 so that they can be combined. The resizing unit CM3-1 designates any one of the three feature quantities as a reference feature quantity, and changes the sizes of the remaining two feature quantities according to the size of the reference feature quantity. For example, the sizes of the feature quantities FM2-1, FM2-2, and FM2-3 are 38×38, 70×70, and 240×320 respectively, and the reference feature quantity is FM2-1. In this case, the resizing unit CM3-1 generates and outputs a feature quantity FM2-2' in which the size of the feature quantity FM2-2 is changed from 70×70 to 38×38. Also, the resizing unit CM3-1 generates and outputs a feature quantity FM2-3' in which the size of the feature quantity FM2-3 is changed from 240×320 to 38×38. Also, the resizing unit CM3-1 does not change the size of the feature quantity FM2-1 and outputs the feature quantity FM2-1 itself as the feature quantity FM2-1'.
[0043] The combining unit CM3-2 inputs the feature quantities FM2-1', FM2-2', and FM2-3' from the resizing unit CM3-1, and generates and outputs a feature quantity FM3 obtained by combining them. For example, the combining unit CM3-2 inputs the feature quantities FM2-1', FM2-2', and FM2-3' each having a size of 38×38, and generates and outputs a feature quantity FM3 having a size of 38×38×3. In this way, the number of channels (dimensionality) increases by combining the feature quantities.
[0044] The resizing unit CM3-3 inputs the feature quantity FM3 from the combining unit CM3-2, and generates and outputs the feature quantities FM3-1, FM3-2, and FM3-3 that have been changed to sizes corresponding to each task. For example, let the input sizes of the components CM4-1, CM4-2, and CM4-3 be 38×38×3, 70×70×3, and 240×320×3 respectively. In this case, the resizing unit CM3-3 generates the feature quantity FM3-2 in which the size of the feature quantity FM3 is changed from 38×38×3 to 70×70×3, and outputs it to the component CM4-2. Also, the resizing unit CM3-3 generates the feature quantity FM3-3 in which the size of the feature quantity FM3 is changed from 38×38×3 to 240×320×3, and outputs it to the component CM4-3. Further, the resizing unit CM3-3 outputs the feature quantity FM3 itself with a size of 38×38×3 as the feature quantity FM3-1 to the component CM4-1.
[0045] Figure 6 is a configuration diagram showing another example of the component CM3. The component CM3 in this example is configured to include three sub-components CM3A, CM3B, and CM3C.
[0046] The sub-component CM3A is configured to generate and output a feature quantity FM3-1 for the component CM4-1 of the object detection task from the feature quantities FM2-1, FM2-2, and FM2-3. The sub-component CM3A includes a resizing unit CM3A-1 that generates and outputs feature quantities FM2-2' and FM2-3' obtained by changing the sizes of the feature quantities FM2-2 and FM2-3 according to the size of the feature quantity FM2-1, and a combining unit CM3A-2 that generates and outputs a feature quantity FM3-1 obtained by combining the three feature quantities FM2-1, FM2-2', and FM2-3'. For example, let the sizes of the feature quantities FM2-1, FM2-2, and FM2-3 be 38×38, 70×70, and 240×320 respectively, and the input size of the component CM4-1 be 38×38×3. In this case, the resizing unit CM3A-1 generates and outputs a feature quantity FM2-2' with the size of the feature quantity FM2-2 changed from 70×70 to 38×38, and generates and outputs a feature quantity FM2-3' with the size of the feature quantity FM2-3 changed from 240×320 to 38×38. The combining unit CM3A-2 combines the feature quantities FM2-1, FM2-2', and FM2-3' with the same size of 38×38 to generate and output a feature quantity FM3-1 with a size of 38×38×3. Thereby, the degradation of the feature quantity FM2-1 due to the combination can be suppressed.
[0047] The sub-component CM3B is configured to generate and output a feature quantity FM3-2 for the component CM4-2 of the pose estimation task from the feature quantities FM2-1, FM2-2, and FM2-3. The sub-component CM3B includes a resizing unit CM3B-1 that generates and outputs feature quantities FM2-1' and FM2-3' obtained by changing the sizes of the feature quantities FM2-1 and FM2-3 according to the size of the feature quantity FM2-2, and a combining unit CM3B-2 that generates and outputs a feature quantity FM3-2 obtained by combining the three feature quantities FM2-1', FM2-2, and FM2-3'. For example, assume that the sizes of the feature quantities FM2-1, FM2-2, and FM2-3 are 38×38, 70×70, and 240×320 respectively, and the input size of the component CM4-2 is 70×70×3. In this case, the resizing unit CM3B-1 generates and outputs a feature quantity FM2-1' with the size of the feature quantity FM2-1 changed from 38×38 to 70×70, and generates and outputs a feature quantity FM2-3' with the size of the feature quantity FM2-3 changed from 240×320 to 70×70. The combining unit CM3B-2 combines the feature quantities FM2-1', FM2-2, and FM2-3' of the same size of 70×70 to generate and output a feature quantity FM3-2 with a size of 70×70×3. Thereby, the deterioration of the feature quantity FM2-2 caused by the combination can be suppressed.
[0048] The sub-component CM3C is configured to generate and output a feature quantity FM3-3 for the component CM4-3 of the semantic segmentation estimation task from the feature quantities FM2-1, FM2-2, and FM2-3. The sub-component CM3C includes a resizing unit CM3C-1 that generates and outputs feature quantities FM2-1' and FM2-2' obtained by changing the sizes of the feature quantities FM2-1 and FM2-2 according to the size of the feature quantity FM2-3, and a combining unit CM3C-2 that generates and outputs a feature quantity FM3-3 obtained by combining the three feature quantities FM2-1', FM2-2', and FM2-3. For example, let the sizes of the feature quantities FM2-1, FM2-2, and FM2-3 be 38×38, 70×70, and 240×320 respectively, and the input size of the component CM4-3 be 240×320×3. In this case, the resizing unit CM3C-1 generates and outputs a feature quantity FM2-1' obtained by changing the size of the feature quantity FM2-1 from 38×38 to 240×240, and generates and outputs a feature quantity FM2-2' obtained by changing the size of the feature quantity FM2-2 from 70×70 to 240×320. The combining unit CM3C-2 combines the feature quantities FM2-1', FM2-2', and FM2-3 of the same size of 240×320 to generate and output a feature quantity FM3-3 of the size 240×320×3. Thereby, deterioration of the feature quantity FM2-3 due to the combination can be suppressed.
[0049] Referring to FIG. 4 again, the component CM4-1 is configured to input the feature quantity FM3-1 from the component CM3 and estimate and output an estimation result ER1 of the object detection task from the feature quantity FM3-1. The feature quantity FM3-1 includes not only the high-order feature quantity FM2-1 specific to the object detection task, but also the high-order feature quantity FM2-2 specific to the pose estimation task and the high-order feature quantity FM2-3 specific to the semantic segmentation estimation. Therefore, the component CM4-1 can perform learning and estimation considering these three high-order feature quantities. The component CM4-1 may use, for example, an output layer (Detections: 8732 per Class, Non-Maximum Suppression) connected to a special convolutional layer constituting the SSD.
[0050] Here, the component CM4-1 may set the weight for determining the priority of the high-order feature amount FM2-1 specific to the object detection task to be greater than the weight for determining the priority of the second other feature amount. For example, the component CM4-1 may set the weight for determining the priority of the high-order feature amount FM2-1 specific to the object detection task to 0.5 and the weight for determining the priority of the second other feature amount to 0.25. By giving a relatively large weight to the feature amount FM2-1 in this way, it is possible to perform learning and estimation considering the three high-order feature amounts, and the importance of the high-order feature amount FM2-1 specific to the object detection task can be increased.
[0051] Also, the component CM4-1 may perform 1×1 convolution (Channel-Wise Convolution) on the input feature amount FM3-1 to reduce the number of dimensions of the high-order feature amount, for example, from 38×38×3 to 38×38×1. Thereby, as the component CM4-1, the network part that estimates and outputs the estimation result from the high-order feature amount in an existing model such as SSD can be used as it is.
[0052] The component CM4-2 is configured to input the feature amount FM3-2 from the component CM3 and estimate and output the estimation result ER2 of the pose estimation task from the feature amount FM3-2. The feature amount FM3-2 includes not only the high-order feature amount FM2-2 specific to the pose estimation task but also the high-order feature amount FM2-1 specific to the object detection task and the high-order feature amount FM2-3 specific to the semantic segmentation estimation. Therefore, the component CM4-2 enables learning and estimation considering those three high-order feature amounts. The component CM4-2 may use, for example, the network part that estimates the pose estimation result from the OpenPose feature map, which is a component of OpenPose.
[0053] Here, the component CM4-2 may set the weight determining the priority of the high-order feature quantity FM2-2 specific to the pose estimation task to be greater than the weight determining the priority of the second feature quantity other than that. For example, the component CM4-2 may set the weight determining the priority of the high-order feature quantity FM2-2 specific to the pose estimation task to 0.5 and the weight determining the priority of the second feature quantity other than that to 0.25. By giving a relatively large weight to the feature quantity FM2-2 in this way, it is possible to perform learning and estimation considering the three high-order feature quantities, and the importance of the high-order feature quantity FM2-2 specific to the pose estimation task can be increased.
[0054] Further, the component CM4-2 may reduce the dimensionality of the high-order feature quantity from, for example, 70×70×3 to 70×70×1 by performing 1×1 convolution (Channel-Wise Convolution) on the input feature quantity FM3-2. Thereby, as the component CM4-2, the network part that estimates and outputs the estimation result from the high-order feature quantity in an existing model such as OpenPose can be used as it is.
[0055] The component CM4-3 is configured to input the feature quantity FM3-3 from the component CM3 and estimate and output the estimation result ER3 of the semantic segmentation estimation task from the feature quantity FM3-3. The feature quantity FM3-3 includes not only the high-order feature quantity FM2-3 specific to the semantic segmentation estimation task, but also the high-order feature quantity FM2-1 specific to the object detection task and the high-order feature quantity FM2-2 specific to the pose estimation. Therefore, the component CM4-3 enables learning and estimation considering those three high-order feature quantities. The component CM4-3 may use, for example, a softmax layer which is a component of SegNet.
[0056] Here, the component CM4-3 may set the weight for determining the priority of the high-order feature quantity FM2-3 specific to the semantic segmentation estimation task to be greater than the weight for determining the priority of the second feature quantity other than that. For example, the component CM4-3 may set the weight for determining the priority of the high-order feature quantity FM2-3 specific to the semantic segmentation estimation task to be 0.5, and the weight for determining the priority of the second feature quantity other than that to be 0.25. By giving a relatively large weight to the feature quantity FM2-3 in this way, it is possible to perform learning and estimation considering three high-order feature quantities, and at the same time, the importance of the high-order feature quantity FM2-3 specific to the semantic segmentation estimation task can be increased.
[0057] Also, the component CM4-3 may perform 1×1 convolution (Channel-Wise Convolution) on the input feature quantity FM3-3 to reduce the dimensionality of the high-order feature quantity, for example, from 240×320×3 to 240×320×1. Thereby, as the component CM4-3, the network part that estimates and outputs the estimation result from the high-order feature quantity in an existing model such as SegNet can be used as it is.
[0058] Next, the details of the learning unit 162 will be described.
[0059] First, the training data used for the machine learning of the model 153 will be described.
[0060] FIG. 7 shows an example of a list of training data used for the machine learning of the model 153. Referring to FIG. 7, a total of n pieces of training data are registered in this list. Each piece of training data is composed of items such as an ID for uniquely identifying the training data, an image, an object detection label, a pose estimation label, and a semantic segmentation estimation label.
[0061] In the image item, the frame image captured by the camera 18 is set. In the object detection label item, the presence or absence of a label is set, and when there is a label, the class such as a person existing in the image, which is label information, and its position information (rectangle information) are set. In the pose estimation label item, the presence or absence of a label is set, and when there is a label, the joint name (joint ID) of the joints existing in the image and its position information are set. In the semantic segmentation estimation label item, the presence or absence of a label is set, and when there is a label, the class of each pixel of the image is set. Thus, among the training data group, in addition to those in which label information is set for all items of the three labels (object detection label, pose estimation label, semantic segmentation estimation label), those in which label information is set only for some label items may be included.
[0062] The training data as described above may be created, for example, by interactive processing with a user. For example, the learning unit 162 displays the image of the camera 18 acquired by the acquisition unit 161 on the screen display unit 14, and receives label information of the image from the user through the operation input unit 13. Then, the learning unit 162 creates a pair of the displayed image and the received label information as one piece of training data. The learning unit 162 creates a necessary and sufficient number of training data by the same method. However, the method of creating training data is not limited to the above.
[0063] Next, a method for the learning unit 162 to learn the model 153 using the training data will be described.
[0064] FIG. 8 is a flowchart showing an example of the learning process of the learning unit 162. The learning process in this example uses the model 153 having the configuration shown in FIG. 4 as the learning target model. Also, the learning process in this example does not learn the entire model 153 at once, but learns while gradually expanding the network part to be learned. Thereby, stable learning can be performed. Specifically, it goes through the following four learning stages.
[0065] (1) Learning stage 1 In the first learning stage, the learning unit 162 learns only the components CM2-1 and CM4-1, which are the deep-layer network parts related to object detection. At this time, the parameters of the component CM-1, which is the backbone, the components CM2-2 and CM4-2, which are the deep-layer network parts related to pose estimation, and the components CM2-3 and CM4-3, which are the deep-layer network parts related to semantic segmentation estimation, are fixed. (2) Second learning stage In the second learning stage, the learning unit 162 learns only the components CM2-1, CM2-2, CM4-1, and CM4-2, which are the deep-layer network parts related to object detection and pose estimation. At this time, the parameters of the component CM-1, which is the backbone, and the components CM2-3 and CM4-3, which are the deep-layer network parts related to semantic segmentation estimation, are fixed. (3) Third learning stage In the third learning stage, the learning unit 162 learns only the components CM2-1, CM2-2, CM2-3, CM4-1, CM4-2, and CM4-3, which are the deep-layer network parts related to all inference tasks, that is, object detection, pose estimation, and semantic segmentation estimation. At this time, the parameters of the component CM-1, which is the backbone, are fixed. (4) Fourth learning stage In the fourth learning stage, the learning unit 162 learns the entire model, that is, the component CM-1, which is the backbone, and the components CM2-1, CM2-2, CM2-3, CM4-1, CM4-2, and CM4-3, which are the deep-layer network parts related to object detection, pose estimation, and semantic segmentation estimation.
[0066] Referring to FIG. 8, the learning unit 162 creates a training data group to be used in each learning stage from the training data group used for the machine learning of the model 153 (step S21).
[0067] For example, in step S21, the learning unit 162 creates, from the list of training data as described in FIG. 7, the necessary number of training data groups for use in learning stage 3 and the training data groups for use in learning stage 4, respectively. In learning stage 3 and learning stage 4, training data with label information set for all items of the three labels (object detection label, pose estimation label, semantic segmentation estimation label) is required. Therefore, the learning unit 162 creates the training data groups for use in learning stage 3 and the training data groups for use in learning stage 4 by extracting from the list the training data that satisfies such conditions.
[0068] Also, in step S21, the learning unit 162 creates the training data group for use in learning stage 2 from the remaining training data groups in the list. In learning stage 2, training data with label information set for the items of the object detection label and the pose estimation label (regardless of the presence or absence of semantic segmentation estimation label information) is required. Therefore, the learning unit 162 creates the training data group for use in learning stage 2 by extracting from the list the training data that satisfies such conditions.
[0069] Also, in step S21, the learning unit 162 creates the training data group for use in learning stage 1 from the remaining training data groups in the list. In learning stage 1, training data with label information set for the item of the object detection label (regardless of the presence or absence of pose estimation label information and semantic segmentation estimation label information) is required. Therefore, the learning unit 162 creates the training data group for use in learning stage 1 by extracting from the list the training data that satisfies such conditions.
[0070] Next, the learning unit 162 performs learning for each stage in the order of learning stage 1, learning stage 2, learning stage 3, and learning stage 4 until a predetermined end condition is satisfied (steps S22 to S25). In the learning for each stage, an error between the inference result of the inference task obtained as the output of the model 153 when an image included in the training data is input to the model 153 and the label information included in the training data is calculated using a previously given loss function. The loss function exists for each of the object detection task, the pose estimation task, and the semantic segmentation estimation task. The loss function for the object detection task is denoted as L1, the loss function for the pose estimation task is denoted as L2, and the loss function for the semantic segmentation estimation task is denoted as L3, respectively.
[0071] In learning stage 1, the parameters of the components CM2-1 and CM4-1 of the model 153 are learned so as to minimize the loss calculated by the loss function L1. In learning stage 2, the parameters of the components CM2-1, CM2-2, CM4-1, and CM4-2 of the model 153 are learned so as to minimize the sum (for example, weighted sum) of the loss calculated by the loss function L1 and the loss calculated by the loss function L2. In learning stage 3, the parameters of the components CM2-1, CM2-2, CM2-3, CM4-1, CM4-2, and CM4-3 of the model 153 are learned so as to minimize the sum (for example, weighted sum) of the loss calculated by the loss function L1, the loss calculated by the loss function L2, and the loss calculated by the loss function L3. In learning stage 4, the parameters of the components CM1, CM2-1, CM2-2, CM2-3, CM4-1, CM4-2, and CM4-3 of the model 153 are learned so as to minimize the sum (for example, weighted sum) of the loss calculated by the loss function L1, the loss calculated by the loss function L2, and the loss calculated by the loss function L3. In each learning, for example, the gradient descent method and the error backpropagation method may be used.
[0072] An example of a method for training the model 153 using the training data has been described above. However, the learning method applicable to the present invention is not limited to the above example. For example, the following learning method may be used. That is, first, only the components CM2-1 and CM4-1 related to object detection are learned (the parameters of the other components CM1, CM2-2, CM2-3, CM4-2, and CM4-3 are fixed). Next, only the components CM2-2 and CM4-2 related to pose estimation are learned (the parameters of the other components CM1, CM2-1, CM2-3, CM4-1, and CM4-3 are fixed). Next, only the components CM2-3 and CM4-3 related to semantic segmentation estimation are learned (the parameters of the other components CM1, CM2-1, CM2-3, CM4-1, and CM4-3 are fixed). Next, only the components CM2-1 to CM2-3 and CM3-1 to CM3-3 related to all inference tasks are learned (the parameter of the component CM1 is fixed). Next, the components CM1, CM2-1 to CM2-3, and CM4-1 to CM4-3 of the entire model are learned.
[0073] As described above, according to the image processing apparatus 10 according to the present embodiment, high-order feature amounts specific to tasks can be mutually used among a plurality of tasks. Therefore, in each of the plurality of tasks, learning and estimation considering the high-order feature amounts specific to the task and the high-order feature amounts specific to other tasks become possible.
[0074] Subsequently, a modification example of the present embodiment will be described.
[0075] <Modification Example 1> In the above-described embodiment, the model 153 was configured to perform semantic segmentation estimation. However, the model 153 may be configured to perform instance semantic segmentation estimation instead of semantic segmentation estimation. In this case, for example, a component for performing object detection from the feature amount FM3-3 may be added between the component CM3 and the component CM4-3 of the multi-task model 153 shown in FIG. 4, and the component CM4-3 may be configured to estimate the class in pixel units for each rectangle of the individually detected object classes.
[0076] <Modification Example 2> In the above-described embodiment, the model 153 was configured to perform three inference tasks of object detection, pose estimation, and semantic segmentation estimation. However, the model 153 may be configured to perform only any two of the inference tasks of object detection, pose estimation, and semantic segmentation estimation. Alternatively, the inference tasks performed by the model 153 are not limited to object detection, pose estimation, and semantic segmentation estimation, and may be other tasks.
[0077] [Second Embodiment] FIG. 9 is a block diagram of an image processing system 20 according to a second embodiment of the present invention. Referring to FIG. 9, the image processing system 20 includes a learning unit 21 and a learned model 22.
[0078] The learning unit 21 is configured to generate a learned model 22 that performs a plurality of different inference tasks from an image. The learning unit 21 can be configured in the same manner as, for example, the learning unit 162 in FIG. 1, but is not limited thereto.
[0079] The learned model 22 is configured to include: a first component that extracts a first feature amount common to the plurality of inference tasks from the above-mentioned image; a second component that is provided corresponding to the inference task and extracts a second feature amount specific to the corresponding inference task from the first feature amount; a third component that combines the second feature amounts extracted for each inference task to generate a third feature amount; and a fourth component that is provided corresponding to the inference task and outputs an inference result of the corresponding inference task from the third feature amount.
[0080] The image processing system 20 configured as described above operates as follows. That is, the learning unit 21 generates a learned model 22 that performs a plurality of different inference tasks from an image. In the above generation, the learning unit 21 causes the learned model 22 to extract a first feature amount common to the plurality of inference tasks from the image, and then, for each inference task, causes the learned model 22 to extract a second feature amount specific to the corresponding inference task from the first feature amount. Next, for each inference task, the second feature amounts extracted for each inference task are combined to generate a third feature amount. Next, for each inference task, the learned model 22 outputs an inference result of the corresponding inference task from the third feature amount.
[0081] According to the image processing system 20 configured and operating as described above, it is possible to mutually utilize task-specific feature amounts among a plurality of inference tasks. The reason is that the image processing system 20 is configured to combine the second feature amounts extracted for each inference task to generate a third feature amount and output an inference result of the corresponding inference task from the third feature amount. Therefore, in each of the plurality of inference tasks, learning and estimation considering the feature amount specific to the task and the feature amounts specific to other tasks become possible.
[0082] [Third Embodiment] FIG. 10 is a block diagram of an image processing system 30 according to the third embodiment of the present invention. Referring to FIG. 10, the image processing system 30 includes an estimation unit 31 and a learned model 32.
[0083] The estimation unit 31 is configured to output inference results of a plurality of different inference tasks from an image using the learned model 32. The estimation unit 31 can be configured in the same way as, for example, the estimation unit 163 in FIG. 1, but is not limited thereto.
[0084] The learned model 32 includes a first component that extracts a first feature amount common to the plurality of inference tasks from the image, a second component that is provided corresponding to the inference task and extracts a second feature amount specific to the corresponding inference task from the first feature amount, a third component that combines the second feature amounts extracted for each inference task to generate a third feature amount, and a fourth component that is provided corresponding to the inference task and outputs an inference result of the corresponding inference task from the third feature amount.
[0085] The image processing system 30 configured as described above operates as follows. That is, the estimation unit 31 estimates and outputs inference results of a plurality of different inference tasks from an image using the learned model 32. In the above estimation, the learned model 32 is first made to extract a first feature amount common to a plurality of inference tasks from the image, then for each inference task, a second feature amount specific to the corresponding inference task is extracted from the first feature amount, then the second feature amounts extracted for each inference task are combined to generate a third feature amount, and then for each inference task, an inference result of the corresponding inference task is output from the third feature amount.
[0086] According to the image processing system 30 configured and operating as described above, it is possible to mutually utilize task-specific feature amounts among a plurality of inference tasks. The reason is that the image processing system 30 is configured to combine the second feature amounts extracted for each inference task to generate a third feature amount and output an inference result of the corresponding inference task from the third feature amount. Therefore, in each of the plurality of inference tasks, learning and estimation considering the feature amount specific to the task and the feature amounts specific to other tasks become possible.
[0087] The present invention has been described with reference to the above-described embodiments, but the present invention is not limited to the above-described embodiments. Various modifications that can be understood by those skilled in the art can be made to the configuration and details of the present invention within the scope of the present invention.
Industrial Applicability
[0088] The present invention can be used in all fields that perform a plurality of inference tasks such as object detection, pose estimation, and semantic segmentation estimation from images such as camera images.
[0089] Some or all of the above embodiments can be described as follows in the appended claims, but are not limited thereto. [Appended Claim 1] A learning unit that generates a learned model that performs a plurality of different inference tasks from an image, The learned model includes a first component that extracts a first feature amount common to the plurality of inference tasks from the image, a second component that is provided corresponding to the inference task and extracts a second feature amount specific to the corresponding inference task from the first feature amount, a third component that combines the second feature amounts extracted for each of the inference tasks to generate a third feature amount, a fourth component that is provided corresponding to the inference task and outputs an inference result of the corresponding inference task from the third feature amount, and an image processing system including the above. [Appended Claim 2] The third component uses one of the plurality of second feature amounts as a reference feature amount, changes the sizes of the second feature amounts other than the reference feature amount to match the size of the reference feature amount, combines the second feature amounts other than the reference feature amount after the size change and the reference feature amount to generate the third feature amount, and for each inference task, changes the size of the third feature amount to match the input size of the fourth component and outputs it to the fourth component. The image processing system according to Appended Claim 1. [Appended Claim 3] The third component includes sub-components corresponding to the inference tasks. The sub-components use the second feature amount of the corresponding inference task as a reference feature amount, change the size of the second feature amounts other than the corresponding inference task to match the size of the reference feature amount, combine the second feature amounts other than the corresponding inference task after the size change and the reference feature amount to generate the third feature amount, and output it to the fourth component. The image processing system according to Supplementary Note 1. [Supplementary Note 4] The learning unit performs learning of the learned model in multiple learning stages. The multiple learning stages include at least a first learning stage in which any one of the multiple inference tasks is set as a learning target task, the parameters of the first component, the second component, and the third component related to the inference tasks other than the learning target task are fixed, and the parameters of the second component and the third component related to the learning target task are learned; a second learning stage in which the parameters of the first component are fixed, and the parameters of the second component and the third component related to all the inference tasks are learned. The image processing system according to any one of Supplementary Notes 1 to 3. [Supplementary Note 5] The fourth component provided corresponding to the inference task makes the weight for determining the priority of the second feature amount of the corresponding inference task among the multiple second feature amounts constituting the third feature amount larger than the weight for determining the priority of the other second feature amounts. The image processing system according to any one of Supplementary Notes 1 to 4. [Supplementary Note 6] The fourth component provided corresponding to the inference task reduces the number of dimensions of the third feature amount by performing 1×1 convolution on the input third feature amount. The image processing system according to any one of Supplementary Notes 1 to 5. [Appendix 7] The plurality of inference tasks includes an object detection task, a pose estimation task, and a semantic segmentation estimation task. The image processing system according to any one of Appendices 1 to 6. [Appendix 8] The image processing system further includes an inference unit that outputs inference results of the plurality of inference tasks from an image using the learned model. The image processing system according to any one of Appendices 1 to 7. [Appendix 9] The image processing system includes an inference unit that outputs inference results of a plurality of different inference tasks from an image using a learned model. The learned model includes: a first component that extracts a first feature amount common to the plurality of inference tasks from the image; a second component that is provided corresponding to the inference task and extracts a second feature amount specific to the corresponding inference task from the first feature amount; a third component that combines the second feature amounts extracted for each inference task to generate a third feature amount; a fourth component that is provided corresponding to the inference task and outputs an inference result of the corresponding inference task from the third feature amount. The image processing system including the above. [Appendix 10] A learned model for performing a plurality of different inference tasks from an image is generated. In the generation, the learned model is caused to: extract a first feature amount common to the plurality of inference tasks from the image; extract a second feature amount specific to the corresponding inference task from the first feature amount for each inference task; combine the second feature amounts extracted for each inference task to generate a third feature amount; output an inference result of the corresponding inference task from the third feature amount for each inference task. The image processing method. [Appendix 11] Using a learned model, estimate and output inference results of a plurality of different inference tasks from an image, In the estimation, the learned model is caused to extract a first feature amount common to the plurality of inference tasks from the image, for each inference task, extract a second feature amount specific to the corresponding inference task from the first feature amount, combine the second feature amounts extracted for each inference task to generate a third feature amount, for each inference task, output an inference result of the corresponding inference task from the third feature amount, Image processing method. [Appendix 12] A program for causing a computer to perform a process of generating a learned model that performs a plurality of different inference tasks from an image, In the generation, the learned model is caused to extract a first feature amount common to the plurality of inference tasks from the image, for each inference task, extract a second feature amount specific to the corresponding inference task from the first feature amount, combine the second feature amounts extracted for each inference task to generate a third feature amount, for each inference task, output an inference result of the corresponding inference task from the third feature amount, A computer-readable recording medium having the program recorded thereon. [Appendix 13] A program for causing a computer to perform a process of estimating and outputting inference results of a plurality of different inference tasks from an image using a learned model, In the estimation, the learned model is caused to extract a first feature amount common to the plurality of inference tasks from the image, for each inference task, extract a second feature amount specific to the corresponding inference task from the first feature amount, combine the second feature amounts extracted for each inference task to generate a third feature amount, For each of the inference tasks, output the inference result of the corresponding inference task from the third feature amount. A computer-readable recording medium storing a program.
Explanation of Signs
[0090] 10 Image processing apparatus 11 Camera I / F section 12 Communication I / F section 13 Operation input section 14 Screen display section 15 Storage section 16 Arithmetic processing section 17 Image server 18 Camera 151 Program 152 Image information 153 Model 154 Estimation result information 161 Acquisition section 162 Learning section 163 Estimation section
Claims
1. A learning unit that generates a learned model for performing a plurality of inference tasks that are different from each other from an image, The learned model includes: A first component that extracts a first feature amount common to the plurality of inference tasks from the image; A second component that is provided corresponding to the inference task and extracts a second feature amount specific to the corresponding inference task from the first feature amount; A third component that combines the second feature amounts extracted for each inference task to generate a third feature amount; A fourth component that is provided corresponding to the inference task and outputs an inference result of the corresponding inference task from the third feature amount; An image processing system including the above.
2. The third component uses one of the second feature amounts extracted for each inference task as a reference feature amount, changes the sizes of the second feature amounts other than the reference feature amount to match the size of the reference feature amount, and combines the second feature amounts other than the reference feature amount and the reference feature amount after changing the sizes to match the size of the reference feature amount to generate the third feature amount. For each inference task, the size of the third feature amount is changed to match the input size of the fourth component and output to the fourth component. The image processing system according to Claim 1.
3. The third component includes a sub-component corresponding to the inference task. The sub-component uses the second feature amount of the corresponding inference task as a reference feature amount, changes the sizes of the second feature amounts other than the corresponding inference task to match the size of the reference feature amount, and combines the second feature amounts other than the corresponding inference task and the reference feature amount after changing the sizes to match the size of the reference feature amount to generate the third feature amount and output it to the fourth component. The image processing system according to Claim 1.
4. The learning unit performs learning of the learned model in a plurality of learning stages. The plurality of learning stages include at least: A first learning stage in which any one of the plurality of inference tasks is set as a learning target task, the parameters of the second component and the fourth component related to the inference tasks other than the learning target task and the first component are fixed, and the parameters of the second component and the fourth component related to the learning target task are learned. A second learning stage of fixing the parameters of the first component and learning the parameters of the second component and the fourth component related to each of the plurality of inference tasks. The image processing system according to any one of claims 1 to 3.
5. The fourth component provided corresponding to the inference task makes the weight that determines the priority of the second feature amount of the corresponding inference task among the plurality of second feature amounts constituting the third feature amount larger than the weight that determines the priority of the other second feature amounts. The image processing system according to any one of claims 1 to 4.
6. An inference unit that outputs inference results of a plurality of mutually different inference tasks from an image using a learned model. The learned model is A first component that extracts a first feature amount common to the plurality of inference tasks from the image. A second component provided corresponding to the inference task, which extracts a second feature amount unique to the corresponding inference task from the first feature amount. A third component that combines the second feature amounts extracted for each inference task to generate a third feature amount. A fourth component provided corresponding to the inference task, which outputs an inference result of the corresponding inference task from the third feature amount. An image processing system including
7. An image processing method by a computer, The computer generates a learned model that performs a plurality of mutually different inference tasks from an image. In the generation, the computer causes the learned model to Extract a first feature amount common to the plurality of inference tasks from the image. For each inference task, extract a second feature amount unique to the corresponding inference task from the first feature amount. Combine the second feature amounts extracted for each inference task to generate a third feature amount. For each inference task, output an inference result of the corresponding inference task from the third feature amount. Image processing method.
8. An image processing method by a computer, The computer estimates and outputs inference results of a plurality of mutually different inference tasks from an image using a learned model. In the estimation, the computer causes the learned model to Extract a first feature amount common to the plurality of inference tasks from the image. For each of the inference tasks, extract a second feature amount specific to the corresponding inference task from the first feature amount. Combine the second feature amounts extracted for each of the inference tasks to generate a third feature amount. For each of the inference tasks, output an inference result of the corresponding inference task from the third feature amount. Image processing method.
9. A program for causing a computer to perform a process of generating a learned model that performs a plurality of different inference tasks from an image, In the generation, in the learned model, extract a first feature amount common to the plurality of inference tasks from the image, for each of the inference tasks, extract a second feature amount specific to the corresponding inference task from the first feature amount, combine the second feature amounts extracted for each of the inference tasks to generate a third feature amount, for each of the inference tasks, output an inference result of the corresponding inference task from the third feature amount. Program.
10. A program for causing a computer to perform a process of estimating and outputting inference results of a plurality of different inference tasks from an image using a learned model, In the estimation, in the learned model, extract a first feature amount common to the plurality of inference tasks from the image, for each of the inference tasks, extract a second feature amount specific to the corresponding inference task from the first feature amount, combine the second feature amounts extracted for each of the inference tasks to generate a third feature amount, for each of the inference tasks, output an inference result of the corresponding inference task from the third feature amount. Program.
Citation Information
Patent Citations
Multitask processing device, multitask model learning device, and program
JP2018055377A
Information processing apparatus, information processing method, and program
JP2019192009A
Information processing apparatus and program
JP2021021978A