Information processing device, information processing method, and program

The information processing apparatus employs reinforcement learning to identify objects in images without positional annotations, addressing the inefficiency of existing techniques and enhancing object recognition accuracy.

WO2025109696A1PCT designated stage expired Publication Date: 2025-05-30FAST ACCOUNTING INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2023/041869
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-11-21
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Existing image recognition techniques require annotations indicating the position of objects within images, which is time-consuming to prepare.

Method used

An information processing apparatus that uses reinforcement learning to identify objects in images without relying on positional annotations, by acquiring image data, selecting operations on regions of the image, calculating confidence levels, and rewarding agents based on their performance.

Benefits of technology

Enables the determination of object types within images without prior positional information, improving efficiency and accuracy in logo recognition tasks where object size and position vary.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2023041869_30052025_PF_FP_ABST
    Figure JP2023041869_30052025_PF_FP_ABST
Patent Text Reader

Abstract

An information processing device 1 comprising: an acquisition unit 131 that acquires image data including, as a photographic subject, an object under assessment; a selection unit 132 whereby an agent which performs an operation on a region that includes at least some of the image data, said region being for specifying the object under assessment, is caused to select an operation for modifying the region; a calculation unit 133 that calculates a certainty factor regarding the type of the object under assessment included in the region after the operation selected by the agent has been performed; a reward determination unit 134 that provides the agent with a reward determined on the basis of the certainty factor calculated by the calculation unit; and an output unit 135 that, when the agent has selected an operation for determining the region in which the object under assessment is located, outputs the location of the region at the time when said determination operation was performed, and the type of the object under assessment in the image data.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing device, information processing method, and program

[0001] The present invention relates to an information processing device, an information processing method, and a program.

[0002] Techniques for identifying objects included as subjects in an image are known (for example, Non-Patent Document 1).

[0003] Juan C Caicedo and Svetlana Lazebnik, “Active object localization with deep reinforcement learning,” in Proceedings of the IEEE International Conference on Computer Vision, 2015, pp. 2488-2496.

[0004] Existing image recognition technologies require annotations that indicate the position of the object being judged contained in the image, and there is a problem in that preparing annotations that indicate the position in the image is time-consuming.

[0005] Therefore, the present invention has been made in consideration of these points, and aims to make it possible to identify an object to be determined contained in an image without relying on an annotation indicating the position of the object to be determined.

[0006] An information processing device of a first aspect of the present invention has an acquisition unit that acquires image data including an object to be determined as a subject, a selection unit that allows an agent operating an area that includes at least a portion of the image data and is used to identify the object to be determined to select an operation to change the area, a calculation unit that calculates a certainty factor for the type of object to be determined that will be included in the area after the operation selected by the agent is performed, a reward determination unit that gives the agent a reward determined based on the certainty factor calculated by the calculation unit, and an output unit that, when the agent selects an operation to determine the area in which the object to be determined is located, outputs the position of the area at the time the determination operation is performed and the type of the object to be determined in the image data.

[0007] The calculation unit may calculate the probability that the object to be determined belongs to each of a plurality of predetermined classes based on the operation selected by the agent, and calculate a certainty of the type of the object to be determined included in the area based on the calculated probability.

[0008] The selection unit may cause the agent to select an operation to be performed on the area based on an operation previously selected by the agent and an image feature extracted from the image data.

[0009] The reward determination unit may give a positive reward to the agent when the confidence level after the agent performs the selected operation is higher than the confidence level before the operation, and may give a negative reward to the agent when the confidence level after the determined operation is lower than the confidence level before the operation.

[0010] The reward determination unit may give a positive reward to the agent when the agent determines the final position of the object to be judged in the image data and the confidence level in the final position of the object to be judged is greater than a threshold value.

[0011] The image processing device may further include a learning unit that uses training image data, which is image data for learning that includes the object to be determined as a subject, and the type of object to be determined contained in the training image data as training data, as training data to teach the agent how to determine the area in which the object to be determined is located.

[0012] An information processing method of a second aspect of the present invention includes the steps of: acquiring image data including an object to be determined as a subject, executed by a computer; having an agent operating an area including at least a portion of the image data for identifying the object to be determined select an operation to change the area; calculating a certainty factor for the type of object to be determined included in the area after the operation selected by the agent is performed; giving the agent a reward determined based on the calculated certainty factor; and, when the agent selects an operation to determine the area in which the object to be determined is located, outputting the position of the area at the time the determination operation was performed and the type of the object to be determined in the image data.

[0013] In a third aspect of the program of the present invention, a computer is caused to execute the steps of acquiring image data including an object to be judged as a subject, having an agent that operates an area that includes at least a portion of the image data and is used to identify the object to be judged select an operation to change the area, calculating a degree of certainty of the type of object to be judged that will be included in the area after the operation selected by the agent is performed, giving the agent a reward determined based on the calculated degree of certainty, and when the agent selects an operation to determine the area in which the object to be judged is located, outputting the position of the area at the time the determination operation was performed and the type of the object to be judged in the image data.

[0014] According to the present invention, it is possible to determine the type of an object to be determined contained in an image even when there is no information on the position of the object to be determined in the image.

[0015] FIG. 1 is a diagram for explaining an overview of an information processing system S. FIG. 2 is a block diagram showing the configuration of an information processing device 1. FIG. 3 is a diagram showing a schematic diagram of processing in the information processing device 1. FIG. 4 is a diagram showing an example of a screen displayed by an output unit 135. FIG. 5 is a flowchart showing the flow of processing in the information processing device 1.

[0016] 1 is a diagram illustrating an overview of the information processing system S. The information processing system S is a system for identifying an object included in an image. The information processing system S includes an information processing device 1 and an information terminal 2.

[0017] The information processing device 1 is a device that determines the type of object contained in an image and the position where the object is contained. More specifically, the information processing device 1 acquires an image and outputs the position of the object contained in the acquired image and the class of the object. In this specification, the class indicates the type of object to be determined. For example, if the object to be determined is a logo for identifying a company or service, the class indicates the type of logo. The logo may indicate a business that provides a product or the like, or may represent a product or the like provided by the business. The information processing device 1 is, for example, a server. The number of classes is determined based on the type of object to be determined contained in image data provided as training data.

[0018] The information terminal 2 is a terminal operated by a user. The information terminal 2 is, for example, a smartphone, a tablet, or a personal computer. The information terminal 2 transmits an image to be judged to the information processing device 1, acquires the judgment result of the information processing device 1, and displays the acquired judgment result on a display unit.

[0019] The processing in the information processing system S will be described with reference to Fig. 1(a). The information processing device 1 acquires image data D ((1) in Fig. 1). The image data D includes a determination target object as a subject. The information processing device 1 identifies the position and type of the determination target object included in the acquired image data D ((2) in Fig. 1). The information processing device 1 outputs the identified position and type of the determination target object to the information terminal 2 ((3) in Fig. 1).

[0020] The process for identifying the position and type of the object to be determined will be described with reference to FIG. 1( b). The information processing device 1 uses a reinforcement learning technique as shown below to have agent A select an operation to be performed on region R, and identify the position and type of the object in the image. Region R is an area for identifying the object to be determined in the input image, and is a so-called bounding box. Agent A determines the operation to be performed on region R based on the feature amount of image data D. Agent A is a trained model that has learned the operation to be performed on region R so as to maximize the reward given in response to the operation. As an example, agent A is a trained model generated by training using a known DQN (Deep Q Network).

[0021] Agent A selects an operation to perform on region R based on the feature quantities of the image included in region R in image data D ((2-1) in Figure 1). Agent A predicts the reward to be given if each operation is performed based on the feature quantities of the image included in region R in image data D, and selects an operation based on the predicted reward. As an example, agent A selects an operation that maximizes the reward. The operation to perform on region R is any of enlarging, reducing, changing the aspect ratio of region R, moving vertically or horizontally, and ending the operation. In other words, when agent A performs an operation on region R, the position, size, and shape of region R are changed. Note that Figure 2(b) shows an example in which an operation to reduce region R is performed.

[0022] The information processing device 1 calculates a confidence score for each class of the classification target based on the feature amounts of the image included in the region R after the operation ((2-2) in FIG. 1). When an image is input, the information processing device 1 calculates a prediction probability using a class prediction model that has been trained to output a prediction probability for each class, and calculates a confidence score based on the calculated prediction probability. The prediction probability indicates the probability that the object to be determined is predicted to belong to the class. The confidence score is an index that indicates the uncertainty of the prediction. If the prediction probabilities of multiple classes are high, the confidence score will be small, and if only the prediction probability of a specific class is high, the confidence score will be large. The information processing device 1 determines a reward to be given to agent A based on the calculated confidence score. The method for determining the reward will be described later.

[0023] When the information processing device 1 acquires an image of the object to be judged, it identifies the position and type of the object to be judged by repeating the above-mentioned processes (2-1) to (2-3) until agent A selects to end the operation.

[0024] If region R does not properly capture the object to be determined contained in the image, the certainty factor is considered to be small. In other words, if the certainty factor is large, region R is considered to specify an appropriate position. Therefore, by using the certainty factor, the information processing device 1 can determine the type of object to be determined contained in the image even when there is no annotation indicating the position of the object to be determined. The information processing device 1 is particularly suitable for logo recognition tasks in which logos contained in images vary in size and position.

[0025] 2 is a block diagram showing the configuration of the information processing device 1. The information processing device 1 has a communication unit 11, a storage unit 12, and a control unit 13. The control unit 13 has an acquisition unit 131, a selection unit 132, a calculation unit 133, a reward determination unit 134, an output unit 135, and a learning unit 136.

[0026] Details of the information processing device 1 will be described with reference to Fig. 3. Fig. 3 is a diagram schematically illustrating processing in the information processing device 1. The acquisition unit 131 acquires image data D including a determination target object as a subject. The acquisition unit 131 acquires the image data D from the information terminal 2. The determination target object is an object to be classified and located by image recognition. As an example, the determination target object is a logo mark for identifying a company or a service provided by a company. In other words, the image data D is image data including an object with a logo mark attached as a subject.

[0027] The acquisition unit 131 inputs the acquired image data D to an image encoder and extracts image features. Note that, when acquiring image data D, the acquisition unit 131 may set region R at a predetermined position (for example, a region including the entire image) that is set in advance as an initial position. Note that the acquisition unit 131 may cause the calculation unit 133 to calculate the confidence level at the time when the image data D is acquired, as will be described later.

[0028] The selection unit 132 causes the agent A, which operates an area that includes at least a part of the image data D and is used to identify the object to be determined, to select an operation to change the area R. The selection unit 132 inputs the feature amount included in the image to the agent A, and causes the agent A to output information indicating the operation to be performed on the area R.

[0029] The calculation unit 133 calculates a certainty factor for the type of the object to be determined included in the region R in which the operation selected by agent A was performed. The calculation unit 133 calculates the probability that the object included in the region R of the image data D belongs to each of multiple classes based on the operation selected by agent A, and calculates a certainty factor for the type of the object to be determined included in the region R based on the calculated probabilities. As an example, the storage unit 12 stores a class prediction model, which is a trained model that has been trained to output a predicted probability for each class when feature amounts of an image are input, and the calculation unit 133 inputs the feature amounts of the image into the class prediction model, thereby outputting the predicted probability for each class and calculating a certainty factor. The calculation unit 133 calculates a certainty factor for each class into which the image can be classified based on the feature amounts of the image included in the region R. The calculation unit 133 may normalize the certainty factor for each class.

[0030] The reward determination unit 134 awards the reward determined based on the confidence level calculated by the calculation unit 133 to the agent A. The reward determination unit 134 determines the reward to be awarded to the agent A based on the confidence level based on the region R before the operation and the confidence level based on the region R after the operation. As an example, the reward determination unit 134 may award a positive reward to the agent A when the confidence level based on the region after the operation is higher than the confidence level based on the region before the operation, and may award a negative reward to the agent A when the confidence level based on the region after the operation is lower than the confidence level based on the region before the operation.

[0031] As an example, the reward determination unit 134 determines the reward based on the following formula. Here, Re represents the reward, C' represents the certainty of the target class based on the region R after the operation, and C represents the certainty of the target class based on the region R before the operation. Note that sign(C'-C) is a sign function that is -1 when the sign of the argument is negative, +1 when the sign of the argument is positive, and 0 when the argument is 0. The target class is, for example, the class with the highest certainty among the classes to be determined.

[0032]

[0033] The selection unit 132, the calculation unit 133, and the reward determination unit 134 repeat the above-described processing until agent A selects the end of the operation. When agent A selects the operation to determine the area in which the object to be determined is located, the output unit 135 outputs the position of region R at the time of the determination operation and the type of the object to be determined in the image data D. The output unit 135 may cause the information terminal 2 to display a screen including the position of region R at the time the end of the operation was selected and the class of the determination result. FIG. 4 is a diagram showing an example of a screen displayed by the output unit 135. In the screen shown in FIG. 4, an object indicating the position of region R at the time the end of the operation was selected is superimposed on the image P of the object to be determined, and information indicating the class of the determination result is also displayed. The class of the determination result is, for example, the class with the highest certainty at the time the end of the operation was selected.

[0034] The learning unit 136 updates the parameters of the agent A based on the reward predicted by the agent A and the reward actually given to the agent A as a result of repeating the operation, and causes the agent A to learn.

[0035] In addition, in the processing of the inference stage when agent A's learning is completed, the processing of the reward determination unit 134 may be omitted.

[0036] By configuring the information processing device 1 in this way, it is possible to determine the type of an object to be determined contained in an image even when there is no information on the position of the object to be determined in the image.

[0037] The selection unit 132 may be configured to determine an operation by further using a history of past operations as input data. That is, the selection unit 132 causes agent A to select an operation to be performed on region R based on the operation selected immediately before by agent A and the feature amount of the image extracted from image data D. The selection unit 132 may also cause agent A to select an operation to be performed on region R based on the contents of the past several operations. By configuring the selection unit 132 in this way, it is possible to select a more appropriate operation content, thereby improving the accuracy of the determination.

[0038] The reward determination unit 134 may be configured to determine a reward to be given to agent A based on the confidence level when the operation on region R is completed and the final position of region R is determined. When agent A determines the final position of the object to be determined in image data D, the reward determination unit 134 gives a positive reward to agent A if the confidence level for the final position of the object to be determined is greater than a threshold. As an example, the reward when the position of the final region R is determined is expressed by the following equation:

[0039] Here, Reω denotes the reward when the final position is determined, η denotes a value preset as the reward for the final position, and τ denotes a value preset as a threshold value for the confidence level. The value of η may be set to be larger than the reward given when the operation content is determined. The value of τ is determined based on the recognition accuracy of the object to be determined.

[0040] By configuring the reward determination unit 134 in this way, it is possible to strike a balance between the reward based on the determined operation and the reward based on the final position, thereby improving the accuracy of the determination.

[0041] [Regarding Learning] The information processing device 1 may be configured to perform agent learning in two stages. First, the acquisition unit 131 acquires learning image data and a class to which the subject included in the learning image data belongs. The learning image data is image data for learning that includes a determination target object as a subject. The learning unit 136 uses the learning image data and the class of the determination target object included in the learning image data as training data to have agent A learn an operation to determine the area in which the determination target object is located.

[0042] The learning unit 136 trains a class prediction model using the training image data and the class of the object to be determined contained in the training image data. As an example, the learning unit 136 may update the parameters of the class prediction model based on the cross-entropy error between the class predicted by the class prediction model and the correct class. At this stage, the learning unit 136 trains the class prediction model using training image data in which the object to be determined is prominently displayed in the image.

[0043] Next, the learning unit 136 simultaneously trains the agent A and the class prediction model based on the training image data and the class of the object to be determined included in the training image data. As an example, the learning unit 136 updates the parameters of the agent A based on the reward predicted by the agent A and the reward actually given to the agent A.

[0044] By configuring the learning unit 136 to learn in two stages in this way, it is possible to efficiently and comprehensively learn information about the object to be determined and the optimal operation.

[0045] [Processing Flow in Information Processing Device 1] Fig. 5 is a flowchart showing the processing flow in the information processing device 1. The flowchart shown in Fig. 5 starts when the information processing device 1 receives an instruction to execute the determination process.

[0046] The acquisition unit 131 acquires image data D (S01). The acquisition unit 131 inputs the image data D to an image encoder and extracts image features (S02). The acquisition unit 131 determines the initial position of a region R (S03).

[0047] The calculation unit 133 calculates the certainty factor of each class based on the feature amount of the image and the position of the region R (S04). Specifically, the calculation unit 133 calculates the predicted probability that the object to be determined included in the image data D belongs to each class based on the extracted feature amount, and calculates the certainty factor based on the calculated predicted probability.

[0048] The selection unit 132 inputs the feature amount of the image and the position of the region R to the agent A, and causes the agent A to select an operation for the region R (S05). The selection unit 132 may further input information indicating past operations to the agent A, and cause the agent A to output an operation.

[0049] The selection unit 132 determines whether or not a termination condition is satisfied (S06). The termination condition is that the agent A has selected to end the operation. If the termination condition is satisfied (YES in S06), the information processing device 1 proceeds to S10.

[0050] If the termination condition is not satisfied (NO in S06), the calculation unit 133 calculates the confidence level of each class based on the feature amount of the image and the position of the region R after the operation (S07). The reward determination unit 134 determines the reward to be given to agent A based on the confidence level based on the region R before the operation and the confidence level based on the region R after the operation (S08).

[0051] The information processing device 1 determines whether or not the processes of S05 to S08 have been executed a predetermined number of times (S09). If the number of times the processes of S05 to S08 have been executed is equal to or greater than the predetermined number of times (YES in S09), the information processing device 1 proceeds to the process of S10. If the predetermined number of times has not been reached (NO in S09), the information processing device 1 proceeds to the process of S05. Note that instead of or in addition to whether the processes of S05 to S08 have been executed a predetermined number of times, the information processing device 1 may proceed to the process of S10 if the reward has reached a certain value.

[0052] If the termination condition is met (YES in S06), or if the number of times the processes of S05 to S08 have been executed is equal to or greater than a predetermined number (YES in S09), the calculation unit 133 determines a reward to be given to agent A based on the confidence level when the final position of region R is determined (S10). The output unit 135 outputs the determination result (S11). Specifically, the output unit 135 outputs the type and position of the object to be determined based on the prediction result calculated by the calculation unit 133 and the final position of region R determined by agent A. Then, the information processing device 1 ends the processing.

[0053] [Effects of the Present Embodiment] As described above, the information processing device 1 can determine the type of an object to be determined contained in an image even when there is no information on the position of the object to be determined in the image.

[0054] The present invention has been described above using embodiments, but the technical scope of the present invention is not limited to the scope described in the above embodiments, and various modifications and changes are possible within the scope of the gist of the present invention. For example, all or part of the device can be configured by functionally or physically distributing or integrating in any unit. Furthermore, new embodiments resulting from any combination of multiple embodiments are also included in the embodiments of the present invention. The effects of the new embodiments resulting from the combination also have the effects of the original embodiments.

[0055] REFERENCE SIGNS LIST 1 Information processing device 2 Information terminal 11 Communication unit 12 Storage unit 13 Control unit 131 Acquisition unit 132 Selection unit 133 Calculation unit 134 Reward determination unit 135 Output unit 136 Learning unit

Claims

1. An acquisition unit that acquires image data including the object to be determined as a subject; a selection unit that causes an agent that operates on the region for identifying the object to be determined, which is a region including at least a part of the image data, to select an operation for changing the region; a calculation unit that calculates the confidence level of the type of the object to be determined included in the region after the operation selected by the agent is performed; a reward determination unit that gives a reward determined based on the confidence level calculated by the calculation unit to the agent; and an output unit that outputs the position of the region at the time when the operation of determining the region where the object to be determined is located is selected by the agent and the type of the object to be determined in the image data. An information processing apparatus having the above components.

2. The calculation unit calculates the probability of belonging to each of a plurality of classes predetermined as the type of the object to be determined based on the operation selected by the agent, and calculates the confidence level of the type of the object to be determined included in the region based on the calculated probability. The information processing apparatus according to claim 1.

3. The selection unit causes the agent to select an operation to be performed on the region based on the operation immediately previously selected by the agent and the feature amount of the image extracted from the image data. The information processing apparatus according to claim 1.

4. The reward determination unit gives a positive reward to the agent when the confidence level when the operation selected by the agent is performed is higher than the confidence level before the operation, and gives a negative reward to the agent when the confidence level when the determined operation is performed is lower than the confidence level before the operation. The information processing apparatus according to claim 3.

5. The reward determination unit gives a positive reward to the agent when the confidence level at the final position of the object to be determined is greater than a threshold value when the agent determines the final position of the object to be determined in the image data. The information processing apparatus according to claim 3 or 4.

6. A learning unit that further has learning image data, which is learning image data including the object to be determined as a subject, and the type of the object to be determined included in the learning image data as teacher data, and causes the agent to learn an operation of determining the region where the object to be determined is located. The information processing apparatus according to claim 1.

7. A computer-executed information processing method comprising: a step of acquiring image data including an object to be determined as a subject; a step of causing an agent that operates on a region including at least a part of the image data and that is for specifying the object to be determined to select an operation for changing the region; a step of calculating a confidence level of the type of the object to be determined included in the region after the operation selected by the agent is performed; a step of giving a reward determined based on the calculated confidence level to the agent; and a step of outputting the position of the region at the time when the operation of determining the region where the object to be determined is located is selected by the agent and the type of the object to be determined in the image data.

8. A program for causing a computer to execute: a step of acquiring image data including an object to be determined as a subject; a step of causing an agent that operates on a region including at least a part of the image data and that is for specifying the object to be determined to select an operation for changing the region; a step of calculating a confidence level of the type of the object to be determined included in the region after the operation selected by the agent is performed; a step of giving a reward determined based on the calculated confidence level to the agent; and a step of outputting the position of the region at the time when the operation of determining the region where the object to be determined is located is selected by the agent and the type of the object to be determined in the image data.

Citation Information

Patent Citations

  • Method, system, and computer program

    WO2022124224A1

  • Image processing device, image processing system, image processing method, and program

    WO2022255418A1