Information processing method
The information processing method employs reinforcement learning to determine the optimal order of image processing techniques for target detection, addressing inefficiencies and improving robustness in existing adaptive image filter generation devices.
Patent Information
- Application Number
- JP2023203362
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-30
- Publication Date
- 2025-06-11
AI Technical Summary
Existing generation devices for adaptive image filters require significant manual effort and are inefficient in target detection, lacking robustness and speed.
An information processing method that uses reinforcement learning to generate a trained model that learns the optimal order of applying image processing methods to detect targets efficiently and accurately.
The method reduces manual work, enables quick and accurate target detection, and achieves high robustness by automatically learning the combination and order of image processing techniques.
Smart Images

Figure 2025088578000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to an information processing method.
Background Art
[0002] Conventionally, in a generation device that generates an adaptive image filter adapted to convert a learning input image into a target image, it has been disclosed to select an adaptive image filter based on the degree of fitness between the output image by each of a plurality of image filters and the target image (see, for example, Patent Document 1).
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] The generation device in the above prior art document generates a plurality of output images and calculates the degree of fitness with the target image for each of them. Therefore, there is room for improvement from the viewpoint of efficiency.
[0005] In view of such circumstances, an object of the present disclosure is to reduce manual work in target detection, perform detection quickly and accurately, and achieve high robustness.
Means for Solving the Problems
[0006] An information processing method according to an embodiment of the present disclosure is an information processing method by an information processing apparatus, receiving an input of a first image in which an object including a target is imaged and a target image in which the target is annotated with respect to the first image; Applying each of a plurality of image processing methods targeting the target image to the first image to generate a plurality of second images; Applying each of the plurality of image processing methods to the second images to generate a plurality of third images; Generating a trained model that learns the order in which the image processing methods are applied by reinforcement learning and outputs the order; including.
Advantages of the Invention
[0007] According to an embodiment of the present disclosure, in target detection, manual work can be reduced, detection can be performed quickly and accurately, and high robustness can be realized.
Brief Description of the Drawings
[0008]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Embodiments for Carrying Out the Invention
[0009] FIG. 1 is a schematic diagram of an information processing apparatus 1 according to the present embodiment. The information processing apparatus 1 can communicate with one or more other terminals via a network. The network includes, for example, a mobile communication network, the Internet, or a fixed communication network.
[0010] In FIG. 1, for simplicity of explanation, only one information processing apparatus 1 is shown. However, the number of information processing apparatuses 1 is not limited to this. For example, the processing executed by the information processing apparatus 1 may be executed by a plurality of information processing apparatuses 1 arranged in a distributed manner.
[0011] The information processing device 1 is a computer such as a server belonging to a cloud computing system or other computing system. The information processing device 1 may be installed, for example, in a facility dedicated to an operator or a shared facility including a data center.
[0012] The internal configuration of the information processing device 1 will be described in detail. The information processing device 1 includes a control unit 11, a communication unit 12, and a storage unit 13. Each component of the information processing device 1 is communicably connected to each other.
[0013] The control unit 11 includes, for example, one or more general-purpose processors including a CPU (Central Processing Unit) or an MPU (Micro Processing Unit). The control unit 11 may include one or more dedicated processors specialized for specific processing. Instead of including a processor, the control unit 11 may include one or more dedicated circuits. The dedicated circuit may be, for example, an FPGA (Field-Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit). The control unit 11 may include an ECU (Electronic Control Unit). The control unit 11 controls the communication unit 12 to transmit and receive arbitrary information.
[0014] The communication unit 12 includes a communication module corresponding to one or more wired or wireless LAN (Local Area Network) standards for connecting to a network. The communication unit 12 may include a module corresponding to one or more mobile communication standards including LTE (Long Term Evolution), 4G (4th Generation), or 5G (5th Generation). The communication unit 12 may include a communication module or the like corresponding to one or more short-range communication standards or specifications including Bluetooth (registered trademark), AirDrop (registered trademark), IrDA, ZigBee (registered trademark), Felica (registered trademark), or RFID. The communication unit 12 transmits and receives arbitrary information via the network.
[0015] The storage unit 13 includes, for example, a semiconductor memory, a magnetic memory, an optical memory, or a combination of at least two of these, but is not limited thereto. The semiconductor memory is, for example, a RAM or a ROM. The RAM is, for example, an SRAM or a DRAM. The ROM is, for example, an EEPROM. The storage unit 13 may function as, for example, a main storage device, an auxiliary storage device, or a cache memory. The storage unit 13 may store the information of the result analyzed or processed by the control unit 11. The storage unit 13 may store various information related to the operation or control of the information processing apparatus 1. The storage unit 13 may store a system program, an application program, embedded software, and the like. The storage unit 13 may be provided outside the information processing apparatus 1 and accessed from the information processing apparatus 1. The storage unit 13 includes a reward value table.
[0016] Hereinafter, an outline of the information processing method in the present embodiment will be described. In the present embodiment, an algorithm is disclosed that outputs, by learning, an optimal combination and application order of image processing techniques to be applied to a captured image of an object in order to detect an abnormal part in the object from the captured image. In the present embodiment, a learned model generated by reinforcement learning is used. The learned model outputs, for example, in an order table TB, a combination and application order of image processing techniques to be applied to the input image in order to detect an abnormal part of the object included in the input image. The order table TB shown in FIG. 2 is an example, and is associated with the application order ("STEP1", "STEP2", and "STEP3"), the image processing technique to be applied ("C", "B", and "A"), the function used in the image processing technique ("c", "b", and "a"), and the parameters used in the image processing technique ("γ", "β", and "α"). The order table TB indicates that by applying the three image processing techniques C, B, and A to the input image in this order, an abnormal part of the object included in the input image can be detected.
[0017] The learning method of the model will be described below. As shown in FIG. 2, the control unit 11 of the information processing apparatus 1 receives an input of a first image G1 in which an object including a target TA is imaged, and a target image GT in which the target TA is annotated by annotation with respect to the first image G1. The target image GT may be prepared manually. The target image may be referred to as a correct image. The target image here is an image on which binarization processing has been performed. The target TA here indicates an abnormal part.
[0018] Here, an outline of the model learning method will be described. The control unit 11 determines the first image G1 as the target image. The control unit 11 executes a step (image processing step) of applying each of a plurality of image processing methods targeting the target image to the target image to generate a plurality of candidate images. The control unit 11 uses one candidate image selected from among the plurality of candidate images generated in the immediately preceding image processing step as a new target image, and executes the image processing step one or more times further. Therefore, the image processing step is executed a plurality of times while changing the target image. Note that some or all of the "plurality of image processing methods" applied to the target image may be different depending on the number of executions of the image processing step. Also, the number of the "plurality of image processing methods" may be different depending on the number of executions of the image processing step. Then, the control unit 11 learns the combination and application order of the image processing methods applied until each candidate image is generated from the first image G1 by reinforcement learning, and generates a learned model.
[0019] Hereinafter, the information processing method in the present embodiment will be described in detail. For the sake of simplicity of explanation, an example in which the above-described image processing step is executed a total of two times will be described, but the number of executions of the image processing step may be three or more.
[0020] The control unit 11 executes the first image processing step. Specifically, the control unit 11 applies each of a plurality of image processing methods to a target image (here, the first image G1) to generate a plurality of candidate images (here, a plurality of second images G2). In the example shown in FIG. 3, the control unit 11 applies the image processing method A to the first image G1 to generate the second image G2-A. The control unit 11 applies the image processing method B to the first image G1 to generate the second image G2-B. The control unit 11 applies the image processing method C to the first image G1 to generate the second image G2-C. The number of image processing methods applied is not limited to three, and can be arbitrarily set as long as it is two or more. Note that the "image processing method" may be represented by, for example, a combination of a function and parameters. When at least one of the function and the parameters is different, the image processing method is considered different.
[0021] The control unit 11 applies binarization processing to each of the plurality of second images G2 to generate a plurality of binarized images G2B. In the example shown in FIG. 3, the control unit 11 executes binarization processing on each of the second image G2-A, the second image G2-B, and the second image G2-C to generate the binarized image G2B-A, the binarized image G2B-B, and the binarized image G2B-C.
[0022] The control unit 11 calculates a reward value corresponding to the transition from the first image G1 to each second image G2. Specifically, the control unit 11 calculates the degree of overlap (for example, the degree of coincidence) between each binarized image G2B and the target image GT as the reward value corresponding to the transition from the first image G1 to each second image G2. The method of calculating the overlap is arbitrary. The reward value may be a numerical value between 0 and 1. In this case, the closer the reward value is to 0, the smaller the overlap, and the closer the reward value is to 1, the larger the overlap. For the sake of convenience of explanation, the "reward value corresponding to the transition from the first image G1 to the second image G2" is also referred to as the "reward value of the second image G2" or the "reward value of the binarized image G2B obtained by binarizing the second image G2". As shown in FIG. 3, the reward values of the binarized image G2B-A, the binarized image G2B-B, and the binarized image G2B-C are 0.4, 0.1, and 0.7, respectively.
[0023] As an additional example or an alternative example, in addition to the degree of match, the control unit 11 may calculate a reward value based on at least one of the cost, energy saving effect, and speed of each image processing method. For example, the control unit 11 may calculate a cost score, an energy saving effect score, and a speed score. The higher the cost score, the lower the cost of the image processing method. The higher the energy saving effect score, the greater the energy saving effect of the image processing method (for example, the lower the power consumption required for the execution of the image processing method). The higher the speed score, the faster the processing speed of the image processing method. The control unit 11 may correct the reward value by adding each score to the reward value calculated based on the degree of match. Alternatively, the control unit 11 may correct the reward value so that it approaches 1 as each score is higher. According to such a configuration, for example, when there are two image processing methods with the same reward value based only on the degree of match, the reward value can be increased for the more advantageous image processing method from the viewpoints of cost, energy saving effect, and speed. Note that it is desirable that the correction amount of the reward value by each score is sufficiently small so that the reward value calculated based only on the degree of match (the reward value before correction) is dominant.
[0024] The control unit 11 specifies the maximum reward value among the reward values calculated for each of the plurality of binarized images G2B, and specifies the second image for which the maximum reward value was calculated. In the example shown in FIG. 3, the second image for which the maximum reward value was calculated is the second image G2-C for which a reward value of 0.7 was calculated.
[0025] As shown in row 41 of FIG. 4, the control unit 11 stores, in the reward value table of the storage unit 13, the applied image processing method, the transition of the image due to the application of the image processing method, and the reward value, in association with the target image GT. The reward values corresponding to the image processing method A and the image processing method B are 0.4 and 0.1, respectively, and the conversion to the target image GT fails. However, the control unit 11 may store, in the storage unit 13, information on the transition and the reward value corresponding to the failed image processing method for learning of failure cases.
[0026] As shown in line 42 of FIG. 4, as an additional example, the control unit 11 associates with each of the binarized images G2B-A, G2B-B, and G2B-C, the applied image processing method, the transition of the image due to the application of the image processing method, and the reward value, and stores them in the storage unit 13. The reinforcement learning described later in this case may include reinforcement learning using the Hindsight Experience Replay (HER) algorithm. Each of the binarized images G2B-A, G2B-B, and G2B-C is not the true target image, but the control unit 11 replaces the target image GT with a plurality of binarized images G2B and stores them according to the experience replay technology algorithm. The reward value for the transition to each of the second images G2-A, G2-B, and G2-C is regarded as 1 and stored.
[0027] The control unit 11 executes the second image processing step. Specifically, the control unit 11 applies each of a plurality of image processing methods to the target image (here, the second image G2-C for which the maximum reward value of 0.7 is calculated) to generate a plurality of candidate images (here, a plurality of third images G3). For example, the control unit 11 applies the image processing method A2 to the second image G2-C to generate the third image G3-A. The control unit 11 applies the image processing method B2 to the second image G2-C to generate the third image G3-B. The control unit 11 applies the image processing method C2 to the second image G2-C to generate the third image G3-C. The number of applied image processing methods is not limited to three, and can be arbitrarily set as long as it is two or more. Also, as described above, the image processing methods A2, B2, and C2 in the second image processing step may be the same as or different from the image processing methods A, B, and C in the first image processing step. The control unit 11 calculates the reward value for each of the plurality of third images G3. For the sake of convenience of explanation, the case where the control unit 11 calculates the reward value as follows is explained. Reward value of the third image G3-A generated by the image processing method A2: 0.2 Reward value of the third image G3-B generated by the image processing method B2: 0.8 Reward value of the third image G3-C generated by the image processing method C2: 0.5
[0028] As an additional example or an alternative example, the control unit 11 may calculate a search value for each of a plurality of second images. For example, the control unit 11 determines that there is a search value for a second image G2 whose reward value exceeds a reference value. As an alternative example, the method for determining the search value is arbitrary. The control unit 11 may select a second image for which it is determined that there is a search value, and apply each of a plurality of image processing methods to the selected second image to generate a plurality of third images G3.
[0029] The control unit 11 stores, in the reward value table of the storage unit 13, the applied image processing method (here, image processing methods A2, B2, or C2), the transition of the image due to the application of the image processing method (here, the transition from the second image G2-C to each third image G3), and the reward value, in association with the target image GT. As an additional example, the control unit 11 stores, in the storage unit 13, the applied image processing method, the transition of the image due to the application of the image processing method, and the reward value, in association with each binarized image G3B obtained by binarizing each third image G3. The reinforcement learning described later performed in this case may include reinforcement learning using the Hindsight Experience Replay (HER) algorithm.
[0030] The control unit 11 determines that the reward value of the third image generated by the image processing method B is the maximum, and selects the third image. The control unit 11 stores the following order in the storage unit 13, assuming that the following order is the most optimal as the order for converting the first image G1 into the target image GT. First: Image processing method C Second: Image processing method B2
[0031] The control unit 11 may repeat the image processing step and the above-described application process until a reward value exceeding a reference value (for example, 0.7) is calculated. The control unit 11 stores in the storage unit 13 the order in which the image processing methods were applied until a reward value exceeding the reference value is calculated. The control unit 11 learns the stored order using reinforcement learning. As an alternative example, the control unit 11 may store only the order in which the image processing method was applied until the maximum reward value is calculated, and learn the stored order using reinforcement learning. The reinforcement learning may be reinforcement learning that uses the Monte Carlo Tree Search (MCTS) algorithm in the search. In this case, in the Monte Carlo tree search algorithm, the control unit 11 applies each of a plurality of image processing methods targeting the target image GT to the first image G1, and calculates a plurality of reward values for the plurality of second images generated by applying the plurality of image processing methods.
[0032] Through learning, the control unit 11 generates a learned model. The learned model can output the optimal order for applying an image processing method to convert the first image G1 into the target image GT. The control unit 11 stores the generated learned model in the storage unit 13.
[0033] In FIG. 5, a flowchart showing the operation of the information processing apparatus 1 is described.
[0034] In S1, the control unit 11 receives the input of the first image G1 and the target image GT annotated with the target TA. In S2, the control unit 11 applies each of a plurality of image processing methods targeting the target image GT to the first image G1 to generate a plurality of second images G2. In S3, the control unit 11 calculates and stores a reward value for the second image G2.
[0035] In S4, the control unit 11 applies each of a plurality of image processing methods to the second image G2 for which the maximum reward value has been calculated to generate a plurality of third images G3. In S5, the control unit 11 calculates and stores a reward value for the third image G3.
[0036] In S6, the control unit 11 learns the order in which the image processing methods are applied through reinforcement learning. In S7, the control unit 11 generates a learned model that outputs the order.
[0037] As described above, according to this embodiment, the control unit 11 of the information processing apparatus 1 receives the input of the first image G1 in which the object including the target TA is imaged and the target image GT in which the target TA is annotated with respect to the first image G1, applies each of a plurality of image processing methods targeting the target image GT to the first image G1 to generate a plurality of second images G2, applies each of the plurality of image processing methods to the second images G2 to generate a plurality of third images G3, learns the order in which the image processing methods are applied through reinforcement learning, and generates a learned model that outputs the order. With this configuration, the information processing apparatus 1 can automatically learn the combination of image processing methods, so in the detection of the target TA, manual work can be reduced, the detection can be performed quickly and accurately, and high robustness can be realized. Furthermore, since the information processing apparatus 1 can explicitly represent the order of application of the image processing methods, the interpretability and visualizability can be improved.
[0038] Also, according to this embodiment, the reinforcement learning includes reinforcement learning that uses the Monte Carlo tree search algorithm in the search, and the operation of the control unit 11 includes, in the Monte Carlo tree search algorithm, applying each of a plurality of image processing methods targeting the target image GT to the first image G1 and calculating a plurality of reward values for the plurality of second images G2 generated by the application of the plurality of image processing methods, and storing the reward values. The reward value is calculated based at least on the degree of coincidence between the second image G2 and the target image GT. The reward value may be calculated based on at least one of the cost, energy saving effect, and speed of each image processing method in addition to the degree of coincidence. With this configuration, the information processing apparatus 1 can efficiently and effectively search for combinations of various image processing methods.
[0039] Also according to the present embodiment, the operation of the control unit 11 includes applying binarization processing to each of the plurality of second images G2 after applying a plurality of image processing methods, and calculating a plurality of reward values corresponding to each of the plurality of image processing methods based on the target image GT and each of the plurality of binarized images G2B generated by the binarization processing. With this configuration, the information processing apparatus 1 can calculate the reward value more accurately.
[0040] Also according to the present embodiment, the operation of the control unit 11 includes calculating a search value for each of the plurality of second images G2, and selecting at least one image from the plurality of second images G2 based on the search value, and applying a plurality of image processing methods to the selected at least one image to generate a third image G3. With this configuration, the information processing apparatus 1 can identify an appropriate image processing method for converting the first image G1 into the target image GT, so that the detection of the target TA can be performed more accurately.
[0041] Also according to the present embodiment, the operation of the control unit 11 includes storing information on the transition from the first image G1 to each of the plurality of second images G2 by applying a plurality of image processing methods. The reinforcement learning includes reinforcement learning using an experience replay technique algorithm. The operation of the control unit 11 includes storing, by the experience replay technique algorithm, the reward value of the transition from the first image G1 to each of the plurality of second images G2 as 1 after replacing the target image GT with the plurality of binarized images G2B. With this configuration, the information processing apparatus 1 can generate additional training examples from failed learning experiences by the HER algorithm, so that the problem that it is difficult to reach the target by random search can be solved.
[0042] Although the present disclosure is described based on the drawings and examples, it should be noted that those skilled in the art may make various modifications and alterations based on the present disclosure. Additionally, changes can be made without departing from the spirit of the present disclosure. For example, the functions and the like included in each means or each step can be rearranged so as not to be logically contradictory, and a plurality of means or steps can be combined into one or divided.
[0043] The drawings for explaining the embodiments according to the present disclosure are schematic. The dimensional ratios and the like on the drawings do not necessarily match the actual ones.
[0044] For example, in the above embodiment, the program for executing all or part of the functions or processes of the information processing apparatus 1 can be recorded on a computer-readable recording medium. The computer-readable recording medium includes non-transitory computer-readable media, for example, a magnetic recording device, an optical disk, a magneto-optical recording medium, or a semiconductor memory. The distribution of the program is performed, for example, by selling, transferring, or lending a portable recording medium such as a DVD (Digital Versatile Disc) or a CD-ROM (Compact Disc Read Only Memory) on which the program is recorded. Also, the distribution of the program may be performed by storing the program in the storage of an arbitrary server and transmitting the program from the arbitrary server to other computers. Further, the program may be provided as a program product. As embodiments of the present disclosure, it is also possible to take embodiments as a system, a program, and a storage medium (as an example, an optical disk, a magneto-optical disk, a CD-ROM, a CD-R, a CD-RW, a magnetic tape, a hard disk, or a memory card, etc.) on which the program is recorded.
[0045] The implementation form of the program is not limited to application programs such as object code compiled by a compiler or program code executed by an interpreter, and may be in the form of program modules incorporated into an operating system. Furthermore, the program does not have to be configured such that all processing is performed only on the CPU on the control board. The program may be configured such that part or all of it is performed by another processing unit implemented on an expansion board or expansion unit added to the board as necessary.
Explanation of Signs
[0046] 1 Information processing apparatus 11 Control unit 12 Communication unit 13 Storage unit
Claims
1. An information processing method by an information processing apparatus, comprising: receiving an input of a first image in which an object including a target is imaged and a target image in which the target is annotated with respect to the first image; applying each of a plurality of image processing methods targeting the target image to the first image to generate a plurality of second images; applying each of a plurality of image processing methods targeting the target image to one second image selected from the plurality of second images to generate a plurality of third images; learning, by reinforcement learning, a combination and an application order of image processing methods applied until each of the third images is generated from the first image to generate a learned model; including The learned model outputs a combination and an application order of image processing methods to be applied to the input image in order to detect the target included in the input image. An information processing method.
2. In the information processing method according to claim 1, the reinforcement learning includes reinforcement learning using a Monte Carlo Tree Search (MCTS) algorithm in search, In the Monte Carlo Tree Search algorithm, calculating and storing a reward value corresponding to a transition from the first image to each of the second images; An information processing method including.
3. In the information processing method according to claim 2, each of the reward values corresponding to the transition from the first image to each of the second images is calculated based at least on a degree of coincidence between each of the second images and the target image. An information processing method.
4. In the information processing method according to claim 3, each of the reward values corresponding to the transition from the first image to each of the second images is, in addition to the degree of coincidence between each of the second images and the target image, for each of the image processing methods applied to the first image to generate each of the second images. Calculated based on at least one of cost, energy saving effect, and speed. An information processing method.
5. In the information processing method according to claim 2, including applying binarization processing to each of the plurality of second images to generate a plurality of binarized images, each of the reward values corresponding to the transition from the first image to each of the second images is calculated based at least on a degree of coincidence between each of the binarized images and the target image. An information processing method.
6. In the information processing method according to claim 1, Calculating a search value for each of the plurality of second images; Selecting the one second image from among the plurality of second images based on the search value; including; The plurality of third images are generated by applying each of the plurality of image processing methods to the selected one second image, an information processing method.
7. In the information processing method according to claim 2, An information processing method including storing information on each of the image processing methods applied to the first image to generate each of the second images and information on the transition from the first image to each of the second images.
8. In the information processing method according to claim 5, The reinforcement learning includes reinforcement learning using a hindsight experience replay (HER) algorithm, An information processing method including, by the experience replay algorithm, replacing the target image with each of the binarized images and storing the reward value corresponding to the transition from the first image to each of the second images as 1.
Citation Information
Patent Citations
Generating device, generating method, and generation program
JP2011014051A