Information processing device, information processing method, and program
Patent Information
- Application Number
- JP2025023696
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2026-08-27
AI Technical Summary
【0009】 本開示の一例示的側面によれば、物体を抽出するためのモデルを効率的に学習させることできるという一例示的効果を奏する。
Smart Images

Figure 2026137529000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to an information processing apparatus, an information processing method, and a program.
Background Art
[0002] Techniques for detecting an object from a target image are known. For example, Patent Document 1 discloses a learning apparatus that includes a supervised learning unit and a self-supervised learning unit and learns an object detection network for detecting an object from target image data.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] On the other hand, as a more advanced technology, techniques for extracting an expression of an object (object expression) included in a target image from the target image and making various predictions and estimations by referring to the object expression have also been developed. In such a technique for extracting an object expression, it has been a problem that it takes time to learn a model for extracting an object.
[0005] The present disclosure has been made in view of the above problems, and an exemplary object thereof is to provide a technique capable of efficiently learning a model for extracting an object.
Means for Solving the Problems
[0006] An information processing device relating to an illustrative aspect of this disclosure includes: a first generation means for extracting object representations of each object included in a target image and generating one or more first gaze regions based on said object representations; a second generation means for generating one or more second gaze regions from the target image using a pre-trained model; a correspondence means for estimating the correspondence between the first gaze regions and the second gaze regions; a loss calculation means for calculating a loss by referring to the corresponding first gaze regions and second gaze regions; and a learning means for training the first generation means by referring to the loss.
[0007] An example of an information processing method relating to this disclosure includes one or more processors extracting object representations of each object included in a target image using an extraction model, generating one or more first gaze regions based on said object representations, generating one or more second gaze regions from the target image using a pre-trained model, estimating the correspondence between the first gaze regions and the second gaze regions, calculating a loss by referring to the corresponding first and second gaze regions, and training the extraction model by referring to the loss.
[0008] An exemplary aspect of the present disclosure is a program that causes a computer to function as an information processing device, the computer to function as: a first generation means for extracting object representations of each object included in a target image and generating one or more first gaze regions based on said object representations; a second generation means for generating one or more second gaze regions from the target image using a pre-trained model; a correspondence means for estimating the correspondence between the first gaze regions and the second gaze regions; a loss calculation means for calculating a loss by referring to the first gaze regions and the second gaze regions that correspond to each other; and a learning means for training the first generation means by referring to the loss. [Effects of the Invention]
[0009] One exemplary aspect of this disclosure is that it enables efficient training of a model for extracting objects. [Brief explanation of the drawing]
[0010] [Figure 1] This is a block diagram showing the configuration of the information processing device related to this disclosure. [Figure 2] This is a flowchart showing the flow of the information processing method related to this disclosure. [Figure 3] This is a block diagram showing the configuration of the information processing system related to this disclosure. [Figure 4] This diagram illustrates an example of processing in the information processing system related to this disclosure. [Figure 5] This diagram illustrates the processing flow in the information processing system related to this disclosure. [Figure 6] This diagram illustrates an example configuration and processing example of the information processing system related to this disclosure. [Figure 7] This diagram illustrates an example configuration and processing example of the information processing system related to this disclosure. [Figure 8] This diagram illustrates an example configuration and processing example of the information processing system related to this disclosure. [Figure 9] This diagram illustrates an example configuration and processing example of the information processing system related to this disclosure. [Figure 10] This is a block diagram showing the configuration of the information processing system related to this disclosure. [Figure 11] This is a block diagram showing the configuration of a computer that functions as an information processing device related to this disclosure. [Modes for carrying out the invention]
[0011] The following are examples of embodiments of the present invention. However, the present invention is not limited to the exemplary embodiments shown below, and various modifications are possible within the scope of the claims. For example, embodiments obtained by appropriately combining some or all of the technologies (things or methods) employed in each of the exemplary embodiments shown below may also be included in the scope of the present invention. Furthermore, embodiments obtained by appropriately omitting some of the technologies employed in each of the exemplary embodiments shown below may also be included in the scope of the present invention. In addition, the effects mentioned in each of the exemplary embodiments shown below are examples of effects that can be expected in that exemplary embodiment and do not define the scope of the present invention. That is, embodiments that do not produce the effects mentioned in each of the exemplary embodiments shown below may also be included in the scope of the present invention.
[0012] [First Embodiment] A first exemplary embodiment, which is an example of an embodiment of the present invention, will be described in detail with reference to the drawings. This exemplary embodiment is the basic form for each of the exemplary embodiments described later. The scope of application of each technology adopted in this exemplary embodiment is not limited to this exemplary embodiment. That is, each technology adopted in this exemplary embodiment can also be adopted in other exemplary embodiments included in this disclosure, to the extent that no particular technical problems occur. Furthermore, each technology shown in the drawings referenced to explain this exemplary embodiment can also be adopted in other exemplary embodiments included in this disclosure, to the extent that no particular technical problems occur.
[0013] (Configuration of Information Processing Device 1) The configuration of the information processing device 1 according to this exemplary embodiment will be described with reference to Figure 1. Figure 1 is a block diagram showing the configuration of the information processing device 1. As shown in Figure 1, the information processing device 1 includes a first generation unit 11, a second generation unit 12, a correspondence unit 13, a loss calculation unit 14, and a learning unit 15.
[0014] (First generation unit 11) The first generation unit 11 acquires a target image, extracts the object representations of each object included in the target image, and generates one or more first fixation regions based on the object representations. Here, the object representation is a numerical representation of each object included in the target image, and as an example, it is represented as a vector. Also, as an example, the object representation is derived by referring to a feature map obtained from the target image. The first fixation region is information generated by referring to the object representation, and as an example, it is represented as one or more first fixation region maps. Note that the first generation unit 11, as an example, · an encoder to which the target image is input, · an extraction model that extracts the object representations of each object by referring to the output of the encoder, and can be configured to include, but this does not limit the exemplary embodiment.
[0015] (Second generation unit 12) The second generation unit 12 generates one or more second fixation regions from the target image using a pre-trained model. Here, the second fixation region is information generated by referring to the target image, and as an example, it is represented as one or more second fixation region maps. Also, specific examples of the pre-trained model do not limit this exemplary embodiment, but various pre-trained models such as DINO, ViT (Vision Transformer), SAM (Segment Anything Model), etc. can be used.
[0016] (Association unit 13) The association unit 13 estimates the correspondence between each of the one or more first fixation regions generated by the first generation unit 11 and the one or more second fixation regions generated by the second generation unit 12. As an example, the association unit 13 may be configured to calculate the similarity between each of the one or more first fixation regions and each of the one or more second fixation regions, and estimate the correspondence between the first fixation region and the second fixation region by referring to the similarity. However, this does not limit the exemplary embodiment.
[0017] (Loss calculation section 14) The loss calculation unit 14 calculates the loss by referring to the first and second observation regions, which are associated with each other by the correspondence unit 13. Specific examples of the loss function referenced by the loss calculation unit 14 are not limited to this exemplary embodiment, but as an example, a loss function can be used that is configured such that the smaller the degree of difference between the first and second observation regions, the smaller the value of the loss.
[0018] (Learning Section 15) The learning unit 15 trains the first generation unit 11 by referring to the loss calculated by the loss calculation unit 14. For example, the learning unit 15 updates the parameters of the extraction model used by the first generation unit 11 to extract the object representation of each object included in the target image by referring to the loss calculated by the loss calculation unit 14.
[0019] (Effects of Information Processing Device 1) As described above, in the information processing device 1, A first generation unit 11 extracts the object representation of each object contained in the target image and generates one or more first gaze regions based on said object representation, A second generation unit 12 generates one or more second gaze regions from the target image using a pre-trained model, A correspondence unit 13 that estimates the correspondence between the first gaze region and the second gaze region, A loss calculation unit 14 calculates the loss by referring to a first and second gaze region that correspond to each other, A learning unit 15 that causes the first generation unit 11 to learn by referring to the loss, A configuration is adopted that includes the following. In this way, the information processing device 1 associates the first gaze region generated by the first generation unit 11 with the second gaze region generated by the second generation unit 12, calculates a loss by referring to these gaze regions, and uses the calculated loss to train the first generation unit 11.
[0020] Therefore, with the above configuration, the first generation unit 11 can be efficiently trained using knowledge from a pre-trained model. In other words, with the above configuration, a model for extracting objects can be efficiently trained.
[0021] (Information processing method S1 flow) Next, the flow of the information processing method S1 according to this exemplary embodiment will be explained with reference to Figure 2. Figure 2 is a flowchart showing the flow of the information processing method S1. As shown in Figure 2, the information processing method S1 includes a step (process) S11 for generating a first gaze region, a step (process) S12 for generating a second gaze region, a step (process) S13 for estimating the correspondence relationship, a step (process) S14 for calculating the loss, and a step (process) S15 for training the extraction model.
[0022] (Step S11) In step S11, the first generation unit 11 acquires a target image, extracts the object representation of each object contained in the target image using an extraction model, and generates one or more first gaze regions based on the object representations. A more detailed explanation of the first generation unit 11 has been given above, so it will be omitted here.
[0023] (Step S12) In step S12, the second generation unit 12 generates one or more second gaze regions from the target image using a pre-trained model. A more detailed explanation of the second generation unit 12 has been given above, so it will be omitted here.
[0024] (Step S13) In step S13, the correspondence unit 13 estimates the correspondence between each of the one or more first gaze regions generated by the first generation unit 11 and the one or more second gaze regions generated by the second generation unit 12. A more detailed explanation of the correspondence unit 13 has been given above, so it will be omitted here.
[0025] (Step S14) In step S14, the loss calculation unit 14 calculates the loss by referring to the first and second observation regions, which are associated with each other by the correspondence unit 13. A more detailed explanation of the loss calculation unit 14 has been given above, so it will be omitted here.
[0026] (Step S15) In step S15, the learning unit 15 trains the extraction model by referring to the loss. A more detailed explanation of the learning unit 15 has been given above, so it will be omitted here.
[0027] (Effects of Information Processing Method S1) As described above, in Information Processing Method S1, - The object representation of each object contained in the target image is extracted using an extraction model, and one or more first gaze regions are generated based on the object representation. Using a pre-trained model, one or more second gaze regions are generated from the target image. • Estimate the correspondence between the first gaze region and the second gaze region. • Calculate the loss by referring to the first and second gaze regions, which correspond to each other. • Train the extraction model by referring to the loss. This configuration is employed. According to the above configuration, the same effect as that of the information processing device 1 is achieved.
[0028] [Second Embodiment] A second exemplary embodiment, which is an example of an embodiment of the present invention, will be described in detail with reference to the drawings. Components having the same function as those described in the above-described exemplary embodiment are denoted by the same reference numerals, and their descriptions are omitted as appropriate. The scope of application of each technology adopted in this exemplary embodiment is not limited to this exemplary embodiment. That is, each technology adopted in this exemplary embodiment can also be adopted in other exemplary embodiments included in this disclosure, to the extent that no particular technical problems arise. Furthermore, each technology shown in the drawings referenced to describe this exemplary embodiment can also be adopted in other exemplary embodiments included in this disclosure, to the extent that no particular technical problems arise.
[0029] (Configuration of Information Processing System 100A) The configuration of the information processing system 100A according to this exemplary embodiment will be described with reference to Figure 3. Figure 3 is a block diagram showing the configuration of the information processing system 100A. As shown in Figure 3, the information processing system 100A includes an information processing device 1A and a monitoring device 60 and a group of cameras 70 connected to the information processing device 1A via a network N. Here, the specific configuration of the network N is not limited to this exemplary embodiment, but as an example, a wireless LAN (Local Area Network), a wired LAN, a WAN (Wide Area Network), a public telephone network, a mobile data communication network, or a combination of these networks can be used.
[0030] (Camera group 70) The camera group 70 comprises one or more cameras. For example, if the monitoring device 60 is configured as a traffic monitoring device, these cameras will capture images of roads, vehicles, etc. In another example, if the monitoring device 60 is configured as a security monitoring device inside a building, these cameras will capture images of people, vehicles, etc. inside or outside the building. The images captured by the camera group 70 are supplied to the information processing device 1A, for example.
[0031] (Monitoring device 60) In general terms, the monitoring device 60 acquires object representations and reconstructed images extracted by the information processing device 1A from images captured by the camera group 70, and performs monitoring processing on the target by referring to the acquired object representations and reconstructed images. As shown in Figure 3, the monitoring device 60 includes a control unit 61 and a communication unit 62. The communication unit 62 receives object representations and reconstructed images from the information processing device 1A. The control unit 61 then performs traffic monitoring processing by referring to the object representations and reconstructed images acquired from the information processing device 1A. As another example, the control unit 61 performs security monitoring processing within a building by referring to the object representations and reconstructed images acquired from the information processing device 1A.
[0032] In this exemplary embodiment, the monitoring device 60 is shown as a separate device from the information processing device 1A, but this does not limit this exemplary embodiment. The functions of the control unit 61 of the monitoring device 60 may also be provided by the control unit of the information processing device 1A.
[0033] (Configuration of Information Processing Device 1A) Next, the configuration of the information processing device 1A according to this exemplary embodiment will be described with reference to Figure 3. As shown in Figure 3, the information processing device 1A includes a control unit 10, a storage unit 20, a communication unit 30, and an input / output unit 40.
[0034] (Communications Section 30) The communication unit 30 communicates with devices outside the information processing device 1A. For example, the communication unit 30 communicates with the camera group 70. The communication unit 30 supplies images received from the camera group 70 to the control unit 10. The communication unit 30 also communicates with the monitoring device 60. The communication unit 30 supplies object representations and reconstructed images derived by the control unit 10 to the monitoring device 60.
[0035] (Input / output section 40) The input / output unit 40 is configured to include at least one of the following input / output devices: a keyboard, mouse, display, printer, touch panel, etc. Alternatively, the input / output unit 40 may be configured to have input / output devices such as a keyboard, mouse, display, printer, touch panel, etc. connected to it. In this configuration, the input / output unit 40 receives various types of information from the connected input device to the information processing device 1A. The input / output unit 40 also outputs various types of information to the connected output device under the control of the control unit 10. An interface such as USB (Universal Serial Bus) can be used as an example of the input / output unit 40.
[0036] (Storage unit 20) The storage unit 20 stores various data referenced by the control unit 10, as well as various data generated by the control unit 10. For example, the storage unit 20 stores: • Target image TI ·Object representation OR • First gaze region RI1 • Second gaze region RI2 Reconstructed image RE ·Loss LS • Parameter group PG • Base model BM The following is stored. The target image TI may be, for example, an image supplied from the camera group 70, or it may not be. The target image TI contains, • In the learning phase, the images (training data) used for the learning process by the learning unit 15 described later are • Images (inference data) referenced by the object representation extraction unit 112 in the inference phase to extract object representations. It includes at least one of the following.
[0037] The object representation OR is a representation extracted (generated) by the object representation extraction unit 112, described later, and is a numerical representation of each object included in the target image TI. For example, the object representation OR can be represented as a vector.
[0038] The first gaze region RI1 is information generated by the object representation extraction unit 112 (described later) by referring to the object representation OR, and is represented, for example, as one or more first gaze region maps. The second gaze region RI2 is information generated by the second generation unit 12 (described later) by referring to the target image TI and using the base model BM, and is represented, for example, as one or more second gaze region maps. Specific examples of the first gaze region RI1 and the second gaze region RI2 will be described later.
[0039] The reconstructed image RE is an image generated by the reconstruction unit 113, which will be described later, and is an image reconstructed by referring to the object representation of each object contained in the target image TI. The loss LS is calculated using a loss function by the loss calculation unit 14, which will be described later. Specific examples of the reconstructed image RE and loss LS will be described later.
[0040] The parameter group PG includes parameters that define various models used by the control unit 10. For example, the parameter group PG includes: • One or more parameters that define the encoder used by the object representation extraction unit 112 - An extraction model used by the object representation extraction unit 112, which defines one or more parameters that define the extraction model for extracting object representations from the output of the encoder. • One or more parameters that define the decoder used by the reconstruction unit 113 These include, among others. At least some of these parameters are subject to learning (update) processing by the learning unit 15, which will be described later.
[0041] The base model BM is a pre-trained model used by the second generation unit 12, described later, to generate one or more second gaze regions RI2 from the target image TI. Specific examples of the base model BM are not limited to this exemplary embodiment, but various pre-trained models such as DINO, ViT (Vision Transformer), and SAM (Segment Anything Model) can be used.
[0042] (Control Unit 10) As shown in Figure 3, the control unit 10 includes a first generation unit 11, a second generation unit 12, a correspondence unit 13, a loss calculation unit 14, a learning unit 15, and an output information generation unit 16.
[0043] (First generation unit 11) The first generation unit 11 acquires a target image TI, extracts object representations OR for each object contained in the target image TI, and generates one or more first gaze regions RI1 based on the object representations OR. As shown in Figure 3, the first generation unit 11 comprises an acquisition unit 111, an object representation extraction unit 112, and a reconstruction unit 113.
[0044] (Acquisition part 111) The acquisition unit 111 acquires the target data TI. The acquisition unit 111 may be configured to acquire the target image TI from the camera group 70 via the communication unit 30, or it may be configured to acquire the target image TI stored in the storage unit 20.
[0045] (Object expression extraction unit 112) The object representation extraction unit 112 extracts the object representation OR of each object included in the target image TI. As an example, the object representation extraction unit 112 extracts the object representation OR of each object included in the target image TI. • An encoder that receives a target image TI as input and generates a feature map of the target image TI, - An extraction model that extracts the object representation OR of each object included in the target image TI by referring to the output (feature map) of the encoder. The configuration may include the following. Here, a CNN (Convolutional Neural Network) may be used as the encoder. Also, an attention model, more specifically a SLOT attention model, may be used as the extraction model. However, these examples do not limit this exemplary embodiment. Furthermore, the object representation extraction unit 112 may be configured to generate one or more first gaze regions RI1 by referring to the extracted object representation OR. More specifically, the object representation extraction unit 112 may, • By referring to one or more object representations OR, one or more mask images (attention masks) are generated as one or more first attention region maps. The attention mask represents the one or more first gaze regions RI1. This configuration is also acceptable.
[0046] (Reconstruction part 113) The reconstruction unit 113 reconstructs each of the objects by referring to the object representations of each object extracted by the object representation extraction unit 112. The reconstruction unit 113 generates a reconstructed image RE including the reconstructed objects, as an example. The reconstruction unit 113 may also be configured as a decoder that receives the object representations of each object as input and outputs the reconstructed image RE. Alternatively, the reconstruction unit 113 may be configured to generate one or more first gaze regions RI1 by referring to the object representations OR. More specifically, the reconstruction unit 113, • A mask image (decoder mask) is generated as one or more first gaze region maps, referencing one or more object representations OR. The decoder mask represents the one or more first gaze regions RI1. This configuration is also acceptable.
[0047] (Second generation unit 12) The second generation unit 12 generates one or more second gaze regions RI2 from the target image TI using a pre-trained base model BM. For example, the second generation unit 12 may refer to the target image TI to generate one or more second gaze region maps, and the one or more second gaze regions RI2 may be represented by these second gaze region maps.
[0048] (Matching section 13) The correspondence unit 13 estimates the correspondence between each of the one or more first gaze regions RI1 generated by the first generation unit 11 and the one or more second gaze regions RI2 generated by the second generation unit 12. For example, the correspondence unit 13 may be configured to calculate the similarity between each of the one or more first gaze regions RI1 and each of the one or more second gaze regions RI2, and then estimate the correspondence between the first gaze region RI1 and the second gaze region RI2 by referring to the similarity.
[0049] More specifically, the correspondence unit 13 may calculate the cosine similarity between the K (K is a natural number) gaze region maps generated by the first generation unit 11 and the K' (K' is a natural number) gaze region maps generated by the second generation unit 12, and determine the correspondence between the first gaze region RI1 and the second gaze region RI2 such that the sum of the cosine similarities is maximized. Furthermore, the Hungarian algorithm may be used to derive the correspondence. However, this is not limited to this exemplary embodiment.
[0050] (Loss calculation section 14) The loss calculation unit 14 calculates the loss by referring to the first and second observation regions, which are associated with each other by the correspondence unit 13. As an example, the loss calculation unit 14 calculates the loss using a loss function configured such that the smaller the degree of difference between the first and second observation regions, the smaller the value of the loss.
[0051] (Learning Section 15) The learning unit 15 refers to the loss calculated by the loss calculation unit 14 and trains the first generation unit 11. For example, the learning unit 15: • Encoder used by the object representation extraction unit 112 • Extraction model used by the object representation extraction unit 112 • Decodes that constitute the reconstruction unit 113 The learning unit 15 updates at least one of these parameters by referring to the loss calculated by the loss calculation unit 14. For example, the learning unit 15 updates these parameters so that the loss calculated by the loss calculation unit 14 becomes smaller.
[0052] (Output information generation unit 16) The output information generation unit 16 generates output information including information generated (extracted) by the control unit 10, and outputs the generated output information to the outside via the communication unit 30 or the input / output unit 40. For example, the output information generation unit 16 may be configured to generate output information including object representations extracted by the object representation extraction unit 112, and to supply said output information to the monitoring device 60 via the communication unit 30. The output information may also include reconstructed images RE generated by the reconstruction unit 113, or information generated by other configurations of the control unit 10.
[0053] (Specific configuration example and processing example 1) Next, with reference to Figures 4 and 5, a specific configuration example of the information processing device 1A and processing example 1 will be described. Figure 4 is a diagram showing a specific configuration example of the information processing device 1A and processing example 1 during the learning phase. Figure 5 is a schematic diagram showing an example of the processing flow by the information processing device 1A during the learning phase.
[0054] As shown in Figure 4, first, the target image TI is acquired by the acquisition unit 111 of the first generation unit 11. The acquired target image TI is input to the encoder 1121 of the object representation extraction unit 112. Then, the object representation extraction model (extraction model) 1122 refers to the output of the encoder 1121 and extracts one or more object representations OR. The object representation extraction unit 112 also generates K (K is a natural number) first attention region maps (attention masks in this example) based on the one or more object representations OR, and supplies the generated first attention region maps to the mapping unit 13.
[0055] Furthermore, the one or more object representations OR extracted by the extraction model 1122 are supplied to the decoder (reconstruction unit) 113, and the decoder 113 generates (reconstructs) a reconstructed image RE from the object representations OR.
[0056] On the other hand, the target image TI is also supplied to the second generation unit 12, which uses the base model BM to generate K' (K' is a natural number) second gaze region maps, and supplies the generated second gaze region maps to the correspondence unit 13.
[0057] Step S11 in Figure 5 schematically shows the first gaze regions RI11 and RI12 indicated by the first gaze region map generated by the first generation unit 11. Step S12 in Figure 5 also schematically shows the second gaze regions RI21 and RI22 indicated by the second gaze region map generated by the second generation unit 12.
[0058] Then, the correspondence section 13 shown in Figure 4, • The first gaze regions RI11 and RI12 generated by the first generation unit 11, • The second gaze regions RI21 and RI22 generated by the second generation unit 12 The correspondence between the elements is estimated. As an example, the correspondence unit 13 estimates the correspondence using the cosine similarity and the Hungarian algorithm described above.
[0059] In step S13 of Figure 5, the correspondence unit 13 performs, as an example, The second gaze region RI21 corresponds to the first gaze region RI11. The second gaze region RI22 corresponds to the first gaze region RI12. It has been indicated that this has been identified.
[0060] Then, the gaze region loss calculation unit 141 shown in Figure 4 calculates the loss by referring to the first gaze region and the second gaze region which are associated with each other by the correspondence unit 13. As an example, the loss calculation unit 14, as shown in step S14 of Figure 5, Differences between the first gaze region RI11 and the second gaze region RI21 Differences between the first gaze region RI12 and the second gaze region RI22 The loss (gaze region loss) is calculated using a loss function configured such that the smaller the value of the function, the smaller the loss value. Alternatively, the L2 norm may be used as the loss function. The gaze region loss calculated by the gaze region loss calculation unit 141 is supplied to the parameter update unit (learning unit) 15.
[0061] On the other hand, the reconstruction loss calculation unit 142 shown in Figure 4 calculates the loss based on the difference between the reconstructed image RE generated by the decoder 113 and the target image TI. As an example, the loss (reconstruction loss) is calculated using a loss function configured such that the smaller the difference between the reconstructed image RE and the target image TI, the smaller the loss value.
[0062] Then, the parameter update unit (learning unit) 15, as shown in step S15 of Figure 5, • The gaze area loss calculated by the gaze area loss calculation unit 141 and • Reconstruction loss calculated by reconstruction loss calculation unit 142 and The parameters of each model used by the first generator 11 are updated so that the linear sum of the values becomes smaller.
[0063] Thus, the information processing device 1A is A first generation unit 11 extracts the object representation of each object contained in the target image TI and generates one or more first gaze regions (K gaze region maps) based on said object representation, A second generation unit 12 generates one or more second gaze regions (K' gaze region maps) from the target image TI using a pre-trained model (base model BM), A correspondence unit 13 that estimates the correspondence between the first gaze region and the second gaze region, A loss calculation unit 14 calculates the loss by referring to a first and second gaze region that correspond to each other, A learning unit 15 that causes the first generation unit 11 to learn by referring to the loss, The information processing device 1A is equipped with the following features. In this way, the first gaze region generated by the first generation unit 11 and the second gaze region generated by the second generation unit 12 are associated with each other, and the loss is calculated by referring to these gaze regions, and the first generation unit 11 is trained by referring to the calculated loss.
[0064] Therefore, with the above configuration, the convergence of learning can be accelerated and the first generation unit 11 can be efficiently trained by using the knowledge from the pre-trained base model BM, or in other words, by transferring the knowledge from the base model BM. For example, with the above configuration, an extraction model for extracting objects can be efficiently trained. More specifically, for example, a world model that realizes amodal segmentation (hidden region estimation) and future prediction can be efficiently constructed. Furthermore, by using this example in the monitoring process by the monitoring device 60, it becomes possible to detect and track vehicles and people without manual correction simply by installing surveillance cameras (camera group 70) and collecting images.
[0065] (Specific configuration example and processing example 2) Next, with reference to Figure 6, a specific configuration example of the information processing device 1A and processing example 2 will be described. Figure 6 is a diagram showing a specific configuration example of the information processing device 1A and processing example 2 during the learning phase.
[0066] As shown in Figure 6, in this example, the decoder (reconstruction unit) 113 generates K (K is a natural number) first gaze region maps (decoder masks in this example), and supplies the generated first gaze region maps to the mapping unit 13. On the other hand, in this example, the object representation extraction unit 112 does not supply the first gaze region maps. Other processing is the same as in the configuration example and processing example 1 described above, so redundant explanations are omitted. With the configuration of this example as well, the first generation unit 11 can be efficiently trained using the knowledge from the pre-trained base model BM.
[0067] (Specific configuration example and processing example 3) Next, with reference to Figure 7, a specific configuration example of the information processing device 1A and processing example 3 will be described. Figure 7 is a diagram showing a specific configuration example of the information processing device 1A and processing example 3 during the learning phase.
[0068] As shown in Figure 7, in this example, the second generation unit 12 comprises a base model BM and a gaze region estimation unit 122. Here, the gaze region estimation unit 122 estimates one or more second gaze regions by referring to the output of the base model BM into which the target image TI is input. As an example, the gaze region estimation unit 122 generates K' (K' is a natural number) second gaze region maps by referring to the output from the intermediate layer or output layer of the base model BM into which the target image TI is input, and supplies the generated second gaze region maps to the correspondence unit 13.
[0069] More specifically, the gaze region estimation unit 122 may be configured to estimate K' gaze region maps by dividing the feature map (h × w × d dimensions) obtained from the base model BM into a predetermined number of clusters K' using the K-means method or the like. Alternatively, the gaze region estimation unit 122 may determine the number of clusters K' based on the number of objects K determined in the object representation extraction unit 112. For example, it may be set so that K' = K.
[0070] Other processes are the same as those described in the configuration example and processing example 1, so redundant explanations will be omitted. With this example configuration, the first generation unit 11 can be efficiently trained using the knowledge from the pre-trained base model BM.
[0071] (Specific Configuration Example and Processing Example 4) Next, with reference to Figure 8, a specific configuration example of the information processing device 1A and processing example 4 will be described. Figure 8 is a diagram showing a specific configuration example of the information processing device 1A and processing example 4 during the learning phase.
[0072] As shown in Figure 8, in this example, the second generation unit 12 comprises a base model BM and a prompt generation unit 123. Here, the prompt generation unit 123 generates prompts for input to the base model BM. For example, these prompts are prompts associated with a target image. These prompts may be predetermined prompts, or they may be generated in response to the target image TI or user input. To give a more specific example, the prompts may be: Grid point prompt Language prompt ? Rectangular prompt Mask prompt Any of the various prompts may be used. The base model BM in this example generates one or more second gaze regions by referring to the target image TI and the prompts associated with the target image TI, as an example. Other processing is the same as in the configuration example and processing example 1 described above, so redundant explanations are omitted. With the configuration of this example, the first generation unit 11 can be efficiently trained using the knowledge from the pre-trained base model BM. In addition, in this example, the gaze region can be suitably limited by using prompts appropriate to the purpose.
[0073] (Specific configuration example and processing example 5) Next, with reference to Figure 9, a specific configuration example and processing example 5 of the information processing device 1A will be described. This example is a specific configuration example and processing example of the information processing device 1A in the inference phase. Figure 9 is a diagram showing a specific configuration example and processing example of the information processing device 1A related to this example.
[0074] As shown in Figure 9, in the inference phase, the target image TI to be inferred is first acquired by the acquisition unit 111. This target image TI is, for example, captured by the camera group 70 shown in Figure 3. However, this is not limited to this example.
[0075] The target image TI acquired by the acquisition unit 111 is supplied to the encoder 1121. Then, the object representation extraction model (extraction model) 1122, which has been learned by any of the processing examples 1 to 4 described above, refers to the output of the encoder 1121 and extracts one or more object representations OR. The extracted object representations are supplied to the output information generation unit 16 as an example and included in the output information. The output information including the object representations is supplied to the monitoring device 60 as an example and referred to in the monitoring process. Note that the information output in the inference phase is not limited to the above example, and the output image may also include the reconstructed image RE generated by the decoder 113 by referring to the object representations.
[0076] With the above configuration, object representations can be suitably extracted from the target image TI using an extraction model that has been efficiently trained using knowledge from a pre-trained base model BM.
[0077] (Additional information for exemplary embodiment 2) This exemplary embodiment is not limited to the examples described above. For example, a pre-trained model can be used as at least one of the encoder 1121, the object representation extraction model (extraction model) 1122, and the decoder (reconstruction unit) 113. For example, a model pre-trained with a large amount of training data can be used as at least one of the encoder 1121, the object representation extraction model 1122, and the decoder 113, and then the learning process can be advanced by transferring the knowledge of the base model BM as described in the processing example above. In this way, by starting the learning process with a pre-trained model that can roughly estimate the object region, the correspondence of the gaze region can be accurately identified, and learning can be advanced more efficiently.
[0078] [Third Embodiment] A third exemplary embodiment, which is an example of an embodiment of the present invention, will be described in detail with reference to the drawings. Components having the same function as those described in the above-described exemplary embodiments are denoted by the same reference numerals, and their descriptions are omitted as appropriate. The scope of application of each technology adopted in this exemplary embodiment is not limited to this exemplary embodiment. That is, each technology adopted in this exemplary embodiment can also be adopted in other exemplary embodiments included in this disclosure, to the extent that no particular technical problems arise. Furthermore, each technology shown in the drawings referenced to describe this exemplary embodiment can also be adopted in other exemplary embodiments included in this disclosure, to the extent that no particular technical problems arise.
[0079] (Configuration of Information Processing System 100B) The configuration of the information processing system 100B according to this exemplary embodiment will be described with reference to Figure 10. Figure 10 is a block diagram showing the configuration of the information processing system 100B. As shown in Figure 10, the information processing system 100B comprises an information processing device 1A and a robot 80 connected to the information processing device 1A via a network N. The information processing device 1A and the network N are the same as in the exemplary embodiment 1, so redundant explanations will be omitted.
[0080] (Robot 80) As shown in Figure 10, the robot 80 comprises a control unit 81, a communication unit 82, a camera group 83, and a drive unit 84. The camera group 83 comprises one or more cameras that capture images of the area in which the robot 80 operates. The drive unit 84 comprises a robot arm, wheels, etc., which are driven based on control by the control unit 81.
[0081] The communication unit 82 supplies images captured by the camera group 83 to the information processing device 1A via the network N. The communication unit 82 also receives object representations extracted from the images from the information processing device 1A via the network N and supplies them to the control unit 81.
[0082] The control unit 81 supplies images of one or more objects to the information processing device 1A via the communication unit 82, and controls the drive unit 84 by referring to the object representation extracted by the information processing device 1A from the images. For example, the control unit 81 refers to the object representation acquired via the communication unit 82 to perform hidden area estimation of the world model, future prediction processing, etc., and learns and plans the control of the robot arm.
[0083] (Information Processing Device 1A) As partially described above, in this exemplary embodiment, images captured by the camera group 83 are supplied to the information processing device 1A and are referenced as target images TI in at least one of the learning phase and inference phase described in exemplary embodiment 2.
[0084] Then, in the inference phase, the information processing device 1A extracts object representations from the target image TI by performing inference processing using the image captured by the camera group 70 as the target image TI. The extracted object representations are then supplied to the robot 80 and referenced to control the drive unit 84.
[0085] [Examples of implementation using software] Some or all of the functions of the information processing devices 1,1A (hereinafter also referred to as "each of the above devices") may be implemented by hardware such as integrated circuits (IC chips) or by software.
[0086] In the latter case, each of the above devices is implemented, for example, by a computer that executes instructions for a program, which is software that realizes each function. An example of such a computer (hereinafter referred to as computer C) is shown in Figure 11. Figure 11 is a block diagram showing the hardware configuration of computer C, which functions as each of the above devices.
[0087] Computer C comprises at least one processor C1 and at least one memory C2. Memory C2 stores a program P that causes computer C to operate as each of the above-mentioned devices. In computer C, processor C1 reads program P from memory C2 and executes it, thereby realizing each of the above-mentioned devices.
[0088] For processor C1, for example, a CPU (Central Processing Unit), GPU (Graphic Processing Unit), DSP (Digital Signal Processor), MPU (Micro Processing Unit), FPU (Floating Point Number Processing Unit), PPU (Physics Processing Unit), TPU (Tensor Processing Unit), quantum processor, microcontroller, or a combination thereof can be used. For memory C2, for example, flash memory, HDD (Hard Disk Drive), SSD (Solid State Drive), or a combination thereof can be used.
[0089] Computer C may also be equipped with RAM (Random Access Memory) for loading program P at runtime and for temporarily storing various data. Furthermore, computer C may be equipped with communication interfaces for sending and receiving data with other devices. Additionally, computer C may be equipped with input / output interfaces for connecting input / output devices such as keyboards, mice, displays, and printers.
[0090] Furthermore, program P can be recorded on a non-temporary, tangible recording medium M that is readable by computer C. Such a recording medium M could be, for example, tape, disk, card, semiconductor memory, or programmable logic circuitry. Computer C can acquire program P via such a recording medium M. Program P can also be transmitted via a transmission medium. Such a transmission medium could be, for example, a communication network or broadcast waves. Computer C can also acquire program P via such a transmission medium.
[0091] Furthermore, each of the above functions of each of the above devices may be implemented by a single processor in a single computer, by multiple processors in a single computer working together, or by multiple processors in each of multiple computers working together. In addition, the programs for implementing each of the above functions in each of the above devices may be stored in a single memory in a single computer, distributed and stored in multiple memories in a single computer, or distributed and stored in multiple memories in each of multiple computers.
[0092] [Additional Note A] This disclosure includes the technologies described in the following appendices. However, the present invention is not limited to the technologies described in the following appendices, and various modifications are possible within the scope of the claims.
[0093] (Note A1) A first generation means that extracts object representations of each object contained in the target image and generates one or more first gaze regions based on said object representations, A second generation means for generating one or more second gaze regions from the target image using a pre-trained model, Correspondence means for estimating the correspondence between the first gaze region and the second gaze region, Loss calculation means that calculates the loss by referring to a first and second gaze region that correspond to each other, A learning means that causes the first generation means to learn by referring to the loss An information processing device equipped with the following features.
[0094] (Appendix A2) The first generation means is, An object representation extraction means for extracting object representations of each object by referring to the aforementioned target image, Reconstruction means for reconstructing each of the aforementioned objects by referring to the object representation of each of the aforementioned objects, It is equipped with, The learning means updates at least the parameters of the object representation extraction means by referring to the loss. The information processing device described in Appendix A1.
[0095] (Note A3) The object representation extraction means generates the first gaze region The information processing device described in Appendix A2.
[0096] (Note A4) The reconstruction means generates the first gaze region. The information processing device described in Appendix A2.
[0097] (Note A5) The object representation extraction means is An encoder into which the aforementioned target image is input, An extraction model that extracts object representations of each object by referring to the output of the encoder, It is equipped with The information processing device described in Appendix A2.
[0098] (Note A6) The aforementioned correspondence means is, The similarity between each of the one or more first gaze regions and each of the one or more second gaze regions is calculated, and the correspondence between the first gaze region and the second gaze region is estimated by referring to the similarity. An information processing device as described in any one of the appendices A1 to A5.
[0099] (Note A7) The loss calculated by the loss calculation means is The value of the loss is set to decrease as the degree of difference between the first gaze region and the second gaze region decreases. An information processing device as described in any one of the appendices A1 to A5.
[0100] (Note A8) The second generation means is The system includes a gaze region estimation means that estimates one or more second gaze regions by referring to the output of the pre-trained model into which the target image has been input. An information processing device as described in any one of the appendices A1 to A5.
[0101] (Note A9) One or more processors, The object representation of each object contained in the target image is extracted using an extraction model, and one or more first gaze regions are generated based on these object representations. Using a pre-trained model, one or more second gaze regions are generated from the target image. To estimate the correspondence between the first gaze region and the second gaze region, The loss is calculated by referring to the first and second gaze regions, which correspond to each other. The extraction model is trained by referring to the loss. An information processing method that includes this.
[0102] (Note A10) A program that makes a computer function as an information processing device. The aforementioned computer, A first generation means that extracts object representations of each object contained in the target image and generates one or more first gaze regions based on said object representations, A second generation means for generating one or more second gaze regions from the target image using a pre-trained model, Correspondence means for estimating the correspondence between the first gaze region and the second gaze region, Loss calculation means that calculates the loss by referring to a first and second gaze region that correspond to each other, A learning means that causes the first generation means to learn by referring to the loss A program that makes it function as such.
[0103] (Note A11) The second generation means is The one or more second gaze regions are generated by referring to the target image and the prompt associated with the target image. An information processing device as described in any one of the appendices A1 to A5. [Explanation of Symbols]
[0104] 1. 1A Information Processing Device 100A, 100B Information Processing System 11. First generation unit (first generation means) 12. Second generation unit (second generation means) 13. Correspondence section (correspondence means) 14 Loss calculation section (loss calculation means) 15. Learning Section (Learning Methods)
Claims
1. A first generation means that extracts object representations of each object included in the target image and generates one or more first gaze regions based on said object representations, A second generation means for generating one or more second gaze regions from the target image using a pre-trained model, Correspondence means for estimating the correspondence between the first gaze region and the second gaze region, Loss calculation means that calculates the loss by referring to a first and second gaze region that correspond to each other, A learning means for causing the first generation means to learn by referring to the loss An information processing device equipped with the following features.
2. The first generating means is An object representation extraction means for extracting object representations of each object by referring to the aforementioned target image, Reconstruction means for reconstructing each of the aforementioned objects by referring to the object representation of each of the aforementioned objects, It is equipped with, The learning means updates at least the parameters of the object representation extraction means by referring to the loss. The information processing apparatus according to claim 1.
3. The object representation extraction means generates the first gaze region. The information processing apparatus according to claim 2.
4. The reconstruction means generates the first gaze region. The information processing apparatus according to claim 2.
5. The object representation extraction means is An encoder into which the aforementioned target image is input, An extraction model that extracts object representations of each object by referring to the output of the encoder, It is equipped with The information processing apparatus according to claim 2.
6. The aforementioned correspondence means is, The similarity between each of the one or more first gaze regions and each of the one or more second gaze regions is calculated, and the correspondence between the first gaze region and the second gaze region is estimated by referring to the similarity. The information processing apparatus according to any one of claims 1 to 5.
7. The loss calculated by the loss calculation means is The value of the loss is set to decrease as the degree of difference between the first gaze region and the second gaze region decreases. The information processing apparatus according to any one of claims 1 to 5.
8. The second generation means is, The system includes a gaze region estimation means that estimates one or more second gaze regions by referring to the output of the pre-trained model into which the target image has been input. The information processing apparatus according to any one of claims 1 to 5.
9. One or more processors, The object representation of each object contained in the target image is extracted using an extraction model, and one or more first gaze regions are generated based on said object representation. Using a pre-trained model, one or more second gaze regions are generated from the target image. To estimate the correspondence between the first gaze region and the second gaze region, The loss is calculated by referring to the first and second gaze regions, which correspond to each other. The extraction model is trained by referring to the loss. An information processing method that includes this.
10. A program that makes a computer function as an information processing device. The aforementioned computer, A first generation means that extracts object representations of each object included in the target image and generates one or more first gaze regions based on said object representations, A second generation means for generating one or more second gaze regions from the target image using a pre-trained model, Correspondence means for estimating the correspondence between the first gaze region and the second gaze region, Loss calculation means that calculates the loss by referring to a first and second gaze region that correspond to each other, A learning means for causing the first generation means to learn by referring to the loss A program that makes it function as such.
Citation Information
Patent Citations
Learning apparatus, detecting device, learning system, learning method, learning program, detecting method, and detecting program
JP2023104705A