Information processing device, control method for information processing device, and program

JP7927792B2Active Publication Date: 2026-10-01CANON KK
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
JP2024101468
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-06-24
Publication Date
2026-10-01
Estimated Expiration
2044-06-24

AI Technical Summary

Benefits of technology

【0009】 本発明によれば、遮蔽関係にある物体が画像中に存在するような状況下においても、画像中の各領域がいずれの被写体に属するかをより好適な態様で推定することが可能となる。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007927792000010
    Figure 0007927792000010
  • Figure 0007927792000011
    Figure 0007927792000011
  • Figure 0007927792000012
    Figure 0007927792000012
Patent Text Reader

Abstract

To estimate to which subject each area in an image belongs in a more suitable mode even in a situation where objects in a shielding relationship exist in the image.SOLUTION: The attribute estimation unit 113 estimates to which of a plurality of subjects a first region corresponding to at least a part of a subject in an image in which the plurality of subjects are captured belongs. The machine learning model applied to the attribute estimation unit 113 is trained using a loss function that calculates a loss value based on at least a second region in which two or more subjects among the plurality of subjects overlap.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to an information processing apparatus, a control method for an information processing apparatus, and a program.

Background Art

[0002] Techniques have been proposed for causing a model to execute various recognition tasks by learning the model based on so-called machine learning that uses images as learning data with a machine such as a computer. Examples of the recognition tasks include a detection task of detecting human body parts (head, face, upper body, whole body, etc.) from an image, and a tracking task of searching for and tracking a specific subject in an image. If it becomes possible to specify a region of an object in an image by applying such a technique, for example, in an imaging device, it becomes possible to focus on the region and adjust the exposure of the region to a more suitable state, so improvement in operability can be expected. Further, if an object region can be specified by a recognition task, the technique can be utilized not only for imaging devices but also for various other applications.

[0003] A neural network (NN: Neural Network) is known as a model for executing the above-described recognition tasks. Among NNs, there is Deep NN (DNN: Deep Neural Network) adapted to deeper layers. In particular, Deep Convolutional NN (DCNN: Deep Convolutional Neural Network), which is configured by repeating multiple layers of convolutional layers and pooling layers, has attracted attention for its high performance (detection accuracy, detection performance). In recent years, a technique called VIT (Vision Transformer), which introduces an attention mechanism into image recognition, has also attracted attention.

[0004] When using the autofocus (AF) function of an imaging device, situations may arise where the focus needs to be on a localized area of ​​the subject, such as a human head or upper body, an animal's eye, or the nose of an airplane. Therefore, in recognition tasks applied to AF, a localized area is detected for each of a series of potential candidate objects captured in the image. If multiple objects are captured in the image, a determination is made as to which object the detected localized area belongs. As a method to enable such determination, a method has been proposed that uses DCNN to estimate which object a localized area belongs to using a score map. Patent Document 1 discloses a technique that outputs a map showing the joint positions of a person, along with information indicating which person in the image each detected joint belongs to. [Prior art documents] [Patent Documents]

[0005] [Patent Document 1] Patent No. 7422456 [Patent Document 2] Japanese Patent Publication No. 2022-19339 [Overview of the Initiative] [Problems that the invention aims to solve]

[0006] On the other hand, in the technology disclosed in Patent Document 1, when objects are in an occluding relationship with each other, the proportion of the occluded object that appears in the image is smaller than that of the occluding object, making it difficult to separate it into a score specific to the person. To address this problem, Patent Document 2 discloses a technology that performs object detection considering the occluding relationship when the main subject is in an occluding relationship with other subjects that have similar appearance characteristics and postures. However, the technology disclosed in Patent Document 2 uses DCNN to perform learning that optimizes the score for the primary subject which is in an occlusion relationship. This makes it possible to identify local regions belonging to the primary subject, but it is not configured to optimize the score for other subjects besides the primary subject. Therefore, the technology disclosed in Patent Document 2 is not suitable for matching local regions belonging to objects that are in an occlusion relationship with other objects and have a small area of ​​existence.

[0007] In view of the above problems, the present invention aims to enable the estimation of which subject each region in an image belongs to, in a more suitable manner, even when objects in an occluding relationship are present in the image. [Means for solving the problem]

[0008] The information processing device according to the present invention has estimation means for estimating which of the multiple subjects a first region corresponding to at least a portion of the subjects in an image in which multiple subjects are captured belongs, and the estimation means determines a second region in which two or more of the multiple subjects overlap. The loss value is weighted according to its size or its proportion to a predetermined area. Learning using a loss function Use a pre-trained model It is characterized by the following: [Effects of the Invention]

[0009] According to the present invention, even in situations where objects in an occluding relationship exist in the image, it becomes possible to estimate, in a more suitable manner, which subject each region in the image belongs to. [Brief explanation of the drawing]

[0010] [Figure 1] This diagram shows an example of the configuration of an imaging system. [Figure 2] This flowchart shows an example of the process involved in training a machine learning model. [Figure 3] This diagram shows an example of a method for generating attribute-based correct answer maps. [Figure 4] This is a diagram showing an example of an attribute score map. [Figure 5] This diagram shows an example of a method for generating attribute-based correct answer maps. [Figure 6] This flowchart shows an example of the process involved in training a machine learning model. [Figure 7] This figure shows an example of a method for generating attribute ground truth maps for foreground and background objects. [Figure 8] This flowchart shows an example of the process involved in training a machine learning model. [Figure 9] This figure shows an example of a method for generating a ground truth map of the background region's attributes. [Modes for carrying out the invention]

[0011] The embodiments will be described in detail below with reference to the attached drawings. Note that the following embodiments do not limit the invention as defined in the claims. While the embodiments describe multiple features, not all of these features are essential to the invention, and the features may be combined in any way. Furthermore, in the attached drawings, identical or similar configurations are given the same reference numerals, and redundant descriptions are omitted. Furthermore, in order to make the explanation of each embodiment easier to understand, this disclosure will focus on imaging systems that use an imaging device such as a digital camera to detect the head and face (hereinafter also referred to as local areas), upper body and whole body (hereinafter also referred to as whole body areas) of a person as the subject. However, the subject is not necessarily limited to a person. For example, the nose or body of a vehicle may be used as the local area or whole area, or the head or tail of an animal may be used as the local area or whole area, and the technology related to this disclosure may be applied to these subjects. Furthermore, in this disclosure, when there is no particular distinction between still images and moving images, they may simply be referred to as "images." Therefore, when the term "image" is used, it may apply to both still images and moving images unless otherwise specified.

[0012] <First Embodiment> As a first embodiment of the present disclosure, an example of a method for training a machine learning model that detects regions of parts such as the head and upper body of a person in an input image by performing training based on so-called machine learning will be described. Specifically, in a situation where there is an overlapping region between a plurality of persons in an image, learning is performed by calculating a loss value suitable for occlusion conditions such that the variance of scores associated within the region of the same person including a local region and a whole-body region is reduced. A method for performing such learning will be described.

[0013] With reference to FIG. 1(a), an example of the hardware configuration of the information processing apparatus 110 according to the present embodiment will be described. A CPU (Central Processing Unit) 101 controls the overall operation of the information processing apparatus 110 by expanding and executing a control program stored in a ROM (Read Only Memory) 103. A RAM (Random Access Memory) 102 is used as a temporary storage area such as a main memory and a work area for the CPU 101. The RAM 102 is also used as an area for expanding the control program so that the control program can be executed by the CPU 101. An input unit 105 corresponds to an input interface that receives input from a user, and is implemented by, for example, an input device such as a keyboard or a touch panel, and can also receive image input and the like. A display unit 106 corresponds to an output interface that presents various types of information to a user, is implemented by, for example, an output device such as a liquid crystal display, and can display various types of data and results of various processes. Furthermore, the information processing apparatus 110 can communicate with other apparatuses by being connected to a predetermined network via a communication unit 104. For example, the information processing apparatus 110 may receive image input from another apparatus, acquire a machine learning model, receive an instruction from a user, or the like via the communication unit 104, and may output various processing results to another apparatus. The storage unit 107 is a storage area for storing various data, and may be used, for example, as a storage area for storing a machine learning model according to the present embodiment, the details of which will be described later. The storage unit 107 can be implemented by storage devices such as, for example, an HDD (Hard Disk Drive), flash memory, and various optical media.

[0014] Referring to FIG. 1(b), an example of the functional configuration of the information processing apparatus 110 according to the present embodiment will be described, with particular attention paid to the configuration of a portion that specifies a region belonging to a predetermined object (e.g., a person) in a target image using a trained machine learning model. An image acquisition unit 111 receives input of image data of moving images or still images in accordance with instructions from a user. A position estimation unit 12 uses a machine learning model (hereinafter also referred to as a region detector) trained to detect regions belonging to a predetermined object in an image to detect local regions and whole-body regions from the target image, and estimates the center position and size of the object to be detected. Regarding the region detector to be used, a machine learning model individually trained with a prepared dataset may be applied, or a machine learning model based on a known object detection technique (e.g., YoLo, ViT, etc.) may be applied. An attribute estimation unit 113 estimates a score map for specifying which person each of the local regions and whole-body regions detected by the position estimation unit 112 using the region detector belongs to. An association processing unit 114 specifies local regions and whole-body regions belonging to each person in the target image based on the score map acquired by the attribute estimation unit 113. Note that a known process (for example, the method disclosed in Patent Document 1) may be applied to the processing related to specifying each region in the image. Note that a region belonging to a predetermined object in an image (in other words, at least a partial region of the predetermined object in the image), such as a local region or a whole-body region detected by the position estimation unit 112, corresponds to an example of the first region.

[0015] Referring to Figure 1(c), an example of the functional configuration of the learning device 121 related to the learning of the machine learning model used in the information processing device 110 described above will be explained. The learning device 121 targets the machine learning model applied to the information processing device 110, takes the learning data as input and performs estimation processing, and uses the results to cause the learning unit 122 to learn the machine learning model as shown below.

[0016] The following describes an example of the processing involved in training a machine learning model by the learning unit 122, referring to Figures 2 to 4. This method calculates a loss value according to the size of the overlapping region when two or more people captured in an image overlap.

[0017] First, with reference to Figure 2, an example of the processing related to the learning of a machine learning model applied to the attribute estimation unit 113 according to this embodiment will be described.

[0018] In S200, the learning unit 122 receives the image acquired by the image acquisition unit 111 (hereinafter also referred to as the input image) as the object to be processed. The acquired image may be either an image in which a person is the subject, or an image in which something other than a person is the subject, such as scenery like the sky or mountains, or animals or vehicles. In S201, the learning unit 122 acquires information about the region of a person in the image, which will be used in the map creation process described in S202. Here, the information about the region of a person in the image is assumed to be a mask region in which the pixel value of the region where the person exists, including the whole body region, is "1", and the pixel value of the other regions is "0". Furthermore, the local region and the whole body region are defined as rectangles in which the center coordinates, width, and height are stored as parameters.

[0019] In S202, the learning unit 122 generates a score map (hereinafter also referred to as the ground truth map) from the information acquired in S201, which will be used as the ground truth for the estimation results by the machine learning model when training the machine learning model. Here, with reference to Figure 3, an example of a method for generating a ground truth map will be described. For example, Figure 3(a) shows an example of an image 300 to which a ground truth map is to be generated. Specifically, image 300 shows an example of an image in which two people overlap in the image (a scene in which part of one person is obscured by the other person). Here, an example of a method for generating the ground truth map 307 shown in Figure 3(c) and the overlapping region map 308 shown in Figure 3(d) will be described using the image 300 shown in Figure 3(a) as the input image.

[0020] (Process 1) The learning unit 122 acquires the whole body regions 303 and 305 and the mask regions 302 and 304 shown in Figure 3(a) as human region information. (Process 2) The learning unit 122 calculates the region where the acquired whole-body regions 303 and 305 overlap with each other, and defines this overlapping region as the overlapping region 306 shown in Figure 3(b). (Process 3) The learning unit 122 generates a ground truth map 307, as illustrated in Figure 3(c), for each of the acquired mask regions 302 and 304, such that different map values ​​(pixel values) are assigned to each mask region. In the example shown in Figure 3(c), the ground truth map 307 assigns map values ​​to each region, with the background portion 301 having a value of "0", the mask region 302 of the person in the foreground having a value of "1", and the mask region 304 of the person in the background having a value of "2". The learning unit 122 also generates an overlapping region map 308, as illustrated in Figure 3(d), based on the overlapping region 306. In the example shown in Figure 3(d), the overlapping region map 308 assigns map values ​​to each region, with the overlapping region 306 having a value of "1" and the other regions having a value of "0". Furthermore, an area where two or more subjects overlap, such as overlapping area 306, is an example of a second area.

[0021] In S203, the learning unit 122 causes the machine learning model to be trained (the machine learning model applied to the attribute estimation unit 113) to estimate a score map (hereinafter also referred to as an attribute score map) for identifying which person each region in the image 300 belongs to. Figure 4 shows an example of an attribute score map 401 estimated by the machine learning model applied to the attribute estimation unit 113. In the example shown in Figure 4, the attribute score map 401 assigns a higher score to regions in the image 300 where people exist compared to other regions.

[0022] In S204, the learning unit 122 calculates the loss for the attribute score map 401 estimated in S203. Specifically, the learning unit 122 uses two loss functions to train the model: one to train the score assigned to a person region in the attribute score map 401 to be the same or close in value within the same person region, and another to train the score to be different between different person regions.

[0023] The loss function used to calculate the loss value (PullLoss) for learning to achieve the same or similar values ​​within the same person region is expressed by the following relation, shown as equation (1-1). In equation (1-1), N represents the number of people in the image, and the score est(i) represents the value of pixel position i in the estimated score map. Also, Area j This is the set of the j-th person (1≦j≦N; here, the number of people in the image N=2). The definitions of N and est(i) are common to all embodiments of this disclosure, including this embodiment. The learning unit 122 weights the PullLoss according to the ratio of each person's region to the total person region in the image. With this weighting, the smaller the person's region, the larger the value added to the PullLoss, and the larger the person's region, the smaller the value added to the PullLoss.

[0024]

number

[0025] The loss function used to determine the loss value (PushLoss) required to learn that different person domains have different values ​​is expressed by the following relational expression, Equation (1-2). In Equation (1-2), β is an empirically obtained hyperparameter. OverlapAreaValue is the size of the overlapping region 306, or the ratio of the overlapping region 306 to the map size of the overlapping region map 308. The learning unit 122 weights the loss calculated from the score value by the sum of the above parameters, namely β and OverlapAreaValue. With this weighting, the loss for attribute score maps with overlapping regions becomes larger than the loss for attribute score maps without overlapping regions.

[0026]

number

[0027] Furthermore, the loss function for calculating PushLoss may be expressed by the relationship shown below as equation (1-3). In equation (1-3), W o Like β, is an empirically obtained hyperparameter. The difference between equation (1-3) and equation (1-2) is that OverlapAreaValue is added as PushLoss to the loss calculated from the ground truth map, but the purpose of weighting is essentially the same in both relationships.

[0028]

number

[0029] The loss function applied in S204 is expressed by the relationship shown below as equation (1-4), obtained by summing equation (1-1) with equation (1-2) or equation (1-3). In equation (1-4), α is an empirically obtained hyperparameter.

[0030]

number

[0031] In S205, the learning unit 122 updates the interlayer coupling weight coefficients (parameters) of the target machine learning model based on the loss shown in equation (1-4). For parameter updates, for example, Momentum SGD is used and performed based on backpropagation.

[0032] In S206, the learning unit 122 determines whether or not predetermined conditions related to the completion of learning the machine learning model have been met. If the learning unit 122 determines in S206 that the predetermined conditions are not met, it proceeds to S200. In this case, the learning unit 122 acquires the next input image in S200 and executes the processing from S201 onwards again, targeting that input image. Furthermore, if the learning unit 122 determines in S206 that a predetermined condition has been met, it terminates the series of processes shown in Figure 2, namely the processes related to learning the machine learning model. Furthermore, the conditions for terminating the training of the machine learning model in S206 may include, for example, conditions specified in advance by the user. For instance, the condition for terminating training may be set as the machine learning model's estimation accuracy exceeding a threshold. Other examples include the training process being executed multiple times or exceeding a threshold, or the training process time exceeding a threshold. Furthermore, in this embodiment, for the sake of simplicity, images are used one by one for learning. However, this does not limit the processing of the learning device 121. For example, it is also possible to use multiple images as a mini-batch for learning. In this case, for example, the learning unit 122 can update the learning parameters based on the sum of the losses of the multiple images.

[0033] In this embodiment, we have described a method for training a model that performs attribute score map estimation, which uses a loss function that calculates a loss value according to the area of ​​overlapping regions of people. By using this method, it is possible to calculate the loss for people whose presence area is small due to overlapping regions, in the same way as for people who are occluded. As a result, for example, when applying a machine learning model to estimate which person a local region and a whole region belong to, it is expected that the estimation accuracy in intersection scenes where overlapping regions occur will be improved compared to when the model is trained using conventional methods.

[0034] (Modified version of the first embodiment) In the embodiment described above, the processing of S202 shown in Figure 2 was explained focusing on the case where the information obtained in the processing of S201 is "whole body region and mask region". In contrast, as a modification of the first embodiment, an example of processing when the information obtained in the processing of S201 is "local region and whole body region" will be described below, with particular attention to the processing of S202.

[0035] (Process 1) The learning unit 122 acquires local regions 500 and 502 and whole-body regions 501 and 503 as human region information, targeting the image 515 shown in Figure 5(a). (Process 2) As shown in Figure 5(b), the learning unit 122 generates a local region circle 514 with a radius of an arbitrary number of surrounding pixels, based on the local region center position 513 within the local region 500. The learning unit 122 also generates a whole region circle 504 with a radius of an arbitrary number of surrounding pixels, based on the whole region center position 505 within the whole region 501. Then, the learning unit 122 calculates the common tangent line 506 of the local region circle 514 and the whole region circle 504, and defines the area enclosed by the local region circle 514, the whole region circle 504, and the common tangent line 506 as the part mask region 507. Alternatively, the learning unit 122 may define a line segment region 508 of arbitrary thickness connecting the local region center position 513 and the whole region center position 505 instead of the part mask region 507. In the following explanation, the part mask region 507 will be used. (Process 3) The learning unit 122 performs substantially the same processing on the local region 502 and the whole-body region 503 as it did on the local region 500 and the whole-body region 501, and defines the region mask region 509. (Process 4) The learning unit 122 generates a ground truth map 510, as illustrated in Figure 5(d), so that different map values ​​(pixel values) are assigned to each of the acquired part mask regions 507 and 509. The learning unit 122 also generates an overlapping region map 512, as illustrated in Figure 5(d), from the overlapping region 511 where the part mask regions 507 and 509 overlap each other. Furthermore, the subsequent processing (processing from S203 onwards shown in Figure 2) will be substantially the same as that of the embodiment described above.

[0036] <Second Embodiment> In the first embodiment, if each subject has a different map value, it is possible to estimate which subject each region in the image belongs to without limiting it to a specific value. Therefore, the model is trained to reduce the variance of the scores of each subject's region without setting a ground truth value. On the other hand, training without setting a ground truth value can be difficult to converge. In light of this situation, the second embodiment of this disclosure describes an example of a method for training a machine learning model so that the map value within the region of a person converges to an arbitrary range of values.

[0037] Referring to Figure 6, an example of the processing related to the learning of a machine learning model applied to the attribute estimation unit 113 according to this embodiment will be described. Note that the processing S600 to S602, S604, S606, and S607 shown in Figure 6 are substantially the same as the processing S200 to S203, S205, and S206 shown in Figure 2, so a detailed explanation will be omitted. Hereafter, the processing S603 and S604 will be described.

[0038] In S603, the learning unit 122 generates a foreground object attribute score map corresponding to the area where an object that occludes other objects (hereinafter also referred to as an occluding object) or an object that is not occluded exists (hereinafter also referred to as a foreground object area). The learning unit 122 also generates a background object area attribute score map corresponding to the area where an object occluded by another object that does not belong to the foreground area exists (hereinafter also referred to as a background object area). Here, focusing on the case where a mask area as exemplified in Figure 7(a) is acquired as person area information in S601, an example of the map generation method in the processing of S603 will be explained.

[0039] (Processing 1) The learning unit 122 generates a foreground object attribute score map 704, as illustrated in Figure 7(b), from the mask regions 701 and 702 of the person, which are foreground objects, among the person region information. In the example shown in Figure 7(b), the foreground object attribute score map 704 assigns map values ​​to each region, with the values ​​in the mask regions 701 and 702 being "1" and the values ​​in the other regions being "0". (Process 2) The learning unit 122 generates a rear object exclusion region 707 from the intersection of mask regions 702 and 703. Then, the learning unit 122 uses the mask region 703 and the rear object exclusion region 707 to generate a rear object attribute score map 705 as illustrated in Figure 7(c). In the example shown in Figure 7(c), the rear object attribute score map 705 assigns map values ​​to each region, with the values ​​in the mask region 706 that does not include the rear object exclusion region 707 being set to "1", and the values ​​in the other regions being set to "0".

[0040] In S604, the learning unit 122 calculates the loss for the foreground object attribute score map and the background object attribute score map, in addition to calculating the PushLoss and PullLoss.

[0041] The loss function for calculating the loss value (FrontLoss) related to the foreground object attribute score map is expressed by the following relational equation (1-5). In equation (1-5), Score_th indicates the threshold value of the map value to be subject to loss and is predetermined. In FrontLoss, the average of the absolute difference between the map value and the threshold in the region where the map value is below the threshold in the foreground object attribute score map is calculated as the loss. This makes it possible to converge the learning of the machine learning model so that the map value in the region where foreground objects exist becomes greater than the threshold value.

[0042]

number

[0043] The loss function used to calculate the loss value (BackLoss) for the back object attribute score map is expressed by the following relation (1-6). In BackLoss, the average difference between the map value and the threshold in the region where the map value is greater than the threshold is calculated as the loss. This makes it possible to converge the learning of the machine learning model so that the map value in the region where back objects exist, excluding overlapping regions, becomes less than the threshold.

[0044]

number

[0045] The loss function applied in S604 is expressed by the following relation, equation (1-7), as the sum of PushLoss, PullLoss, equations (1-5) and (1-6). In equation (1-7), α, γ, and σ are empirically obtained hyperparameters.

[0046]

number

[0047] In this embodiment, we have described an example of a method for adding a loss function using a threshold for map values ​​for foreground and background objects when training a model that performs attribute score map estimation. As described above, in this method, when calculating the loss value applied when training the machine learning model, the loss for occluded objects or objects that are not occluded (FrontLoss) and the loss for occluded objects (BackLoss) are taken into account. By applying such a method, it becomes possible to train the machine learning model so that the map values ​​converge to a specific range, and thus it is expected that the convergence of training will be accelerated.

[0048] <Third Embodiment> In a third embodiment of this disclosure, an example of a method is described in which, in addition to the loss functions described in the first and second embodiments, a loss function relating to the attribute score map of an area where no people exist (hereinafter referred to as the background area) is applied.

[0049] Referring to Figure 8, an example of the processing related to the learning of a machine learning model applied to the attribute estimation unit 113 according to this embodiment will be described. Note that the processing S800 to S803, S805, S807, and S808 shown in Figure 8 are substantially the same as the processing S600 to S604, S606, and S607 shown in Figure 6, so a detailed explanation will be omitted. Hereafter, the processing S804 and S806 will be described.

[0050] In S804, the learning unit 122 generates an attribute score map of the background region. Here, we will explain an example of the map generation method in the processing of S804, focusing on the case where the mask region shown in Figure 9(a) is obtained as person region information in S801.

[0051] The learning unit 122 generates a background region attribute score map 905, as exemplified in Figure 9(b), from the background region 904, which corresponds to the areas in image 900 other than the mask regions 901 to 903 of the person. In the example shown in Figure 9(b), the background region attribute score map 905 assigns map values ​​to each region, with the values ​​in the background region 904 being "1" and the values ​​in the mask regions 901 to 903 being "0".

[0052] In S806, the learning unit 122 calculates PushLoss, PullLoss, FrontLoss, and BackLoss, as well as a loss related to the background region attribute map (BackGroundLoss).

[0053] The loss function used to calculate BackGroundLoss is expressed by the following relation, shown as equation (1-8). In equation (1-8), value_th is a threshold value that corresponds to the value at which the map value in the background region converges, and is predetermined. In BackGroundLoss, the average of the absolute difference between value_th and the map value is calculated as the loss, so that the map value in the background region converges to value_th. This makes it possible to converge the map value in the background region to a fixed value. Therefore, for example, by training a machine learning model so that the map value of the person region becomes a value other than the map value of the background region, it becomes possible to separate the attributes of the person region and the background region. Note that value_th may be the same value as Score_th in the second embodiment.

[0054]

number

[0055] The loss function applied in S806 is expressed as the sum of PushLoss, PullLoss, FrontLss, BackLoss, and equation (1-8), and is shown below as equation (1-9). In equation (1-9), α, γ, σ, and ε are empirically obtained hyperparameters.

[0056]

number

[0057] In this embodiment, we have described, as an example, a method for adding a loss function that converges the map values ​​of the background region to a unique value during the training of a model that performs attribute score map estimation. In this method, as described above, when calculating the loss value applied during the training of the machine learning model, it is taken into account that the scores in areas where no subject exists converge to a pre-set score. By applying such a method, it becomes possible to encourage the machine learning model to train to converge the map values ​​of the background region to a specific value, and thus it is expected to have the effect of promoting the separation of map values ​​between the person region and other regions.

[0058] <Other Embodiments> The present invention can also be realized by supplying a program that implements one or more of the functions of the above-described embodiments to a system or device via a network or storage medium, and by having one or more processors in the computer of that system or device read and execute the program. It can also be realized by a circuit (e.g., an ASIC) that implements one or more functions.

[0059] Furthermore, the disclosure of this embodiment includes the following configurations, methods, and programs. (Configuration 1) The configuration includes estimation means for estimating which of the multiple subjects a first region corresponding to at least a portion of the subjects in an image in which multiple subjects are captured, The estimation means is characterized in that learning is performed using a loss function that calculates a loss value based on at least a second region in which two or more of the multiple subjects overlap. (Configuration 2) The information processing apparatus according to Configuration 1, characterized in that the subject is a person, and the first region is a region that includes at least some of the multiple parts of the person. (Configuration 3) The information processing apparatus according to Configuration 1 or 2, wherein the estimation means is trained to output different scores for each of the multiple subjects. (Configuration 4) The information processing apparatus according to Configuration 3, wherein the estimation means is trained to output different scores between two or more first regions belonging to each of the two or more overlapping subjects. (Configuration 5) An information processing device comprising: estimation means for estimating which of the multiple subjects a first region corresponding to at least a portion of the subjects in an image in which multiple subjects are captured belongs; and learning means for learning the estimation means using a loss function that calculates a loss value based on at least a second region in which two or more of the multiple subjects overlap. (Configuration 6) The information processing device according to Configuration 5, wherein the loss function is characterized by weighting the loss value according to the size of the second region. (Configuration 7) The information processing apparatus according to Configuration 5 or 6, wherein the loss function is characterized by weighting the loss value such that the difference between the scores output for each of the two or more first regions belonging to a common subject becomes smaller. (Configuration 8) The information processing apparatus according to any one of Configurations 5 to 7, characterized in that the loss function calculates the loss value by taking into account the loss for a subject not obscured by other subjects and the loss for a subject obscured by other subjects. (Configuration 9) The information processing apparatus according to any one of Configurations 3 to 6, wherein the loss function calculates the loss value taking into account that the score in the region where the subject does not exist converges to a pre-set score. (Method 1) A method for controlling an information processing device, comprising an estimation step of estimating which of the multiple subjects a first region corresponding to at least a portion of the subjects in an image in which multiple subjects are captured belongs, and characterized in that the machine learning model applied in the estimation step is trained using a loss function that calculates a loss value based on at least a second region in which two or more of the multiple subjects overlap. (Method 2) A method for controlling an information processing device, comprising: an estimation step of estimating which of the multiple subjects a first region corresponding to at least a portion of the subjects in an image in which multiple subjects are captured belongs; and a learning means of learning a machine learning model applied in the estimation step using a loss function that calculates a loss value based on at least a second region in which two or more of the multiple subjects overlap. (Program 1) A program for causing a computer to function as an information processing device, having estimation means for estimating which of the multiple subjects a first region corresponding to at least a portion of the subjects in an image in which multiple subjects are captured belongs, wherein the estimation means is trained using a loss function that calculates a loss value based on at least a second region in which two or more of the multiple subjects overlap. (Program 2) A program for causing a computer to function as an information processing device, comprising: estimation means for estimating which of the multiple subjects a first region corresponding to at least a portion of the subjects in an image in which multiple subjects are captured belongs; and learning means for learning the estimation means using a loss function that calculates a loss value based on at least a second region in which two or more of the multiple subjects overlap. [Explanation of Symbols]

[0060] 110 Information Processing Device 113 Attribute estimation part 121 Learning device 122 Learning Department

Claims

1. The system has estimation means for estimating which of the multiple subjects a first region corresponding to at least a portion of the subjects in an image in which multiple subjects are captured belongs. The estimation means uses a trained model that has been trained using a loss function that weights the loss value according to the size of a second region where two or more of the multiple subjects overlap or its proportion to a predetermined region. An information processing device characterized by the following features.

2. The subject in question is a person. The first region is a region that includes at least some of the multiple body parts of the person. The information processing apparatus according to claim 1, characterized in that

3. The information processing apparatus according to claim 1, characterized in that the estimation means is trained to output different scores for each of the multiple subjects.

4. The information processing apparatus according to claim 3, characterized in that the estimation means is trained to output different scores in two or more first regions belonging to each of the two or more overlapping subjects.

5. An estimation means for estimating which of the multiple subjects a first region corresponding to at least a portion of the subjects in an image in which multiple subjects are captured belongs, A learning means that learns the estimation means using a loss function that weights the loss value according to the size of a second region where two or more subjects from the plurality of subjects overlap or as a percentage of a predetermined region, An information processing device characterized by having

6. The information processing apparatus according to claim 1 or 5, characterized in that the loss function weights the loss values ​​such that the difference between the scores output for each of the two or more first regions belonging to a common subject becomes smaller.

7. The information processing apparatus according to claim 3, characterized in that the loss function calculates the loss value taking into account that the score in the region where the subject does not exist converges to a pre-set score.

8. The loss function comprises a first loss corresponding to the region of existence of a subject not obscured by other subjects and a subject obscuring other subjects, A second loss corresponding to the region of existence of the subject that is obscured by the subject that obscures the other subject, The information processing apparatus according to claim 1 or 5, characterized in that it calculates the loss value using the method described above.

9. The information processing apparatus according to claim 8, characterized in that the second loss is calculated based on the region obtained by excluding the region that overlaps with the occluding object from the region where the occluding object exists.

10. A method for controlling an information processing device, The method includes an estimation step of estimating which of the multiple subjects a first region corresponding to at least a portion of the subjects in an image in which multiple subjects are captured belongs. The machine learning model applied in the estimation step is trained using a loss function that weights the loss value according to the size of a second region where two or more of the multiple subjects overlap or their proportion to a predetermined region. A control method for an information processing device, characterized by the features described herein.

11. A method for controlling an information processing device, An estimation step of estimating which of the multiple subjects a first region corresponding to at least a portion of the subjects in an image in which multiple subjects are captured belongs, A learning step in which the machine learning model applied in the estimation step is trained using a loss function that weights the loss value according to the size of a second region where two or more subjects from the plurality of subjects overlap or as a percentage of a predetermined region, A control method for an information processing device, characterized by having the following features.

12. A program for causing a computer to execute the control method of the information processing device described in Claim 10.

13. A program for causing a computer to execute the control method of the information processing device described in Claim 11.

Citation Information

Patent Citations

  • Double-task pedestrian detection method with head information

    CN113642520A

  • Image classification method based on online knowledge distillation

    CN116206327A

  • Target detection method and device, network equipment and storage medium

    CN116612279A

  • Object detection device, object detection method, program, and recording medium

    JP2021117530A

  • Information processing apparatus, information processing method, and program

    JP2022019339A