Person learning device and person estimation device and programs therefor
The person learning and estimation devices use a model trained on person area data with time/position information to accurately estimate individuals in videos, addressing the challenge of faceless person estimation in existing technologies.
Patent Information
- Application Number
- JP2023196363
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-20
- Publication Date
- 2025-05-30
AI Technical Summary
Existing methods for estimating persons in videos struggle when faces are not visible, particularly when multiple individuals wear similar clothing, and there is uncertainty about the accuracy of statistical data-based methods.
A person learning device and a person estimation device that utilize a person area feature amount calculation model learned from person area data, including time/position information, to accurately estimate persons even when their faces are not shown in the video.
The proposed solution enables accurate person estimation in videos, even when faces are obscured, by leveraging time/position information, thereby improving metadata attachment and video understanding tasks.
Smart Images

Figure 2025082869000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a person learning device, a person estimation device, and programs thereof.
Background Art
[0002] Broadcasting stations have a large number of videos such as footage videos shot daily and programs produced. In order to effectively utilize a large number of videos for program production and the like, it is necessary to attach metadata to the videos.
[0003] As one of the metadata, there is a label (hereinafter, person label) for distinguishing the persons moving in the video. This person label is an important element for understanding the video, and plays an important role in a wide range of tasks such as video search necessary for program production and generation of video description texts.
[0004] On the premise of attaching a person label, it is necessary to accurately estimate the persons shown in the video. Therefore, conventionally, a method of estimating a person based on face feature amounts has been proposed (Patent Document 1). In the method described in this Patent Document 1, when a person's face is not shown in the video such as when the person has their back to the camera, the person cannot be estimated because the face feature amounts cannot be obtained.
[0005] On the other hand, in order to deepen the understanding of the program, information on persons whose faces are not shown is also important. Therefore, a method of estimating a person even when the person's face is not shown in the video has been proposed. For example, Patent Document 2 proposes a method of estimating a person based on clothing when the person's face is not shown in the video. Also, Patent Document 3 proposes a method of estimating a person using statistical data.
Prior Art Documents
Patent Documents
[0006]
Patent Document 1
Patent Document 2
[0007] However, in the method described in Patent Document 2, when multiple people are wearing the same clothing (e.g., uniforms), it is difficult to accurately estimate the person labels. Also, in the method described in Patent Document 3, the specific content of the statistical data is not shown, and it is not certain whether the person labels can be accurately estimated.
[0008] In view of the above problems, an object of the present invention is to provide a person learning device, a person estimation device, and programs thereof that can accurately estimate a person even when no face is shown in the video. [Means for Solving the Problems]
[0009] To solve the above problems, a person learning device according to the present invention is a person learning device that learns a person area feature amount calculation model for estimating a person included in a video using person area data representing a first feature amount, time / position information, and a person label of a person area included in each frame of the video, and includes a data integration unit, a feature amount calculation unit, a loss calculation unit, and an optimization unit.
[0010] According to such a configuration, the data integration unit receives the person area data and integrates the first feature amount and the time / position information for each frame in the input person area data for each person area. The feature amount calculation unit calculates a series of second feature amounts corresponding to the person area by inputting a series of the first feature amount and the time / position information integrated by the data integration unit into the person area feature amount calculation model.
[0011] The loss calculation unit calculates a loss such that the distances between the second feature amounts corresponding to the same person label are close and the distances between the second feature amounts corresponding to different person labels are far. The optimization unit optimizes the parameters of the person region feature amount calculation model so that the loss calculated by the loss calculation unit is minimized.
[0012] For example, in TV programs and movies, the way of showing people and the position of the camera are devised so that viewers can easily understand the content, and there is a connection between the previous and subsequent shots. This connection between the shots is considered to be reflected in the position and time of the person region included in the video. Therefore, the person learning device generates a person region feature amount calculation model that can accurately estimate a person even when no face is shown in the video by also learning the time and position information of the person region.
[0013] Also, in order to solve the above problems, a person estimation device according to the present invention is a person estimation device that estimates a person included in a video using person region data representing a first feature amount, time / position information, and a person label of a person region included in each frame of the video, and a person region feature amount calculation model learned by the person learning device, and includes a data integration unit, a feature amount calculation unit, and a person estimation unit.
[0014] According to such a configuration, the data integration unit receives person region data corresponding to frames in which a person's face is shown and frames in which a person's face is not shown, and integrates the first feature amount and time / position information for each frame in the input person region data for each person region.
[0015] The feature amount calculation unit calculates a series of second feature amounts representing a person corresponding to the person region by inputting the series of the first feature amount and time / position information integrated by the data integration unit into the person region feature amount calculation model. The person estimation unit calculates the distance between the second feature amount of the person region data corresponding to the frame in which the person's face is shown and the second feature amount of the person region data corresponding to the frame in which the person's face is not shown, and outputs the person label that minimizes the calculated distance.
[0016] For example, in TV programs and movies, the way of showing people and the position of the camera are devised so that viewers can easily understand the content, and there is a connection between the previous and subsequent shots. This connection between the shots can be considered to be reflected in the position and time of the person area included in the video. Therefore, the person estimation device can accurately estimate a person even when no face is shown in the video by using a person area feature amount calculation model that has also learned the time and position information of the person area.
[0017] Note that the present invention can also be realized by a program for causing a computer to function as the above-described person learning device or person estimation device.
Effect of the Invention
[0018] According to the present invention, a person can be accurately estimated even when no face is shown in the video.
Brief Description of the Drawings
[0019]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Mode for Carrying Out the Invention
[0020] Hereinafter, embodiments of the present invention will be described with reference to the drawings. However, each of the embodiments described below is for embodying the technical idea of the present invention, and the present invention is not limited to the following unless otherwise specified. Also, the same means may be denoted by the same reference numerals, and the description thereof may be omitted.
[0021] [Configuration of Person Learning Device] With reference to FIG. 1, the configuration of the person learning device 1 according to the embodiment will be described. The person learning device 1 learns a person region feature amount calculation model for estimating a person included in a video using person region data representing a first feature amount, time / position information, and a person label of a person region included in each frame of the video.
[0022] The person region is a region in the video that includes a person, that is, a rectangular region (bounding box) that surrounds the person region. This person region can be generated by performing known person detection and tracking on a moving image. At this time, the person region may directly use the results of known person detection and tracking, or the results may be manually corrected. Also, the person region may be manually generated from the video. Hereinafter, the person region data of the same person existing in consecutive frames may be described as "person-track".
[0023] As described below, the person region data includes a first feature amount f of the person region i j and time / position information c i j and a person label l i which are pre-assigned. This person label l i is a label that can uniquely identify each person. Note that the subscript i is a number assigned to each person-track, that is, a unique number for each person region (however, i is an integer of 1 or more). Also, the subscript j is the number of each frame (however, j is an integer of 1 or more).
[0024] The first feature amount f i jis a vector representing the image features of the person area and can be obtained by any method. Also, the first feature quantity f i j can be obtained using a general CNN (Convolutional Neural Network) or a neural network using an attention mechanism (e.g., Vision Transformer). For example, the first feature quantity f i j is 256 - dimensional. Additionally, the first feature quantity f i j may be a feature quantity designed manually such as SIFT (Scale - Invariant Feature Transform).
[0025] The time - position information c i j is information representing the time and position of the person area. Specifically, the time - position information c i j includes the frame number F of each frame j and the coordinates (x 1,i j , y 1,i j , x 2,i j , y 2,i j ) of the bounding box of each frame. That is, the time - position information c i j is a 5 - dimensional vector (F j , x 1,i j , y 1,i j , x 2,i j , y 2,i j ). Also, the coordinates (x 1,i j , y 1,i j ) represent the upper - left coordinates of the bounding box, and the coordinates (x 2,i j , y 2,i j ) represent the lower - right coordinates of the bounding box.
[0026] In this embodiment, the data integration unit 10 integrates the first feature amount f for each frame i j and the time / position information c i j by averaging for each person area. Specifically, the data integration unit 10 calculates the average value of the first feature amounts f of each frame i 1 ,…,f i ni as the first feature amount f for each person area i (where ni is the number of data included in the i-th person-track).
[0027] In the example of FIG. 2, the video includes a plurality of shots T (T 1 , T 2 ,…, T n ). This shot T represents a sequence of frames F where the screen does not switch. In FIG. 2, for ease of viewing, only one representative frame F of each shot T is shown, but each shot T is composed of a plurality of frames F
[0028] The video includes a total of three persons, which are respectively referred to as person "Alice", person "Bob", and person "Charlie". In FIG. 2, the person area data of person "Alice" is labeled with symbol D 2 , the person area data of person "Bob" is labeled with symbol D 1 , and the person area data of person "Charlie" is labeled with symbol D 3
[0029] Here, since the faces of persons "Bob" and "Charlie" are shown in each shot T, person estimation is easy. On the other hand, in shot T 2 , person "Alice" has her back turned and her face is not shown, but it is considered that there is a connection with the previous shot T 1 . For example, the front and rear shots T 1 , T 2 In this case, the position of the person "Alice" is considered to be close. Therefore, by using the connection of such shots, we estimate the position of the person "Alice" whose face is not shown.
[0030] As shown in FIG. 1, the person learning device 1 includes a data integrating unit 10, a feature amount calculating unit 11, a loss calculating unit 12, and an optimizing unit 13. The data integration unit 10 receives person area data and integrates the first feature amount and time / position information for each frame in the input person area data for each person area.
[0031] In this embodiment, the data integration unit 10 extracts the time and position information c i 1 ,…,c i ni From the time and location information for each person area, c i For example, the data integration unit 10 calculates the time and position information c i As the starting frame number F 1 and the end frame number F ni , and the average coordinates of each vertex of the bounding box in each frame (x 1,i mean ,y 1,i mean ,x 2,i mean ,y 2,i mean In this case, the time and location information c i is a six-dimensional vector (F 1 ,F ni ,x 1,i mean ,y 1,i mean ,x 2,i mean ,y 2,i mean )
[0032] Next, the data integration unit 10 calculates the first feature value f i and time and location information c i Then, the data integration unit 10 arranges the first feature amounts fi and time·position information c i obtains a series of length k therefrom. At this time, the data integration unit 10 arranges the first feature amount f in time series i and time·position information c i sets a window of length k for them, and slides the window at a constant interval s from the beginning to the end of the time series to sequentially obtain the series of the window part (for example, k = 16, s = 8). Note that the series of the first feature amount is f = {f i} i=1 k is set, and the series of time·position information is c = {c i} i=1 k is set.
[0033] Thereafter, the data integration unit 10 outputs the obtained series f of the first feature amount and the series c of the time·position information to the feature amount calculation unit 11.
[0034] The feature amount calculation unit 11 calculates a series z of the second feature amount corresponding to the person area by inputting the series f and c of the first feature amount and the time·position information integrated by the data integration unit 10 into the person area feature amount calculation model M.
[0035] As shown in FIG. 1, the feature amount calculation unit 11 includes a person area feature amount calculation model M. For example, the feature amount calculation unit 11 stores the person area feature amount calculation model M in an internal memory (not shown).
[0036] The person area feature amount calculation model M can use a neural network model that can handle series. For example, the person area feature amount calculation model M may be a Transformer encoder or a Vision Transformer that applies a Transformer to image processing. In the example of FIG. 2, as the person area feature amount calculation model M, a Transformer encoder M2 having an attention mechanism and a fully connected layer (PE: Positional Embedding) M1 is used.
[0037] Here, the feature amount calculation unit 11 inputs the series f of the first feature amount and the series c of the time and position information into the person region feature amount calculation model M after combining them. The methods of combining the series f and c include any of sum, product, or concatenation, or a combination of sum, product, or concatenation and fully connected layers. The sum of the series f and c is represented by f + c. Also, the product of the series f and c is represented by the element-wise product f * c. Further, the concatenation of the series f and c is represented by the vector concatenation [f, c].
[0038] The first feature amount f i and the time and position information c i If the number of dimensions is different, it is advisable to use a fully connected layer. For example, the combination of the sum of the series f and c and a fully connected layer is represented by f + FC(c) (where FC is the fully connected layer). Here, to control the influence of the position and time information c i when using a fully connected layer, a learnable weight may be used, or the output of the fully connected layer may be limited by the maximum value of the norm. Specifically, when using a learnable weight, the combination of the sum of the series f and c and a fully connected layer is represented by f + w * FC(c) (where w is the learnable weight). Also, when limiting the maximum value of the norm, if the norm of FC(c) exceeds a preset maximum value (for example, 0.5), a scalar is multiplied so that it becomes that maximum value. For example, when the maximum value is 0.5, FC(c) is multiplied by 1 / max(1, |FC(c)| / 0.5). Note that max is a function that returns the maximum value.
[0039] In the example of FIG. 2, the number of dimensions of the first feature amount f i (256 dimensions) is different from the number of dimensions of the time and position information c i (6 dimensions). Therefore, the feature amount calculation unit 11 inputs the series c of the time and position information into the fully connected layer M1 and then adds it to the series f of the first feature amount. Then, the feature amount calculation unit 11 inputs the addition result into the Transformer encoder M2. Then, the Transformer encoder M2 outputs the series z of the second feature amount (z 1 , …, z k ).
[0040] Note that in the example of FIG. 2, in the feature space α, the second feature amount z 1, z 3 of the set, and the second feature quantity z 2 , z 4 is in an attracting relationship (Attract). Also, the second feature quantity z 1 , z 2 of the set, the second feature quantity z 1 , z 4 of the set, the second feature quantity z 2 , z 3 of the set, and the second feature quantity z 3 , z 4 of the set is in a repelling relationship (Repel).
[0041] In this way, the feature quantity calculation unit 11 calculates a series z of the second feature quantity corresponding to the person area. Then, the feature quantity calculation unit 11 outputs the calculated series z of the second feature quantity to the loss calculation unit 12.
[0042] The loss calculation unit 12 calculates a loss such that the distances between the second feature quantities corresponding to the same person label are close and the distances between the second feature quantities corresponding to different person labels are far. Note that the series of correct person labels input to the loss calculation unit 12 is l = {l} i=1 k Let's assume.
[0043] Here, the smaller the loss value, the closer the distances between the second feature quantities with the same person label. Also, the smaller the loss value, the farther the distances between the second feature quantities with different person labels. For example, the loss calculation unit 12 calculates a loss L defined by the following equations (1) and (2). Note that δ ij is determined from the series l of the correct person labels, and δ i = 1 when l j = l ij , and δ i = 0 when l j ≠ l ij . Also, g is the distance margin of the second feature quantities of different person labels.
[0044]
Equation
[0045] After that, the loss calculation unit 12 outputs the calculated loss L to the optimization unit 13.
[0046] The optimization unit 13 optimizes the parameters (weights, biases) of the person area feature amount calculation model M so that the loss L calculated by the loss calculation unit 12 is minimized. That is, the optimization unit 13 learns the parameters of the person area feature amount calculation model M that minimize the loss L. For example, the optimization unit 13 uses a known optimization method such as the gradient descent method or Adam to optimize the parameters of the person area feature amount calculation model M.
[0047] [Operation of the person learning device] With reference to FIG. 3, the operation of the person learning device 1 will be described. As shown in FIG. 3, in step S1, the data integration unit 10 integrates the first feature amount and the time / position information for each frame for each person area. In step S2, the feature amount calculation unit 11 calculates a series z of second feature amounts by inputting the series f, c of the first feature amount and the time / position information integrated in step S1 into the person area feature amount calculation model M.
[0048] In step S3, the loss calculation unit 12 calculates a loss L such that the distances between the second feature amounts corresponding to the same person label are close and the distances between the second feature amounts corresponding to different person labels are far. In step S4, the optimization unit 13 optimizes the parameters of the person area feature amount calculation model M so that the loss L calculated in step S3 is minimized.
[0049] In step S5, the optimization unit 13 determines whether to end the learning (optimization) of the person area feature amount calculation model M. For example, when the parameters of the person area feature amount calculation model M converge, the optimization unit 13 determines to end the learning.
[0050] If the learning of the person area feature amount calculation model M is not ended (No in step S5), the person learning device 1 returns to the process of step S1. When the learning of the person area feature amount calculation model M is completed (Yes in step S5), the person learning device 1 proceeds to the process of step S6.
[0051] In step S6, the optimization unit 13 stores the parameters of the learned person area feature amount calculation model M in the internal memory and ends the process. In this way, the person learning device 1 can generate a person area feature amount calculation model M that can accurately estimate a person even when no face is reflected in the video.
[0052] [Configuration of Person Estimation Device] With reference to FIG. 4, the configuration of the person estimation device 2 according to the embodiment will be described. The person estimation device 2 estimates a person included in a video by using person area data representing a first feature amount, time / position information, and a person label of a person area included in each frame of the video, and the person area feature amount calculation model M learned by the person learning device 1 in FIG. 1. As shown in FIG. 4, the person estimation device 2 includes a data integration unit 20, a feature amount calculation unit 21, and a person estimation unit 22.
[0053] The data integration unit 20 is input with person area data corresponding to frames in which a person's face is reflected and frames in which a person's face is not reflected, and integrates the first feature amount and time / position information for each frame in the input person area data for each person area.
[0054] Here, among the person area data (person-track), those including frames in which a person's face is reflected may be described as "body-track", and those not including frames in which a person's face is not reflected may be described as "back-track". The object for which the person estimation device 2 estimates a person is back-track. For body-track, the correct person label is given in advance. That is, the person label of back-track is estimated on the premise that the correct person label is given to body-track.
[0055] In this embodiment, a person-track including body-track and back-track is input to the data integration unit 20. Since the data integration unit 20 performs the same processing as the data integration unit 10 in FIG. 1, its description is omitted.
[0056] After that, the data integration unit 20 outputs the series f' of the first feature amounts of body-track and back-track respectively, and the series c' of time and position information to the feature amount calculation unit 21.
[0057] The feature amount calculation unit 21 calculates the series z' of the second feature amounts corresponding to the person area by inputting the series f', c' of the first feature amounts and time and position information integrated by the data integration unit 20 into the person area feature amount calculation model M. This person area feature amount calculation model M is learned by the person learning device 1 in FIG. 1.
[0058] Here, the feature amount calculation unit 21 inputs the series f', c' of the first feature amounts and time and position information of body-track and back-track respectively into the person area feature amount calculation model M, and calculates the second feature amounts z' of body-track and back-track respectively. In other respects, since the feature amount calculation unit 21 performs the same processing as the feature amount calculation unit 11 in FIG. 1, its description is omitted. After that, the feature amount calculation unit 21 outputs the series z' of the second feature amounts of body-track and back-track respectively to the person estimation unit 22.
[0059] The person estimation unit 22 calculates the distance between the second feature amount of the person area data corresponding to the frame in which the person's face is reflected and the second feature amount of the person area data corresponding to the frame in which the person's face is not reflected, and outputs the person label that minimizes the calculated distance.
[0060] The person estimation unit 22 is input with a series l' of person labels of each person included in body-track. For example, the user of the person estimation device 2 manually inputs the series l of person labels to the person estimation unit 22. Further, the person estimation unit 22 is input with a series z' of the second feature amounts of body-track and back-track respectively from the person estimation unit 22.
[0061] Normally, people shown in the same frame are considered not to be the same person. That is, when a plurality of people are included in the same frame, those people are considered to be different. Therefore, the person estimation unit 22 may exclude body-track that is not estimated to be the same person as the back-track of the estimation target.
[0062] Specifically, the person estimation unit 22 may narrow down, as candidate body-track, what remains after excluding the following (1) and (2) from body-track. (1) body-track included in the same frame as the back-track of the estimation target (2) body-track having the same person label as the body-track in (1)
[0063] Note that there may be cases where the same person is shown in the same frame, such as when the same person is reflected in a mirror. In such a case, the person estimation unit 22 may not perform the narrowing down of the candidate body-track described above and may use all body-track as candidate body-track.
[0064] Next, the person estimation unit 22 calculates the distance between the second feature amount of the back-track of the estimation target and the second feature amount of the candidate body-track. For example, the person estimation unit 22 can calculate an arbitrary distance such as the distance between feature vectors. Then, the person estimation unit 22 obtains the candidate body-track for which the distance between the second feature amounts is the minimum with respect to the back-track of the estimation target.
[0065] In the example of FIG. 5, the video includes a plurality of shots T (T 1 ,…,T m-1,T m ,…,T n ) is included (however, 1 < m < n). Also, assume that a person "Alice" whose face is not shown in shot T m is the target of estimation. In the feature space α, the second feature z 9 of the person "Alice" to be estimated is the second feature of the back-track. Also, in the feature space α, the second feature z 1 , z 7 of the person "Alice", the second feature z 8 , z 10 of the person "Bob", and the second feature z 16 of the person "Charlie" are included as the second features of the candidate body-track.
[0066] First, the person estimation unit 22 excludes the second features z m of the person "Bob" included in shot T 8 , z 10 from the candidate body-track for the person "Alice" to be estimated. As a result, the second features z 1 , z 7 of the person "Alice" and the second features z 16 of the person "Charlie" are narrowed down as candidate body-tracks.
[0067] Next, the person estimation unit 22 calculates the distances between the second feature z 9 of the person to be estimated and the second features z 1 , z 7 , z 16 of the candidate body-track, respectively. Here, for the second feature z 9 of the person to be estimated, the second feature z 7 of the person "Alice" has the closest distance. Therefore, the person estimation unit 22 acquires the person label corresponding to the person "Alice" from the series l´ of person labels and outputs it externally as the estimation result.
[0068] [Operation of the Person Estimation Device] Referring to FIG. 6, the operation of the person estimation device 2 will be described. As shown in FIG. 6, in step S10, for each of body-track and back-track, the data integration unit 20 integrates the first feature amount and time-position information for each frame for each person area. In step S11, the feature amount calculation unit 21 inputs the series f´, c´ of the first feature amount and time-position information of each of body-track and back-track into the person area feature amount calculation model M, thereby calculating the series z´ of the second feature amount of each of body-track and back-track.
[0069] In step S12, the person estimation unit 22 narrows down the candidate body-track, calculates the distance between the second feature amount of the back-track and the second feature amount of the candidate body-track, and outputs the person label with the minimum calculated distance. In this way, by using the person area feature amount calculation model M that has learned the time-position information of the person area, the person estimation device 2 can accurately estimate a person even when no face is shown in the video.
[0070] [Operation and Effect] For example, in a TV program or a movie, in order to make it easier for viewers to understand the content, the way of showing people and the position of the camera are devised, and there is a connection between the front and rear shots. This connection of the shots can be considered to be reflected in the position and time of the person area included in the video.
[0071] Therefore, the person learning device 1 generates a person area feature amount calculation model M that can accurately estimate a person even when no face is shown in the video by learning the time-position information of the person area. Furthermore, by using the person area feature amount calculation model M that has also learned the time-position information of the person area, the person estimation device 2 can accurately estimate a person even when no face is shown in the video. Thereby, by automatically assigning metadata to a person whose face is not shown in the video, it is useful for video search and automatic generation of video description texts.
[0072] Although the embodiments have been described in detail above, the present invention is not limited to the above-described embodiments, and includes design changes and the like within a range not departing from the gist of the present invention.
[0073] In the above-described embodiment, the data integration unit has been described as integrating by using the average value of the first feature amount and the time / position information. However, the present invention is not limited to this. For example, the data integration unit may integrate using the median value or the total value of the first feature amount and the time / position information.
[0074] In the above-described embodiment, the person learning device and the person estimation device have been described as independent devices. However, the present invention is not limited to this. That is, the person learning device and the person estimation device may be integrated into one device.
[0075] In the above-described embodiment, the person learning device and the person estimation device have been described as independent hardware. However, the present invention is not limited to this. For example, the present invention can also be realized by a program for causing hardware resources such as a CPU, a memory, and a hard disk provided in a computer to function as the above-described person learning device or person estimation device. This program may be distributed via a communication line, or may be written on a recording medium such as a CD-ROM or a flash memory and distributed.
Explanation of Reference Numerals
[0076] 1 Person learning device 10 Data integration unit 11 Feature amount calculation unit 12 Loss calculation unit 13 Optimization unit 2 Person estimation device 20 Data integration unit 21 Feature amount calculation unit 22 Person estimation unit
Claims
1. A person learning device that learns a person area feature amount calculation model for estimating a person included in the video using person area data representing a first feature amount, time / position information, and a person label of a person area included in each frame of the video, a data integration unit that inputs the person area data and integrates the first feature amount and time / position information for each frame in the input person area data for each person area; a feature amount calculation unit that calculates a series of second feature amounts corresponding to the person area by inputting a series of the first feature amount and time / position information integrated by the data integration unit into the person area feature amount calculation model; a loss calculation unit that calculates a loss such that the distances between the second feature amounts corresponding to the same person label are close and the distances between the second feature amounts corresponding to different person labels are far; an optimization unit that optimizes the parameters of the person area feature amount calculation model so that the loss calculated by the loss calculation unit is minimized; A person learning device characterized by comprising:
2. The person area feature amount calculation model is a Transformer encoder, and the person learning device according to claim 1 is characterized in that.
3. The data integration unit integrates the first feature amount and time / position information for each frame by averaging for each person area, and the person learning device according to claim 1 is characterized in that.
4. A person estimation device that estimates a person included in the video using person area data representing a first feature amount, time / position information, and a person label of a person area included in each frame of the video, and a person area feature amount calculation model learned by the person learning device according to claim 1, a data integration unit that inputs the person area data corresponding to the frame in which the face of the person is reflected and the frame in which the face of the person is not reflected, and integrates the first feature amount and time / position information for each frame in the input person area data for each person area; a feature amount calculation unit that calculates a series of second feature amounts corresponding to the person area by inputting a series of the first feature amount and time / position information integrated by the data integration unit into the person area feature amount calculation model; a person estimation unit that calculates the distance between the second feature amount of the person area data corresponding to the frame in which the face of the person is reflected and the second feature amount of the person area data corresponding to the frame in which the face of the person is not reflected, and outputs a person label that minimizes the calculated distance; A person estimation device characterized by comprising
5. The person area feature amount calculation model is a Transformer encoder, and the person estimation device according to claim 4, characterized in that.
6. The data integration unit integrates the first feature amount and time / position information for each frame for each person area by averaging, and the person estimation device according to claim 4, characterized in that.
7. A program for causing a computer to function as the person learning device according to claim 1.
8. A program for causing a computer to function as the person estimation device according to claim 4.
Citation Information
Patent Citations
Person estimation device, person estimation method and program
JP2015232759A
Person estimation program, person estimation method, and person estimation device
JP2023109299A
Apparatus estimation device and method, and computer program
JP4439523B2