Attribute identification device and attribute identification method

By introducing a multi-layer learning model and similarity judgment module into the object recognition device, the problem of long-term object attribute recognition time in the prior art is solved, and a faster attribute recognition process is realized.

JP2025074393APending Publication Date: 2025-05-14NEC CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2023185157
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-10-30
Publication Date
2025-05-14

AI Technical Summary

Technical Problem

The prior art requires a long calculation time when identifying object properties, especially when processing bounding boxes for each frame.

Method used

By introducing modules such as object detection, object tracking, feature acquisition, attribute extraction and similarity judgment into the object recognition device, the multi-layer output data in the learning model is used as feature quantities, and the attribute features of past frames are reused through tracking ID and similarity judgment to reduce the attribute recognition time of each frame.

Benefits of technology

It effectively shortens the time for object attribute recognition, and reduces the computational burden per frame by reusing attributes of similar features.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025074393000001_ABST
    Figure 2025074393000001_ABST
Patent Text Reader

Abstract

To provide an attribute identification device that can reduce a time required for identifying an attribute.SOLUTION: Object detection means, when given a frame, derives a bounding box of a detection target object from within the frame based on a learning model including multiple layers, and defines output data of one layer included in the learning model or the frame itself as a feature value of the frame. Feature value acquisition means acquires a feature value of the bounding box of an attribute identification target object from the feature value of the frame. Similarity degree determination means determines a similarity degree between the feature value of the bounding box of the attribute identification target object and a feature value of a past bounding box corresponding to a tracking ID of the bounding box, and when the similarity degree is equal to or greater than a predetermined threshold, uses an attribute corresponding to the feature value of the past bounding box to identify an attribute of the attribute identification target object in the bounding box.SELECTED DRAWING: Figure 8
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present disclosure relates to an attribute specifying device, an attribute specifying method, and an attribute specifying program. [Background technology]

[0002] Methods for identifying attributes of objects in images are described in Patent Document 1 and Non-Patent Document 1. Patent Document 1 describes identifying attributes from feature quantities related to traffic lines. Non-Patent Document 1 describes a technology for extracting object attributes from images.

[0003] Taking the attributes of a "human" as an example, the attributes include the color and type of clothes, the color of hair, and the presence or absence of accessories. However, the attributes of a "human" are not limited to these. In addition, the items that become attributes change depending on the type of object. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] JP 2016-57998 A [Non-patent literature]

[0005] [Non-Patent Document 1] Dangwei Li et al., “Multi-attribute Learning for Pedestrian Attribute Recognition in Surveillance Scenarios”, [Retrieved September 14, 2023], Internet<URL :https: / / dangweili.github.io / misc / pdfs / acpr15-att.pdf > Summary of the Invention [Problem to be solved by the invention]

[0006] When the techniques described in Patent Document 1 and Non-Patent Document 1 are used to find attributes corresponding to each bounding box for each frame, it takes a long time to perform calculations.

[0007] Therefore, an object of the present invention is to provide an attribute specifying device, an attribute specifying method, and an attribute specifying program that can reduce the time required to specify an attribute. [Means for solving the problem]

[0008] The attribute identification device according to the present disclosure is characterized in that it includes: object detection means for, when a frame is given, deriving a bounding box of a detection object from within the frame based on a learning model including a plurality of layers, and defining output data of one layer included in the learning model or the frame itself as a feature of the frame; object tracking means for adding a tracking ID to the bounding box; feature acquisition means for acquiring features of the bounding box of the attribute-specific object from the features of the frame; attribute extraction means for extracting attributes of the attribute-specific object from the bounding box of the attribute-specific object; storage means for storing a combination of the frame ID, the tracking ID, the features of the bounding box of the attribute-specific object, and the attributes of the attribute-specific object; and similarity determination means for determining a similarity between the features of the bounding box of the attribute-specific object and the features of a past bounding box corresponding to the tracking ID of the bounding box, and if the similarity is equal to or greater than a predetermined threshold, stopping extraction of the attributes of the attribute-specific object and reusing the attributes corresponding to the features of the past bounding box to identify the attributes of the attribute-specific object in the bounding box.

[0009] The attribute identification method according to the present disclosure is characterized in that, when a frame is given to a computer, the computer derives a bounding box of a detection object from within the frame based on a learning model including a plurality of layers, defines output data of one layer included in the learning model or the frame itself as a feature of the frame, adds a tracking ID to the bounding box, obtains a feature of a bounding box of the attribute-identified object from the feature of the frame, extracts attributes of the attribute-identified object from the bounding box of the attribute-identified object, stores a combination of the frame ID, the tracking ID, the feature of the bounding box of the attribute-identified object, and the attributes of the attribute-identified object, calculates a similarity between the feature of the bounding box of the attribute-identified object and a feature of a past bounding box corresponding to the tracking ID of the bounding box, and if the similarity is equal to or greater than a predetermined threshold, stops extraction of the attributes of the attribute-identified object, and identifies the attributes of the attribute-identified object in the bounding box by reusing the attribute corresponding to the feature of the past bounding box.

[0010] An attribute identification program according to the present disclosure causes a computer to execute, when a frame is given, an object detection process of deriving a bounding box of a detection object from within the frame based on a learning model including a plurality of layers, and defining output data of one layer included in the learning model or the frame itself as a feature of the frame, an object tracking process of adding a tracking ID to the bounding box, a feature acquisition process of acquiring features of a bounding box of an attribute-specific object from the features of the frame, an attribute extraction process of extracting attributes of the attribute-specific object from the bounding box of the attribute-specific object, a storage process of storing a combination of the frame ID, the tracking ID, the features of the bounding box of the attribute-specific object, and the attributes of the attribute-specific object in a storage device, and a similarity determination process of determining a similarity between the features of the bounding box of the attribute-specific object and the features of a past bounding box corresponding to the tracking ID of the bounding box, and if the similarity is equal to or greater than a predetermined threshold, stopping extraction of the attributes of the attribute-specific object and reusing the attributes corresponding to the features of the past bounding box to identify the attributes of the attribute-specific object in the bounding box. Effect of the Invention

[0011] According to the present disclosure, it is possible to reduce the time required to identify attributes. [Brief description of the drawings]

[0012] [Figure 1] 1 is a block diagram showing a configuration example of an attribute specifying device according to the present disclosure. [Diagram 2] FIG. 2 is a schematic diagram illustrating an example of a bounding box identified within a frame. [Diagram 3] FIG. 1 is a schematic diagram showing examples of bounding boxes containing a common human being derived from three consecutive frames. [Figure 4] 10 is a flowchart illustrating an example of a process progress of an attribute identifying device according to the present disclosure. [Diagram 5] 10 is a flowchart illustrating an example of a process progress of an attribute identifying device according to the present disclosure. [Figure 6] 10 is a flowchart illustrating an example of a process progress of an attribute identifying device according to the present disclosure. [Figure 7] FIG. 2 is a schematic block diagram showing an example of the configuration of a computer related to an attribute specifying device. [Figure 8] 1 is a block diagram showing an overview of an attribute specifying device according to the present disclosure. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0013] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings.

[0014] "Object detection" is an operation of deriving a bounding box of a detected object, a score indicating the reliability of the bounding box, and a class indicating the type of the detected object, based on a given frame.

[0015] A "detection target" is an object for which a bounding box, score, and class are to be derived from a frame. For example, a "human" or a "vehicle" is an example of a detection target. However, the detection target is not limited to these.

[0016] An "attribute-specific object" is an object of a detection object whose attribute should be specified. For example, when the detection objects are a "human" and a "vehicle" and the attribute of the human is to be specified, the "human" corresponds to the attribute-specific object. In addition, each of the detection objects may correspond to an attribute-specific object.

[0017] 1 is a block diagram showing an example of the configuration of an attribute identification device according to the present disclosure. The attribute identification device 1 includes an object detection unit 2, an object tracking unit 3, a feature acquisition unit 4, an attribute extraction unit 5, a storage unit 6, and a similarity determination unit 7.

[0018] When a frame is given, the object detection unit 2 derives a bounding box of the detection object from within the frame based on a learning model including multiple layers. Specifically, the object detection unit 2 derives a bounding box of the detection object, a score indicating the reliability of the bounding box, and a class indicating the type of the detection object based on the frame and the learning model. FIG. 2 is a schematic diagram showing an example of a bounding box identified within a frame. FIG. 2 shows a bounding box 11 derived from a frame 10. FIG. 2 also illustrates an example in which only one human being, the detection object, is captured within the frame 10. Note that a frame ID is set for the frame.

[0019] Furthermore, a learning model including multiple layers is, for example, a deep neural network, but the learning model is not limited to a deep neural network.

[0020] In addition, the object detection unit 2 determines the output data of one layer included in the learning model or the given frame itself as the feature of the frame. That is, not only the output data of one layer obtained by sequentially applying layers to a frame, but also the result of applying a layer to a frame zero times (i.e., the frame itself) are included in the concept of the feature of the frame. An example of the output data of a layer is, for example, the result of a convolution operation using a layer, but the output data of a layer is not limited to this. The feature of the frame is expressed as a tensor.

[0021] When the multiple layers included in the learning model are divided into a first half and a second half, the object detection unit 2 preferably defines the output data of the first half layer as the feature amount of the frame, because the output data of the first half layer contains information on spatial features and the contour of the detection target.

[0022] The object detection unit 2 inputs the derived bounding boxes, scores and classes to the object tracking unit 3.

[0023] Moreover, the object detection unit 2 inputs the frame ID, the bounding box of the attribute identification object, and the feature amount of the frame to the feature amount acquisition unit 4. The attribute identification object is designated in advance by a user of the attribute identification device 1. In the following, an example will be described in which the attribute identification object is a human being.

[0024] The object tracking unit 3 adds a tracking ID to the input bounding box. The object tracking unit 3 refers to past bounding boxes and tracking IDs, and adds the tracking ID to the bounding box so that bounding boxes including a common detection target have a common tracking ID.

[0025] Furthermore, the object tracking unit 3 adds a new tracking ID to the bounding box of a newly appeared detected object.

[0026] The information indicating the tracking result is the bounding box, score, and class derived by the object detection unit 2 plus the tracking ID of the bounding box. The object tracking unit 3 may output the information indicating the tracking result to the outside.

[0027] The object tracking unit 3 inputs the bounding box of the attribute-specific object and its tracking ID to the attribute extraction unit 5.

[0028] Furthermore, the object tracking unit 3 inputs the tracking ID of the bounding box of the attribute-specific object to the feature acquisition unit 4.

[0029] The feature amount acquiring unit 4 acquires the feature amount of the bounding box of the attribute-identifying object from the feature amount of the frame. The feature amount of the bounding box of the attribute-identifying object is also expressed as a tensor. Specifically, the feature amount acquiring unit 4 acquires the feature amount of the bounding box by cutting out a portion determined by the coordinates (position) of the bounding box from the feature amount of the frame. In other words, the feature amount of the bounding box is data obtained by cutting out a portion determined by the position of the bounding box from the feature amount of the frame.

[0030] Furthermore, the feature amount acquiring unit 4 converts the feature amount (tensor) of the bounding box of the attribute-identified object into, for example, a matrix with a predetermined number of rows or a vector with a predetermined number of elements. Here, an example will be described in which the feature amount acquiring unit 4 converts the feature amount of the bounding box of the attribute-identified object into a matrix with a predetermined number of rows.

[0031] It should be noted that the result of transforming the bounding box feature amount can also be considered as the bounding box feature amount.

[0032] The feature amount acquiring unit 4 stores in the storage unit 6 a combination of the frame ID, the result of the conversion of the feature amount of the bounding box of the attribute-identifying object, and the tracking ID of the bounding box.

[0033] Furthermore, the feature amount acquiring unit 4 inputs the result of converting the feature amount of the bounding box of the attribute identification object and the tracking ID of the bounding box to the similarity determining unit 7.

[0034] The attribute extraction unit 5 extracts the attributes of the attribute-specific object from the bounding box of the attribute-specific object. The method for extracting the attributes may be a known method. For example, the attribute extraction unit 5 may extract the attributes from the bounding box of the attribute-specific object using the technique described in Non-Patent Document 1.

[0035] However, the attribute extraction unit 5 extracts attributes of an attribute-specific object from the bounding box of the attribute-specific object only when the tracking ID of the bounding box is a new tracking ID and when the similarity determination unit 7 does not reuse attributes obtained in the past as the attributes of the attribute-specific object.

[0036] When the attribute extraction unit 5 extracts an attribute of the attribute-identified object, it refers to the tracking ID input from the object tracking unit 3 and adds the extracted attribute to the combination of the frame ID, the conversion result of the feature quantity of the bounding box of the attribute-identified object, and the tracking ID of the bounding box stored in the memory unit 6.

[0037] The storage unit 6 is a storage device that stores a combination of a frame ID, a tracking ID, a feature amount of a bounding box of an attribute-identifying object, and an attribute of the attribute-identifying object. More specifically, the storage unit 6 stores the conversion result of the feature amount as the feature amount of the bounding box of the attribute-identifying object.

[0038] The similarity determination unit 7 obtains a similarity between the feature amount of the bounding box of the attribute-specified object and the feature amount of the past bounding box corresponding to the tracking ID of the bounding box. At this time, the similarity determination unit 7 receives the conversion result of the feature amount of the bounding box of the attribute-specified object and the tracking ID of the bounding box from the feature amount acquisition unit 4. The similarity determination unit 7 obtains the conversion result of the feature amount of the past bounding box corresponding to the tracking ID from the storage unit 6. These two conversion results are represented by a matrix with a predetermined number of rows. The similarity determination unit 7 obtains a centered kernel alignment (CKA) based on the conversion result (matrix with a predetermined number of rows) of the feature amount of the bounding box of the attribute-specified object input from the feature amount acquisition unit 4 and the conversion result (matrix with a predetermined number of rows) of the feature amount of the past bounding box. Then, the similarity determination unit 7 determines the CKA as the above-mentioned similarity. Note that the similarity obtained based on the two matrices is not limited to the CKA, and other values ​​obtained based on the two matrices may be the above-mentioned similarity. The above past is, for example, the most recent past (the previous frame), but is not limited to the most recent past.

[0039] The similarity being equal to or greater than a predetermined threshold means that the bounding box of the attribute-identifying object is similar to the past bounding box corresponding to the tracking ID of the bounding box. Therefore, if the similarity is equal to or greater than a predetermined threshold, the similarity determination unit 7 stops the attribute extraction by the attribute extraction unit 5 and identifies the attribute of the attribute-identifying object in the current bounding box of interest by using the attribute corresponding to the feature amount of the past bounding box (more specifically, the result of the conversion of the feature amount).

[0040] The similarity determination unit 7 then adds the identified attribute to the combination of the frame ID, the conversion result of the feature amount of the bounding box of the attribute-identified object, and the tracking ID of the bounding box, which is stored in the memory unit 6.

[0041] If the similarity is less than the threshold, the attribute extraction unit 5 extracts attributes of the attribute-identified object from the bounding box of the attribute-identified object as described above. Then, the attribute extraction unit 5 adds the extracted attributes to the combination of the frame ID, the conversion result of the feature amount of the bounding box of the attribute-identified object, and the tracking ID of the bounding box, which are stored in the storage unit 6.

[0042] The object detection unit 2, the object tracking unit 3, the feature acquisition unit 4, the attribute extraction unit 5, and the similarity determination unit 7 are realized, for example, by a CPU (Central Processing Unit) of a computer that operates according to an attribute specification program. In this case, the CPU reads the attribute specification program from a program recording medium such as a program storage device of the computer, and operates as the object detection unit 2, the object tracking unit 3, the feature acquisition unit 4, the attribute extraction unit 5, and the similarity determination unit 7 according to the attribute specification program.

[0043] The storage unit 6 is realized, for example, by a storage device provided in the above-mentioned computer.

[0044] 3 is a schematic diagram showing an example of bounding boxes derived from three consecutive frames that contain a common human. Since the three bounding boxes 21, 22, and 23 contain a common human, a common tracking ID is added to the three bounding boxes 21, 22, and 23.

[0045] The bounding box 21 is assumed to be a bounding box derived from a frame in which a human 25 appears for the first time. In this case, the object tracking unit 3 adds a new tracking ID to the bounding box 21. Then, the object tracking unit 3 inputs the bounding box 21 and its tracking ID to the attribute extraction unit 5. The object tracking unit 3 also inputs the tracking ID to the feature acquisition unit 4.

[0046] The feature amount acquisition unit 4 receives the frame ID, the bounding box 21, and the features of the frame from the object detection unit 2. The feature amount acquisition unit 4 acquires the features of the bounding box 21 from the features of the frame, and converts the features into a matrix with a predetermined number of rows. The feature amount acquisition unit 4 then stores in the storage unit 6 a combination of the frame ID, the conversion result of the features of the bounding box 21, and the tracking ID of the bounding box 21.

[0047] In this case, since the tracking ID of the bounding box 21 is a new tracking ID, the attribute extraction unit 5 extracts the attributes of the person 25 from the bounding box 21. Then, the attribute extraction unit 5 adds the extracted attributes to the combination of the frame ID, the conversion result of the feature amount of the bounding box 21, and the tracking ID of the bounding box 21, which includes the tracking ID. As a result, the storage unit 6 stores the combination of the frame ID, the conversion result of the feature amount of the bounding box 21, the tracking ID of the bounding box 21, and the attributes of the person 25.

[0048] Suppose a second frame is input, and a bounding box 22 is derived from that frame. A bag belonging to a person 25 is captured in the bounding box 22. The object tracking unit 3 adds a tracking ID to the bounding box 22 that is the same as that of the bounding box 21. The object tracking unit 3 then inputs the bounding box 22 and its tracking ID to the attribute extraction unit 5. The object tracking unit 3 also inputs the tracking ID to the feature acquisition unit 4.

[0049] The feature acquisition unit 4 receives the frame ID of the second frame, the bounding box 22, and the features of that frame from the object detection unit 2. The feature acquisition unit 4 acquires the features of the bounding box 22 from the features of the frame, and converts the features into a matrix with a predetermined number of rows. The feature acquisition unit 4 then stores a combination of the frame ID, the conversion result of the features of the bounding box 22, and the tracking ID of the bounding box 22 in the storage unit 6. The feature acquisition unit 4 also inputs the conversion result of the features of the bounding box 22, and the tracking ID of the bounding box 22 to the similarity determination unit 7.

[0050] The similarity determination unit 7 acquires from the storage unit 6 the past conversion result of the feature amount of the bounding box 21 corresponding to the tracking ID. The similarity determination unit 7 calculates the CKA based on the conversion result of the feature amount of the bounding box 22 and the past conversion result of the feature amount of the bounding box 21. The similarity determination unit 7 then determines the CKA as the similarity between the feature amount of the bounding box 22 and the feature amount of the bounding box 21. In this example, it is determined that this similarity is less than a predetermined threshold value.

[0051] In this case, the attribute extraction unit 5 extracts the attributes of the person 25 from the bounding box 22. Then, the attribute extraction unit 5 adds the extracted attributes to a combination of the frame ID of the second frame, the result of the conversion of the feature amount of the bounding box 22, and the tracking ID of the bounding box 22, which includes the tracking ID input from the object tracking unit 3. As a result, the storage unit 6 stores the combination of the frame ID of the second frame, the result of the conversion of the feature amount of the bounding box 22, the tracking ID of the bounding box 22, and the attributes of the person 25.

[0052] It is assumed that a third frame is input, and a bounding box 23 is derived from the frame. The bounding box 23 includes a bag of a person 25, as in the bounding box 22. The object tracking unit 3 adds a tracking ID common to the bounding boxes 21 and 22 to the bounding box 23. Then, the object tracking unit 3 inputs the bounding box 23 and its tracking ID to the attribute extraction unit 5. The object tracking unit 3 also inputs the tracking ID to the feature acquisition unit 4.

[0053] The feature acquisition unit 4 receives the frame ID of the third frame, the bounding box 23, and the features of that frame from the object detection unit 2. The feature acquisition unit 4 acquires the features of the bounding box 23 from the features of the frame, and converts the features into a matrix with a predetermined number of rows. The feature acquisition unit 4 then stores a combination of the frame ID, the conversion result of the features of the bounding box 23, and the tracking ID of the bounding box 23 in the storage unit 6. The feature acquisition unit 4 also inputs the conversion result of the features of the bounding box 23, and the tracking ID of the bounding box 23 to the similarity determination unit 7.

[0054] The similarity determination unit 7 acquires from the storage unit 6 the conversion result of the feature amount of the bounding box 22 in the past corresponding to the tracking ID. The similarity determination unit 7 calculates a CKA based on the conversion result of the feature amount of the bounding box 23 and the conversion result of the feature amount of the bounding box 22 in the past. The similarity determination unit 7 then determines the CKA as the similarity between the feature amount of the bounding box 23 and the feature amount of the bounding box 22. In this example, it is determined that this similarity is equal to or greater than a predetermined threshold value.

[0055] In this case, the similarity determination unit 7 stops the extraction of attributes by the attribute extraction unit 5, and identifies the attributes of the person 25 in the current bounding box 23 by reusing the attributes corresponding to the conversion result of the feature amount of the past bounding box 22. Then, the similarity determination unit 7 adds the identified attributes to the combination of the frame ID of the third frame, the conversion result of the feature amount of the bounding box 23, and the tracking ID of the bounding box 23, which includes the tracking ID input from the feature acquisition unit 4. As a result, the storage unit 6 stores the combination of the frame ID of the third frame, the conversion result of the feature amount of the bounding box 23, the tracking ID of the bounding box 23, and the attributes of the person 25.

[0056] In this way, when the similarity between the feature amount of the current bounding box and the feature amount of the past bounding box is equal to or greater than a predetermined threshold, the similarity determination unit 7 stops the attribute extraction by the attribute extraction unit 5 and identifies the attribute of the attribute-specific object in the current bounding box by reusing the attribute corresponding to the conversion result of the feature amount of the past bounding box. Therefore, for one attribute-specific object, the attribute extraction unit 5 does not necessarily extract the attribute of the attribute-specific object in the bounding box every frame, and the similarity determination unit 7 may reuse the attribute corresponding to the current bounding box from the attribute corresponding to the feature amount of the past bounding box. And, the operation of reusing an already existing attribute as the current attribute does not take time. Therefore, the time required for identifying the attribute can be shortened.

[0057] In the above description, the feature acquisition unit 4 converts the feature (tensor) of the bounding box of the attribute-specific object into a matrix with a predetermined number of rows. In this case, the similarity determination unit 7 uses CKA as the similarity between the feature of the current bounding box and the feature of the past bounding box.

[0058] The feature amount acquiring unit 4 may convert the feature amount (tensor) of the bounding box of the attribute-identified object into a vector with a predetermined number of elements. Even in this case, the feature amount acquiring unit 4 stores a combination of the frame ID, the conversion result of the feature amount of the bounding box of the attribute-identified object, and the tracking ID of the bounding box in the storage unit 6. In addition, the feature amount acquiring unit 4 inputs the conversion result of the feature amount of the bounding box of the attribute-identified object and the tracking ID of the bounding box to the similarity determining unit 7.

[0059] When calculating the similarity between the bounding box feature of the attribute-specified object and the feature of the past bounding box corresponding to the tracking ID of the bounding box, the similarity determination unit 7 may calculate the cosine similarity between the transformation result (vector) of the bounding box feature of the attribute-specified object and the transformation result (vector) of the feature of the past bounding box. This cosine similarity can also be used as the similarity between the feature of the current bounding box and the feature of the past bounding box.

[0060] This is the same as the above description, except that the feature acquisition unit 4 converts the feature (tensor) of the bounding box of the attribute-specific object into a vector with a predetermined number of elements, and the similarity determination unit 7 calculates the cosine similarity as the similarity between the feature of the current bounding box and the feature of a past bounding box.

[0061] Next, the process will be described. Figures 4, 5, and 6 are flowcharts showing an example of the process of the attribute specification device according to the present disclosure. Detailed description of items that have already been described will be omitted.

[0062] First, the object detection unit 2 receives one frame (step S1).

[0063] Then, the object detection unit 2 derives a set of a bounding box, a score, and a class based on the frame, and determines the feature amount of the frame (step S2). The object detection unit 2 inputs the set of the bounding box, the score, and the class to the object tracking unit 3. The object detection unit 2 also inputs the frame ID, the bounding box of the attribute-specified object, and the feature amount of the frame to the feature amount acquisition unit 4.

[0064] After step S2, the object tracking unit 3 adds a tracking ID to the bounding box (step S3). The object tracking unit 3 inputs the bounding box of the attribute-specific object and its tracking ID to the attribute extraction unit 5. The object tracking unit 3 also inputs the tracking ID of the bounding box of the attribute-specific object to the feature acquisition unit 4.

[0065] Following step S3, the feature amount acquiring unit 4 acquires the feature amount of the bounding box of the attribute identification object from the feature amount of the frame, and converts the feature amount (step S4).

[0066] Next, the feature amount acquiring unit 4 stores a combination of the frame ID, the result of the conversion of the feature amount of the bounding box of the attribute-identified object, and the tracking ID of the bounding box in the storage unit 6 (step S5). The feature amount acquiring unit 4 also inputs the result of the conversion of the feature amount of the bounding box of the attribute-identified object and the tracking ID of the bounding box to the similarity determining unit 7.

[0067] Next, the object tracking unit 3 judges whether or not the tracking ID added in step S3 is a new tracking ID (step S6).

[0068] If the tracking ID is a new tracking ID (Yes in step S6), the process proceeds to step S7 (see FIG. 5).

[0069] In step S7, the attribute extraction unit 5 extracts the attributes of the attribute identification object from the bounding box of the attribute identification object.

[0070] Next, the attribute extraction unit 5 adds the attribute extracted in step S7 to the combination stored in step S5 (step S8). In step S8, the process ends.

[0071] If the tracking ID is not a new tracking ID, the process proceeds to step S9 (see FIG. 6) (No in step S6).

[0072] In step S9, the similarity determination unit 7 determines the similarity between the feature amount of the bounding box of the attribute-identifying object and the feature amount of a past bounding box corresponding to the tracking ID of that bounding box.

[0073] Next, the similarity determination unit 7 determines whether or not the similarity is equal to or greater than a predetermined threshold value (step S10).

[0074] If the similarity is less than the predetermined threshold (No in step S10), the above-mentioned steps S7 and S8 (see FIG. 5) are executed and the process ends. Steps S7 and S8 have already been described, so a description thereof will be omitted here.

[0075] If the similarity is equal to or greater than the predetermined threshold value (Yes in step S10), the process proceeds to step S11.

[0076] In step S11, the similarity determination unit 7 identifies the attributes of the attribute-identification target object in the current bounding box of interest by utilizing the attributes corresponding to the conversion result of the feature amount of the past bounding box.

[0077] The similarity determination unit 7 then adds the attribute identified in step S11 to the combination stored in step S5 (step S12). The process ends in step S12.

[0078] When the next frame is sent, the attribute identification device 1 may repeat the operations from step S1 onwards.

[0079] As described above, for one attribute-specific object, the attribute extraction unit 5 does not necessarily extract the attributes of the attribute-specific object in the bounding box for each frame, and the similarity determination unit 7 may reuse the attributes corresponding to the current bounding box from the attributes corresponding to the feature amount of the past bounding box. And, the operation of reusing the existing attributes as the current attributes does not take time. Therefore, the time required for identifying the attributes can be shortened.

[0080] In the above embodiment, the predetermined threshold value to be compared with the similarity degree may be specified in advance by the user of the attribute identification device 1.

[0081] In the above embodiment, the feature acquisition unit 4 converts the features and stores the converted results of the features in the storage unit 6. The feature acquisition unit 4 may not convert the features, and may store the features before conversion in the storage unit 6. When the similarity determination unit 7 determines the similarity between the features of the bounding box of the attribute-specific object and the features of a past bounding box corresponding to the tracking ID of the bounding box, the similarity determination unit 7 may convert each feature into a matrix with a predetermined number of rows or a vector with a predetermined number of elements, and use the conversion result to determine the CKA or cosine similarity.

[0082] In the above embodiment, the feature amount acquiring unit 4 acquires the feature amount of the bounding box of the attribute identification object from the feature amount of the frame. The feature amount acquiring unit 4 may calculate the feature amount of the bounding box of the attribute identification object based on the frame.

[0083] 7 is a schematic block diagram showing an example of the configuration of a computer related to an attribute specification device. The computer 2000 includes, for example, a CPU 2001, a main memory device 2002, an auxiliary memory device 2003, and an interface 2004.

[0084] The attribute specification device according to the present disclosure is realized, for example, by a computer 2000. The operation of the attribute specification device is stored in the form of a program (attribute specification program) in an auxiliary storage device 2003. A CPU 2001 reads out the program from the auxiliary storage device 2003, loads the program in a main storage device 2002, and executes the processing described in the above embodiment in accordance with the program.

[0085] The auxiliary storage device 2003 is an example of a non-transient tangible medium. Other examples of non-transient tangible media include a magnetic disk, a magneto-optical disk, a CD-ROM (Compact Disk Read Only Memory), a DVD-ROM (Digital Versatile Disk Read Only Memory), a semiconductor memory, and the like, which are connected via an interface 2004.

[0086] Next, an overview of the attribute identification device according to the present disclosure will be described. Fig. 8 is a block diagram showing an overview of the attribute identification device according to the present disclosure. The attribute identification device includes an object detection means 72, an object tracking means 73, a feature amount acquisition means 74, an attribute extraction means 75, a storage means 76, and a similarity determination means 77.

[0087] When a frame is given to the object detection means 72 (e.g., object detection unit 2), it derives a bounding box of the object to be detected from within the frame based on a learning model including multiple layers, and determines the output data of one layer included in the learning model or the frame itself as the feature of the frame.

[0088] The object tracking means 73 (for example, the object tracking unit 3) adds a tracking ID to the bounding box.

[0089] The feature amount acquiring means 74 (for example, the feature amount acquiring unit 4) acquires the feature amount of the bounding box of the attribute identification object from the feature amount of the frame.

[0090] The attribute extraction means 75 (for example, the attribute extraction unit 5) extracts the attribute of the attribute identification object from the bounding box of the attribute identification object.

[0091] The storage means 76 (for example, the storage unit 6) stores a combination of a frame ID, a tracking ID, a feature amount of a bounding box of an attribute-specific object, and an attribute of the attribute-specific object.

[0092] A similarity determination means 77 (e.g., similarity determination unit 7) calculates the similarity between the features of the bounding box of the attribute-identified object and the features of a past bounding box corresponding to the tracking ID of the bounding box, and if the similarity is equal to or greater than a predetermined threshold, stops extracting the attributes of the attribute-identified object and identifies the attributes of the attribute-identified object in the bounding box by reusing the attributes corresponding to the features of the past bounding box.

[0093] Such an arrangement can reduce the time required to identify attributes.

[0094] The above-described embodiment of the present invention can be described as follows, but is not limited to the following.

[0095] (Appendix 1) an object detection means for, when a frame is given, deriving a bounding box of a detection target object from within the frame based on a learning model including a plurality of layers, and defining output data of one layer included in the learning model or the frame itself as a feature of the frame; an object tracking means for adding a tracking ID to the bounding box; A feature amount acquisition means for acquiring a feature amount of a bounding box of an attribute-specific object from the feature amount of the frame; an attribute extraction means for extracting attributes of the attribute identification object from a bounding box of the attribute identification object; a storage means for storing a combination of a frame ID, the tracking ID, a feature amount of a bounding box of the attribute identification object, and an attribute of the attribute identification object; a similarity determination means for determining a similarity between a feature amount of the bounding box of the attribute-specific object and a feature amount of a past bounding box corresponding to the tracking ID of the bounding box, and if the similarity is equal to or greater than a predetermined threshold, stopping the extraction of the attribute of the attribute-specific object and reusing the attribute corresponding to the feature amount of the past bounding box to determine the attribute of the attribute-specific object in the bounding box. An attribute specification device comprising:

[0096] (Appendix 2) The object detection means includes: When the multiple layers included in the learning model are divided into a first half and a second half, the output data of the first half layer is defined as the feature amount of the frame. 2. An attribute determination device as described in claim 1.

[0097] (Appendix 3) The similarity determination means, when determining the similarity between the feature amount of the bounding box of the attribute-specified object and the feature amount of a past bounding box corresponding to the tracking ID of the bounding box, determines a centered kernel alignment (CKA) based on a conversion result obtained by converting each of the two feature amounts into a matrix having a predetermined number of rows, and sets the CKA as the similarity between the two feature amounts. 3. An attribute identification device according to claim 1 or 2.

[0098] (Appendix 4) The similarity determination means, when determining the similarity between the feature amount of the bounding box of the attribute-specified object and the feature amount of a past bounding box corresponding to the tracking ID of the bounding box, determines a cosine similarity based on a conversion result obtained by converting each of the two feature amounts into a vector having a predetermined number of elements, and sets the cosine similarity as the similarity between the two feature amounts. 3. An attribute identification device according to claim 1 or 2.

[0099] (Appendix 5) The computer When a frame is given, a bounding box of a detection object is derived from within the frame based on a learning model including multiple layers, and output data of one layer included in the learning model or the frame itself is defined as a feature of the frame; Attaching a tracking ID to the bounding box; Obtaining a feature amount of a bounding box of an attribute-specific object from the feature amount of the frame; Extracting attributes of the attribute-specific object from a bounding box of the attribute-specific object; storing a combination of a frame ID, the tracking ID, a feature amount of a bounding box of the attribute identification object, and an attribute of the attribute identification object; A similarity between the feature amount of the bounding box of the attribute-specific object and the feature amount of a past bounding box corresponding to the tracking ID of the bounding box is calculated, and if the similarity is equal to or greater than a predetermined threshold, the extraction of the attribute of the attribute-specific object is stopped, and the attribute corresponding to the feature amount of the past bounding box is reused to identify the attribute of the attribute-specific object in the bounding box. 13. A method for identifying attributes, comprising:

[0100] (Appendix 6) On the computer, an object detection process in which, when a frame is given, a bounding box of a detection target object is derived from within the frame based on a learning model including multiple layers, and output data of one layer included in the learning model or the frame itself is defined as a feature of the frame; an object tracking process that adds a tracking ID to the bounding box; A feature amount acquisition process for acquiring a feature amount of a bounding box of an attribute-specific object from the feature amount of the frame; an attribute extraction process for extracting attributes of the attribute identification object from a bounding box of the attribute identification object; a storage process for storing a combination of a frame ID, the tracking ID, a feature amount of a bounding box of the attribute-identifying object, and an attribute of the attribute-identifying object in a storage device; and A similarity determination process for determining the similarity between the feature amount of the bounding box of the attribute-specific object and the feature amount of a past bounding box corresponding to the tracking ID of the bounding box, and if the similarity is equal to or greater than a predetermined threshold, stopping the extraction of the attribute of the attribute-specific object and reusing the attribute corresponding to the feature amount of the past bounding box to determine the attribute of the attribute-specific object in the bounding box. Attribute specific program for executing.

[0101] Some or all of the configurations described in Supplementary Note 2 to Supplementary Note 4 that are dependent on Supplementary Note 1 above may also be dependent on Supplementary Note 5 and Supplementary Note 6 in the same dependent relationship as Supplementary Note 2 to Supplementary Note 4. Furthermore, not limited to Supplementary Note 1, Supplementary Note 5, and Supplementary Note 6, various hardware, software, various recording means for recording software, or systems may similarly be made to be dependent on some or all of the configurations described as supplementary notes within the scope of the above-mentioned embodiment.

[0102] Although the present disclosure has been described above with reference to the embodiments, the present disclosure is not limited to the above-described embodiments. Various modifications that can be understood by a person skilled in the art can be made to the configuration and details of the present disclosure within the scope of the present disclosure. [Industrial Applicability]

[0103] The present invention can be suitably applied to an attribute specifying device that specifies the attribute of an attribute specifying object in a frame. [Explanation of symbols]

[0104] 1 Attribute identification device 2. Object detection section 3. Object Tracking Unit 4. Feature Acquisition Unit 5 Attribute extraction part 6 Memory section 7 Similarity determination section

Claims

1. an object detection means for, when a frame is given, deriving a bounding box of a detection target object from within the frame based on a learning model including a plurality of layers, and defining output data of one layer included in the learning model or the frame itself as a feature of the frame; an object tracking means for adding a tracking ID to the bounding box; A feature amount acquisition means for acquiring a feature amount of a bounding box of an attribute-specific object from the feature amount of the frame; an attribute extraction means for extracting attributes of the attribute identification object from a bounding box of the attribute identification object; a storage means for storing a combination of a frame ID, the tracking ID, a feature amount of a bounding box of the attribute identification object, and an attribute of the attribute identification object; a similarity determination means for determining a similarity between a feature amount of the bounding box of the attribute-specific object and a feature amount of a past bounding box corresponding to the tracking ID of the bounding box, and if the similarity is equal to or greater than a predetermined threshold, stopping the extraction of the attribute of the attribute-specific object and reusing the attribute corresponding to the feature amount of the past bounding box to determine the attribute of the attribute-specific object in the bounding box. An attribute specification device comprising:

2. The object detection means includes: When the multiple layers included in the learning model are divided into a first half and a second half, the output data of the first half layer is defined as the feature amount of the frame. The attribute specification device according to claim 1 .

3. The similarity determination means, when determining the similarity between the feature amount of the bounding box of the attribute-specified object and the feature amount of a past bounding box corresponding to the tracking ID of the bounding box, determines a centered kernel alignment (CKA) based on a conversion result obtained by converting each of the two feature amounts into a matrix having a predetermined number of rows, and sets the CKA as the similarity between the two feature amounts. The attribute specification device according to claim 1 or 2.

4. The similarity determination means, when determining the similarity between the feature amount of the bounding box of the attribute-specified object and the feature amount of a past bounding box corresponding to the tracking ID of the bounding box, determines a cosine similarity based on a conversion result obtained by converting each of the two feature amounts into a vector having a predetermined number of elements, and sets the cosine similarity as the similarity between the two feature amounts. The attribute specification device according to claim 1 or 2.

5. The computer When a frame is given, a bounding box of a detection object is derived from within the frame based on a learning model including a plurality of layers, and output data of one layer included in the learning model or the frame itself is defined as a feature of the frame; Attaching a tracking ID to the bounding box; Obtaining a feature amount of a bounding box of an attribute-specific object from the feature amount of the frame; Extracting attributes of the attribute-specific object from a bounding box of the attribute-specific object; storing a combination of a frame ID, the tracking ID, a feature amount of a bounding box of the attribute identification object, and an attribute of the attribute identification object; A similarity between the feature amount of the bounding box of the attribute-specific object and the feature amount of a past bounding box corresponding to the tracking ID of the bounding box is calculated, and if the similarity is equal to or greater than a predetermined threshold, the extraction of the attribute of the attribute-specific object is stopped, and the attribute corresponding to the feature amount of the past bounding box is reused to identify the attribute of the attribute-specific object in the bounding box.

13. A method for identifying attributes, comprising:

6. On the computer, an object detection process in which, when a frame is given, a bounding box of a detection target object is derived from within the frame based on a learning model including multiple layers, and output data of one layer included in the learning model or the frame itself is defined as a feature of the frame; an object tracking process that adds a tracking ID to the bounding box; A feature amount acquisition process for acquiring a feature amount of a bounding box of an attribute-specific object from the feature amount of the frame; an attribute extraction process for extracting attributes of the attribute identification object from a bounding box of the attribute identification object; a storage process for storing a combination of a frame ID, the tracking ID, a feature amount of a bounding box of the attribute identification object, and an attribute of the attribute identification object in a storage device; and A similarity determination process for determining a similarity between the feature amount of the bounding box of the attribute-specific object and the feature amount of a past bounding box corresponding to the tracking ID of the bounding box, and if the similarity is equal to or greater than a predetermined threshold, stopping the extraction of the attribute of the attribute-specific object and reusing the attribute corresponding to the feature amount of the past bounding box to determine the attribute of the attribute-specific object in the bounding box. Attribute specific program for executing.

Citation Information

Patent Citations

  • Object identification method

    JP2016057998A