Person identification device

The person identification device uses imaging distance and aspect ratio to enhance accuracy in identifying known and unknown individuals, while the tracking system efficiently tracks individuals across locations by forming location-specific galleries and integrating person IDs.

JP2026055600APending Publication Date: 2026-03-31今井 龙一 +3
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-09-18
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Conventional person identification systems face challenges in accurately determining whether a person is known or unknown and tracking their movement between different locations, often resulting in large and inefficient galleries due to the inclusion of images from varied locations.

Method used

A person identification device that considers imaging distance and aspect ratio, along with similarity, to make accurate determinations, and a tracking system that forms location-specific galleries and integrates person IDs across multiple locations using additional judgment information.

Benefits of technology

Enables precise identification of known and unknown individuals and tracks their movement across locations without overwhelming the gallery generation process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026055600000001_ABST
    Figure 2026055600000001_ABST
Patent Text Reader

Abstract

To provide an identification device capable of accurately identifying individuals. [Solution] The imaging unit 2 captures the target person and outputs it as a whole image. The whole image acquisition means 4 acquires the whole image that includes the target person. The target person image extraction means 6 extracts the target person image from the whole image. The additional information acquisition means 8 estimates the distance between the imaging unit 2 and the target person, i.e., the imaging distance, based on the whole image. The determination means 10 acquires the similarity between the person in the person image and the most similar person registered in the person DB 12. Furthermore, the determination means 10 acquires the imaging distance of the person image (additional determination information) and, based on the similarity and imaging distance, determines whether the person in the person image is a person registered in the person DB 12. If registered, the person in the person image is identified as the most similar person in the person DB 12; if not registered, the person is newly registered in the person DB 12.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to an apparatus for identifying a person.

Background Art

[0002] An apparatus for identifying a person in a captured image has been proposed. For example, in Patent Document 1, an apparatus is disclosed that calculates the feature vectors (feature quantities) of a plurality of person images in advance, registers them as a gallery, and identifies a person based on the degree of similarity with the feature vector of the person image to be judged.

[0003] If the degree of similarity with a person registered in the gallery exceeds a threshold value, the person is identified as such, and if the degree of similarity is less than the threshold value, the person is registered in the gallery as a new person.

[0004] Thereby, while accurately performing person identification, a new person can be registered in the gallery.

Prior Art Documents

Patent Documents

[0005]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0006] However, conventional apparatuses such as Patent Document 1 have had the following problems.

[0007] First, simply judging by a threshold value of the feature vector of a person image and determining whether the person is already registered in the gallery, the accuracy of known and unknown persons is not necessarily high.

[0008] Secondly, conventional technology forms a gallery of images of people taken at the same location for identification, and it was not intended to identify the same person taken at different locations. If we were to try to generate a gallery that includes people taken at different locations as well as those taken at the same location, the number of people registered in the gallery would become enormous, which would be an undesirable problem in terms of operation.

[0009] The purpose of this invention is to solve at least one of the above-mentioned problems and to provide an identification device that can accurately determine whether a person is known or unknown, or a system that can track the movement of a person between different locations without placing an excessive burden on gallery generation. [Means for solving the problem]

[0010] The following lists several independent features of this invention. Each of these features is independent and does not necessarily need to be combined, but can be combined as desired.

[0011] (1)(2) The person identification device according to the present invention comprises: an imaging unit that captures an overall image including a target person; a recording unit that records images of multiple people to be identified as a gallery; a target person image acquisition means that acquires an overall image from the imaging unit; a target person image extraction means that extracts a target person image from the overall image; an additional judgment information acquisition means that estimates the imaging distance between the target person and the imaging unit based on the overall image, the target person image, or both, and uses this as additional judgment information; and a judgment means that determines whether or not the target person is a person registered in the gallery, taking into account not only the similarity between the target person image and the person images in the gallery, but also the influence based on the additional judgment information.

[0012] By considering not only similarity but also the imaging distance to determine whether or not a person is registered in the gallery, it is possible to make a more accurate determination of known and unknown individuals.

[0013] (3) The person identification device according to this invention is characterized in that the additional information acquisition means acquires the aspect ratio of the target person image as additional information in place of or in addition to the imaging distance, and the determination means makes a determination using the aspect ratio of the target person image as additional information in place of or in addition to the imaging distance.

[0014] Therefore, by considering not only similarity but also aspect ratio to determine whether or not a person is registered in the gallery, it is possible to make a more accurate determination of known and unknown individuals.

[0015] (4) The person identification device according to this invention is characterized in that the imaging unit captures multiple overall images in succession, and the determination means makes a determination on each of the multiple target person images extracted from the multiple overall images that are presumed to be of the same person, and integrates these to make a final determination.

[0016] Because the decision is based on multiple images, it can more accurately distinguish between known and unknown information.

[0017] (5) The person identification device according to this invention is characterized in that, if the determination means determines that the person in the target person image is not a person registered in the gallery, it registers the person as a new person in the gallery.

[0018] Therefore, you can add and update the gallery with new characters.

[0019] (6)(9)(10) The tracking system according to this invention comprises: an imaging unit installed at multiple locations to image a target person; a location gallery forming means that forms a gallery for each location based on the target person image of the person captured at each location, assigning the same ID to the same person; a person ID integration means that acquires each location gallery and generates correspondence data that associates the person ID in each location gallery with the same person; an identification means that identifies the person in the person image captured by the imaging unit at the location by referring to the location gallery; and a tracking means that associates the person identified at each location with the person identified at each location based on the correspondence data and tracks the movement of the person between locations.

[0020] Therefore, it is possible to track individuals across multiple locations without forming a single, large gallery.

[0021] (7) The tracking system according to this invention comprises a location person identification device and an integration device, wherein the location person identification device comprises the imaging unit, the location gallery forming means, and the identification means, and the integration device comprises the person ID integration means and the tracking means.

[0022] Therefore, a system can be constructed using a location-based person identification device and an integration device.

[0023] (8) The tracking system according to the present invention further comprises a location person identification device, which includes a target person image acquisition means for acquiring an overall image from the imaging unit, a target person image extraction means for extracting a target person image from the overall image, and an additional judgment information acquisition means for estimating the imaging distance between the target person and the imaging unit based on the overall image, the target person image, or both, and using this as additional judgment information, wherein the identification means identifies the target person by considering not only the similarity between the target person image and the person images in the gallery, but also the influence based on the additional judgment information.

[0024] Therefore, it becomes possible to identify and track individuals more accurately.

[0025] (11) The tracking system according to this invention is characterized in that the person ID integration means determines whether each person registered in a location gallery is the same person as a person registered in another location gallery for each location gallery, and generates corresponding data.

[0026] Therefore, it is possible to generate correspondence data with high accuracy.

[0027] In this embodiment, step S2 corresponds to the "means for acquiring the image of the target person."

[0028] In this embodiment, step S6 corresponds to the "means for acquiring additional judgment information".

[0029] In this embodiment, the "determination means" corresponds to step S10.

[0030] In this embodiment, step S24 corresponds to the "point gallery formation means".

[0031] In the embodiment, step S52 corresponds to the "person ID integration means".

[0032] In this embodiment, the "identification means" corresponds to step S23.

[0033] In this embodiment, step S53 corresponds to the "tracking means".

[0034] The term "device" is a concept that includes not only devices composed of a single computer, but also devices composed of multiple computers connected via a network or the like. Therefore, if the means (or even a part of the means) of the present invention are distributed across multiple computers, these multiple computers constitute the device.

[0035] The term "program" is a concept that includes not only programs that can be directly executed by the CPU, but also source code programs, compressed programs, encrypted programs, and programs that work in conjunction with the operating system to perform their functions. [Brief explanation of the drawing]

[0036] [Figure 1] This is a functional configuration diagram of a person identification device according to one embodiment of the present invention. [Figure 2] This is the hardware configuration of a person identification device. [Figure 3] This is a diagram showing feature vectors. [Figure 4] Display a database of individuals. [Figure 5] This diagram illustrates the features in a person database. [Figure 6] This is a flowchart of a person identification program. [Figure 7] This is a flowchart of a person identification program. [Figure 8] This is an example of an image captured. [Figure 9] This is an example of extracting a person's image from a captured image. [Figure 10] This is a diagram illustrating the learning process of a random forest. [Figure 11] This figure shows the relationship between the similarity of unknown and known individuals and the imaging distance. [Figure 12] This is the result of determining whether each person image is known or unknown. [Figure 13] This is an example of determining known / unknown information by focusing on a single individual. [Figure 14] This is a functional configuration diagram of the tracking system according to the second embodiment. [Figure 15] This is the system configuration. [Figure 16] This is an example of the arrangement of point devices T1 to Tn. [Figure 17] This is the hardware configuration of the location device T. [Figure 18] This is the hardware configuration of the tracking device S. [Figure 19] This is a flowchart of the tracking process. [Figure 20] This is a description of the generated random forest. [Figure 21] This is a diagram to explain the learning process of a random forest. [Figure 22] This is a detailed flowchart for generating corresponding data. [Figure 23] This is the result of the estimation that the same person was identified at two different locations. [Figure 24] This is an example of the correspondence data obtained through integration. [Figure 25] This figure shows the configuration of a person identification device in another example. [Figure 26] This is the functional configuration of a person identification device according to another embodiment. [Figure 27] This is the hardware configuration of a person identification device. [Figure 28] This is a diagram illustrating the trained deep metric learning model FlipReID. [Figure 29] This is a person's image for registration in the Person Database 38. [Figure 30A] This is a feature vector recorded in the person database 38. [Figure 30B] This is a feature vector recorded in the person database 38. [Figure 31] This is a flowchart of the person identification program 36. [Figure 32] Figure 7A shows the captured image, and Figure 7B shows the person extracted using YOLO. [Figure 33] Figure 8A shows the subject's image, and Figure 8B shows the image with a* reduced. [Figure 34] This diagram shows the grouping of human images used for pre-analysis, based on whether or not accuracy improved when a* was reduced. [Figure 35] This figure shows the images of people from groups 2 and 3 plotted in a multidimensional space based on their image features. [Figure 36] This figure shows the number of matches when k is varied in the preliminary analysis. [Figure 37] This is the functional configuration of the person identification device according to the second embodiment. [Figure 38] This is a diagram showing the person database 38. [Figure 39] This is a flowchart of the person identification program 36. [Figure 40] This diagram shows the state of person recognition and possession recognition in captured images. [Modes for carrying out the invention]

[0037] 1. First Embodiment 1.1 Functional Configuration Figure 1 shows a functional configuration diagram of a person identification device according to one embodiment of the present invention. The person DB 12 stores feature vectors 14 of images of multiple people.

[0038] The imaging unit 2 captures the target person and outputs it as a whole image. The whole image acquisition means 4 acquires the whole image that includes the target person. The target person image extraction means 6 extracts the target person's image from the whole image. The additional information acquisition means 8 estimates the distance between the imaging unit 2 and the target person, i.e., the imaging distance, based on the whole image.

[0039] The determination means 10 obtains the similarity between the person in the person image and the most similar person registered in the person database 12. Furthermore, the determination means 10 obtains the imaging distance (additional determination information) of the person image and determines whether or not the person in the person image is a person registered in the person database 12 based on the similarity and imaging distance.

[0040] If the person in the image is already registered, the person in the image will be identified as the most similar person in the person database 12. If the person is not already registered, they will be newly registered in the person database 12.

[0041] As described above, in this embodiment, not only similarity but also imaging distance is considered to determine whether the subject is a person registered in the person database 12. Therefore, it is possible to determine more accurately whether the person is known or unknown.

[0042] 1.2 Hardware Configuration Figure 2 shows the hardware configuration of the person identification device. The CPU 20 is connected to memory 22, display 24, SSD 26, DVD-ROM drive 28, keyboard / mouse 30, and communication circuit 32.

[0043] The communication circuit 32 is for acquiring images captured by the camera 42. The camera 42 is installed, for example, on a university campus or on a street to capture images of people. Therefore, the camera 42 can capture images of students, professors, and other people moving around on a university campus or on a street corner.

[0044] SSD26 contains the operating system 34, a person identification program 36, and a person database 38. The person identification program 36 works in conjunction with the operating system 34 to perform its functions. These programs were originally recorded on DVD-ROM 40 and installed on SSD26 via DVD-ROM drive 28. Alternatively, they could have been installed from a server device on the internet.

[0045] 1.3 Person Identification Process The following describes the person identification process using the person identification device. In this embodiment, multiple people are registered in a person database in advance, and the system determines which of the pre-registered people the person in the target person image is.

[0046] Therefore, in this embodiment, it is necessary to record the person database 38, which is a gallery, on the SSD 26 in advance. First, let's describe this person database 38. (1) Person DB38 Person DB38 is a database containing numerous feature vectors of images of multiple known individuals who are the target of identification. During person identification, the feature vector of an unknown person's image is calculated, and it is determined which of the person's feature vectors registered in Person DB38 is closest to it, which is then used to identify the person.

[0047] In this embodiment, to construct the person DB38, a deep distance learning model is used in which multiple feature quantities of images of a large number of people (people other than the person being identified above) (including images from multiple angles and poses for each person) are trained to not approximate images of different people, but to approximate images of the same person. For example, FlipReID can be used.

[0048] Even if the images are of the same person, differences in angles (front, back, side, etc.) and posture (head up, head down, etc.) will result in different features (edges, etc.) extracted from the images. Deep metric learning learns that the feature vector, which is represented by multiple features, approximates the same person but does not approximate different people.

[0049] Figure 3 schematically illustrates this. When feature vectors, which are represented by multiple feature quantities, are placed in a multidimensional space (shown as a two-dimensional space in Figure 3 for simplicity), the system learns to select features so that the feature vectors of the same person α (person β) are close together, and learns the weights of each feature quantity.

[0050] Therefore, when the FlipReID deep distance learning model, which has been trained in this manner, is given images of multiple known individuals to be identified, as shown in Figure 4, it will output an approximate feature vector and similarity score for each individual. The person database 38 is a record of these feature vectors for each image, associated with the individuals. Figure 5 shows an example of the person database 38. Figure 3, mentioned above, is a plot of this database in a multidimensional space (shown as a two-dimensional space in the figure).

[0051] The person database 38 contains many registered individuals, each assigned a person ID. When a person image is input, its feature vector is calculated, and the person ID with the feature vector most similar to the input feature vector is output along with its similarity score.

[0052] (2) Person identification process The flowchart for the person identification program 36 is shown in Figures 6 and 7. The CPU 20 acquires the image captured by the camera 42 (step S1). Figure 8 shows an example of the acquired image. The CPU 20 extracts people from this image (step S2). This extraction process can be performed using a trained object detection model (for example, YOLO). Figure 9 shows the result of person extraction. In this example, four people are enclosed in frames and their images have been extracted.

[0053] The CPU 20 provides the extracted person image to the trained deep distance learning model FlipReID and calculates a feature vector (step S4). Furthermore, it finds the feature vector closest to this person image's feature vector from the person DB 38. That is, as shown in Figure 3, in a multidimensional space composed of multiple features, it finds the feature vector closest to the extracted person image's feature vector from the person DB 38.

[0054] The CPU20 also obtains the similarity score at this time. If this similarity score is higher than a predetermined value (i.e., it is similar), it can be identified that the person depicted in the extracted person image is the person with the closest feature vector. If it is lower than the predetermined value (i.e., it is not similar), it can be determined that the person is not registered in the person database38.

[0055] However, in this embodiment, the determination of whether a person is known or unknown is not made simply by making the judgment described above. The determination of whether a person is known or unknown is made through the following process.

[0056] The CPU 20 calculates the imaging distance at the time the extracted person image was captured as additional judgment information (step S6). For example, the CPU 20 calculates the distance from the camera to the person based on the on-screen coordinates (foot position) of the center point 50 of the lower edge of the detection frame shown in Figure 9. This can be achieved, for example, by preparing a correspondence table between on-screen coordinates and the distance from the camera to the person.

[0057] Next, CPU20 provides the pre-trained random forest with the similarity score and distance from the camera to estimate whether the person in the person image is known or unknown (step S7).

[0058] The random forest used here is trained using the similarity between images of people registered in the aforementioned person database 38 (known) and people not registered in the database (unknown), and the distance from the camera at the time the image was captured, as training data.

[0059] Figure 10 shows the learning process of Random Forest 60. The training / ground truth data 64 consists of similarity and imaging distance data for a large number of known and unknown person images. Ground truth data indicating whether the person is known or unknown is provided for each of these training data points.

[0060] The learning program 62 modifies the parameters of the random forest 60 by referring to the discrepancies between the known and unknown estimation results obtained by providing the training data to the random forest 60 and the ground truth data.

[0061] Using the random forest trained in the manner described above, it will be possible to record whether a person in a person image is registered in the person database 38 (known) or not (unknown).

[0062] In this embodiment, the imaging distance is also taken into consideration when determining whether a person is known or unknown for the following reasons. Figure 11 shows the relationship between the similarity distance (the closer the similarity distance, the higher the similarity) when an unknown person is input, and the similarity distance and imaging distance when a known person is input. As is clear from this figure, the similarity required to distinguish between known and unknown people changes depending on the imaging distance. Therefore, in this embodiment, the imaging distance is also taken into consideration when determining whether a person is known or unknown.

[0063] The CPU 20 repeats the above process for all the person images extracted in step S2, from step S4 onward, and records known and unknown information as shown in Figure 12. In step S2, the CPU 20 assigns a tracking ID to each extracted person image.

[0064] Next, the CPU 20 acquires the image captured in the next frame (step S1). Furthermore, it extracts the person image from this captured image (step S2). In this embodiment, the same tracking ID is assigned to the same person based on the position and image characteristics of the person image in the previous frame, and tracking is performed accordingly.

[0065] By repeating the above process for consecutively captured frames, the known / unknown determination for each frame will be recorded for individuals assigned the same tracking ID.

[0066] When a person with the same tracking ID has passed by the camera after being captured by the camera (step S9), the CPU 20 obtains a known or unknown determination for that person with the tracking ID, as shown in Figure 13. The CPU 20 counts the number of frames in which the person is determined to be known or unknown, and determines whether the person is known or unknown by majority vote (step S10). In the example in Figure 13, since there are more known frames, it is determined that the person is known.

[0067] If the CPU determines that the person is known, it identifies them as the person with the highest similarity determined in step S5 and outputs the person ID (step S12). If the person with the highest similarity differs in each frame, it is preferable to decide by majority vote.

[0068] If the CPU determines that the person is unknown, it registers the person's image (the number of images corresponding to the number of frames) in the person DB38 (step S14). In other words, it learns FlipReID as a new person. As a result, the person in the person image is newly registered in the person DB38, and from the next time onward, it can be identified as a known person.

[0069] As described above, in this embodiment, the determination of whether a person is known or unknown is made by considering not only the similarity but also the imaging distance from the camera. Therefore, it is possible to determine whether a person is known or unknown with greater accuracy.

[0070] 1.4 Other Variations (1) In the above embodiment, images are taken using cameras installed on the university campus or on the streets. However, images may also be taken using cameras installed inside buildings or facilities.

[0071] (2) In the above embodiment, FlipReID is used for person identification. However, a deep metric learning model other than FlipReID may be used. Alternatively, a deep learning model other than a deep metric learning model may be used. Or, a person may be identified using a logical structure based on the person's features captured in the image.

[0072] (3) In the above embodiment, a random forest is used to estimate known and unknown. However, other deep learning models may be used. Alternatively, the determination of known and unknown may be made using a logical structure.

[0073] (4) In the above embodiment, the imaging distance is estimated based on the position of the center point of the lower edge of the detection frame on the image. However, it may also be estimated based on the position of other points. Alternatively, the distance from the camera to the person may be measured by radar ranging or the like and used.

[0074] (5) In the above embodiment, the imaging distance is used as additional judgment information to determine whether the image is known or unknown. However, any of the following features may be used instead of, or in addition to, this for determining whether the image is known or unknown.

[0075] i) Size of the person's image (height, width) ii) Aspect ratio of the person's image (height ÷ width) iii) Average brightness of the person's image iv) Average value of a* in the person image v) Average value of b* in the person image Furthermore, since the size of the human image is roughly proportional to the imaging distance, it can be used as a substitute for the imaging distance.

[0076] Here, the average value of a* is the average of the a* values ​​of each pixel in the CIELAB color space across the entire image of the person. The average value of b* is the average of the b* values ​​of each pixel in the CIELAB color space across the entire image of the person.

[0077] (6) In the above embodiment, the determination of known or unknown is made by integrating the decisions in multiple frames. However, the determination in a single frame may be used as is.

[0078] (7) In the above embodiment, FlipReID is not trained on human images that are determined to be known. However, human images that are determined to be known may also be used to train FlipReID.

[0079] (8) In the above embodiment, when identifying a person's image, the determination is made based on all people in the person DB38. However, it is clear that the multiple person images extracted from the captured image acquired in step S1 are of different people. Therefore, the identification may be performed on the premise that the multiple person images extracted from the same captured image are not of the same person.

[0080] For example, if multiple images of people extracted from the same captured image are presumed to be of the same person, it is preferable to adopt the presumption with the highest similarity and, for the presumptions with low similarity, to presume that the person is the second most similar.

[0081] (9) In the above embodiment, known and unknown estimations are made using a random forest set up at each location. However, even at the same location, random forests may be set up according to changes in the captured image, such as a random forest for the morning and a random forest for the evening.

[0082] (10) In the above embodiment, a device for identifying a person was described. However, a similar method can also be applied to a device for identifying animals, robots, machines, etc.

[0083] (11) The above modifications can be combined with each other. Furthermore, the above embodiments and their modifications can be implemented in combination with other embodiments and their modifications.

[0084] 2. Second Embodiment 2.1 Functional Configuration Figure 14 shows a functional configuration diagram of the tracking system according to the second embodiment. In this embodiment, point devices T1, T2, etc. are provided at each location, and a tracking device S is provided.

[0085] The imaging unit 2 of the location device T1 captures an image of the target person and provides it to the location gallery forming means 74. Based on the captured images of the person, the location gallery forming means 74 forms a location gallery 78 in which the same person is assigned the same person ID.

[0086] The identification means 76 receives the captured image of a person, refers to the location gallery 78, identifies the captured person, and outputs a person ID.

[0087] The location device T2... performs the same processing as described above. However, since the location gallery 78 is formed for each location, they are not interchangeable.

[0088] The person integration means 84 of the tracking device S extracts the same person from the location gallery 78 of each location device T1, T2, etc., and generates correspondence data that associates each person ID. The tracking means 82 obtains the identified person ID from each location device T1, T2, etc., and tracks the movement of the same person between locations by referring to the correspondence data.

[0089] In this embodiment, correspondence data is generated to associate the same person with each location gallery formed at each point. Therefore, it is possible to track the movement of the same person at different locations with a simple configuration.

[0090] 2.2 System Configuration and Hardware Configuration Figure 15 shows the system configuration of the tracking system. Location devices T1, T2...Tn are installed at their respective locations. These location devices T1, T2...Tn are connected to the internet. The integrated tracking device S is also connected to the internet and is configured to communicate with the location devices T1, T2...Tn via the internet.

[0091] Figure 16 shows point devices T1, T2...Tn placed on street corners. Each point device T1, T2...Tn is equipped with a camera capable of capturing images of passersby. Note that each point only needs to be equipped with at least one camera, and the main body of the point device may be located in a different place.

[0092] Figure 17 shows the hardware configuration of location device T. Note that location devices T1, T2, ..., Tn have similar hardware configurations and will be described as location device T.

[0093] The CPU 20 is connected to memory 22, display 24, SSD 26, DVD-ROM drive 28, keyboard / mouse 30, communication circuit 32, and camera 42, which is the imaging unit.

[0094] In this embodiment, the camera 42 is connected without a communication circuit, but it may also be connected via a communication circuit, as in the first embodiment. The camera 42 is installed, for example, on a university campus or on a street to capture images of people. Therefore, the camera 42 can capture images of students, professors, or other people moving around on a university campus or on a street corner.

[0095] The communication circuit 32 is for connecting to the internet. The SSD 26 stores the operating system 34, a person identification program 36, and a person database 38, which is a location gallery. The person identification program 36 works in cooperation with the operating system 34 to perform its functions. These programs were originally recorded on a DVD-ROM 40 and installed on the SSD 26 via the DVD-ROM drive 28. Alternatively, they could have been installed from a server device on the internet.

[0096] Figure 18 shows the hardware configuration of the tracking device S. The CPU 120 is connected to the memory 122, SSD 126, DVD-ROM drive 128, and communication circuit 132.

[0097] The communication circuit 132 is for connecting to the internet. The tracking device S can communicate with the location device T via the communication circuit 132.

[0098] SSD126 contains the operating system 134 and the tracking program 136. The tracking program 136 works in conjunction with the operating system 134 to perform its functions. These programs were originally recorded on DVD-ROM 140 and installed on SSD126 via DVD-ROM drive 128. Alternatively, they could have been installed from a server device on the internet.

[0099] 2.3 Person Tracking Process Figure 19 shows a flowchart of the person tracking process. It shows the flowchart of the person identification program 36 of the location device and the flowchart of the tracking program 136 of the tracking device.

[0100] The CPU 20 of the location device T1 (hereinafter sometimes abbreviated as location device T1) acquires captured images from the camera 42 (step S21). From these captured images, a person image is extracted using YOLO or the like, as shown in Figure 9 (step S22).

[0101] Next, the location device T1 refers to the person DB38 and identifies the person depicted in each extracted person image (step S23). This process can be performed using, for example, the process in the first embodiment. In this embodiment, the person DB38 also records the imaging distance at the time the person image was captured, along with the feature vector of the person image for each person.

[0102] Next, if the location device T1 determines that the person image is of a new person (an unknown person), it processes the image to register it as a new person in the person database 38 (step S24). Even if the person in the person image has been identified, the person database 38 is further trained based on that person image (step S24).

[0103] The update process for these person databases (DB38) may be performed all at once after a predetermined number of person images have been accumulated.

[0104] Subsequently, the location device T1 repeatedly executes steps S21 to S24 to identify the imaged person and update the person database 38. In this way, the person database 38 of the location device T1 is updated and built up to be specific to that location.

[0105] The above describes the process for location device T1, but the other location devices T2 to Tn will similarly identify the person being imaged and update the person database 38. Therefore, a person database 38 specific to each location will be formed for location devices T1 to Tn.

[0106] The CPU 120 of the tracking device S (hereinafter sometimes abbreviated as tracking device S) acquires the respective person DB 38 from each location device T1 to Tn (step S51). Based on the acquired person DB 38 from each location, the tracking device S generates correspondence data that shows the correspondence between the people in each person DB 38 (step S52).

[0107] In this embodiment, in order to generate the corresponding data described above, a random forest for determining the identity of the same person among the person databases 38 at each location is pre-trained and formed, as shown in Figure 20. Figure 20 illustrates the case where there are three locations.

[0108] A random forest RF12 is provided to determine whether a person registered in the person database 38 at location 1 is known or unknown in the person database 38 at location 2. In other words, the person at location 1 is used as a query, and the person at location 2 is used as a gallery to determine whether they are known or unknown.

[0109] Conversely to the above, a random forest RF12 is provided to determine whether a person registered in the person database 38 at location 2 is known or unknown to a person registered in the person database 38 at location 1. In other words, it is a random forest RF21 that uses the people at location 2 as a query and the people at location 3 as a gallery to determine whether they are known or unknown.

[0110] The same random forests described above are also provided for location 1 and location 3, and location 2 and location 3. Therefore, in this embodiment, six random forests RF12, 12, 13, 31, 23, and 32 are provided. Note that if there are n locations, n P n-1 This will result in the creation of several random forests.

[0111] For example, the generation of the Random Forest RF12 is performed as shown in Figure 21. The first data (feature vector and imaging distance) of the first person (let's call it ID=1001) in the person DB38 at location 1 is used as training data. This feature vector is then given to the person DB32 at location 2, and the feature vector with the highest similarity (minimum similarity distance) is calculated. This similarity is then used as training data. Furthermore, data indicating whether or not person 1001 is registered in the person DB38 at location 2 (known or unknown) is used as ground truth data.

[0112] The learning program provides the feature vector, imaging distance, and similarity of person 1001 to the random forest RF12 and performs known / unknown estimation. That is, it estimates whether person 1001 is known or unknown in the person DB38 at location 2.

[0113] The learning program learns the parameters of the random forest RF12 based on the error between the estimated results and the ground truth data.

[0114] Similarly to the above, training is performed based on the next imaging distance, similarity, and known / unknown data for person (ID=1001). Once training is complete for all feature vectors of person (ID=1001), training is then performed similarly based on the data for person (ID=1002).

[0115] In this way, a trained random forest RF12 can be obtained.

[0116] Similarly, random forests RF21, RF13, RF31, RF23, and RF32 can be trained and generated in the same manner.

[0117] Figure 22 shows the details of the data generation process for the corresponding data in Figure 19 (step S52). Based on the person DB38 acquired from each point device T1 to Tn, the tracking device S uses the trained random forest RF described above to estimate whether all people registered in the person DB38 at a given point are known or unknown in the person DBs at other points (step S72).

[0118] For example, for all individuals registered in the person database 38 of location device T1, it is estimated whether they are known or unknown in the person database 38 of location device T2. In this case, a random forest RF12 is used. The random forest RF12 is given the imaging distance of the person image (feature vector) registered in the person database of location device T1 and the similarity (minimum similarity distance) when that feature vector is given to location device T2, and it is estimated whether they are known or unknown. If they are known, it is estimated that the person registered in location device T1 and the person registered in location device T2 are the same person.

[0119] If multiple feature vectors (i.e., person images) of the same person are registered in the person DB38 of location device T1, the known / unknown judgments for each feature vector are aggregated, and the determination of whether it is known or unknown is made by majority vote.

[0120] The tracking device S performs this for all combinations of point devices (steps S71, S73). This allows a correspondence between each point to be obtained as shown in Figure 23.

[0121] Random Forest RF12 estimation indicates that Person 1001 in Person DB38 at location 1 (location device T1) is the same person as Person 2004 in Person DB38 at location 2 (location device T2). Similarly, Person 1003 and Person 2014 are shown to be the same person.

[0122] The tracking device S generates correspondence data as shown in Figure 24 based on a table of identical individuals as shown in Figure 23 (step S74). The tracking device S determines that two individuals are the same person if it is determined that they are the same person in relation to one person DB 38 from the other person DB 38, and vice versa.

[0123] For example, in Figure 23, Random Forest RF12 determines that individuals 1003 and 2014 are the same person. Random Forest RF21 also determines that individuals 2014 and 1003 are the same person. Therefore, individuals 1003 and 2014 are considered the same person from both perspectives. Consequently, in the corresponding data in Figure 24, 1003 and 2014 are shown as the same person.

[0124] On the other hand, while Random Forest RF12 identifies individuals 1002 and 2004 as the same person, Random Forest RF21 does not. Therefore, in the corresponding data in Figure 24, 1002 and 2004 are not treated as the same person.

[0125] Similarly, in Random Forest RF23, individuals 2003 and 3046 are determined to be the same person. On the other hand, in Random Forest RF32, individuals 3046 and 2002 are determined to be the same person. In this case as well, the corresponding data in Figure 24 does not treat individuals 2003 and 3046, or individuals 3046 and 2002, as the same person.

[0126] In this way, the correspondence between each person registered in the person database 38 at each location can be obtained. In step S53 of Figure 19, the tracking device S can refer to the correspondence data and track the movement of the same person between locations.

[0127] Furthermore, the person database 38 at each location device T1, T2...Tn is updated as needed. Therefore, it is preferable that the corresponding data generated by the tracking device S be updated and regenerated at the shortest possible time intervals (for example, every hour).

[0128] 2.4 Other variations (1) In the above embodiment, the imaging distance is used to determine whether a person in the person DB38 is known or unknown. However, the determination may be made based solely on similarity (minimum similarity distance) without using the imaging distance.

[0129] (2) In the above embodiment, FlipReID is used as the person database. However, a deep metric learning model other than FlipReID may be used. Alternatively, a deep learning model other than a deep metric learning model may be used. Or, a person may be identified using a logical structure based on the person's features captured in the image.

[0130] (3) In the above embodiment, the tracking system is constructed using location devices T1, T2...T3 and a tracking device S. However, as shown in Figure 25, an imaging unit (camera) 72 may be placed at each location, and the tracking device S may be used to form a person database 38 for each location.

[0131] (4) The above modifications can be combined with each other. Furthermore, the above embodiments and their modifications can be implemented in combination with other embodiments and their modifications.

[0132] 3. Third Embodiment 3.1 Functional Configuration Figure 26 shows a functional configuration diagram of the person identification device according to the third embodiment. The recording unit 212 has feature vectors 214 of person images of multiple people to be identified pre-recorded.

[0133] The target person image acquisition means 202 acquires a target person image. The feature acquisition means 204 analyzes and acquires the image features of the target person image. The change determination means 208 determines, based on the image features, whether it is better to change at least one element in the color space of the target person image in order to improve the identification accuracy of the identification means 210.

[0134] If the change determination means 208 determines that a change should be made, the identification image determination means 206 changes at least one element in the color space of the target person image to generate an identification target person image. If the change determination means 208 determines that no change should be made, the target person image remains unchanged and is used as the identification target person image.

[0135] The identification means 210 calculates the feature vector of the target person image for identification and compares it with the feature vector of multiple person images that are the target of identification recorded in the recording unit to identify the person in the target person image.

[0136] In this embodiment, it is determined whether changing at least one element in the color space of the target person image improves identification accuracy, and if the identification accuracy improves, the element is changed. Therefore, even if the target person image was captured under different circumstances, the person can be identified with high accuracy.

[0137] 3.2 Hardware Configuration Figure 27 shows the hardware configuration of the person identification device. The CPU 220 is connected to the memory 222, display 224, SSD 226, DVD-ROM 228, keyboard / mouse 230, and communication circuit 232.

[0138] The communication circuit 232 is a circuit for connecting to the internet. In this example, cameras 242 installed inside and outside the building are connected via the internet. The CPU 220 can acquire images captured by the cameras 242 via the internet using the communication circuit 232.

[0139] SSD226 contains the operating system 234, the person identification program 236, and the person database 238. The person identification program 236 works in conjunction with the operating system 234 to perform its functions. These programs were originally recorded on DVD-ROM 240 and installed on SSD226 via DVD-ROM drive 228.

[0140] The person database 238 contains pre-registered feature vectors of person images of multiple people to be identified. The person identification program 236 calculates the feature vector of the target person image acquired from the camera 242, compares it with the feature vector in the person database 238, and identifies which person it is.

[0141] 3.3 Person Identification Process The following describes the person identification process using the person identification device. In this embodiment, multiple people are registered in a person database in advance, and the system determines which of the pre-registered people the person in the target person image is.

[0142] Therefore, in this embodiment, it is necessary to record the person database 238 on the SSD 226 in advance. First, let's describe this person database 238.

[0143] (1) Person DB238 Person DB238 is a database containing numerous feature vectors of images of multiple known individuals who are the target of identification. During person identification, the feature vector of an unknown person's image is calculated, and it is determined which of the person's feature vectors registered in Person DB238 is closest to it, which is then used to identify the person.

[0144] In this embodiment, to construct the person DB238, a deep distance learning model is used that trains multiple feature quantities of images of a large number of people (people other than the person being identified above) (including images from multiple angles and poses for each person) so that images of the same person do not approximate images of different people, but do approximate images of the same person. For example, FlipReID can be used.

[0145] Even if the images are of the same person, differences in angles (front, back, side, etc.) and posture (head up, head down, etc.) will result in different features (edges, etc.) extracted from the images. Deep metric learning learns that the feature vector, which is represented by multiple features, approximates the same person but does not approximate different people.

[0146] Figure 28 schematically illustrates this. When the feature vectors, which are represented by multiple feature quantities, are placed in a multidimensional space (shown as a two-dimensional space in Figure 28 for simplicity), the feature vectors of person α (person β) are trained to be close together.

[0147] Therefore, when the FlipReID deep distance learning model, which has been trained in this manner, is given images of multiple known individuals to be identified, as shown in Figure 29, it will output an approximate feature vector for each individual. The person database 238 is a record of these feature vectors for each image, associated with the individuals. An example of the person database 238 is shown in Figure 30A. Figure 30B shows this plotted in a multidimensional space (shown as a two-dimensional space in the figure).

[0148] (2) Person identification process Figure 31 shows a flowchart of the person identification program 236. The CPU 220 acquires the images captured by the camera 242 (step S201). While the images from the camera 242 may be acquired individually each time, all images from the camera 242 may be acquired together, recorded to the SSD 226, and then read sequentially from the SSD 226. Here, we will explain assuming that images like those shown in Figure 32A have been acquired.

[0149] The CPU220 extracts the person from the captured image and surrounds it with a boundary box BX (step S2). The area surrounded by the boundary box BX is extracted as the target person image. Figure 33A shows an example of the extracted person image.

[0150] Furthermore, to extract people from captured images, a deep learning model trained to extract various objects, including people, can be used. For example, a deep learning model with the YOLO algorithm, trained on images of many different types of objects, can be used. In this case, the training data consists of the captured image and the coordinate values ​​(coordinate positions in the captured image) of the boundary box surrounding the object (e.g., a person) depicted in the image, with the type of object (e.g., a person) attached as the ground truth data.

[0151] Next, CPU220 decides whether to perform identification using the extracted person images as they are, or to change the image quality of the extracted person images before performing identification.

[0152] Here, we will first explain the process of performing identification using the extracted person images as they are.

[0153] The CPU 220 provides the extracted person image (target person image) to the trained deep distance learning model FlipReID, which was used to generate the person DB 238, and calculates a feature vector (step S9). The CPU 220 finds the feature vector closest to the feature vector of this person image from the person DB 238 (see Figure 30). That is, as shown in Figure 30B, in a multidimensional space composed of multiple features, the CPU 220 finds the feature vector closest to the feature vector U of the target person image from the person DB 238. In the example of Figure 30B, the feature vector closest to feature vector U is person 202, so the person in the target person image is identified as person 202.

[0154] In this way, the person depicted in the target person image can be identified in the person DB238 as the person associated with the feature vector (step S210).

[0155] As described above, CPU220 identifies individuals by referring to the person database238.

[0156] As mentioned above, before performing person identification, CPU220 decides whether to use the extracted person image (target person image) as is for identification, or to modify the image quality of the extracted person image (target person image) (change elements in the color space) before identification. CPU220 performs this decision based on the image characteristics of the target person image. The process is explained below.

[0157] First, the CPU 220 calculates the image features of the target person image (Figure 33A) (step S204). In this embodiment, the average brightness of the person image, the median brightness, and the average a * a * Median, mean b * , b * The system calculates nine image features: the median, image width, image height, and sharpness.

[0158] The average brightness is the average value of the brightness of each pixel over the entire image. The median brightness is the median value of the brightness of each pixel over the entire image. Average a * is the a in the CIELAB color space of each pixel * averaged over the entire image. The median of a * is the median value of a for each pixel * over the entire image. Average b * is the b in the CIELAB color space of each pixel * averaged over the entire image. The median of b * is the median value of b for each pixel * over the entire image. The horizontal width of the image is the number of pixels in the horizontal direction of the image. The vertical height of the image is the number of pixels in the vertical direction of the image. The sharpness is obtained by grayscale converting the image, applying a Laplacian filter to this, and calculating the variance of the pixel values. If the variance is large, it can be said that there are many edge parts and the image is sharp.

[0159] Subsequently, it is advisable for the CPU 220 to calculate the feature vector using the person image as it is and perform identification by referring to the person DB 238, or to calculate the feature vector using an image (an image with the image quality changed) in which the a * of the person image is overall decreased (10% decrease in this example) and perform identification by referring to the person DB 238 to make a determination (step S5).

[0160] This determination is made using the k-nearest neighbor algorithm (k-nn) based on data prepared in advance as follows.

[0161] First, prepare person images of a plurality of persons for pre-analysis (images of persons different from the persons registered in the above person DB 38). Apply these plurality of person images to the learned deep distance learning model FlipReID to generate a person DB for the persons for pre-analysis.

[0162] Next, a large number of images of individuals for pre-analysis (images with different angles and poses than the registration images mentioned above) are prepared and fed into the trained deep distance learning model FlipReID to calculate feature vectors. The individuals are then identified by referring to the person database. Since all of the individuals for pre-analysis are known individuals, it is possible to determine whether the identification result is correct or not. That is, the images can be divided into those that were successfully identified and those that were not.

[0163] Next, regarding the subject images for preliminary analysis, a * The overall image quality is reduced (by 10% in this example) to obtain an image of a person for identification. This image of a person for identification is fed into a pre-trained deep distance learning model, FlipReID, to calculate feature vectors, and the person is identified by referring to a person database. In this case as well, the images of people for identification can be divided into those that were successfully identified and those that were not.

[0164] Therefore, a * The person's image that was not changed, and a * When the images of people whose quality was reduced are presented in a table of successes and failures, they can be classified into four groups as shown in Figure 34.

[0165] In this embodiment, a * Images of people in group 2, for which a decrease in is deemed undesirable for improving identification accuracy, and a * Using the person images of Group 3, for which it is deemed preferable to reduce the brightness of each person image to improve identification accuracy, the respective image features (average brightness of the person image, median brightness, average a) were used. * a * Median, mean b * , b * Calculate the median, image width, image height, and sharpness.

[0166] Then, the images of each person from Group 2 and Group 3 are placed in a multidimensional space based on these image features. Figure 35A shows the placed images. In this example, there are nine image features, so the placement would be in a nine-dimensional space, but for simplicity, Figure 35A displays it in a two-dimensional space.

[0167] Figure 35A shows that, based on the image features, a * This plots whether or not the value should be reduced. Therefore, by examining which group the image features of the unknown person image (Figure 33A) calculated in step S204 are closest to, a * It is possible to decide whether or not to lower it.

[0168] In this embodiment, k=3 and the K-nn method is used for determination. Figure 35B shows plot T of image features of a person image where the person is unknown. Since k=3, three plots closest to T are selected. The selected plots are in group 2(a * Two groups (a) should not be reduced. * Since there is only one that should be reduced, by majority vote, a * It can be decided that it should not be reduced.

[0169] In this embodiment, k is set to an odd number, but if k is even, it may not be possible to decide by majority vote. In that case, it is best to decide based on which group the plot closest to plot T belongs to.

[0170] As described above, in step S205, CPU220 a * Should it be reduced, a * Decide whether or not it should be reduced.

[0171] Next, CPU220 is a * If it is determined that the value should not be reduced, the person image extracted in step S202 (Figure 33A) is used as the person image for identification (step S207). Also, the CPU 220 a* If it is determined that the value should be reduced, the overall value of the person image extracted in step S202 (Figure 33A) should be reduced. * The image quality is reduced and used as the person image for identification (step S208). * Figure 33B shows the identification image of the person obtained by reducing the filter ratio. It is difficult to distinguish in the black and white image, but it has a slightly bluish tint.

[0172] Next, the CPU 220 provides the person image for identification to the trained deep distance learning model FlipReID, which was used when generating the person DB 238, and calculates a feature vector (step S209). The CPU 220 finds the feature vector closest to the feature vector of this person image from the person DB 238 (see Figure 30). That is, as shown in Figure 30B, in a multidimensional space composed of multiple features, the CPU 220 finds the feature vector closest to the feature vector of the target person image from the person DB 238 (step S210). In this way, the person is identified.

[0173] In this embodiment, a * Would generating identification images of people without reducing the quality improve identification accuracy, or * The system determines, based on the image characteristics of the target person's image, whether reducing the image quality to generate an identification image of the person improves identification accuracy. Therefore, the image quality can be changed based on the image characteristics of the target person's image to improve identification accuracy.

[0174] The CPU 220 repeats the processes in steps S204 to S210 described above for the number of people captured in the image acquired in step S201 (steps S203, S211).

[0175] Once all the human images extracted from the captured images have been identified, the CPU 220 performs a tracking process (step S212). That is, it considers the human identification results in images captured at a similar time by other nearby cameras to determine where the identified person has moved to.

[0176] Next, the CPU 220 acquires images from other cameras 242 and repeatedly executes steps S202 and below. Once processing is complete for all images from cameras 242, the process returns to acquiring images from the first camera 242 and repeats.

[0177] In this way, each camera 242 can perform person identification on the captured images taken at each time point.

[0178] (3) Determination of k in the K-nn method Step S205(a) described above * In the determination of whether or not to decrease the value, we explained using k=3. It is preferable to determine the value of k as follows.

[0179] Using known images of people for preliminary analysis, a * When the value is reduced, * Figure 34 shows the results of investigating whether the identification accuracy changes when the value is not reduced (this point has already been mentioned). This group 2 (a * (should not be reduced) and group 3 (a * Regarding the image characteristics (which should be reduced), the average brightness of the human image, median brightness, average a * a * Median, mean b * , b * Figure 35A shows the data plotted on a multidimensional interval (represented as a two-dimensional space in the figure) based on the median of the image, the image width, the image height, and the sharpness.

[0180] As mentioned above, the image features T of an unknown person's image are plotted in this multidimensional space (see Figure 35B), and it is determined which group three adjacent plots (k=3) belong to, and a majority vote is taken. * They are deciding whether or not to lower it.

[0181] The question then arises as to what value of k is appropriate. In this embodiment, the value of k is determined as follows.

[0182] First, we prepared a large number of other images of people whose identities were already known for preliminary analysis, and for each of them, we used the original image of the person and a * Using reduced-resolution images of people, the trained deep distance learning model FlipReID is fed to calculate feature vectors, and people are identified by referring to a person database.

[0183] As a result, these images of people are grouped as shown in Figure 34. Of these, only the images of people belonging to groups 2 and 3 are extracted. The images of people belonging to groups 2 and 3 are a * It is known whether or not it is better to lower it.

[0184] Therefore, for each person's image, the image features are calculated and arranged in a multidimensional space as plot T, as shown in Figure 35B. First, with k=1, the previously known a * We then determine whether it is appropriate to reduce the value. Next, we repeat this process with k=2, k=3, etc.

[0185] Figure 36 shows an example of the results. In this example, the number of matches was recorded when k was varied for 20 human images. The number of matches was highest when k=3, so this value is selected.

[0186] As described above, the value of k is determined in this way, so the optimal value of k may vary depending on the camera 242 and the imaging conditions.

[0187] 3.4 Variations (Other) (1) In the above embodiment, the camera 242 is connected to the person identification device via the Internet. However, it may also be connected via a LAN or a dedicated line. Alternatively, the captured images recorded by the camera 242 may be recorded on a portable recording medium and read by the person identification device.

[0188] (2) In the above embodiment, the person database 238 is recorded in the person identification device. However, the person database 38 may be recorded in another device connected via a network such as the Internet or LAN.

[0189] (3) In the above embodiment, a pre-trained deep learning model using the YOLO algorithm is used for object recognition. However, other object recognition algorithms, such as YOLACT, may also be used.

[0190] In YOLO, as shown in Figure 33A, the target person image includes the background image, and feature vectors and image features are calculated including the background image. In YOLACT, however, a boundary box can be used that follows the outline of the extracted person instead of a rectangle, so feature vectors and image features can be calculated for person images without a background.

[0191] (4) In the above embodiment, the image features include the average brightness of the person image, the median brightness, and the average a * a * Median, mean b * , b * The median, image width, image height, and sharpness are used. However, any combination of these image features may also be used.

[0192] For example, average brightness ÷ image size can be used as an image feature. In this case, the features would be whether it is small and bright. Also, (average b * × sharpness) ÷ average a * This can also be used as an image feature. In this case, the features would be whether it is vivid and has a greenish tint.

[0193] Furthermore, at a minimum, the feature that changes the image quality (a above) * ) related to image features (average a above) * a * It is preferable to use the median of the two.

[0194] (5) In the above embodiment, a* The image quality is changed by reducing a predetermined percentage. However, a * It is also possible to increase b by a predetermined percentage. * Ya L * You may also change these. Furthermore, you may change multiple of these. You may also change R, G, B, etc.

[0195] (6) In the above embodiment, in step S209, the person with the closest feature vector is selected as the person to be identified. However, it is also possible to extract multiple adjacent feature vectors and perform the identification based on which person's features are most numerous (if the number is the same, the person with the closest feature vector is identified).

[0196] (7) In the above embodiment, the person identification device is configured as a standalone PC. However, it may also be configured as a server device. In this case, the identification results will be transmitted to other devices such as terminal devices.

[0197] (8) In the above embodiment, the trained deep metric model FlipReID is given the person image to be identified to calculate the feature vector, which is then registered as the person DB238. However, an untrained deep metric model may be trained using the person image to be identified, the person image to be identified is given to this model to calculate the feature vector, and this is then registered as the person DB238.

[0198] (9) The above modifications can be implemented in combination with each other. They can also be implemented in combination with other embodiments.

[0199] 4. Fourth Embodiment 4.1 Overall configuration Figure 37 shows the functional configuration of the person identification device according to the fourth embodiment. The target person image acquisition means 202 extracts a person image from the captured image. The extracted target person image is provided to the feature acquisition means 204, and a feature vector is calculated.

[0200] The possession information acquisition means 218 extracts possessions from the captured image and estimates whether or not possessions are present or what type of possessions they are.

[0201] The person database 212 records the feature vectors 214 of each image of the person to be identified. If the image at the time of registration includes possessions, the type of those possessions is also recorded.

[0202] The identification means 210 identifies a person by referring to the person database 212 based on the feature vectors of the target person's image and the information of the person's possessions.

[0203] In this embodiment, identification is performed by also taking into account information about the possessions, thereby improving the accuracy of identification.

[0204] 4.2 Hardware Configuration The hardware configuration is the same as in Figure 27 of the third embodiment.

[0205] 4.3 Person Identification Process (1) Person DB238 Person DB238 has the same configuration as in the third embodiment. However, when registering the feature vectors of the person image to be identified, YOLO is used to determine whether the person possesses any possessions, and if so, the type of possession identified by YOLO is also registered.

[0206] An example of the person database 238 generated in this way is shown in Figure 38. If the person's belongings are visible in the image used for registration, the type of belongings, such as "bag" or "backpack," is also recorded.

[0207] (2) Person identification process Figure 39 shows a flowchart of the person identification program 236. The CPU 220 acquires images from the camera 242 (step S201). Furthermore, it extracts a person image from these images using YOLO (step S202). It also estimates the type of belongings the person is carrying using YOLO (step S220).

[0208] Figure 40 shows the extraction status of person images and possessions. YOLO labels the extracted boundary boxes with their type (person, bag, backpack, handbag, etc.), allowing for identification of whether they are person images or possessions. It also allows for estimation of the type of possessions.

[0209] Furthermore, the system determines that an item whose boundary box has the greatest overlap with the person's boundary box belongs to that person.

[0210] The CPU 220 calculates the feature vector of the extracted person image (step S9). Similar to the first embodiment, identification is performed by referring to the person DB 238 based on this feature vector (step S210).

[0211] However, in this embodiment, if the subject person has possessions, the registration data that records the same type of possession is selected from the person DB238. For example, if the subject person has a "bag," only the data in the person DB238 shown in Figure 38 that records "bag" as an possession is selected.

[0212] CPU220 identifies the person whose feature vector is most different from the selected data.

[0213] Thus, in this embodiment, the type of possession is taken into consideration when performing identification, which improves the accuracy of the identification.

[0214] 4.4 Variations (Other) (1) In the above embodiment, data from the person DB238 is selected based on the type of possessions of the person. However, data from the person DB238 may be selected based on whether or not the person possesses possessions, regardless of the type of possessions.

[0215] (2) The identification method that takes into account the possessions according to the above embodiment can be implemented in combination with the first to third embodiments.

[0216] (3) The above modifications can be implemented in combination with each other. They can also be implemented in combination with other embodiments.

Claims

1. An imaging unit that captures an overall image including the person in question, A recording unit that records images of multiple people who are the target of identification as a gallery, A target person image acquisition means that acquires an overall image from the aforementioned imaging unit, A target person image extraction means for extracting a target person's image from the overall image, An additional judgment information acquisition means estimates the imaging distance between the target person and the imaging unit based on the overall image, the target person image, or both, and uses this as additional judgment information. A determination means for determining whether or not the person in question is a person registered in the gallery, taking into account not only the similarity between the image of the person in question and the images of people in the gallery, but also the influence based on the additional judgment information, A person identification device equipped with this device.

2. A person identification program for realizing a person identification device using a computer, wherein the computer A means for acquiring an image of a target person, which acquires the overall image from an imaging unit that captures an overall image including the target person, A target person image extraction means for extracting a target person's image from the overall image, An additional judgment information acquisition means estimates the imaging distance between the target person and the imaging unit based on the overall image, the target person image, or both, and uses this as additional judgment information. A person identification program that functions as a means of determining whether or not a target person is a person registered in the gallery, taking into account not only the similarity between the target person image and the person images in the gallery in which multiple person images to be identified are registered, but also the influence based on the additional judgment information.

3. In the apparatus of claim 1 or the program of claim 2, The additional information acquisition means acquires the aspect ratio of the target person image as additional information, in lieu of or in addition to the imaging distance. The determination means is characterized in that it performs the determination using the aspect ratio of the target person image as additional information, instead of or in addition to the imaging distance.

4. In the apparatus of claim 1 or the program of claim 2, The imaging unit captures multiple images of the whole picture in succession. The determination means is characterized by making a determination for each of the multiple images of a target person that are presumed to be the same person, based on the multiple images of a target person extracted from the multiple overall images, and then integrating these to make a final determination.

5. In the apparatus of claim 1 or the program of claim 2, The device or program is characterized in that, if the determination means determines that the person in the target person image is not a person registered in the gallery, it registers the person as a new person in the gallery.

6. It consists of an imaging unit installed at multiple locations to capture images of the target person, A location gallery formation means that forms a gallery for each location where the same person is assigned the same ID based on the target person image taken at each location, A person ID integration means that acquires gallery data for each location and generates correspondence data that associates the person ID in each gallery for the same person, An identification means for identifying a person in a person image captured by the imaging unit at a given location, by referring to the aforementioned location gallery, A tracking means that tracks the movement of a person between locations by associating them with the person identified at each location based on the aforementioned correspondence data, A tracking system equipped with [features / equipment].

7. In the tracking system of claim 6, The tracking system includes a location-based person identification device and an integration device. The device comprises the aforementioned location person identification device, the aforementioned imaging unit, the aforementioned location gallery forming means, and the aforementioned identification means. The tracking system is characterized in that the integration device comprises the person ID integration means and the tracking means.

8. In the tracking system of claim 7, The aforementioned location person identification device, furthermore, A target person image acquisition means that acquires an overall image from the aforementioned imaging unit, A target person image extraction means for extracting a target person's image from the overall image, The system includes an additional judgment information acquisition means that estimates the imaging distance between the target person and the imaging unit based on the overall image, the target person image, or both, and uses this as additional judgment information. The tracking system is characterized in that the identification means identifies the target person by considering not only the similarity between the target person image and the person images in the gallery, but also the influence based on the additional judgment information.

9. A person ID integration means that acquires a gallery of each location formed based on the person imaged at each location, and generates correspondence data that associates the person ID in each location gallery with the same person, An identification information acquisition means that acquires identification information by identifying a person in a person image captured by the imaging unit at the location by referring to the aforementioned location gallery, A tracking means that tracks the movement of a person between locations by associating them with the person identified at each location based on the aforementioned correspondence data, An integrated device equipped with [the necessary components].

10. An integrated program for realizing an integrated device using a computer, wherein the computer A person ID integration means that acquires a gallery of each location formed based on the person imaged at each location, and generates correspondence data that associates the person ID in each location gallery with the same person, An identification information acquisition means that acquires identification information by identifying a person in a person image captured by the imaging unit at the location by referring to the aforementioned location gallery, An integrated program that, based on the aforementioned correspondence data, associates individuals identified at each location and functions as a tracking means for tracking the movement of those individuals between locations.

11. In the system of claim 6, the apparatus of claim 9, or the program of claim 10, The person ID integration means is a system, apparatus, or program characterized by determining, for each location gallery, whether each person registered in that location gallery is the same person as a person registered in another location gallery, and generating corresponding data.

Citation Information

Patent Citations

  • Information processing system, information processing method, and program

    JP2019186955A