Person identification device, and person identification method

The person identification device and method address the challenge of comprehensive space coverage by integrating position data from multiple cameras using a 3D model and machine learning, enhancing identification and tracking accuracy.

JP2025173540APending Publication Date: 2025-11-28NTT DOCOMO INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024079098
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-05-15
Publication Date
2025-11-28

AI Technical Summary

Technical Problem

Comprehensive coverage of real space using a single camera is difficult due to the angle of view and obstacles, necessitating multiple cameras, which require integration of estimation results from individual images.

Method used

A person identification device and method that utilizes a 3D model of the target space, integrating position information from multiple cameras by determining and comparing positions of individuals across frames and between cameras using machine learning and conversion information to a common coordinate system.

Benefits of technology

Enables accurate identification and tracking of individuals across multiple cameras by integrating position data, improving coverage and accuracy in person identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025173540000001_ABST
    Figure 2025173540000001_ABST
Patent Text Reader

Abstract

To enable integration of estimation results of a position of a person reflected in a captured image of each of a plurality of cameras.SOLUTION: A person identification device 20 comprises a first determination unit 230a, a second determination unit 230b, and an integration unit 230c. The first determination unit 230a determines a first position at time T1 and a second position at time T2, related to a person of a camera 30A. The first determination unit 230a determines a third position at the time T1 and a fourth position at the time T2, related to a person of a camera 30B. The second determination unit 230b determines whether to have supplemented a locus of the position about the person of the camera 30A on the basis of the first position and the second position, and determines whether to have supplemented the locus of the position about the person of the camera 30B on the basis of the third position and the fourth position. The integration unit 230c, when it is determined that the person captured by the camera 30A and the person captured by the camera 30B are the same, integrates the first position and the third position, and integrates the second position and the fourth position.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to a person identification device and a person identification method. [Background technology]

[0002] Patent Document 1 discloses a technique for estimating the position of a subject included in an image captured by a single camera. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Publication No. 2024-2538 Summary of the Invention [Problem to be solved by the invention]

[0004] In order to estimate the position information of a person in real space using images captured by a camera, it is necessary to comprehensively capture the real space within the imaging range. Comprehensive coverage is difficult with a single camera due to the angle of view and obstacles, so imaging with multiple cameras is required. When multiple cameras are used, it is necessary to integrate the estimation results based on the images captured by each camera. [Means for solving the problem]

[0005] A person identification device according to an embodiment of the present disclosure includes a first determination unit that, based on a first image in which a target area is viewed from a first direction, determines, for a person appearing in the first image, a first position on a 3D model of the target area in the first frame and a second position on the 3D model in a second frame subsequent to the first frame, and, based on a second image in which the target area is viewed from a second direction different from the first direction, determines, for the person appearing in the second image, a third position on the 3D model in the first frame and a fourth position on the 3D model in the second frame; and, when it is determined based on the first position and the second position that the person appearing in the first image is the same between the first frame and the second frame, determines that the person appearing in the first image is an existing person whose position trajectory has been captured, and when it is determined based on the first position and the second position that the person appearing in the first image is different between the first frame and the second frame, determines that the person appearing in the first image is an existing person whose position trajectory has been captured a second determination unit that determines a person appearing in the second image to be a new person whose positional trajectory has not been captured, and, if it is determined based on the third position and the fourth position that the person appearing in the second image is the same between the first frame and the second frame, determines the person appearing in the second image to be the existing person, and, if it is determined based on the third position and the fourth position that the person appearing in the second image is different between the first frame and the second frame, determines the person appearing in the second image to be the new person; and an integration unit that, if it is determined based on the first position and the third position that the person appearing in the first image and the person appearing in the second image are the same in the first frame, integrates the first position and the third position, and, if it is determined based on the second position and the fourth position that the person appearing in the first image and the person appearing in the second image are the same in the second frame, integrates the second position and the fourth position.

[0006] Further, a person identification method according to an embodiment of the present disclosure includes: determining, based on a first image in which a target area is viewed from a first direction, a first position on a 3D model of the target area in the first frame and a second position on the 3D model in a second frame subsequent to the first frame, for a person appearing in the first image; determining, based on a second image in which the target area is viewed from a second direction different from the first direction, a third position on the 3D model in the first frame and a fourth position on the 3D model in the second frame, for the person appearing in the second image; and, when it is determined based on the first position and the second position that the person appearing in the first image is the same between the first frame and the second frame, determining that the person appearing in the first image is an existing person whose position trajectory has been captured; and, when it is determined based on the first position and the second position that the person appearing in the first image is different between the first frame and the second frame, determining that the person appearing in the first image is an existing person whose position trajectory has been captured. determining that a person appearing in one image is a new person whose positional trajectory has not been captured; if it is determined based on the third position and the fourth position that the person appearing in the second image is the same between the first frame and the second frame, determining that the person appearing in the second image is the existing person; if it is determined based on the third position and the fourth position that the person appearing in the second image is different between the first frame and the second frame, determining that the person appearing in the second image is the new person; if it is determined based on the first position and the third position that the person appearing in the first image and the person appearing in the second image are the same in the first frame, merging the first position and the third position; and if it is determined based on the second position and the fourth position that the person appearing in the first image and the person appearing in the second image are the same in the second frame, merging the second position and the fourth position. [Effects of the Invention]

[0007] According to the present disclosure, by using a 3D model of the target space, it is possible to identify a person between cameras by comparing the position information of the person in a model coordinate system based on images of the person captured by multiple cameras. [Brief explanation of the drawings]

[0008] [Figure 1] FIG. 1 is a block diagram showing a configuration of a person identification device 20 according to an embodiment of the present disclosure. [Figure 2] 10 is a diagram for explaining the imaging ranges of cameras 30A and 30B. FIG. [Figure 3] FIG. 2 is a diagram for explaining the position of the center of gravity of the foot. [Figure 4] 10 is a flowchart showing the flow of a person identification method executed by the processing device 230 in accordance with a program PR1. [Figure 5] 10 is a diagram for explaining an example of the operation of the person identification device 20. FIG. [Figure 6] 10 is a diagram for explaining an example of the operation of the person identification device 20. FIG. [Figure 7] 10 is a diagram for explaining an example of the operation of the person identification device 20. FIG. DETAILED DESCRIPTION OF THE INVENTION

[0009] (A. Embodiment) Fig. 1 is a block diagram showing an example of the configuration of a person identification device 20. The person identification device 20 is a device that identifies whether or not the people appearing in each captured image are the same (i.e., whether they are the same person) based on each captured image obtained by capturing images of people present in a real space such as a brick-and-mortar store, and estimates the positions in real space of people identified as the same person based on each captured image. As shown in Fig. 1, the person identification device 20 includes a communication device 210, a storage device 220, a processing device 230, and a bus 240 that interconnects these devices.

[0010] The communication device 210 has an interface to which other devices can be connected, and communicates with the other devices connected to the interface via wireless or wired communication. Specific examples of other devices connected to the communication device 210 include cameras 30A and 30B installed in a target area RS where the person identification device 20 identifies and estimates the position of a person. The person identification device 20 may be connected to the cameras 30A and 30B via a communication network. In this case, the cameras 30A and 30B may be installed in a physical store or the like, and the person identification device 20 may be installed remotely from the physical store or the like. Each of the cameras 30A and 30B may be, for example, a monocular camera, which captures images of its respective imaging range at a predetermined interval and transmits image data representing the captured images to the communication device 210. In this embodiment, the cameras 30A and 30B preferably capture images synchronously with each other. That is, in this embodiment, it is preferable that the time when the camera 30A captures an image and the time when the camera 30B captures an image are approximately the same. However, cameras 30A and 30B do not necessarily need to capture images synchronously. In the following description, when there is no need to distinguish between cameras 30A and 30B, cameras 30A and 30B will be referred to as "camera 30." Furthermore, each captured image captured by camera 30 at a predetermined cycle will be referred to as a frame.

[0011] FIG. 2 is a diagram illustrating the imaging range of the camera 30 in the target area RS. FIG. 2 illustrates a bird's-eye view of the target area RS in which the camera 30 is installed. In this embodiment, the target area RS is a real space such as, for example, a brick-and-mortar store. As shown in FIG. 2, a door DR is provided on a wall defining the target area RS, through which people enter and exit the target area RS. In this embodiment, the cameras 30A and 30B are disposed at different positions in the target area RS so that their imaging ranges partially overlap in the target area RS, and so that their imaging directions (the directions of the optical axes of the imaging lenses) are different, in order to comprehensively capture the interior of the target area RS. Specifically, the camera 30A is disposed in the far right corner of the target area RS as viewed from the position of the door DR, with its imaging direction facing the center of the wall on which the door DR is provided. The camera 30B is disposed in the far left corner of the target area RS as viewed from the position of the door DR, with its imaging direction facing the center of the wall on which the door DR is provided. Camera 30A is an example of a first camera in the present disclosure. The imaging direction of camera 30A is an example of a first direction in the present disclosure. Camera 30B is an example of a second camera in the present disclosure. The imaging direction of camera 30B is an example of a second direction in the present disclosure.

[0012] In FIG. 2, boundary lines L1 and L2 that define the imaging range of camera 30A are shown with dashed lines, and the area between boundary lines L1 and L2 is the imaging range of camera 30A. Therefore, areas A1, A2, and A4 in FIG. 2 are included in the imaging range of camera 30A. Also in FIG. 2, boundary lines L3 and L4 that define the imaging range of camera 30B are shown with dashed lines, and the area between boundary lines L3 and L4 is the imaging range of camera 30A. Therefore, areas A2, A3, and A5 in FIG. 2 are included in the imaging range of camera 30B. In this embodiment, two cameras 30 are installed in target area RS, but three or more cameras 30 may be installed. This is because the larger the target area RS, the more cameras 30 are required to comprehensively capture the interior of the target area RS.

[0013] The storage device 220 is a recording medium readable by the processing device 230. The storage device 220 includes, for example, a nonvolatile memory and a volatile memory. The nonvolatile memory is, for example, a read-only memory (ROM), an erasable programmable read-only memory (EPROM), and an electrically erasable programmable read-only memory (EEPROM). The volatile memory is, for example, a random access memory (RAM). The storage device 220 stores the machine learning model MDL, conversion information D1, and program PR1.

[0014] The machine learning model MDL functions as an AI (artificial intelligence) for estimating the skeleton of a person appearing in an image captured by the camera 30, based on the image. The machine learning model MDL is generated by performing machine learning such as deep learning using training data that pairs a training image (an image containing a person) with the name and relative position of each feature point in the skeleton of the person appearing in the image. The name and relative position of each feature point serve as a ground truth label. The relative position of a feature point refers to the distance and direction from the camera that captured the training image to the feature point, and this distance is measured using, for example, a depth sensor. Specific examples of feature points include joints such as shoulders, elbows, wrists, and ankles, and extremities such as toes and heels. Examples of such machine learning models include models using MMPose, OpenPose, etc. When an image containing a person is input to the machine learning model, the machine learning model outputs skeleton position information indicating the names and relative positions of feature points in the skeleton of the person appearing in the input image, and an index indicating the reliability (likelihood) of the estimation result. The skeleton position information indicates the relative position of each feature point based on the position of the camera that captured the image from which the skeleton position information was generated. In this embodiment, the index indicating the reliability of the estimation result is a value greater than 0 and less than 1, and the closer the index value is to 1, the higher the reliability of the relative position of each feature point.

[0015] The conversion information D1 is information for converting the relative position of each feature point indicated by the skeleton position information obtained by inputting the captured image of the camera 30 into the machine learning model MDL into a position in a coordinate system (e.g., a three-dimensional coordinate system with a predetermined position in the 3D model as the origin; hereinafter, referred to as a model coordinate system) that defines the position in the 3D model of the target area RS. The 3D model of the target area RS is formed by defining planes or curved surfaces corresponding to each surface (floor, wall, and ceiling) that partitions the target area RS, for example, using CAD (Computer Aided Design) or the like. The conversion information D1 in this embodiment includes information indicating the installation position of the camera 30A in the model coordinate system and information indicating the installation position of the camera 30B in the model coordinate system. For example, for a feature point whose relative position from the camera 30A is represented by the skeleton position information generated based on the captured image of the camera 30A, the position in the model coordinate system is determined by translating the installation position of the camera 30A in the model coordinate system according to the relative position. The 3D model may be generated based on the captured image output from the camera 30A and the captured image output from the camera 30B.

[0016] The processing device 230 includes one or more central processing units (CPUs). The one or more CPUs are an example of one or more processors. Each of the processor and the CPU is an example of a computer. The processing device 230 reads a program PR1 from the storage device 220. The processing device 230 executes the read program PR1 to function as a first determination unit 230a, a second determination unit 230b, and an integration unit 230c.

[0017] Based on an image captured by camera 30A, the first determination unit 230a determines, for a person appearing in the captured image, the position in the model coordinate system in an arbitrary frame (hereinafter referred to as the first frame) and the position in the model coordinate system in a second frame, which is a frame subsequent to the first frame. Note that if a person is not appearing in the image captured by camera 30A, the position in the model coordinate system is not determined. The position of the person may be any position that represents the person. In this embodiment, the first determination unit 230a determines, for example, the position of the center of gravity of the foot skeleton as the position of the person. When a person moves, the torso moves up, down, left, and right, but the feet spend a long time in contact with the floor and are stable. For this reason, the position of the center of gravity of the foot skeleton is suitable as a position that represents the person. Hereinafter, the position of the person in the model coordinate system determined based on the first frame in the image captured by camera 30A will be referred to as the first position. Also, below, the position of the person in the model coordinate system determined based on the second frame in the image captured by camera 30A will be referred to as the second position.

[0018] Furthermore, based on the image captured by camera 30B, first determination unit 230a determines the position in the model coordinate system in the first frame and the position in the model coordinate system in the second frame of a person appearing in the captured image. Hereinafter, the position of the person in the model coordinate system determined based on the first frame of the image captured by camera 30B will be referred to as the third position. Hereinafter, the position of the person in the model coordinate system determined based on the second frame of the image captured by camera 30B will be referred to as the fourth position.

[0019] More specifically, the first determination unit 230a inputs a captured image representing a first frame of the image captured by the camera 30A into the machine learning model MDL, thereby acquiring skeletal position information indicating the positions of skeletal feature points of the person appearing in the captured image and the aforementioned indices. Hereinafter, the indices obtained by inputting the first frame of the image captured by the camera 30A into the machine learning model MDL will be referred to as first indices. Next, the first determination unit 230a determines the relative positions of the centers of gravity of the feet of the person appearing in the captured image, relative to the position of the camera 30A, based on the feature points of the foot skeleton (in this embodiment, the toes and heels of both feet) among the feature points whose relative positions are indicated by the skeletal position information.

[0020] FIG. 3 is an explanatory diagram of the center of gravity of the foot. FIG. 3 shows the toe RT and heel RE, which are characteristic points of the right foot skeleton RB, and the toe LT and heel LE, which are characteristic points of the left foot skeleton LB. As shown in FIG. 3, the center of gravity of the foot refers to the position of the midpoint CC of the line segment connecting the midpoint RM of the line segment connecting the toe RT and heel RE of the right foot and the midpoint LM of the line segment connecting the toe LT and heel LE of the left foot. The first determination unit 230a converts the relative position of the center of gravity determined in the above manner into a position in the model coordinate system using the conversion information D1, thereby determining the position in the model coordinate system of the person appearing in the first frame of the image captured by camera 30A, i.e., the first position.

[0021] The first determination unit 230a similarly determines the second, third, and fourth positions using the machine learning model MDL and conversion information D1. Hereinafter, an index obtained by inputting the second frame of the image captured by camera 30A into the machine learning model MDL will be referred to as the second index. An index obtained by inputting the first frame of the image captured by camera 30B into the machine learning model MDL will be referred to as the third index. An index obtained by inputting the second frame of the image captured by camera 30B into the machine learning model MDL will be referred to as the fourth index.

[0022] The second determination section 230b includes a first determination section 230b1 and a second determination section 230b2. The first determination unit 230b1 performs person identification between frames. Person identification between frames refers to determining whether a person appearing in each of multiple frames captured by a single camera at different times is the same. In this embodiment, the first determination unit 230b1 determines whether a person appearing in a first frame of an image captured by the camera 30A is the same as a person appearing in a second frame of the image captured by the camera 30A, based on the first and second positions determined by the first determination unit 230a. For example, if the distance between the first and second positions is less than a threshold, the first determination unit 230b1 determines that the person appearing in the first frame of the image captured by the camera 30A is the same as the person appearing in the second frame of the image captured by the camera 30A. Here, the threshold may be determined according to the time difference between the capture times of the first and second frames. If either or both of the first position and the second position have not been determined, inter-frame person identification for the image captured by camera 30A is not performed. Similarly, the first determination unit 230b1 determines whether the person appearing in the first frame of the image captured by camera 30B is the same as the person appearing in the second frame of the image captured by camera 30B, based on the third and fourth positions determined by the first determination unit 230a. If either or both of the third and fourth positions have not been determined, inter-frame person identification for the image captured by camera 30B is not performed.

[0023] The first determination unit 230b1 may refer to the skeletal position information used to determine the first position and the skeletal position information used to determine the second position and determine whether the person appearing in the first frame of the image captured by the camera 30A is the same as the person appearing in the second frame of the image captured by the camera 30A, without converting to a model coordinate system. For example, if the difference in distance between specific feature points (e.g., the distance from the toe to the heel or the distance from the knee to the ankle) is less than a predetermined threshold, the first determination unit 230b1 may determine that the person appearing in the first frame of the image captured by the camera 30A is the same as the person appearing in the second frame of the image captured by the camera 30A. Similarly, it is also possible to determine whether the person appearing in the first frame of the image captured by the camera 30B is the same as the person appearing in the second frame of the image captured by the camera 30B, without converting to a model coordinate system.

[0024] The second determination unit 230b2 performs inter-camera person identification. Inter-camera person identification refers to determining whether the person appearing in each frame captured at the same time by multiple different cameras is the same. In this embodiment, the second determination unit 230b2 determines whether the person appearing in the first frame of the image captured by camera 30A is the same as the person appearing in the first frame of the image captured by camera 30B. Note that if a person is not captured in either or both of the first frame of the image captured by camera 30A and the first frame of the image captured by camera 30B, inter-camera person identification for the first frame is not performed. In this embodiment, the second determination unit 230b2 determines whether the person appearing in the first frame of the image captured by camera 30A is the same as the person appearing in the first frame of the image captured by camera 30B, based on the foot orientation in the model coordinate system of the person appearing in the first frame of the image captured by camera 30A, and the foot orientation, first position, and third position in the model coordinate system of the person appearing in the first frame of the image captured by camera 30B. An example of the person's foot orientation is the direction from the heel to the toe. Information representing the foot orientation in the model coordinate system is obtained based on the position of the heel transformed into the model coordinate system and the position of the toe transformed into the model coordinate system. In this embodiment, the foot orientation refers to, for example, the direction of the right foot, and refers to the direction from the heel RE of the right foot to the toe RT of the right foot in the model coordinate system. In this embodiment, person identification between cameras is performed taking into account the direction of the person's feet, which improves the accuracy of person identification between cameras compared to when person identification between cameras is performed without taking into account the direction of the feet. However, person identification between cameras may also be performed based on the first and third positions without taking into account the direction of the feet.

[0025] In this embodiment, the second determination unit 230b2 determines the orientation of the feet of a person appearing in a first frame of an image captured by the camera 30A based on skeletal position information obtained by inputting the first frame into the machine learning model MDL and transformation information D1. More specifically, the second determination unit 230b2 acquires the relative positions of the right toe RT and the right heel RE from the skeletal position information obtained by inputting the first frame of the image captured by the camera 30A into the machine learning model MDL. Next, the second determination unit 230b2 calculates the positions of the right toe RT and the right heel RE in the model coordinate system based on the relative positions of the right toe RT and the right heel RE and the transformation information D1. The second determination unit 230b2 then determines the direction from the heel RE of the right foot to the toe RT of the right foot in the model coordinate system as the foot orientation of the person appearing in the first frame of the image captured by camera 30A. Similarly, the second determination unit 230b2 determines the foot orientation of the person appearing in the first frame of the image captured by camera 30B based on the skeleton position information and transformation information D1 obtained by inputting the first frame into the machine learning model MDL. The second determination unit 230b2 then determines that the two persons are the same person if the angle formed between the foot orientation of the person appearing in the first frame of the image captured by camera 30A and the foot orientation of the person appearing in the first frame of the image captured by camera 30B is equal to or smaller than a first angle and the distance between the first position and a third position is equal to or smaller than a threshold. The first angle may be determined taking into account system error. The first angle may be, for example, 45 degrees, and more preferably 15 degrees.

[0026] In addition, the second determination unit 230b2 may determine whether the person appearing in the first frame of the image captured by camera 30A is the same as the person appearing in the first frame of the image captured by camera 30B, for example, by comparing the distances between specific feature points in the skeleton.

[0027] Similarly, the second determination unit 230b2 determines whether the person appearing in the second frame of the image captured by camera 30A and the person appearing in the second frame of the image captured by camera 30B are the same person based on the orientation of the feet in the model coordinate system of the person appearing in the second frame of the image captured by camera 30A, the orientation of the feet in the model coordinate system of the person appearing in the second frame of the image captured by camera 30B, the second position, and the fourth position. That is, the second determination unit 230b2 determines that the two are the same person when the angle formed between the orientation of the feet of the person appearing in the second frame of the image captured by camera 30A and the orientation of the feet of the person appearing in the second frame of the image captured by camera 30B is equal to or smaller than the first angle and the distance between the second position and the fourth position is equal to or smaller than a threshold. If no person is captured in either or both of the second frame in the image captured by camera 30A and the second frame in the image captured by camera 30B, person identification between the cameras regarding the second frame is not performed.

[0028] When the first determination unit 230b1 determines that the person appearing in the first frame of the image captured by the camera 30A and the person appearing in the second frame of the image captured by the camera 30A are the same person, the second determination unit 230b determines that the person captured by the camera 30A from the first frame to the second frame is an existing person whose positional trajectory in the model coordinate system has been captured. On the other hand, when the first determination unit 230b1 determines that the person appearing in the first frame of the image captured by the camera 30A and the person appearing in the second frame of the image captured by the camera 30A are not the same person, the second determination unit 230b determines that the person captured by the camera 30A is a new person whose positional trajectory in the model coordinate system has not been captured.

[0029] Similarly, if the second determination unit 230b determines that the person appearing in the first frame of the image captured by camera 30B and the person appearing in the second frame of the image captured by camera 30B are the same person, it determines that the person captured by camera 30B from the first frame to the second frame is an existing person. Conversely, if the second determination unit 230b determines that the person appearing in the first frame of the image captured by camera 30B and the person appearing in the second frame of the image captured by camera 30B are not the same person, it determines that the person captured by camera 30B is a new person.

[0030] When the second determination unit 230b2 determines that the person appearing in the first frame of the image captured by the camera 30A is the same as the person appearing in the first frame of the image captured by the camera 30B, the integrating unit 230c integrates the first position and the third position. If either the first position or the third position has not been determined, the integration of the first position and the third position is not performed. In this embodiment, the integrating unit 230c integrates the first position and the third position by a weighted average using the first index described above as the weight of the first position and the third index described above as the weight of the third position. For example, if the first position is (x1, y1, z1) and the third position is (x3, y3, z3), the integrating unit 230c calculates the integrated position (X, Y, Z) according to the following formulas (1) to (3). In formulas (1) to (3), a1 is the first index and a3 is the third index. X=(a1×x1+a3×x3) / (a1+a3)...(1) Y=(a1×y1+a3×y3) / (a1+a3)···(2) z=(a1×z1+a3×z3) / (a1+a3)...(3)

[0031] By integrating the first and third positions using a weighted average with the first index as the weight for the first position and the third index as the weight for the third position, the trajectory of the position captured for the existing person becomes smoother. In addition, the more likely of the first and third positions contributes more to the integrated position, improving the accuracy of the integrated position.

[0032] Similarly, when the second determination unit 230b2 determines that the person appearing in the second frame of the image captured by camera 30A is the same as the person appearing in the second frame of the image captured by camera 30B, the integrating unit 230c integrates the second position and the fourth position by a weighted average in which the second index is used as the weight for the second position and the fourth index is used as the weight for the fourth position. Note that when either the second position or the fourth position has not been determined, the second position and the fourth position are not integrated.

[0033] Furthermore, the processing device 230 operates in accordance with the program PR1 to repeatedly execute a person identification method that prominently exhibits the features of the present disclosure at a predetermined interval. Figure 4 is a flowchart showing the flow of processing in this person identification method.

[0034] In the first determination process S100, the processing device 230 functions as a first determination unit 230a. In the first determination process S100, the processing device 230 determines a first position and a second position for a person appearing in an image captured by camera 30A, based on the image captured by the camera 30A. Also, in the first determination process S100, the processing device 230 determines a third position and a fourth position for a person appearing in an image captured by camera 30B, based on the image captured by the camera 30B.

[0035] In steps S110, S120, S130, S150, and S160, the processing device 230 functions as the second determination unit 230b. More specifically, in step S110, the processing device 230 determines whether or not there is inter-frame person identification for at least one of the cameras 30A and 30B. If there is inter-frame person identification for at least one of the cameras 30A and 30B, the determination result in step S110 is "Yes." On the other hand, if there is no inter-frame person identification for the camera 30A and no inter-frame person identification for the camera 30B, the determination result in step S110 is "No." If the determination result in step S110 is "No," the processing device 230 executes the process of step S120. If the determination result in step S110 is "Yes," the processing device 230 executes the process of step S150.

[0036] In step S120, which is executed if the determination result in step S110 is "No," the processing device 230 determines a person appearing in an image captured by the camera 30 as a new person. The determination result in step S110 is "No" if a person cannot be identified between frames, which means that the person appearing in the past first frame and the person appearing in the current second frame are different.

[0037] In step S130 following step S120, the processing device 230 determines whether or not there is person identification between the cameras. If there is person identification between the cameras, the determination result in step S130 is "Yes." If there is no person identification between the cameras, the determination result in step S130 is "No." If the determination result in step S130 is "Yes," the processing device 230 executes the integration process S140 and terminates the person identification method. If the determination result in step S130 is "No," the processing device 230 terminates the person identification method without executing the integration process S140. In the integration process S140, the processing device 230 functions as the integration unit 230c, and integrates the first position and the third position determined in the first determination process S100 using the weighted average described above, while integrating the second position and the fourth position determined in the first determination process S100 using the weighted average described above.

[0038] In step S150, which is executed if the determination result in step S110 is "Yes," the processing device 230 determines that a person appearing in an image captured by the camera 30 is an existing person. In step S160 following step S150, the processing device 230 determines whether or not there has been person identification between the cameras, as in step S130. If there has been person identification between the cameras, the determination result in step S160 is "Yes." If there has not been person identification between the cameras, the determination result in step S160 is "No." If the determination result in step S160 is "Yes," the processing device 230 executes integration processing S170 and ends the person identification method. If the determination result in step S160 is "No," the processing device 230 ends the person identification method without executing integration processing S170. In the integration process S170, the processing device 230 functions as the integration unit 230c, as in the integration process S140, and integrates the first position and the third position determined in the first determination process S100 using the weighted average described above, while integrating the second position and the fourth position determined in the first determination process S100 using the weighted average described above.

[0039] For example, as shown in FIGS. 5, 6, and 7, assume that a user U who entered the target area RS through the door DR stays in area A1 from time T1 to time T2 (see FIG. 5), then moves to area A2 and stays in area A2 from time T3 to time T4 (see FIG. 6), and then moves to area A3 and stays in area A3 from time T5 to time T6 (see FIG. 7). The user U depicted with a solid line in FIGS. 5, 6, and 7 represents the user U who appears in the image captured by camera 30A, and the user U depicted with a dotted line represents the user U who appears in the image captured by camera 30B. The image captured at time T1 is an example of a first frame, and the image captured at time T2 is an example of a second frame. The image captured at time T3 is an example of a first frame, and the image captured at time T4 is an example of a second frame. The image captured at time T5 is an example of a first frame, and the image captured at time T6 is an example of a second frame.

[0040] As shown in FIG. 5, area A1 is included in the imaging range of camera 30A but is not included in the imaging range of camera 30B. Therefore, user U appears in the images captured by camera 30A during the period from time T1 to time T2, but user U does not appear in the images captured by camera 30B. Therefore, in the first determination process S100, a first position P1 and a second position P2 are determined based on the first and second frames of the images captured by camera 30A, but a third position and a fourth position are not determined. Because user U does not appear in the images captured by camera 30B during the period from time T1 to time T2, there is no person identification between frames for camera 30B. Furthermore, because user U appears in the first frame of the images captured by camera 30A at time T1, there is no person identification between frames for camera 30A until time T2. Therefore, the determination result in step S110 is "No", the processing of step S120 is executed, and the person appearing in the image captured by camera 30A is determined to be a new person. As described above, while user U appears in the image captured by camera 30A during the period from time T1 to time T2, user U does not appear in the image captured by camera 30B, so there is no person identification between the cameras, and the determination result in step S130 is "No". Therefore, integration processing S140 is not executed during the period from time T1 to time T2.

[0041] As shown in FIG. 6, area A2 is included in the imaging range of camera 30A and also in the imaging range of camera 30B. Therefore, user U appears in the images captured by camera 30A during the period from time T3 to time T4, and user U also appears in the images captured by camera 30B. Therefore, in the first determination process S100, the first position P1 and the second position P2 are determined based on the images captured by camera 30A at times T3 and T4, respectively, while the third position P3 and the fourth position P4 are determined based on the images captured by camera 30B at times T3 and T4, respectively. At time T3, person identification between frames has been performed for camera 30A, so the determination result of step S110 is “Yes,” and the process of step S150 is executed. That is, the person appearing in the images captured by camera 30A is determined to be an existing person. Furthermore, in the period from time T3 to time T4, the first position P1, the second position P2, the third position P3, and the fourth position P4 are determined, and person identification between the cameras is performed, so the determination result in step S160 becomes "Yes" and the integration process S170 is executed. That is, the first position P1 and the third position P3 determined in the first determination process S100 are integrated, and the second position P2 and the fourth position P4 are integrated.

[0042] As shown in FIG. 7, area A3 is included in the imaging range of camera 30B but is not included in the imaging range of camera 30A. Therefore, user U appears in the images captured by camera 30B during the period from time T5 to time T6, but user U does not appear in the images captured by camera 30A. Therefore, in the first determination process S100, the third position P3 and the fourth position P4 are determined based on the images captured by camera 30B at times T5 and T6, respectively, but the first position and the second position are not determined. During the period from time T5 to time T6, person identification between frames is performed, and the determination result in step S110 is "Yes." Therefore, the processes from step S150 onward are executed, and the person appearing in the images captured by camera 30B is determined to be a known person. During the period from time T5 to time T6, the first position and the second position are not determined, and person identification between cameras is not performed, so the determination result in step S160 is "No." Therefore, the integration process S170 is not executed during the period from time T5 to time T6.

[0043] As described above, according to this embodiment, it is possible to determine whether a person appearing in a first frame of an image captured by camera 30A and a person appearing in a second frame are the same person based on the first and second positions. Similarly, it is possible to determine whether a person appearing in a first frame of an image captured by camera 30B and a person appearing in a second frame are the same person based on the third and fourth positions. In other words, it is possible to identify people between frames.

[0044] Furthermore, according to this embodiment, it is possible to determine whether a person appearing in a first frame of an image captured by camera 30A is the same as a person appearing in a first frame of an image captured by camera 30B, and to determine whether a person appearing in a second frame of an image captured by camera 30A is the same as a person appearing in a second frame of an image captured by camera 30B. That is, it is possible to identify people captured by camera 30A and camera 30B at the same time. Then, if it is determined that the person captured by camera 30A is the same as the person captured by camera 30B, the position of the person determined based on the image captured by camera 30A can be integrated with the position of the person determined based on the image captured by camera 30B.

[0045] (B: Transformation) The above embodiment can be modified as follows. (B-1) The integrating unit 230c may integrate the first position and the third position by selecting either the first position or the third position. Similarly, for integrating the second position and the fourth position, the integrating unit 230c may integrate the second position and the fourth position by selecting either the second position or the fourth position. According to this aspect, the processing load related to the integration of positions is lighter than when integrating positions using weighted averaging.

[0046] Various modes are possible for selecting either the first position or the third position, and either the second position or the fourth position. For example, when selecting either the first position or the third position, a mode can be considered in which the position closer to the camera 30 is selected. According to this mode, if the distance between the camera 30A and the first position is shorter than the distance between the camera 30B and the third position, the integrating unit 230c selects the first position. Conversely, if the distance between the camera 30B and the third position is shorter than the distance between the camera 30A and the first position, the integrating unit 230c selects the third position. This is because a position determined based on an image captured by the camera 30 at a shorter distance is considered to be more accurate. Similarly, when selecting either the second position or the fourth position, the integrating unit 230c selects the position closer to the camera 30.

[0047] Another possible mode for selecting either the first position or the third position is to select the one with a higher degree of certainty based on a first index indicating the certainty of the first position and a third index indicating the certainty of the third position. In this mode, if the first index is greater than the third index, the integrating unit 230c selects the first position. Conversely, if the first index is greater than the third index, the integrating unit 230c selects the third position. This is because a position with a larger index indicating certainty is considered to be more accurate. Similarly, when selecting either the second position or the fourth position, the integrating unit 230c selects the one with a higher degree of certainty.

[0048] (B-2: Variation 2) In the above-described embodiment, the program PR1 is stored in the storage device 220 of the person identification device 20, but the program PR1 may be manufactured or sold as a standalone program. When selling the program PR1, the program PR1 may be provided to a purchaser by writing the program PR1 to a computer-readable recording medium such as a flash ROM and distributing it, or by downloading it via a telecommunications line.

[0049] (C:Other) (1) In the above-described embodiment, ROM and RAM are exemplified as storage device 220, but storage device 220 may also be a flexible disk, a magneto-optical disk (e.g., a compact disk, a digital versatile disk, a Blu-ray (registered trademark) disc), a smart card, a flash memory device (e.g., a card, a stick, a key drive), a CD-ROM (Compact Disc-ROM), a register, a removable disk, a hard disk, a floppy (registered trademark) disk, a magnetic strip, a database, a server, or other suitable storage medium.

[0050] (2) In the above-described embodiments, the described information, signals, etc. may be represented using any of a variety of different technologies. For example, data, instructions, commands, information, signals, bits, symbols, chips, etc. that may be referred to throughout the above description may be represented by voltages, currents, electromagnetic waves, magnetic fields or magnetic particles, optical fields or photons, or any combination thereof.

[0051] (3) In the above-described embodiment, input and output information may be stored in a specific location (for example, memory) or may be managed using a table. Input and output information may be overwritten, updated, or added to. Output information may be deleted. Input information may be transmitted to another device.

[0052] (4) In the above-described embodiment, the determination may be made based on a value represented by one bit (0 or 1), a Boolean value (true or false), or a comparison of numerical values ​​(e.g., comparison with a predetermined value).

[0053] (5) The order of the process procedures, sequences, flowcharts, etc. illustrated in the above-described embodiments may be rearranged unless inconsistent. For example, the methods described in this disclosure present elements of various steps using an example order, and are not limited to the particular order presented.

[0054] (6) Each function illustrated in FIG. 1 is realized by any combination of at least one of hardware and software. Furthermore, the method for realizing each functional block is not particularly limited. That is, each functional block may be realized using a single device that is physically or logically coupled, or may be realized using two or more physically or logically separated devices that are connected directly or indirectly (for example, by wire, wirelessly, etc.) and these multiple devices. A functional block may also be realized by combining software with the single device or the multiple devices.

[0055] (7) The programs exemplified in the above-described embodiments should be broadly construed to mean instructions, instruction sets, code, code segments, program code, programs, subprograms, software modules, applications, software applications, software packages, routines, subroutines, objects, executable files, threads of execution, procedures, functions, etc., regardless of whether they are called software, firmware, middleware, microcode, hardware description language, or by other names.

[0056] Software, instructions, information, etc. may also be transmitted or received over a transmission medium. For example, if software is transmitted from a website, server, or other remote source using wired technologies (such as coaxial cable, fiber optic cable, twisted pair, Digital Subscriber Line (DSL)), and / or wireless technologies (such as infrared, microwave), then these wired and / or wireless technologies are included within the definition of transmission media.

[0057] (8) In each of the foregoing embodiments, the terms "system" and "network" are used interchangeably.

[0058] (9) The information, parameters, etc. described in this disclosure may be expressed using absolute values, relative values ​​from a predetermined value, or corresponding other information.

[0059] (10) In the above-described embodiments, the terms "connected," "coupled," or any variation thereof refers to any direct or indirect connection or coupling between two or more elements, and may include the presence of one or more intermediate elements between two elements that are "connected" or "coupled" to each other. The coupling or connection between elements may be physical, logical, or a combination thereof. For example, "connected" may be read as "access." As used in this disclosure, two elements may be considered to be "connected" or "coupled" to each other using at least one of one or more wires, cables, and printed electrical connections, as well as electromagnetic energy having wavelengths in the radio frequency range, microwave range, and optical (both visible and invisible) range, as some non-limiting and non-exhaustive examples.

[0060] (11) In the above embodiments, the phrase "based on" does not mean "based only on," unless otherwise specified. In other words, the phrase "based on" means both "based only on" and "based at least on."

[0061] (12) As used in this disclosure, the terms "determining" and "determining" may encompass a wide variety of actions. "Determining" and "determining" may include, for example, judging, calculating, computing, processing, deriving, investigating, looking up, searching, inquiring (e.g., searching in a table, database, or other data structure), ascertaining, and the like. "Determining" and "determining" may also include receiving (e.g., receiving information), transmitting (e.g., sending information), input, output, accessing (e.g., accessing data in memory), and the like. Furthermore, "judgment" and "decision" can include regarding resolving, selecting, choosing, establishing, comparing, etc. as having been "judgment" or "decision." In other words, "judgment" and "decision" can include regarding some action as having been "judgment" or "decision." Furthermore, "judgment (decision)" can be interpreted as "assuming," "expecting," "considering," etc.

[0062] (13) In the above embodiments, when "include," "including," and variations thereof are used, these terms are intended to be inclusive, similar to the term "comprising." Furthermore, the term "or" as used in this disclosure is not intended to be an exclusive or.

[0063] (14) In this disclosure, where articles are added by translation, such as a, an, and the in English, this disclosure may include the nouns following these articles being plural.

[0064] (15) In this disclosure, the term "A and B are different" may mean "A and B are different from each other." The term may also mean "A and B are each different from C." Terms such as "separate" and "combined" may also be interpreted in the same way as "different."

[0065] (16) Each aspect / embodiment described in this disclosure may be used alone, in combination, or switched depending on the implementation. Notification of predetermined information (e.g., notification that "X is true") is not limited to explicit notification, but may be implicit (e.g., not notifying the predetermined information).

[0066] (D: Aspects understood from the above-described embodiments or modifications) Although the present disclosure has been described in detail above, it is clear to those skilled in the art that the present disclosure is not limited to the embodiments described herein. The present disclosure can be implemented in modified and altered forms without departing from the spirit and scope of the present disclosure as defined by the claims. Therefore, the description of the present disclosure is intended to be illustrative and does not have any limiting meaning on the present disclosure. The following aspects can be understood from at least one of the above-described embodiments or modifications.

[0067] A person identification device according to a first aspect of the present disclosure includes a first determination unit that, based on a first image in which a target area is viewed from a first direction, determines, for a person appearing in the first image, a first position on a 3D model of the target area in the first frame and a second position on the 3D model in a second frame subsequent to the first frame, and, based on a second image in which the target area is viewed from a second direction different from the first direction, determines, for the person appearing in the second image, a third position on the 3D model in the first frame and a fourth position on the 3D model in the second frame; and, when it is determined based on the first position and the second position that the person appearing in the first image is the same between the first frame and the second frame, determines that the person appearing in the first image is an existing person whose position trajectory has been captured, and when it is determined based on the first position and the second position that the person appearing in the first image is different between the first frame and the second frame, determines that the person appearing in the first image is an existing person whose position trajectory has been captured The image capturing device is provided with a second determination unit that determines a person appearing in an image to be a new person whose positional trajectory has not been captured, and if it is determined based on the third position and the fourth position that the person appearing in the second image is the same between the first frame and the second frame, determines the person appearing in the second image to be the existing person, and if it is determined based on the third position and the fourth position that the person appearing in the second image is different between the first frame and the second frame, determines the person appearing in the second image to be the new person; and an integration unit that, if it is determined based on the first position and the third position that the person appearing in the first image and the person appearing in the second image are the same in the first frame, integrates the first position and the third position, and if it is determined based on the second position and the fourth position that the person appearing in the first image and the person appearing in the second image are the same in the second frame, integrates the second position and the fourth position. According to the person identification device of the first aspect, it is possible to identify a person who appears in a first image captured by a first camera and a person who appears in a second image captured by a second camera.According to the first aspect, it is possible to identify a person who appears in a first image captured by a first camera and a person who appears in a second image captured by a second camera.

[0068] A person identification device according to a second aspect (an example of the first aspect) of the present disclosure may include a first determination unit that determines, based on the first position and the second position, whether a person appearing in the first image is the same between the first frame and the second frame, and determines, based on the third position and the fourth position, whether a person appearing in the second image is the same between the first frame and the second frame; and a second determination unit that determines, based on the first position and the third position, whether a person appearing in the first image and a person appearing in the second image are the same in the first frame, and determines, based on the second position and the fourth position, whether a person appearing in the first image and a person appearing in the second image are the same in the second frame. According to the second aspect of the person identification device, in addition to being able to identify the person appearing in the first image captured by the first camera and the person appearing in the second image captured by the second camera, it is also possible to determine whether the person appearing in the first image is the same between the first frame and the second frame, and whether the person appearing in the second image is the same between the first frame and the second frame.

[0069] The second determination unit in the person identification device according to a third aspect (an example of the second aspect) of the present disclosure may determine, in the first frame, whether the person appearing in the first image and the person appearing in the second image are the same, based on the orientation of the feet of the person appearing in the first frame of the first image, the orientation of the feet of the person appearing in the first frame of the second image, the first position, and the third position, and may determine, in the second frame, whether the person appearing in the first image and the person appearing in the second image are the same, based on the orientation of the feet of the person appearing in the second frame of the first image, the orientation of the feet of the person appearing in the second frame of the second image, the second position, and the fourth position. According to the person identification device of the third aspect, identification accuracy is improved by taking into account the orientation of the feet.

[0070] The second determination unit in the person identification device according to a fourth aspect (an example of the third aspect) of the present disclosure may determine that the person appearing in the first image and the person appearing in the second image are the same in the first frame if an angle formed between an orientation of the feet of the person appearing in the first frame of the first image and an orientation of the feet of the person appearing in the first frame of the second image is equal to or smaller than a first angle and a distance between the first position and the third position is equal to or smaller than a threshold, and may determine that the person appearing in the first image and the person appearing in the second image are the same in the second frame if an angle formed between an orientation of the feet of the person appearing in the second frame of the first image and an orientation of the feet of the person appearing in the second frame of the second image is equal to or smaller than the first angle and a distance between the second position and the fourth position is equal to or smaller than the threshold. According to the person identification device of the fourth aspect, identification accuracy is improved by taking into account the orientation of the feet.

[0071] The first determination unit in the person identification device according to a fifth aspect (an example of the first, second, third, or fourth aspect) of the present disclosure may determine the first position and the second position based on a foot skeleton of a person appearing in the first image, and may determine the third position and the fourth position based on a foot skeleton of a person appearing in the second image. According to the fifth aspect, the first position and the second position can be determined based on the foot skeleton of a person appearing in the first image. Furthermore, according to the fifth aspect, the third position and the fourth position can be determined based on the foot skeleton of a person appearing in the second image.

[0072] In the person identification device according to a sixth aspect (an example of the first, second, third, fourth, or fifth aspect) of the present disclosure, the first determination unit may determine the first position, the second position, the third position, and the fourth position using a machine learning model. The machine learning model may output a first index indicating a likelihood of the first position, a second index indicating a likelihood of the second position, a third index indicating a likelihood of the third position, and a fourth index indicating a likelihood of the fourth position. The integrating unit may integrate the first position and the third position by a weighted average in which the first index is a weight for the first position and the third index is a weight for the third position, and integrate the second position and the fourth position by a weighted average in which the second index is a weight for the second position and the fourth index is a weight for the fourth position. According to the person identification device of the sixth aspect, a position trajectory is smoothed.

[0073] The integrating unit in the person identification device according to a seventh aspect (an example of the first, second, third, fourth, or fifth aspect) of the present disclosure may integrate the first position and the third position by selecting either the first position or the third position, and may integrate the second position and the fourth position by selecting either the second position or the fourth position. According to the seventh aspect, the processing load related to the integration is lighter than in the sixth aspect.

[0074] In the person identification device according to an eighth aspect (an example of the seventh aspect) of the present disclosure, the first determination unit may determine the first position, the second position, the third position, and the fourth position using a machine learning model. The machine learning model may output a first index indicating a likelihood of the first position, a second index indicating a likelihood of the second position, a third index indicating a likelihood of the third position, and a fourth index indicating a likelihood of the fourth position. The integrating unit may select either the first position or the third position based on the first index and the third index, and select either the second position or the fourth position based on the first index and the third index. According to the eighth aspect, the first position and the third position can be integrated by selecting the one with a higher likelihood between the first position and the third position, and the second position and the fourth position can be integrated by selecting the one with a higher likelihood between the second position and the fourth position.

[0075] A person identification method according to a ninth aspect of the present disclosure includes determining, based on a first image in which a target area is viewed from a first direction, a first position on a 3D model of the target area in the first frame and a second position on the 3D model in a second frame subsequent to the first frame, for a person appearing in the first image; determining, based on a second image in which the target area is viewed from a second direction different from the first direction, a third position on the 3D model in the first frame and a fourth position on the 3D model in the second frame, for the person appearing in the second image; if it is determined based on the first position and the second position that the person appearing in the first image is the same between the first frame and the second frame, determining that the person appearing in the first image is an existing person whose position trajectory has been captured; if it is determined based on the first position and the second position that the person appearing in the first image is different between the first frame and the second frame, determining that the person appearing in the first image is an existing person whose position trajectory has been captured; determining that a person appearing in an image is a new person whose positional trajectory has not been captured; if it is determined based on the third position and the fourth position that the person appearing in the second image is the same between the first frame and the second frame, determining that the person appearing in the second image is the existing person; if it is determined based on the third position and the fourth position that the person appearing in the second image is different between the first frame and the second frame, determining that the person appearing in the second image is the new person; if it is determined based on the first position and the third position that the person appearing in the first image and the person appearing in the second image are the same in the first frame, merging the first position and the third position; and if it is determined based on the second position and the fourth position that the person appearing in the first image and the person appearing in the second image are the same in the second frame, merging the second position and the fourth position. According to the ninth aspect, similarly to the first aspect, it is possible to identify a person who appears in a first image captured by a first camera and a person who appears in a second image captured by a second camera. [Explanation of symbols]

[0076] 20...person identification device, 210...communication device, 220...storage device, 230...processing device, 230a...first decision unit, 230b...second decision unit, 230b1...first judgment unit, 230b2...second judgment unit, 230c...integration unit, 240...bus, PR1...program, MDL...machine learning model, D1...conversion information.

Claims

1. a first determination unit that determines, based on a first image in which a target area is viewed from a first direction, a first position on a 3D model of the target area in a first frame and a second position on the 3D model in a second frame subsequent to the first frame, with respect to a person appearing in the first image, and that determines, based on a second image in which the target area is viewed from a second direction different from the first direction, a third position on the 3D model in the first frame and a fourth position on the 3D model in the second frame, with respect to the person appearing in the second image; a second determination unit that, when it is determined based on the first position and the second position that a person appearing in the first image is the same between the first frame and the second frame, determines that the person appearing in the first image is an existing person whose positional trajectory is being captured; when it is determined based on the first position and the second position that the person appearing in the first image is different between the first frame and the second frame, determines that the person appearing in the first image is a new person whose positional trajectory is not being captured; when it is determined based on the third position and the fourth position that the person appearing in the second image is the same between the first frame and the second frame, determines that the person appearing in the second image is the existing person; and when it is determined based on the third position and the fourth position that the person appearing in the second image is different between the first frame and the second frame, determines that the person appearing in the second image is the new person; an integration unit that integrates the first position and the third position when it is determined based on the first position and the third position that the person appearing in the first image and the person appearing in the second image in the first frame are the same, and that integrates the second position and the fourth position when it is determined based on the second position and the fourth position that the person appearing in the first image and the person appearing in the second image in the second frame are the same; A person identification device comprising:

2. a first determination unit that determines whether a person appearing in the first image is the same between the first frame and the second frame based on the first position and the second position, and determines whether a person appearing in the second image is the same between the first frame and the second frame based on the third position and the fourth position; a second determination unit that determines, based on the first position and the third position, whether a person appearing in the first image and a person appearing in the second image in the first frame are the same as each other, and that determines, based on the second position and the fourth position, whether a person appearing in the first image and a person appearing in the second image in the second frame are the same as each other; The person identification device according to claim 1 , comprising:

3. The second determination unit determining whether the person appearing in the first image and the person appearing in the second image are the same in the first frame based on the orientation of the feet of the person appearing in the first frame of the first image, the orientation of the feet of the person appearing in the first frame of the second image, the first position, and the third position; determining whether the person appearing in the first image and the person appearing in the second image are the same in the second frame based on the orientation of the feet of the person appearing in the second frame of the first image, the orientation of the feet of the person appearing in the second frame of the second image, the second position, and the fourth position; The person identification device according to claim 2 .

4. The second determination unit when an angle formed between a direction of a foot of a person appearing in a first frame of the first image and a direction of a foot of a person appearing in a first frame of the second image is equal to or smaller than a first angle and a distance between the first position and the third position is equal to or smaller than a threshold, determining that the person appearing in the first image and the person appearing in the second image are the same in the first frame; When an angle formed between a direction of a foot of the person appearing in the second frame of the first image and a direction of a foot of the person appearing in the second frame of the second image is equal to or smaller than the first angle, and a distance between the second position and the fourth position is equal to or smaller than the threshold value, it is determined that the person appearing in the first image and the person appearing in the second image are the same in the second frame. The person identification device according to claim 3 .

5. the first determination unit determines the first position, the second position, the third position, and the fourth position using a machine learning model; the machine learning model outputs a first index indicating a likelihood of the first position, a second index indicating a likelihood of the second position, a third index indicating a likelihood of the third position, and a fourth index indicating a likelihood of the fourth position; the integrating unit integrates the first position and the third position by a weighted average in which the first index is a weight for the first position and the third index is a weight for the third position, and integrates the second position and the fourth position by a weighted average in which the second index is a weight for the second position and the fourth index is a weight for the fourth position. The person identification device according to claim 1 .

6. The integration unit 2. The person identification device according to claim 1, wherein the first location and the third location are integrated by selecting either the first location or the third location, and the second location and the fourth location are integrated by selecting either the second location or the fourth location.

7. the first determination unit determines the first position, the second position, the third position, and the fourth position using a machine learning model; the machine learning model outputs a first index indicating a likelihood of the first position, a second index indicating a likelihood of the second position, a third index indicating a likelihood of the third position, and a fourth index indicating a likelihood of the fourth position; 7. The person identification device according to claim 6, wherein the integration unit selects either the first position or the third position based on the first index and the third index, and selects either the second position or the fourth position based on the first index and the third index.

8. Based on a first image in which a target area is viewed from a first direction, determining a first position on a 3D model of the target area in a first frame and a second position on the 3D model in a second frame subsequent to the first frame, with respect to a person appearing in the first image; and based on a second image in which the target area is viewed from a second direction different from the first direction, determining a third position on the 3D model in the first frame and a fourth position on the 3D model in the second frame, with respect to the person appearing in the second image; When it is determined based on the first position and the second position that the person appearing in the first image is the same between the first frame and the second frame, the person appearing in the first image is determined to be an existing person whose positional trajectory is being captured; when it is determined based on the first position and the second position that the person appearing in the first image is different between the first frame and the second frame, the person appearing in the first image is determined to be a new person whose positional trajectory is not being captured; when it is determined based on the third position and the fourth position that the person appearing in the second image is the same between the first frame and the second frame, the person appearing in the second image is determined to be the existing person; and when it is determined based on the third position and the fourth position that the person appearing in the second image is different between the first frame and the second frame, the person appearing in the second image is determined to be the new person; When it is determined based on the first position and the third position that the person appearing in the first image and the person appearing in the second image in the first frame are the same, the first position and the third position are integrated; and when it is determined based on the second position and the fourth position that the person appearing in the first image and the person appearing in the second image in the second frame are the same, the second position and the fourth position are integrated; A person identification method comprising:

Citation Information

Patent Citations

  • Object position estimation device

    JP2024002538A