Identification method, identification program, and information processing device
By generating positional correspondences between image elements between cameras, the problem of identifying unmarked targets in surveillance camera terminals was solved, achieving high-precision identification of unmarked persons.
Patent Information
- Application Number
- CN202380098718.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-01
- Publication Date
- 2025-12-30
AI Technical Summary
Existing surveillance camera terminals have difficulty identifying targets without markings in overlapping areas.
By comparing the feature information of people in the images across multiple cameras, the positional correspondence of elements between the images from each camera is generated, thereby enabling the re-identification of people.
No prior marking is required, reducing marking time, lowering processing load, improving identification accuracy, and reducing errors caused by changes in viewpoint and lighting.
Smart Images

Figure CN121241376A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to an identification method, an identification program, and an information processing apparatus. BACKGROUND
[0002] In various utilization scenarios such as prevention, marketing analysis, and customer action analysis, tracking technology using images of video cameras is flexibly applied. As the utilization scenarios of the tracking technology are thus expanded, the importance of the task of Re-ID (Re-Identification) that identifies the identity of a target photographed by different video cameras is gradually increasing.
[0003] As one of the technologies related to Re-ID, a surveillance camera terminal is proposed as follows (for example, refer to Patent Literature 1). For example, four points marked on the ground in an overlapping region monitored by a plurality of surveillance camera terminals at the time of setting the surveillance camera terminal are marked as a reference, and the positions of targets are detected from frame images respectively photographed by the plurality of surveillance camera terminals. Further, the positions of the targets detected for each frame image of the respective surveillance camera terminals are converted to positions in a common coordinate system in accordance with coordinate conversion parameters calculated using the four points marked. Based on the distances between the targets converted to the positions in the common coordinate system in this way, the targets located in the overlapping region are identified.
[0004] PRIOR ART DOCUMENTS
[0005] PATENT LITERATURE
[0006] Patent Literature 1: International Publication No. 2011 / 010490 SUMMARY
[0007] PROBLEMS TO BE SOLVED BY THE INVENTION
[0008] However, in the related art represented by the above-described surveillance camera terminal, there is a problem that it is difficult to identify a target in the case where no mark is marked in the overlapping region.
[0009] In one aspect, an object is to provide an identification method, an identification program, and an information processing apparatus capable of achieving identification of a target without marking.
[0010] MEANS FOR SOLVING PROBLEMS
[0011] The identification method of one way is executed by a computer as follows: in the case where a first person is detected from a first image taken using a first camera and a second person is detected from a second image taken using a second camera, relationship information associating a position in the first image where the first person is detected and a position in the second image where the second person is detected is generated, and the first person and the second person are identified based on feature information of the first person and feature information of the second person.
[0012] Inventive Effects
[0013] According to one embodiment, identification of a target can be achieved without a marker. BRIEF DESCRIPTION OF DRAWINGS
[0014] Figure 1 is a block diagram showing an example of a functional configuration of an information processing apparatus.
[0015] Figure 2 is a diagram showing an example of person re-identification.
[0016] Figure 3 is a diagram showing an example of a configuration of a camera.
[0017] Figure 4 is a diagram showing an example of person detection.
[0018] Figure 5 is a diagram showing an example of generation of a correspondence relationship between positions of image elements.
[0019] Figure 6 is a diagram showing an example of a configuration of a camera.
[0020] Figure 7 is a diagram showing an example of generation of a correspondence relationship between positions of image elements.
[0021] Figure 8 is a diagram showing an example of a configuration of a camera.
[0022] Figure 9 is a diagram showing an example of person re-identification.
[0023] Figure 10 is a diagram showing an example of a configuration of a camera.
[0024] Figure 11 is a flowchart showing steps of an overall process of an information processing apparatus.
[0025] Figure 12 is a flowchart showing steps of a second person re-identification process.
[0026] Figure 13 is a diagram showing an example of person detection.
[0027] Figure 14 It is a diagram representing the generated example of the positional correspondence between image elements.
[0028] Figure 15 This is a diagram illustrating inter-block interpolation.
[0029] Figure 16 This is a diagram representing a structural example of corresponding relational data.
[0030] Figure 17 This is a flowchart illustrating the steps involved in the re-identification process for the second person.
[0031] Figure 18 This is a diagram illustrating an example of login in an application.
[0032] Figure 19 This is a flowchart illustrating the steps of a person tracking process based on an application example.
[0033] Figure 20 This is a diagram illustrating an example of a hardware structure. Detailed Implementation
[0034] Hereinafter, embodiments of the identification method, identification procedure, and information processing apparatus of this application will be described with reference to the accompanying drawings. In each embodiment, only one example or side view is shown; through such illustration, the range of values, functions, and application scenarios are not limited. Furthermore, the embodiments can be appropriately combined without contradicting the processing content.
[0035] <Example 1>
[0036] <Overall Structure>
[0037] Figure 1 This is a block diagram illustrating an example of the functional structure of an information processing device. Figure 1 The information processing device 10 shown provides a multi-camera tracking function that uses cameras 30A to 30N to track targets.
[0038] Information processing device 10 is an example of a computer that provides the aforementioned multi-camera tracking function. For example, information processing device 10 can be implemented to provide the aforementioned multi-camera tracking function to a locally deployed server. In addition, information processing device 10 can also provide the aforementioned multi-camera tracking function as a cloud service by implementing it as a PaaS (Platform as a Service) or SaaS (Software as a Service) application.
[0039] like Figure 1As illustrated, the information processing apparatus 10 is connected with the cameras 30A to 30N via a network NW. The network NW can be realized by any kind of communication network such as the Internet, a LAN (Local Area Network), or the like, regardless of whether it is wired or wireless.
[0040] The cameras 30A to 30N are imaging devices that capture images. Hereinafter, the cameras 30A to 30N can be referred to as "cameras 30" without distinguishing the individual cameras 30A to 30N.
[0041] The cameras 30A to 30N can be configured such that the entire region to be tracked as a support object is included in the imaging ranges of the cameras 30A to 30N by the multi-camera tracking function described above. In this case, the cameras 30 can be arranged so that the imaging ranges of the cameras 30 overlap each other. For example, the imaging range of one camera 30 can overlap the imaging ranges of one or more other cameras 30.
[0042] Further, the actions of the cameras 30A to 30N to capture images can be synchronized among the cameras 30A to 30N. In this way, the images captured by the cameras 30A to 30N can be transmitted to the information processing apparatus 10 in frame units.
[0043] <Multi-camera tracking>
[0044] Hereinafter, as a target to be tracked by the multi-camera tracking function described above, a "person" is exemplified, but the target is not limited to this and can be any object such as a moving body. For example, the target can be a living body other than a person such as a horse or a dog, or a non-living body such as a vehicle with two or four wheels or a flying body such as a drone.
[0045] For example, as use cases of the multi-camera tracking function, there can be mentioned monitoring, marketing analysis, and action analysis of customers, targeting public facilities such as stations, commercial facilities such as shopping centers, or complex facilities.
[0046] In such use cases, a person to be tracked can move within the imaging ranges of the plurality of cameras 30A to 30N. In this case, in order to suppress the interruption of the activity route of the person by tracking, in a case where a person mapped in a certain camera is mapped in another camera, the viewpoints of the cameras for tracking can be switched by associating the persons with each other as the same person.
[0047] <Person re-identification>
[0048] Thus, as a part of the above-described multi-camera tracking function, the information processing apparatus 10 performs person re-identification that identifies the identity of a person photographed by the plurality of cameras 30A to 30N, so-called Person Re-ID.
[0049] Figure 2 is a schematic diagram that represents an example of person re-identification. In Figure 2 , for ease of explanation, a top view of the facility that is the subject of photographing by the camera 30A and the camera 30B is schematically shown, and the field of view angle of the camera 30A and the camera 30B is schematically shown in the top view. Further, Figure 2 a bounding box (Bounding Box) of a target of which the category is "Person" is shown as a result of target detection for the image 20A captured by the camera 30A. Hereinafter, the bounding box is sometimes labeled as "Bbox". Further, a Bbox of a target of which the category is "Person" is shown as a result of target detection for the image 20B captured by the camera 30B.
[0050] As Figure 2 indicated, Person Re-ID is a task of associating the same person, that is, the person represented by the same hatching, photographed by different plurality of cameras 30, that is, the camera 30A and the camera 30B.
[0051] As one example, person re-identification can be achieved by collating feature information extracted from images corresponding to the regions of the Bboxes. Hereinafter, the image corresponding to the region of the Bbox is sometimes labeled as "Bbox image". As an example of the feature information, for example, a feature amount (feature vector) obtained by embedding the Bbox image into a feature space can be cited. By evaluating the similarity or distance of pairs of such feature information, person re-identification can be achieved.
[0052] For example, in the example shown in Figure 2 , the Bbox 21 detected from the image 20A and the Bbox 24 detected from the image 20B are identified as the same person. In this case, the same ID "1" is assigned to the Bbox 21 and the Bbox 24. Further, the Bbox 22 detected from the image 20A and the Bbox 23 detected from the image 20B are identified as the same person. In this case, the same ID "2" is assigned to the Bbox 22 and the Bbox 23.
[0053] <One aspect of the problem>
[0054] However, as explained in the above Background Art section, in person re-identification typified by the above-described monitoring camera terminal, there is an aspect that it is difficult to perform identification of a target without a marker placed in an overlapping region.
[0055] <One aspect of the solution to the problem>
[0056] Therefore, the information processing apparatus 10 of the present embodiment generates the correspondence of the positions of the elements between the images of the respective cameras based on the detected positions of the person detected from the images of the respective cameras, in the context of performing person re-identification by collating the feature information of the person within the images between the plurality of cameras. In addition, the "element" referred to here can be a pixel or a block of a collection of pixels as an image element.
[0057] Figure 3 is a diagram showing a configuration example of the cameras 30. In Figure 3 , the camera 30A and the camera 30B are schematically shown as top views of the facilities that are the photographic subjects, and the field angles of the camera 30A and the camera 30B are shown in the top views.
[0058] In such a configuration of the cameras 30, in Figure 3 , an example is shown in which a certain person moves from the position in frame t (hatched) to the position in frame t+1 (dotted hatching of the point). In this case, in each of the frames of frame t and frame t+1, the detection result of the person shown in Figure 4 is obtained for each of the camera 30A and the camera 30B, and the correspondence of the positions of the image elements shown in Figure 5 is generated.
[0059] Figure 4 is a diagram showing an example of detection of a person. Figure 5 is a diagram showing an example of generation of the correspondence of the positions of the image elements. In Figure 4 , the detected positions of the person in each of the frames of frame t and frame t+1 are shown for each of the camera 30A and the camera 30B. Also, in Figure 4 , the middle point of the bottom side of the Bbox contained in the four sides is plotted with a white circular mark as the detected position of the person.
[0060] As shown in Figure 4 , as an example, the image photographed by the camera 30 is divided into 24 blocks of four rows and six columns. Such a block is taken as an element of the image of the camera 30, and the correspondence of the positions of the elements to each other is recorded between the camera 30A and the camera 30B in units of frames.
[0061] For example, in the image 20A photographed by the camera 30A in frame t, the detected position of the person becomes the block with the block number "15", on the other hand, in the image 20B photographed by the camera 30B in frame t, the detected position of the person becomes the block with the block number "14". In this case, as shown in Figure 5As shown, the block number "15" of the image 20A of the camera 30A is recorded as the correspondence between the block number "14" of the image 20B of the camera 30B.
[0062] Furthermore, in image 21A of frame t+1 captured by camera 30A, the detected position of the person is in block number "21". On the other hand, in image 21B of frame t+1 captured by camera 30B, the detected position of the person is in block number "13". In this case, as... Figure 5 As shown, the block number "21" of the image 21A of the camera 30A is recorded as the correspondence between the block number "13" of the image 21B of the camera 30B.
[0063] Thus, in Figure 4 The example shown illustrates the detection positions of the characters in frame t and frame t+1, but it is clear that by overlapping with further time after frame t+2, more correspondences between the positions of the elements can be accumulated.
[0064] In addition, Figure 3 The example shown illustrates two cameras 30, camera 30A and camera 30B, but the number of cameras 30 is not limited to two; there can be three or more cameras 30.
[0065] Figure 6 This is a diagram illustrating an example configuration of camera 30. Figure 6 For ease of explanation, a top view of the facility where the three cameras 30A to 30C are the subjects of the photographs is shown, and the field of view of the cameras 30A to 30C is shown in the top view.
[0066] Figure 7 This is a graph representing a generation example of the correspondence between images from multiple cameras. Figure 7 In the diagram, the detection positions of the person in frame t are shown, categorized by camera positions 30A to 30C. Furthermore, in... Figure 7 In the middle, the midpoint of the bottom edge of the four sides contained in the Bbox is drawn as the detection position of the character.
[0067] like Figure 7 As shown, as an example, the image captured by camera 30 is divided into 24 blocks in four rows and six columns. These blocks are used as elements of the image from camera 30, and the positional correspondence between the elements is recorded frame by frame between cameras 30A and 30C.
[0068] For example, in the image 20A taken by the camera 30A at frame t, the detected position of the person becomes a block with block number "15". In addition, in the image 20B taken by the camera 30B at frame t, the detected position of the person becomes a block with block number "14". Also, in the image 20C taken by the camera 30C at frame t, the detected position of the person becomes a block with block number "7".
[0069] In this case, the correspondence relation of the block number "15" as an element of the image 20A of the camera 30A, the block number "14" as an element of the image 20B of the camera 30B, and the block number "7" as an element of the image 20C of the camera 30C is recorded.
[0070] Further, in Figure 6 the detected position of the person in frame t is exemplified, but by overlapping further time lapses after frame t+1, more correspondence relation data 13B of positions of elements each other can be accumulated.
[0071] Thus, in the context of the person re-identification that collates the feature information of the person within the images between the plurality of cameras, the correspondence relation of the positions of the elements each other between the images of the cameras 30 can be generated. According to such correspondence relation, the person re-identification can be performed independently of the matching based on the feature information, and the person re-identification can be realized in a different logic.
[0072] Figure 8 is a diagram that shows an example of the arrangement of the cameras 30. In Figure 8 , for convenience of explanation, the three cameras 30A to 30C are schematically shown as a plan view of the facility that is the object of photographing, and the field of view angles of the cameras 30A to 30C are schematically shown in the plan view.
[0073] In such arrangement of the cameras 30, in Figure 8 , an example in which two persons, person A and person B, exist in the photographing ranges of the three cameras 30A to 30C is shown. In this case, by collating the combination of the detected positions of the persons obtained from each of the images 20A to 20C taken by the three cameras 30A to 30C with the correspondence relation data 13B shown in Figure 7 , the person re-identification can be performed.
[0074] Figure 9 is a diagram that shows an example of the person re-identification. In Figure 9 , the images 20A to 20C taken by the cameras 30A to 30C, respectively, and the detection results of the persons are shown. Also, in Figure 9 , as the detected positions of the persons, the middle points of the bottom sides of the Bboxes of the targets of the class "Person" are plotted with white circular marks.
[0075] As shown in FIG. 10, in the image 20A captured by the camera 30A, as the detection position d11 of the person, the block number "7" is detected, and as the detection position d12 of the person, the block number "15" is detected. In addition, in the image 20B captured by the camera 30B, as the detection position d21 of the person, the block number "11" is detected, and as the detection position d22 of the person, the block number "14" is detected. Further, in the image 20C captured by the camera 30C, as the detection position d31 of the person, the block number "7" is detected, and as the detection position d32 of the person, the block number "20" is detected. Figure 9 In this case, the combination of the detection position d12 of the person of the camera 30A, the detection position d21 of the person of the camera 30B, and the detection position d31 of the person of the camera 30C coincides with the data entry of the first line of the correspondence relationship data 13B. Thus, it is determined that the person A appearing in the block number "15" of the camera 30A, the block number "14" of the camera 30B, and the block number "7" of the camera 30C is the same person.
[0076] In addition, the combination of the detection position d11 of the person of the camera 30A, the detection position d22 of the person of the camera 30B, and the detection position d32 of the person of the camera 30C coincides with the data entry of the second line of the correspondence relationship data 13B. Thus, it is determined that the person B appearing in the block number "7" of the camera 30A, the block number "11" of the camera 30B, and the block number "20" of the camera 30C is the same person.
[0077] As described above, the information processing apparatus 10 according to the present embodiment does not use a marker such as a tape that needs to be set in advance, but generates a correspondence relationship of positions of elements to each other between images of the respective cameras 30 using the target as a clue.
[0078] Therefore, according to the information processing apparatus 10 of the present embodiment, person re-identification can be achieved without a marker. Therefore, the effort of setting a marker such as a tape in advance in a space to be a shooting target can be reduced. In addition, in a case where person re-identification is performed using a correspondence relationship of positions of elements to each other between images of the respective cameras 30, it is not necessary to perform coordinate conversion to a common coordinate system as in the person re-identification performed in the surveillance camera terminal described above, so the processing load can be reduced. Furthermore, in the person re-identification performed by the surveillance camera terminal described above, as the position of the person as a subject moves away from the 4-point marker, the error of the coordinate conversion increases, but it is possible to suppress the generation of such an error and the increase of the error.
[0079]
[0080] In addition, the information processing apparatus 10 according to the present embodiment can achieve robust person re-identification as compared to person re-identification based on comparison of feature information. That is, in a case where person re-identification is performed by comparison of feature information, as the change in viewpoint between the cameras 30, the change in lighting conditions becomes larger, the accuracy of person re-identification decreases and becomes larger, and as the resolution of the cameras 30 becomes lower, the accuracy of person re-identification decreases and becomes larger. On the other hand, in a case where person re-identification is performed using the correspondence relation of positions of elements between images of the respective cameras 30, as with the information processing apparatus 10 according to the present embodiment, the person re-identification is less likely to be affected by the change in viewpoint, the change in lighting conditions, the resolution of the cameras, and the like, and thus it is possible to suppress the decrease in accuracy of person re-identification.
[0081] <Structure of the information processing apparatus 10>
[0082] Figure 1 Blocks related to the multi-camera tracking function of the information processing apparatus 10 are schematically shown. As shown in Figure 1 , the information processing apparatus 10 has a communication control section 11, a storage section 13, and a control section 15. In addition, in the present embodiment, only functional sections related to the above-described multi-camera tracking function are shown, and functional sections other than those shown can also be provided in the information processing apparatus 10. Figure 1
[0083] The communication control section 11 is a functional section that controls communication with other apparatuses such as the cameras 30A to 30N. As one example, the communication control section 11 can be implemented by a network interface card such as a LAN card. As one aspect, the communication control section 11 can accept images from the cameras 30 in units of frames, or accept dynamic images for a certain period of time. As another aspect, the communication control section 11 can also output the result of multi-camera tracking to an arbitrary external apparatus.
[0084] The storage section 13 is a functional section that stores various data. As one example, the storage section 13 is implemented by a memory inside, outside, or auxiliary to the information processing apparatus 10. For example, the storage section 13 stores ID information 13A and correspondence relation data 13B. In addition, the explanation of the ID information 13A and the correspondence relation data 13B is given in the context of a scenario in which the reference, generation, or registration of the ID information 13A and the correspondence relation data 13B is performed.
[0085] The control section 15 is a functional section that performs overall control of the information processing apparatus 10. For example, the control section 15 can be implemented by a hardware processor. In addition, the control section 15 can also be implemented by hardwired logic. Figure 1 As shown, the control section 15 has an acquisition section 15A, a person detection section 15B, a feature extraction section 15C, a tracking section 15D, a first re-authentication section 15E, a correspondence relationship generation section 15F, and a second re-authentication section 15G.
[0086] The acquisition section 15A is a processing section that acquires images. As one example, the acquisition section 15A can acquire images captured by the camera 30 via the network NW. At this time, the acquisition section 15A can also acquire a dynamic image containing an arbitrary amount of images if it can acquire the images of each frame in real time for each frame.
[0087] The person detection section 15B is a processing section that detects a person from each image captured by the camera 30. As one example, the person detection section 15B can be implemented by a machine learning model that outputs a target Box by inputting an image. Such a machine learning model can also be implemented by a model combining a Transformer and a CNN (Convolutional Neural Network), in addition to YOLO (You Only Look Once), SSD (Single Shot MultiBox Detector), and the like.
[0088] The feature extraction section 15C is a processing section that extracts a feature amount of each person detected from the images captured by the camera 30. As one example, the feature extraction section 15C can be implemented by a machine learning model that outputs a feature amount (feature vector) representing an appearance feature by inputting an image, i.e., a Bbox image. Such a machine learning model can be implemented by a CNN or the like that embeds an input image in a feature space.
[0089] The tracking section 15D is a processing section that tracks persons detected from the images captured by the camera 30. As one example, the tracking section 15D can be implemented by a machine learning model that associates targets between frames, such as MOT (Multiple-Object Tracking), SORT (Simple Online and RealTime Tracking), Deep SORT, and the like.
[0090] More specifically, the tracking section 15D can perform the following processing in parallel for each camera 30. For example, the tracking section 15D assigns the person ID stored in the ID information 13A to the person Bbox detected in the frame of the image acquired by the acquisition section 15A. Here, the ID information 13A can be data associating the person Bbox and the person ID detected from the image of the frame with each frame. At this time, the tracking section 15D calculates the similarity or the distance between the feature vector of the person Bbox extracted by the feature extraction section 15C and the feature vector of the person Bbox detected in the previous frame. Then, the tracking section 15D assigns the person ID associated with the person Bbox having the greatest similarity or the smallest distance among the persons detected in the previous frame to the person Bbox extracted by the feature extraction section 15C. At this time, in the case where the greatest similarity is below a threshold value or the smallest distance is above a threshold value, the tracking section 15D can assign a new person ID to the person Bbox extracted by the feature extraction section 15C and assign the new person ID. Then, the tracking section 15D registers the assignment result of the person ID of the person Bbox as the assignment result of the person ID of the current frame of the ID information 13A.
[0091] Further, here, only the previous frame is taken as the object of calculation, but any number of past frames can also be taken as the object of calculation of the similarity or the distance. In addition, here, an example of performing tracking by evaluation of the similarity or the distance of the appearance features is given, but tracking can also be performed based on the degree of overlap of the Boxes with each other between frames.
[0092] The first re-identification section 15E is a processing section that performs person re-identification, so-called Person Re-ID, by collating the feature vectors of the persons within the images between the plurality of cameras 30A to 30N.
[0093] More specifically, the first re-identification section 15E collates the feature vector of the Bbox image detected from the image of one camera 30 and the feature vector of the Bbox image detected from the image of the other camera 30 for each pair of 2 cameras 30. For example, the first re-identification section 15E assigns the same person ID to the pair of Bboxes having the greatest similarity between the feature vectors or the pair of Bboxes having the smallest distance between the feature vectors. Here, as an example, it is assumed that the person ID that is assigned first among the paired person IDs assigned by the tracking section 15D, that is, the old person ID, is commonly assigned. At this time, in the case where the greatest similarity is below a threshold value or the smallest distance is above a threshold value, the same person ID can also not be assigned, and each of the paired person IDs assigned by the tracking section 15D is maintained. After that, the assignment result of the person ID of the person Bbox by the first re-identification section 15E is updated as the assignment result of the person ID of the current frame of the ID information 13A.
[0094] The correspondence generation section 15F is a processing section that generates a correspondence of positions of elements to each other between images of each camera 30. As one example, the correspondence generation section 15F starts processing in a case where the correspondence data 13B is not completed. The "not completed" referred to here means a state where the correspondence of positions, which can be used for person re-identification by the second re-identification section 15G, is not accumulated in the correspondence data 13B. The completion or non-completion of such correspondence data 13B can be determined with reference to the following criteria.
[0095] As one aspect, the correspondence generation section 15F determines that the correspondence data 13B is not completed in a case where the elapsed time from the start time of the generation of the correspondence data 13B is less than a threshold value, and determines that the correspondence data 13B is completed in a case where the elapsed time is equal to or more than the threshold value.
[0096] As another aspect, the correspondence generation section 15F determines that the correspondence data 13B is not completed in a case where the total number of data entries registered in the correspondence data 13B is less than a threshold value, and determines that the correspondence data 13B is completed in a case where the total number of data entries is equal to or more than the threshold value.
[0097] As a further aspect, the correspondence generation section 15F calculates a coverage rate of blocks from the data entries registered in the correspondence data 13B. The "coverage rate" referred to here means a proportion of blocks that encompass the correspondence of positions among all blocks, and is, for example, a value obtained by dividing the number of types of blocks in which the correspondence of positions is registered in the correspondence data 13B by the total number of blocks included in the image. In a case where the coverage rate of blocks is less than a threshold value, the correspondence data 13B is determined to be not completed, and in a case where the coverage rate of blocks is equal to or more than the threshold value, the correspondence data 13B is determined to be completed.
[0098] Then, in a case where the correspondence data 13B is not completed, the correspondence generation section 15F determines whether the number of persons present in the space of the photographic subject is less than a threshold value, for example, "2". Such number determination can be achieved by determining whether the number of targets of the class name "Person" detected from all images photographed by the cameras 30A to 30N is less than a threshold value, for example, "2".
[0099] Further, the above-described number determination is not limited to being achieved by image processing, and can be achieved by all other means. For example, it is also possible to accept user input of a point in time or a period in which the number of persons is less than a threshold value, for example, designation of a frame number, and the like, via a user interface not shown. Further, the number of terminals that receive beacons from a beacon receiver arranged in the space of the photographic subject from a user terminal can be determined to be less than a threshold value.
[0100] Here, in a case where the number of persons existing in the space of the photographic subject is smaller than a threshold value, the correspondence relationship generating section 15F performs the following processing. That is, the correspondence relationship generating section 15F adds, to the correspondence relationship data 13B, a data entry including a combination of block numbers corresponding to the detection positions of the persons detected by the person detecting section 15B for each camera 30 in the frame of the image of each camera 30 acquired by the acquiring section 15A. Further, the generation of the correspondence relationship of the positions can be performed as already described using Figures 3-7 .
[0101] An example of the saving of such correspondence relationship data 13B will be described. For example, the correspondence relationship data 13B can also be saved in a list form. In this case, as already described using Figure 5 , Figure 7 and Figure 9 , the correspondence relationship data 13B saves, as a data entry, a combination of block numbers corresponding to the detection positions of the persons in the images of each camera 30.
[0102] Further, the correspondence relationship data 13B can also be stored in a matrix form. Figure 10 is a diagram showing an example of the structure of the correspondence relationship data 13B. Figure 10 The correspondence relationship data 13B1 in which the correspondence relationship of the positions of the elements of each other between the images of the two cameras 30, the camera 30A and the camera 30B, is saved in a list form is shown in
[0103] As shown in Figure 10 , the correspondence relationship data 13B1 in the list form can be converted into the correspondence relationship data 13B2 in a matrix form. The column in the correspondence relationship data 13B2 in the matrix form refers to the block number of the camera 30A, and the row refers to the block number of the camera 30B. A binary of "0" or "1" is held in each element of the correspondence relationship data 13B2 in the matrix form, indicating whether there is a correspondence relationship. For example, in a case where there is a correspondence relationship, the value of "1" is stored, and on the other hand, in a case where there is no correspondence relationship, the value of "0" is stored.
[0104] For example, the data entry of the first row of the correspondence relationship data 13B1 in the list form, that is, the data entry of the slanting hatching, indicates that there is a correspondence relationship with the block number "1" of the camera 30A and the block number "0" of the camera 30B. The recording of the correspondence relationship equivalent thereto is realized by storing "1" in the element of the second column of the correspondence relationship data 13B2 in the matrix form.
[0105] Further, the data entry of the fourth row of the list-form correspondence relation data 13B1, that is, the data entry of the black-and-white reverse display indicates that there is a correspondence relation with the block number "21" of the camera 30A and the block number "23" of the camera 30B. The recording of the correspondence relation equivalent thereto is realized by storing "1" in the element of the twenty-third row and the twenty-first column of the matrix-form correspondence relation data 13B2.
[0106] Further, in Figure 10 , the two cameras 30, the camera 30A and the camera 30B are exemplified, but the correspondence relation of the positions with respect to three or more cameras 30 can also be similarly stored. That is, regardless of the number of cameras 30, as long as the correspondence relation of each pair of storage positions is stored. For example, in the case where the number of cameras 30 is N, as long as N x (N-1) sets of Figure 10 correspondence relation data 13B2 shown in the matrix form can be generated.
[0107] Returning to Figure 1 , the second re-identification section 15G is a processing section that performs the re-identification of the person by collating the combination of the detection positions of the person in the images of the respective cameras 30 with the combinations of the image elements between the plurality of cameras registered in the correspondence relation data 13B.
[0108] As an example, the second re-identification section 15G can be activated all the time for each frame, and can also be activated in the case where a certain condition is satisfied, for example, in the case where a person is detected in the image of any one of the cameras 30A to 30N. In this way, as an example, the mode of activation all the time or conditional activation can be selected according to the utilization scenario of the performance of the processor, the memory, or the time period of the information processing apparatus 10 mounted.
[0109] After the activation of the second re-identification section 15G, the second re-identification section 15G performs the following processing from the N cameras 30 of the cameras 30A to 30N in the amount of the combination total number K of pairs of two cameras 30. Hereinafter, one camera 30 included in the kth pair is identified as camera i, and the other camera 30 is identified as camera j.
[0110] For example, the second re-identification section 15G retrieves the combination of the detection positions of the Bbox from the correspondence relation data 13B for each combination of the number L of Bbox detected from the image of the camera i of the kth pair and the number M of Bbox detected from the image of the camera j of the kth pair.
[0111] More specifically, the second re-identification unit 15G retrieves, from the correspondence relation data 13B, a combination of the block number kil corresponding to the detected position of the 1st Bbox of the camera i and the block number kjm corresponding to the detected position of the mth Bbox of the camera j.
[0112] At this time, in a case where the combination of the block number kil and the block number kjm hits, it can be identified that the person of the 1st Bbox of the camera i and the person of the mth Bbox of the camera j are the same person.
[0113] In this case, the second re-identification unit 15G determines whether the person re-identification result of the second re-identification unit 15G and the person re-identification result of the first re-identification unit 15E are inconsistent. Hereinafter, the person re-identification result of the first re-identification unit 15E will be sometimes referred to as the "first person re-identification result", and the person re-identification result of the second re-identification unit 15G will be sometimes referred to as the "second person re-identification result".
[0114] For example, in a case where different person IDs are assigned to the 1st Bbox of the camera i and the mth Bbox of the camera j by the first re-identification unit 15E, it is determined that the first person re-identification result and the second person re-identification result are inconsistent.
[0115] In this case, the second re-identification unit 15G assigns the same person ID to the 1st Bbox of the camera i and the mth Bbox of the camera j. For example, in a case where the same person ID is assigned to the Bbox containing either one of the 1st Bbox of the camera i or the mth Bbox of the camera j by the pair of the other camera 30, the second re-identification unit 15G preferentially assigns the person ID. Further, the second re-identification unit 15G preferentially assigns the person ID of the earlier number among the person IDs assigned by the tracking unit 15, that is, the old person ID, or the person ID for which the number of frames of tracking continues is large. Thereby, the assignment result of the person ID of the person Bbox by the second re-identification unit 15G is updated to the assignment result of the person ID of the current frame of the ID information 13A.
[0116] On the other hand, in a case where the combination of the block number kil and the block number kjm does not hit, it can be recognized that the person of the 1st Bbox of the camera i and the person of the mth Bbox of the camera j are not the same person.
[0117] In this case, the second re-identification unit 15G determines whether the person re-identification result of the second re-identification unit 15G and the person re-identification result of the first re-identification unit 15E are inconsistent.
[0118] For example, in a case where the same person ID is assigned to the 1st Bbox of the camera i and the mth Bbox of the camera j by the first re-identification section 15E, it is determined that the first person re-identification result and the second person re-identification result are inconsistent.
[0119] In this case, the second re-identification section 15G assigns different person IDs to the 1st Bbox of the camera i and the mth Bbox of the camera j. For example, the second re-identification section 15G restores the person IDs assigned to the 1st Bbox of the camera i and the mth Bbox of the camera j, respectively, by the tracking section 15D before the same person ID is assigned by the first re-identification section 15E. Thereby, the assignment result of the person IDs to the person Bbox by the second re-identification section 15G is updated to the assignment result of the person IDs of the current frame of the ID information 13A.
[0120] The ID information 13A thus obtained is nothing more than an example, and can be output to a software that performs processing of monitoring, marketing analysis, action analysis of customers, and the like, or a back-end that provides a service, and the like.
[0121] <PROCESS FLOW>
[0122] Next, the processing flow of the information processing apparatus 10 of the present embodiment will be described. Here, after (1) the overall processing performed by the information processing apparatus 10 is described, (2) the second person re-identification processing is described.
[0123] (1) Overall Processing
[0124] Figure 11 is a flowchart showing the steps of the overall processing of the information processing apparatus 10. Figure 11 The processing shown in is nothing more than an example, and can be performed in a frame unit. For example, Figure 11 As shown in, the acquisition section 15A acquires images captured by each camera 30 via the network NW (step S101).
[0125] Next, the person detection section 15B detects a person from each of the images captured by each camera 30 (step S102). Then, the feature extraction section 15C extracts a feature amount of each person detected from the images captured by the camera 30 (step S103).
[0126] Thereafter, the tracking section 15D assigns a person ID saved in the ID information 13A to each person detected from the images captured by the camera 30 by tracking the person detected from the images captured by the camera 30 between frames (step S104).
[0127] Next, the first re-identification section 15E assigns the same person ID to the same person by performing person re-identification by collating the feature amounts of the person in the image between the plurality of cameras 30A to 30N (step S105).
[0128] Then, in a case where the correspondence relationship data 13B is not completed (step S106 is NO), the correspondence relationship generation section 15F determines whether the number of persons existing in the space of the photographic subject is less than a threshold value, for example, "2" (step S107).
[0129] At this time, in a case where the number of persons existing in the space of the photographic subject is less than the threshold value (step S107 is YES), the correspondence relationship generation section 15F performs the following processing. That is, the correspondence relationship generation section 15F appends a data entry including a combination of the block numbers corresponding to the detection positions of the persons detected by each camera 30 in the current frame to the correspondence relationship data 13B (step S108).
[0130] On the other hand, in a case where the correspondence relationship data 13B is completed (step S106 is NO), the second re-identification section 15G performs the following processing. That is, the second re-identification section 15G performs person re-identification by collating the combination of the detection positions of the persons in the images of the respective cameras 30 and the combination of the image elements between the plurality of cameras registered in the correspondence relationship data 13B (step S109).
[0131] After the assignment results of the person IDs assigned by the tracking section 15D, the first re-identification section 15E, and the second re-identification section 15G per person are output to an arbitrary output destination (step S110), the processing is ended.
[0132] (2) Second Person Re-identification Processing
[0133] Figure 12 is a flowchart showing the steps of the second person re-identification processing. Figure 12 The processing shown corresponds to Figure 11 the processing of step S109. As Figure 12 shown, the second re-identification section 15G performs a loop processing 1 that repeatedly performs the processing from step S301 to step S309 in an amount of the combination total number K of the combinations of the combinations of 2 cameras 30 from the N cameras 30 of the cameras 30A to 30N. Further, in Figure 12 an example of repeatedly performing the processing from step S301 to step S309 is shown, but the processing can also be performed in parallel.
[0134] Further, the second re-identification section 15G executes a loop process 2 that repeatedly performs the process from step S301 to step S308 in the amount of the number L of the Box detected from the image of the camera i of the kth pair. Further, in the case where the process from step S301 to step S308 is repeated, the process can also be executed in parallel. Figure 12 In the case where the process from step S301 to step S308 is repeated, the process can also be executed in parallel.
[0135] Further, the second re-identification section 15G executes a loop process 2 that repeatedly performs the process from step S301 to step S307 in the amount of the number M of the Box detected from the image of the camera j of the kth pair. Further, in the case where the process from step S301 to step S307 is repeated, the process can also be executed in parallel. Figure 12 In the case where the process from step S301 to step S307 is repeated, the process can also be executed in parallel.
[0136] That is, the second re-identification section 15G retrieves, from the correspondence relation data 13B, a combination of the block number kil corresponding to the detection position of the 1st Box of the camera i and the block number kjm corresponding to the detection position of the mth Box of the camera j (step S301).
[0137] At this time, in the case where the combination of the block number kil and the block number kjm hits (step S302 is), it can be identified that the person of the 1st Box of the camera i and the person of the mth Box of the camera j are the same person.
[0138] In this case, the second re-identification section 15G determines whether the second person re-identification result obtained through the branch of step S302 is is inconsistent with the first person re-identification result obtained in step S105 (step S303).
[0139] Then, in the case where the first person re-identification result and the second person re-identification result are inconsistent (step S303 is), the second re-identification section 15G allocates the same person ID to the 1st Box of the camera i and the mth Box of the camera j (step S304). Further, in the case where the first person re-identification result and the second person re-identification result are consistent (step S303 No), the process of step S304 is skipped.
[0140] On the other hand, in the case where the combination of the block number kil and the block number kjm does not hit (step S302 No), it can be recognized that the person of the 1st Box of the camera i and the person of the mth Box of the camera j are not the same person.
[0141] In this case, the second re-identification unit 15G determines whether the second person re-identification result obtained through the NO branch of Step S302 is inconsistent with the first person re-identification result obtained in Step S105 (Step S305).
[0142] Then, in a case where the first person re-identification result and the second person re-identification result are inconsistent (YES in Step S305), the second re-identification unit 15G assigns different person IDs to the 1th Bbox of the camera i and the mth Bbox of the camera j (Step S306).
[0143] After that, the loop counter m that counts the M Bboxes detected from the image of the camera j of the kth pair is incremented, and the loop processing 3 is repeated until the loop counter m exceeds M.
[0144] By repeating such loop processing 3, each of the 1th Bbox of the L Bboxes detected from the image of the camera i of the kth pair is compared with each of the M Bboxes detected from the image of the camera j of the kth pair.
[0145] After that, the loop counter 1 that counts the L Bboxes detected from the image of the camera i of the kth pair is incremented, and the loop processing 2 is repeated until the loop counter 1 exceeds L.
[0146] By repeating such loop processing 2, all combinations of the L Bboxes detected from the image of the camera i of the kth pair and the M Bboxes detected from the image of the camera j of the kth pair are compared.
[0147] After that, the loop counter k that counts the total number of combinations of pairs of two cameras 30 combined from the N cameras 30 of the cameras 30A to 30N is incremented, and the loop processing 1 is repeated until the loop counter k exceeds K.
[0148] By repeating such loop processing 1, the second person re-identification processing ends for all combinations of pairs of two cameras 30 combined from the N cameras 30 of the cameras 30A to 30N.
[0149] <One aspect of the effect>
[0150] As described above, the information processing apparatus 10 according to the present embodiment generates a correspondence relationship between elements between images of each camera based on the detection positions of persons detected from the images of each camera in the context of performing person re-identification by comparing feature information of persons in images with a plurality of cameras.
[0151] Therefore, the information processing apparatus 10 of the present embodiment does not need the tape or the like set in advance, but generates the correspondence relation of the positions of the elements to each other between the images of the respective cameras 30 using the target as a clue to be tracked.
[0152] Therefore, according to the information processing apparatus 10 of the present embodiment, person re-identification can be implemented without a marker. Therefore, the work of setting a marker such as a tape in advance in a space to be photographed can be reduced. Also, in a case where person re-identification is performed using the correspondence relation of the positions of the elements to each other between the images of the respective cameras 30, as in person re-identification performed in the surveillance camera terminal described above, coordinate conversion into a common coordinate system is not performed, so the processing load can be reduced. Further, in person re-identification performed in the surveillance camera terminal described above, as the position of the person as a subject departs from the marker at 4 o'clock, the error of coordinate conversion increases, but the generation of such an error and the increase in the error can be suppressed.
[0153] In addition, according to the information processing apparatus 10 of the present embodiment, robust person re-identification can be implemented compared to person re-identification based on the comparison of feature information. That is, in a case where person re-identification is performed by comparison of feature information, as the change in the viewpoint between the cameras 30 and the change in the lighting condition become larger, the accuracy of person re-identification decreases and becomes larger, and as the resolution of the cameras 30 becomes lower, the accuracy of person re-identification decreases and becomes larger. On the other hand, in a case where person re-identification is performed using the correspondence relation of the positions of the elements to each other between the images of the respective cameras 30 as in the information processing apparatus 10 of the present embodiment, the change in the viewpoint, the change in the lighting condition, the resolution of the cameras, and the like do not easily affect, so the accuracy of person re-identification can be suppressed from decreasing.
[0154] <Embodiment 2>
[0155] In addition, the embodiments related to the apparatus disclosed so far have been described, but the present application can be implemented in various different ways other than the above-described embodiments. Therefore, other embodiments included in the present application will be described below.
[0156] <Orientation of a person>
[0157] In Embodiment 1 described above, an example in which the correspondence relation of the positions of the blocks to each other between the plurality of cameras is registered in the correspondence relation data 13B is presented, but other data can be further registered. As an example, in the correspondence relation data 13B, the orientation of a person, for example, the moving direction, can be further registered for each camera 30.
[0158] Figure 13 is a diagram showing an example of detection of a person. Figure 14is a drawing of a generation example of a correspondence relation indicating positions between image elements. In Figure 13 , in the case of the camera arrangement and the person movement shown in Figure 3 , the detected positions of the person in each of frames t and t+1 are shown for camera 30A and camera 30B. Also, in Figure 13 , as the detected position of the person, the middle point of the bottom edge among the four edges included in the Box is plotted in white. Also, in Figure 13 , as the orientation of the person, the moving direction calculated from the difference between the detected position of the person in the previous frame and the detected position of the person in the current frame is plotted with an arrow.
[0159] As an example, as shown in Figure 13 , the image captured by camera 30 is divided into 24 blocks of four rows and six columns. Such a block is taken as an element of the image of camera 30, and the correspondence relation of the positions of the elements is recorded in frame units between camera 30A and camera 30B.
[0160] For example, in the image 20A captured by camera 30A in frame t, the detected position of the person becomes the block with block number "15", while in the image 20B captured by camera 30B in frame t, the detected position of the person becomes the block with block number "14". In this case, as shown in Figure 14 , the correspondence relation of the block number "15" as an element of the image 20A of camera 30A and the block number "14" as an element of the image 20B of camera 30B is recorded. Also, as shown in Figure 13 , in the image 20A captured by camera 30A in frame t, the arrow in the lower left direction, which is the moving direction, can be calculated from the difference between the detected position of the person in frame t-1 and the detected position of the person in frame t. Therefore, as shown in Figure 14 , the orientation "arrow in the lower left direction" is recorded in correspondence with the data entry of camera 30A in frame t. Also, as shown in Figure 13 , in the image 20B captured by camera 30B in frame t, the arrow in the left direction, which is the moving direction, can be calculated from the difference between the detected position of the person in frame t-1 and the detected position of the person in frame t. Therefore, as shown in Figure 14 , the orientation "arrow in the left direction" is recorded in correspondence with the data entry of camera 30B in frame t.
[0161] Also, as shown in Figure 13 , in the image 21A captured by camera 30A in frame t+1, the detected position of the person becomes the block with block number "21", while in the image 21B captured by camera 30B in frame t+1, the detected position of the person becomes the block with block number "13". In this case, as shown in Figure 14As shown, the block number "21" that is an element of the image 21A of the camera 30A and the block number "13" that is an element of the image 21B of the camera 30B are recorded in correspondence with each other. Further, as shown, the arrow toward the left direction is recorded in correspondence with the camera 30A of the data entry of the frame t+1. Further, as shown, the arrow toward the left direction is recorded in correspondence with the camera 30B of the data entry of the frame t+1. Figure 13 As shown, in the image 21A of the frame t+1 shot by the camera 30A, the arrow toward the left direction can be calculated from the difference between the detection position of the person of the frame t and the detection position of the person of the frame t+1. Therefore, as shown, the arrow toward the left direction is recorded in correspondence with the camera 30A of the data entry of the frame t+1. Further, as shown, in the image 21B of the frame t+1 shot by the camera 30B, the arrow toward the left direction can be calculated from the difference between the detection position of the person of the frame t and the detection position of the person of the frame t+1. Therefore, as shown, the arrow toward the left direction is recorded in correspondence with the camera 30B of the data entry of the frame t+1. Figure 14 As shown, in the image 21A of the frame t+1 shot by the camera 30A, the arrow toward the left direction can be calculated from the difference between the detection position of the person of the frame t and the detection position of the person of the frame t+1. Therefore, as shown, the arrow toward the left direction is recorded in correspondence with the camera 30A of the data entry of the frame t+1. Further, as shown, in the image 21B of the frame t+1 shot by the camera 30B, the arrow toward the left direction can be calculated from the difference between the detection position of the person of the frame t and the detection position of the person of the frame t+1. Therefore, as shown, the arrow toward the left direction is recorded in correspondence with the camera 30B of the data entry of the frame t+1. Figure 13 As shown, in the image 21A of the frame t+1 shot by the camera 30A, the arrow toward the left direction can be calculated from the difference between the detection position of the person of the frame t and the detection position of the person of the frame t+1. Therefore, as shown, the arrow toward the left direction is recorded in correspondence with the camera 30A of the data entry of the frame t+1. Further, as shown, in the image 21B of the frame t+1 shot by the camera 30B, the arrow toward the left direction can be calculated from the difference between the detection position of the person of the frame t and the detection position of the person of the frame t+1. Therefore, as shown, the arrow toward the left direction is recorded in correspondence with the camera 30B of the data entry of the frame t+1. Figure 14 As shown, in the image 21A of the frame t+1 shot by the camera 30A, the arrow toward the left direction can be calculated from the difference between the detection position of the person of the frame t and the detection position of the person of the frame t+1. Therefore, as shown, the arrow toward the left direction is recorded in correspondence with the camera 30A of the data entry of the frame t+1. Further, as shown, in the image 21B of the frame t+1 shot by the camera 30B, the arrow toward the left direction can be calculated from the difference between the detection position of the person of the frame t and the detection position of the person of the frame t+1. Therefore, as shown, the arrow toward the left direction is recorded in correspondence with the camera 30B of the data entry of the frame t+1.
[0162] Thus, in the example shown, the detection positions of the person in the frame t and the frame t+1 are exemplified as an example, but it is obvious that more correspondence of the positions of the elements can be accumulated by overlapping further time lapses after the frame t+2. Figure 13 For example, as shown, by registering the directions of the person in each camera 30 in the correspondence data 13B for each frame, in the case where there are a plurality of persons in the same block, if the directions are different, the persons can be distinguished according to the directions. That is, by collating the directions between the frames, the persons can be distinguished. Thus, in the case where the second re-identification section 15G uses the direction of the person, the update of the correspondence data 13B can be continued after the completion of the correspondence data 13B.
[0163] Figure 14
[0164] <Speed of person>
[0165] As described above, the example where the directions of the person in each camera 30 are registered in the correspondence data 13B for each frame is exemplified, but further other data can be registered. For example, in the correspondence data 13B, the speed or speed ratio of the person, for example, the moving speed can be further registered for each camera 30. Such moving speed can be calculated by converting the moving distance calculated from the detection position of the person of the previous frame and the detection position of the person of the current frame into each unit time. In this case, in the case where there are a plurality of persons in the same block, the persons can be distinguished according to the difference of the speed. That is, the persons can be distinguished by collating the speed between the frames. Thus, in the case where the second re-identification section 15G uses the speed of the person, the update of the correspondence data 13B can be continued after the completion of the correspondence data 13B.
[0166] <Interpolation between blocks>
[0167] For example, when the correspondence generating section 15F registers the combination of the block numbers of each camera 30 in time series to the correspondence data 13B, in a case where the arrangement of the block numbers corresponding to the detection positions of the person detected in time series by each camera 30 is an arrangement moving in any one of the eight directions of up, down, left, right, upper left, lower left, upper right, and lower right, skipping a certain number, for example, one block, the combination of the block numbers sandwiched by the arrangement of the block numbers is additionally registered to the correspondence data 13B.
[0168] Figure 15 is a diagram illustrating the interpolation between blocks. In Figure 15 , an example is shown in which the image 20A and the image 20B captured by the camera 30A and the camera 30B are divided into a total of 24 blocks of four rows and six columns. Also, in Figure 15 , for each block of the image 20A and the image 20B, the presence or absence of the person detection is represented by two values of "0" or "1". As Figure 15 indicated, an example is shown in which the arrangement of the block numbers corresponding to the detection positions of the person detected in time series by the camera 30A is the block number "9", the block number "19", and the arrangement of the block numbers corresponding to the detection positions of the person detected in time series by the camera 30B is the block number "16", the block number "14". In this case, the arrangement of the block numbers of the camera 30A becomes an arrangement skipping one in the lower left direction. Further, the arrangement of the block numbers of the camera 30B is an arrangement skipping one in the right direction. In this case, the combination of the block numbers sandwiched by the arrangement of the block numbers of the camera 30A and the camera 30B, that is, the combination of the block number "14" and the block number "15", is additionally registered in the correspondence data 13B.
[0169] <Frequency>
[0170] In the above-described embodiment 1, an example is shown in which the presence or absence of the positional correspondence is registered in the correspondence data 13B, but the frequency of the positional correspondence can also be maintained by updating the number of times of the positional correspondence in a case where the same combination of the block numbers is detected at the time of generation of the correspondence data 13B. Figure 16 is a diagram showing a structure example of the correspondence data. Figure 16The column in the correspondence data 13B2 in the matrix form indicates the block number of the camera 30A, and the row indicates the block number of the camera 30B. The frequency of the correspondence is held in each element of the correspondence data 13B2 in the matrix form. For example, the element of the first row and the second column of the correspondence data 13B2 in the matrix form indicates that the frequency of the combination of the block number "1" of the camera 30A and the block number "0" of the camera 30B is "5". The frequency can be a measured value of the number of detections of the combination of the block numbers, or can be a value after normalization. For example, the maximum value of the number of detections is determined in advance, and the frequency can be normalized according to the following equation (1). In this way, by holding the frequency of the correspondence of the positions in the correspondence data 13B, in the case where the first person re-identification result and the second person re-identification result are not consistent, the correction that narrows down in the case where the frequency of the combination of the block numbers exceeds the threshold value to make the first person re-identification result consistent with the second person re-identification result can be performed.
[0171] Frequency = number of detections / maximum value of number of detections (1)
[0172] Figure 17 is a flowchart showing the steps of the second person re-identification processing. In Figure 17 , steps in which steps different from the flowchart shown in Figure 12 are performed are given different step numbers.
[0173] As shown in Figure 17 , the processing performed in the branch of step S302 is different from the flowchart shown in Figure 12 . That is, in the case where the combination of the block number kil and the block number kjm is hit (step S302 Yes), the second re-identification section 15G determines whether the frequency of the combination of the block number kil and the block number kjm exceeds the threshold value (step S501).
[0174] Here, in the case where the frequency of the combination of the block number kil and the block number kjm exceeds the threshold value (step S501 Yes), the processing of step S303 is moved to. On the other hand, in the case where the frequency of the combination of the block number kil and the block number kjm does not exceed the threshold value (step S501 No), the processing of step S305 is moved to.
[0175] By adding such processing of step S501, the correction from the first person re-identification result to the second person re-identification result can be performed in the case where the frequency is high, and as a result, the accuracy of the Re-ID can be improved.
[0176] <Dispersal and merging>
[0177] Furthermore, the constituent elements of the devices illustrated do not necessarily need to be physically configured as shown. That is, the specific method of distributing / combining the devices is not limited to the method shown, and they can be configured, in any unit, functionally or physically, according to various loads or usage conditions. For example, the acquisition unit 15A, the person detection unit 15B, the feature extraction unit 15C, the tracking unit 15D, the first re-identification unit 15E, the correspondence generation unit 15F, or the second re-identification unit 15G can be connected to the information processing device 10 as external devices via a network. Alternatively, other devices can each have the acquisition unit 15A, the person detection unit 15B, the feature extraction unit 15C, the tracking unit 15D, the first re-identification unit 15E, the correspondence generation unit 15F, or the second re-identification unit 15G, and can be connected and cooperate via a network to realize the function of the information processing device 10.
[0178] <Application Example>
[0179] Next, use Figure 18 The corresponding use case will be described. The information processing device 10 can analyze the actions of the recorded person using images captured by cameras 30A to 30N. Facility 1 is a railway facility, airport, shop, etc. In addition, turnstiles G1 configured in facility 1 are configured at the entrance of shops, railway facilities, airport boarding gates, etc.
[0180] First, let's look at examples where the login target is a railway facility or an airport. At railway facilities and airports, gate G1 is installed at ticket checks, counters, or inspection stations. In this case, if the biometric information of a person has been pre-registered as that of a passenger on the railway or airplane, the information processing device 10 determines that authentication based on the person's biometric information is successful.
[0181] Next, let's take the example of registering a store. In the case of a store, the gate G1 is installed at the store's entrance. At this time, as a login, when a person's biometric information is registered as a member of the store, the information processing device 10 determines that authentication based on the person's biometric information is successful.
[0182] Here, the details of the login process are explained. Authentication is performed by acquiring a vein image, such as that obtained by a vein sensor, from the biosensor 31. This determines the ID, name, and other information of the person logging in.
[0183] At this time, the information processing device 10 uses cameras 30A to 30N to acquire images of the registered user. Next, the information processing device 10 detects people from the images. The tracking unit 15D of the information processing device 10 tracks the people detected from the images captured by the cameras 30 between frames. The information processing device 10 associates the registered user's ID and name with the tracked person.
[0184] Here, use Figure 19 The application example is illustrated using a store as an example. Upon login, the information processing device 10 obtains the biometric information of the person passing through the gate G1 located at a predetermined position within the store (step S601). Specifically, the information processing device 10 obtains, for example, a vein image obtained from the biometric sensor 31 mounted on the gate G1 at the store entrance, and performs authentication. At this time, the information processing device 10 determines the user's ID, name, etc., based on the biometric information.
[0185] In addition, a biosensor 31 is mounted on a turnstile G1 located at a designated position in the facility, and detects the biometric information of people passing through the turnstile G1. Furthermore, cameras 30A to 30N are installed on the ceiling of the shop.
[0186] Alternatively, the information processing device 10 can replace the biosensor 31 to obtain biometric information from facial images captured by a camera mounted on the gate G1 installed at the entrance of the store, and then perform authentication.
[0187] Next, the information processing device 10 determines whether the authentication based on the person's biometric information is successful (step S602). If the authentication is successful (step S602 Yes), the process proceeds to step S603. On the other hand, if the authentication fails (step S602 No), the process proceeds to step S601.
[0188] Then, the information processing device 10 analyzes the image containing the person passing through the gate G1 and identifies the person in the image as a person registered to facility 1 (step S603). The information processing device 10 stores the identification information of the person determined based on biometric information in the storage unit in a corresponding manner. Specifically, the information processing device 10 stores the ID and name of the person to be registered in a corresponding manner with the identified person.
[0189] Then, the information processing device 10 analyzes the images acquired by the cameras 30A to 30N, and uses the results of identifying the first person and the second person to track the identified person (step S604). That is, the information processing device 10 identifies the identity of the person captured by the multiple cameras 30A to 30N. Then, the information processing device 10 determines the trajectory of the identified person within the facility 1 by determining the route for tracking the identified person.
[0190] Thus, after the person logs in, by determining whether the logged-in person has acquired a product arranged in the store, it is possible to analyze actions related to the person's purchase. Here, the actions related to the person's purchase are described. The information processing apparatus 10 generates the person's skeletal information by analyzing the image containing the tracked person. Then, the information processing apparatus 10 identifies the action of the tracked person acquiring a product using the generated skeletal information. That is, the information processing apparatus 10 determines whether the person has acquired a certain product from among the plurality of products arranged in the store from the time the person enters the store to the time the person leaves the store after the person enters the store. Then, the information processing apparatus 10 stores the result of whether the product was acquired in association with the ID and name of the person who logged in.
[0191] Specifically, the information processing apparatus 10 determines the customer staying in the store and the product arranged in the store from the captured image by the cameras 30A to 30N using an existing object detection technique. In addition, the information processing apparatus 10 estimates the position and posture of each joint of the person using the skeletal information of the determined person generated from the image captured by the cameras 30A to 30N using an existing skeletal estimation technique. Furthermore, the information processing apparatus 10 detects the action of holding the product, the action of putting the product in the basket or the cart, and the like based on the positional relationship between the skeletal information and the product. For example, the information processing apparatus 10 determines that the product is held when the skeletal information at the position of the person's arm overlaps with the region of the product.
[0192] In addition, the existing object detection algorithm is, for example, an object detection algorithm using deep learning such as Faster R-CNN (Convolutional Neural Network). In addition, it can also be an object detection algorithm such as YOLO (You Only Look Once), SSD (Single Shot Multibox Detector), and the like. Furthermore, the existing skeletal estimation algorithm is, for example, a skeletal estimation algorithm using deep learning such as DeepPose, OpenPose, and the like Human Pose Estimation.
[0193] In addition, here, as an example of the living body information, a vein image is exemplified, but the living body information can also be a face image, a fingerprint image, an iris image, and the like. In addition, in the above-described embodiment 1, an example in which the feature amount of the image captured by the camera 30 embedded in the feature space by the feature extraction unit 15C is used for the first person re-identification by the first re-identification unit 15E is exemplified, but it is not limited thereto. For example, the living body information detected by the cameras 30A to 30N or the feature amount extracted from the living body information can also be used for the first person re-identification.
[0194] <Hardware structure>
[0195] In addition, the various processes explained in the above-described embodiments can be realized by executing a program prepared in advance by a computer such as a personal computer or a workstation. Therefore, hereinafter, using Figure 20 , an example of a computer that executes an authentication program having the same functions as those of Embodiment 1 and Embodiment 2 will be explained.
[0196] Figure 20 is a diagram showing an example of a hardware structure. As shown in Figure 20 , the computer 100 has an operation section 110a, a speaker 110b, a camera 110c, a display 120, and a communication section 130. In addition, the computer 100 includes a CPU 150, a ROM 160, an HDD 170, and a RAM 180. Each of these 110 to 180 is connected via a bus 140.
[0197] The HDD 170 stores, as shown in Figure 20 , an authentication program 170a that functions in the same way as the acquisition section 15A, the person detection section 15B, the feature extraction section 15C, the tracking section 15D, the first re-authentication section 15E, the correspondence relationship generation section 15F, and the second re-authentication section 15G shown in Figure 1 . This authentication program 170a can also be integrated or separated in the same way as each of the constituent elements of the acquisition section 15A, the person detection section 15B, the feature extraction section 15C, the tracking section 15D, the first re-authentication section 15E, the correspondence relationship generation section 15F, and the second re-authentication section 15G shown in Figure 1 . That is, in the HDD 170, all of the data shown in Embodiment 1 described above can not necessarily be stored, and data for processing can be stored in the HDD 170.
[0198] In such an environment, the CPU 150 expands the authentication program 170a read out from the HDD 170 to the RAM 180. As a result, as shown in Figure 20 , the authentication program 170a functions as an authentication process 180a. This authentication process 180a expands various data read out from the HDD 170 in a region assigned to the authentication process 180a in a storage region possessed by the RAM 180, and executes various processes using the expanded various data. For example, as an example of a process executed by the authentication process 180a, processes shown in Figure 11 , Figure 12 , Figure 17 and the like are included. In addition, in the CPU 150, all of the processing sections shown in Embodiment 1 described above can not necessarily act, and processing sections corresponding to the processes as execution targets can be virtually realized.
[0199] Further, the above-described authentication program 170a can not necessarily be initially stored in the HDD 170, the ROM 160. For example, each program can be stored in a "portable physical medium" such as a floppy disk, a so-called FD, a CD-ROM, a DVD disk, a magneto-optical disk, an IC card, or the like, which is inserted into the computer 100. Further, the computer 100 can acquire each program from such a portable physical medium and execute it. Alternatively, each program can be stored in another computer or a server device, or the like, which is connected to the computer 100 via a public line, the Internet, a LAN, a WAN, or the like, and the computer 100 can acquire each program from them and execute it.
[0200] Mark Description
[0201] 10 Information processing apparatus
[0202] 11 Communication control section
[0203] 13 Storage section
[0204] 13A ID information
[0205] 13B Correspondence data
[0206] 15 Control section
[0207] 15A Acquisition section
[0208] 15B Character detection section
[0209] 15C Feature extraction section
[0210] 15D Tracking section
[0211] 15E First re-authentication section
[0212] 15F Correspondence generation section
[0213] 15G Second re-authentication section
[0214] 30A, 30B, ··· 30N Camera
Claims
1. A method of identification, characterized in that, The computer executes the following process: In a case where a first person is detected from a first image captured using a first camera and a second person is detected from a second image captured using a second camera, relationship information is generated, the relationship information being information associating a position within the first image at which the first person is detected and a position within the second image at which the second person is detected, The first person and the second person are identified based on feature information of the first person and feature information of the second person.
2. The identification method according to claim 1, wherein The generating process includes the following process: In a case where the number of persons detected from the first image and the number of persons detected from the second image is less than a threshold value, the relationship information is generated.
3. The identification method according to claim 1, wherein The computer further executes the following process: The first person and the second person are identified based on whether a combination of the position within the first image at which the first person is detected and the position within the second image at which the second person is detected is registered in the relationship information, In a case where a first identification result using the feature information and a second identification result using the relationship information are not consistent, the first identification result is corrected to the second identification result.
4. The identification method according to claim 3, wherein The process of identifying using the feature information includes the following process: In a case where an elapsed time from the start of the generation of the relationship information is above a threshold value, the identification of the first person and the second person is performed.
5. The identification method according to claim 3, wherein The process of identifying using the feature information includes the following process: In a case where the number of associations of the position within the first image and the position within the second image in the relationship information is above a threshold value, the identification of the first person and the second person is performed.
6. The identification method according to claim 3, wherein The process of identifying using the feature information includes the following process: In a case where a coverage rate of a combination of the position within the first image and the position within the second image in the relationship information is above a threshold value, the identification of the first person and the second person is performed.
7. The identification method according to claim 1, wherein The generating process includes the following process: The orientation of the first person detected inter-frame of the first image and the orientation of the second person detected inter-frame of the second image are further associated.
8. The identification method according to claim 1, wherein The generating process includes the following process: The speed of the first person detected inter-frame of the first image and the speed of the second person detected inter-frame of the second image are further associated.
9. The identification method according to claim 1, wherein A sensor or a camera is mounted to a gate arranged at a predetermined position of a facility, and based on detection of biological information of a person passing through the gate by the sensor or the camera, biological information of the person is acquired, Upon success of authentication based on the acquired biological information of the person, an image including the person passing through the gate is analyzed, whereby the person included in the image is recognized as a person who entered the facility, Recognition information of the person determined from the biological information and the recognized person are stored in correspondence with each other, The recognized person is tracked using a result of identifying the first person and the second person.
10. The identification method according to claim 9, wherein the facility is a store, the gate is arranged at an entrance of the store, upon the acquired biological information of the person being registered as an object of a member of the store, success of authentication based on the biological information of the person is determined, by tracking the person moving in the store, an action related to purchase of the person in the store from entry into the store to exit therefrom is determined.
11. The identification method according to claim 10, wherein skeletal information of the person is generated by analyzing an image including the tracked person, whether or not the tracked person performed an action of acquiring a product arranged in the store in the store is recognized as an action related to purchase using the generated skeletal information.
12. The identification method according to claim 9, wherein the facility is any one of a railway facility and an airport, the gate is arranged at a ticket gate of the railway facility, a counter or a check station of the airport, upon the acquired biological information of the person being registered as an object of a passenger of a train or an airplane in advance, success of authentication based on the biological information of the person is determined.
13. An identification procedure, characterized by The computer is caused to execute: in a case where a first person is detected from a first image captured using a first camera and a second person is detected from a second image captured using a second camera, relationship information is generated, the relationship information being information associating a position within the first image at which the first person is detected and a position within the second image at which the second person is detected, the first person and the second person are identified based on feature information of the first person and feature information of the second person.
14. The identification program according to claim 13, wherein a sensor or a camera is mounted to a gate arranged at a predetermined position of a facility, and based on detection of biological information of a person passing through the gate by the sensor or the camera, biological information of the person is acquired, upon success of authentication based on the acquired biological information of the person, an image including the person passing through the gate is analyzed, whereby the person included in the image is recognized as a person who entered the facility, recognition information of the person determined from the biological information and the recognized person are stored in correspondence with each other, the recognized person is tracked using a result of identifying the first person and the second person.
15. The authentication procedure according to claim 14, wherein the facility is a store, the gate is disposed at an entrance of the store, when the acquired biometric information of the person is registered as an object of a member of the store, it is determined that authentication based on the biometric information of the person is successful, by tracking the person moving in the store, it is determined that the person has an action related to purchase from entering the store to leaving the store in the store.
16. The authentication procedure according to claim 15, wherein by analyzing an image including the tracked person, skeletal information of the person is generated, using the generated skeletal information, it is identified whether the tracked person has an action of acquiring a product disposed in the store as an action related to the purchase.
17. The authentication procedure according to claim 14, wherein the facility is any one of a railway facility and an airport, the gate is disposed at a boarding gate of the railway facility or the airport, when the acquired biometric information of the person is registered as an object of a boarding passenger of a train or an airplane in advance, it is determined that authentication based on the biometric information of the person is successful.
18. An information processing apparatus comprising: has a control section, the control section performs the following processing: when a first person is detected from a first image captured using a first camera and a second person is detected from a second image captured using a second camera, relationship information is generated, the relationship information being information associating a position within the first image where the first person is detected and a position within the second image where the second person is detected, based on feature information of the first person and feature information of the second person, the first person and the second person are authenticated.
Citation Information
Patent Citations
Surveillance camera terminal
WO2011010490A1