Object identification system

By building an object determination system, automatically selecting effective features and calculating similarity, the determination and authentication failure caused by camera setting conditions is solved, and the determination and authentication accuracy of people is improved.

CN115410221BActive Publication Date: 2025-08-08HITACHI LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210541818.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-05-27
Filing Date
2022-05-17
Publication Date
2025-08-08
Estimated Expiration
2042-05-17

AI Technical Summary

Technical Problem

In the prior art, it is difficult to effectively capture the characteristics for human determination and authentication according to the camera setting status, resulting in a high possibility of determination or authentication failure.

Method used

By constructing a target object determination system, the camera information extraction unit, feature extraction unit, camera setting estimate unit, feature selection unit and comprehensive matching unit are used to automatically select effective features, and the similarity degree is calculated to improve the accuracy of determination and authentication.

Benefits of technology

According to the camera setting environment, effective features are selected, thereby improving the accuracy of human determination and authentication and reducing the possibility of determination or authentication failure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115410221B_ABST
    Figure CN115410221B_ABST
Patent Text Reader

Abstract

A system for identifying an object automatically selects features effective for identification based on camera settings, and automatically selects a method for calculating similarity used for identification based on the selected features. The system includes a camera information extraction unit that acquires image information from an image captured by a camera; a feature extraction unit that acquires image feature information from the image information acquired by the camera information extraction unit; a camera setting estimation unit that calculates camera setting information using the image information acquired by the camera information extraction unit; a feature selection unit that selects features using the image feature information acquired by the feature extraction unit and the camera setting information acquired by the camera setting estimation unit; a comprehensive matching unit that matches the features selected by the feature selection unit with features of the person to be identified; a comprehensive similarity calculation unit that calculates similarity using the matching results of the comprehensive matching unit and outputs an identification result; and a result output unit that outputs the identification result of the comprehensive similarity calculation unit.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an object identifying system for identifying and authenticating an object, such as a person, based on a captured image. Background Art

[0002] Identifying and authenticating a person as an object, for example, is a process required to improve the security and convenience of various systems. It has been widely used in situations where security is particularly required, such as ATMs in financial institutions and entrances and exits of important facilities.

[0003] Human identification technology varies depending on the usage scenario, but a technology that is frequently used in the field of security assurance is a technology that uses image features such as faces and clothing to identify and authenticate people.

[0004] When identifying a person, this system compares features extracted from the area of the person in the image with pre-stored features of the person to be identified, calculates a similarity, and determines that they are the same person if the similarity is high. For example, Patent Document 1 discloses a method for performing authentication by comparing facial and clothing features extracted from a two-dimensional image with features of the person to be authenticated that have been pre-registered in an authentication registration database.

[0005] Furthermore, recent advances in sensor technology have increased the number of human characteristics that can be acquired beyond images. Examples of these characteristics include positional information about a person in three-dimensional space and walking characteristics obtained from stereo cameras. Positional information refers to the location of a person's standing position in three-dimensional space and the positions of their joints (such as their arms, head, and waist). Furthermore, walking characteristics include walking speed, stride length, acceleration, and other characteristics measured in three-dimensional space.

[0006] According to Patent Document 1, it is possible to recognize a person using a camera image. However, it is difficult to fully utilize this recognition function in the following scenarios.

[0007] An example of a difficult recognition scenario is when using features extracted from images, such as color, grayscale, brightness, shape, texture, edges, and outlines, to identify a person within a two-dimensional image. Depending on the camera's placement, it may be impossible to obtain the necessary features, leading to identification failure. For example, when extracting facial information from a camera installed at the entrance to an area where a person is to be identified, it is preferable to install the camera directly in front of the entrance. However, if the camera is installed to the side or rear of the person's face, the camera may not be able to obtain facial information, potentially leading to identification or authentication failure.

[0008] In other examples of difficult recognition scenarios, if the camera's angle of installation deviates from its initial setting due to factors such as aging, building vibration, or other factors, the features extracted from the captured image, such as faces, may deviate from the registered information. Comparing these deviated features with those pre-registered in the authentication system could result in identification or authentication failure.

[0009] These difficult scenarios and factors can include issues between the person in the image and the surrounding environment, or camera settings caused by the camera parameters used to capture the image. Depending on the camera's environment, features used for identification may not be captured, or the camera's location may be offset, potentially leading to failures in identification and authentication.

[0010] Patent Document 1: Japanese Patent Application Laid-Open No. 2020-124367 Summary of the Invention

[0011] Therefore, an object of the present invention is to provide an object identification system that automatically selects features effective for identification and authentication based on camera installation conditions and calculates similarity based on the selected features.

[0012] Based on the above, in the present invention, a person identification system includes: a camera information extraction unit that obtains image information from an image captured by a camera; a feature extraction unit that obtains image feature information from the image information obtained by the camera information extraction unit; a camera setting inference unit that uses the image information obtained by the camera information extraction unit to determine camera setting information; a feature selection unit that uses the image feature information obtained by the feature extraction unit and the camera setting information obtained by the camera setting inference unit to select features; a comprehensive matching unit that uses the features selected by the feature selection unit to match features of a person to be identified; a comprehensive similarity calculation unit that uses the matching results of the comprehensive matching unit to determine similarity and output an identification result; and a result output unit that outputs the identification result of the comprehensive similarity calculation unit.

[0013] According to the object identification system of the present invention, effective features captured by a camera are selected based on the installation environment of the camera actually installed in an everyday environment, thereby improving the accuracy of human identification and authentication. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 This is a diagram showing a configuration example of the object identification system according to the first embodiment.

[0015] Figure 2 1 is a diagram showing a configuration example of a case where the camera installation estimating unit 11 is manually operated.

[0016] Figure 3 1 is a diagram showing a configuration example of a case where the camera installation estimating unit 11 is automatically executed.

[0017] Figure 4 It is a figure showing a specific image example.

[0018] Figure 5a 1 is a diagram showing an example of processing of a single feature by the integrated matching unit 16 .

[0019] Figure 5b 1 and 2 are diagrams showing examples of processing of a plurality of features by the integrated matching unit 16 .

[0020] Figure 6a 1 is a diagram showing an example of processing of a single feature by the integrated similarity calculation unit 17 .

[0021] Figure 6b 1 is a diagram showing an example of processing of integrating a plurality of features by the similarity calculation unit 17 .

[0022] Figure 7 This is a diagram showing a configuration example of an object identification system according to the second embodiment.

[0023] Figure 8 This figure shows an example of feature selection using 2D features, camera parameter information, and environmental information.

[0024] Figure 9 This is a diagram showing a configuration example of an object identification system according to the third embodiment.

[0025] Figure 10 This figure shows an example of feature selection using 3D features, camera parameter information, and environmental information.

[0026] Figure 11 This is a diagram showing an example of a processing flow of a target object identification method according to an embodiment of the present invention.

[0027] Explanation of symbols

[0028] 11 camera setting estimation unit, 12 camera information extraction unit, 13 2D feature extraction unit, 14 3D feature extraction unit, 15 feature selection unit, 16 comprehensive matching unit, 17 comprehensive similarity calculation unit, 18 result output unit, 1a camera setting information, 1b 2D feature database, 1c 3D feature database, 2 stereo camera, 2a monocular camera, 21 camera parameter acquisition unit, 22 environmental information acquisition unit, 23 camera parameter information, 24 environmental information, 31 image extraction unit, 32 camera calibration unit, 33 image segmentation unit, 71 2D comprehensive matching unit, 72 2D comprehensive similarity calculation unit, 91 3D comprehensive matching unit, 92 3D comprehensive similarity calculation unit. DETAILED DESCRIPTION

[0029] Hereinafter, an embodiment of the object identification system of the present invention will be described in detail with reference to the accompanying drawings. In the following description, a case where the object is a person is taken as an example, but the object is of course not limited to a person.

[0030] [Example 1]

[0031] Figure 1 This is a diagram showing a configuration example of an object identification system according to the first embodiment of the present invention. Figure 1 The object determination system 10 is composed of multiple databases 1 (1a, 1b, 1c), a camera unit 2, a camera setting inference unit 11 as various functional units processed by program execution in the operation unit of a computer, a camera information extraction unit 12, a 2D feature extraction unit 13, a 3D feature extraction unit 14, a feature selection unit 15, a comprehensive matching unit 16, a comprehensive similarity calculation unit 17 and a result output unit 18.

[0032] The present invention will be described in detail below. Briefly, when identifying a person based on their features, the camera setting estimation unit 11 sets a value for the feature based on conditions determined by the camera setting estimation unit 11 and provides this value to subsequent processing. Specifically, if the feature value is low, its weight in subsequent judgments is set low, or it is not used in the judgment.

[0033] In the embodiment of the present invention, in order to achieve the above functions, the Figure 1 The image is captured by a stereo camera 2 composed of two or more monocular cameras 2a. The image is input to the camera information extraction unit 12, which obtains a two-dimensional image and three-dimensional information. Based on this two-dimensional image and three-dimensional information, the 2D feature extraction unit 13 extracts 2D features, and the 3D feature extraction unit 14 extracts 3D features, and each feature is stored in the 2D feature database 1b and the 3D feature database 1c, respectively. In other embodiments of the present invention, the camera 2 only needs to be a monocular camera or a stereo camera, or both. Therefore, the feature information can use either 2D features or 3D features, or both. In Example 1, the use of both is exemplified.

[0034] Furthermore, the camera setting estimation unit 11 obtains information output from the stereo camera 2 and information extracted by the camera information extraction unit 12, and uses these information to estimate camera setting information. The camera setting information estimated by the camera setting estimation unit 11 is stored in the camera setting information database 1a.

[0035] Next, the feature selection unit 15 inputs camera setting information, 2D features, and 3D features from the camera setting information database 1a, 2D feature database 1b, and 3D feature database 1c, respectively, and selects features effective for identifying a person.

[0036] On this basis, the selected effective features are matched with the features of the person to be identified in the comprehensive matching unit 16, and the result is input to the comprehensive similarity calculation unit 17. Then, the comprehensive similarity calculation unit 17 calculates the comprehensive similarity based on the input matching result and outputs the result to the result output unit 18.

[0037] In addition, Figure 1 In FIG, the stereo camera 2 is a camera having a pair of monocular cameras 2a built therein, and simultaneously captures two-dimensional images from left and right viewpoints to generate three-dimensional information including depth.

[0038] Furthermore, the camera setting estimation unit 11 estimates camera setting information. In this case, the camera setting information includes parameter information of the stereo camera 2 and environmental information captured by the stereo camera 2 .

[0039] Here, the camera parameters include internal parameters such as the focal length of the camera, and external parameters such as the height of the camera, the distance to the subject, etc. By using the internal and external parameters, the camera's shooting area information can be estimated.

[0040] Furthermore, the environmental information indicates layout information of passages and places where people move in the area captured by the stereo camera 2. This information (stereo camera 2 parameter information and environmental information captured by the stereo camera 2) is stored as camera setting information in the camera setting information database 1a.

[0041] Next, use Figure 2 、 Figure 3 A detailed configuration example of the camera installation estimation unit 11 is described. Figure 2 This is a structure example for manual implementation. Figure 3 This shows an example of a structure where automatic execution is performed.

[0042] Figure 2 The camera setting estimation unit 11 is composed of a camera parameter acquisition unit 21 and an environment information acquisition unit 22. The camera parameter acquisition unit 21 acquires the camera parameters (internal parameters and external parameters) and stores them in the camera setting information database 1a as camera parameter information 23. The environment information acquisition unit 22 acquires the environment information and stores it in the camera setting information database 1a as environment information 24. Figure 2The above is an example of manual setting, so this data should be acquired and saved when the system is built or at an appropriate time thereafter, such as when the operator performs inspections.

[0043] Regarding specific data acquisition, for example, in the case of a stereo camera fixed at the entrance of an area where people are to be identified, internal parameters are determined based on specifications, while external parameters are determined by measuring the height and angle of the camera. Therefore, this information is input into the camera parameter acquisition unit 21. Furthermore, layout information of passageways, walls, and other areas in the area where people are to be identified is acquired in advance and input into the environmental information acquisition unit 22. Thus, camera parameters and environmental information are determined based on the camera's specifications and installation status, as well as information about the area where people are to be identified. This information is then stored in the camera installation information database 1a as camera parameter information 23 and environmental information 24.

[0044] Figure 3 The camera setup estimation unit 11 for automatic acquisition consists of an image extraction unit 31, a camera calibration unit 32, and an image segmentation unit 33. The image extraction unit 31 extracts a two-dimensional image and camera specification information from the camera information extraction unit 12 and the stereo camera 2. The camera calibration unit 32 uses the extracted two-dimensional image and camera specifications to automatically calculate the camera's internal and external parameters. The calculated camera parameters are stored as camera parameter information 23 in the camera setup information database 1a. Furthermore, the image segmentation unit 33 analyzes the two-dimensional image to extract layout information such as aisles and walls within the image. The extracted layout information is stored as environmental information 24 in the camera setup information database 1a.

[0045] Alternatively, the camera calibration unit 32 may automatically calculate the camera's intrinsic and extrinsic parameters using the two-dimensional image extracted by the camera information extraction unit 12. In this case, an automatic calibration method using patterns on the acquired two-dimensional image and the positional relationships between the patterns is used for the calculations.

[0046] Here, return to Figure 1 , and other components are explained. The 2D feature extraction unit 13 extracts features that identify a person from the input two-dimensional image. Features extracted by the 2D feature extraction unit 13 include facial information such as the facial outline, the positional relationship between the eyes and mouth, information obtained by digitizing the person's silhouette (outline), and clothing information. When clothing information is used as a feature, color, grayscale, brightness, shape, texture, edges, etc. are used.

[0047] The 3D feature extraction unit 14 uses the input 2D images and 3D information to extract 3D features. 3D features include positional information in three-dimensional space and walking characteristics obtained from a stereo camera. Positional information includes the position of a person's standing position in three-dimensional space and the positions of a person's joints (such as arms, head, and waist). Walking characteristics include walking speed, stride length, acceleration, and other features measured in three-dimensional space. These features are predefined for the subject to be identified or authenticated and registered in the 3D feature database 1c.

[0048] As described above, when referring to the camera parameter information 23 and the environmental information 24 as the camera setting information stored in the database 1a, some features related to a person may be obtained under conditions that are not suitable for image determination.

[0049] For example, in the camera parameter information 23, there may be situations where a feature, such as a face, cannot be clearly seen because it is too far away, blurred, or obscured by other body parts. Furthermore, in the environmental information 24, there may be obstacles such as hedges and trees in the foreground of the image, making it difficult to detect the movement of a person's features, or a person may be receding, making it difficult to see the face.

[0050] Return to Figure 1 The feature selection unit 15 uses the camera setting information (camera parameter information 23, environmental information 24) output from the camera setting inference unit 11, the 2D features output from the 2D feature extraction unit 13 and stored in the 2D feature database 1b, and the 3D features output from the 3D feature extraction unit 14 and stored in the 3D feature database 1c to select features that are effective for determination.

[0051] During selection, the confidence levels of 2D and 3D features are calculated based on the camera configuration information (camera parameter information 23 and environmental information 24). The confidence level is a numerical value ranging from 0 to 1. A value of 0 indicates that the feature cannot be obtained from the configured camera. A value of 1 indicates that the feature can be reliably obtained from the configured camera. Furthermore, a confidence level threshold is pre-set. If the calculated confidence level of a feature exceeds the threshold, the feature is selected as valid.

[0052] On the other hand, if it is lower than the threshold, the feature is not selected. Figure 4 , examples of these feature selections are explained. Figure 4 This is an example in which a passage 41 extends from the front to the back of the screen, and walls 42 are provided on both sides of the passage 41. A scene in which a person 44 as a specific target moves along a path 43 from the front to the back of the screen is shown.

[0053] In this case, layout information for walls 42, passages 41, and the like can be obtained from environmental information 23 stored in camera installation information database 1a. By dividing passages 41 and walls 42, the direction of movement of person 44 can be estimated as being from the front to the back, or from the back to the front. Therefore, when person 44 appears in the frame, if they move from the front, i.e., from bottom to top, their back will be facing the camera, making it impossible to capture a facial image. On the other hand, if they move from the back, i.e., from top to bottom, it is expected that a facial image can be captured.

[0054] Therefore, it is possible to detect a person within the image and determine whether a facial image can be acquired based on whether the person is moving from bottom to top or from top to bottom. If the person whose facial image can be acquired is expected to be moving from top to bottom, the confidence level of information such as the facial outline, eye, and mouth positional relationship is set high. If the person whose facial image cannot be acquired is expected to be moving from bottom to top, the confidence level of information such as the facial outline, eye, and mouth positional relationship is set low.

[0055] In addition to this example, in scenarios where there are obstructions or where a person's facial image cannot be obtained due to the camera's installation location or angle, the confidence level of 3D features such as the person's position information and walking information is set high.

[0056] In addition, Figure 4 , an example of setting the confidence level based on the environmental information 24 is described. However, using the camera parameter information 23, when the image is blurred or too small in the camera parameters and is clearly not suitable for evaluation, the confidence level of the feature can be set low.

[0057] Return to Figure 1 The images input to the comprehensive matching unit 16 are 2D feature images or 3D feature images stored in the databases 1b and 1c. These images are assigned confidence information determined based on the camera setting information stored in the database 1a. Alternatively, images with low confidence are not initially used in the processing of the image determination unit 10d.

[0058] The comprehensive matching unit 16 matches the features of the person to be identified using the features selected by the feature selection unit 15. Furthermore, the comprehensive similarity calculation unit 17 calculates the final similarity using the matching results calculated by the comprehensive matching unit 16. These processing examples are described using the accompanying drawings. Figure 5a and Figure 5b is a diagram showing an example of processing by the comprehensive matching unit 16. Figure 6a and Figure 6b 2 is a diagram showing an example of processing performed by the comprehensive similarity calculation unit 17 .

[0059] in, Figure 5a and Figure 6a An example of processing for each single feature is shown. Figure 5a In the example, feature 1 is input to the comprehensive matching unit 16, and the feature 1 matching unit 16-1 compares the feature of the person input as feature 1 with the feature of the person to be identified, and uses it as matching result 1. Figure 6b In the example, the matching result 1 is directly output as the final similarity of the comprehensive similarity calculation unit 17. Therefore, even when other features are input at other timings, the final similarity for each feature is output.

[0060] then, Figure 5b and Figure 6b Indicates an example of processing multiple features. Figure 5b In the embodiment, a plurality of images related to features 1 to N are input to the comprehensive matching unit 16. The features of the person input as feature 1 are compared with the features of the person to be identified in the feature 1 matching unit 16-1. Similarly, the features of the person input as feature N are compared with the features of the person to be identified in the feature N matching unit 16-N. These are set as matching results 1 to matching results N. Figure 6b Matching results 1 to N are aggregated into one and output as the final similarity by the comprehensive similarity calculation unit 17 .

[0061] If the processing of the comprehensive matching unit 16 is described in more detail, Figure 5a In the case where only feature 1 is selected as the feature for matching, the feature 1 matching unit is implemented corresponding to feature 1. Figure 5b When multiple features are selected, the feature matching unit (feature 1 matching unit, feature 2 matching unit, feature N matching unit) performs processing corresponding to each feature, and then outputs the matching results (matching result 1, matching result 2, matching result N).

[0062] The matching results are output as numerical data. When performing matching, the person to be matched can be pre-determined and their features registered. Alternatively, if the person to be matched has not been registered in advance, the initially extracted person can be registered as the target. The feature matching method is selected appropriately based on the selected features. For example, when images are selected as features, methods such as SIFT, SURF, and AKZE can be applied. Matching using deep learning can also be implemented.

[0063] Next, if the processing of the comprehensive similarity calculation unit 17 is described in more detail, the final similarity is calculated using the matching result calculated by the comprehensive matching unit 16. However, at this time, the features selected by the feature selection unit 15 are as follows: Figure 5aIn the case of one, the matching result becomes one, and its value is as follows Figure 6a That becomes the final similarity.

[0064] On the other hand, in the case of selected features such as Figure 5b In the case of multiple features, use the matching results of each feature (matching result 1, matching result 2, matching result N), such as Figure 6b The final similarity is calculated in this way. If the selected feature types and matching methods differ, each matching result is normalized according to its output range, then weighted according to the feature's confidence level is added to create the final similarity. If the calculated final similarity exceeds a pre-set threshold, the two individuals are determined to be the same person.

[0065] The result output unit 18 outputs the result of the comprehensive similarity calculation unit 17. The output method can be changed according to the user who determines the application target of the system. For example, when managing people by ID, the ID is output, and when managing people by name, the name is output.

[0066] In this way, the similarity determination in the present invention can be performed individually or by combining multiple features.

[0067] According to the first embodiment described above, effective features captured by the camera are selected based on the installation environment of the camera actually installed in the daily environment, thereby improving the accuracy of person identification and authentication.

[0068] [Example 2]

[0069] Next, use Figure 7 Next, an object identification system according to a second embodiment of the present invention will be described. The difference from the first embodiment is that the output destinations of the camera information extraction unit 12 are only two: the 2D feature extraction unit 13 and the camera installation estimation unit 11 .

[0070] Therefore, the feature selection unit 15 selects features using the 2D features, camera parameter information stored in the camera setting information database 1a, and environmental information. The selected features are input to the 2D comprehensive matching unit 71. The matching results are then input to the 2D comprehensive similarity calculation unit 72, which calculates the comprehensive similarity. The result output unit 18 then outputs the comprehensive similarity results.

[0071] As described above, in Example 2, 2D features, camera parameter information, and environmental information are used to perform feature selection. Figure 8 This example will be described. Figure 8This depicts a scene where channel 82 extends left and right within the frame, with non-channel areas 81 existing above and below. In this situation, it can be inferred that persons 83 and 84 are moving left and right within the frame. Therefore, it can be inferred that a frontal facial image cannot be captured. Furthermore, in the case of person 83 moving from right to left within the frame, as indicated by trajectory 85, it can be inferred that the left side of the body is captured in the camera image.

[0072] On the other hand, in the case of person 84 moving from left to right within the frame as shown by trajectory 86, it can be inferred that the right side of the body is captured in the camera image. Therefore, when identifying person 83, the confidence level of 2D features obtained from the left side of the body and features with significant differences on the left side is also increased. Similarly, when identifying person 84, the confidence level of features obtained from the right side of the body and features with significant differences on the right side is increased. In this way, by setting the confidence level using movement information that can be inferred from environmental information, the accuracy of person identification can be improved.

[0073] Thus, according to the monitoring target area, Figure 8 As shown, even when there is no depth or the depth is small, it can be realized only by 2D information without using 3D information, and a correspondingly low-cost system can be achieved.

[0074] [Example 3]

[0075] Next, use Figure 9 The person identification system according to the third embodiment of the present invention will be described. The difference from the first embodiment is that the output destinations of the camera information extraction unit 12 are only two: the 3D feature extraction unit 14 and the camera installation estimation unit 11.

[0076] Therefore, the feature selection unit 15 selects features using the 3D features, camera parameter information stored in the camera setting information database 1a, and environmental information. The selected features are input to the 3D comprehensive matching unit 91. The matching results are then input to the 3D comprehensive similarity calculation unit 92, which calculates the comprehensive similarity. The result output unit 18 then outputs the comprehensive similarity results.

[0077] As described above, in the third embodiment, 3D features, camera parameter information, and environmental information are used to perform feature selection. Figure 10 This example will be described. Figure 10 The example of a stereo camera being set up in a passage 101 is shown. The passage 101 extends to the left and right in the image, and there are non-passage areas 102 above and below it. When the camera is set up in this way, a person 103 is Figure 10 The head, represented by the black dot, is closest to the camera, but the body (in Figure 10The hollow rectangle in the figure shows only the shoulders, Figure 4 、 Figure 8 Compared to 2D images, the information that can be obtained is reduced, and therefore the number of 2D features that can be obtained is reduced. Furthermore, since the image is taken from above, it is difficult to obtain facial images. In such cases, it is possible to identify a person by obtaining 3D features such as height, positional information, shoulder position, and arm and leg movement.

[0078] Furthermore, in the above-mentioned embodiments 1 to 3, a case where a stereo camera is used as a camera has been described, but different sensors such as an RGB sensor, an RGBD sensor, and a distance sensor may be used in combination.

[0079] In addition, in the object identification systems described in the above-mentioned embodiments 1 to 3, the scene where only one or two people exist in the screen is used as an example. However, when there are multiple people, multiple people can be identified by performing the same identification process on each person.

[0080] Thus, according to the monitoring target area, Figure 10 As shown, when only images from above are available and depth is difficult to obtain, 3D information can be used for implementation, without using 2D information that cannot obtain information in the height direction, and a correspondingly low-cost system can be achieved.

[0081] [Example 4]

[0082] The object identification system shown in the first, second and third embodiments is implemented using a computer. However, the processing flow in this case can be as follows, for example: Figure 11 Like that.

[0083] In the processing flow of the fourth embodiment, the camera image data is obtained in the first processing step S1 and is taken into the computer. Then, in the processing step S2, the camera setting information is obtained, but this can be obtained as follows. Figure 2 It can be stored and used as pre-prepared offline data, or it can be performed at an appropriate timing when the image is acquired. Figure 3 Camera calibration and camera segmentation processing are performed to obtain the latest camera setting information.

[0084] Images are accumulated through the processing to date. During the image processing stage, camera information is extracted in step S3, and 2D and 3D features are extracted in step S5. In parallel with these processes, camera setting information (environmental information 24 and camera parameter information 23) is extracted in step S4.

[0085] Thereafter, the process moves to the comprehensive matching process in step S7 , the comprehensive similarity calculation process in step S8 , and the result output process in step S9 .

Claims

1. A person identification system, characterized in that: have: a camera information extraction unit that obtains image information from an image captured by the camera; a feature extraction unit that obtains image feature information from the image information obtained by the camera information extraction unit; a camera setting estimation unit for obtaining camera setting information using the image information obtained by the camera information extraction unit; a feature selection unit that selects features using the image feature information obtained by the feature extraction unit and the camera setting information obtained by the camera setting estimation unit; a comprehensive matching unit that uses the features selected by the feature selection unit to match the features of the person to be identified; a comprehensive similarity calculation unit that calculates similarity using the matching result of the comprehensive matching unit and outputs a determination result; and A result output unit outputs the determination result of the comprehensive similarity calculation unit.

2. The person identification system according to claim 1, characterized in that The camera setting estimation unit receives input of camera setting information. The feature selection unit selects a feature using the camera setting information received by the camera setting estimation unit.

3. The person identification system according to claim 2, characterized in that The camera setting information includes at least one of internal parameters such as a focal length and a distortion coefficient of the camera, external parameters such as an installation height and an installation angle of the camera, and layout information of a place where the camera is installed.

4. The person identification system according to claim 3, characterized in that The feature selection unit calculates a degree of certainty using the camera setting information and the image feature information, and sets the image feature information obtained from the feature extraction unit as a valid feature when the degree of certainty is greater than a predetermined value.

5. The person identification system according to claim 4, characterized in that The camera is a stereo camera, The feature extraction unit obtains two-dimensional image features and three-dimensional image features, The camera placement estimating unit obtains camera placement information based on two-dimensional image features and three-dimensional image features.

Citation Information

Patent Citations

  • Body health condition image analysis device, method, and system

    JP2020124367A

  • Method and apparatus for calibration of binocular camera and binocular camera

    CN103606149A

  • Biometric authentication system

    CN110945520A