Object identification device, object identification program, and object identification method

The object identification device uses color and depth image data with estimation models to derive feature vectors, addressing low-light and distance limitations, enabling robust object recognition across varied conditions.

WO2026154882A1PCT designated stage Publication Date: 2026-07-23SOFTBANK CORPORATION +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
SOFTBANK CORPORATION
Filing Date
2025-12-16
Publication Date
2026-07-23

AI Technical Summary

Technical Problem

Existing 3D recognition technologies struggle with identifying objects in low-light conditions and limited shooting distance ranges, making it difficult to recognize objects in various environments.

Method used

An object identification device that utilizes a combination of color and depth image data, employing estimation models to derive color, depth, and three-dimensional feature vectors, allowing identification through different feature vectors based on lighting conditions and distance.

Benefits of technology

Enables object identification in diverse environments by using appropriate feature vectors, ensuring recognition in low-light conditions and beyond typical shooting distances.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025043881_23072026_PF_FP_ABST
    Figure JP2025043881_23072026_PF_FP_ABST
Patent Text Reader

Abstract

The present invention identifies an object in various environments. A communication terminal 2 comprises: a color feature estimation unit (40) that uses a machine-trained first estimation model (30) to estimate a color feature vector from data of a color image; a depth feature estimation unit (41) that uses a machine-trained second estimation model (31) to estimate a depth feature vector from data of a depth image; a three-dimensional feature estimation unit (42) that uses a machine-trained third estimation model (32) to estimate a three-dimensional feature vector including the color image and the depth image from the color feature vector and the depth feature vector; and an authentication unit (24) the uses one of the color feature vector, the depth feature vector, and the three-dimensional feature vector to perform face authentication.
Need to check novelty before this filing date? Find Prior Art

Description

Object Identification Device, Object Identification Program, and Object Identification Method

[0001] The present disclosure relates to an object identification device, an object identification program, and an object identification method.

[0002] In the face recognition method described in Patent Document 1, first, a first face image captured by an RGB (red, green, blue) camera and a second face image captured by a NIR (near-infrared) camera with a different shooting direction from the RGB camera are acquired. Next, while acquiring a face fusion image obtained by image fusion of the first face image and the second face image, three-dimensional depth information is acquired by stereo matching based on the first face image and the second face image. Then, face recognition is performed based on the face fusion image and the three-dimensional depth information, and a value characterizing the face recognition reliability is acquired.

[0003] Also, the camera "Intel (registered trademark) RealSense (trademark) Depth Camera D415" manufactured by Intel has a dedicated RGB sensor, two IR sensors, and an IR projector. By performing stereo matching on two IR images from the two IR sensors, a depth image, which is a two-dimensional image, is acquired.

[0004] Japanese Patent Application Laid-Open No. 2024-005748

[0005] An object identification device according to Aspect 1 of the present disclosure is an object identification device that identifies an object included in a color image, which is a two-dimensional image, and a depth image, which is a two-dimensional image, and includes an acquisition unit that acquires data of the color image and data of the depth image, a color feature estimation unit that estimates a color feature vector indicating a feature of the object included in the color image using a first estimation model from the data of the color image, a depth feature estimation unit that estimates a depth feature vector indicating a feature of the object included in the depth image using a second estimation model from the data of the depth image, a three-dimensional feature estimation unit that estimates a three-dimensional feature vector indicating a feature of the object included in a three-dimensional image including the color image and the depth image using a third estimation model from the color feature vector and the depth feature vector, and an identification unit that identifies the object using any one of the color feature vector, the depth feature vector, and the three-dimensional feature vector.

[0006] The object identification device according to Embodiment 2 of the present disclosure further comprises a selection unit that selects one of the color feature vector, the depth feature vector, and the three-dimensional feature vector based on the color image data and the depth image data, and the identification unit may identify the object using the feature vector selected by the selection unit.

[0007] The object identification device according to embodiment 3 of the present disclosure may further include a storage unit for storing the first estimation model, the second estimation model, and the third estimation model, as in embodiment 1 or 2.

[0008] The object identification device according to aspect 4 of the present disclosure may include, in aspects 1 to 3, a fully connected layer that reduces the dimensionality of a vector obtained by combining the color feature vector and the depth feature vector, a bias layer that corrects the bias from the center of the dimensionality-reduced vector, and a normalization layer that generates the three-dimensional feature vector by normalizing the corrected vector.

[0009] In the object identification device according to aspect 5 of the present disclosure, in aspects 1 to 4, the color image may be an image of the region of the object extracted from the original color image by object detection, and the depth image may be an image of the region corresponding to the region of the object extracted from the original depth image.

[0010] In the object identification device according to embodiment 6 of this disclosure, the object may be a human face, as described in embodiments 1 to 5 above.

[0011] The object identification program according to aspect 7 of this disclosure is an object identification program for causing a computer to function as an object identification device as described in aspects 1 to 6 above, and may be an object identification program for causing a computer to function as a color feature estimation unit, a depth feature estimation unit, a three-dimensional feature estimation unit, and an identification unit.

[0012] A method for identifying an object according to aspect 8 of the present disclosure is a method for identifying an object contained in a two-dimensional color image and a two-dimensional depth image, comprising: an acquisition step of acquiring data of the color image and data of the depth image; a color feature estimation step of estimating a color feature vector indicating the characteristics of the object contained in the color image using a first estimation model from the data of the color image; a depth feature estimation step of estimating a depth feature vector indicating the characteristics of the object contained in the depth image using a second estimation model from the data of the depth image; a three-dimensional feature estimation step of estimating a three-dimensional feature vector indicating the characteristics of the object contained in a three-dimensional image including the color image and the depth image using a third estimation model from the color feature vector and the depth feature vector; and an identification step of identifying the object using any of the color feature vector, the depth feature vector, and the three-dimensional feature vector.

[0013] This is a schematic diagram showing the configuration of an authentication system according to one embodiment of the present disclosure. This is a block diagram showing the configuration of a communication terminal in the above authentication system. This is a flowchart showing an example of the authentication process flow in the above communication terminal. This is a block diagram showing specific examples of the color feature estimation unit, depth feature estimation unit, and three-dimensional feature estimation unit in the above communication terminal. This is a block diagram showing the configuration of a learning unit that machine-learns the FC layer and bias layer in the above three-dimensional feature estimation unit. This is a schematic diagram showing the configuration of an authentication system according to another embodiment of the present disclosure. This is a block diagram showing the configuration of an access control device in the above authentication system.

[0014] Hereinafter, one embodiment of this disclosure will be described in detail with reference to the drawings. For ease of understanding, the background and challenges of this disclosure will be described first, followed by a detailed description of the disclosure.

[0015] In the following, the distance from the camera to the object will be referred to as the "shooting distance." Conventionally, 3D recognition technology is known that uses machine learning models to identify objects contained in color images and depth images.

[0016] However, with the above 3D identification technology, it is difficult to identify the object if the amount of light incident on the RGB camera that creates the color image is significantly too little or too much. Furthermore, the range of shooting distances in which depth images can be created is limited, depending on the sensor that creates the depth image. For this reason, it is difficult to identify the object if it is not present within the range of the shooting distance. This disclosure aims to identify objects in various environments.

[0017] The object identification device according to this disclosure is an object identification device that identifies an object contained in a two-dimensional color image and a two-dimensional depth image, and comprises: an acquisition unit that acquires data from the color image and data from the depth image; a color feature estimation unit that estimates the color feature vector from the color image data using a first estimation model; a depth feature estimation unit that estimates the depth feature vector from the depth image data using a second estimation model; a three-dimensional feature estimation unit that estimates a three-dimensional feature vector including the color image and the depth image using a third estimation model from the color feature vector and the depth feature vector; and an identification unit that identifies the object using any of the color feature vector, the depth feature vector, and the three-dimensional feature vector.

[0018] With the above configuration, if the amount of light incident on the RGB camera that creates the color image is significantly low or significantly high, it is possible to identify the object using the depth feature vector estimated from the depth image. Also, if the object is not present within the range of the shooting distance, it is possible to identify the object using the color feature vector estimated from the color image. Furthermore, if the amount of light incident on the RGB camera is appropriate and the object is present within the range of the shooting distance, it is possible to identify the object using the three-dimensional feature vector estimated from the depth feature vector and the color feature vector. Therefore, it is possible to identify objects in various environments.

[0019] [Embodiment 1] <Outline Configuration of Authentication System 1> Hereinafter, an authentication system 1 according to one embodiment of the present disclosure will be described with reference to Figures 1 to 3. Figure 1 is a block diagram showing the overall configuration of the authentication system 1. The authentication system 1 is a system that identifies and authenticates a human face as the object. As shown in Figure 1, the authentication system 1 includes a communication terminal 2 and a server device 3.

[0020] The communication terminal 2 and the server device 3 are connected via a network N such as the Internet. As indicated by arrow R, the communication terminal 2 has a function to perform facial recognition of the user M using the communication terminal 2.

[0021] Network N may be other than the Internet, such as a LAN (Local Area Network), mobile communication systems such as 4G, 5G, and 6G, LTE (Long Term Evolution), or Wi-Fi (registered trademark).

[0022] [Configuration of Communication Terminal 2] Next, the configuration of communication terminal 2 will be explained with reference to Figure 2. Figure 2 is a block diagram showing the configuration of communication terminal 2. Communication terminal 2 is a smartphone, tablet, PC (Personal Computer), etc., used by user M. As shown in Figure 2, communication terminal 2 includes a control unit 10, a storage unit 11, a camera unit 12, a communication unit 13, an operation unit 14, and a display unit 15.

[0023] The control unit 10 comprehensively controls the operation of various components of the communication terminal 2 and is composed of a computer including, for example, a CPU (Central Processing Unit) and memory. The operation control of these various components is performed by having the computer execute a control program. Further details of the control unit 10 will be described later.

[0024] The storage unit 11 is for recording information and is composed of storage devices such as a hard disk or flash memory. Details of the storage unit 11 will be described later.

[0025] The camera unit 12 is used for taking photographs and is built into the communication terminal 2. The camera unit 12 has lenses on both the side of the communication terminal 2 that is on the same side as the operation unit 14 and on the side opposite to the operation unit 14. The camera unit 12 is configured to capture both video and still images.

[0026] In this embodiment, the camera unit 12 creates a three-dimensional image (hereinafter abbreviated as "three-dimensional image") that includes a two-dimensional RGB color image and a two-dimensional depth image. The depth image can be created using known methods such as the stereo method, the ToF (Time of Flight) method, or the method using an infrared sensor.

[0027] The communication unit 13 is connected to the network N by wire or wireless connection and communicates with the server device 3 and other communication terminals via the network N. The communication unit 13 is composed of a NIC (Network Interface Card), an antenna, and the like.

[0028] The operation unit 14, for example, consists of a touch panel and accepts various operations from user M. The operation unit 14 creates input data corresponding to the accepted operations and transmits it to the control unit 10. The operation unit 14 may also have buttons for inputting characters, numbers, etc.

[0029] The display unit 15 is a display device for displaying various information from the control unit 10, and is, for example, a liquid crystal display (LCD) or an organic electroluminescent display (OLED).

[0030] (Details of the control unit and storage unit) As shown in Figure 2, the control unit 10 includes an acquisition unit 20, a feature estimation unit 21, a selection unit 22, a registration unit 23, and an authentication unit 24 (identification unit). The storage unit 11 stores the first estimation model 30, the second estimation model 31, the third estimation model 32, and the registered data 33.

[0031] The first estimation model 30 is a learning model that estimates a color feature vector CF representing the facial features contained in a color image CI from the color image CI data. The second estimation model 31 is a learning model that estimates a depth feature vector DF representing the facial features contained in a depth image DI from the depth image DI data. The third estimation model 32 is a learning model that estimates a three-dimensional feature vector TF representing the facial features contained in a three-dimensional image TI including the color image CI and depth image DI from the color feature vector CF and depth feature vector DF. The first estimation model 30, the second estimation model 31, and the third estimation model 32 are pre-trained by the server device 3 and stored in the memory unit 11.

[0032] The registered data 33 is data registered for the facial recognition of user M, and includes a color feature vector CF, a depth feature vector DF, and a three-dimensional feature vector TF that represent the feature quantities of user M's face.

[0033] The acquisition unit 20 acquires color image CI and depth image DI data from the camera unit 12. The acquisition unit 20 sends the acquired color image CI and depth image DI data to the feature estimation unit 21 and the selection unit 22.

[0034] The acquisition unit 20 may also perform face detection (object detection) on the original color image acquired from the camera unit 12 to identify the face region, extract the identified face region from the original color image, and use the extracted face region image as the color image CI used by the feature estimation unit 21. Alternatively, the acquisition unit 20 may extract the region corresponding to the face region from the original depth image acquired from the camera unit 12, and use the extracted region image as the depth image DI used by the feature estimation unit 21.

[0035] In this case, the sizes of the color image CI and the depth image DI can be reduced compared to the sizes of the original color image and the original depth image, respectively. As a result, the processing load on the feature estimation unit 21 is reduced, and / or the processing speed is improved. Note that the pixel region of the face may include the pixels surrounding that pixel region.

[0036] The feature estimation unit 21 includes a color feature estimation unit 40, a depth feature estimation unit 41, and a three-dimensional feature estimation unit 42. The color feature estimation unit 40 estimates a color feature vector CF from color image CI data from the acquisition unit 20 using a first estimation model 30 of the storage unit 11. The depth feature estimation unit 41 estimates a depth feature vector DF from depth image DI data from the acquisition unit 20 using a second estimation model 31 of the storage unit 11. The three-dimensional feature estimation unit 42 estimates a three-dimensional feature vector TF from the color feature vector CF and the depth feature vector DF using a third estimation model 32 of the storage unit 11.

[0037] The selection unit 22 selects an image usable for face recognition from the color image CI and depth image DI data based on the color image CI and depth image DI data from the acquisition unit 20. Examples of color image CIs that cannot be used for face recognition include images with uneven brightness due to sunlight, dark images, overexposed images, and backlit images. Examples of depth image DIs that cannot be used for face recognition include images in which the depth included in the depth image DI falls outside the range of depths usable for face recognition.

[0038] If the selection unit 22 selects only the color image CI, it instructs the feature estimation unit 21 to send the color feature vector CF estimated by the color feature estimation unit 40 to the authentication unit 24. If the selection unit 22 selects only the depth image DI, it instructs the feature estimation unit 21 to send the depth feature vector DF estimated by the depth feature estimation unit 41 to the authentication unit 24. If the selection unit 22 selects both the color image CI and the depth image DI, it instructs the feature estimation unit 21 to send the three-dimensional feature vector TF estimated by the three-dimensional feature estimation unit 42 to the authentication unit 24.

[0039] Furthermore, when registering a user for facial recognition, the selection unit 22 instructs the feature estimation unit 21 to send the color feature vector CF, the depth feature vector DF, and the three-dimensional feature vector TF to the registration unit 23 only when both the color image CI and the depth image DI are selected.

[0040] The registration unit 23 performs user registration for facial recognition. Specifically, the registration unit 23 stores the color feature vector CF, depth feature vector DF, and three-dimensional feature vector TF from the feature estimation unit 21 in the registration data 33 of the storage unit 11.

[0041] The authentication unit 24 performs face authentication (identification) of user M by comparing one of the feature vectors from the feature estimation unit 21 (color feature vector CF, depth feature vector DF, and three-dimensional feature vector TF) with the corresponding feature vector contained in the registered data 33 of the storage unit 11. Specifically, the authentication unit 24 calculates the similarity between the feature vector from the feature estimation unit 21 and the feature vector from the registered data 33, and determines that the face authentication was successful if the similarity is above a predetermined threshold. At this time, the authentication unit 24 displays information indicating that authentication was successful via the display unit 15 and accepts operations from user M via the operation unit 14.

[0042] <Authentication Process> Figure 3 is a flowchart showing an example of the authentication process flow by the communication terminal 2 with the above configuration. As shown in Figure 3, first, the acquisition unit 20 uses the image of the face region extracted by face detection from the original color image acquired from the camera unit 12 as the face image CI to be used in this process, and the image of the region corresponding to the face region extracted from the original depth image acquired from the camera unit 12 as the depth image DI to be used in this process (S10). Next, the selection unit 22 determines whether the color image CI is an image that can be used for face authentication based on the data of the color image CI (S11).

[0043] If the color image CI is an image that can be used for facial recognition (YES in S11), the selection unit 22 determines whether or not the depth image DI is an image that can be used for facial recognition based on the data of the depth image DI (S12).

[0044] When the depth image DI is an image available for face authentication (YES in S12), the authentication unit 24 performs face authentication (S13) using the three-dimensional feature vector TF estimated by the three-dimensional feature estimation unit 42 from the color feature vector CF estimated by the color feature estimation unit 40 from the data of the color image CI and the depth feature vector DF estimated by the depth feature estimation unit 41 from the data of the depth image DI. Then, the above authentication process ends.

[0045] On the other hand, when the depth image DI is not an image available for face authentication (NO in S12), the authentication unit 24 performs face authentication (S14) using the color feature vector CF estimated by the color feature estimation unit 40 from the data of the color image CI. Then, the above authentication process ends.

[0046] On the other hand, when the color image CI is not an image available for face authentication (NO in S11), the selection unit 22 determines whether the depth image DI is an image available for face authentication based on the data of the depth image DI (S15).

[0047] When the depth image DI is an image available for face authentication (YES in S15), the authentication unit 24 performs face authentication (S16) using the depth feature vector DF estimated by the depth feature estimation unit 41 from the data of the depth image DI. Then, the above authentication process ends.

[0048] On the other hand, when the depth image DI is not an image available for face authentication (NO in S15), the authentication unit 24 displays and outputs information indicating that face authentication could not be performed with the acquired color image CI and depth image DI via the display unit 15 (S17). Then, the above authentication process ends.

[0049] As described above, the communication terminal 2 of this embodiment can perform facial recognition using a three-dimensional feature vector TF when the color image CI and depth image DI are images usable for facial recognition. Furthermore, if the color image CI is an image usable for facial recognition but the depth image DI is an image unusable for facial recognition, facial recognition can be performed using a color feature vector CF. Furthermore, if the depth image DI is an image usable for facial recognition but the color image CI is an image unusable for facial recognition, facial recognition can be performed using a depth feature vector DF. Therefore, facial recognition can be performed in various environments.

[0050] Furthermore, it is not necessary to obtain the first estimated model 30, the second estimated model 31, and the third estimated model 32 from the server device 3 via the communication network N. Therefore, the communication terminal of this embodiment can perform facial recognition even if it is not connected to the communication network N.

[0051] <Example> Figure 4 is a block diagram showing specific examples of the color feature estimation unit 40, the depth feature estimation unit 41, and the three-dimensional feature estimation unit 42.

[0052] As shown in the upper part of Figure 4, the color feature estimation unit 40 includes a CNN (convolutional neural network) layer 50 and a normalization layer 51. The CNN layer 50 corresponds to the first estimation model 30 and is a learning model based on known architectures such as ArcFace, SphereFace, and CosFace. For example, the data of a 112 x 112 pixel color image CI is converted into a 128-dimensional vector by the CNN layer 50 and normalized to a unit hypersphere by the normalization layer 51, thereby converting it into a 128-dimensional color feature vector CF.

[0053] As shown in the middle section of Figure 4, the depth feature estimation unit 41 includes a CNN layer 52 and a normalization layer 53. The CNN layer 52 corresponds to the second estimation model 31 and is a learning model based on the known architecture described above. The CNN layer 50 of the color feature estimation unit 40 and the CNN 52 of the depth feature estimation unit 41 are machine-learned separately. For example, data from a 112 x 112 pixel depth image DI is converted into a 128-dimensional vector by the CNN layer 52 and normalized to a unit hypersphere by the normalization layer 53, thereby converting it into a 128-dimensional depth feature vector DF.

[0054] As shown in the lower part of Figure 4, the three-dimensional feature estimation unit 42 comprises an FC (fully connected) layer 54, a bias layer 55, and a normalization layer 56. The FC layer 54 and the bias layer 55 correspond to the third estimation model 32. For example, a 256-dimensional feature vector obtained by combining a 128-dimensional color feature vector CF and a 128-dimensional depth feature vector DF is dimensionally compressed to a 128-dimensional feature vector by the FC layer 54, the bias of the dimensionally compressed feature vector from the center is corrected by the bias layer 55, and it is normalized to a unit hypersphere by the normalization layer 56, thereby being converted into a 128-dimensional three-dimensional feature vector TF. Thus, the color feature vector CF, the depth feature vector DF, and the three-dimensional feature vector TF may have the same number of dimensions. In this case, the FC layer 54 only needs to dimensionally compress the feature vector obtained by combining the color feature vector CF and the depth feature vector DF to the above number of dimensions.

[0055] Figure 5 is a block diagram showing the configuration of a learning unit 60 that performs machine learning on the FC layer 54 and bias layer 55 of the three-dimensional feature estimation unit 42 shown in Figure 4. The learning unit 60 may be provided in the server device 3 or in the communication terminal 2.

[0056] As shown in Figure 5, the learning unit 60 comprises a first FC layer 61, a batch center layer 62, a normalization layer 63, a second FC layer 64, and a loss layer 65. The first FC layer 61 and the normalization layer 63 are the same as the FC layer 54 and the normalization layer 56 shown in Figure 4, respectively.

[0057] The batch center layer 62 learns bias data that shows the deviation from the center of the dimensionality-reduced feature vector. The second FC layer receives the 128-dimensional feature vector from the normalization layer 63 and the representative vector for each class as input, and calculates the similarity (e.g., cosine similarity) between the feature vector and the representative vector for each class.

[0058] The loss layer 65 obtains a loss function (e.g., Arcface loss function) from the similarity of each class from the second FC layer, and acquires weighting data to input to the first FC layer 61 and bias data to input to the batch center layer 62, which are used for class classification training to reduce the intra-class variance and increase the inter-class variance of the feature vectors. These weighting data and bias data are then input to the FC layer 54 and bias layer 55 of the three-dimensional feature estimation unit 42, respectively.

[0059] In this embodiment, facial recognition using a three-dimensional feature vector (TF) showed a significant improvement in accuracy, with the recognition rate reduced to half compared to facial recognition using a color feature vector (CF).

[0060] [Embodiment 2] <Outline Configuration of Authentication System 1A> Next, an authentication system 1A according to another embodiment of the present disclosure will be described with reference to Figures 6 and 7. For the sake of convenience of explanation, components having the same function as those described in Embodiment 1 will be denoted by the same reference numerals, and their descriptions will not be repeated.

[0061] Figure 6 is a schematic diagram showing the configuration of the authentication system 1A. As shown in Figure 6, the authentication system 1A includes an authentication device 2A, a server device 3A, cameras 4A and 4B, an access control device 5A, and a gate G. Note that the gate G is optional.

[0062] The authentication device 2A, cameras 4A and 4B, access control device 5A, and server device 3A are connected via network N. Furthermore, the access control device 5A and gate G are connected by wired or wireless means.

[0063] The access control device 5A may also have an authentication device 2A built into it. Alternatively, the authentication device 2A and the server device 3A may be configured as a single information processing device.

[0064] Authentication device 2A is installed, for example, inside a building such as a company, and is used to authenticate the faces of multiple people M1, M2, and M3 passing through gate G, which is located at the entrance of the company. Authentication device 2A is used, for example, to manage the entry and exit of company employees. Note that authentication device 2A has the same configuration as communication terminal 2 shown in Figure 2, so its explanation will be omitted.

[0065] Cameras 4A and 4B can be installed near gate G, for example, above gate G. Camera 4A captures images of the faces of people attempting to enter through gate G. Camera 4B captures images of the faces of people attempting to exit through gate G. The placement locations and number of cameras 4A and 4B can be changed as needed.

[0066] Cameras 4A and 4B may be built into the authentication device 2A. Cameras 4A and 4B may also have zoom lenses. The acquisition unit 20 of the authentication device 2A may acquire images from cameras 4A and 4B that have been magnified by the zoom lenses of the cameras 4A and 4B. This makes it possible to recognize the faces of people who are located far from gate G.

[0067] Figure 7 is a block diagram showing the configuration of the access control device 5A. As shown in Figure 7, the access control device 5A has a CPU (Central Processing Unit) 71, a ROM (Read Only Memory) 72, a RAM (Random Access Memory) 73, a communication unit 74, an input unit 75, and an output unit 76, which are connected by an internal bus.

[0068] The CPU 71 controls the overall operation of the access control device 5A. The ROM 72 stores programs and other data for the CPU 71 to control various operations. The RAM 73 is used as a storage area to temporarily record data and signals used by the CPU 71 when executing the above programs, or as a work area for data processing. The CPU 71 controls the opening and closing operations of the gate G, etc., based on the programs read from the ROM 72.

[0069] The communication unit 74 communicates with the authentication device 2A via the network N. The communication unit 74 receives, for example, information regarding the results of the authentication process by the authentication device 2A. The input unit 75 is an input interface for receiving user input operations. The input unit 75 has a touch panel and buttons, etc. The output unit 76 outputs a signal to open and close the gate G based on control by the CPU 71.

[0070] Server device 3A stores the color feature vector CF, depth feature vector DF, and three-dimensional feature vector TF of the faces of employees of companies, etc., which are the objects of authentication by authentication device 2A, and associates them with an ID that identifies each employee.

[0071] When a person approaches gate G and moves to position M3, the authentication device 2A recognizes and tracks the person included in the color image CI from the color image CI data. When the person approaches gate G and moves to position M2, and the depth included in the depth image DI enters the depth range usable for facial recognition, the authentication device 2A estimates the depth feature vector DF from the depth image DI data.

[0072] When the person approaches gate G further and moves to position M3, and the color image CI becomes an image usable for facial recognition, the authentication device 2A estimates a color feature vector CF from the color image CI data. Then, the authentication device 2A estimates a three-dimensional feature vector TF from the depth feature vector DF and the color feature vector CF. Therefore, the authentication device 2A can authenticate the face of the person using at least one of the estimated depth feature vector DF, color feature vector CF, and three-dimensional feature vector TF.

[0073] [Modification] In the above embodiment, the authentication device 2A is used to authenticate people passing through gate G located inside the building, but the embodiment is not limited to this. For example, the authentication device 2A may also be used to authenticate animals such as pets passing through gate G.

[0074] Furthermore, an authentication device 2A may be used to identify the license plate of a vehicle passing through gate G from a color image, and to identify the vehicle type from the color image and / or depth image, in order to authenticate the vehicle associated with that vehicle type and license plate. In this case, it is possible to detect suspicious vehicles with replaced license plates. Thus, the object to be authenticated (identified) is not limited to a human face.

[0075] This disclosure demonstrates that by utilizing AI (Artificial Intelligence), it will be possible to perform high-speed authentication of human faces and other features, forming an innovative technological foundation for the telecommunications business and contributing to the achievement of Sustainable Development Goal 9, "Build resilient infrastructure, promote inclusive and sustainable industrialization and foster innovation."

[0076] <Example of implementation using software> The functions of each device (hereinafter referred to as "device") that constitutes the authentication systems 1 and 1A can be realized by a program that causes a computer to function as the device, and by a program that causes a computer to function as each control block of the device (particularly each part included in the control unit 10).

[0077] In this case, the device includes a computer having at least one control device (e.g., a processor) and at least one storage device (e.g., memory) as hardware for executing the program. By executing the program using this control device and storage device, the functions described in each of the embodiments are realized.

[0078] The above program may be recorded on one or more computer-readable recording media, not temporary ones. These recording media may or may not be provided by the above device. In the latter case, the program may be supplied to the above device via any wired or wireless transmission medium.

[0079] Furthermore, some or all of the functions of each of the above control blocks can also be realized by logic circuits. For example, an integrated circuit in which logic circuits functioning as each of the above control blocks are formed is also included in the scope of the present invention. In addition, it is also possible to realize the functions of each of the above control blocks by, for example, a quantum computer.

[0080] Furthermore, each of the processes described in the above embodiments may be performed by AI (Artificial Intelligence). In this case, the AI ​​may operate on the control device described above, or it may operate on other devices (for example, an edge computer or a cloud server).

[0081] (Additional Notes) This disclosure is not limited to the embodiments described above, and various modifications are possible within the scope of the claims. Embodiments obtained by appropriately combining the technical means disclosed in different embodiments are also included in the technical scope of this disclosure.

[0082] 1, 1A Authentication system 2 Communication terminal (object identification device) 2A Authentication device (object identification device) 3, 3A Server device 4A, 4B Camera 5A Access control device 10 Control unit 11 Storage unit 12 Camera unit 13 Communication unit 14 Operation unit 15 Display unit 20 Acquisition unit 21 Feature estimation unit 22 Selection unit 23 Registration unit 24 Authentication unit (identification unit) 30 First estimation model 31 Second estimation model 32 Third estimation model 33 Registered data 40 Color feature estimation unit 41 Depth feature estimation unit 42 Three-dimensional feature estimation unit 51, 53, 56 Normalization layer 50, 52 CNN layer 54 FC layer 55 Bias layer 60 Learning unit 61 First FC layer 62 Batch center layer 63 Normalization layer 64 2nd FC layer 65 loss layer

Claims

1. An object identification device for identifying an object contained in a two-dimensional color image and a two-dimensional depth image, comprising: an acquisition unit for acquiring data from the color image and data from the depth image; a color feature estimation unit for estimating a color feature vector indicating the characteristics of the object contained in the color image using a first estimation model from the color image data; a depth feature estimation unit for estimating a depth feature vector indicating the characteristics of the object contained in the depth image using a second estimation model from the depth image data; a three-dimensional feature estimation unit for estimating a three-dimensional feature vector indicating the characteristics of the object contained in a three-dimensional image including the color image and the depth image using a third estimation model from the color feature vector and the depth feature vector; and an identification unit for identifying the object using any of the color feature vector, the depth feature vector, and the three-dimensional feature vector.

2. The object identification device according to claim 1, further comprising a selection unit that selects one of the color feature vector, the depth feature vector, and the three-dimensional feature vector based on the color image data and the depth image data, wherein the identification unit identifies the object using the feature vector selected by the selection unit.

3. The object identification device according to claim 1, further comprising a storage unit for storing the first estimation model, the second estimation model, and the third estimation model.

4. The object identification device according to claim 1, wherein the third estimation model includes a fully connected layer that reduces the dimensionality of a vector obtained by combining the color feature vector and the depth feature vector; a bias layer that corrects the deviation of the dimensionality-reduced vector from the center; and a normalization layer that generates the three-dimensional feature vector by normalizing the corrected vector.

5. The object identification device according to claim 1, wherein the color image is an image of the region of the object extracted from the original color image by object detection, and the depth image is an image of the region corresponding to the region of the object extracted from the original depth image.

6. The object identification device according to claim 1, wherein the object is a human face.

7. An object identification program for causing a computer to function as an object identification device according to claim 1, comprising a color feature estimation unit, a depth feature estimation unit, a three-dimensional feature estimation unit, and an identification unit, the object identification program for causing the computer to function as an identification unit.

8. An object identification method for identifying an object contained in a two-dimensional color image and a two-dimensional depth image, comprising: an acquisition step of acquiring data of the color image and data of the depth image; a color feature estimation step of estimating a color feature vector indicating the characteristics of the object contained in the color image using a first estimation model from the data of the color image; a depth feature estimation step of estimating a depth feature vector indicating the characteristics of the object contained in the depth image using a second estimation model from the data of the depth image; a three-dimensional feature estimation step of estimating a three-dimensional feature vector indicating the characteristics of the object contained in a three-dimensional image including the color image and the depth image using a third estimation model from the color feature vector and the depth feature vector; and an identification step of identifying the object using any of the color feature vector, the depth feature vector, and the three-dimensional feature vector.