Object identification device, object identification program, and object identification method
The object identification device uses color and depth image data through feature estimation models to address lighting and distance limitations, enhancing recognition accuracy and adaptability across varied conditions.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- SOFTBANK CORPORATION
- Filing Date
- 2025-01-16
- Publication Date
- 2026-05-26
Smart Images

Figure 0007866085000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to an object identification device, an object identification program, and an object identification method.
Background Art
[0002] In the face recognition method described in Patent Document 1, first, a first face image captured by an RGB (red, green, blue) camera and a second face image captured by a NIR (near-infrared) camera with a different shooting direction from the RGB camera are acquired. Next, while acquiring a face fusion image obtained by image fusion of the first face image and the second face image, three-dimensional depth information is acquired by stereo matching based on the first face image and the second face image. Then, face recognition is performed based on the face fusion image and the three-dimensional depth information, and a value characterizing the face recognition reliability is obtained.
[0003] Also, the camera "Intel (registered trademark) RealSense (trademark) Depth Camera D415" manufactured by Intel has a dedicated RGB sensor, two IR sensors, and an IR projector. By performing stereo matching on two IR images from the two IR sensors, a depth image, which is a two-dimensional image, is acquired.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Summary of the Invention
Means for Solving the Problems
[0005] An object identification device according to one embodiment of the present disclosure is an object identification device that identifies an object contained in a two-dimensional color image and a two-dimensional depth image, comprising: an acquisition unit that acquires data of the color image and data of the depth image; a color feature estimation unit that estimates a color feature vector indicating the characteristics of the object contained in the color image using a first estimation model from the data of the color image; a depth feature estimation unit that estimates a depth feature vector indicating the characteristics of the object contained in the depth image using a second estimation model from the data of the depth image; a three-dimensional feature estimation unit that estimates a three-dimensional feature vector indicating the characteristics of the object contained in a three-dimensional image including the color image and the depth image using a third estimation model from the color feature vector and the depth feature vector; and an identification unit that identifies the object using any of the color feature vector, the depth feature vector, and the three-dimensional feature vector.
[0006] An object identification program according to one embodiment of this disclosure is an object identification program that causes a computer to function as a color feature estimation unit, a depth feature estimation unit, a three-dimensional feature estimation unit, and an identification unit.
[0007] An object identification method according to one embodiment of the present disclosure is an object identification method for identifying objects contained in a two-dimensional color image and a two-dimensional depth image, comprising: an acquisition step of acquiring data of the color image and data of the depth image; and using a first estimation model from the data of the color image 、 The aforementioned color image The color that indicates the characteristics of the object included in the above-mentioned object. A color feature estimation step in which feature vectors are estimated, and a second estimation model using the depth image data. 、 The aforementioned depth image Depth indicating the characteristics of the object included A depth feature estimation step for estimating feature vectors, and the above Color characteristics The characteristic vector and the depth degree characteristic Using a third estimation model from the characteristic vector, a three-dimensional image including the color image and the depth image is generated. Three-dimensional representation showing the characteristics of the object included in the aforementioned object. A three-dimensional feature estimation step for estimating feature vectors, and the aforementioned Color characteristics Character vector, the depth degree characteristic The characteristic vector, and the cubic GentokuThe method includes an identification step of identifying the object using one of the characteristic vectors. [Brief explanation of the drawing]
[0008] [Figure 1] This is a schematic diagram showing the configuration of an authentication system according to one embodiment of this disclosure. [Figure 2] This block diagram shows the configuration of the communication terminals in the authentication system described above. [Figure 3] This flowchart shows an example of the authentication process flow in the above-mentioned communication terminal. [Figure 4] This block diagram shows specific examples of the color feature estimation unit, depth feature estimation unit, and three-dimensional feature estimation unit in the above-mentioned communication terminal. [Figure 5] This block diagram shows the configuration of the learning unit that performs machine learning on the FC layer and bias layer in the three-dimensional feature estimation unit described above. [Figure 6] This is a schematic diagram showing the configuration of an authentication system according to another embodiment of the present disclosure. [Figure 7] This block diagram shows the configuration of the access control device in the authentication system described above. [Modes for carrying out the invention]
[0009] Hereinafter, one embodiment of this disclosure will be described in detail with reference to the drawings. For ease of understanding, the background and challenges of this disclosure will be described first, followed by a detailed description of the disclosure.
[0010] In the following, the distance from the camera to the object will be referred to as the "shooting distance." Conventionally, 3D recognition technology is known that uses machine learning models to identify objects contained in color images and depth images.
[0011] However, in the above 3D identification technology, when the amount of light incident on the RGB camera that creates the above color image is extremely small or extremely large, it is difficult to identify the above object. Also, depending on the sensor that creates the depth image, the range of shooting distances in which the depth image can be created is limited. Therefore, when the above object does not exist within the range of the above shooting distance, it is difficult to identify the above object. The present disclosure aims to identify an object in various environments.
[0012] An object identification device according to the present disclosure is an object identification device that identifies an object included in a color image, which is a two-dimensional image, and a depth image, which is a two-dimensional image, and includes an acquisition unit that acquires data of the color image and data of the depth image, and a first estimation model is used from the data of the color image to 、 the color image The color that indicates the characteristics of the object included in the above-mentioned object. a color feature estimation unit that estimates a feature vector, and a second estimation model is used from the data of the depth image to 、 the depth image Depth indicating the characteristics of the object included a depth feature estimation unit that estimates a feature vector, and the Color characteristics feature vector and the depth degree characteristic feature vector, and a third estimation model is used to estimate a feature vector of a three-dimensional image including the color image and the depth image Three-dimensional representation showing the characteristics of the object included in the aforementioned object. a three-dimensional feature estimation unit, and the Color characteristics feature vector, the depth degree characteristic feature vector, and the three-dimensional Gentoku feature vector, and an identification unit that identifies the object using any of them.
[0013] According to the above configuration, when the amount of light incident on the RGB camera that creates the above color image is extremely small or extremely large, it is possible to identify the object using the depth degree characteristic feature vector estimated from the depth image. Also, when the above object does not exist within the range of the above shooting distance, it is possible to identify the object using the Color characteristics feature vector estimated from the color image. Also, when the amount of light incident on the above RGB camera is appropriate and the above object exists within the range of the above shooting distance, the depth degree characteristicFeature vector and the above Color characteristics Estimated from the feature vector and the above three dimensional It is possible to identify an object using the feature vector. Therefore, it is possible to identify an object in various environments.
[0014] 〔Embodiment 1〕 <Schematic configuration of authentication system 1> Hereinafter, an authentication system 1 according to an embodiment of the present disclosure will be described with reference to FIGS. 1 to 3. FIG. 1 is a block diagram showing the overall configuration of the authentication system 1. The authentication system 1 is a system that identifies and authenticates a human face as an object. As shown in FIG. 1, the authentication system 1 includes a communication terminal 2 and a server device 3.
[0015] The communication terminal 2 and the server device 3 are connected via a network N such as the Internet. The communication terminal 2 has a function of performing face authentication of the user M who uses the communication terminal 2, as indicated by the arrow R.
[0016] In addition to the Internet, the network N may be, for example, a LAN (Local Area Network), a mobile communication system such as 4G, 5G, 6G, LTE (Long Term Evolution), or Wi-Fi (registered trademark).
[0017] [Configuration of communication terminal 2] Next, the configuration of the communication terminal 2 will be described with reference to FIG. 2. FIG. 2 is a block diagram showing the configuration of the communication terminal 2. The communication terminal 2 is a smartphone, a tablet terminal, a PC (Personal Computer), etc. used by the user M. As shown in FIG. 2, the communication terminal 2 includes a control unit 10, a storage unit 11, a camera unit 12, a communication unit 13, an operation unit 14, and a display unit 15.
[0018] The control unit 10 comprehensively controls the operation of various components of the communication terminal 2 and is composed of a computer including, for example, a CPU (Central Processing Unit) and memory. The operation control of these various components is performed by having the computer execute a control program. Further details of the control unit 10 will be described later.
[0019] The memory unit 11 is for recording information and is composed of storage devices such as a hard disk or flash memory. Details of the memory unit 11 will be described later.
[0020] The camera unit 12 is used for taking photographs and is built into the communication terminal 2. The camera unit 12 has lenses on both the side of the communication terminal 2 that is on the same side as the operation unit 14 and on the side opposite to the operation unit 14. The camera unit 12 is configured to capture both video and still images.
[0021] In this embodiment, the camera unit 12 creates a three-dimensional image (hereinafter abbreviated as "three-dimensional image") that includes a two-dimensional RGB color image and a two-dimensional depth image. The depth image can be created using known methods such as stereo, ToF (Time of Flight), or infrared sensor methods.
[0022] The communication unit 13 is connected to the network N by wire or wireless connection and communicates with the server device 3 and other communication terminals via the network N. The communication unit 13 consists of a NIC (Network Interface Card), an antenna, and the like.
[0023] The operation unit 14, for example, consists of a touch panel and accepts various operations from user M. The operation unit 14 creates input data corresponding to the accepted operations and transmits it to the control unit 10. The operation unit 14 may also have buttons or the like for inputting characters, numbers, etc.
[0024] The display unit 15 is a display device for displaying various information from the control unit 10, and is, for example, a liquid crystal display (LCD) or an organic electroluminescent display (OLED).
[0025] (Details of the control unit and memory unit) As shown in Figure 2, the control unit 10 includes an acquisition unit 20, a feature estimation unit 21, a selection unit 22, a registration unit 23, and an authentication unit 24 (identification unit). The storage unit 11 stores the first estimation model 30, the second estimation model 31, the third estimation model 32, and the registered data 33.
[0026] The first estimation model 30 is a learning model that estimates a color feature vector CF, which represents the facial features contained in the color image CI, from the color image CI data. The second estimation model 31 is a learning model that estimates a depth feature vector DF, which represents the facial features contained in the depth image DI, from the depth image DI data. The third estimation model 32 is a learning model that estimates a three-dimensional feature vector TF, which represents the facial features contained in the three-dimensional image TI, which includes the color image CI and depth image DI, from the color feature vector CF and depth feature vector DF. The first estimation model 30, the second estimation model 31, and the third estimation model 32 are pre-trained on the server device 3 and stored in the memory unit 11.
[0027] The registered data 33 is data registered for facial recognition of user M, and includes a color feature vector CF, a depth feature vector DF, and a three-dimensional feature vector TF that represent the features of user M's face.
[0028] The acquisition unit 20 acquires color image CI and depth image DI data from the camera unit 12. The acquisition unit 20 sends the acquired color image CI and depth image DI data to the feature estimation unit 21 and the selection unit 22.
[0029] The acquisition unit 20 may also perform face detection (object detection) on the original color image acquired from the camera unit 12 to identify the face region, extract the identified face region from the original color image, and use the extracted face region image as the color image CI used by the feature estimation unit 21. Alternatively, the acquisition unit 20 may extract the region corresponding to the face region from the original depth image acquired from the camera unit 12, and use the extracted region image as the depth image DI used by the feature estimation unit 21.
[0030] In this case, the sizes of the color image CI and the depth image DI can be reduced compared to the sizes of the original color image and the original depth image, respectively. As a result, the processing load on the feature estimation unit 21 is reduced, and / or the processing speed is improved. Note that the pixel region of the face may include the pixels surrounding that pixel region.
[0031] The feature estimation unit 21 includes a color feature estimation unit 40, a depth feature estimation unit 41, and a three-dimensional feature estimation unit 42. The color feature estimation unit 40 estimates the color feature vector CF from the color image CI data from the acquisition unit 20 using the first estimation model 30 of the storage unit 11. The depth feature estimation unit 41 estimates the depth feature vector DF from the depth image DI data from the acquisition unit 20 using the second estimation model 31 of the storage unit 11. The three-dimensional feature estimation unit 42 estimates the three-dimensional feature vector TF from the color feature vector CF and the depth feature vector DF using the third estimation model 32 of the storage unit 11.
[0032] The selection unit 22 selects images usable for face recognition from the color image CI and depth image DI data based on the color image CI and depth image DI data from the acquisition unit 20. Examples of color image CIs that cannot be used for face recognition include images with uneven brightness due to sunlight, dark images, overexposed images, and backlit images. Examples of depth image DIs that cannot be used for face recognition include images in which the depth included in the depth image DI falls outside the range of depths usable for face recognition.
[0033] If the selection unit 22 selects only the color image CI, it instructs the feature estimation unit 21 to send the color feature vector CF estimated by the color feature estimation unit 40 to the authentication unit 24. If the selection unit 22 selects only the depth image DI, it instructs the feature estimation unit 21 to send the depth feature vector DF estimated by the depth feature estimation unit 41 to the authentication unit 24. If the selection unit 22 selects both the color image CI and the depth image DI, it instructs the feature estimation unit 21 to send the three-dimensional feature vector TF estimated by the three-dimensional feature estimation unit 42 to the authentication unit 24.
[0034] Furthermore, when registering a user for facial recognition, the selection unit 22 instructs the feature estimation unit 21 to send the color feature vector CF, the depth feature vector DF, and the three-dimensional feature vector TF to the registration unit 23 only when both the color image CI and the depth image DI are selected.
[0035] The registration unit 23 performs user registration for facial recognition. Specifically, the registration unit 23 stores the color feature vector CF, depth feature vector DF, and three-dimensional feature vector TF from the feature estimation unit 21 in the registration data 33 of the storage unit 11.
[0036] The authentication unit 24 performs face authentication (identification) of user M by comparing one of the feature vectors from the feature estimation unit 21 (color feature vector CF, depth feature vector DF, and three-dimensional feature vector TF) with the corresponding feature vector contained in the registered data 33 of the storage unit 11. Specifically, the authentication unit 24 calculates the similarity between the feature vector from the feature estimation unit 21 and the feature vector from the registered data 33, and determines that the face authentication was successful if the similarity is above a predetermined threshold. At this time, the authentication unit 24 displays information indicating that authentication was successful via the display unit 15 and accepts operations from user M via the operation unit 14.
[0037] <Authentication process> Figure 3 is a flowchart showing an example of the authentication process flow by the communication terminal 2 with the above configuration. As shown in Figure 3, first, the acquisition unit 20 uses the image of the face region extracted by face detection from the original color image acquired from the camera unit 12 as the face image CI to be used in this process, and the image of the region corresponding to the face region extracted from the original depth image acquired from the camera unit 12 as the depth image DI to be used in this process (S10). Next, the selection unit 22 determines whether the color image CI is an image that can be used for face authentication based on the data of the color image CI (S11).
[0038] If the color image CI is an image that can be used for face recognition (YES in S11), the selection unit 22 determines whether or not the depth image DI is an image that can be used for face recognition based on the depth image DI data (S12).
[0039] If the depth image DI is an image that can be used for face recognition (YES in S12), the authentication unit 24 performs face recognition using the three-dimensional feature vector TF estimated by the three-dimensional feature estimation unit 42 from the color feature vector CF estimated by the color feature estimation unit 40 from the color image CI data and the depth feature vector DF estimated by the depth feature estimation unit 41 from the depth image DI data (S13). After that, the authentication process is terminated.
[0040] On the other hand, if the depth image DI is not an image that can be used for face recognition (NO in S12), the authentication unit 24 performs face recognition using the color feature vector CF estimated by the color feature estimation unit 40 from the color image CI data (S14). After that, the authentication process is terminated.
[0041] On the other hand, if the color image CI is not an image that can be used for face recognition (NO in S11), the selection unit 22 determines whether or not the depth image DI is an image that can be used for face recognition based on the depth image DI data (S15).
[0042] If the depth image DI is an image that can be used for facial recognition (YES in S15), the authentication unit 24 performs facial recognition using the depth feature vector DF estimated by the depth feature estimation unit 41 from the depth image DI data (S16). After that, the authentication process ends.
[0043] On the other hand, if the depth image DI is not an image that can be used for facial recognition (NO in S15), the authentication unit 24 outputs information via the display unit 15 indicating that facial recognition could not be performed with the acquired color image CI and depth image DI (S17). After that, the authentication process is terminated.
[0044] Based on the above, the communication terminal 2 of this embodiment can perform facial recognition using a three-dimensional feature vector TF when the color image CI and depth image DI are images that can be used for facial recognition. Furthermore, if the color image CI is an image that can be used for facial recognition, but the depth image DI is an image that cannot be used for facial recognition, facial recognition can be performed using a color feature vector CF. Furthermore, if the depth image DI is an image that can be used for facial recognition, but the color image CI is an image that cannot be used for facial recognition, facial recognition can be performed using a depth feature vector DF. Therefore, facial recognition can be performed in various environments.
[0045] Furthermore, it is not necessary to obtain the first estimated model 30, the second estimated model 31, and the third estimated model 32 from the server device 3 via the communication network N. Therefore, the communication terminal of this embodiment can perform facial recognition even if it is not connected to the communication network N.
[0046] <Examples> Figure 4 is a block diagram showing specific examples of the color feature estimation unit 40, the depth feature estimation unit 41, and the three-dimensional feature estimation unit 42.
[0047] As shown in the upper part of Figure 4, the color feature estimation unit 40 comprises a CNN (Convolutional Neural Network) layer 50 and a normalization layer 51. The CNN layer 50 corresponds to the first estimation model 30 and is a learning model based on known architectures such as ArcFace, SphereFace, and CosFace. For example, the data of a 112×112 pixel color image CI is converted into a 128-dimensional vector by the CNN layer 50 and normalized to a unit hypersphere by the normalization layer 51, thereby converting it into a 128-dimensional color feature vector CF.
[0048] As shown in the middle section of Figure 4, the depth feature estimation unit 41 comprises a CNN layer 52 and a normalization layer 53. The CNN layer 52 corresponds to the second estimation model 31 and is a learning model based on the known architecture described above. The CNN layer 50 of the color feature estimation unit 40 and the CNN 52 of the depth feature estimation unit 41 are machine-learned separately. For example, 112 × 112 pixel depth image DI data is converted into a 128-dimensional vector by the CNN layer 52 and normalized to a unit hypersphere by the normalization layer 53, thereby converting it into a 128-dimensional depth feature vector DF.
[0049] As shown in the lower part of Figure 4, the three-dimensional feature estimation unit 42 comprises an FC (fully connected) layer 54, a bias layer 55, and a normalization layer 56. The FC layer 54 and the bias layer 55 correspond to the third estimation model 32. For example, a 256-dimensional feature vector obtained by combining a 128-dimensional color feature vector CF and a 128-dimensional depth feature vector DF is dimensionally compressed to a 128-dimensional feature vector by the FC layer 54, the bias of the dimensionally compressed feature vector from the center is corrected by the bias layer 55, and it is normalized to a unit hypersphere by the normalization layer 56, thereby being converted into a 128-dimensional three-dimensional feature vector TF. Thus, the color feature vector CF, the depth feature vector DF, and the three-dimensional feature vector TF may have the same number of dimensions. In this case, the FC layer 54 only needs to dimensionally compress the feature vector obtained by combining the color feature vector CF and the depth feature vector DF to the above number of dimensions.
[0050] Figure 5 is a block diagram showing the configuration of a learning unit 60 that performs machine learning on the FC layer 54 and bias layer 55 of the three-dimensional feature estimation unit 42 shown in Figure 4. The learning unit 60 may be provided in the server device 3 or in the communication terminal 2.
[0051] As shown in Figure 5, the learning unit 60 comprises a first FC layer 61, a batch center layer 62, a normalization layer 63, a second FC layer 64, and a loss layer 65. The first FC layer 61 and the normalization layer 63 are the same as the FC layer 54 and the normalization layer 56 shown in Figure 4, respectively.
[0052] The batch center layer 62 learns bias data that shows the deviation from the center of the dimensionality-reduced feature vector. The second FC layer receives the 128-dimensional feature vector from the normalization layer 63 and the representative vector for each class as input, and calculates the similarity (e.g., cosine similarity) between the feature vector and the representative vector for each class.
[0053] The loss layer 65 calculates a loss function (e.g., Arcface loss function) from the similarity of each class from the second FC layer and learns it as a classification to obtain weighting data to input to the first FC layer 61 and bias data to input to the batch center layer 62, which reduce the intra-class variance and increase the inter-class variance of the feature vectors. These weighting data and bias data are then input to the FC layer 54 and bias layer 55 of the three-dimensional feature estimation unit 42, respectively.
[0054] In this embodiment, facial recognition using a three-dimensional feature vector (TF) showed a significant improvement in accuracy, with the recognition rate reduced to half compared to facial recognition using a color feature vector (CF).
[0055] [Embodiment 2] <Outline configuration of authentication system 1A> Next, an authentication system 1A according to another embodiment of this disclosure will be described with reference to Figures 6 and 7. For the sake of convenience, components having the same function as those described in Embodiment 1 above will be denoted by the same reference numerals, and their descriptions will not be repeated.
[0056] Figure 6 is a schematic diagram showing the configuration of the authentication system 1A. As shown in Figure 6, the authentication system 1A includes an authentication device 2A, a server device 3A, cameras 4A and 4B, an access control device 5A, and a gate G. Note that gate G is optional.
[0057] The authentication device 2A, cameras 4A and 4B, access control device 5A, and server device 3A are connected via network N. Access control device 5A and gate G are connected by wired or wireless connection.
[0058] The access control device 5A may also have an authentication device 2A built into it. Alternatively, the authentication device 2A and the server device 3A may be configured as a single information processing device.
[0059] Authentication device 2A is installed, for example, inside a building such as a company, and is used to authenticate the faces of multiple people M1, M2, and M3 passing through gate G, which is located at the entrance of the company. Authentication device 2A is used, for example, to manage the entry and exit of company employees. Note that authentication device 2A has the same configuration as communication terminal 2 shown in Figure 2, so its explanation will be omitted.
[0060] Cameras 4A and 4B can be installed near gate G, for example, above gate G. Camera 4A captures images of the faces of people attempting to enter through gate G. Camera 4B captures images of the faces of people attempting to exit through gate G. The placement locations and number of cameras 4A and 4B can be changed as needed.
[0061] Cameras 4A and 4B may be built into the authentication device 2A. Cameras 4A and 4B may also have zoom lenses. The acquisition unit 20 of the authentication device 2A may acquire images from cameras 4A and 4B that have been magnified by the zoom lenses of the cameras 4A and 4B. This makes it possible to recognize the faces of people who are located far from gate G.
[0062] Figure 7 is a block diagram showing the configuration of the access control device 5A. As shown in Figure 7, the access control device 5A has a CPU (Central Processing Unit) 71, a ROM (Read Only Memory) 72, a RAM (Random Access Memory) 73, a communication unit 74, an input unit 75, and an output unit 76, which are connected by an internal bus.
[0063] The CPU 71 controls the overall operation of the access control device 5A. The ROM 72 stores programs and other data for the CPU 71 to control various operations. The RAM 73 is used as a storage area to temporarily record data and signals used by the CPU 71 when executing the above programs, or as a work area for data processing. The CPU 71 controls the opening and closing operations of gate G and other operations based on the programs read from the ROM 72.
[0064] The communication unit 74 communicates with the authentication device 2A via the network N. The communication unit 74 receives, for example, information regarding the results of the authentication process by the authentication device 2A. The input unit 75 is an input interface for receiving user input operations. The input unit 75 has a touch panel and buttons, etc. The output unit 76 outputs a signal to open and close the gate G based on control by the CPU 71.
[0065] Server device 3A stores the color feature vector CF, depth feature vector DF, and three-dimensional feature vector TF of the faces of employees of companies, etc., which are the objects of authentication by authentication device 2A, and associates them with an ID that identifies each employee.
[0066] When a person approaches gate G and moves to position M3, the authentication device 2A recognizes and tracks the person included in the color image CI from the color image CI data. When the person approaches gate G and moves to position M2, and the depth included in the depth image DI enters the depth range usable for facial recognition, the authentication device 2A estimates the depth feature vector DF from the depth image DI data.
[0067] When the person approaches gate G further and moves to position M3, and the color image CI becomes an image usable for facial recognition, the authentication device 2A estimates the color feature vector CF from the color image CI data. Then, the authentication device 2A estimates the three-dimensional feature vector TF from the depth feature vector DF and the color feature vector CF. Thus, the authentication device 2A can authenticate the person's face using at least one of the estimated depth feature vector DF, the color feature vector CF, and the three-dimensional feature vector TF.
[0068] [Variation] In the above embodiment, the authentication device 2A is used to authenticate people passing through gate G located inside the building, but the embodiment is not limited to this. For example, the authentication device 2A may also be used to authenticate animals such as pets passing through gate G.
[0069] Furthermore, authentication device 2A may be used to identify the license plate of a vehicle passing through gate G from a color image, and to identify the vehicle type from the color image and / or depth image, in order to authenticate the vehicle associated with that vehicle type and license plate. In this case, suspicious vehicles with replaced license plates can be detected. Thus, the object to be authenticated (identified) is not limited to a human face.
[0070] This disclosure demonstrates that by utilizing AI (Artificial Intelligence), it will be possible to perform high-speed authentication of human faces and other features, forming an innovative technological foundation for the telecommunications business and contributing to the achievement of Sustainable Development Goal 9, "Build resilient infrastructure, promote inclusive and sustainable industrialization and foster innovation."
[0071] <Examples of implementation using software> The functions of each device (hereinafter referred to as "device") that constitutes the authentication systems 1 and 1A are programs that cause a computer to function as the device, and these programs can be realized by programs that cause a computer to function as each control block of the device (especially each part included in the control unit 10).
[0072] In this case, the device includes a computer having at least one control device (e.g., a processor) and at least one storage device (e.g., memory) as hardware for executing the program. By executing the program using this control device and storage device, the functions described in each of the embodiments are realized.
[0073] The above program may be recorded on one or more computer-readable recording media, not temporary ones. These recording media may or may not be provided by the above device. In the latter case, the program may be supplied to the above device via any wired or wireless transmission medium.
[0074] Furthermore, some or all of the functions of each of the above control blocks can also be realized by logic circuits. For example, an integrated circuit in which logic circuits functioning as each of the above control blocks are formed is also included in the scope of the present invention. In addition, it is also possible to realize the functions of each of the above control blocks by, for example, a quantum computer.
[0075] Furthermore, each process described in the above embodiments may be performed by AI (Artificial Intelligence). In this case, the AI may operate on the control device described above, or it may operate on other devices (for example, an edge computer or a cloud server). <Summary> This disclosure includes at least the following aspects:
[0076] Object identification device according to aspect 1 of the present disclosure, object identification device for identifying an object contained in a two-dimensional color image and a two-dimensional depth image, comprising: an acquisition unit for acquiring data of the color image and data of the depth image; a color feature estimation unit for estimating a color feature vector indicating the characteristics of the object contained in the color image using a first estimation model from the data of the color image; a depth feature estimation unit for estimating a depth feature vector indicating the characteristics of the object contained in the depth image using a second estimation model from the data of the depth image; a three-dimensional feature estimation unit for estimating a three-dimensional feature vector indicating the characteristics of the object contained in a three-dimensional image including the color image and the depth image using a third estimation model from the color feature vector and the depth feature vector; and an identification unit for identifying the object using any of the color feature vector, the depth feature vector, and the three-dimensional feature vector.
[0077] According to the above configuration, if the color image and depth image are images usable for facial recognition, the object can be identified using the three-dimensional feature vector. Furthermore, if the color image is usable for facial recognition but the depth image is not, the object can be identified using the color feature vector. Also, if the depth image is usable for facial recognition but the color image is not, the object can be identified using the depth feature vector. Therefore, objects can be identified in various environments.
[0078] The object identification device according to Embodiment 2 of this disclosure, in Embodiment 1, based on the color image data and the depth image data, Color characteristicsCharacter vector, the depth degree characteristic The characteristic vector, and the cubic Gentoku The identification unit may further include a selection unit that selects one of the feature vectors, and the identification unit may identify the object using the feature vector selected by the selection unit.
[0079] The object identification device according to embodiment 3 of the present disclosure may further include a storage unit for storing the first estimation model, the second estimation model, and the third estimation model, in embodiment 1 or 2.
[0080] According to the above configuration, there is no need to obtain the first estimation model, the second estimation model, and the third estimation model from the server via a communication network. Therefore, the object identification device can identify objects even without being connected to a communication network.
[0081] The object identification device according to Embodiment 4 of this disclosure, in Embodiments 1 to 3, the third estimation model is Color characteristics The characteristic vector and the aforementioned depth degree characteristic A fully connected layer that reduces the dimensionality of the vector formed by combining the characteristic vector, a bias layer that corrects the bias from the center of the dimensionality-reduced vector, and the cubic vector that normalizes the corrected vector. Gentoku It may also include a normalization layer that generates a characteristic vector.
[0082] In the object identification device according to aspect 5 of the present disclosure, in aspects 1 to 4, the color image may be an image of the region of the object extracted from the original color image by object detection, and the depth image may be an image of the region corresponding to the region of the object extracted from the original depth image.
[0083] According to the above configuration, the sizes of the color image and the depth image can be reduced compared to the sizes of the original color image and the original depth image, respectively. As a result, the processing load in the color feature estimation unit, the depth feature estimation unit, and the three-dimensional feature estimation unit is reduced, and / or the processing speed is improved.
[0084] In the object identification device according to aspect 6 of this disclosure, the object may be a human face in aspects 1 to 5.
[0085] The object identification program according to aspect 7 of this disclosure is an object identification program for causing a computer to function as an object identification device as described in aspects 1 to 6 above, and may be an object identification program for causing a computer to function as a color feature estimation unit, a depth feature estimation unit, a three-dimensional feature estimation unit, and an identification unit.
[0086] Object identification method according to aspect 8 of the present disclosure is an object identification method for identifying objects contained in a two-dimensional color image and a two-dimensional depth image, comprising: an acquisition step of acquiring data of the color image and data of the depth image; and using a first estimation model from the data of the color image 、 The aforementioned color image The color that indicates the characteristics of the object included in the above-mentioned object. A color feature estimation step in which feature vectors are estimated, and a second estimation model using the depth image data. 、 The aforementioned depth image Depth indicating the characteristics of the object included A depth feature estimation step for estimating feature vectors, and the above Color characteristics The characteristic vector and the depth degree characteristic Using a third estimation model from the characteristic vector, a three-dimensional image including the color image and the depth image is generated. Three-dimensional representation showing the characteristics of the object included in the aforementioned object. A three-dimensional feature estimation step for estimating feature vectors, and the aforementioned Color characteristics Character vector, the depth degree characteristic The characteristic vector, and the cubic Gentoku The method includes an identification step of identifying the object using one of the characteristic vectors.
[0087] The method described above produces the same effects as in Embodiment 1.
[0088] (Additional notes) This disclosure is not limited to the embodiments described above, and various modifications are possible within the scope of the claims. Embodiments obtained by appropriately combining the technical means disclosed in different embodiments are also included in the technical scope of this disclosure. [Explanation of Symbols]
[0089] 1. 1A Authentication System 2. Communication terminal (object identification device) 2A Authentication device (object identification device) 3. 3A Server Equipment 4A, 4B Camera 5A Access Control System 10 Control Unit 11 Storage section 12 Camera section 13 Communications Department 14 Control section 15 Display section 20 Acquisition Department 21 Feature Estimation Unit 22 Selection Section 23 Registration Department 24 Authentication Unit (Identification Unit) 30. First Estimated Model 31. Second Estimated Model 32 Third Estimated Model 33. Registered Data 40 Color Feature Estimation Unit 41 Depth Feature Estimation Unit 42 Three-dimensional feature estimation unit 51, 53, 56 normalization layer 50, 52 CNN layers 54 FC layer 55 Bias Layer 60 Learning Department 61 1st FC layer 62 Batch Center Layer 63 Normalization layer 64 2nd FC layer 65 loss layer
Claims
1. An object identification device that identifies objects contained in a two-dimensional color image and a two-dimensional depth image, An acquisition unit that acquires the data of the color image and the data of the depth image, A color feature estimation unit estimates color feature vectors that represent the characteristics of the object included in the color image using a first estimation model from the data of the color image, A depth feature estimation unit estimates depth feature vectors that represent the characteristics of the object included in the depth image using a second estimation model from the data of the depth image, A three-dimensional feature estimation unit estimates a three-dimensional feature vector representing the features of the object included in a three-dimensional image including the color image and the depth image, using a third estimation model based on the color feature vector and the depth feature vector. A selection unit that selects at least one image from the color image and depth image that can be used to identify the object, based on the data of the color image and the data of the depth image, An object identification device comprising: an identification unit that identifies an object using at least one feature vector from among the color feature vector, the depth feature vector, and the three-dimensional feature vector that corresponds to an image selected by the selection unit.
2. The object identification device according to claim 1, further comprising a storage unit for storing the first estimation model, the second estimation model, and the third estimation model.
3. The third estimation model described above is: A fully connected layer that reduces the dimensionality of a vector obtained by combining the color feature vector and the depth feature vector, A bias layer that corrects the bias from the center of the dimensionality-reduced vector, The object identification device according to claim 1, comprising a normalization layer that generates the three-dimensional feature vector by normalizing the corrected vector.
4. The object identification device according to claim 1, wherein the color image is an image of the region of the object extracted from the original color image by object detection, and the depth image is an image of the region corresponding to the region of the object extracted from the original depth image.
5. The object identification device according to claim 1, wherein the object is a human face.
6. An object identification program for causing a computer to function as an object identification device according to claim 1, wherein the computer functions as a color feature estimation unit, a depth feature estimation unit, a three-dimensional feature estimation unit, a selection unit, and an identification unit.
7. An object identification method for identifying objects contained in a two-dimensional color image and a two-dimensional depth image, An acquisition step of acquiring the data of the color image and the data of the depth image, A color feature estimation step in which a color feature vector indicating the features of the object included in the color image is estimated from the data of the color image using a first estimation model, A depth feature estimation step in which a depth feature vector indicating the features of the object included in the depth image is estimated from the depth image data using a second estimation model, A three-dimensional feature estimation step in which a third estimation model is used to estimate a three-dimensional feature vector that indicates the features of the object included in the three-dimensional image including the color image and the depth image, from the color feature vector and the depth feature vector, A selection step of selecting at least one image from the color image and depth image that can be used to identify the object, based on the data of the color image and the data of the depth image. An object identification method comprising: an identification step of identifying an object using at least one feature vector from among the color feature vector, the depth feature vector, and the three-dimensional feature vector that corresponds to the image selected in the selection step.