Information processing method, information processing device, and information processing program
Patent Information
- Application Number
- JP2023556217
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2022-04-05
- Filing Date
- 2022-09-27
- Publication Date
- 2026-09-30
- Estimated Expiration
- 2042-09-27
Smart Images

Figure 0007914730000001 
Figure 0007914730000002 
Figure 0007914730000003
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to a technology for recognizing a target person. [Background Art]
[0002] Non-Patent Document 1 discloses a technology for recognizing a target person by comparing face images with each other and comparing audio data with each other between a registered person and the target person.
[0003] However, in Non-Patent Document 1, although the comparison result between face images has high accuracy, in a case where the comparison result between audio data has low accuracy, it is not considered that the comparison result between face images is affected by the comparison result between audio data, which rather reduces recognition accuracy, so there is a need for further improvement. [Prior Art Literature] [Non-Patent Literature]
[0004] [Non-Patent Document 1] Jesus Villalba, Daniel Garcia-Romero, Nanxin Chen, Gregory Sell, Jonas Borgstrom, Alan McCree, L. Paola Garcia-Perera1, Saurabh Kataria, Phani Sankar Nidadavolu,Pedro A. Torres-Carrasquillo, Najim Dehak , “Advances in Speaker Recognition for Telephone and Audio-Visual Data: the JHU-MIT Submission for NIST SRE19” , Odyssey 2020 The Speaker and Language Recognition Workshop1-5 November 2020, Tokyo, Japan [Summary of the Invention]
[0005] This disclosure aims to solve these problems and provides a technology that can recognize a target person with high accuracy, regardless of the accuracy of the voice data, when using voice data and facial images to recognize a target person.
[0006] An information processing method in one aspect of the present disclosure is an information processing method in a computer, comprising: obtaining a face similarity score indicating the similarity between the face of a first person and the face of a second person; obtaining a voice similarity score indicating the similarity between the voice of the first person and the voice of a second person; if the face similarity score falls within an integration range that includes a threshold used when determining whether the first person and the second person are the same person, calculating an integrated similarity score by integrating the face similarity score and the voice similarity score, determining the integrated similarity score as the final similarity score; and if the face similarity score does not fall within the integration range, calculating the face similarity score as the final similarity score and outputting the final similarity score.
[0007] According to this disclosure, when recognizing a target person using voice data and facial images, the target person can be recognized with high accuracy regardless of the accuracy of the voice data. [Brief explanation of the drawing]
[0008] [Figure 1] This is a block diagram showing an example of the overall configuration of the information processing system in Embodiment 1 of this disclosure. [Figure 2] A flowchart illustrating an example of processing in the information processing device according to Embodiment 1. [Figure 3] This graph shows the relationship between weight coefficients and facial similarity. [Figure 4] This is a diagram to explain the problems of the comparative example. [Figure 5] This is a diagram illustrating the effects of the information processing device in Embodiment 1. [Figure 6] This table summarizes the experimental results of the information processing device in Embodiment 1 and the comparative example. [Figure 7]This figure shows an example of the overall configuration of the information processing system in Embodiment 2 of the present disclosure. [Figure 8] This flowchart shows an example of the process by which the information processing device in Embodiment 2 of this disclosure determines the integration scope. [Figure 9] This diagram illustrates the first method for determining the scope of integration. [Figure 10] This diagram illustrates the second method for determining the scope of integration. [Figure 11] This diagram illustrates the third method for determining the scope of integration. [Figure 12] This figure shows an example of the overall configuration of the information processing system in Embodiment 3 of this disclosure. [Modes for carrying out the invention]
[0009] (Knowledge forming the basis of this disclosure) In recent years, in order to further improve the accuracy of recognizing target individuals, multimodal recognition technology that uses facial images in addition to voice data to recognize target individuals has been investigated (for example, Non-Patent Document 1). In such multimodal recognition technology, an integrated similarity score is calculated by integrating voice similarity, which is the similarity between the voice data of the target individual and the voice data of a registered individual, and facial similarity, which is the similarity between the facial image of the target individual and the facial image of a registered individual. The calculated integrated similarity score is then compared with a threshold to determine whether or not the target individual is a registered individual.
[0010] However, even if the acquired facial image of the target person is highly accurate, if the acquired audio data of the target person is of low accuracy due to the influence of noise or other factors, the high facial similarity value may be affected by the low audio similarity value, causing the combined similarity to fall below the threshold and potentially leading to the misidentification that the target person is not a registered person. Therefore, in such cases, using audio similarity in addition to facial similarity can actually lead to a decrease in the accuracy of target person recognition.
[0011] This disclosure was made to address these issues.
[0012] (1) An information processing method in one aspect of the present disclosure is an information processing method in a computer, which obtains a face similarity score indicating the similarity between the face of a first person and the face of a second person, obtains a voice similarity score indicating the similarity between the voice of the first person and the voice of a second person, calculates an integrated similarity score by integrating the face similarity score and the voice similarity score if the face similarity score falls within an integration range that includes a threshold used when determining whether the first person and the second person are the same person, determines the integrated similarity score as the final similarity score, and if the face similarity score does not fall within the integration range, calculates the face similarity score as the final similarity score and outputs the final similarity score.
[0013] In this configuration, if the facial similarity is within the integration range that includes the threshold used to determine whether the first person is the same person as the second person, the integrated similarity is calculated by integrating the facial similarity and voice similarity, and this integrated similarity is determined as the final similarity. On the other hand, if the facial similarity is not within the integration range, the facial similarity is determined as the final similarity. Thus, in this configuration, if the facial similarity is near the threshold and determination is difficult using facial similarity alone, it is possible to recognize the target person using the integrated similarity, which is the result of integrating facial similarity and voice similarity. On the other hand, if the facial similarity is not near the threshold and determination is easy using facial similarity alone, it is possible to recognize the target person using only facial similarity. As a result, the target person can be recognized with high accuracy regardless of the accuracy of the voice data.
[0014] (2) In the information processing method described in (1) above, distribution information is further obtained that includes a first distribution showing the relationship between the facial similarity and the frequency of facial similarity for the same person, and a second distribution showing the relationship between the facial similarity and the frequency of facial similarity for different people, and the integrated range is calculated based on the first and second distributions.
[0015] According to this configuration, the integration range is calculated based on distribution information including a first distribution indicating the relationship between face similarity and the frequency of face similarity for the same person, and a second distribution indicating the relationship between face similarity and the frequency of face similarity for different persons, so the integration range can be determined with high accuracy.
[0016] (3) In the information processing method according to (2) above, the integration range may be determined based on a width from the minimum value of the face similarity in the first distribution to the maximum value of the face similarity in the second distribution.
[0017] According to this configuration, the integration range is determined based on the width from the minimum value of the face similarity in the first distribution to the maximum value of the face similarity in the second distribution, so the integration range can be determined with high accuracy.
[0018] (4) In the information processing method according to any one of (2) or (3) above, the integration range may be determined based on a first width from the threshold value to the minimum value of the face similarity in the first distribution, and a second width from the threshold value to the maximum value of the face similarity in the second distribution.
[0019] According to this configuration, the integration range is determined based on the first width from the threshold value to the minimum value of the face similarity in the first distribution, and the second width from the threshold value to the maximum value of the face similarity in the second distribution, so the integration range can be determined with high accuracy.
[0020] (5) In the information processing method according to any one of (2) to (4) above, the integration range has a width centered on the threshold value, and the width may be determined based on a third width from the minimum value of the face similarity to the maximum value of the face similarity in the first distribution and the second distribution, and the accuracy of the distribution information.
[0021] According to this configuration, the integration range has a width centered on a threshold, and this width is determined based on a third width from the minimum value to the maximum value of the face similarity across the entire range of the first and second distributions, and the accuracy of the distribution information, so that the integration range can be determined with high accuracy.
[0022] (6) In the information processing method described in any one of (1) to (5) above, the integrated similarity may have a value obtained by weighting the face similarity and the voice similarity by a predetermined weight coefficient and averaging them.
[0023] With this configuration, the integrated similarity score is obtained by weighting the facial similarity score and the voice similarity score by a predetermined weight coefficient and averaging them, thus allowing for the appropriate integration of facial similarity and voice similarity scores.
[0024] (7) In the information processing method described in (6) above, the predetermined weight coefficient may be a fixed value.
[0025] With this configuration, the integrated similarity score is calculated using fixed weight coefficients, making it easy to calculate the integrated similarity score.
[0026] (8) In the information processing method described in (6) above, the predetermined weight coefficient may be set to a value such that the ratio of the voice similarity to the face similarity increases as the face similarity approaches the center of the integration range.
[0027] With this configuration, the integrated similarity is calculated using a weighting coefficient whose value decreases as the face similarity approaches the center of the integration range, thus enabling accurate calculation of integrated similarity.
[0028] (9) In the information processing method described in any one of (1) to (8) above, the integrated similarity may be the sum of the value obtained by multiplying the voice similarity by a weighting coefficient and the face similarity.
[0029] With this configuration, the voice similarity, whose value has been adjusted by a weighting coefficient, is added to the face similarity to calculate the combined similarity. Therefore, the combined similarity can be accurately calculated based on the face similarity.
[0030] (10) In the information processing method described in any one of (1) to (9) above, the following steps may be taken: a face image of the first person is acquired; the facial feature quantities of the first person are calculated from the face image of the first person; the facial feature quantities of the second person are acquired; and the facial similarity is calculated from the facial feature quantities of the first person and the facial feature quantities of the second person, and in the acquisition of the facial similarity, the calculated facial similarity may be acquired.
[0031] With this configuration, if the first person is the target person and the second person is the registered person, it is possible to calculate whether or not the first person is the registered person.
[0032] (11) In the information processing method described in any of (1) to (10) above, voice data of the first person is acquired, the voice features of the first person are calculated from the voice data of the first person, the voice features of the second person are acquired, and the voice similarity is calculated from the voice features of the first person and the voice features of the second person, and in acquiring the voice similarity, the calculated voice similarity is acquired.
[0033] With this configuration, if the first person is the target person and the second person is the registered person, it is possible to determine whether or not the first person is the registered person.
[0034] (12) In the information processing method described in any of (1) to (11) above, if the final similarity exceeds the threshold, it is determined that the first person and the second person are the same person; if the final similarity is less than the threshold, it is determined that the first person and the second person are different people; and the determination result of whether or not the first person and the second person are the same person may be output.
[0035] This configuration allows us to determine whether the first person and the second person are the same person using the final similarity score.
[0036] (13) An information processing device in another aspect of the present disclosure includes: a first acquisition unit that acquires a face similarity score indicating the similarity between the face of a first person and the face of a second person; a second acquisition unit that acquires a voice similarity score indicating the similarity between the voice of a first person and the voice of a second person calculated based on the voice features of the first person and the voice features of the second person; an integration unit that, if the face similarity score is within the integration range, calculates an integrated similarity score by integrating the face similarity score and the voice similarity score and determines the integrated similarity score as the final similarity score; and an output unit that outputs the final similarity score.
[0037] This configuration provides an information processing device that can recognize a target person with high accuracy, regardless of the accuracy of the audio data.
[0038] (14) An information processing program in another aspect of the present disclosure causes a computer to perform the following processes: obtain a face similarity score indicating the similarity between the face of a first person and the face of a second person; obtain a voice similarity score indicating the similarity between the voice of a first person and the voice of a second person calculated based on the voice features of the first person and the voice features of the second person; if the face similarity score is within the integration range, calculate an integrated similarity score by integrating the face similarity score and the voice similarity score, determine the integrated similarity score as the final similarity score; if the face similarity score is not within the integration range, calculate the face similarity score as the final similarity score, and output the final similarity score.
[0039] This configuration allows for the provision of an information processing program that can accurately recognize a target person regardless of the accuracy of the audio data.
[0040] This disclosure can also be implemented as an information processing system operated by such an information processing program. Furthermore, it goes without saying that such a computer program can be distributed via computer-readable, non-temporary recording media such as CD-ROMs or via communication networks such as the Internet.
[0041] The embodiments described below are all specific examples of this disclosure. The numerical values, shapes, components, steps, and order of steps shown in the following embodiments are examples only and are not intended to limit this disclosure. Furthermore, among the components in the following embodiments, those not described in the independent claim representing the highest-level concept will be described as optional components. In addition, the contents of each embodiment can be combined.
[0042] (Embodiment 1) Figure 1 is a block diagram showing an example of the overall configuration of the information processing system 100 in Embodiment 1 of this disclosure. The information processing system 100 is a system that uses voice data and facial images to determine whether a target person to be recognized is the same person as a pre-registered registered person. The target person is an example of a first person, and the registered person is an example of a second person. The information processing system 100 is applied, for example, to an access control system that manages entry and exit of a target person to a management area. The management area is, for example, a building, residence, office, etc. However, the application examples of the information processing system 100 are not limited to this, and it may also be applied to a personal authentication system that performs personal authentication using voice data and facial images.
[0043] The information processing system 100 includes an information processing device 1, a camera 2, a microphone 3, and a display 4. The information processing device 1 is a computer. The information processing device 1 includes a processor 11 and memory 12. The processor 11 is, for example, a CPU (Central Processing Unit). The memory 12 is, for example, a non-volatile rewritable storage device such as flash memory.
[0044] The processor 11 includes a face feature calculation unit 111, a face similarity calculation unit 112, a first acquisition unit 113, a voice feature calculation unit 114, a voice similarity calculation unit 115, a second acquisition unit 116, an integration unit 117, a recognition unit 118, and an output unit 119. The face feature calculation unit 111 to the output unit 119 may be implemented, for example, by the processor 11 executing an information processing program, or they may be composed of dedicated hardware circuits such as ASICs.
[0045] The face feature calculation unit 111 acquires a face image of the target person captured by the camera 2 and calculates face features, which are the facial features of the target person, from the acquired face image. A face image is an image that includes the face of the target person. A face image is digital image data in which pixel data is arranged in a predetermined row × predetermined column. A face image may be a monochrome image or a color image having three color components: R, G, and B. Face features are, for example, vectors that represent the features of the face.
[0046] The face feature calculation unit 111 can calculate face features by inputting face images into the face recognition model. The face recognition model is, for example, a pre-trained model created by machine learning on a large number of datasets in which face images are used as explanatory variables and face features are used as the target variable.
[0047] The face recognition model is pre-stored in memory 12, for example. However, this is just one example; the face feature calculation unit 111 may calculate the face features by sending a face image to an external server that stores the face recognition model and obtaining the face features sent back as a response from the external server.
[0048] The face similarity calculation unit 112 obtains the facial features of the registered person and calculates face similarity, which is the similarity between the obtained facial features of the registered person and the facial features of the target person. Since the facial features of the registered person are stored in memory 12 beforehand, the face similarity calculation unit 112 only needs to obtain the facial features of the registered person from memory 12. The facial features of the registered person are calculated in advance by inputting the facial image of the registered person into the face recognition model. Therefore, the facial features of the registered person have the same number of dimensions as the facial features of the target person.
[0049] Memory 12 may store multiple registered individuals' facial features in association with multiple individual IDs. In this case, the facial similarity calculation unit 112 only needs to calculate the facial similarity between the facial features of the registered individuals corresponding to the individual ID entered by the target individual and the facial features of the target individual. In this case, the target individual can input the individual ID via the operating device (not shown in the diagram).
[0050] Any index capable of evaluating the similarity between vectors may be used for facial similarity. In this embodiment, the facial similarity is set to be larger the closer the facial similarity of the registered person is to the facial similarity of the target person. The facial similarity has a normalized value within a predetermined range (e.g., 0 to 100, 0 to 200, etc.). For example, the facial similarity is calculated by normalizing the Euclidean distance or cosine similarity within a predetermined range such that the value increases as the similarity between the facial similarity of the target person and the facial similarity of the registered person increases.
[0051] The first acquisition unit 113 acquires the face similarity calculated by the face similarity calculation unit 112.
[0052] The voice feature calculation unit 114 acquires the voice data of the target person picked up by the microphone 3 and calculates voice features, which are the characteristics of the target person's voice, from the acquired voice data. The voice data is, for example, digital voice data obtained by A / D conversion of the analog voice data of the target person picked up by the microphone 3. Voice features are vectors that represent the characteristics of the voice. Examples of voice features include the x vector and the i vector.
[0053] The speech feature calculation unit 114 can calculate speech features by inputting speech data into the speech recognition model. The speech recognition model is, for example, a pre-trained model created by machine learning on a large number of datasets in which speech data is the explanatory variable and speech features are the target variable.
[0054] The speech recognition model is pre-stored in memory 12. However, this is just one example; the speech feature calculation unit 114 may calculate the speech features by sending speech data to an external server that stores the speech recognition model and obtaining the speech features sent back as a response from the external server.
[0055] The voice similarity calculation unit 115 acquires the voice features of the registered person and calculates voice similarity, which is the similarity between the acquired voice features of the registered person and the voice features of the target person. Since the voice features of the registered person are stored in memory 12 beforehand, the voice similarity calculation unit 115 only needs to acquire the voice features of the registered person from memory 12. The voice features of the registered person are calculated in advance by inputting the registered person's voice data into the voice recognition model. Therefore, the voice features of the registered person have the same number of dimensions as the voice features of the target person.
[0056] Memory 12 may store multiple registered individuals' voice features in association with multiple individual IDs. In this case, the voice similarity calculation unit 115 only needs to calculate the voice similarity between the voice features of the registered individuals corresponding to the individual IDs entered by the target individual via the operating device and the voice features of the target individual.
[0057] The second acquisition unit 116 acquires the voice similarity calculated by the voice similarity calculation unit 115 and inputs the voice similarity to the integration unit 117.
[0058] The integration unit 117 calculates an integrated similarity by integrating the face similarity and voice similarity if the face similarity acquired by the first acquisition unit 113 falls within the integration range, and determines the integrated similarity as the final similarity. On the other hand, if the face similarity acquired by the first acquisition unit 113 does not fall within the integration range, the face similarity acquired by the first acquisition unit 113 is determined as the final similarity. The integration range is a range that includes the threshold T1 used to determine whether the target person is the same person as the registered person, and is stored in memory 12 in advance. The calculation method for the integrated similarity will be described later.
[0059] The recognition unit 118 compares the final similarity calculated by the integration unit 117 with a threshold T1 to determine whether the target person is the same person as the registered person, that is, whether the target person is the same person or someone else. The threshold T1 is pre-stored in the memory 12. For example, if the final similarity is greater than the threshold T1, the recognition unit 118 determines that the target person is the same person as the registered person. On the other hand, if the final similarity is less than or equal to the threshold T1, the recognition unit 118 determines that the target person is a different person from the registered person.
[0060] The output unit 119 generates judgment result information indicating the judgment result by the recognition unit 118, and outputs the generated judgment result information to the display 4.
[0061] Memory 12 stores the face recognition model, voice recognition model, integration range, and threshold T1.
[0062] Camera 2 is a camera installed, for example, at the entrance / exit of a controlled area. When a person is detected by a motion sensor (not shown) attempting to enter the controlled area, Camera 2 captures a facial image of the person. Alternatively, if the person enters a person ID via an operating device (not shown), Camera 2 captures a facial image of the person. Camera 2 inputs the captured facial image to Processor 11. If a person ID is entered, Camera 2 only needs to associate the facial image with the person ID and input it to Processor 11.
[0063] Microphone 3 is a sound-collecting device installed, for example, at the entrance / exit of a management area. When a person is detected by a motion sensor (not shown) attempting to enter the management area, microphone 3 collects the voice data of that person. Alternatively, when the person enters their person ID via an operating device (not shown), microphone 3 collects the voice data of that person. Microphone 3 inputs the collected voice data to processor 11.
[0064] Display 4 is a display device installed, for example, at the entrance / exit of a management area. Display 4 displays the judgment result information output by the output unit 119. If the recognition unit 118 determines that the person in question is the same person as the registered person, Display 4 displays first judgment result information indicating that the person is the registered person. On the other hand, if the recognition unit 118 determines that the person in question is a different person from the registered person, Display 4 displays second judgment result information indicating that the person is a different person. The first judgment result information may indicate that the person in question is permitted to enter the management area. The second judgment result information may indicate that the person in question is denied entry to the management area.
[0065] Next, we will explain the processing of the information processing device 1. Figure 2 is a flowchart showing an example of the processing of the information processing device 1 in Embodiment 1.
[0066] (Step S1) The facial feature calculation unit 111 acquires a facial image of the target person from the camera 2.
[0067] (Step S2) The facial feature calculation unit 111 calculates the facial features of the target person by inputting the facial image into the facial recognition model.
[0068] (Step S3) The face similarity calculation unit 112 obtains the facial features of the registered person from the memory 12.
[0069] (Step S4) The face similarity calculation unit 112 calculates face similarity, which is the similarity between the face features of the target person calculated by the face feature calculation unit 111 and the face features of the registered person. The first acquisition unit 113 acquires the face similarity calculated by the face similarity calculation unit 112 and inputs the acquired face similarity to the integration unit 117.
[0070] (Step S5) The voice feature calculation unit 114 acquires voice data from the microphone 3.
[0071] (Step S6) The voice feature calculation unit 114 calculates the voice features of the target person by inputting the voice data into the voice recognition model.
[0072] (Step S7) The voice similarity calculation unit 115 obtains the voice features of the registered person from the memory 12.
[0073] (Step S8) The voice similarity calculation unit 115 calculates voice similarity, which is the similarity between the voice features of the target person calculated by the voice feature calculation unit 114 and the voice features of the registered person. The second acquisition unit 116 acquires the voice similarity calculated by the voice similarity calculation unit 115 and inputs the acquired voice similarity to the integration unit 117.
[0074] (Step S9) The integration unit 117 determines whether the face similarity input from the first acquisition unit 113 falls within the integration range. If it is determined that the face similarity falls within the integration range (YES in step S9), the process proceeds to step S10. On the other hand, if it is determined that the face similarity falls outside the integration range (NO in step S9), the process proceeds to step S11.
[0075] (Step S10) The integration unit 117 calculates an integrated similarity score by integrating the face similarity score and the voice similarity score, and determines the integrated similarity score as the final similarity score. The integrated similarity score is calculated by, for example, the following three methods. When the processing in step S10 is completed, the process proceeds to step S12.
[0076] (1st method) The integration unit 117 calculates the integrated similarity by weighting the facial similarity and voice similarity using fixed weight coefficients and averaging them. Specifically, the integration unit 117 calculates the integrated similarity using the following equation (1).
[0077] s = α·sv + (1-α)·sf (1)
[0078] s is the combined similarity score. α is a fixed weight coefficient, between 0 and 1. sv is the speech feature. sf is the facial feature.
[0079] (Second method) The integration unit 117 calculates the integrated similarity by weighting the facial similarity and voice similarity with a variable weight coefficient and averaging them. Specifically, the integration unit 117 calculates the integrated similarity using the following equation (2).
[0080] s = α·sv + (1-α)·sf (2)
[0081] The weighting coefficient α is set to a value such that the ratio of voice similarity sv to face similarity sf increases as face similarity sf approaches the center of the integration range.
[0082] Figure 3 is graph G1, which shows the relationship between the weighting coefficient α and the face similarity score sf. In graph G1, the vertical axis represents the weighting coefficient α, and the horizontal axis represents the face similarity score sf. p is the minimum value of the integration range, and q is the maximum value of the integration range. c is the center of the integration range, and is expressed as c = (p + q) / 2.
[0083] Based on the above, the weighting coefficient α is expressed by the following equations (3) and (4).
[0084] α = (sf - p) / (cp) (sf ≤ c) (3) α = (q - sf) / (qc) (c <sf) (4)
[0085] When the face similarity score sf is less than or equal to the center c, the weighting coefficient α increases linearly as the face similarity score sf approaches the center c, as shown in equation (3). On the other hand, when the face similarity score sf is greater than the center c, the weighting coefficient α decreases linearly as the face similarity score sf moves away from the center c, as shown in equation (4). When the face similarity score sf is at the center c, the weighting coefficient α is 1, as shown in equation (3) or equation (4).
[0086] Thus, in the second method, the weight coefficient α is set to approach 1 as the face similarity sf approaches the center c. Therefore, the face similarity sf and the voice similarity sv are weighted and averaged using a weight coefficient that changes linearly, such that the proportion of voice similarity sv to face similarity sf increases as the face similarity sf approaches the center c. On the other hand, the weight coefficient α is set to approach 0 as the face similarity sf moves away from the center c. Therefore, the face similarity sf and the voice similarity sv are weighted and averaged using a weight coefficient that changes linearly, such that the proportion of voice similarity sv to face similarity sf decreases as the face similarity sf approaches the minimum value p or maximum value q from the center c.
[0087] (3rd method) The integration unit 117 calculates the integrated similarity by adding the value obtained by multiplying the voice similarity sv by the weight coefficient α and the face similarity sf. Specifically, the integration unit 117 calculates the integrated similarity using the following formula (5).
[0088] s = α·sv + sf (5)
[0089] α is a fixed weighting coefficient, between 0 and 1. Thus, in this third method, the combined similarity s is calculated by adding the voice similarity sv, which is weighted by the weighting coefficient α, to the face similarity sf. Therefore, the combined similarity can be accurately calculated while using face similarity as the basis.
[0090] (Step S11) The integration unit 117 determines the face similarity calculated by the face similarity calculation unit 112 as the final similarity.
[0091] (Step S12) The recognition unit 118 determines whether the final similarity is greater than the threshold T1. If the final similarity is greater than the threshold T1 (YES in step S12), the process proceeds to step S13. On the other hand, if the final similarity is less than or equal to the threshold T1 (NO in step S12), the process proceeds to step S14.
[0092] (Step S13) The recognition unit 118 determines that the target person is the same person as the registered person, that is, the person in question.
[0093] (Step S14) The recognition unit 118 determines that the target person is a different person from the registered person, that is, a different person.
[0094] (Step S15) The output unit 119 generates judgment result information indicating the judgment result by the recognition unit 118 and outputs the judgment result information to the display 4. As a result, the display 4 displays either first judgment result information indicating that the target person has been determined to be the person in question, or second judgment result information indicating that the target person has been determined to be someone else. As a result, the target person can be notified of the judgment result.
[0095] Furthermore, if the information processing device 1 determines that the person in question is indeed the person in question, it may send a control signal to the automatic door at the entrance / exit of the management area to open the automatic door. On the other hand, if the information processing device 1 determines that the person in question is not the person in question, it may refrain from sending a control signal to the automatic door to open the automatic door.
[0096] Next, the effects of the information processing device 1 will be explained in comparison with the comparative example. Figure 4 is a diagram illustrating the problems of the comparative example. In the distribution information D1 shown in Figure 4, the vertical axis represents frequency and the horizontal axis represents face similarity sf. Distribution information D1 includes a first distribution D101 and a second distribution D102. The first distribution D101 is a hypothetical distribution of face similarity sf that is expected to be obtained when a large number of trials are conducted comparing the face features of the target person and the registered person, when the target person is the same person as the registered person. The second distribution D102 is a hypothetical distribution of face similarity sf that is expected to be obtained when a large number of trials are conducted comparing the face features of the target person and the registered person, when the target person is a different person from the registered person. The first distribution D101 is distributed on the side of face similarity sf that is higher than that of the second distribution D102. A part of the leftmost region of the first distribution D101 overlaps with a part of the rightmost region of the second distribution D102. In the comparative example, the face similarity value sf (=70) at the center of this overlapping region is used as the threshold T1.
[0097] In the comparative example, the combined similarity s is compared to the threshold T1 (=70) regardless of whether the face similarity sf is within the combined range or not. In the comparative example, the combined similarity s is calculated as s = (sf + sv) / 2.
[0098] Here, let's consider the case where the facial similarity score sf is 100 and the voice similarity score sv is 20. In this case, the facial similarity score sf is 100, which is significantly larger than the threshold T1 (=70), so there is a high probability that the subject is indeed the person in question.
[0099] However, in the comparative example, the combined similarity score s is calculated to be 60 (=(100+20) / 2), and since the combined similarity score s falls below the threshold T1 (=70), the subject is judged not to be the person in question. Thus, in the comparative example, since the determination of whether or not the subject is the person in question is made using only the combined similarity score s, there is a possibility of misjudgment if a low voice similarity score sv is obtained, even in cases where judgment using face similarity score sf would be easy. Such a low voice similarity score sv occurs when there is a lot of noise in the environment surrounding microphone 3, or when the subject speaks in a direction other than microphone 3. In this case, using voice similarity score sv actually reduces the recognition accuracy.
[0100] Therefore, the information processing device 1 calculates the integrated similarity when the face similarity sf is within the integration range and it is difficult to determine whether or not the target person is the person in question based solely on the face similarity sf.
[0101] Figure 5 is a diagram illustrating the effect of the information processing device 1 in Embodiment 1. The distribution information D1 shown in Figure 5 is the same as in Figure 4. In the example in Figure 5, the integration range W1 has a face similarity sf value in the range of 60 or more and 80 or less. Now, let's consider the case where the face similarity sf is 100 and the voice similarity sv is 20. In this case, in Embodiment 1, since the face similarity sf is 100 and is not within the integration range W1, the face similarity sf is determined as the final similarity. Therefore, the final similarity exceeds the threshold T1, and the target person is determined to be the person in question.
[0102] On the other hand, in this embodiment, if the face similarity sf is within the integration range W1 and it is difficult to make a judgment based solely on the face similarity sf, the integrated similarity s is calculated as the final similarity. Therefore, Embodiment 1 can improve the accuracy of determining whether or not the target person is the person in question.
[0103] Figure 6 is a table summarizing the experimental results of the information processing device 1 in Embodiment 1 and the comparative example. EER (%) is an error rate evaluation scale commonly used in speaker recognition, with a smaller value indicating higher performance. minC is a cost defined by NIST (National Institute of Standards and Technology), with a smaller value indicating higher performance.
[0104] As shown in Figure 6, the EER (%) was 0.406 in the comparative example, compared to 0.381 in Embodiment 1. Similarly, the minC was 0.021 in the comparative example, compared to 0.012 in Embodiment 1. Therefore, it was confirmed that the method of Embodiment 1 showed higher performance in both EER (%) and minC compared to the method of the comparative example.
[0105] In this embodiment 1, if the facial similarity is near the threshold and determination is difficult using facial similarity alone, it becomes possible to recognize the target person using an integrated similarity that combines facial similarity and voice similarity. On the other hand, if the facial similarity is not near the threshold and determination is easy using facial similarity alone, it becomes possible to recognize the target person using only facial similarity. As a result, the target person can be recognized with high accuracy regardless of the accuracy of the voice data.
[0106] (Embodiment 2) Embodiment 2 calculates the integration range based on distribution information. Figure 7 shows an example of the overall configuration of the information processing system 100 in Embodiment 2 of this disclosure. In Figure 7, the difference from Figure 1 is that the processor 11A of the information processing device 1A further has an integration range determination unit 120. In Embodiment 2, the same reference numerals are used for components that are the same as in Embodiment 1, and their descriptions are omitted.
[0107] The integration range determination unit 120 acquires distribution information including a first distribution showing the relationship between face similarity and the frequency of face similarity for the same person, and a second distribution showing the relationship between face similarity and the frequency of face similarity for different people. Based on the first and second distributions, the integration range determination unit 120 calculates the integration range and stores the calculated integration range in the memory 12.
[0108] Figure 8 is a flowchart showing an example of the process by which the information processing device 1A determines the integration range in Embodiment 2 of this disclosure.
[0109] (Step S30) The integration range determination unit 120 acquires training data for determining the integration range. Here, the integration range determination unit 120 can acquire the training data from an external terminal (not shown in the figure). The external terminal is, for example, a desktop computer.
[0110] The training data includes first training data and second training data. The first training data includes a large number of face similarity scores obtained by performing a large number of trials comparing the face features of the target person and the registered person when the target person and the registered person are the same person. In this trial, the target person may be multiple people or a single person. The second training data includes a large number of face similarity scores obtained by performing a large number of trials comparing the face features of the target person and the registered person when the target person and the registered person are different people.
[0111] (Step S31) The integrated range determination unit 120 calculates distribution information from the acquired training data. This allows the integrated range determination unit 120 to obtain distribution information. Here, the integrated range determination unit 120 calculates the first distribution by classifying the facial features contained in the first training data into multiple classes and determining the frequency of the facial features in each class. Furthermore, the integrated range determination unit 120 calculates the second distribution by classifying the facial features contained in the second training data into multiple classes and determining the frequency of the facial features in each class. This completes the calculation of distribution information.
[0112] (Step S32) The integration range determination unit 120 determines the integration range based on the first distribution and the second distribution. The integration range is determined using the following three methods.
[0113] (1st determination method) Figure 9 illustrates the first method for determining the integration range W1. The distribution information D10 shown in Figure 9 includes the first distribution D11 and the second distribution D12. In the distribution information D10, the vertical axis represents frequency and the horizontal axis represents face similarity sf. The first distribution D11 is distributed on the side where face similarity sf is higher than that of the second distribution D12. A portion of the leftmost region of the first distribution D11 overlaps with a portion of the rightmost region of the second distribution D12. The threshold T1 is, for example, the face similarity sf value at the center of this overlapping region.
[0114] The integration range determination unit 120 determines the integration range W1 based on the width W2 from the minimum value A1 of the face similarity sf in the first distribution D11 to the maximum value A2 of the face similarity sf in the second distribution D12.
[0115] Specifically, the integration range determination unit 120 calculates the length of the integration range W1 by multiplying the width W2 by a predetermined coefficient (for example, 1.1) to allow for a margin in the width W2. The integration range determination unit 120 also determines the position of the integration range W1 such that its center lies at the center of the width W2. Note that 1.1 is just an example of the coefficient, and other appropriate values such as 1.05 or 1.15 may be used.
[0116] (Second determination method) Figure 10 illustrates the second method for determining the integration range W1. The integration range determination unit 120 determines the integration range W1 based on a first width W21 from the threshold T1 to the minimum value A1 of the face similarity sf in the first distribution D11, and a second width W22 from the threshold T1 to the maximum value A2 of the face similarity sf in the second distribution D12.
[0117] Specifically, the integrated range determination unit 120 calculates the first width W31 by multiplying the first width W21 by a predetermined coefficient (e.g., 1.1) to provide a margin, and calculates the second width W32 by multiplying the second width W22 by a predetermined coefficient (e.g., 1.1) to provide a margin. Then, the integrated range determination unit 120 calculates the integrated range W1 by connecting the first width W31 and the second width W32. Note that 1.1 is just an example of the coefficient, and other appropriate values such as 1.05 or 1.15 may be used.
[0118] (3rd determination method) Figure 11 illustrates the third method for determining the integration range W1. The integration range determination unit 120 determines the width of the integration range W1 based on the third width W3, which ranges from the minimum value B1 of the face similarity sf to the maximum value B2 of the face similarity sf in the first distribution D11 and the second distribution D12, and the accuracy of the distribution information.
[0119] The accuracy of the distribution information D10 is, for example, the average of the accuracy rates of the first distribution D11 and the second distribution D12. The accuracy rate of the first distribution D11 is, for example, the ratio of the number of trials in the first distribution D11 that are above the threshold T1 to the total number of trials in the first distribution D11. The accuracy rate of the second distribution D12 is, for example, the ratio of the number of trials in the second distribution D12 that are below the threshold T1 to the total number of trials in the second distribution D12. Note that the accuracy rate of the first distribution D11 may also be, for example, the ratio of the area of the region in the first distribution D11 that is above the threshold T1 to the total area of the first distribution. The accuracy rate of the second distribution D12 may also be, for example, the ratio of the area of the region in the second distribution D12 that is below the threshold T1 to the total area of the second distribution D12.
[0120] The accuracy of the distribution information D10 may be, for example, the average of the error rates of the first distribution D11 and the second distribution D12. The error rate of the first distribution D11 is, for example, the ratio of the number of trials in the first distribution D11 that are below the threshold T1 to the total number of trials in the first distribution D11. The error rate of the second distribution D12 is, for example, the ratio of the number of trials in the second distribution D12 that are above the threshold T1 to the total number of trials in the second distribution D12. The error rate of the first distribution D11 may also be, for example, the ratio of the area of the region in the first distribution D11 that is below the threshold T1 to the total area of the first distribution D11. The error rate of the second distribution D12 may also be, for example, the ratio of the area of the region in the second distribution D12 that is above the threshold T1 to the total area of the second distribution D12.
[0121] The integration range determination unit 120 should determine the width of the integration range W1 by reducing the width W3 as the accuracy of the distribution information D10 increases. Then, the integration range determination unit 120 should determine the position of the integration range W1 such that its center lies at the threshold T1.
[0122] The integration unit 117 can then determine whether or not to calculate the integrated similarity by comparing the integrated range W1 determined in this way with the face similarity sf.
[0123] Thus, according to Embodiment 2, since the integration range is determined based on distribution information calculated based on actual cases, the integration range can be determined with high accuracy.
[0124] (Embodiment 3) Embodiment 3 applies the information processing system 100 of Embodiment 1 to a network. Figure 12 shows an example of the overall configuration of the information processing system 100 in Embodiment 3 of this disclosure.
[0125] The information processing system 100 comprises an information processing device 1B and a terminal 200. The information processing device 1B and the terminal 200 are connected via a network for communication. The network is, for example, a wide-area communication network such as the Internet.
[0126] The information processing device 1B is, for example, a cloud server including one or more computers, and is further equipped with a communication unit 13 in addition to the information processing device 1. The communication unit 13 is a communication device that connects the information processing device 1B to a network. The communication unit 13 receives facial images and voice data transmitted from the terminal 200. The communication unit 13 transmits judgment result information indicating the judgment result by the recognition unit 118 to the terminal 200.
[0127] Terminal 200 may be a tablet computer or a mobile device such as a smartphone, or it may be a desktop computer. Terminal 200 is equipped with a camera 2A, a microphone 3A, a display 4A, and a communication unit 5A. Camera 2A captures a facial image of the target person. Microphone 3A captures audio data of the target person. Display 4A displays the judgment result information. Communication unit 5A transmits the facial image captured by camera 2A and the audio data captured by microphone 3A to the information processing device 1B. Communication unit 5A receives the judgment result information transmitted from the information processing device 1B.
[0128] The information processing system 100 in Embodiment 3 is a system that uses a terminal 200 to determine whether or not a target person is the person in question. When the target person speaks to the terminal 200, a facial image of the target person is captured by the camera 2A, and the spoken voice data is picked up by the microphone 3A. The captured facial image and the picked-up voice data are then transmitted from the terminal 200 to the information processing device 1B. The information processing device 1B, having received the facial image and voice data, determines whether or not the target person is the person in question using the method described in Embodiment 1, and transmits determination result information indicating whether or not the target person is the person in question to the terminal 200. The terminal 200, having received the determination result information, displays the determination result information on the display 4A. This allows the determination result to be presented to the target person.
[0129] The following modifications may be adopted for this disclosure.
[0130] (1) In Embodiment 2, the integrated range determination unit 120 was described as calculating distribution information based on learning data acquired from an external terminal (not shown), but this disclosure is not limited thereto. The integrated range determination unit 120 may acquire distribution information from an external terminal (not shown).
[0131] (2) In Embodiment 3, the information processing device 1A shown in Embodiment 2 may be applied.
[0132] (3) In the information processing devices 1, 1A, and 1B, the face feature calculation unit 111, the face similarity calculation unit 112, the voice feature calculation unit 114, and the voice similarity calculation unit 115 may be provided on an external device. The external device is, for example, a terminal 200. In this case, the first acquisition unit 113 will acquire face similarity from the external device, and the second acquisition unit 116 will acquire voice similarity from the external device.
[0133] (4) In the information processing devices 1, 1A, and 1B, the recognition unit 118 may be located in an external device (not shown). In this case, the output unit 119 should transmit the final similarity calculated by the integration unit 117 to the external device. Furthermore, in this case, the recognition unit 118 of the external device should determine whether or not the person in question is the person in question by comparing the final similarity with a threshold.
[0134] (5) Camera 2 may input facial images to the information processing device 1 at predetermined intervals. Microphone 3 may input audio data to the information processing device 1 at predetermined intervals. In this case, the information processing device 1 only needs to periodically determine whether or not the person being processed is the person in question.
[0135] (6) In Figure 2, the processing set of steps S1 to S4 and the processing set of steps S5 to S8 may be executed in parallel. [Industrial applicability]
[0136] This disclosure is useful in the field of technology for identifying whether or not a person is who they claim to be.
Claims
1. A method of information processing in a computer, A facial similarity score is obtained, which indicates the degree of similarity between the face of the first person and the face of the second person, calculated based on facial feature quantities that represent the facial features of the first person as vectors, calculated from the facial image of the first person, and facial feature quantities that represent the facial features of the second person as vectors. Based on the voice features of the first person calculated from the voice data of the first person and the voice features of the second person calculated as vectors, a voice similarity score is obtained that indicates the degree of similarity between the voice of the first person and the voice of the second person. If the facial similarity falls within an integration range that includes a threshold used to determine whether the first person is the same person as the second person, the integrated similarity is calculated by integrating the facial similarity and the voice similarity, and the integrated similarity is determined as the final similarity. If the facial similarity does not fall within the integration range, the facial similarity is calculated as the final similarity. Output the final similarity score. Information processing methods.
2. Furthermore, distribution information is obtained that includes a first distribution calculated by classifying multiple facial similarity scores in multiple classes when the first person and the second person are the same person, and determining the frequency of the facial similarity scores in each class, and a second distribution calculated by classifying a large number of facial similarity scores in multiple classes when the first person and the second person are different people, and determining the frequency of the facial similarity scores in each class. The aforementioned integration range is calculated based on the first distribution and the second distribution. The information processing method according to claim 1.
3. The integration range is determined based on the width from the minimum value of the face similarity in the first distribution to the maximum value of the face similarity in the second distribution. The information processing method according to claim 2.
4. The integration range is determined based on a first width from the threshold to the minimum value of the face similarity in the first distribution and a second width from the threshold to the maximum value of the face similarity in the second distribution. The information processing method according to claim 2.
5. The aforementioned integration range has a width centered on the threshold, The width is determined based on the third width, which ranges from the minimum value to the maximum value of the face similarity in the first and second distributions, and the accuracy of the distribution information. The information processing method according to claim 2.
6. The aforementioned integrated similarity has a value obtained by weighting the facial similarity and the voice similarity by a predetermined weight coefficient and averaging them. The information processing method according to claim 1.
7. The predetermined weight coefficient is a fixed value. The information processing method according to claim 6.
8. The predetermined weighting coefficient is set to a value such that the ratio of the voice similarity to the face similarity increases as the face similarity approaches the center of the integration range. The information processing method according to claim 6.
9. The aforementioned integrated similarity is the sum of the value obtained by multiplying the voice similarity by a weighting coefficient and the face similarity. The information processing method according to claim 1.
10. Furthermore, if the final similarity exceeds the threshold, it is determined that the first person and the second person are the same person, and if the final similarity is less than the threshold, it is determined that the first person and the second person are different people. Furthermore, the system outputs the result of determining whether the first person and the second person are the same person. The information processing method according to claim 1.
11. A first acquisition unit that acquires a face similarity score indicating the similarity between the face of the first person and the face of the second person, calculated based on a face feature quantity that represents the facial features of the first person as a vector calculated from the face image of the first person and a face feature quantity that represents the facial features of the second person as a vector, A second acquisition unit that acquires a voice similarity score indicating the similarity between the voice of the first person and the voice of the second person, calculated based on a voice feature quantity that represents the characteristics of the voice of the first person calculated from the voice data of the first person as a vector and a voice feature quantity that represents the characteristics of the voice of the second person as a vector, If the face similarity is within the integration range, the integration unit calculates the integrated similarity by integrating the face similarity and the voice similarity, and determines the integrated similarity as the final similarity; if the face similarity is not within the integration range, the integration unit determines the face similarity as the final similarity. The system includes an output unit that outputs the final similarity score. Information processing device.
12. On the computer, A facial similarity score is obtained, which indicates the degree of similarity between the face of the first person and the face of the second person, calculated based on facial feature quantities that represent the facial features of the first person as vectors, calculated from the facial image of the first person, and facial feature quantities that represent the facial features of the second person as vectors. Based on the voice features of the first person calculated from the voice data of the first person and the voice features of the second person calculated as vectors, a voice similarity score is obtained that indicates the degree of similarity between the voice of the first person and the voice of the second person. If the face similarity is within the integration range, the integrated similarity is calculated by integrating the face similarity and the voice similarity, and the integrated similarity is determined as the final similarity. If the face similarity is not within the integration range, the face similarity is calculated as the final similarity. The process is executed to output the final similarity score. program.
Citation Information
Patent Citations
Personal authentication system
JP2000148985A