Identity verification system and method thereof
The identity authentication system uses voice-based speaker and semantic identification to securely authenticate users without requiring specific devices, addressing challenges with forgotten or cracked keys and enhancing convenience for all users, including those with visual impairments.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2026-04-02
AI Technical Summary
Existing identity verification methods face challenges such as the inability to obtain a one-time password when keys are forgotten, cracked, or specific devices are not available, and are inconvenient for individuals with visual impairments.
An identity authentication system and method that utilizes voice signals for authentication by sampling a first voice signal, generating authentication prompts, and performing speaker and semantic identification to obtain an identity authentication result without requiring specific equipment.
Enables secure and convenient identity authentication by allowing users to complete challenges accurately using voice inputs, eliminating the need for separate devices and overcoming issues with forgotten or cracked keys.
Smart Images

Figure 2026057469000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an authentication system and its method, and particularly to an identity authentication system and its method.
Background Art
[0002] With the development of science and technology, electronic devices are widely applied in the technical fields of manufacturing, communication, transportation, healthcare, commerce, community communication, and entertainment. For example, servers providing distributed cluster computing functions, embedded devices incorporated into automation equipment in multiple different technical fields, personal electronic devices for realizing portable video viewing, convenient access to smart digital assistants, etc. Furthermore, currently, multiple types of consumer electronic devices such as smartphones and tablets have been developed and have become essential items in people's lives and work. However, with the popularization of consumer electronic devices, the emphasis on the security of the devices is increasing.
[0003] Particularly, regarding identity authentication on applications, passwords are essential for device security and are formally called keys. As the complexity of security improves, it even derives to encryption keys and is further divided into public keys and private keys for handshaking between applications, and is applied to public handshakes and private handshakes.
[0004] However, encryption technology for identity verification is no longer a problem the public faces. The problem the public faces is how to set relatively good keys in applications for identity verification. Therefore, identity verification service providers have created various technologies to help users set keys in applications. However, if a user forgets the key being used or if the key is cracked, various problems arise. For example, it was necessary to find the key, set up a new account, or create a new account to solve the problem of forgotten or cracked keys.
[0005] Currently, technologies have emerged that allow for the acquisition of one-time passwords by using separate electronic devices. For example, one-time password information can be received on a mobile phone, one-time passwords can be generated using a smartphone application, or one-time passwords can be provided by smart security locks. However, if the electronic device providing the one-time password is lost or not adequately protected, the one-time password will still be unusable or can be cracked. Alternatively, if the mobile phone is out of range, the user cannot acquire a one-time password using the electronic device and therefore cannot use it for identity verification.
[0006] Furthermore, the current method of identity verification involves users completing the process using gestures, touching a touchscreen with their fingers, or pressing buttons. This makes identity verification extremely inconvenient for people with visual impairments or those who cannot see the screen.
[0007] In view of the problems of the prior art described above, the present invention provides an identity authentication system and method that can improve situations in which a key is forgotten or a one-time password cannot be used. [Overview of the project] [Problems that the invention aims to solve]
[0008] One of the objectives of the present invention is to provide an identity authentication system and method that improves security and enables identity authentication without the need for specific equipment by constructing signal sampling data based on a first voice signal from the user, inputting a second voice signal according to authentication prompt information, and performing speaker identification and semantic identification to obtain an identity authentication result. [Means for solving the problem]
[0009] The present invention provides an identity authentication method. The identity authentication method is an identity authentication method applied when an arithmetic processor inputs a first audio signal to the arithmetic processor through an audio input element, wherein the arithmetic processor samples and identifies the first audio signal by executing a speaker identification module and generates corresponding signal sampling data, the signal sampling data includes at least one first signal segment, the identity authentication method first generates authentication prompt information randomly using the arithmetic processor and transmits it to an output element, causing the output element to output the authentication prompt information, the authentication prompt information includes at least one prompt object and object prompt information, the object prompt information corresponds to the at least one prompt object, and subsequently the audio The second voice signal is input to the arithmetic processor through the voice input element according to the authentication prompt information, and the arithmetic processor is instructed to execute the speaker identification module to sample at least one second signal segment from the second voice signal and identify the at least one second signal segment based on the at least one first signal segment of the previously generated signal sampling data, thereby identifying the speaker of the second voice signal. The arithmetic processor is then instructed to execute the semantic identification module to identify the second voice signal and generate semantic object data, and the arithmetic processor is instructed to compare the object prompt information based on the semantic object data to generate an identity authentication result. As a result, a person who is correctly authenticated can complete the challenge accurately, and the user can complete identity authentication quickly without binding any device.
[0010] In one embodiment of the present invention, the registration prompt information and the authentication prompt information are video or audio.
[0011] In one embodiment of the present invention, the arithmetic processor is made to execute the speaker identification module to sample at least one sampling segment of the second speech signal, to identify the at least one sampling segment based on the at least one first signal segment, and to execute the semantic identification module to identify the second speech signal and generate semantic object data. In this step, the at least one second signal segment corresponds to at least one second speaker feature parameter, and the semantic identification unit executes the semantic identification module to perform feature extraction, obtain feature values of the second speech signal, and convert them into semantic object data by combining the at least one second speaker feature parameter of the second speech signal with the feature extraction results.
[0012] In one embodiment of the present invention, the processing unit executes the speaker identification module to convert the first speech signal into a plurality of word vectors, encodes an order into these word vectors, extracts features from these word vectors to obtain a plurality of feature vectors, and generates the signal sampling data by performing a normalization operation on these feature vectors.
[0013] In one embodiment of the present invention, the arithmetic processor is instructed to execute the speaker identification module to sample at least one second signal segment from the second voice signal, to identify the at least one second signal segment based on the at least one first signal segment, and to execute the semantic identification module of the identity authentication program to identify the second voice signal and generate semantic object data. In this step, first, the arithmetic processor samples at least one second signal segment based on the second voice signal, then compares the at least one second signal segment based on the at least one first signal segment to identify the second voice signal, and if the at least one second signal segment matches the at least one first signal segment, the arithmetic processor can determine that the speaker of the first voice signal and the speaker of the second voice signal are the same person, and then executes the semantic identification module to identify the second voice signal and generate the semantic object data.
[0014] In one embodiment of the present invention, in the step of running the speaker identification module to sample at least one second signal segment from the second speech signal, identifying the at least one second signal segment based on the at least one first signal segment, and running the semantic identification module of the identity authentication program to identify the second speech signal and generate semantic object data, the arithmetic processor runs the speaker identification module to convert the second speech signal into a plurality of word vectors, encode the order of these word vectors, extract features from these word vectors to obtain a plurality of second feature vectors, and generate the at least one second signal segment by performing a normalization operation on these second feature vectors.
[0015] In one embodiment of the present invention, the speaker identification module is a WavLM module, a SpeakerNet module, or a TitaNet module, and the semantic identification module is a TranSformer module, a wav2Vec2.0 module, or a LAS module.
[0016] The present invention further provides an identity authentication system. The identity authentication system comprises an arithmetic processor and an output element, the arithmetic processor having a voice input element connected to it, receiving a first voice signal through the voice input element to randomly generate authentication prompt information, executing a speaker identification module to sample and identify the first voice signal, and generating corresponding signal sampling data, the signal sampling data comprising at least one first signal segment, the output element being connected to the arithmetic processor, the output element outputting the authentication prompt information during the authentication stage, the authentication prompt information comprising at least one prompt object and object prompt information, the object prompt information corresponding to the at least one prompt object. The arithmetic processor then receives a second audio signal through the audio input element, and by executing the speaker identification module, the arithmetic processor samples at least one second signal segment from the second audio signal and identifies the at least one second signal segment based on the at least one first signal segment, thereby determining whether the speaker of the second audio signal and the speaker of the first audio signal are the same person. The arithmetic processor also executes the semantic identification module to identify the second audio signal and generate semantic object data, which includes a plurality of semantic objects. The arithmetic processor then generates an identity authentication result by comparing the object prompt information based on the semantic object data. As a result, a correctly authenticated person can complete the challenge accurately, and the user can complete identity authentication quickly without binding any device.
[0017] In another embodiment of the present invention, the registration prompt information and the authentication prompt information of the identity authentication system are video information or audio information.
[0018] In another embodiment of the present invention, the arithmetic processor executes the speaker identification module to convert the first speech signal into a plurality of word vectors, encode an order into these word vectors, extract features from these word vectors to obtain a plurality of feature vectors, and generate the signal sampling data by performing a normalization operation on these feature vectors.
[0019] In another embodiment of the present invention, the arithmetic processor identifies the second audio signal by comparing the at least one second signal segment based on the at least one first signal segment, and if the at least one second signal segment matches the at least one first signal segment, the arithmetic processor identifies the second audio signal and generates the semantic object data by executing the semantic identification module.
[0020] In another embodiment of the present invention, the computing processor executes the speaker identification module to convert the second speech signal into a plurality of word vectors, encode an order into these word vectors, extract features from these word vectors to obtain a plurality of second feature vectors, and generate the at least one second signal segment by performing a normalization operation on these second feature vectors.
[0021] In another embodiment of the present invention, the at least one second signal segment corresponds to at least one second speaker feature parameter, and the arithmetic processor executes the semantic identification module to extract features from the second speech signal and converts the feature extraction results of the second speech signal into semantic object data by combining them with the at least one second speaker feature parameter.
[0022] In another embodiment of the present invention, the speaker identification module of the authentication system is a WavLM module, a SpeakerNet module or a TitaNet module, and the semantic identification module of the authentication system is a TranSformer module, a wav2Vec2.0 module or a LAS module.
Brief Description of Drawings
[0023] [Figure 1A] FIG. 1A is a flowchart for obtaining signal sampling data according to an embodiment of the present invention. [Figure 1B] FIG. 1B is a flowchart for authentication according to an embodiment of the present invention. [Figure 2A] FIG. 2A is a schematic diagram for obtaining signal sampling data according to an embodiment of the present invention. [Figure 2B] FIG. 2B is a schematic diagram for randomly generating authentication prompt information according to an embodiment of the present invention. [Figure 2C] FIG. 2C is a schematic diagram for inputting an audio signal according to an embodiment of the present invention. [Figure 2D] FIG. 2D is a schematic diagram for comparing a first signal segment and an audio signal according to an embodiment of the present invention. [Figure 2E] FIG. 2E is a schematic diagram for obtaining semantic object data according to an embodiment of the present invention. [Figure 2F] FIG. 2F is a schematic diagram for obtaining and outputting an authentication result according to an embodiment of the present invention. [Figure 3] FIG. 3 is a schematic diagram of registration prompt information according to an embodiment of the present invention. [Figure 4A] FIG. 4A is a schematic diagram of authentication prompt information according to an embodiment of the present invention. [Figure 4B] FIG. 4B is a schematic diagram of authentication prompt information according to an embodiment of the present invention. [Figure 4C]Figure 4C is a schematic diagram of authentication prompt information relating to one embodiment of the present invention. [Figure 5] Figure 5 is a flowchart of identity authentication according to another embodiment of the present invention. [Figure 6A] Figure 6A is a schematic diagram illustrating the acquisition of signal sampling data according to another embodiment of the present invention. [Figure 6B] Figure 6B is a schematic diagram illustrating the random generation of authentication prompt information according to another embodiment of the present invention. [Figure 6C] Figure 6C is a schematic diagram showing an input of an audio signal according to another embodiment of the present invention. [Figure 6D] Figure 6D is a schematic diagram comparing a first signal segment and an audio signal according to another embodiment of the present invention. [Figure 6E] Figure 6E is a schematic diagram illustrating the acquisition of semantic object data according to another embodiment of the present invention. [Figure 6F] Figure 6F is a schematic diagram illustrating the acquisition and output of identity authentication results according to another embodiment of the present invention. [Figure 7A] Figure 7A is a schematic diagram illustrating the acquisition of signal sampling data according to another embodiment of the present invention. [Figure 7B] Figure 7B is a schematic diagram illustrating the random generation of authentication prompt information according to another embodiment of the present invention. [Figure 7C] Figure 7C is a schematic diagram showing an input of an audio signal according to another embodiment of the present invention. [Figure 7D] Figure 7D is a schematic diagram comparing a first signal segment and an audio signal according to another embodiment of the present invention. [Figure 7E] Figure 7E is a schematic diagram illustrating the acquisition of semantic object data according to another embodiment of the present invention. [Figure 7F] Figure 7F is a schematic diagram illustrating the acquisition and output of identity authentication results according to another embodiment of the present invention. [Figure 8A] Figure 8A is a schematic diagram illustrating the acquisition of signal sampling data according to another embodiment of the present invention. [Figure 8B]Figure 8B is a schematic diagram illustrating the random generation of authentication prompt information according to another embodiment of the present invention. [Figure 8C] Figure 8C is a schematic diagram showing an input of an audio signal according to another embodiment of the present invention. [Figure 8D] Figure 8D is a schematic diagram comparing a first signal segment and an audio signal according to another embodiment of the present invention. [Figure 8E] Figure 8E is a schematic diagram illustrating the acquisition of semantic object data according to another embodiment of the present invention. [Figure 8F] Figure 8F is a schematic diagram illustrating the acquisition and output of identity authentication results according to another embodiment of the present invention. [Modes for carrying out the invention]
[0024] To help examiners more clearly understand and identify the features and effects of the present invention, optimal examples and a detailed description are provided below.
[0025] Current identity authentication technologies have problems such as the inability to obtain a one-time password unless the key is forgotten, the key is cracked, or a specific device is bound. Therefore, the present invention provides an identity authentication system and method, which inputs a corresponding first voice signal through a voice input element, identifies it and converts it into signal sampling data, inputs a second voice signal based on authentication prompt information and performs authentication, and if the second voice signal matches at least one first signal segment of the signal sampling data, performs semantic identification to obtain semantic object data and obtains an identity authentication result by comparing it with the authentication prompt information. By correctly authenticating a person, identity authentication is completed by completing the correct challenge, and the problem of not being able to obtain a one-time password unless the key is forgotten, the key is cracked, or a specific device is bound is solved.
[0026] The identity verification system and its methods are as follows:
[0027] Refer to Figure 1A. Figure 1A is a flowchart for acquiring signal sampling data according to one embodiment of the present invention. In this embodiment, the identity authentication method of the present invention first acquires signal sampling data for speaker identification and includes the following steps.
[0028] Step S10: Inputs the first audio signal to the processing processor through the audio input element.
[0029] Step S12: The arithmetic processor executes the speaker identification module, samples and identifies the first speech signal, and generates the corresponding signal sampling data.
[0030] Please refer to Figure 1B. Figure 1B is a flowchart of identity authentication according to one embodiment of the present invention. In this embodiment, the identity authentication method of the present invention continues to perform the following steps based on the signal sampling data obtained in step S12 above.
[0031] Step S20: The arithmetic processor is used to execute the identity authentication program, randomly generate authentication prompt information, transmit it to the output element, and cause the output element to output the authentication prompt information.
[0032] Step S30: The second audio signal is input to the processing processor via the audio input element according to the authentication prompt information.
[0033] Step S40: The arithmetic processor is instructed to execute the speaker identification module of the identity authentication program to identify the second voice signal based on the first signal segment, and also to execute the semantic identification module of the identity authentication program to identify the second voice signal and generate semantic object data.
[0034] Step S60: The arithmetic processor is instructed to compare object prompt information based on semantic object data and generate an identity authentication result.
[0035] Refer to Figures 2A to 2F together. Figures 2A to 2F are schematic diagrams illustrating the acquisition of signal sampling data, random generation of authentication prompt information, input of an audio signal, comparison of a first signal segment with the audio signal, acquisition of semantic object data, and acquisition of an identity authentication result according to one embodiment of the present invention. As shown in the figures, the identity authentication method of the present invention is applied to an identity authentication system 10, which comprises an electronic device 12 and a host 14. In this embodiment, an arithmetic processor 142 is provided in the host 14, the electronic device 12 is communicably connected to the host 14, and the electronic device 12 is provided with an audio input element 122 and an output element 124. For example, the electronic device 12 is connected to the host 14 via a wireless network, the electronic device 12 is a smartphone, and the host 14 is a remote server. The electronic device 12 and the host 14 can transmit data through a transmission protocol such as the Hypertext Transfer Protocol or the Transmission Control Protocol (TCP). The arithmetic processor 142 executes an identity authentication program 1422, which comprises a speaker identification unit 1422A and a semantic identification unit 1422B. The speaker identification unit 1422A executes the speaker identification module 14222A, and the semantic identification unit 1422B executes the semantic identification module 14222B.
[0036] In step S10, as shown in Figure 2A, user U inputs their first voice U1 to the electronic device 12 through the voice input element 122, and the electronic device 12 transmits it to the host 14 as a first voice signal VOC1. That is, the first voice signal VOC1 is input to the arithmetic processor 142 through the voice input element 122. In the subsequent step S12, the arithmetic processor 142 executes the speaker identification module 14222A to sample and identify the first voice signal VOC1 and generate the corresponding signal sampling data SD. That is, from the first voice signal VOC1, after identification, SD1 is cut and extracted into at least one first signal segment corresponding to at least one first speaker embedding parameter SE1 and stored as the signal sampling data SD. Here, the processing unit 142 executes the speaker identification module 14222A to convert the first speech signal VOC1 into multiple word vectors, encode the order of these word vectors, extract features from these word vectors to obtain multiple feature vectors, and generate the signal sampling data SD by performing a normalization operation on these word vectors. Since the speaker identification module 14222A itself is a prior art such as the WavLM module, SpeakerNet module, or TitaNet module, a description of the speaker identification module 14222A is omitted here.
[0037] Here, the user of the electronic device 12 constructs a corresponding signal sampling data SD on the host 14, and the host 14 may store the signal sampling data SD on an internal recording medium such as a conventional hard disk, solid-state hard disk, or memory, or it may store the signal sampling data SD on an external physical database or cloud database such as a NAS system or Google Cloud Hard Disk.
[0038] Furthermore, as shown in Figure 3, the identity authentication program 1422 of the present invention further generates registration prompt information 1424 and sends it to the electronic device 12, which can output the registration prompt information 1424 as video information or audio information through the output element 124. In this embodiment, the registration prompt information 1424 is shown as video information through the output element 124, and therefore, in this embodiment, the output element 124 is a display element such as a liquid crystal display. The registration prompt information 1424 of this embodiment includes a plurality of registration prompt characters (for example, Ming Dynasty short stories), and the user can input the corresponding first audio signal VOC1 according to the registration prompt information 1424. Therefore, the speaker identification module 14222A executed by the arithmetic processor 142 identifies and samples the first audio signal VOC1 and generates the corresponding signal sampling data SD. Furthermore, the registration prompt information 1424 may further include at least one prompt object, so that in addition to displaying video information through the output element 124, it is also possible to play audio information. By having the user U input the first audio signal VOC1 for more than 5 seconds in accordance with the registration prompt information 1424, relatively effective signal sampling data SD can be obtained.
[0039] Refer again to Figure 1A. In step S20, as shown in Figure 2B, the host 14 randomly generates authentication prompt information 1426 by the authentication program 1422 executed by the arithmetic processor 142, and as shown in Figure 4A, the authentication prompt information 1426 displayed by the output element 124 includes a plurality of first prompt objects H1, a plurality of second prompt objects H2, and object prompt information H3, the object prompt information H3 corresponds to these first prompt objects H1 or these second prompt objects H2, in particular to the shape and quantity of these first prompt objects H1 or these second prompt objects H2, for example, these first prompt objects H1 are five triangles and these second prompt objects H2 are four circles. In addition, as shown in Figure 4B, the authentication prompt information 1426 displayed by the output element 124 may also be calculation prompt information, that is, it includes a calculation object H4 and object prompt information H3, for example, "1+99=?" as the calculation object and object prompt information H3 is "Please answer the following calculation." Furthermore, as shown in Figure 4C, the present invention may also use the above-mentioned registration prompt information 1424 as authentication prompt information 1426, that is, the authentication prompt information 1426 displayed by the output element 124 includes a character object H5 and object prompt information H3, for example, object prompt information H3 is "Please read the following literally" and character object H5 is "It's a nice day today."
[0040] In step S30, as shown in Figure 2C, when the user U of the electronic device 12 emits a second voice U2, the second voice signal VOC2 is input to the electronic device 12 through the voice input element 122 in accordance with the authentication prompt information 1426 displayed in step S20, and transmitted to the host 14. In other words, the arithmetic processor 142 receives the second voice signal VOC2 through the voice input element 122, and subsequently the speaker identification module 14222A in the identity authentication program 1422 makes a determination.
[0041] In step S40, as shown in Figure 2D, the arithmetic processor 142 executes the speaker identification module 14222A in the identity authentication program 1422, reading the previously stored signal sampling data SD and sampling at least one second signal segment VOC21 from the second voice signal VOC2, and identifying the second voice signal VOC2 by determining through the speaker identification module 14222A whether the at least one second signal segment VOC21 of the second voice signal VOC2 matches the at least one first signal segment based on the signal sampling data SD.
[0042] Here, the arithmetic processor 142 executes the speaker identification module 14222A to convert the second speech signal VOC2 into multiple word vectors, encode the order of these word vectors, extract features from these word vectors to obtain multiple second feature vectors, and normalize these second feature vectors to generate at least one second signal segment VOC21. This calculation method is an example of conventional speaker identification techniques and is therefore omitted from explanation. Similar to the at least one first signal segment SD1, the at least one second signal segment VOC21 corresponds to at least one second speaker feature parameter.
[0043] Simultaneously, in step S40, referring to Figure 2E, the arithmetic processor 142 identifies the second speech signal VOC2 by executing the semantic identification module 14222B of the semantic identification unit 1422B and generates the corresponding semantic object data 144. Here, the arithmetic processor 142 extracts features from the second speech signal VOC2 by executing the semantic identification module 14222B, and combines the at least one second speaker feature parameter SE2 of the second speech signal VOC2 with the feature extraction result FE to convert it into semantic object data 144. For example, the semantic identification module is the TranSformer module, the wav2Vec2.0 module, or the LAS module, all of which first extract features and then combine the feature extraction result with speaker features to convert it into semantic data. Since these are mature technologies, the explanation of the semantic identification module 14222B is omitted here.
[0044] In step S60, as shown in Figure 2E, the arithmetic processor 142 compares the object prompt information H3 based on the semantic object data 144 and generates an identity authentication result 146. That is, if the arithmetic processor 142 determines that the semantic object data 144 matches the object prompt information H3, the identity authentication result 146 is displayed as authentication passed, and if the arithmetic processor 142 determines that the semantic object data 144 does not match the object prompt information H3, the identity authentication result 146 is displayed as authentication failed. The present invention solves the problem that a one-time password cannot be obtained unless the key is forgotten, the key is cracked, or a specific device is bound, by completing identity authentication by having a correctly authenticated person complete the correct challenge.
[0045] In the embodiment described above, the arithmetic processor 142 executes the speaker identification module 14222A and the semantic identification module 14222B simultaneously in the same step. Alternatively, they may be executed separately, and the details of this are as follows.
[0046] Refer to Figure 5. Figure 5 is a flowchart of identity authentication according to another embodiment of the present invention. As shown in the figure, the identity authentication method of the present invention may execute the speaker identification module 14222A and the semantic identification module 14222B separately at the authentication stage. The step of acquiring the signal sampling data SD in this embodiment is the same as steps S10 to S12 in the above-described embodiment, and no further illustrations or explanations are provided in this embodiment. Steps S20 to S30 and step S60 in this embodiment are the same as in the above-described embodiment, and their explanation is omitted here.
[0047] In step S42, as shown in Figure 2D, the arithmetic processor 142 executes the speaker identification module 14222A of the identity authentication program 1422 to obtain the previously stored signal sampling data SD and to sample at least one second signal segment VOC21 from the second voice signal VOC2. In the following step S44, the arithmetic processor 142, through the speaker identification module 14222A, determines, based on the signal sampling data SD, whether at least one second signal segment VOC21 of the second voice signal VOC2 matches at least one first signal segment SD1. If it is determined that they match, the arithmetic processor 142 then executes step S46 to read the second voice signal VOC2 into the semantic identification unit 1422B. If it is determined that they do not match, the arithmetic processor 142 then executes step S30 to have the electronic device 12 transmit another second voice signal VOC2 to the host 14. In the following step S42, the arithmetic processor 142 samples the at least one second signal segment VOC21, makes a determination in step S44, and then decides whether to execute step S30 or step S46.
[0048] In step S46, as shown in Figure 2E, the arithmetic processor 142 executes the semantic identification module 14222B of the semantic identification unit 1422B to identify the second voice signal VOC2 and generate the corresponding semantic object data 144. Here, the arithmetic processor 142 performs feature extraction on the second voice signal VOC2 by executing the semantic identification module 14222B, combines the feature extraction results of the second voice signal VOC2 with at least one second speaker feature parameter SE2 corresponding to at least one second signal segment VOC21, and converts it into semantic object data 144.
[0049] The embodiments described so far have taken the example that the arithmetic processor 142 of the host 14 executes the identity authentication program 1422, but the present invention may also be another identity authentication system 20 that includes an electronic device 22 that directly executes the identity authentication program 1422, and the details are as follows.
[0050] Refer further to Figures 6A to 6F. Figures 6A to 6F are schematic diagrams illustrating another embodiment of the present invention, in which signal sampling data is acquired, authentication prompt information is randomly generated, an audio signal is input, a first signal segment is compared with the audio signal, semantic object data is acquired, and an identity authentication result is acquired. In this embodiment, the identity authentication system 20 of the present invention comprises an electronic device 22, which comprises an audio input element 222, an output element 224, and an arithmetic processor 226, the arithmetic processor 226 executing the identity authentication program 1422. Therefore, the difference between this embodiment and the previously described embodiment is that in this embodiment, the electronic device 22 directly performs all the steps described in the previously described embodiment.
[0051] Refer again to Figure 1A. In step S10, as shown in Figure 6A, the first audio signal VOC1 is input to the electronic device 22 through the audio input element 222, and the first audio signal VOC1 is input to the arithmetic processor 226 through the audio input element 222. As a result, in step S12, the arithmetic processor 226 executes the speaker identification module 1422A to sample and identify the first audio signal VOC1 and generate the corresponding signal sampling data SD. That is, it extracts the at least one first signal segment SD1 corresponding to the at least one first speaker feature parameter SE1 from the first audio signal VOC1 and stores it as the signal sampling data SD.
[0052] In step S20, as shown in Figure 6B, the electronic device 22 generates authentication prompt information 1426 through the identity authentication program 1422 executed by the arithmetic processor 226.
[0053] In step S30, as shown in Figure 6C, the electronic device 22 receives the second audio signal VOC2 through the audio input element 222. That is, the arithmetic processor 226 receives the second audio signal VOC2 through the audio input element 222, and then the speaker identification module 14222A in the identity authentication program makes a determination.
[0054] In step S42, as shown in Figure 6D, the arithmetic processor 226 executes the speaker identification module 14222A of the identity authentication program 1422 to read the previously stored signal sampling data SD, sample at least one second signal segment VOC21 from the second voice signal VOC2, and determine through the signal sampling data SD whether the first signal segment SD1 matches the second signal segment VOC21 of the second voice signal VOC2. If they match, the arithmetic processor 226 then executes step S46; if they do not match, the arithmetic processor 226 executes step S30 again to transmit another second voice signal VOC2 to the electronic device 22, and then proceeds to steps S42 and S44 to determine whether the speaker of the other second voice signal VOC2 is the same as the speaker of the first voice signal VOC1.
[0055] In step S46, as shown in Figure 6E, the arithmetic processor 226 executes the semantic identification module 14222B to identify the second audio signal VOC2 and generate the corresponding semantic object data 144. Here, the arithmetic processor 226 performs feature extraction on the second audio signal VOC2 by executing the semantic identification module 14222B, and combines the feature extraction results of the second audio signal VOC2 with the second speaker patent parameters to convert them into semantic object data 144. For example, the semantic identification module is the TranSformer module, the wav2Vec2.0 module, or the LAS module, so the description of the semantic identification module 14222B is omitted.
[0056] In step S60, as shown in Figure 6F, the arithmetic processor 226 compares the object prompt information H3 based on the semantic object data 144 and generates the identity authentication result 146. That is, if the arithmetic processor 226 determines that the semantic object data 144 matches the object prompt information H3, the identity authentication result 146 is displayed as authentication passed, and if the arithmetic processor 226 determines that the semantic object data 144 does not match the object prompt information H3, the identity authentication result 146 is displayed as authentication failed. The present invention solves the problem that a one-time password cannot be obtained unless the key is forgotten, the key is cracked, or a specific device is bound, by completing identity authentication by having a correctly authenticated person complete the correct challenge.
[0057] From the embodiments described above, it can be seen that the identity authentication system and method of the present invention have the advantage that the WiTime password is difficult to crack, and that the user does not need to prepare any separate device or software during authentication. Furthermore, the authentication prompt information 1426 described above may include non-personal data and non-pure numerical arrays.
[0058] Furthermore, the authentication method of the present invention is also applicable to purely electrical circuit operation methods, and as shown in Figures 7A to 7F, the identity authentication system 30 comprises a control processing circuit 321, an audio input element 322, an output element 324, a speaker identification calculation processor 326, and a semantic identification calculation processor 328, and is equipped with a first memory element RAM1 and a second memory element RAM2, the control processing circuit 321, the first memory element RAM1 and the second memory element RAM2 correspond to the calculation processor 142 described in the above embodiment, and the control processing circuit 321, the speaker identification calculation processor 326 and the semantic identification calculation processor The operation of the decoder 328 corresponds to the operation of the identity authentication program 1422 in the embodiment described above, the speaker identification arithmetic processor 326 and the semantic identification arithmetic processor 328 execute the speaker identification module 14222A and the semantic identification module 14222B respectively, the control processing circuit 321, the speaker identification arithmetic processor 326 and the semantic identification arithmetic processor 328 can each be implemented by an FPGA circuit, a SOC circuit or other integrated circuit having logic operation capabilities, and the first memory element RAM1 and the second memory element RAM2 are used for temporary storage of data. The other execution methods are the same as in the embodiment described above, so their description is omitted.
[0059] Furthermore, the identity authentication method of the present invention is applicable to a purely electrical circuit operation method, and the speaker identification module 14222A and the semantic identification module 14222B can share an arithmetic processor. That is, as shown in Figures 8A to 8F, the difference between Figures 7A to 7F and Figures 8A to 8F is that in Figures 7A to 7F, the speaker identification module 14222A and the semantic identification module 14222B are provided in the speaker identification arithmetic processor 326 and the semantic identification arithmetic processor 328, respectively, but in Figures 8A to 8F, both the speaker identification module 14222A and the semantic identification module 14222B are provided in the identification arithmetic processor 330, meaning that the identification arithmetic processor 330 has the arithmetic functions of both the aforementioned speaker identification arithmetic processor 326 and the semantic identification arithmetic processor 328. Other execution methods are the same as in the embodiments described above, so their explanation is omitted.
[0060] Furthermore, according to the embodiments described above, the identity authentication system and method of the present invention include a registration step in which a first voice signal is input to acquire corresponding signal sampling data, thereby registering the user's voice sample. The present invention also simultaneously includes an authentication step, in which authentication prompt information is output randomly to prompt the user to input a corresponding second voice signal according to the authentication prompt information, to identify whether the speaker is a user of the signal sampling data, and after confirming that the user is a user of the signal sampling data, semantic identification is performed to acquire semantic object data, and finally the semantic object data is compared with the authentication prompt information to generate a corresponding identity authentication result, thereby solving the problem of not being able to obtain a one-time password unless the key is forgotten, the key is cracked, or a specific device is bound.
[0061] Therefore, the present invention possesses novelty, inventiveness, and practicality, and complies with the patent application requirements of the Taiwan Patent Law. We hereby file an invention patent application in accordance with the law and hope that the patent will be granted promptly.
[0062] The above represents the optimal embodiment of the present invention and does not limit the scope of implementation of the present invention. Various changes and improvements based on the shape, structure, features and concept described in the claims of the present invention are all included within the scope of the claims of the present invention. [Explanation of Symbols]
[0063] 10, 20: Identity Verification System 12, 22: Electronic equipment 122, 222: Audio input elements 124, 224: Output elements 226: Arithmetic Processor 14: Host 142: Arithmetic Processor 1422: Identity Verification Program 1422A: Speaker identification unit 14222A: Speaker identification module 1422B: Semantic Identification Unit 14222B: Semantic Identification Module 1424: Registration prompt information 1426: Authentication prompt information 144: Semantic Object Data 146: Identity Verification Results 322: Audio input element 324: Output element 326: Speaker Identification Processing Processor 328: Semantic Identification Processor 330: Identification and processing processor H1: First prompt object H2: Second prompt object H3: Object Prompt Information SD: Signal Sampling Data SD1: First signal segment SE1: Speaker 1 characteristic parameters SE2: Second speaker characteristic parameters U:User U1: First audio U2: Second audio VOC1: First audio signal VOC2: Second audio signal VOC21: Second signal segment S10~S60: Step
Claims
1. An identity authentication method applied when an arithmetic processor inputs a first audio signal to the arithmetic processor through an audio input element, wherein the arithmetic processor, by executing a speaker identification module, samples and identifies the first audio signal and generates corresponding signal sampling data, the signal sampling data includes at least one first signal segment, and the identity authentication method is The arithmetic processor randomly generates authentication prompt information and transmits it to the output element, causing the output element to output the authentication prompt information, the authentication prompt information includes at least one prompt object and object prompt information, the object prompt information includes a step corresponding to the at least one prompt object, The steps include inputting a second audio signal to the arithmetic processor through the audio input element in accordance with the authentication prompt information, The steps include: causing the arithmetic processor to execute the speaker identification module to sample at least one second signal segment from the second speech signal, identifying the at least one second signal segment based on the at least one first signal segment, and causing the semantic identification module to execute to identify the second speech signal and generate semantic object data; A method for authenticating identity, comprising the step of causing the arithmetic processor to compare the object prompt information based on the semantic object data to generate an identity authentication result.
2. The arithmetic processor is used to randomly generate authentication prompt information and transmit it to the output element, causing the output element to output the authentication prompt information, the authentication prompt information includes at least one prompt object and object prompt information, and the object prompt information corresponds to the step of at least one prompt object, The identity authentication method according to claim 1, wherein the host transmits the authentication prompt information generated by the processing processor to an electronic device, the electronic device outputs the authentication prompt information through the output element, and the authentication prompt information is video information or audio information.
3. In the step of inputting a second audio signal to the arithmetic processor in accordance with the authentication prompt information through the audio input element, The identity authentication method according to claim 2, wherein the electronic device receives the second voice signal through the voice input element in accordance with the authentication prompt information and transmits it to the host, thereby inputting the second voice signal to the processing processor.
4. In the step of having the arithmetic processor execute the speaker identification module to sample at least one second signal segment from the second speech signal, identifying the at least one second signal segment based on the at least one first signal segment, and having the semantic identification module execute to identify the second speech signal and generate semantic object data, The identity authentication method according to claim 1, wherein the at least one second signal segment corresponds to at least one second speaker feature parameter, and the arithmetic processor executes the semantic identification module to extract features from the second speech signal and combines the feature extraction results of the second speech signal with the at least one second speaker feature parameter to convert them into semantic object data.
5. The identity authentication method according to claim 4, wherein the processing unit executes the speaker identification module to convert the first speech signal into a plurality of word vectors, encodes an order into these word vectors, extracts features from these word vectors to obtain a plurality of feature vectors, and generates the signal sampling data by performing a normalization operation on these feature vectors.
6. The steps of having the arithmetic processor execute the speaker identification module to sample at least one second signal segment from the second audio signal, identify the at least one second signal segment based on the at least one first signal segment, and have the semantic identification module of the identity authentication program execute to identify the second audio signal and generate semantic object data are as follows: The arithmetic processor samples the at least one second signal segment based on the second audio signal. The arithmetic processor compares the at least one second signal segment based on the at least one first signal segment and identifies the second audio signal. The identity authentication method according to claim 1, wherein if the at least one second signal segment matches the at least one first signal segment, the arithmetic processor executes the semantic identification module to identify the second voice signal and generate the semantic object data.
7. In the step of having the arithmetic processor execute the speaker identification module to sample at least one second signal segment from the second voice signal, identify the at least one second signal segment based on the at least one first signal segment, and have the semantic identification module of the identity authentication program execute to identify the second voice signal and generate semantic object data, The identity authentication method according to claim 6, wherein the processing unit executes the speaker identification module to convert the second speech signal into a plurality of word vectors, encode an order into these word vectors, extract features from these word vectors to obtain a plurality of second feature vectors, and generate the at least one second signal segment by performing a normalization operation on these second feature vectors.
8. The identity authentication method according to claim 1, wherein the speaker identification module is a WaveLM module, a SpeakerNet module, or a TitaNet module, and the semantic identification module is a Transformer module, a wave2Vec2.0 module, or a LAS module.
9. The identity authentication method according to claim 1, wherein the arithmetic processor comprises a speaker identification processor and a semantic identification processor, the speaker identification processor executes the speaker identification module, and the semantic identification processor executes the semantic identification module.
10. The identity authentication method according to claim 9, wherein the speaker identification processor and the semantic identification processor are integrated into an identification arithmetic processor and capable of simultaneously executing the speaker identification module and the semantic identification module.
11. It is an identity verification system, A voice input element is connected, randomly generates authentication prompt information, receives a first voice signal through the voice input element, and executes a speaker identification module to sample and identify the first voice signal, generate corresponding signal sampling data, and the signal sampling data is processed by an arithmetic processor including at least one first signal segment. The system is connected to the aforementioned processing processor and outputs the authentication prompt information during the authentication phase, wherein the authentication prompt information includes at least one prompt object and object prompt information, and the object prompt information includes an output element corresponding to the at least one prompt object, An identity authentication system comprising: an arithmetic processor receiving a second audio signal through the audio input element; an arithmetic processor executing the speaker identification module to sample at least one second signal segment from the second audio signal, identifying the at least one second signal segment based on the at least one first signal segment, and executing the semantic identification module to identify the second audio signal and generate semantic object data; and an arithmetic processor generating an identity authentication result by comparing the object prompt information based on the semantic object data.
12. The identity authentication system according to claim 11, wherein the processing processor is provided in a host, the output element is provided in an electronic device, the host transmits the generated authentication prompt information to the electronic device, the electronic device outputs the authentication prompt information through the output element, and the authentication prompt information is video information or audio information.
13. The identity authentication system according to claim 12, wherein the voice input element is provided in the electronic device, and the electronic device receives the second voice signal through the voice input element in accordance with the authentication prompt information and transmits it to the host, thereby inputting the second voice signal to the processing processor.
14. The identity authentication system according to claim 11, wherein the processing unit executes the speaker identification module to convert the first speech signal into a plurality of word vectors, encode an order into these word vectors, extract features from these word vectors to obtain a plurality of first feature vectors, and generate the signal sampling data by performing a normalization operation on these first feature vectors.
15. The identity authentication system according to claim 11, wherein the arithmetic processor identifies the second audio signal by comparing the at least one second signal segment based on the at least one first signal segment, and if the at least one second signal segment matches the at least one first signal segment, the arithmetic processor executes the semantic identification module to identify the second audio signal and generate the semantic object data.
16. The identity authentication system according to claim 11, wherein the processing unit executes the speaker identification module to convert the second speech signal into a plurality of word vectors, encode an order into these word vectors, extract features from these word vectors to obtain a plurality of second feature vectors, and generate the at least one second signal segment by performing a normalization operation on these second feature vectors.
17. The identity authentication system according to claim 11, wherein the arithmetic processor performs feature extraction on the second audio signal by executing the semantic identification module and converts the feature extraction results of the second audio signal into semantic object data.
18. The identity authentication system according to claim 11, wherein the speaker identification module is a WaveLM module, a SpeakerNet module, or a TitaNet module, and the semantic identification module is a Transformer module, a wave2Vec2.0 module, or a LAS module.
19. The identity authentication system according to claim 11, wherein the arithmetic processor comprises a speaker identification processor and a semantic identification processor, the speaker identification processor executes the speaker identification module, and the semantic identification processor executes the semantic identification module.
20. The identity authentication system according to claim 19, wherein the speaker identification processor and the semantic identification processor are integrated into an identification arithmetic processor and capable of simultaneously executing the speaker identification module and the semantic identification module.