A multi-factor internet of things device authentication method and system
By leveraging the microphone and speaker systems of IoT devices, and utilizing acoustic characteristics to extract FBank and MFCC features as well as spatial distance features, a multi-factor authentication model is constructed. This solves the security and universality issues of device authentication in the power distribution IoT environment, and achieves efficient legality verification.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-06
- Publication Date
- 2026-03-24
AI Technical Summary
Existing device authentication methods are insufficient in security in the power distribution IoT environment, cannot effectively distinguish legitimate devices, and lack universality and convenience.
By leveraging the acoustic characteristics of the built-in microphone and speaker systems in IoT devices, FBank and MFCC features are extracted through audio signal acquisition and processing. Combined with the spatial distance features of the speaker and microphone, a multi-factor authentication model is constructed to verify the legitimacy of the device.
It enables universal, secure, convenient, and efficient authentication of IoT devices in power distribution network environments, resists replay attacks and injection attacks, and improves the security of device authentication.
Smart Images

Figure CN119652609B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of smart grids, and more specifically to a multi-factor Internet of Things (IoT) device authentication method and system. Background Technology
[0002] The construction of smart grids is a crucial strategic deployment in my country's energy sector, gradually driving changes in the country's basic power production model and representing the development direction of the power industry. The Internet of Things (IoT) has become a vital technological means to promote smart grid development. However, the ubiquitous, comprehensive sensing, reliable transmission, and intelligent processing characteristics of IoT networks also expose distribution IoT to greater security threats. One of the fundamental mechanisms for protecting the security of distribution IoT is device legitimacy verification, which verifies the identity of IoT entities and controls their access. Therefore, in the face of security threats in the distribution IoT environment, how to verify the legitimacy of IoT devices connected to distribution network terminals has become an urgent problem to be solved.
[0003] With the continuous development of technology and the increase of security threats, traditional device authentication solutions, such as key authentication, MAC address filtering, and IP address filtering, lack effective protection schemes and countermeasures against new intelligent devices accessing the network and emerging malicious attacks targeting intelligent devices. At present, some novel device authentication schemes utilize physical layer characteristics to enhance device identity authentication. Physical layer device identification schemes can be roughly divided into three categories: software fingerprinting, channel fingerprinting, and hardware fingerprinting. (1) In terms of software characteristics, many browser configuration information can be used to distinguish devices, such as user agents, installed fonts, plugin information, benchmark tests, etc. Software fingerprints are generated by the current configuration of the system, so software fingerprints are not static, but are likely to change over time. Moreover, software-based protocols cannot distinguish different physical devices running the same software. (2) Device authentication schemes based on channel fingerprints usually use radio strength signals (RSS) and channel state information (CSI). Although RSS-based mechanisms have been studied in depth, their efficiency is still limited because RSS can only provide coarse-grained information and is highly correlated with transmission power and distance. CSI-based mechanisms are difficult to promote because they rely on dedicated hardware and cannot be widely deployed on mobile devices. (3) Hardware fingerprinting schemes are designed based on reflections of hardware defects or unique responses of hardware to transmission waveforms, relying on some static characteristic sources. Much work has been dedicated to identifying devices by utilizing subtle differences in signals generated by hardware components. For example, wireless network cards can be distinguished by utilizing the characteristics of radio frequency signals emitted by a transmitter. However, these methods cannot be extended to Internet tracking. Data collected from accelerometers can also be used to distinguish users without active stimulation. Photos taken by cameras can also be distinguished by patterns and noise.
[0004] Therefore, the above-mentioned device authentication methods have various shortcomings and are not applicable to the power distribution network environment. Thus, it is necessary to explore and research new device authentication technologies for power distribution networks to achieve universality, security, convenience, and efficiency in verifying the identity of IoT devices in the power distribution network, thereby establishing the capability to verify the legitimacy of IoT devices in the power distribution network. Summary of the Invention
[0005] The technical problem to be solved by this invention is to provide a multi-factor IoT device authentication method and system that addresses the above-mentioned problems in the prior art. This method utilizes acoustic characteristics to capture manufacturing defects and spatial distance of the device's built-in microphone-speaker system to distinguish whether the device is legitimate and extracts an acoustic fingerprint for device authentication. This enables universal, secure, convenient, and efficient verification of the legitimacy of terminal devices and IoT devices in a power distribution network environment.
[0006] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is as follows:
[0007] A multi-factor authentication method for IoT devices includes the following steps:
[0008] The registration audio is played through the speaker of the IoT device and collected through the microphone of the IoT device. The registration audio collected by each microphone is processed to obtain the corresponding FBank features and MFCC features. At the same time, the spatial distance features between the speaker and the microphone of the IoT device are calculated based on the registration audio collected by each microphone. The spatial distance features and the FBank features and MFCC features corresponding to the registration audio collected by each microphone are used as sound fingerprints to train the authentication model.
[0009] The authentication audio is played through the speaker of the IoT device and collected by the microphone of the IoT device. The authentication audio collected by each microphone is processed to obtain the corresponding FBank features and MFCC features. At the same time, the spatial distance features between the speaker and the microphone of the IoT device are calculated based on the authentication audio collected by each microphone. The spatial distance features and the FBank features and MFCC features corresponding to the authentication audio collected by each microphone are used as the sound fingerprint input to the trained authentication model to obtain the authentication result.
[0010] Furthermore, before playing the registration audio through the speaker of the IoT device, and before playing the authentication audio through the speaker of the IoT device, it also includes: collecting ambient sound through the microphone of the IoT device.
[0011] Furthermore, the signal processing for both the registration audio captured by each microphone and the authentication audio captured by each microphone includes:
[0012] The sound signal captured by the current microphone is denoised. The denoised sound signal is then pre-emphasized, framed, and windowed. The signal is then subjected to a fast Fourier transform to obtain the frequency domain signal. The frequency domain spectrum is obtained. The Mel spectrum is obtained by using a filter bank on the frequency domain spectrum. The obtained Mel spectrum is then paired to calculate the FBank feature. The FBank feature is then subjected to a discrete cosine transform to obtain the MFCC feature.
[0013] Furthermore, the IoT device is equipped with two microphones. When calculating the spatial distance characteristics between the IoT device's speaker and microphone based on the registration audio collected by each microphone, and when calculating the spatial distance characteristics between the IoT device's speaker and microphone based on the authentication audio collected by each microphone, the expression for the spatial distance characteristics is as follows:
[0014]
[0015] in, and These represent the sound distortion produced by the first and second microphones of the IoT device, respectively. and This indicates the spatial distance between the speaker and the two microphones of an IoT device. This indicates the attenuation that occurs during sound propagation. It is distance and frequency The function.
[0016] Furthermore, when using the spatial distance features and the FBank and MFCC features corresponding to the registered audio collected by each microphone as sound fingerprints to train the authentication model, specifically, the sound fingerprints are used as sample features to train an unsupervised learning vector machine to obtain the authentication model.
[0017] This invention also proposes a multi-factor IoT device authentication system, comprising paired IoT devices, an authentication server, and a power distribution network terminal device, wherein:
[0018] The IoT device is used to send a registration request or authentication request to the verification server, and after receiving the instruction from the verification server, it generates the corresponding registration audio or authentication audio, plays the registration audio or authentication audio through a speaker, and simultaneously collects the played registration audio or authentication audio through a microphone and sends it to the verification server.
[0019] The verification server, upon receiving a registration request from an IoT device, sends a corresponding registration audio generation instruction to the IoT device. This instruction plays the registration audio through the IoT device's speaker and collects the played registration audio through the IoT device's microphone. The server then acquires the registration audio collected by each microphone of the IoT device and performs signal processing to obtain the corresponding FBank and MFCC features. Simultaneously, the server calculates the spatial distance features between the IoT device's speaker and microphone based on the registration audio collected by each microphone. The spatial distance features, along with the FBank and MFCC features corresponding to the registration audio collected by each microphone, are used as sound fingerprints to train the authentication model.
[0020] The verification server is also used to send a corresponding authentication audio generation instruction to the IoT device after receiving the authentication request from the IoT device, so as to play the authentication audio through the speaker of the IoT device and collect the played authentication audio through the microphone of the IoT device. Then, the authentication audio collected by each microphone of the IoT device is obtained and the signal is processed to obtain the corresponding FBank feature and MFCC feature. At the same time, the spatial distance feature between the speaker and the microphone of the IoT device is calculated based on the authentication audio collected by each microphone. The spatial distance feature and the FBank feature and MFCC feature corresponding to the authentication audio collected by each microphone are used as the sound fingerprint input to the trained authentication model to obtain the authentication result and send it to the distribution network terminal equipment.
[0021] The power distribution network terminal equipment is used to obtain and parse the authentication result. If the authentication result of the IoT device is a legitimate device, the IoT device is allowed to access; otherwise, the IoT device is blocked from accessing.
[0022] The present invention also proposes a multi-factor Internet of Things (IoT) device authentication apparatus, comprising a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of any of the multi-factor IoT device authentication methods described herein.
[0023] The present invention also proposes a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the multi-factor Internet of Things device authentication method described in any one of the claims.
[0024] The present invention also proposes a computer program product, including a computer program that, when executed by a processor, implements the steps of the multi-factor Internet of Things device authentication method described in any one of the claims.
[0025] Compared with the prior art, the advantages of the present invention are as follows:
[0026] (1) This invention utilizes the built-in microphone and speaker of the Internet of Things device to complete device authentication without the need to introduce additional hardware or sensors. It can be applied to devices that lack dedicated hardware or sensors in device authentication between distribution network terminal equipment and device configuration.
[0027] (2) The present invention utilizes the spatial distance characteristics between the speaker and microphone of the Internet of Things device as an added factor for device authentication, which can resist replay attacks and injection attacks and improve the security of device authentication. Attached Figure Description
[0028] Figure 1 This is a flowchart of a method according to an embodiment of the present invention.
[0029] Figure 2 This is a schematic diagram illustrating the interaction between power grid terminal equipment and Internet of Things (IoT) devices.
[0030] Figure 3 This is a schematic diagram of the sound propagation path in IoT devices. Detailed Implementation
[0031] The present invention will be further described below with reference to the accompanying drawings and specific preferred embodiments, but this does not limit the scope of protection of the present invention.
[0032] Example 1
[0033] With the widespread use of audio communication, many existing IoT devices have built-in speakers and microphones. During manufacturing, due to variations in hardware production processes and other uncontrollable factors, the analog circuitry of these components exhibits unique defects and characteristics. Existing research shows that these defects make each microphone and speaker unique. Furthermore, the spatial distance between the speaker and microphone in an IoT device is fixed. By recording the audio response of the speaker-microphone system and performing power spectrum calculations on the recorded audio stream to obtain audio characteristics and speaker-microphone spatial distance characteristics, these acoustic properties can be used to detect manufacturing defects and spatial distance variations in the device's built-in microphone-speaker system to distinguish legitimate devices and extract acoustic fingerprints for device authentication.
[0034] Based on the aforementioned characteristics of IoT devices, this embodiment proposes a multi-factor IoT device authentication method, which can achieve universal, secure, convenient, and efficient verification of the legitimacy of terminal devices and IoT devices in a power distribution network environment, including the following steps:
[0035] During the registration phase, such as Figure 1 As shown, it includes the following steps:
[0036] S101) Play the registration audio through the speaker of the IoT device and collect the played registration audio through the microphone of the IoT device;
[0037] S102) Perform signal processing on the registered audio collected by each microphone to obtain the corresponding FBank features and MFCC features;
[0038] S103) Calculate the spatial distance features between the speaker and microphone of the IoT device based on the registered audio collected by each microphone, and use the spatial distance features and the FBank features and MFCC features corresponding to the registered audio collected by each microphone as sound fingerprint training authentication model.
[0039] During the certification phase, such as Figure 1 As shown, it includes the following steps:
[0040] S201) Play authentication audio through the speaker of the IoT device and capture the played authentication audio through the microphone of the IoT device;
[0041] S202) Perform signal processing on the authentication audio collected by each microphone to obtain the corresponding FBank features and MFCC features;
[0042] S203) Calculate the spatial distance features between the speaker and microphone of the IoT device based on the authentication audio collected by each microphone, and use the spatial distance features and the FBank and MFCC features corresponding to the authentication audio collected by each microphone as sound fingerprint inputs to the trained authentication model to obtain the authentication result.
[0043] The following is a detailed explanation of each step.
[0044] The method in this embodiment is applied to Figure 2 The system shown includes an IoT device and a verification server. In step S101, the IoT device sends a registration request and receives an instruction from the verification server, generates an audio for registration, plays the registration audio using its built-in speaker, and simultaneously uses its built-in microphone to capture the played registration audio as a response audio and store it in the device's storage unit before sending it to the verification server.
[0045] In this embodiment, the built-in microphone is used to sense and collect the response audio. The obtained result is a mixed sound signal containing ambient sound and response audio. Therefore, before playing the registration audio through the speaker of the IoT device, the ambient sound is also collected through the microphone of the IoT device to facilitate the noise reduction operation in subsequent steps.
[0046] In step S102 of this embodiment, the verification server performs signal processing on the acquired response audio, removes noise, and calculates FBank and MFCC features. Specifically, the acquired mixed sound signal is divided into noise and signal, the MMSE algorithm is used for sound denoising, the denoised sound signal is pre-emphasized, framed, and windowed, and then subjected to Fast Fourier Transform to obtain the frequency domain signal. The frequency domain spectrum is obtained, a filter bank is used on the frequency domain spectrum to obtain the Mel spectrum, and the FBank and MFCC features corresponding to each microphone are calculated. Therefore, the signal processing of the registered audio acquired by each microphone includes the following steps:
[0047] (S1021) Noise reduction is performed on the sound signal currently acquired by the microphone. Specifically, the sound signal is divided into noise and signal parts according to its time length. Then, the MMSE algorithm is used to reduce noise in the signal part based on the noise part. The MMSE algorithm is an algorithm that improves signal quality and clarity by estimating the statistical characteristics of the signal. Noise is estimated for the noise part, and the statistical characteristics of the noise are used to reduce noise in the signal part, thereby reducing the influence of noise and obtaining a denoised signal. The noise reduction principle of the MMSE algorithm is well known to those skilled in the art and will not be described in detail in this embodiment.
[0048] (S1022) The denoised audio signal is pre-emphasized, framed, and windowed, then subjected to a Fast Fourier Transform to obtain the frequency domain signal. The frequency domain spectrum is obtained, and a filter bank is used to obtain the Mel spectrum. Specifically, the frequency domain spectrum is passed through a set of Mel-scale triangular filters to smooth the spectrum, eliminate harmonics, and highlight the original speech formants. Using the output of each filter within its frequency range as weights, the corresponding energies at corresponding frequencies in the frequency domain spectrum are weighted and summed to obtain the Mel spectrum, as shown in the following expression:
[0049]
[0050] in, Describes the first triangular filter in a set of Mel-scaled triangle filters. One filter, In the Fourier transform, the first... One point, Indicates the center frequency of the filter. It is the 1st Fourier transform The output obtained by passing each point through the filter In the Fourier transform, the first... The energy of each point. After this calculation, for one frame, we will get... One output.
[0051] S1023) Perform pairwise calculations on the obtained Mel spectrum to obtain the Mel filter bank coefficients, i.e., the FBank characteristics, as shown in the following expression:
[0052]
[0053] in, Describes the first triangular filter in a set of Mel-scaled triangle filters. One filter, Indicates the first Frame data, Indicates the first The first frame of data List, Indicates the first Frame data after the first Mel spectrum values obtained from the filters.
[0054] S1024) Performing Discrete Cosine Transform (DCT) on the FBank features yields the Mel-frequency cepstral coefficients, i.e., the MFCC features, expressed as follows:
[0055]
[0056] in, Describes the first triangular filter in a set of Mel-scaled triangle filters. One filter, This represents the total number of a set of Mel-scale triangular filters. Indicates the first Frame data, Indicates the first The first frame of data List.
[0057] Mel filter bank coefficients (FBank) and Mel cepstral coefficients (MFCC) are two parameters extracted in the Mel-scale frequency domain. The Mel scale describes the nonlinear characteristics of human ear frequencies, and compared to the normal frequency mechanism, Mel values are closer to the human ear's auditory mechanism. Their relationship with frequency can be approximated by the following formula:
[0058]
[0059] in, The frequency of the sound signal is expressed in Hz.
[0060] In step S103 of this embodiment, when calculating the spatial distance characteristics of the speaker and microphone of the IoT device based on the registered audio collected by each microphone, the spatial distance characteristics are obtained by using mathematical methods to extract parameters related to spatial distance based on the sound propagation model of the noise-reduced sound signal constructed.
[0061] Specifically, when playing audio using the built-in speaker and simultaneously capturing audio using two built-in microphones, the sound propagation paths of the two channels are as follows: Figure 3 As shown, the following model is constructed:
[0062]
[0063]
[0064] in and These represent the audio signals from the first and second audio channels, respectively. This indicates the authentication audio being played. This indicates sound distortion produced by the speaker. and This indicates the sound distortion produced by the first and second microphones. It is distance and frequency The propagation of sound waves follows a power-law decay function, therefore This indicates the attenuation that occurs during sound propagation. and This indicates the spatial distance between the speaker and the two microphones.
[0065] From the above sound propagation expression, we can obtain the coefficients related to the spatial distance between the speaker and microphone of the IoT device, that is... The expression is as follows:
[0066]
[0067] in and This indicates the sound distortion produced by the first and second microphones. and This indicates the spatial distance between the speaker and the two microphones.
[0068] In step S103 of this embodiment, when using the spatial distance features and the FBank and MFCC features corresponding to the registered audio collected by each microphone as sound fingerprints to train the authentication model, specifically, the sound fingerprints are used as sample features to train an unsupervised learning vector machine to obtain a classifier model as the authentication model. How to use sample features to train an unsupervised learning vector machine is well known to those skilled in the art, and will not be described in detail in this embodiment.
[0069] It should be noted that due to the influence of sound echoes, the collected sound signals will have slight differences under different conditions. To avoid the influence of sound echoes, this embodiment collects response audio (i.e., registration audio) multiple times during the registration phase. Then, through steps S102 and S103, the corresponding FBank features, MFCC features, and spatial distance features are obtained. This allows for the use of a large number of samples to train the unsupervised vector machine, enhancing the model's recognition ability and reducing errors caused by sound echoes.
[0070] Step S201 in this embodiment is similar to step S101. After the IoT device sends an authentication request and receives an instruction from the verification server, it generates audio for authentication and plays the authentication audio using its built-in speaker. Simultaneously, it uses its built-in microphone to capture the played authentication audio as response audio and stores it in the device's storage unit before sending it to the verification server. Furthermore, before playing the authentication audio through the IoT device's speaker, it also captures ambient sound through the IoT device's microphone to facilitate noise reduction operations in subsequent steps.
[0071] Step S202 in this embodiment is the same as step S102, which involves signal processing of the acquired response audio. The acquired mixed sound signal is divided into noise and signal, and the MMSE algorithm is used for sound denoising. The denoised sound signal is then pre-emphasized, framed, and windowed, and then subjected to Fast Fourier Transform to obtain the frequency domain signal. The frequency domain spectrum is obtained, and a filter bank is used on the frequency domain spectrum to obtain the Mel spectrum. The FBank features and MFCC features of the two microphone channels are then calculated. Further details are omitted here.
[0072] In this embodiment, step S203, which calculates the spatial distance characteristics of the IoT device's speaker and microphone based on the authentication audio collected by each microphone, is the same as step S102, which calculates the spatial distance characteristics of the IoT device's speaker and microphone based on the registration audio collected by each microphone. It also uses a two-channel sound propagation model to calculate and extract parameters related to the spatial distance between the speaker and microphone to obtain the spatial distance characteristics of the device's speaker and microphone. Further details will not be elaborated here.
[0073] In this embodiment, the trained authentication model in step S203 is the classifier model obtained by training the unsupervised vector machine with a large number of samples in step S103. When the spatial distance features and the FBank and MFCC features corresponding to the authentication audio collected by each microphone are used as sound fingerprints input into the trained authentication model, the three features are integrated into a device sound fingerprint, which is input into the trained classifier model for classification, and the classification result is obtained to determine whether it is a legitimate device or an abnormal device. In this embodiment, when determining whether a device is legitimate or abnormal based on the classification result, test accuracy is used as a verification method. An accuracy threshold is set. When the test accuracy of the classifier model reaches the threshold, the authentication is passed and the device is recognized as a legitimate device; otherwise, the authentication fails.
[0074] Example 2
[0075] This embodiment proposes a multi-factor IoT device authentication system, including paired IoT devices, an authentication server, and power distribution network terminal equipment, wherein:
[0076] The IoT device is used to send a registration request or authentication request to the verification server, and after receiving the instruction from the verification server, it generates the corresponding registration audio or authentication audio, plays the registration audio or authentication audio through a speaker, and simultaneously collects the played registration audio or authentication audio through a microphone and sends it to the verification server.
[0077] The verification server, upon receiving a registration request from an IoT device, sends a corresponding registration audio generation instruction to the IoT device. This instruction plays the registration audio through the IoT device's speaker and collects the played registration audio through the IoT device's microphone. The server then acquires the registration audio collected by each microphone of the IoT device and performs signal processing to obtain the corresponding FBank and MFCC features. Simultaneously, the server calculates the spatial distance features between the IoT device's speaker and microphone based on the registration audio collected by each microphone. The spatial distance features, along with the FBank and MFCC features corresponding to the registration audio collected by each microphone, are used as sound fingerprints to train the authentication model.
[0078] The verification server is also used to send a corresponding authentication audio generation instruction to the IoT device after receiving the authentication request from the IoT device, so as to play the authentication audio through the speaker of the IoT device and collect the played authentication audio through the microphone of the IoT device. Then, the authentication audio collected by each microphone of the IoT device is obtained and the signal is processed to obtain the corresponding FBank feature and MFCC feature. At the same time, the spatial distance feature between the speaker and the microphone of the IoT device is calculated based on the authentication audio collected by each microphone. The spatial distance feature and the FBank feature and MFCC feature corresponding to the authentication audio collected by each microphone are used as the sound fingerprint input to the trained authentication model to obtain the authentication result and send it to the distribution network terminal equipment.
[0079] The power distribution network terminal equipment is used to obtain and parse the authentication result. If the authentication result of the IoT device is a legitimate device, the IoT device is allowed to access; otherwise, the IoT device is blocked from accessing.
[0080] Example 3
[0081] This embodiment proposes a multi-factor IoT device authentication apparatus, including a memory, a processor, and a computer program stored in the memory. The processor executes the computer program to implement the steps of the multi-factor IoT device authentication method described in Embodiment 1.
[0082] This embodiment also proposes a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the multi-factor IoT device authentication method described in Embodiment 1.
[0083] This embodiment also proposes a computer program product, including a computer program that, when executed by a processor, implements the steps of the multi-factor IoT device authentication method described in Embodiment 1.
[0084] The above description is merely a preferred embodiment of the present invention. The scope of protection of the present invention is not limited to the above embodiments. All technical solutions falling within the scope of the present invention's concept are within the scope of protection of the present invention. It should be noted that for those skilled in the art, any improvements and modifications made without departing from the principles of the present invention should also be considered within the scope of protection of the present invention.
Claims
1. A multi-factor authentication method for Internet of Things (IoT) devices, characterized in that, Includes the following steps: The registration audio is played through the speaker of the IoT device and collected through the microphone of the IoT device. The registration audio collected by each microphone is processed to obtain the corresponding FBank features and MFCC features. At the same time, the spatial distance features between the speaker and the microphone of the IoT device are calculated based on the registration audio collected by each microphone. The spatial distance features and the FBank features and MFCC features corresponding to the registration audio collected by each microphone are used as sound fingerprints to train the authentication model. The authentication audio is played through the speaker of the IoT device and collected through the microphone of the IoT device. The authentication audio collected by each microphone is processed to obtain the corresponding FBank features and MFCC features. At the same time, the spatial distance features between the speaker and the microphone of the IoT device are calculated based on the authentication audio collected by each microphone. The spatial distance features and the FBank features and MFCC features corresponding to the authentication audio collected by each microphone are used as the sound fingerprint input to the trained authentication model to obtain the authentication result. The IoT device uses two built-in microphones to collect audio. When calculating the spatial distance characteristics between the IoT device's speaker and microphone based on the registration audio collected by each microphone, and when calculating the spatial distance characteristics between the IoT device's speaker and microphone based on the authentication audio collected by each microphone, the spatial distance characteristics are expressed as follows: in, and These represent the audio signals from the first and second audio channels, respectively. and These represent the sound distortion produced by the first and second microphones of the IoT device, respectively. and This indicates the spatial distance between the speaker and the two microphones of an IoT device. This indicates the attenuation that occurs during sound propagation. It is distance and frequency The function.
2. The multi-factor IoT device authentication method according to claim 1, characterized in that, Before playing registration audio through the speaker of the IoT device, and before playing authentication audio through the speaker of the IoT device, it also includes: collecting ambient sound through the microphone of the IoT device.
3. The multi-factor IoT device authentication method according to claim 1, characterized in that, Signal processing for both the registration audio captured by each microphone and the authentication audio captured by each microphone includes: The sound signal captured by the current microphone is denoised. The denoised sound signal is then pre-emphasized, framed, and windowed. The signal is then subjected to a fast Fourier transform to obtain the frequency domain signal. The frequency domain spectrum is obtained. The Mel spectrum is obtained by using a filter bank on the frequency domain spectrum. The obtained Mel spectrum is then paired to calculate the FBank feature. The FBank feature is then subjected to a discrete cosine transform to obtain the MFCC feature.
4. The multi-factor IoT device authentication method according to claim 1, characterized in that, When using the spatial distance features and the FBank and MFCC features corresponding to the registered audio collected by each microphone as sound fingerprints to train the authentication model, specifically, the sound fingerprints are used as sample features to train an unsupervised learning vector machine to obtain the authentication model.
5. A multi-factor Internet of Things (IoT) device authentication system, characterized in that, This includes interconnected IoT devices, testing servers, and power distribution network terminal equipment, among which: The IoT device is used to send a registration request or authentication request to the verification server, and after receiving the instruction from the verification server, it generates the corresponding registration audio or authentication audio, plays the registration audio or authentication audio through a speaker, and simultaneously collects the played registration audio or authentication audio through a microphone and sends it to the verification server. The verification server, upon receiving a registration request from an IoT device, sends a corresponding registration audio generation instruction to the IoT device. This instruction plays the registration audio through the IoT device's speaker and collects the played registration audio through the IoT device's microphone. The server then acquires the registration audio collected by each microphone of the IoT device and performs signal processing to obtain the corresponding FBank and MFCC features. Simultaneously, the server calculates the spatial distance features between the IoT device's speaker and microphone based on the registration audio collected by each microphone. The spatial distance features, along with the FBank and MFCC features corresponding to the registration audio collected by each microphone, are used as sound fingerprints to train the authentication model. The verification server is also used to send a corresponding authentication audio generation instruction to the IoT device after receiving the authentication request from the IoT device, so as to play the authentication audio through the speaker of the IoT device and collect the played authentication audio through the microphone of the IoT device. Then, the authentication audio collected by each microphone of the IoT device is obtained and the signal is processed to obtain the corresponding FBank feature and MFCC feature. At the same time, the spatial distance feature between the speaker and the microphone of the IoT device is calculated based on the authentication audio collected by each microphone. The spatial distance feature and the FBank feature and MFCC feature corresponding to the authentication audio collected by each microphone are used as the sound fingerprint input to the trained authentication model to obtain the authentication result and send it to the distribution network terminal equipment. The IoT device uses two built-in microphones to collect audio. When calculating the spatial distance characteristics between the IoT device's speaker and microphone based on the registration audio collected by each microphone, and when calculating the spatial distance characteristics between the IoT device's speaker and microphone based on the authentication audio collected by each microphone, the spatial distance characteristics are expressed as follows: in, and These represent the audio signals from the first and second audio channels, respectively. and These represent the sound distortion produced by the first and second microphones of the IoT device, respectively. and This indicates the spatial distance between the speaker and the two microphones of an IoT device. This indicates the attenuation that occurs during sound propagation. It is distance and frequency The function; The power distribution network terminal equipment is used to obtain and parse the authentication result. If the authentication result of the IoT device is a legitimate device, the IoT device is allowed to access; otherwise, the IoT device is blocked from accessing.
6. A multi-factor Internet of Things (IoT) device authentication apparatus, comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the steps of the multi-factor Internet of Things device authentication method according to any one of claims 1 to 4.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the steps of the multi-factor Internet of Things device authentication method according to any one of claims 1 to 4.
8. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program implements the steps of the multi-factor Internet of Things device authentication method according to any one of claims 1 to 4.