RFID-based voice perception method

By attaching RFID tags to ordinary glasses and combining deep learning networks and contrastive learning techniques, dynamic facial and voice features are extracted, solving the problems of device customization and high cost in existing technologies. This achieves low-cost, natural, and robust voice perception, suitable for various scenarios.

CN117194897BActive Publication Date: 2025-11-11SHANGHAI JIAOTONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311173181.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-12
Publication Date
2025-11-11
Estimated Expiration
2043-09-12

AI Technical Summary

Technical Problem

Existing voice perception technologies require customized equipment and are costly, limiting their application scenarios and making it difficult to achieve natural and robust voice perception on ordinary devices.

Method used

By attaching RFID tags that do not require a power supply to ordinary glasses, and combining conditional denoising autoencoder networks and deep learning networks to extract facial motion, skeletal vibration, and air vibration features, and using a contrastive learning network to remove user specificity, a speech perception model is constructed.

Benefits of technology

It achieves low-cost, natural, and robust voice perception on ordinary glasses, suitable for voice perception and voice assistant applications in noisy environments, with high accuracy and wide applicability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117194897B_ABST
    Figure CN117194897B_ABST
Patent Text Reader

Abstract

A voice perception method based on RFID, through a conditional denoising autoencoder network (CDAE), the collected RF signals are preprocessed to eliminate the interference of body movement, and then the preprocessed RF signals are extracted through a recurrent neural network (RNN), a convolutional recurrent neural network (CRNN) and a deep residual shrinkage network (DRSN) to obtain facial movement, skeletal vibration and air vibration features, and after fusion, the comparative learning network is input to remove the user-specificity in the facial voice dynamic feature, obtain the facial voice dynamic feature related to the sound content, and then construct a facial voice dynamic model according to the facial voice dynamic feature, and realize voice perception. The application can perceive the facial voice dynamics when people speak by pasting a RFID tag on ordinary glasses without continuous power supply and low cost, and can realize voice perception in a natural and robust manner, and has wide application scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a technology in the field of radio frequency tags, specifically a voice sensing method based on RFID. Background Technology

[0002] With the rapid development of IoT technology, voice perception has become crucial for various IoT applications. Some studies have proposed facial dynamics-based voice perception technologies. These technologies use measuring electrodes attached to the face, motion sensors on AR / VR headsets, or vibration sensors on custom glasses to sense the dynamics of a speaking face to achieve voice perception. However, these methods are either only applicable to specific scenarios and require customized equipment, or are costly and have limited application scenarios. Summary of the Invention

[0003] This invention addresses the shortcomings of existing technologies, which are cumbersome to use and limited in application scenarios. It proposes an RFID-based voice perception method that uses a low-cost RFID tag attached to ordinary glasses to perceive the dynamic facial speech of a person speaking. This method achieves voice perception in a natural and robust manner and has a wide range of applications.

[0004] This invention is achieved through the following technical solution:

[0005] This invention relates to an RFID-based voice perception method. The method preprocesses the collected RF signals using a Conditional Denoising Autoencoder Network (CDAE) to eliminate interference from body motion. The preprocessed RF signals are then processed by a Recurrent Neural Network (RNN), a Convolutional Recurrent Neural Network (CRNN), and a Deep Residual Shrinking Network (DRSN) to extract facial motion, skeletal vibration, and air vibration features. These features are then fused and input into a Contrast Learning Network to remove user-specific characteristics from the facial voice dynamic features. This process yields facial voice dynamic features that are independent of the user and only relevant to the sound content. Finally, a facial voice dynamic model is constructed based on these features to achieve voice perception.

[0006] The preprocessing refers to generating the RF signal spectrum of a person speaking without body movement interference by using a conditional denoising autoencoder network based on the RF signal spectrum x corresponding to the person speaking with body movement. The generation process uses the RF signal spectrum collected without body movement interference (Ground truth) as a reference standard to remove the interference caused by human body movement on the RF signal, that is, the interference caused by other body movements when a person speaks on the RF signal perception of facial speech dynamics.

[0007] The conditional denoising autoencoder network comprises an encoder and a decoder, wherein: the encoder performs feature encoding on the signal spectrum x with body motion interference to obtain the latent representation z of the spectrum; the decoder performs feature decoding on the latent representation z and the signal spectrum without body motion interference (ground truth) to obtain the output spectrum.

[0008] The user specificity refers to the uniqueness of each user's facial and voice dynamics when they say the same content.

[0009] This invention relates to a system for implementing the above method, comprising: a recurrent neural network unit, a convolutional recurrent neural network unit, and a deep residual shrinking network unit, wherein: the recurrent neural network unit extracts features based on signals corresponding to facial movements to obtain facial movement features; the convolutional recurrent neural network unit extracts features based on signals corresponding to skeletal vibrations to obtain skeletal vibration features; and the deep residual shrinking network unit extracts features based on signals corresponding to air vibrations to obtain air vibration features.

[0010] Technical effect

[0011] Compared to existing technologies, this invention uses a thin, low-cost RFID tag attached to ordinary glasses to sense facial speech dynamics during speech, achieving a natural and robust speech perception. In system construction, this invention designs deep learning networks to extract facial movements, skeletal vibrations, and air vibrations during speech. Then, an adaptive feature fusion method is used to fuse these three features, and user-specific features are removed to construct a speech perception model, ultimately achieving accurate and robust speech perception. The facial speech dynamics-based speech perception technology developed in this invention requires no customized equipment, is low-cost, and has wide applicability. Furthermore, by capturing subtle facial speech dynamics, this invention can be applied to speech perception in noisy environments, voice assistants, and other applications. Attached Figure Description

[0012] Figure 1 This is a flowchart of the present invention;

[0013] Figure 2 This is a schematic diagram of an RF signal propagation model.

[0014] Figure 3 Phase of the RF signal when the user speaks while wearing glasses;

[0015] Figure 4 For conditional denoising autoencoder network architecture;

[0016] Figure 5 This is a schematic diagram of the feature extraction network architecture based on RNN, CRNN, and DRSN;

[0017] Figure 6 A schematic diagram of a user-specific removal architecture based on contrastive learning;

[0018] Figure 7 A schematic diagram of the confusion matrix for digital perception;

[0019] Figure 8 This is a schematic diagram of the confusion matrix for hot word perception. Detailed Implementation

[0020] like Figure 1 As shown, this embodiment relates to an RFID-based voice perception method. By attaching an RFID tag to glasses, the RFID tag can detect subtle facial voice dynamics when the glasses are in direct contact with the face and the user speaks. Since facial voice dynamics are closely related to the content of the speech, voice perception can be achieved by using the RFID tag attached to the user's glasses to detect facial voice dynamics. The specific steps include:

[0021] Step 1) Data collection: such as Figure 2 As shown, when the RF signal transmitting antenna and receiving antenna are located in the same position, the signal transmitting antenna T x First, an electromagnetic signal is emitted to activate the RFID tag affixed to the glasses. Once activated, the tag reflects and scatters the RF signal back to the signal receiving antenna R. x When the user speaks, the eyeglass frame rotates around its dynamic center O, meaning the position of the frame in contact with the ear changes, and correspondingly, the label moves from position A to position B. Therefore, the distance difference in the signal propagation path from position A to position B can be approximately represented as |l o -l|≈Δd=r(1-cos(2πft)), where: l is the distance from the RFID antenna to location A, l0 is the distance from the RFID antenna to location B, Δd is the change in signal propagation distance as the tag moves from location A to location B, r is the length between the dynamic center O and the RFID tag, t is time, and f is the frequency of the facial voice dynamics. Due to the backscattering characteristics of the RF signal, the actual signal transmission distance from the antenna to point B can be described as 2(l+Δd). Therefore, the phase of the RF signal at B can be approximately expressed as... Where λ is the wavelength of the RF signal. Combining the above equations, we obtain... because Since f is a constant, the RF signal caused by facial speech dynamics satisfies a cosine property that varies with time. Therefore, when a user speaks, the frequency f of the facial speech dynamics can be inferred from the phase θ(t) of the received RF signal, which can be used for live speech perception.

[0022] Because the wavelengths of RF signal harmonics are shorter than the fundamental frequency, they are more sensitive to tag displacement. Therefore, the USRP N210 was chosen to monitor the third harmonic signal at a frequency of 2761.89MHz, and the phase signal was extracted at a sampling frequency of 2MHz to improve the accuracy of voice perception. After signal amplification, the RF signals collected when the user wears glasses with an RFID tag and speaks different speech content are as follows: Figure 3 As shown in the figure, significant phase fluctuations in the RF signal can be observed when the user says the words "one," "two," "three," "hello," and "excellent." These results indicate that the RF signal is highly sensitive to human facial speech dynamics.

[0023] Step 2) Signal Preprocessing: Besides facial speech dynamics, human motion can also be perceived by RF signals, which affects the perception of facial speech dynamics. To achieve accurate speech perception, it is necessary to remove human body motion interference from the received RF signals. This embodiment designs a conditional denoising autoencoder (CDAE) based on signal generation technology to generate an RF signal spectrogram without body motion interference from the collected RF signals. Specifically, when the user speaks (without body movement), the spectrogram of RF signals including facial movement, bone conduction vibration, and air conduction vibration is collected as the ground truth for correction. The spectrogram of RF signals with body movement is collected as input to the CDAE, including the signal responses of body movement, facial movement, bone vibration, and air vibration. In the CDAE, the spectrogram with body movement is processed by an encoder and decoder to generate a new spectrogram including the signal responses of facial movement, bone vibration, and air vibration. The generation process aims at the ground truth spectrogram without body movement. Therefore, CDAE can eliminate the effects of various body movements, including small body movements (such as nodding / shaking the head) and large body movements (such as walking / gesturing).

[0024] Because the impact of RF signals on facial speech dynamics (such as facial movements, skeletal vibrations, and air vibrations) varies depending on the speech content, this embodiment uses a self-attention mechanism for feature fusion. This mechanism can fuse three types of facial speech dynamic features using different adaptive weights for different speech content. Specifically, this embodiment first uses filters and source separation techniques to separate the signals corresponding to facial movements, skeletal vibrations, and air vibrations. Then, networks based on RNN, CRNN, and DRSN are used to extract facial movement, skeletal vibration, and air vibration features, respectively.

[0025] Step 3) Facial Motion Feature Extraction: For facial motion, human speech behavior drives facial movements, causing facial muscles to contract and relax randomly during speech. Facial muscles encode speech features (such as phonemes, rhythm, and volume) and biometric features (such as speech behavior, facial morphological changes, and tissue movement). Since speech and biometric features are time-varying, this embodiment uses a recurrent neural network (RNN) model to extract features corresponding to facial motion.

[0026] like Figure 5 As shown in (a), the recurrent neural network includes: 10 identical RNN models, each RNN model containing several interconnected GRU units, wherein: each RNN model extracts features based on each temporal statistical feature to obtain one facial motion feature vector, specifically including: calculating 10 temporal statistical features (maximum, minimum, average, variance, range, root mean square, median, quartiles, skewness, and kurtosis) corresponding to the phase of the RF signal corresponding to the facial motion to describe the facial motion characteristics; for each statistical feature, the one RNN model is used to extract the temporal features of the facial motion; a GRU-based RNN model is used as the basis for temporal feature extraction, and after each temporal statistical feature is processed by each RNN model, a facial motion feature vector is output; finally, all 10 feature vectors are combined into a feature matrix, representing the facial motion features extracted when the user speaks.

[0027] Step 4) Bone vibration feature extraction: Since bone vibration exists in a high-frequency form, this embodiment uses a CRNN network model to extract the time-frequency features of bone vibration.

[0028] like Figure 5 As shown in (b), the CRNN model includes: a CNN structure for extracting the frequency features of bone vibration and a bidirectional LSTM structure for extracting the time-domain features of bone vibration, wherein: the CNN structure obtains the bone vibration feature sequence representing the frequency domain features of bone vibration based on the spectrum of the RF signal corresponding to bone vibration; the bidirectional LSTM structure extracts the time-domain features of bone vibration and integrates all the extracted feature vectors into a feature matrix to represent the time-frequency features extracted from the bone vibration signal.

[0029] The CNN structure includes: a batch normalization layer, three convolutional layers, and three max pooling layers. The batch normalization layer is used to accelerate training and prevent overfitting. The convolutional layers are used to extract high-frequency skeletal vibration feature contours. Each convolutional layer is supplemented with a max pooling layer for dimensionality reduction.

[0030] The bidirectional LSTM model includes two bidirectional LSTM units, which extract the temporal features of the bone vibration signal from two opposite directions.

[0031] Step 5) Air vibration feature extraction: For air vibration, this embodiment uses a DRSN network to extract air vibration features from the RF signal spectrum including wind noise.

[0032] like Figure 5 As shown in (c), the DRSN network architecture includes two convolutional-pooling units and three ResNet blocks. Each convolutional-pooling unit comprises one convolutional layer and one pooling layer. When the spectrogram of the RF signal with wind noise is input into the convolutional-pooling unit, the relationship between time and frequency is learned. Due to the large size of the spectrogram, a max-pooling layer is added to each convolutional layer of the convolutional-pooling unit to reduce dimensionality. The output of the convolutional-pooling unit is a feature representation of the spectrogram with wind noise. This feature representation is then input into the three ResNet blocks for wind noise removal and feature extraction. Finally, a feature matrix is ​​obtained, representing the air vibration features extracted after wind noise removal.

[0033] Step 6) Feature Fusion: The three facial speech dynamic features extracted above reflect human facial movement, skeletal vibration, and air vibration characteristics, respectively. To comprehensively describe human speech dynamic features, this embodiment fuses these three facial speech dynamic features to construct a speech perception model. The impact of these three facial speech dynamic features differs for different speech content. Therefore, this embodiment uses a self-attention mechanism for feature fusion, which utilizes different adaptive weights to fuse the three facial speech dynamic features for different speech content.

[0034] Step 7) User-specific removal: Since human facial speech dynamics include user-specific features such as speech behavior, facial shape and tissue characteristics, this embodiment uses a contrastive learning network to remove user-specific features from the extracted facial speech dynamic features, obtaining features that are only related to the speech content.

[0035] like Figure 6 As shown, the contrastive learning network includes an encoder and a projector, wherein: the encoder f θ (·) Obtain each input word pair x i and x j Feature embedding h i and h j After that, the projector g θ (·) Based on the feature embedding, the corresponding feature representation z is obtained. Then, a contrastive loss function is used to guide the facial speech dynamic features of the same word to attract each other in the feature space, while the features of different speech words repel each other. Therefore, the extracted features have intra-class compactness and inter-class discriminability, x iIt is the input to the contrastive learning network, which is a set of RF signal features from the same spoken word by different users.

[0036] When the contrastive loss function converges, the resulting feature embeddings are user-independent and only related to the speech content.

[0037] Step 8) Speech Perception Model Construction. Based on the obtained feature embeddings, a facial speech dynamic model is further trained using the cross-entropy loss function to achieve speech perception. The cross-entropy loss function is described as follows: Where p(x) i q(x) represents the distribution of the actual labels. i ) represents the distribution of the predicted output, and c is the number of sample categories. Based on the cross-entropy loss function described above, the model is iterated until the loss function converges, thus constructing the speech perception model.

[0038] Step 9) Speech Perception Implementation: During speech perception implementation, the constructed speech perception model and the corresponding RF signals from the speaker are used to infer the possible speech content. Specifically, if k words have been trained in the model, then the constructed speech perception model will have k templates corresponding to these k words, i.e., w1, w2, ..., w... k In the process of speech perception, when the subject utters the word "w", the features of the word "w" are first extracted from the received RF signal. i Then calculate w respectively i The Euclidean distance between the word and k templates in the speech perception model is calculated. Based on these k distances, the word w with the smallest distance is the perceived word. Finally, all the perceived words are connected to form the complete speech perception content.

[0039] In this practical experiment, the Impinj Speedway R420 commercial RFID reader was used to transmit RF signals, while a USRP N210 with an SBX daughterboard was used to capture harmonic RF signals. The COTS reader and USRP device used antennas with a gain of 15dBi operating at 920MHz and a gain of 16dBi operating at 2.4GHz, respectively. To sense the dynamics of human speech, a passive RFID tag was affixed to an eyeglass frame. The Impinj RFID reader was set to query the RFID tag at a carrier frequency of 920.63MHz, while the USRP device listened for harmonic echo signals at a sampling frequency of 2761.89MHz at 2MHz. The RF harmonic signals were captured and preprocessed by a computer. Twelve volunteers wearing glasses were recruited to participate in the experiment, and a selection of numbers and hot words was chosen to evaluate the performance of the invention. Each time data was collected, volunteers were required to read a series of paragraphs containing the aforementioned numbers and hot words according to their own habits, with each paragraph being read 10 times. The performance of the invention in sensing the selected numbers and hot words was then evaluated.

[0040] like Figure 7 The image shows the confusion matrix for digit perception, where each element represents the overall accuracy of digit perception across all volunteers. Each row of the confusion matrix represents the recognition result, and each column represents the true result. The i-th element of the matrix... th row and j th The elements of a column describe the probability that a sample is identified as class i but is actually class j, i.e., the accuracy. As can be observed from the figure, this invention achieves an average accuracy of 91.23% for digital perception, with all digital perception accuracies exceeding 86%.

[0041] like Figure 8 The figure shows the confusion matrix for hot word perception, where each element represents the overall accuracy of hot word perception across all volunteers. The figure illustrates that the present invention achieves an average accuracy of 89.32% in hot word perception, with all hot word perception accuracies exceeding 85%. These results demonstrate that the present invention can effectively perceive words (i.e., numbers and hot words).

[0042] Compared to existing technologies, this invention only requires attaching a thin, low-cost RFID tag to ordinary glasses to sense facial and vocal dynamics during speech, achieving accurate, natural, and robust voice perception. This method does not require costly and complex customized equipment, thus having a wider range of applications. This is something that existing inventions have never been able to achieve, representing a valuable supplement and significant improvement to existing voice perception methods.

[0043] The above-described specific implementations can be partially adjusted by those skilled in the art in different ways without departing from the principles and purpose of the present invention. The scope of protection of the present invention is defined by the claims and is not limited to the above-described specific implementations. All implementation schemes within the scope of the claims are bound by the present invention.

Claims

1. A voice perception method based on RFID, characterized in that, After preprocessing the collected RF signals to eliminate body motion interference using a Conditional Denoising Autoencoder Network (CDAE), the preprocessed RF signals are then processed by a Recurrent Neural Network (RNN), a Convolutional Recurrent Neural Network (CRNN), and a Deep Residual Shrinking Network (DRSN) to extract facial motion, skeletal vibration, and air vibration features. These features are then fused and input into a Contrast Learning Network to remove user-specific features from the facial speech dynamic features. After obtaining facial speech dynamic features that are independent of the user and only related to the sound content, a facial speech dynamic model is constructed based on these features to achieve speech perception. The recurrent neural network includes: 10 identical RNN models, each RNN model containing several interconnected GRU units, wherein: each RNN model extracts features based on each temporal statistical feature to obtain a facial motion feature vector, specifically including: calculating 10 temporal statistical features corresponding to the phase of the RF signal corresponding to the facial motion to describe the facial motion characteristics. The time-domain statistical features include: maximum value, minimum value, average value, variance, range, root mean square, median, quartiles, skewness, and kurtosis; For each statistical feature, a separate RNN model is used to extract the temporal features of facial motion. A GRU-based RNN model is used as the basis for temporal feature extraction. After each temporal statistical feature is processed by each RNN model, a facial motion feature vector is output. Finally, all 10 feature vectors are combined into a feature matrix, representing the facial motion features extracted when the user speaks. The CRNN model includes: a CNN structure for extracting frequency features of bone vibration and a bidirectional LSTM structure for extracting time-domain features of bone vibration. The CNN structure obtains a bone vibration feature sequence representing the frequency domain features of bone vibration based on the spectrum of the RF signal corresponding to bone vibration. The bidirectional LSTM structure extracts the time-domain features of bone vibration and integrates all extracted feature vectors into a feature matrix, representing the time-frequency features extracted from the bone vibration signal. The DRSN network architecture includes two convolutional-pooling units and three ResNet blocks. The convolutional-pooling units receive the spectrum of the RF signal with wind noise and learn the relationship between time and frequency, then output a feature representation of the spectrum with wind noise. The three ResNet blocks perform wind noise removal and feature extraction based on the feature representation, respectively, to obtain a feature matrix of air vibration features extracted after removing wind noise. Each convolutional-pooling unit consists of a convolutional layer and a max-pooling layer; The contrastive learning network includes an encoder and a projector, wherein the encoder... Get the input word pairs respectively and Feature embedding and Then, the projector Obtain the corresponding feature representation based on feature embedding. Then, the contrastive loss function is used to guide the facial speech dynamic features of the same word to attract each other in the feature space, while the features of different speech words repel each other. The CNN structure includes: a batch normalization layer, three convolutional layers and three max pooling layers. The batch normalization layer is used to accelerate training and prevent overfitting. The convolutional layers are used to extract high-frequency skeletal vibration feature contours. Each convolutional layer is supplemented with a max pooling layer for dimensionality reduction. The bidirectional LSTM model includes two bidirectional LSTM units, which extract the temporal features of the bone vibration signal from two opposite directions.

2. The RFID-based voice perception method according to claim 1, characterized in that, The preprocessing refers to: using a conditional denoising autoencoder network to analyze the RF signal spectrum corresponding to human speech accompanied by body movements. Generated RF signal spectrum when a human is speaking without any body movement interference. The generation process uses the RF signal spectrum (Ground truth) collected when there is no body movement interference as a reference standard, thereby removing the interference caused by human body movement on the RF signal, that is, the interference caused by other body movements when a person speaks on the RF signal perception of facial speech dynamics.

3. The RFID-based voice perception method according to claim 1, characterized in that, The conditional denoising autoencoder network includes an encoder and a decoder, wherein the encoder bases the signal spectrum with body motion interference. Feature encoding is performed to obtain the latent representation of the spectrum. The decoder is based on the latent representation. Feature decoding is performed on the signal spectrum without body motion interference (ground truth) to obtain the output spectrum. ; The user specificity refers to the uniqueness of each user's facial and voice dynamics when they say the same content.

4. The RFID-based voice perception method according to any one of claims 1-3, characterized in that, specifically include: Step 1) Data Collection: When the RF signal transmitting antenna and receiving antenna are located in the same position, the signal transmitting antenna... First, an electromagnetic signal is emitted to activate the RFID tag attached to the glasses. Once activated, the tag reflects and scatters the RF signal back to the signal receiving antenna. ; When the user speaks, the glasses frame revolves around the dynamic center. Rotation, meaning the position of the eyeglass frame in contact with the ear changes, correspondingly moves the tag from position A to position B; the distance difference in the signal propagation path from position A to position B is approximately expressed as... in: Let A be the distance from the RFID antenna to location A. The distance from the RFID antenna to point B. This represents the change in signal propagation distance as the tag moves from position A to position B. For dynamic center The length between the RFID tag and the RFID tag For time, The frequency of facial voice dynamics; the actual signal transmission distance from the antenna to point B is... RF signal phase at point B in: The wavelength of the RF signal is used to obtain... because It is a constant; that is, the phase received from the RF signal when the user speaks. Frequency of facial speech dynamics inferred Used for live speech perception; Step 2) Signal preprocessing, specifically including: 2.1) A conditional denoising autoencoder (CDAE) based on signal generation technology is used to generate an RF signal spectrum map free from body motion interference from the collected RF signals. When the user is not involved in speaking with body movements, the spectrum map of RF signals including facial movements, bone conduction vibrations, and air conduction vibrations is collected as the correction basis. The encoder and decoder of the conditional denoising autoencoder process the spectrum map with body movements, and a new spectrum map is generated with the correction basis spectrum map without body movements as the target. This spectrum map includes the signal responses of facial movements, bone vibrations, and air vibrations to eliminate the influence of various body movements. 2.2) The signals corresponding to facial motion, skeletal vibration and air vibration are separated using filters and source separation techniques. Then, the features of facial motion, skeletal vibration and air vibration are extracted using RNN, CRNN and DRSN-based networks, respectively. Step 3) Facial motion feature extraction: Use a recurrent neural network (RNN) model to extract features corresponding to facial motion; Step 4) Bone vibration feature extraction: Extract the time-frequency features of bone vibration using a CRNN network model; Step 5) Air vibration feature extraction: Use a DRSN network to extract air vibration features from the RF signal spectrum including wind noise; Step 6) Feature Fusion: Use a self-attention mechanism to fuse the above three facial and speech dynamic features to construct a speech perception model; Step 7) User-specific removal: A contrastive learning network is used to remove user-specific features from the extracted facial speech dynamic features, obtaining features that are only related to the speech content; Step 8) Speech Perception Model Construction: Based on the obtained feature embeddings, a facial speech dynamic model is trained using the cross-entropy loss function to achieve speech perception. The cross-entropy loss function is described as follows: ,in It is the actual distribution of labels. It is the distribution of the predicted output. This is the number of sample categories. Based on the cross-entropy loss function mentioned above, the model is iterated until the loss function converges, thus constructing the speech perception model. Step 9) Speech perception implementation: Use the constructed speech perception model and the corresponding RF signal when the subject speaks to infer the possible speech content, assuming the model has already been trained. If there are 100 words, then the constructed speech perception model will contain 100 words. This template and this Each word corresponds to one of the words, that is In the process of speech perception, when the person being perceived says a word... At that time, the words are first extracted from the received RF signal. Features Then calculate separately In speech perception models The Euclidean distance between the templates is based on this. Given a distance, find the word with the smallest distance. These are the perceived words. Finally, connecting all the perceived words together gives the complete speech perception content.

5. A system for implementing the RFID-based voice perception method according to any one of claims 1-4, characterized in that, include: The system comprises a recurrent neural network unit, a convolutional recurrent neural network unit, and a deep residual shrinking network unit. Specifically: the recurrent neural network unit extracts facial motion features based on signals corresponding to facial motion; the convolutional recurrent neural network unit extracts skeletal vibration features based on signals corresponding to skeletal vibration; and the deep residual shrinking network unit extracts air vibration features based on signals corresponding to air vibration.

Citation Information

Patent Citations

  • Identity authentication method fusing user multi-source sound production characteristics, storage medium and equipment

    CN112116742A

  • Robust speaker recognition method based on spectrogram denoising and adversarial learning

    CN116469394A