Cross-modal silent speech reconstruction method and system based on ear canal air pressure micro-motion perception

By using ear canal air pressure micro-motion sensing technology, the problem of reconstructing high-frequency acoustic signals from low-frequency mechanical signals in silent voice interaction has been solved. This has enabled high-fidelity voice reconstruction and recognition, improved recognition accuracy and privacy protection, reduced power consumption, enhanced environmental robustness and user adaptability, and supported real-time interaction.

CN121687060BActive Publication Date: 2026-06-23DONGHUA UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
DONGHUA UNIV
Filing Date
2026-02-12
Publication Date
2026-06-23

Smart Images

  • Figure CN121687060B_ABST
    Figure CN121687060B_ABST
Patent Text Reader

Abstract

The application discloses a cross-modal silent speech reconstruction method and system based on ear canal air pressure micro-motion sensing, and belongs to the technical field of human-computer interaction and wearable computing. The method uses a micro-pressure sensing unit placed in an in-ear earphone to collect a non-acoustic air pressure sequence caused by the movement of a sound-producing organ; through adaptive baseline drift suppression and rhythm perception data enhancement processing, a robust feature space is constructed; further, an end-to-end deep neural network containing domain adversarial adaptation, cross-modal semantic alignment, coarse-grained mel-spectrogram generation and residual detail correction is used to map the TPVS to a high-fidelity acoustic mel spectrum. The application effectively breaks through the technical bottleneck of the lack of high-frequency acoustic features in low-frequency mechanical signals, realizes high-precision silent speech command analysis in a mobile and noisy scene, and introduces a coupled quality evaluation gate and trigger-based start / stop control at the inference end to suppress invalid inference and reduce power consumption when wearing is poor or there is no trigger condition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of human-computer interaction and wearable computing technology, and in particular to a cross-modal silent speech reconstruction method and system based on the sensing of ear canal air pressure micro-motion. Background Technology

[0002] With the rapid development of the Internet of Things (IoT) and wearable computing technologies, smart headphones have gradually become the next generation of core interactive terminals after smartphones. Traditional voice interaction relies heavily on sound wave signals, which presents an irreconcilable contradiction in practical applications: on the one hand, in quiet environments such as libraries and conference rooms, spoken commands can cause social interference and leak user privacy; on the other hand, in high-noise environments such as subways and factories, environmental noise can severely reduce the accuracy of speech recognition. To overcome the limitations of acoustic signals, Silent Speech Interfaces (SSIs) have emerged. Their core concept is to decode user intent by capturing non-acoustic motion signals from vocal organs (such as lips, tongue, and jaw).

[0003] Currently, the mainstream SSI technology solutions mainly include the following three categories, but all of them have significant technical bottlenecks:

[0004] The first category is computer vision-based solutions. These solutions utilize cameras to capture minute movements of the lips or facial muscles, using lip-reading algorithms to infer speech content. However, their drawbacks are as follows: First, they are extremely sensitive to ambient lighting conditions, with recognition rates dropping sharply in low-light or backlit environments; second, analyzing lip movements via video streams requires significant computing power, making real-time operation on resource-constrained headsets difficult; and most importantly, the introduction of cameras raises serious privacy concerns for users, limiting their routine use in daily life.

[0005] The second category is based on active sensing solutions. These solutions typically utilize ultrasonic, millimeter-wave radar, or laser sensors to emit active detection signals towards the face or ear canal and receive the echoes to analyze the movement of the vocal organs. While offering high accuracy, their limitations include continuous power consumption from actively emitting modulated signals, significantly shortening the battery life of wearable devices; furthermore, the large size of radar or ultrasonic modules, when integrated into in-ear headphones, significantly increases the device's weight and size, impacting wearing comfort.

[0006] The third category comprises contact-based sensor solutions. These solutions utilize patch-type electromyography (sEMG) sensors to capture electrical signals from facial muscles, or inertial measurement units (IMUs) to capture vibrations of the mandible. However, SEMG signals are highly susceptible to changes in skin sweat and electrode contact impedance, and their characteristics vary significantly among different users, making them difficult to reuse. While IMU solutions are easy to integrate, they are extremely sensitive to minute shifts in wearing position and are easily affected by motion artifacts such as walking and head rotation, resulting in poor robustness in mobile scenarios.

[0007] In recent years, research has focused on the emerging approach of "ear canal pressure sensing." Anatomically, the human ear canal is located adjacent to the temporomandibular joint. When a user speaks (even silently), the movement of the mandible compresses the anterior wall of the ear canal, causing micrometer-level dynamic deformation of the ear canal volume, which in turn causes pressure fluctuations within the closed ear canal. Compared to the methods mentioned above, pressure sensing has inherent advantages: it is a passive sensing method with extremely low power consumption; and the sealed ear canal naturally isolates it from external acoustic noise and light interference.

[0008] However, there are significant physical gaps and technical challenges in applying non-acoustic barometric sequences (TPVS) to speech reconstruction:

[0009] The significant difference in sampling rate and frequency band: TPVS is essentially a low-frequency mechanical signal caused by organ movement, with an effective frequency band typically below 20Hz, requiring a sampling rate of only around 100Hz. Human speech, on the other hand, is a high-frequency acoustic signal, containing abundant high-frequency harmonics and formants, typically requiring a sampling rate of 16kHz or higher. Recovering a 16kHz acoustic signal from a 100Hz mechanical signal presents a huge challenge.

[0010] Lack of acoustic features: TPVS only reflects the motion envelope of the vocal organs, completely losing the fundamental frequency and fine spectral structure generated by vocal cord vibration. Existing technologies cannot conjure realistic speech spectra out of thin air without prior acoustic knowledge.

[0011] Weak signals and individual differences: The air pressure changes caused by ear canal deformation are extremely weak (usually at the Pascal level) and are easily drowned out by environmental air pressure drift (such as altitude changes); in addition, the ear canal geometry and pronunciation habits of different users are very different, making it difficult for the model to generalize among different users.

[0012] Therefore, how to break through the modal barrier between low-frequency mechanical signals and high-frequency acoustic signals, reconstruct a high-fidelity Mel spectrum from limited air pressure characteristics while removing environmental interference, and achieve robust adaptation to different users and different speech rates is a key technical problem that urgently needs to be solved in this field. Summary of the Invention

[0013] This application provides a cross-modal silent speech reconstruction method and system based on ear canal air pressure micro-motion sensing, which solves the following core technical bottlenecks in the existing silent speech interaction and wearable human-computer interaction technology fields:

[0014] 1. The physical gap in low-frequency mechanical-high-frequency acoustic cross-modal reconstruction: TPVS is essentially a low-frequency mechanical deformation signal caused by the temporomandibular joint compressing the anterior wall of the ear canal (effective bandwidth is usually below 10Hz, sampling rate is only about 100Hz), completely lacking the high-frequency harmonics and formant structures generated by vocal cord vibration. Existing technologies struggle to establish an effective nonlinear mapping from narrowband, non-acoustic mechanical signals to broadband, high-dimensional acoustic spectra; direct reconstruction often leads to speech ambiguity, low intelligibility, and loss of detail.

[0015] 2. The challenge of cross-modal semantic alignment under weak supervision: In silent speech scenarios, there is a lack of synchronized acoustic signals as "real labels," and there is no explicit linear or frame-level temporal correspondence between the topology of the barometric pressure waveform and the standard phonemes. Traditional supervised learning methods based on strong frame-level alignment fail in such weakly supervised tasks, making it difficult for the model to capture linguistic features in TPVS;

[0016] 3. Generalization Bottleneck Due to Individual Differences and Dynamic Speech Rate Variations: TPVS is a product of the coupling of vocalization movement, ear canal geometry, and wearing seal. Differences in ear canal structure among different users can lead to drastic changes in signal baseline and fluctuation amplitude, and the vocalization rate of the same user exhibits random dynamics. Traditional fixed-parameter models struggle to decouple semantic features from identity features, resulting in a sharp decline in recognition performance across users or in scenarios with varying speech rates.

[0017] 4. Deterioration of signal-to-noise ratio of weak signals in dynamic environments: Micro-barometric sensors are extremely sensitive to changes in environmental altitude (such as elevator movement) and non-vocal body movements (such as vibrations from walking). The effective deformation signal at the micrometer level is easily overwhelmed by environmental baseline drift or motion artifacts, limiting the robustness of the system in mobile scenarios.

[0018] To achieve the above objectives, the technical solution of this invention is as follows:

[0019] In a first aspect, embodiments of the present invention provide a cross-modal silent speech reconstruction method based on ear canal pressure micro-motion sensing, comprising: acquiring non-acoustic air pressure changes caused by the movement of the vocal organs through at least one air pressure sensor installed in the user's ear canal to obtain an original air pressure change time series; performing baseline drift suppression and amplitude enhancement on the original air pressure change time series to highlight articulation-related fluctuations to obtain a preprocessed sequence; determining the start and end boundaries of silent articulation events based on the temporal energy characteristics of the preprocessed sequence, extracting effective articulation segments, and performing time normalization on the effective articulation segments; inputting the effective articulation segments into a pre-trained spectrogram reconstruction model to generate a coarse-grained acoustic Mel spectrogram; the spectrogram reconstruction model is configured to extract semantic features related to the articulation content from the air pressure micro-motion signal; supplementing and correcting high-frequency details of the coarse-grained acoustic Mel spectrogram based on a residual learning mechanism to obtain a high-fidelity target Mel spectrum; inputting the target Mel spectrum into an automatic speech recognition model to output the corresponding silent speech recognition result.

[0020] In some possible implementations, baseline drift suppression includes at least one of sliding window mean subtraction, bandpass filtering, or adaptive baseline estimation; amplitude enhancement includes at least one of squaring, absolute value, logarithmic compression, or normalization of the preprocessed sequence to improve the separability of articulation-related deformation fluctuations.

[0021] In some possible implementations, the start and end boundaries of silent phonation events are determined based on the temporal energy characteristics of the preprocessed sequence, including: calculating the sliding window short-time energy curve of the preprocessed sequence; detecting the energy peak under the condition of satisfying the minimum interval constraint; and searching forward and backward with the energy peak as the center until the energy or energy fluctuation rate is lower than a preset threshold to determine the start and end boundaries of the phonation event.

[0022] In some possible implementations, time normalization involves interpolating and resampling the effective speech segments to normalize the duration or number of frames of the effective speech segments to a preset length.

[0023] In some possible implementations, the spectrogram reconstruction model includes at least an encoder and a semantic latent space mapping unit; wherein the encoder is configured to extract content representations from valid speech segments and suppress representation components related to user identity or wearing differences through a domain adversarial learning mechanism; the domain adversarial learning mechanism includes an identity discrimination branch and gradient inversion or an equivalent min-max optimization mechanism.

[0024] In some possible implementations, the semantic latent space mapping unit maps content representations to the semantic latent space through contrastive learning or matching learning; wherein, the positive sample pairs of contrastive learning include content representations and text representations corresponding to the same statement or the same semantic label.

[0025] In some possible implementations, a high-fidelity target Mel spectrum is obtained by supplementing and correcting high-frequency details of a coarse-grained acoustic Mel spectrum based on a residual learning mechanism. This includes: inputting the coarse-grained acoustic Mel spectrum into a residual network; the residual network predicting and outputting a residual spectrum corresponding to the missing high-frequency details in the coarse-grained acoustic Mel spectrum; and superimposing the residual spectrum with the coarse-grained acoustic Mel spectrum to obtain the high-fidelity target Mel spectrum. The residual network is optimized using a hybrid loss function during the training phase. This hybrid loss function includes at least one of phoneme-level CTC loss, frequency domain structure consistency loss, and prosodic correlation loss based on the signal envelope.

[0026] Secondly, embodiments of the present invention provide a training method for a spectrogram reconstruction model, used to train the spectrogram reconstruction model of the first aspect. The training method includes: acquiring a training sample set containing original barometric pressure change time series and supervision information, wherein the supervision information includes text, acoustic Mel spectra, or a combination of both; performing time scaling on the acoustic signals in the training samples while keeping the pitch constant, and performing corresponding time scaling or resampling on the original barometric pressure change time series synchronized with it to generate multi-speed training samples; jointly training an encoder with semantic classification loss and identity adversarial loss so that the encoder outputs content representations that are insensitive to user identity; applying contrastive learning constraints to the content representations and text representations to obtain semantic latent space mapping units; training a generative model to output coarse-grained Mel spectra, and training a residual learning network to output detailed residuals under the condition of freezing the generative model.

[0027] In some possible implementations, the encoder is jointly trained with semantic classification loss and identity adversarial loss so that the encoder outputs content representations that are insensitive to user identity. This includes: using a domain adversarial neural network framework so that the content representations output by the encoder are simultaneously input into the semantic classifier and the identity discriminator, and using a gradient reversal layer or an equivalent min-max optimization mechanism to achieve adversarial training of the identity discriminator.

[0028] Thirdly, embodiments of the present invention provide a cross-modal silent speech reconstruction system based on ear canal pressure micro-motion sensing, comprising: a sensing module, used to acquire non-acoustic air pressure changes caused by the movement of the vocal organs through at least one air pressure sensor installed in the user's ear canal, to obtain an original air pressure change time series; a preprocessing module, used to suppress baseline drift and enhance amplitude of the original air pressure change time series to highlight articulation-related fluctuations, to obtain a preprocessed sequence; and an event detection module, used to determine the start and end boundaries of silent articulation events based on the temporal energy characteristics of the preprocessed sequence, and to extract... The system obtains valid speech fragments and performs time normalization on them. A spectrogram reconstruction module inputs the valid speech fragments into a pre-trained spectrogram reconstruction model to generate a coarse-grained acoustic Mel spectrum. The spectrogram reconstruction model is configured to extract semantic features related to the speech content from barometric pressure fluctuation signals. A correction module supplements and corrects high-frequency details of the coarse-grained acoustic Mel spectrum based on a residual learning mechanism to obtain a high-fidelity target Mel spectrum. A decoding module inputs the target Mel spectrum into an automatic speech recognition model and outputs the corresponding silent speech recognition result.

[0029] One or more technical solutions provided in the embodiments of the present invention have at least the following technical effects or advantages:

[0030] 1. Achieved high-quality cross-modal reconstruction from low-frequency mechanical signals to high-frequency acoustic spectra: This invention utilizes a two-stage strategy of semantic alignment and residual correction to reconstruct a high-fidelity acoustic Mel spectrum containing rich harmonics and formant structures from a low-frequency ear canal pressure signal of only about 100Hz. This method overcomes the physical bottleneck of traditional techniques that cannot recover speech details due to signal bandwidth and feature loss, significantly improving the accuracy and intelligibility of silent speech recognition.

[0031] 2. Excellent privacy protection features: This invention relies on passively sensing non-acoustic air pressure changes in the ear canal without collecting any sound waves or image information. It eliminates the risk of leakage of biometric features such as voiceprints and lip shapes from the physical source, solves the inherent privacy risks of visual and acoustic solutions, and is suitable for interactive scenarios with extremely high privacy requirements.

[0032] 3. Strong environmental robustness and anti-interference capability: By utilizing the physical isolation characteristics of the quasi-sealed cavity of the ear canal, combined with adaptive baseline drift suppression and narrowband pass filtering, this invention can effectively suppress external high-decibel environmental noise, low-frequency drift caused by altitude changes, and high-frequency artifacts generated by head movement, thereby maintaining stable signal quality and recognition performance in dynamic environments such as movement and noise.

[0033] 4. Significantly reduce system power consumption and improve battery life: The passive micro-pressure sensor eliminates the need for active signal transmission. Furthermore, through trigger-based start-stop control and invalid inference suppression mechanisms, the system only operates when necessary, greatly reducing overall power consumption and enabling long-term operation in wearable devices with limited battery capacity.

[0034] 5. Possesses excellent user generalization and speech rate adaptation capabilities: By introducing rhythm-aware data enhancement during the training phase and combining it with a domain adversarial feature decoupling mechanism, the model can learn the air pressure morphology invariant under different speech rates and remove the identity features brought about by differences in user ear canal structure, thereby achieving robust adaptation to unseen users and dynamic speech rates, improving the system's practicality and accessibility.

[0035] 6. Supports low-latency, edge-deployable real-time interaction: Through model quantization, lightweight network architecture design and edge inference optimization, the entire method can achieve end-to-end processing latency of less than 200ms on mobile devices, meeting the needs of real-time voice interaction and providing a feasible solution for privacy-sensitive applications in offline environments. Attached Figure Description

[0036] To more clearly illustrate the embodiments of the present invention, the accompanying drawings used in the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0037] Figure 1 This is a schematic diagram comparing the physical characteristics of the ear canal air pressure micro-motion signal and the standard acoustic speech signal in an embodiment of the present invention;

[0038] Figure 2 A schematic flowchart of an embodiment of a cross-modal silent speech reconstruction method based on ear canal air pressure micro-motion sensing provided by the present invention;

[0039] Figure 3 This is a breakdown of the physical structure of the hardware prototype and an experimental scenario diagram in an embodiment of the present invention.

[0040] Figure 4 This is a schematic diagram of the TPVS signal waveform after bandpass filtering in an embodiment of the present invention;

[0041] Figure 5 This is a schematic diagram of the event capture results based on short-time energy peak detection in an embodiment of the present invention;

[0042] Figure 6 This is a schematic diagram of the architecture of the three-stage Mel spectrum reconstruction pipeline in an embodiment of the present invention;

[0043] Figure 7This is a network architecture diagram of the encoder in an embodiment of the present invention;

[0044] Figure 8 This is a schematic diagram of the adversarial pre-training strategy of the encoder in an embodiment of the present invention;

[0045] Figure 9 This is a schematic diagram illustrating the semantic alignment principle of the S-Former module in an embodiment of the present invention;

[0046] Figure 10 This is a comparison diagram of the Mel spectrum generated at different training stages and the real spectrum in an embodiment of the present invention;

[0047] Figure 11 This is a schematic diagram of a cross-modal silent speech reconstruction system based on ear canal air pressure micro-motion sensing in an embodiment of the present invention. Detailed Implementation

[0048] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.

[0049] In the relevant descriptions of this embodiment, the terms "including," "containing," and "possessing" are all open terms and are generally understood to include but not be limited to; the term "at least one" is generally understood to mean one or more, where "multiple" refers to two or more; the term "at least one of the following" or similar expressions refer to any combination of these items, including any combination of single or multiple items, for example, "at least one of a, b, or c", or "at least one of a, b, and c", which can all mean: a, b, c, ab (i.e., a and b), ac, bc, or abc, where a, b, and c can be single or multiple; the symbol "A / B" is used to describe the selection relationship of associated objects, generally indicating an "or" relationship.

[0050] In the following description of the embodiments, the terminology used in the embodiments of this application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The singular forms "a" and "the" as used in the embodiments of this application and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise.

[0051] Those skilled in the art should understand that, in the following description of the embodiments of this application, the sequence of numbers does not imply the order of execution. Some or all steps may be executed in parallel or sequentially. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0052] Those skilled in the art will understand that the numerical ranges in the embodiments of this application should be understood to specifically disclose each intermediate value between the upper and lower limits of the range. Any stated value or intermediate value within a stated range, as well as any other stated value or each smaller range between intermediate values ​​within a range, are also included within this invention. The upper and lower limits of these smaller ranges may be independently included or excluded from the range.

[0053] Unless otherwise stated, the technical / scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains. While this application describes only preferred methods and materials, any methods and materials similar or equivalent to those described herein may be used in the implementation or testing of this application. All references to this specification are incorporated by way of citation to disclose and describe the methods and / or materials associated with those references. In the event of any conflict with any incorporated reference, the content of this specification shall prevail.

[0054] With the rapid development of the Internet of Things (IoT) and wearable computing technologies, smart headphones have gradually become the next generation of core interactive terminals after smartphones. Traditional voice interaction relies heavily on sound wave signals, which presents an irreconcilable contradiction in practical applications: on the one hand, in quiet environments such as libraries and conference rooms, spoken commands can cause social interference and leak user privacy; on the other hand, in high-noise environments such as subways and factories, environmental noise can severely reduce the accuracy of speech recognition. To overcome the limitations of acoustic signals, Silent Speech Interfaces (SSIs) have emerged. Their core concept is to decode user intent by capturing non-acoustic motion signals from vocal organs (such as lips, tongue, and jaw). Currently, mainstream SSI (Speech-Based Intuition) technologies mainly fall into three categories: First, computer vision-based solutions utilize cameras to capture minute movements of the lips or facial muscles, using lip-reading algorithms to infer speech content. Second, active sensing-based solutions typically employ ultrasonic, millimeter-wave radar, or laser sensors to emit active detection signals towards the face or ear canal, receiving echoes to analyze the displacement of the vocal organs. Third, contact sensor-based solutions use patch-type electromyography (sEMG) sensors to capture electrical signals from facial muscles, or inertial measurement units (IMUs) to capture vibrations of the mandible. In recent years, research has focused on the emerging pathway of ear canal pressure sensing. Anatomically, the human ear canal is adjacent to the temporomandibular joint. When a user speaks (even silently), the movement of the mandible compresses the anterior wall of the ear canal, causing micrometer-level dynamic deformation of the ear canal volume, which in turn causes pressure fluctuations within the closed ear canal. Figure 1 As shown, Figure 1 This is a schematic diagram comparing the physical characteristics of the ear canal air pressure micro-motion signal and the standard acoustic speech signal in an embodiment of the present invention, demonstrating the non-acoustic characteristics of the TPVS signal, which lacks high-frequency harmonics and formants.

[0055] Applying ear canal pressure micro-motion sequences to speech reconstruction presents a huge physical gap and technical challenges. TPVS is essentially a low-frequency mechanical signal caused by organ movement, with an effective frequency band typically below 20Hz and a sampling rate of only about 100Hz. In contrast, human speech is a high-frequency acoustic signal containing rich high-frequency harmonics and formants, and the sampling rate typically needs to reach above 16kHz. Recovering a 16kHz acoustic signal from a 100Hz mechanical signal is a huge gap.

[0056] Based on this, the embodiments of this application provide a cross-modal silent speech reconstruction method and system based on ear canal air pressure micro-motion sensing, which solves the technical problems of existing technologies such as the physical limits of low-frequency mechanical signal reconstruction of high-frequency acoustic signals, insufficient generalization ability across users and speech rates, poor privacy protection and noise resistance performance, and high power consumption.

[0057] To illustrate the technical solution of the present invention, specific embodiments are described below.

[0058] Figure 2 A schematic flowchart illustrating an embodiment of a cross-modal silent speech reconstruction method based on ear canal pressure micro-motion sensing provided by this invention is shown below. Figure 2 As shown, the above method may include:

[0059] S101, by using at least one air pressure sensor installed in the user's ear canal, collects non-acoustic air pressure changes caused by the movement of the vocal organs, and obtains the original air pressure change time series;

[0060] For example, Figure 3 For a breakdown of the physical structure and experimental scenario of the hardware prototype in this embodiment of the invention, please refer to [link / reference]. Figure 3 As shown, the sensor can be integrated into the cavity of a self-developed in-ear headphone or custom earbud, which can be obtained through 3D printing. The sensor can be configured as follows: employing a micro-MEMS (Micro-Electro-Mechanical Systems) pressure sensor (such as the Bosch BMP390), with a sensor size not exceeding 3mm × 3mm × 3mm, and a sensing accuracy at the Pascal level, capable of capturing minute pressure fluctuations caused by micron-level ear canal deformation; the sensor's detection end faces the inside of the ear canal, and is externally wrapped with a flexible sealing material, such as memory foam ear tips, to ensure a tight fit against the ear canal wall. The ear canal pressure micro-motion timing signal acquired by the ear canal pressure sensor... A measure of the temporal changes in the motor state of the speech organs related to silent phonation The two follow a non-linear mapping model, which is expressed as:

[0061] ;

[0062] in, It can be a one-dimensional or multi-dimensional vector, and it is not required to be directly measured; It represents a nonlinear mapping function from the movement of the vocal organs to the slight changes in air pressure in the ear canal, and is used to characterize the motion-deformation-air pressure coupling effect during the vocalization process; This represents individual ear canal parameters, used to characterize the influence of individual differences in ear canal geometry, fit tightness, skin / soft tissue elasticity, etc., on the mapping relationship; This is an additive noise term used to characterize non-noise-producing factors such as environmental pressure disturbances, pressure fluctuations caused by slight head movements, and sensor measurement noise; t represents the time variable.

[0063] By using flexible acoustic seals that fit the user's ear canal, such as Figure 3As shown, multiple sizes such as S / M / L are available. A flexible acoustic seal creates a quasi-sealed air chamber between the earphone / earplug and the ear canal wall, isolating external air pressure interference and acoustic noise. This ensures that the sensor only captures the dynamic deformation of the ear canal volume caused by speech-related movements (movements of the speech organs dominated by the temporomandibular joint, such as jaw opening and closing, tongue movement, and lip movements), which is then converted into air pressure change signals.

[0064] The data acquisition process supports various sound scenarios, including audible and silent sounds, without requiring active transmission of detection signals. It only needs to passively sense and capture non-acoustic air pressure changes, thus avoiding privacy leaks and additional power consumption. The sensor continuously collects air pressure data at a preset sampling rate, ranging from 1Hz to 200Hz, with 100Hz to 200Hz being preferred. This sampling rate range fully covers the effective frequency band of temporomandibular joint movement (1Hz-6Hz) while suppressing high-frequency irrelevant noise. During sampling, it can communicate with a microcontroller (such as an Arduino Nano 33 BLE Sense) via an I2C interface to read sensor data in real time.

[0065] In some embodiments, the collected raw air pressure data carries a timestamp and can be transmitted to a processing terminal, such as a mobile terminal or headphone main control chip, via Bluetooth Low Energy (BLE) protocol or serial port (115200bps baud rate), and temporarily stored as a raw air pressure change time series. This series can directly reflect the temporal characteristics of the movement of the vocal organs, providing raw data support for subsequent preprocessing.

[0066] In some embodiments, after obtaining the raw time series of pressure changes (hereinafter referred to as TPVS), the above method may further include coupling quality assessment and gating.

[0067] Specifically, after the user wears the earbuds and during inference, a coupling quality score is constructed based on the static pressure level, low-frequency drift amplitude, short-time energy stability, and / or signal-to-noise ratio of TPVS. When the score is lower than a preset threshold, the user is prompted to adjust the earbud fit and discard the corresponding data segment or pause inference.

[0068] For example, with a length of The sliding window is used to evaluate TPVS. Take a time frame of 0.5s-1.5s; calculate within each window:

[0069] a) Static pressure offset This is used to reflect the degree of sealing.

[0070] b) Low-frequency drift amplitude It can be measured by the peak-to-peak value or standard deviation of the 0-0.5Hz component;

[0071] c) Short-time energy stability It can be measured by the coefficient of variation of the energy sequence;

[0072] d) Signal-to-noise ratio estimation It can be measured by the ratio of the energy of the effective frequency band (e.g., 1-6Hz) to the energy of the non-voice band;

[0073] The coupling quality score is obtained by normalizing the above indicators and then summing them using a weighted average. and set a threshold (For example, 0.6). Among them, Let be the weighting coefficient of the i-th coupling quality sub-index. is the normalized score of the i-th coupling quality sub-index.

[0074] S102, baseline drift suppression is performed on the original air pressure change time series, and amplitude enhancement is performed to highlight the articulation-related fluctuations, resulting in a preprocessed sequence;

[0075] In some embodiments, baseline drift suppression includes at least one of sliding window mean subtraction, bandpass filtering, or adaptive baseline estimation; amplitude enhancement includes at least one of squaring, absolute value, logarithmic compression, or normalization of the preprocessed sequence to improve the separability of articulation-related deformation fluctuations.

[0076] Understandable, the original pressure sequence It is susceptible to factors such as altitude changes and body temperature fluctuations, resulting in nonlinear baseline drift. For example, in this embodiment of the invention, a sliding window mean subtraction method is used for low-frequency trend terms caused by altitude or temperature changes. The calculation formula is as follows:

[0077] ;

[0078] in, The original signal; The signal is after drift removal; N is the window length (range 0.5s-5s, preferably 1s-3s); a is the half-width of the window, satisfying... , where i is the relative offset index within the window. This represents the local mean term centered at t. This step centers the signal to zero mean by tracking the local mean in real time and removing low-frequency trend terms.

[0079] In some embodiments, when a drastic baseline change (such as a rapid change in altitude) is detected, the system can automatically switch to a dynamic window adjustment mode, shortening the window length to 0.5s to improve the real-time performance of drift suppression; for slow drift scenarios, a 3s long window is used to ensure signal smoothness.

[0080] In some embodiments, bandpass filtering is used to filter effective frequency band features. This invention designs a targeted filtering scheme based on the effective frequency band characteristics of vocal organ movement (primarily temporomandibular joint):

[0081] For example, a second-order Butterworth bandpass filter is used, with the passband range set to 0.5Hz-10Hz; the low cutoff frequency (0.5Hz) is used to filter out sensor background noise and extremely slow changes in ambient air pressure, while the high cutoff frequency (10Hz) is used to suppress high-frequency body motion noise caused by walking, head turning, etc. Figure 4 This is a schematic diagram of the TPVS signal waveform after bandpass filtering in an embodiment of the present invention. Figure 4 As shown, this frequency band effectively preserves the motion characteristics of TMJ while filtering out sensor background noise and intense body movement noise caused by walking.

[0082] In some embodiments, this is performed after baseline drift suppression to avoid interference from drift signals on the filtering effect. A zero-phase filtering algorithm is used during the filtering process to prevent signal timing distortion and ensure that the temporal characteristics of the articulation event are not destroyed. After filtering, only the core frequency band signal of the temporomandibular joint movement (1Hz-6Hz) can be retained. This frequency band is directly related to the movement of the articulatory organs, laying the foundation for subsequent feature extraction.

[0083] In some embodiments, amplitude enhancement is used to improve the separability of vocal undulations.

[0084] Understandably, the amplitude of the de-drifted and filtered air pressure signal is weak, and the distinction between sound-related fluctuations and irrelevant noise is low. Amplification processing is needed to highlight the effective features. One or more of the following combined strategies can be used for enhancement:

[0085] Squaring operation: squaring the signal point by point. It amplifies the amplitude of deformation fluctuations related to pronunciation while suppressing small, irrelevant noise.

[0086] Absolute value conversion: Taking the absolute value of negative fluctuations unifies the direction of signal fluctuations, which facilitates subsequent energy characteristic calculations;

[0087] Normalization processing: The enhanced signal is mapped to the interval [0, 1] or [-1, 1] to eliminate amplitude differences between different users and different wearing states, and to unify the feature scale.

[0088] After the above three-step processing steps of baseline drift suppression and bandpass filter amplitude enhancement, a standardized preprocessed sequence is output, which can meet the following conditions:

[0089] The signal baseline is stable, fluctuating around the zero mean, with no obvious low-frequency drift;

[0090] The effective frequency band energy ratio is ≥85%, and the energy ratio of high-frequency noise and low-frequency interference is ≤15%.

[0091] The amplitude discrimination of articulation-related fluctuations is improved, and the signal-to-noise ratio with irrelevant noise is ≥10dB. The preprocessed sequence is directly used in subsequent event detection steps, providing high-quality data support for accurately locating the boundaries of articulation events.

[0092] In some embodiments, after preprocessing, the above method can further perform coupling quality assessment and gating on the preprocessed sequence.

[0093] For example, with a length of Sliding time window ( (With a step size of 0.1s-0.5s, four types of features are calculated and linearly fused on the preprocessed TPVS to obtain:)

[0094] ;

[0095] in, For signal-to-noise ratio related metrics, the following can be taken: , The energy range is 0.5Hz-8Hz. It is the sum of the energy in the 0Hz-0.3Hz and above frequency bands; It is a complement of the frequency band energy ratio or spectral flatness, used to eliminate flat noise caused by leakage / loosening; It is a peak structure consistency index, which can be obtained by the number of peaks falling within the expected range and the small variance of the interval between adjacent peaks; For multi-sensor consistency metrics, in a binaural scenario, the normalized maximum cross-correlation value or coherence of the left and right TPVS can be used. Weights - satisfy =1 (e.g., 0.25 for each), gate threshold A value of 0.55-0.85 (preferably 0.65-0.75) is acceptable. continuous A window ( If the value is 3-8, then pause the reasoning and prompt the user to adjust the wearing method; if continuous A window ( If the result is 1-3, then the reasoning is restored.

[0096] S103, Based on the temporal energy characteristics of the preprocessed sequence, determine the start and end boundaries of the silent sound event, extract the effective sound segment, and perform time normalization on the effective sound segment;

[0097] In some embodiments, determining the start and end boundaries of the silent sounding event based on the temporal energy characteristics of the preprocessed sequence in step S103 above may include:

[0098] Calculate the sliding window short-time energy curve of the preprocessed sequence; detect the energy peak under the condition of satisfying the minimum interval constraint; search forward and backward with the energy peak as the center until the energy or energy fluctuation rate is lower than the preset threshold to determine the start and end boundaries of the pronunciation event.

[0099] Understandably, the amplitude of pronunciation-related fluctuations in the preprocessed sequence is still relatively weak, and temporal energy calculation is needed to highlight the temporal characteristics of pronunciation events and provide a basis for boundary localization.

[0100] Specifically, firstly, for the preprocessed sequence after baseline drift suppression, its short-time energy curve or envelope is calculated. Typically, the signal is squared or its absolute value is taken to enhance articulation-related fluctuations, and then the movement energy is calculated within a set short-time window (e.g., 20 ms to 50 ms). On the short-time energy curve, local energy peaks that satisfy minimum interval constraints are detected, for example, a minimum interval greater than 100 ms, to avoid misinterpreting a single phonation as multiple phonations. This peak point corresponds to the moment when jaw movement is most significant during articulation.

[0101] Then, using the detected energy peak as the center, a search is performed in both forward (left) and backward (right) directions until a stopping condition is met. The stopping condition can be set as follows: the energy value at the current point decays to below a preset percentage of the peak energy (e.g., 3% to 10%), or the fluctuation rate of energy change is below a certain threshold. The two points determined in this way are the start and end boundaries of this silent vocalization event. Based on the determined start and end boundaries, the corresponding data segments are extracted from the original preprocessed sequence as valid vocalization segments. This operation filters out signals generated by prolonged silence or invalid body movements, focusing on valuable vocalization intervals.

[0102] like Figure 5 As shown, Figure 5 This is a schematic diagram of the event capture result based on short-time energy peak detection in an embodiment of the present invention. When an energy peak is detected to exceed a preset threshold, the system searches both sides of the time axis until the energy decays to less than 3% of the peak value, thereby accurately capturing silent sound segments.

[0103] In some embodiments, time normalization includes interpolating and resampling the effective speech segment to normalize the effective speech segment's time length or number of frames to a preset length.

[0104] Understandably, given the inherent differences in pronunciation duration between different users and commands, directly using variable-length segments is detrimental to stable model training and inference. Therefore, it is necessary to normalize the time dimension of the extracted effective pronunciation segments. Unifying the time length (or total number of frames) of all effective pronunciation segments to a preset fixed length can be achieved using interpolation resampling techniques. Specifically, the original segments of unequal length are treated as discrete sequences on the time axis, and through algorithms such as linear interpolation and spline interpolation, they are resampled to a sequence with a fixed number of sampling points while maintaining the basic signal morphology.

[0105] By performing time normalization, the impact of speech rate differences on the model is eliminated, ensuring the consistency of all input segments on the time scale. This is a key preprocessing step to improve the robustness and generalization ability of the model.

[0106] S104, The effective speech fragments are input into the pre-trained spectrogram reconstruction model to generate coarse-grained acoustic Mel spectrograms; the spectrogram reconstruction model is configured to extract semantic features related to the speech content from the barometric pressure micro-motion signal;

[0107] In this embodiment of the invention, the spectral reconstruction model is an end-to-end deep neural network designed for cross-modal mapping of barometric pressure micro-motion signals, semantic features, and acoustic spectra. It consists of a three-stage cascaded structure: a domain adversarial encoder (Baro-Encoder), a semantic latent space mapping unit (S-Former), and a generative coarse spectrum builder (MS-GAN generator). Its architecture is as follows: Figure 6 As shown, Figure 6 This is a schematic diagram of the architecture of the three-stage Mel spectrum reconstruction pipeline in an embodiment of the present invention, which includes three cascaded stages: semantic encoding, coarse-grained generation, and residual correction.

[0108] The encoder is configured to extract content representations from the effective pronunciation segments and suppress representation components related to user identity or wearing differences through a domain adversarial learning mechanism; the domain adversarial learning mechanism includes an identity discrimination branch and a gradient inversion or equivalent min-max optimization mechanism.

[0109] Specifically, encoders such as Figure 7 As shown, Figure 7 The network architecture diagram of the encoder in this embodiment of the invention adopts parallel one-dimensional convolution (processing the original air pressure waveform) and two-dimensional convolution (processing the TPVS short-time Fourier transform spectrum), and fuses multi-scale time-frequency features through the Inception module.

[0110] Domain adversarial learning (DAL) is a technique for decoupling content representation from user identity and wearability differences. Through a collaborative design involving identity discrimination branches, gradient inversion layers, and mini-max optimization, it forces the encoder to actively suppress irrelevant representation components. For example, it aims to eliminate individual differences. Figure 8 As shown, Figure 8 This diagram illustrates the adversarial pre-training strategy of the encoder in this embodiment of the invention. The network incorporates a Gradient Reversal Layer (GRL) for joint training with the identity discriminator, forcing the encoder to output content representations insensitive to user identity, ear canal geometry, and wearing differences. During training, the GRL inverts the gradient sign of the identity classifier, with the optimization objective being to minimize the command classification loss. Simultaneously maximize the loss of identity classification This forces the encoder to extract user-independent general features.

[0111] The core function of the semantic latent space mapping unit is to map the content representation refined by the encoder to a latent space shared with the text semantics, achieving cross-modal alignment between barometric signal representation and text semantic representation, thus ensuring the semantic validity of the content representation. For example... Figure 9 As shown, Figure 9 This is a schematic diagram illustrating the semantic alignment principle of the S-Former module in this embodiment of the invention. Through a contrastive learning mechanism, it aligns air pressure features... Mapped to the text embeddings generated by the pre-trained text encoder Shared semantic latent space. Loss function. The InfoNCE approach aims to shorten the distance between semantically consistent sample pairs and push away mismatched sample pairs.

[0112] The generator uses a U-Net architecture that removes skip connections, such as Figure 6 As shown in step two, a generative adversarial network is used to generate the basic Mel spectrum from the semantic vector. The generator employs a U-Net architecture that eliminates skip connections to prevent low-frequency air pressure noise from directly penetrating to the output.

[0113] In some embodiments, residual correction is as follows: Figure 6 As shown in step three, a residual network is introduced to predict high-frequency details. Output the final spectrum This module is optimized using a hybrid loss function, expressed as:

[0114] ;

[0115] in, This is the overall loss function for the residual correction module; Phoneme-level CTC loss is used to constrain the final predicted spectrum. The decoded phoneme sequence is consistent with the target phoneme sequence; This is the spectral convergence loss, used to constrain the difference in overall structure between the predicted spectrum and the true spectrum; Envelope correlation loss is used to constrain the consistency between the predicted spectrum and the actual spectrum in terms of energy envelope time series. These are non-negative weighting coefficients used to balance the contributions of each loss term. For example, Figure 10 This is a comparison diagram of the Mel spectrum generated at different training stages and the real spectrum in an embodiment of the present invention. See [link / reference] Figure 10 As shown, residual correction significantly improves the high-frequency texture clarity of the spectrum.

[0116] Understandably, current silent speech reconstruction models face multiple structural bottlenecks in training: First, the supervision signal is highly sparse, making it impossible to simultaneously acquire high-quality acoustic Mel spectra in real silent speech scenarios, resulting in a lack of strong supervision labels for generative modeling; second, the speech rate generalization ability is weak, the model has poor robustness to fast and slow speech rates, and the cross-speech rate recognition error rate fluctuates greatly; third, identity bias is serious, the differences in ear canal geometry, wearing seal, and jaw movement patterns among different users lead to strong individual specificity in TPVS, and if not decoupled, the encoder is prone to misjudging identity-related artifacts as semantic features, causing a precipitous decline in cross-user performance; fourth, multi-stage joint optimization is difficult, the spectrogram reconstruction module includes four coupled sub-modules: encoding, semantic mapping, coarse-grained generation, and residual correction, and end-to-end joint training is prone to gradient conflicts and convergence instability. These problems collectively restrict the feasibility and generalization reliability of cross-modal silent speech reconstruction methods based on ear canal pressure micro-motion sensing in practical wearable devices.

[0117] Based on this, embodiments of the present invention also provide a training method for a spectral reconstruction model, used to train the spectral reconstruction model described in the above embodiments. The training method includes:

[0118] S1, Obtain a training sample set containing the original time series of air pressure changes and supervision information, including text, acoustic Mel spectrum or a combination of both;

[0119] The original air pressure change time series refers to the TPVS signal collected by the air pressure sensor obtained through the method described in the above embodiments. The supervision information has several optional configuration forms: the first is pure text supervision, which is suitable for TPVS-text paired samples collected in silent speech scenarios. The text is converted into character / word / phoneme level token sequences after word segmentation, and the length is truncated or padded to a fixed dimension. The second is pure acoustic Mel spectrum supervision, which is suitable for TPVS-Mel spectrum paired samples collected synchronously in speech scenarios. The Mel spectrum is extracted by short-time Fourier transform. The third is hybrid supervision, in which the same TPVS sample is simultaneously associated with text labels and corresponding Mel spectra to form triples, which are used to jointly optimize semantic alignment and spectrum reconstruction objectives. The construction of this sample set supports the fusion of multi-source heterogeneous data. It can include high-quality synchronous data collected in a professional recording studio under controlled laboratory conditions, as well as weakly labeled on-site data transmitted back via Bluetooth from mobile terminals in the user's daily wear state.

[0120] S2 performs time stretching on the acoustic signals in the training samples while keeping the pitch constant, and performs corresponding time stretching or resampling on the original air pressure change time series synchronized with it to generate multi-speed training samples.

[0121] Specifically, to address the issue of inconsistent speaking speeds among users, this invention introduces a rhythm variation strategy during the training phase, including:

[0122] 1) Audio rhythm transformation: using a phase vocoder to transform the spoken speech signals in the training set. Perform timeline stretching or compression to generate speech rate scaling factor (Step size 0.1) speech variant .

[0123] 2) Barometric Pressure Synchronous Resampling: This involves resampling the barometric pressure sequence acquired synchronously with the audio. First, upsampling is performed to match the audio sampling rate, followed by scaling with the same scaling factor. Perform linear interpolation resampling. The interpolation formula is:

[0124] ;

[0125] in, This represents the original discrete barometric pressure sequence obtained synchronously with the i-th speech sample; This represents the value at the j-th sampling point of the original barometric pressure sequence after time-axis stretching / compression and resampling under a speech rate scaling factor β; j is the discrete sampling index of the resampled sequence. ; This represents the real index position on the original sequence that maps the sampling point to; This indicates the floor function. This indicates the rounding up operation, used to determine the two adjacent original sampling points on which the interpolation depends; These are the interpolation weights (ranging from [0, 1]), used in... and Linear interpolation is performed between them.

[0126] 3) Construct an enhanced sample set: Combine the generated synchronization sample pairs Adding it to the training set enables the model to learn pressure morphological invariants at different time scales.

[0127] S3 uses semantic classification loss and identity adversarial loss to jointly train the encoder, so that the encoder outputs content representations that are insensitive to user identity.

[0128] In some embodiments, step S3 above may include: employing a domain adversarial neural network framework, so that the content representation output by the encoder is simultaneously input into the semantic classifier and the identity discriminator, and implementing adversarial training of the identity discriminator through a gradient reversal layer or an equivalent min-max optimization mechanism.

[0129] The semantic classification loss uses cross-entropy loss. The supervision target is a predefined command category, and the classifier is a fully connected layer. The identity adversarial loss is implemented through a domain adversarial neural network framework: after the encoder output content representation, a gradient reversal layer is connected, followed by an identity discriminator, whose supervision target is the user ID, and the identity classification loss is... Also representing cross-entropy; the joint optimization objective is:

[0130] ;

[0131] in, To counteract the weighting coefficients, used to adjust the strength of identity invariance constraints.

[0132] During training, GRL inverts the gradient sign of the identity classifier and optimizes by minimizing the command classification loss. Simultaneously maximize the loss of identity classification This forces the encoder to extract features. satisfy and This mathematically isolates user-irrelevant general features, such as ear canal geometry, wearing tightness, and jaw muscle strength, while retaining motion envelope features strongly correlated with the vocal content. This adversarial training mechanism allows the encoder's output to exhibit clear semantic clustering and fuzzy identity distribution in visualization, thus eliminating the interference of individual biases on subsequent spectrum generation. This means updating the encoder parameters during training. To minimize semantic classification loss, thereby improving the ability of features to distinguish command categories; This means updating the encoder parameters during training. To maximize the identity classification loss, thereby reducing the amount of information in the features that can be used to distinguish user identities, and prompting the encoder to output general features that are insensitive to user identities; This is the set of parameters for the encoder (feature extraction network).

[0133] S4, apply contrastive learning constraints to content representation and text representation to obtain semantic latent space mapping units;

[0134] Content representation refers to the content output by the encoder in step S3, and text representation refers to the embedding vector generated by the pre-trained text encoder from the corresponding instruction text; the two are dimensionally aligned. The contrastive learning constraint uses the InfoNCE loss function, expressed as:

[0135] ;

[0136] in For cosine similarity, The value is the temperature coefficient, the denominator iterates through all negative samples in the batch, and B is the training batch size. The contrastive learning loss function (InfoNCE loss) is used to compare the barometric pressure features of the same semantic sample. With text embedding Zoom in, and With other non-matching items in the batch Push away to achieve implicit alignment.

[0137] This constraint will include pressure characteristics. With text embedding Pull closer, push away at the same time Other This mapping project maps the two to a shared semantic latent space. In this space, the TPVS features of the same semantic instruction are highly similar to the text embedding in terms of Euclidean distance, while the distance between different semantic instructions increases significantly. This semantic latent space mapping unit is the S-Former module; after training, this unit can operate independently without text supervision, generating semantically aligned latent vectors with only TPVS input, supporting zero-shot transfer learning.

[0138] S5 trains the generative model to output coarse-grained Mel spectrograms, and trains the residual learning network to output detailed residuals under the condition of freezing the generative model.

[0139] The generative model uses the MS-GAN architecture, with its generator employing a variant of U-Net: the encoder path includes 4 levels of downsampling, the decoder path has no skip connections, and the final output is an 80-channel convolutional layer that produces a coarse-grained Mel-ray spectrogram; the discriminator uses a PatchGAN structure, outputting a T×80 true / false probability map; training employs a hybrid loss mechanism, with the loss function expressed as:

[0140] ;

[0141] in, For phoneme-level CTC loss; For spectral convergence loss; Envelope correlation loss, jointly constraining the accuracy, clarity, and prosodic naturalness of the reconstructed speech; S is the true Mel spectrum. The predicted Mel spectrum is output by the residual correction module. It is the Frobenius norm (a form of matrix 2 norm used to measure the overall difference of a spectral matrix). The correlation coefficient operator, preferably the Pearson correlation coefficient, is used to measure the similarity in shape between two sequences. This represents the energy envelope sequence extracted from the true spectrum S. Indicates the predicted spectrum The extracted energy envelope sequence; The weighting coefficients are used to balance the contributions of the spectral convergence term and the envelope correlation to the overall optimization. The residual learning network is a residual learning-based Acoustic Details Correction (PEARL) module, whose input is... Output high-frequency detail residual The structure is a zero-initialized 3-layer dilated convolutional network to ensure that the residual prediction is initially zero, avoiding damage to the coarse-grained basic structure; during training, the generative model is frozen, meaning all parameters of MS-GAN are fixed, and only the parameters of PEARL are updated. This phased training strategy avoids the problem of gradient direction conflict between the generator and the residual network in end-to-end joint training: MS-GAN focuses on learning low-frequency formants and syllable contours, while PEARL focuses on repairing high-frequency harmonics and transient details.

[0142] Through the above steps, this invention achieves a systematic reconstruction of the training process for the spectrogram reconstruction model: Step S1 constructs a mixed sample set covering both spoken and silent scenarios, and strong and weak supervision, providing multi-level supervision signals for the model; Step S2 injects speech rate invariant priors into the data level through rhythm enhancement driven by a phase vocoder, enabling the model to inherently possess speech rate adaptation capabilities; Step S3 utilizes a gradient inversion mechanism to achieve content-identity decoupling, eliminating individual bias from the feature source; Step S4 uses contrastive learning to establish weakly supervised semantic anchors between TPVS and text, overcoming the bottleneck of no speech labels in silent scenarios; Step S5 adopts a phased optimization paradigm of freezing the generator and fine-tuning the residuals to ensure the convergence stability and functional specificity of each sub-module. By constructing a multi-source heterogeneous supervision sample set, the technical problem of sparse supervision signals in silent scenarios is solved; by introducing rhythm-aware data enhancement and adversarial decoupling mechanisms, the technical problems of insufficient speech rate diversity and severe identity bias are solved; by adopting a phased training strategy and multi-dimensional loss constraints, the technical problem of difficult multi-stage joint optimization is solved.

[0143] S105, based on the residual learning mechanism, high-frequency details are supplemented and corrected on the coarse-grained acoustic Mel spectrum to obtain a high-fidelity target Mel spectrum;

[0144] In some embodiments, step S105 may specifically include:

[0145] S1051, inputs coarse-grained acoustic Mel spectrum into the residual network;

[0146] The residual network is a lightweight convolutional neural network that uses zero-initialized weights and a stacked residual block structure, with input dimensions and coarse-grained spectrum. Figure 1 The output dimension is the same, but it focuses on high-frequency incremental modeling. For example, its network architecture can be: 3 layers of dilated convolution with channel attention, or a U-Net-style encoder-decoder structure without skip connections to avoid repeated injection of low-frequency information. During the training phase, the network is constrained to learn only the local differences in the spectral amplitude domain and does not change the global energy distribution.

[0147] S1052, the residual network predicts and outputs a residual spectrum corresponding to the high-frequency details missing in the coarse-grained acoustic Mel spectrum;

[0148] Specifically, the residual network takes the coarse-grained spectrum output in step S1051 as input and predicts a detailed residual spectrum through multi-layer nonlinear transformation. This residual spectrum may be positive or negative in value, and its physical meaning is to locally adjust and supplement the frequency components that are too strong or too weak, missing or incorrect in the coarse-grained spectrum.

[0149] S1053, the residual spectrum is superimposed with the coarse-grained acoustic Mel spectrum to obtain a high-fidelity target Mel spectrum;

[0150] The superposition operation is an element-wise addition operation. This operation keeps the low-frequency base of the coarse-grained spectrum unchanged and injects learned and calibrated detail increments only in the high-frequency region, avoiding the high-frequency overshoot or artifact amplification problems inherent in generative models.

[0151] In some embodiments, in order to guide the residual network to learn truly effective detail corrections, a multi-constraint hybrid loss function is used for optimization during the training phase to ensure that the corrected spectrum approximates real speech at multiple levels.

[0152] The hybrid loss function includes at least one of phoneme-level CTC loss, frequency domain structure consistency loss, and prosodic correlation loss based on signal envelope.

[0153] It should be noted that phoneme-level CTC loss is an end-to-end loss function designed for sequence modeling tasks. It does not rely on frame-level alignment annotation and is suitable for situations where there is no explicit time synchronization relationship between TPVS and phoneme sequences in silent speech scenarios. Frequency domain structural consistency loss refers to a normalized distance metric used to constrain the consistency of the energy distribution pattern between the predicted spectrum and the real spectrum at the Mel scale. Prosodic correlation loss based on signal envelope refers to using the time domain envelope extracted by Hilbert transform as a carrier to measure the consistency between the reconstructed spectrogram and the real spectrogram at the suprasegmental prosodic feature level.

[0154] S106, input the target Mel spectrum into the automatic speech recognition model, and output the corresponding silent speech recognition result.

[0155] Specifically, the input to the above steps is the high-fidelity target Mel spectrum obtained after the aforementioned series of processing steps. This spectrum already possesses a feature structure in the time-frequency domain that is highly similar to real spoken speech. The automatic speech recognition model can be an end-to-end ASR system based on architectures such as Transformer. The ASR model receives the Mel spectrum as input, and through its internal acoustic encoder and language decoder, analyzes the spectral features frame by frame or based on an attention mechanism, ultimately outputting the most probable text sequence, such as text control commands like playing music or opening navigation, or it can be directly mapped to predefined control command codes.

[0156] To achieve low latency and privacy protection, the model is preferably deployed on a local terminal and optimized using techniques such as streaming processing and model quantization. Its working cycle is managed by the system's trigger control unit and is activated only during a valid inference session. The final recognition result will be delivered to the operating system or application to perform corresponding operations, thereby completing the full mapping from physical signals to user intent.

[0157] For example, the trained model is exported in ONNX format, and its size is compressed to below 5MB using FP16 half-precision quantization to adapt to mobile computing power. On the mobile device (such as an Android smartphone), the NPU is invoked for acceleration via an inference framework (e.g., ONNXRuntime Mobile). The system acquires TPVS in real time, preprocesses it, inputs it into the model, streams Mel spectra, and inputs them into a locally running automatic speech recognition engine for decoding into text. Experiments show that the total inference latency is less than 200ms. Terminal inference is initiated by trigger events, including wake words, key presses, touch input, or application invocation. Coupling quality assessment gating is performed before or during inference; if the coupling quality falls below a threshold, inference is paused and a prompt to adjust the wearing device is displayed. After outputting text or control commands, spectrogram reconstruction and decoding are stopped, and the system waits for the next trigger event.

[0158] In some embodiments, terminal inference is initiated by a triggering event, which includes, but is not limited to, wake words, key / touch operations, application programming interface calls, or peripheral instructions; after triggering, an inference session is entered, and the trigger debouncing time is set. (e.g., 0.3s-1.0s) to ignore repeated triggering within a short period. At the start of the inference session, it is preferable to first acquire a steady-state baseline (e.g., 0.2s-0.5s) for updating. And coupling quality score.

[0159] The current inference session will terminate and enter a waiting state if any of the following conditions are met:

[0160] a) Text / control commands have been output;

[0161] b) Exceeding the session duration limit (e.g., 2s-6s);

[0162] c) The duration of continuous detection of no valid sound events exceeds (e.g., 0.8s-2.0s);

[0163] d) Coupling quality score Continuously below the threshold Exceed One window;

[0164] e) Changes in wearing status or earpiece dislodgement (which can be determined by static pressure surges, contact detection, or other wearable sensors). Optionally, when the decoding confidence level is below a threshold, the terminal prompts the user to repeat the input without executing the control command.

[0165] In some embodiments, the trigger control unit can be implemented using a finite state machine, and its states include at least: waiting state S0 (waiting for a trigger event); baseline calibration state S1 (acquiring a steady-state baseline and updating the drift estimate); acquisition state S2 (continuously acquiring TPVS and buffering); preprocessing and event detection state S3 (filtering / enhancing and locating event boundaries); spectrum reconstruction state S4 (generating a coarse spectrum and performing residual refinement); decoding and output state S5 (ASR decoding and outputting instructions); and termination state S6 (stopping inference and returning to S0).

[0166] Specifically, when the triggering event arrives, the system transitions from S0 to S1; when the coupling quality score is reached... And when a valid event is detected, proceed from S3 to S4; when Continue to exceed One window or exceeding the session duration limit When the output instruction is completed, the system enters S6 and returns to S0; when the output instruction is completed, it enters S6 from S5 and enters the waiting state.

[0167] In some embodiments, the trigger time can be 100ms-800ms (preferably 200ms-400ms); maximum session duration. The output time can be 2s-20s (preferably 4s-10s); the waiting time after output can be 0.2s-2s; when a change in wearing status or a coupling quality that is consistently below the threshold is detected... The session can be ended early to reduce false triggers and power consumption.

[0168] This invention provides a cross-modal silent speech reconstruction method based on the sensing of minute air pressure movements in the ear canal. By using at least one air pressure sensor placed in the user's ear canal to collect non-acoustic air pressure changes caused by the movement of the vocal organs, a raw air pressure change time series is obtained. This achieves passive sensing of minute air pressure fluctuations within the ear canal, avoiding the power consumption problems associated with active signal transmission. Baseline drift suppression processing of the raw air pressure change time series effectively eliminates low-frequency interference caused by environmental air pressure changes while preserving effective frequency band signals related to speech. Furthermore, amplitude enhancement highlights speech-related fluctuations, improving signal separability and providing high-quality input for subsequent event detection. The start and end boundaries of silent speech events are determined based on the temporal energy characteristics of the preprocessed sequence, extracting effective speech segments. Time normalization of these effective speech segments solves the problem of inconsistent speech segment lengths at different speech rates, providing standardized input for subsequent processing. Effective vocal segments are input into a pre-trained spectrogram reconstruction model to generate a coarse-grained acoustic Mel spectrum. This model is configured to extract semantic features related to the vocal content from the air pressure micro-motion signal, achieving an initial mapping from low-frequency mechanical signals to acoustic features. Based on a residual learning mechanism, high-frequency details are supplemented and corrected in the coarse-grained acoustic Mel spectrum to obtain a high-fidelity target Mel spectrum. A residual network predicts and outputs residual spectra corresponding to the missing high-frequency details in the coarse-grained acoustic Mel spectrum. These residual spectra are then superimposed on the coarse-grained acoustic Mel spectrum, effectively overcoming the technical bottleneck of reconstructing high-frequency details from low-frequency signals. Finally, the target Mel spectrum is input into an automatic speech recognition model, outputting the corresponding silent speech recognition result, completing the full conversion from ear canal air pressure micro-motion to voice commands. This scheme, through the combination of domain adversarial pre-training and semantic alignment training stages, enables the encoder to extract semantic content features independent of user identity and maps air pressure signal features and text semantics to a shared semantic space, solving the problem of insufficient cross-user generalization ability. The residual network is optimized by a hybrid loss function, which includes phoneme-level CTC loss, frequency domain structure consistency loss, and prosodic correlation loss based on signal envelope. This ensures the accuracy, clarity, and naturalness of the reconstructed speech. The solution achieves high-precision silent speech command parsing in mobile and noisy scenarios, while eliminating the risk of user voiceprint privacy leakage from the physical source. It has advantages such as low power consumption, high noise immunity, and strong privacy protection.

[0169] Based on the same inventive concept, embodiments of this application also provide a cross-modal silent speech reconstruction system based on the sensing of ear canal air pressure micro-motion. Figure 11 This is a schematic diagram of a cross-modal silent speech reconstruction system based on ear canal air pressure micro-motion sensing, as described in an embodiment of the present invention. (See attached diagram.) Figure 11 As shown, the cross-modal silent speech reconstruction system based on ear canal pressure micro-motion sensing may include:

[0170] The sensing module is used to collect non-acoustic air pressure changes caused by the movement of the vocal organs through at least one air pressure sensor installed in the user's ear canal, and obtain the original air pressure change time series.

[0171] The preprocessing module is used to suppress baseline drift and enhance amplitude to highlight articulation-related fluctuations in the original time series of air pressure changes, resulting in a preprocessed sequence.

[0172] The event detection module is used to determine the start and end boundaries of silent sound events based on the temporal energy characteristics of the preprocessed sequence, extract valid sound segments, and perform time normalization on the valid sound segments;

[0173] The spectrogram reconstruction module is used to input effective speech segments into a pre-trained spectrogram reconstruction model to generate coarse-grained acoustic Mel spectrograms; the spectrogram reconstruction model is configured to extract semantic features related to the speech content from barometric pressure micro-motion signals;

[0174] The correction module is used to supplement and correct high-frequency details of the coarse-grained acoustic Mel spectrum based on the residual learning mechanism to obtain a high-fidelity target Mel spectrum.

[0175] The decoding module is used to input the target Mel spectrum into the automatic speech recognition model and output the corresponding silent speech recognition result.

[0176] The sensing module includes:

[0177] like Figure 3 As shown, a prototype device was constructed in this embodiment of the invention. The absolute air pressure signal within the ear canal was continuously acquired at a sampling rate of approximately 108Hz using a MEMS barometric pressure sensor (e.g., Bosch BMP390) ​​built into the in-ear headphones. The sensing accuracy reaches the Pascal level, meeting the needs of ear canal micro-movement sensing.

[0178] The flexible acoustic seal allows the earphones to fit snugly against the ear canal wall using memory foam ear tips, forming a quasi-sealed air chamber. This ensures the capture of micron-level ear canal deformation signals. The TPVS (Transmission Throat Spectroscopy) and the movement of the vocal organs follow a nonlinear mapping model: air pressure signal... Movement of the vocal organs The two follow a non-linear mapping model: .

[0179] The data transmission unit integrates a microcontroller or development board (such as Arduino Nano 33 BLE Sense), reads air pressure data via an I2C interface, and transmits the raw data with timestamps to the processing terminal via Bluetooth Low Energy (BLE) protocol or serial port (such as 115200bps).

[0180] The preprocessing module includes:

[0181] The DC drift removal unit executes a sliding window mean subtraction algorithm, with the following formula: To eliminate low-frequency interference caused by changes in altitude or body temperature;

[0182] Parameter range: The air pressure sampling rate can be 1Hz-200Hz, preferably 100Hz-200Hz; the drift suppression window can be 0.5s-5s, preferably 1s-3s; the filtering can use a high-pass filter of 0.1Hz-1Hz to suppress slow drift and a low-pass filter of 8Hz-30Hz to suppress high-frequency artifacts, or an equivalent bandpass filter of 0.2Hz-20Hz.

[0183] The filtering unit integrates a second-order Butterworth bandpass filter with a passband range locked at 0.5Hz-10Hz to preserve the TMJ motion characteristics of 1-6Hz and filter out sensor noise floor and body noise.

[0184] The event detection module includes:

[0185] The event capture unit calculates the square of the signal amplitude. And short-term energy (window length 10ms-50ms), locate the energy peak and search to both sides until the energy decays to below a preset percentage (e.g. 3%) of the peak, and extract the pronunciation segment.

[0186] Parameter range: The short-time energy window length can be 10ms-80ms, preferably 20ms-50ms; the energy smoothing window can be 50ms-300ms; the minimum peak interval can be 100ms-500ms; the energy decay ratio of the boundary search can be 1%-10%, preferably 2%-6%; after truncation, linear interpolation / spline interpolation can be used to normalize the fragment to a fixed length (e.g., within the range of 0.8s-3s).

[0187] In some embodiments, the system may further include a rhythm adaptation enhancement module for addressing the problem of inconsistent user speech rates, which includes:

[0188] Audio rhythm transformation: using a phase vocoder to process spoken speech in the training set. Perform time stretching / compression without changing the pitch to generate a scaling factor. (Step size 0.1) speech variants;

[0189] Parameter range (rhythm enhancement): time scaling factor The value can be 0.6-1.5, preferably 0.7-1.3; the rate step size can be 0.05-0.2; the synchronous speed change of TPVS can be achieved by resampling, while maintaining the consistency of event boundary labels.

[0190] TPVS synchronous resampling: for synchronously acquired barometric pressure sequences First, upsampling is performed, followed by applying the same scaling factor. Perform linear interpolation resampling to generate pseudo samples ;

[0191] Sample expansion and construction of enhanced sample pairs This is used to train downstream models.

[0192] The spectral reconstruction module includes:

[0193] 1) Domain adversarial encoder: such as Figure 7 As shown, it includes a parallel one-dimensional convolutional branch (processing TPVS waveforms) and a two-dimensional convolutional branch (processing the spectrograms corresponding to TPVS). The two branches perform feature fusion through the Inception module and the cross-attention mechanism.

[0194] Identity decoupling layer: such as Figure 8 As shown, a gradient inversion layer (GRL) and an identity classifier are connected after the feature extraction layer. This layer is configured to invert the sign of the identity classification gradient during backpropagation training, with the optimization objective being:

[0195] ;

[0196] This allows us to extract general semantic features that are independent of user identity.

[0197] 2) Semantic alignment unit:

[0198] Cross-modal mapping layer: such as Figure 9 As shown, this unit will convert the TPVS latent feature vector Mapped to text embedding vectors generated by a pre-trained text encoder (such as Sentence-BERT) Same semantic space.

[0199] Contrastive Learning Optimizer: Utilizing Contrastive Learning Loss Minimize semantically consistent sample pairs ( The distance between them enables implicit alignment under weak supervision.

[0200] 3) Generative Spectrum Construction Unit:

[0201] U-Net generator: such as Figure 6 As shown in step two, an encoder-decoder structure is adopted, with a multi-head self-attention layer integrated at the encoder end to capture the global context. To prevent low-frequency noise from penetrating, the decoder path does not contain skip connections and directly reconstructs the Mel spectrum from semantic features.

[0202] Discriminator: A convolutional neural network using the PatchGAN architecture is used to determine the authenticity of the generated spectrum.

[0203] Hybrid Loss Controller: Combines pixel-level L2 reconstruction loss and adversarial loss, with weighting coefficients... Set to 0.2.

[0204] 4) Residual detail correction unit:

[0205] Residual prediction networks: such as Figure 6 As shown in step three, a zero-initialized convolutional layer is used to receive the coarse-grained spectrum and predict the high-frequency detail residuals. Output the final spectrum .

[0206] Multidimensional Constrained Optimizer: Configured to compute and minimize the sum of the following three losses:

[0207] a. Phoneme-level CTC loss: Constrain the accuracy of the phoneme sequence after spectrum decoding; This indicates that at time (frame) t, in the posterior probability of the phonemes output by the model, the alignment path π takes the label at that time. The probability value; The label (which can be a specific phoneme label or a blank symbol) represents the CTC alignment path π in frame t, and is used to describe the alignment relationship from the "frame-level label sequence" to the "target phoneme sequence". Represents the target phoneme sequence The t-th symbol / label (obtained from text or forced alignment) or used to represent the target sequence constraint corresponding to path π;

[0208] b. Weighted spectral convergence loss: Constrain the consistency of frequency domain capability distribution;

[0209] c. Envelope correlation loss: The goal is to maximize the Pearson correlation coefficient between the reconstructed signal and the Hilbert envelope of the real signal.

[0210] The decoding module is responsible for the system's final output and applications, including:

[0211] Deployment and Inference: The trained model is exported to ONNX format, and the model size is compressed to below 5MB using FP16 half-precision quantization. On mobile devices (such as smartphones), inference is accelerated by calling the NPU / DSP.

[0212] Interactive control: The system receives the TPVS transmitted by module 2 in real time, reconstructs it into a high-fidelity Mel spectrum by correction module 1105, and streams it to the locally running Automatic Speech Recognition (ASR) engine to output text commands and execute corresponding device control operations (such as answering calls and playing music). Experiments show that the total latency of the system in mobile scenarios is less than 200ms.

[0213] Interactive control unit: Used to implement trigger-based start and stop control. After receiving a wake word / key / touch / application call or other trigger event, it starts an inference session. In the session, it performs coupling quality gating and debouncing control, and stops spectrum reconstruction and decoding and enters a waiting state when conditions such as output command, timeout or insufficient coupling are met.

[0214] The various embodiments in this specification are described in a progressive manner. For the same or similar parts between the various embodiments, please refer to each other. Each embodiment focuses on describing the differences from other embodiments.

[0215] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit this application. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of this application.

Claims

1. A cross-modal silent speech reconstruction method based on ear canal pressure micro-motion sensing, characterized in that, include: By using at least one air pressure sensor installed in the user's ear canal, non-acoustic air pressure changes caused by the movement of the vocal organs are collected to obtain the original air pressure change time series. The original time series of air pressure changes was subjected to baseline drift suppression and amplitude enhancement to highlight articulation-related fluctuations, resulting in a preprocessed sequence. Calculate the sliding window short-time energy curve of the preprocessed sequence; detect the energy peak under the condition of satisfying the minimum interval constraint; search forward and backward with the energy peak as the center until the energy or energy fluctuation rate is lower than the preset threshold, determine the start and end boundaries of the silent sound event, extract the effective sound segment, and perform time normalization on the effective sound segment; time normalization includes interpolating and resampling the effective sound segment to normalize the time length or frame number of the effective sound segment to the preset length; The effective vocal segments are input into a pre-trained spectrogram reconstruction model to generate coarse-grained acoustic Mel spectrograms. The spectrogram reconstruction model is an end-to-end deep neural network designed for cross-modal mapping of barometric pressure signals, semantic features, and acoustic spectrograms. The spectrogram reconstruction model includes at least an encoder and a semantic latent space mapping unit. The encoder is configured to extract content representations from the effective vocal segments and suppress representation components related to user identity or wearing differences through a domain adversarial learning mechanism. The domain adversarial learning mechanism forces the encoder to actively suppress irrelevant representation components through a collaborative design of identity discrimination branch, gradient inversion layer, and min-max optimization. The spectrogram reconstruction model introduces a gradient inversion layer and an identity discriminator for joint training, forcing the encoder to output content representations that are insensitive to user identity, ear canal geometry, and wearing differences. During training, the gradient inversion layer inverts the gradient sign of the identity classifier, with the optimization objective being to minimize the command classification loss while maximizing the identity classification loss, forcing the encoder to extract user-irrelevant general features. The high-frequency details of the coarse-grained acoustic Mel spectrum are supplemented and corrected based on a residual learning mechanism to obtain a high-fidelity target Mel spectrum. This includes: inputting the coarse-grained acoustic Mel spectrum into a residual network; the residual network predicting and outputting a residual spectrum corresponding to the missing high-frequency details in the coarse-grained acoustic Mel spectrum; and superimposing the residual spectrum with the coarse-grained acoustic Mel spectrum to obtain the high-fidelity target Mel spectrum. In the residual learning mechanism, the residual network is optimized using a hybrid loss function during the training phase. The hybrid loss function includes phoneme-level CTC loss, frequency domain structure consistency loss, and prosodic correlation loss based on signal envelope. The target Mel spectrum is input into the automatic speech recognition model, and the corresponding silent speech recognition result is output.

2. The method according to claim 1, characterized in that, The baseline drift suppression includes at least one of sliding window mean subtraction, bandpass filtering, or adaptive baseline estimation; the amplitude enhancement includes at least one of squaring, absolute value, logarithmic compression, or normalization of the preprocessed sequence to improve the separability of articulation-related deformation fluctuations.

3. The method according to claim 2, characterized in that, The semantic latent space mapping unit maps the content representation to the semantic latent space through contrastive learning or matching learning; wherein, the positive sample pairs of contrastive learning include the content representation and text representation corresponding to the same statement or the same semantic tag.

4. A method for training a spectral reconstruction model, used to train the spectral reconstruction model according to any one of claims 1 to 3, characterized in that, The training method includes: Obtain a training sample set containing the original time series of air pressure changes and supervision information, wherein the supervision information includes text, acoustic Mel spectrum, or a combination of both; The acoustic signals in the training samples are time-scaled while keeping the pitch constant, and the original time series of air pressure changes synchronized with them are time-scaled or resampled accordingly to generate multi-speed training samples. The encoder is jointly trained with semantic classification loss and identity adversarial loss, so that the encoder outputs content representations that are insensitive to user identity. A contrastive learning constraint is applied to the content representation and the text representation to obtain semantic latent space mapping units; The generative model is trained to output coarse-grained Mel spectra, and a residual learning network is trained under the condition of freezing the generative model to output detailed residuals.

5. The method according to claim 4, characterized in that, The method of jointly training the encoder with semantic classification loss and identity adversarial loss, so that the encoder outputs content representations that are insensitive to user identity, includes: A domain adversarial neural network framework is adopted, in which the content representation output by the encoder is simultaneously input into the semantic classifier and the identity discriminator, and the adversarial training of the identity discriminator is achieved through a gradient reversal layer or an equivalent min-max optimization mechanism.

6. A cross-modal silent speech reconstruction system based on ear canal air pressure micro-motion sensing, characterized in that, include: The sensing module is used to collect non-acoustic air pressure changes caused by the movement of the vocal organs through at least one air pressure sensor installed in the user's ear canal, and obtain the original air pressure change time series. The preprocessing module is used to suppress baseline drift and enhance amplitude to highlight articulation-related fluctuations in the original air pressure change time series, resulting in a preprocessed sequence. The event detection module is used to calculate the sliding window short-time energy curve of the preprocessed sequence; detect the energy peak under the condition of satisfying the minimum interval constraint; search forward and backward with the energy peak as the center until the energy or energy fluctuation rate is lower than the preset threshold, determine the start and end boundaries of the silent sound event, extract the effective sound segment, and perform time normalization on the effective sound segment; the time normalization includes interpolation and resampling of the effective sound segment to normalize the time length or frame number of the effective sound segment to the preset length; The spectrogram reconstruction module is used to input the effective vocal segments into a pre-trained spectrogram reconstruction model to generate coarse-grained acoustic Mel spectrograms. The spectrogram reconstruction model is an end-to-end deep neural network designed for cross-modal mapping of barometric pressure signals, semantic features, and acoustic spectrograms. The spectrogram reconstruction model includes at least an encoder and a semantic latent space mapping unit. The encoder is configured to extract content representations from the effective vocal segments and suppress representation components related to user identity or wearing differences through a domain adversarial learning mechanism. The domain adversarial learning mechanism, through a collaborative design of identity discrimination branch, gradient inversion layer, and min-max optimization, forces the encoder to actively suppress irrelevant representation components. The spectrogram reconstruction model introduces a gradient inversion layer and an identity discriminator for joint training, forcing the encoder to output content representations insensitive to user identity, ear canal geometry, and wearing differences. During training, the gradient inversion layer inverts the gradient sign of the identity classifier, with the optimization objective being to minimize the command classification loss while maximizing the identity classification loss, forcing the encoder to extract user-irrelevant general features. A correction module is used to supplement and correct high-frequency details of the coarse-grained acoustic Mel spectrogram based on a residual learning mechanism to obtain a high-fidelity target Mel spectrum. This includes: inputting the coarse-grained acoustic Mel spectrogram into a residual network; the residual network predicting and outputting a residual spectrum corresponding to the missing high-frequency details in the coarse-grained acoustic Mel spectrogram; and superimposing the residual spectrum with the coarse-grained acoustic Mel spectrogram to obtain the high-fidelity target Mel spectrum. In the residual learning mechanism, the residual network is optimized using a hybrid loss function during the training phase. The hybrid loss function includes phoneme-level CTC loss, frequency domain structure consistency loss, and prosodic correlation loss based on the signal envelope. The decoding module is used to input the target Mel spectrum into the automatic speech recognition model and output the corresponding silent speech recognition result.