A human voice compensation method based on laser-induced graphene

By combining a laser-induced graphene microphone with an artificial intelligence speech restoration model, the problems of poor adaptability and narrow frequency band coverage of existing auxiliary speech technologies are solved, achieving efficient and smooth speech compensation effects.

CN122120691APending Publication Date: 2026-05-29XI AN JIAOTONG UNIV

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
XI AN JIAOTONG UNIV
Filing Date
2026-02-11
Publication Date
2026-05-29

Smart Images

  • Figure CN122120691A_ABST
    Figure CN122120691A_ABST
Patent Text Reader

Abstract

The application discloses a human voice compensation method based on laser-induced graphene, relates to the cross field of biomedical engineering, flexible sensing technology and artificial intelligence, and aims to solve the problems of poor adhesion, narrow frequency band coverage and low speech recognition degree of traditional auxiliary sound production equipment. The method is designed in an integrated manner of "customized LIG pickup preparation - wide frequency band vibration collection - artificial intelligence speech restoration", acquires the vibration signals of the neck of a patient and divides the 100Hz-8kHz full frequency band to establish a personalized feature library, customizes the porous, fibrous or gradient LIG microstructure according to the frequency band requirement and optimizes the laser parameters, prepares a flexible pickup to ensure close adhesion with the neck, then trains a personalized speech restoration model based on a CNN+LSTM hybrid model, finally assembles and debugs the system to realize accurate collection of full frequency band signals and natural speech restoration with a semantic accuracy of greater than or equal to 98%, which is suitable for different patient requirements and supports dynamic optimization, and meets the daily communication needs of patients with difficulty in sound production.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of biomedical engineering, flexible sensing technology and artificial intelligence, and specifically to a method for human voice compensation based on laser-induced graphene. Background Technology

[0002] Dysphonia is a type of speech dysfunction caused by factors such as post-laryngeal surgery, vocal cord damage, and neurogenic lesions. Patients often face difficulties such as unclear pronunciation, distorted tone, and even inability to communicate normally, severely impacting their daily lives and social interactions. Speech is essentially a mechanical wave formed by the vibration of laryngeal tissues transmitted through the air. When a person speaks, the skin surface of the neck generates vibration signals highly correlated with speech information. Capturing and accurately reproducing these vibration signals is the core research direction for assistive voice technology. With the increasing aging population and rising incidence of laryngeal diseases, the clinical demand for efficient, non-invasive, and highly adaptable assistive voice technologies is becoming increasingly urgent.

[0003] Currently, assisted speech technologies are mainly divided into two categories: invasive and non-invasive. Invasive technologies, represented by vocal cord prostheses, involve surgically implanting a prosthesis into the larynx to replace the function of the vocal cords. While this can partially restore the ability to produce sound, it presents problems such as surgical trauma, postoperative infection risks, and poor adaptability. Furthermore, it cannot be flexibly adjusted according to the individual vocal characteristics of the patient, limiting its applicability. Among non-invasive technologies, artificial larynxes are widely used. They generate sound through bone conduction or electromagnetic excitation, but these devices are mostly rigid structures, making it difficult to fit closely to the contours of the patient's neck. This leads to distortion in vibration signal collection, resulting in a mechanical and unrecognizable output voice. Traditional air conduction microphones rely on air transmission to capture sound, making them susceptible to environmental noise interference. They cannot directly acquire the original vibration signals of the neck tissue, and their voice restoration accuracy is less than 60% in quiet environments. Performance deteriorates further in complex environments, making it difficult to meet the needs of daily communication.

[0004] In the field of vibration sensing materials, piezoelectric sensors are used to capture vibration signals due to their contact-based acquisition advantages. However, limited by the physical properties of the material itself, their frequency response range is relatively narrow, mostly concentrated in the 500Hz-5kHz range. The core frequency band of human voice is 100Hz-8kHz, encompassing key speech components such as low pitch and voiced sounds. This results in the loss of low-frequency signals and attenuation of high-frequency signals, leading to problems such as missing tones and blurred details during speech reconstruction. Furthermore, existing sensing devices mostly employ a single structural design, which cannot achieve accurate capture of vibration signals across the entire frequency range, further limiting the improvement of auxiliary sound generation effects.

[0005] Laser-induced graphene (LIG), a novel flexible carbon material, possesses excellent mechanical flexibility, high sensitivity, and customizable microstructure properties. Its fabrication process does not require a high-temperature, high-vacuum environment; it can be directly generated on a polyimide (PI) substrate using lasers, and the microstructure morphology can be flexibly controlled through laser parameters. Studies have shown that LIG microstructures (porous, fibrous, and graphite structures) have different intrinsic frequencies. Porous structures, due to their high specific surface area and elastic properties, are sensitive to high-frequency vibrations, while fibrous structures, with their large aspect ratio and good flexibility, are suitable for capturing low-frequency vibration signals. This provides a novel material basis for wide-band vibration acquisition. However, the application of current LIG technology in assisted sound generation is still lacking. Microstructure optimization design for the human voice frequency band is not yet implemented, there is a lack of technical solutions for accurately matching the frequency response of LIG microstructures with the human voice frequency band, and the possibility of achieving full-band vibration acquisition through gradient microstructure design has not been explored. Therefore, the advantages of LIG materials in flexible sensing have not been fully utilized.

[0006] In terms of speech restoration technology, existing solutions mostly rely on simple electroacoustic conversion of single vibration signals, without incorporating artificial intelligence algorithms for in-depth processing and personalized correction of multi-frequency vibration signals. Due to significant individual differences in neck contours and vocal habits among patients, a single conversion model struggles to adapt to diverse needs, resulting in low fluency and significant semantic deviations in the restored speech, failing to meet the naturalness requirements of daily communication. Furthermore, existing technologies lack an integrated design of "vibration acquisition - frequency band matching - intelligent restoration," with each stage operating independently. This leads to signal loss and distortion during transmission, further reducing the overall performance of the assisted speech system.

[0007] In summary, existing assisted speech technologies suffer from several significant drawbacks: invasive methods are highly traumatic and have poor adaptability, while non-invasive methods lack adequate fit, have narrow frequency coverage, and weak anti-interference capabilities. Sensing materials fail to achieve precise matching with the human voice frequency band, making it impossible to efficiently capture vibration signals across the entire frequency range. Speech reconstruction lacks intelligent and personalized correction mechanisms, resulting in fluency and recognizability that fall short of practical needs. The core reason for these problems lies in the failure to integrate the microstructure design of flexible sensing materials with the vibration characteristics of the human voice, and the lack of an integrated technical approach that achieves multi-stage collaborative optimization. Therefore, developing a flexible, wide-band, and intelligent voice compensation method based on LIG (Liquidity-Induced Geometric Governance) to achieve accurate vibration signal acquisition and natural speech reconstruction is crucial for solving communication problems for patients with speech difficulties, and has significant clinical application value and market prospects. Summary of the Invention

[0008] To address the problems existing in the prior art, the present invention aims to provide a laser-induced graphene-based voice compensation method. Through a collaborative design of the entire process—from customized LIG microphone fabrication to broadband vibration acquisition and AI-based voice restoration—it enables patients with speech difficulties to achieve fluent pronunciation. A method for human voice compensation based on laser-induced graphene, comprising the following steps: Step 1: Acquisition and frequency band analysis of patient vocal characteristics Step 1-1: Preset speech sample library construction: Select commonly used sentences covering different tones, speaking speeds and semantics to form a preset speech library containing several sets of samples. Each set of samples is a few seconds long, ensuring coverage of the core frequency band of human voice from 100Hz to 8kHz. Steps 1-2: Select a laser vibrometer with an accuracy of 0.01μm as a reference sensor, place it against the area of ​​strongest sound vibration on both sides of the thyroid cartilage in the patient's neck, fix the sensor position and mark it; guide the patient to read the sentences in the preset speech library in sequence, repeat each sentence multiple times, and the sensor collects the neck skin vibration signal at a sampling frequency of 20kHz, and simultaneously records the patient's breathing state and neck muscle tension during the pronunciation. Steps 1-3: Transmit the collected vibration signals to the MATLAB data analysis platform, perform spectrum analysis using Fast Fourier Transform, and plot the "frequency-vibration amplitude" curve; Steps 1-4: Extract the main vibration frequency bands and divide them into the low frequency band (100Hz-500Hz) corresponding to voiced sounds and low tones, the mid frequency band (500Hz-3kHz) corresponding to basic vowels and consonants, and the high frequency band (3kHz-8kHz) corresponding to voiceless consonants and tonal changes. Steps 1-5: Record the peak frequency, vibration amplitude range, and duration parameters for each frequency band to establish a personalized "frequency band-vibration characteristic" database for the patient; Step 2: LIG microstructure design and parameter optimization Step 2-1: Select microstructure scheme based on coverage frequency band Step 2-2: If Step 2-1 only needs to cover a single frequency band: for the high frequency band 3kHz-8kHz, use porous graphene microstructures to improve high frequency response sensitivity by utilizing their high specific surface area and elastic modulus; for the low frequency band 100Hz-500Hz, use graphene fiber microstructures to enhance low frequency capture capability by utilizing their large aspect ratio and flexibility. Step 2-3: If step 2-1 needs to cover the entire frequency band: design a gradient LIG microstructure array, draw N parallel LIG lines, and increase the laser scanning speed from the 1st line to the Nth line in a gradient of 20mm / s-500mm / s, so that each LIG line forms a microstructure with a gradient of inherent frequency, which together cover the entire frequency band from 100Hz to 8kHz. Step 2-4: If Step 2-1 requires coverage of mixed frequency bands: Divide the LIG pickup area according to the proportion of each frequency band, with the high frequency area accounting for 30%, the mid frequency area accounting for 50%, and the low frequency area accounting for 20%, corresponding to porous graphene microstructure, porous and fiber hybrid structure, and fiber microstructure, respectively. Steps 2-5: Using PI film, three types of samples are processed: porous graphene, graphene fiber, and gradient array. Steps 2-6: Use a vibration table to simulate vibrations of different frequencies and amplitudes within the range of 100Hz-8kHz, record the sample response amplitude using a laser vibrometer, and plot the "frequency-response sensitivity" curve. Steps 2-7: If the response sensitivity of a certain frequency band is lower than the preset value, adjust the laser parameters of the corresponding microstructure and repeat the test until the response sensitivity of each frequency band is not lower than the preset value. Establish a correspondence table of "frequency band-microstructure-laser parameters". Step 3: LIG pickup fabrication and packaging Step 3-1: Select a flexible PI film and cut it to fit the size of an adult's neck; Step 3-2: Wipe with alcohol, clean with ultrasound, and dry in sequence to ensure the surface of the PI film substrate is clean; Step 3-3: Transmittance Detection: Detect the transmittance of the PI film to a 1064nm laser; the transmittance should not be less than 90%. Steps 3-4: LIG microstructure fabrication: Using a 1064nm near-infrared laser, porous graphene, graphene fibers, or gradient arrays are fabricated according to the determined parameters. Steps 3-5: Electrode fabrication: Electrodes are formed at both ends of the LIG microstructure using silver paste printing; Steps 3-6: Wire connection: Attach the copper wire to the electrode and encapsulate it with epoxy resin; Steps 3-7: Continuity test: Use a multimeter to test the continuity between the wire and the LIG microstructure; Steps 3-8: Encapsulation: The entire microphone is encapsulated and cured using medical-grade silicone rubber; Steps 3-9: Fit Test: Fit the packaged microphone to the neck model and test its contact area under 5N pressure to ensure that the fit area is not less than 95%; Step 4: Training the AI ​​speech restoration model Step 4-1: Collect paired data of neck vibration signals and speech signals from several groups of normal people as the basic dataset; Step 4-2: Collect several sets of LIG microphone vibration signals and standard speech pairing data from patients to create personalized datasets; Step 4-3: Data preprocessing: Filter and normalize the vibration signal, extract MFCC features from the speech signal, and form training samples; Step 4-4: Using a CNN+LSTM hybrid model, extract the spatial and temporal features of the vibration signal. The fully connected layer outputs the speech MFCC features. Steps 4-5: Model Training and Validation: Train the model according to the set parameters and adjust the model structure based on the validation results; Steps 4-6: Personalized fine-tuning: Fine-tune the model using transfer learning until the accuracy of restoring speech semantics reaches over 98%; Step 5: Assembly and debugging of the voice compensation system Step 5-1: System Connection: Connect the LIG microphone to the data processing module and connect it to the speaker wirelessly; Step 5-2: Model Implantation: Deploy the trained AI speech restoration model to the data processing module; Step 5-3: Performance Testing: Test frequency band coverage, semantic accuracy, and wearing comfort; Step 5-4: System Optimization: If the detection indicators do not meet the standards, return to the corresponding step for adjustment; Step 5-5: Dynamic Updates: The system continuously records data during use and performs periodic optimizations based on changes in the patient's vocalization status.

[0009] Preferably, during steps 1-2, the ambient noise level is kept below 30dB to avoid interfering with the acquisition of vibration signals from the neck skin.

[0010] Preferably, in steps 2-3, N ≥ 30 parallel LIG lines are drawn, with a line width of 10 μm and a line spacing of 5 μm. By setting no less than 30 parallel LIG lines, the number of effective sensing units can be significantly increased, thereby enhancing the stability and repeatability of the overall signal output of the device. At the same time, controlling the line width to 10 μm and the line spacing to 5 μm is beneficial for achieving high-density integration within a limited area. While ensuring electrical isolation between adjacent LIG lines, the number of current-carrying channels and specific surface area are increased, thereby improving the sensitivity and signal-to-noise ratio of the device, while taking into account both microstructure consistency and fabrication controllability.

[0011] Preferably, in steps 2-5, the power density for processing porous graphene is 18000 W / cm², and the scanning speed is 20-150 mm / s; the power density for processing graphene fibers is 28000 W / cm², and the scanning speed is 300-500 mm / s; and the power density for processing gradient arrays is 20000 W / cm², and the scanning speed is 150-300 mm / s. By matching corresponding power density and scanning speed parameters to address the differences in energy input and heat accumulation characteristics of different microstructures, precise control of the graphene microstructure morphology can be achieved. Specifically, using a lower scanning speed and moderate power density for porous graphene is beneficial for forming a uniformly distributed pore structure and increasing the specific surface area; using a higher power density and high-speed scanning for graphene fibers helps to achieve continuous and dense fibrous conductive channels; and the gradient array, through a combination of moderate power density and scanning speed, achieves a smooth transition of structural parameters, thereby improving the consistency of the overall structure and the controllability of the functional gradient, and enhancing the overall performance of the device.

[0012] Preferably, the preset value in steps 2-7 is 0.1 V / g. Setting the preset value to 0.1 V / g can ensure signal resolution while avoiding oversaturation of the output signal, which is beneficial for achieving accurate response to small accelerations or weak external stimuli. This preset value takes into account both sensitivity and dynamic range, so that the device can maintain good linearity and stability under different working conditions, thereby improving the reliability and practicality of the test results.

[0013] Preferably, in steps 3-4, a 1064nm near-infrared laser is used to process porous graphene, graphene fibers, or gradient arrays according to determined parameters. The specific parameter settings are as follows: The parameters for processing porous graphene were set as follows: spot diameter 30μm, power density 18000-22000W / cm², scanning speed 20-150mm / s, and scanning overlap rate 6%. The parameters for processing graphene fibers were set as follows: spot diameter 30μm, power density 25000-30000W / cm², scanning speed 300-500mm / s, and scanning overlap rate 6%. The processing parameters for the gradient array region were set as follows: spot diameter 30μm, power density 20000W / cm², scanning speed 150-300mm / s, and scanning overlap rate 6%; 1064nm near-infrared laser was used for processing.

[0014] 1064 nm near-infrared lasers exhibit excellent energy coupling characteristics with polymer substrates, enabling efficient and controllable laser-induced graphitization processes without the introduction of additional chemical reagents. The use of a uniform 30 μm spot diameter and a 6% scanning overlap ensures uniform energy distribution between adjacent scanning trajectories, preventing structural defects. Through this parameter combination, various structural morphologies can be stably obtained on the same substrate, improving fabrication consistency and overall device performance.

[0015] Preferably, the data processing module in step 5-1 includes a signal amplification unit, a filtering unit, and an analog-to-digital conversion unit.

[0016] Preferably, the detection indicators in step 5-4 include: frequency band coverage integrity, i.e., signal acquisition rate of each frequency band from 100Hz to 8kHz ≥98%; speech restoration accuracy, i.e., semantic accuracy ≥98%; and wearing comfort, i.e., no obvious pressure after wearing continuously for 2 hours.

[0017] Compared with the prior art, the present invention has the following advantages: 1) Traditional microphones, due to their rigid structure, cannot adapt to the curve of the neck, resulting in gaps that cause vibration signal attenuation. The LIG microphone of this invention uses a flexible PI film material to fit closely to the skin, accurately capturing subtle and low-frequency vibrations, improving the acquisition accuracy by more than 30% compared to traditional microphones, and solving the pain point of "signal loss".

[0018] 2) Traditional microphones have narrow frequency coverage (e.g., only 300Hz-800Hz), resulting in the loss of high and low frequency signals. This invention utilizes a customized LIG microstructure: a porous graphene microstructure enhances the response at high frequencies, while a fiber microstructure reduces loss at low frequencies. Combined with microstructure response testing, it covers the 50Hz-1500Hz frequency band, twice as wide as traditional microphones, ensuring a sense of sound layering.

[0019] 3) Traditional techniques rely on fixed algorithms or templates, resulting in mechanical sound reproduction and low recognizability. The AI ​​model of this invention undergoes closed-loop training of "collection-comparison-correction" to adapt to changes in vocalization state, achieving a sound reproduction fluency of 92 points (traditional 65 points) and a recognition rate of over 95%, enabling normal communication.

[0020] 4) Traditional methods rely on multiple microphone combinations, resulting in large size and frequency gaps. This invention uses a "gradient scanning LIG line" design, enabling seamless acquisition across the entire frequency band with a single microphone. Simultaneously, the "spectrum analysis - customization - AI training" process adapts to different patients, increasing the compatibility rate from the traditional 30% to 90%. Attached Figure Description

[0021] Figure 1 Flowchart of patient vocal feature acquisition and frequency band analysis process.

[0022] Figure 2 Schematic diagram of LIG microstructure design and parameter optimization process.

[0023] Figure 3 Schematic diagram of LIG pickup fabrication and packaging process.

[0024] Figure 4 A schematic diagram of the training process for an artificial intelligence speech restoration model.

[0025] Figure 5 Schematic diagram of the assembly and debugging process of the human voice compensation system. Detailed Implementation

[0026] This invention provides a voice compensation method based on laser-induced graphene for patients with speech difficulties. By using a flexible laser-induced graphene (LIG) microphone for precise vibration acquisition, wide-band microstructure adaptation, and artificial intelligence voice restoration technology, it solves the problems of poor fit, narrow frequency band, and low voice recognition of traditional devices. The invention is further described in detail with reference to the accompanying drawings.

[0027] Step 1: Acquisition and frequency band analysis of patient vocal characteristics, such as... Figure 1 As shown, the specific steps are as follows: Step 1-1: Construct a pre-defined speech sample library A speech sample library containing 500 sets of sentences is pre-built. The sentences cover different tones (high level tone, rising tone, falling-rising tone, falling tone), speech speed (slow, medium, fast) and semantic types (daily communication, emotional expression, professional terminology). Each set of sentences is 3 to 5 seconds long, so that the sample speech audio segments completely cover the core frequency band of human voice from 100Hz to 8kHz.

[0028] Step 1-2: Neck vibration signal acquisition A laser vibrometer with an accuracy of 0.01 μm was used and fixed to pre-marked positions on both sides of the thyroid cartilage in the patient's neck. The patient was guided to read aloud sentences from the aforementioned speech sample library, with each sentence repeated three times. Vibration signals from the neck skin surface were acquired at a sampling frequency of 20 kHz, and the patient's respiratory status and neck muscle tension were recorded simultaneously. During the acquisition process, the ambient noise was controlled to be no higher than 30 dB to avoid interfering with the acquisition of vibration signals.

[0029] Steps 1-3: Vibration signal spectrum analysis The collected vibration signals are input into the MATLAB analysis platform, and the frequency-vibration amplitude curve is obtained by performing spectrum analysis on the signals through Fast Fourier Transform (FFT).

[0030] Steps 1-4: Frequency Band Allocation Based on the spectral analysis results, the human voice vibration signal is divided into the low frequency band (100Hz~500Hz, corresponding to voiced sounds and low pitch), the mid frequency band (500Hz~3kHz, corresponding to basic vowels and consonants) and the high frequency band (3kHz~8kHz, corresponding to voiceless consonants and tone changes).

[0031] Steps 1-5: Building a Personalized Feature Database Record the peak frequency, vibration amplitude range, and duration parameters corresponding to each frequency band to establish a personalized "frequency band-vibration characteristic" database for each patient.

[0032] Step 2: LIG microstructure design and parameter optimization, such as Figure 2 As shown, the specific steps are as follows: Step 2-1: Determine the target coverage frequency band Based on the vibration characteristic database established in steps 1-5, determine the frequency band type that the LIG pickup needs to cover, which can be a single frequency band, a mixed frequency band, or a full frequency band.

[0033] Step 2-2: Selection of microstructure for a single frequency band When only high-frequency coverage is required, porous laser-induced graphene microstructures are selected to enhance high-frequency response sensitivity by utilizing their high specific surface area and elastic modulus; when only low-frequency coverage is required, fibrous laser-induced graphene microstructures are selected to enhance low-frequency capture capability by utilizing their large aspect ratio (≥100:1) and flexibility.

[0034] Steps 2-3: Full-band microstructure design When it is necessary to cover the entire frequency band from 100Hz to 8kHz, a gradient LIG microstructure array is designed. No less than 30 parallel LIG lines are drawn on the PI substrate, each with a line width of 10μm and a line spacing of 5μm. The laser scanning speed increases from 5mm / s to 100mm / s along the array direction, thereby forming an inherent frequency gradient distribution.

[0035] Steps 2-4: Mixed frequency band region division When there is a need for mixed frequency bands, the microphone area is divided into high-frequency, mid-frequency and low-frequency areas, with area ratios of 30%, 50% and 20% respectively, and porous graphene microstructures, hybrid structures (porous + fiber) and fiber microstructures are set respectively.

[0036] Steps 2-5: Sample Processing Three types of LIG samples were processed using 100μm thick PI films: porous graphene (power density 18000W / cm², scanning speed 20-150mm / s), graphene fiber (power density 28000W / cm², scanning speed 300-500mm / s), and gradient array (power density 20000W / cm², scanning speed 150-300mm / s).

[0037] Steps 2-6: Frequency Response Test Vibration tables were used to simulate vibrations of different frequencies and amplitudes within the range of 100Hz to 8kHz. The response amplitude of each sample was tested using a laser vibrometer, and frequency-response sensitivity curves were plotted.

[0038] Steps 2-7: Parameter Optimization When the response sensitivity of a certain frequency band is lower than 0.1V / g, adjust the corresponding laser power density or scanning speed (the power density can be increased to 22000W / cm² for porous structures and the scanning speed can be reduced to 80mm / s for fiber structures), repeat the test until the sensitivity of each frequency band is not lower than 0.1V / g, and establish a correspondence table of "frequency band - microstructure - laser parameters".

[0039] Step 3: LIG pickup fabrication and packaging, such as... Figure 3 As shown, the specific steps are as follows: Step 3-1: Substrate Preparation A flexible PI film with a thickness of 80μm was selected and cut into 60mm×40mm dimensions.

[0040] Step 3-2: Cleaning treatment Wipe the surface three times in the same direction with a cotton ball soaked in 95% alcohol to remove oil and dust; place it in a 40kHz ultrasonic cleaner and rinse with deionized water for 15 minutes; place it in a 60℃ drying oven to dry for 30 minutes, ensuring that the surface moisture content is ≤0.1%.

[0041] Step 3-3: Use an infrared pulsed laser, maintain an atmospheric environment during processing, control laser energy fluctuations within ±5% to ensure microstructure consistency, and set processing parameters according to the "frequency band-microstructure-laser parameter" correspondence table.

[0042] Steps 3-4: Set the processing parameters for the porous graphene region as follows: spot diameter 30μm, power density 18000-22000W / cm², scanning speed 20-150mm / s, scanning overlap rate 6%; use 1064nm near-infrared laser for processing.

[0043] Steps 3-5: Set the parameters for processing the graphene fiber area as follows: spot diameter 30μm, power density 25000-30000W / cm², scanning speed 300-500mm / s, scanning overlap rate 6%; use 1064nm near-infrared laser for processing.

[0044] Steps 3-6: Set the processing parameters for the gradient array area as follows: spot diameter 30μm, power density 20000W / cm², scanning speed 150-300mm / s, scanning overlap rate 6%; use 1064nm near-infrared laser for processing.

[0045] Steps 3-7: Fabricate rectangular electrodes (5mm×2mm) at both ends of the LIG microstructure using silver paste printing process, with a thickness of 10μm.

[0046] Steps 3-8: Attach a 0.1mm diameter flexible copper wire to the electrode and seal it with epoxy glue to ensure that the contact resistance between the wire and the electrode is ≤1Ω.

[0047] Steps 3-9: Test continuity: Use a multimeter to test the continuity between the microstructure and the wires to ensure there are no open circuits or short circuits.

[0048] Steps 3-10: Place the microphone against the human neck model and apply 5N pressure. The contact area between the microphone and the model should be ≥95% and there should be no obvious wrinkles.

[0049] Step 4: Training the AI ​​speech restoration model, such as... Figure 4 As shown, the specific steps are as follows: Step 4-1: Building the Basic Dataset We collected 1,000 pairs of neck vibration signal-voice signal data from normal individuals (covering different genders, ages, and accents) as the base dataset.

[0050] Step 4-2: Personalized Data Collection Two hundred sets of paired data of "LIG pickup vibration signal - standard speech" were collected from patients (patients read a preset speech library and recorded standard speech at the same time) as a personalized dataset.

[0051] Step 4-3: Data Preprocessing The vibration signal is filtered (100Hz high-pass + 8kHz low-pass) and normalized; the speech signal is subjected to MFCC feature extraction to form training samples.

[0052] Step 4-4: Model Building A CNN+LSTM hybrid model is constructed. The CNN layer (3 convolutional layers + 2 pooling layers) extracts the spatial features of the vibration signal, the LSTM layer (2 bidirectional LSTM layers) captures the temporal features, and the fully connected layer outputs the speech MFCC features.

[0053] Steps 4-5: Set the training parameters as follows: batch size=32, learning rate=0.001, number of iterations=1500, and the loss function is mean squared error (MSE). Steps 4-6: Validate once every 200 iterations. If the speech similarity of the validation set is less than 85%, adjust the number of neurons in the LSTM layer (from 128 to 256) until the speech similarity of the validation set is ≥90%.

[0054] Steps 4-7: Have the patient wear a LIG microphone and read 100 sentences that were not used in training. Collect vibration signals and input them into the model to obtain preliminary restored speech. Steps 4-8: Compare the initially restored speech with the standard speech, calculate the tone error, speech rate error and semantic accuracy. If the semantic accuracy is less than 95%, fine-tune the model parameters through transfer learning (freeze the CNN layers and train only the LSTM layers and fully connected layers). Steps 4-9: Repeat the correction 3 times to ensure that the final restored speech matches the patient's expected semantics by ≥98%.

[0055] Step 5: Assemble and debug the voice compensation system, such as... Figure 5 As shown, the specific steps are as follows: Step 5-1: Hardware connection: The LIG pickup is connected to the data processing module (including signal amplification unit, filtering unit, and analog-to-digital conversion unit) via a flexible cable. The data processing module is connected to the portable speaker (frequency response range 100Hz-8kHz) via Bluetooth. Step 5-2: Implant the trained AI speech restoration model into the data processing module, and set the signal processing delay to ≤10ms to ensure real-time performance; Step 5-3: The patient wears the LIG microphone (fixed at the marked position on the neck), turns on the system, and reads aloud the preset speech library and daily sentences in sequence, recording the restored speech; Step 5-4: The test indicators include: frequency band coverage integrity (signal acquisition rate of each frequency band from 100Hz to 8kHz ≥98%), voice restoration accuracy (semantic accuracy ≥98%), and wearing comfort (no obvious pressure after continuous wearing for 2 hours). Step 5-5: If a certain indicator fails to meet the standard, make targeted adjustments (e.g., if the frequency band coverage is incomplete, optimize the LIG microstructure parameters; if the speech accuracy is insufficient, supplement the training data) until all indicators meet the requirements. Steps 5-6: The system automatically records the "vibration signal - restored speech" data during the patient's daily use, and initiates dynamic model optimization every 500 sets of data. Steps 5-7: If the patient's vocal state changes (such as postoperative recovery or adjustment of condition), return to step 1 to re-acquire vibration characteristics, adjust the LIG microstructure design and model parameters, and ensure that the compensation effect remains stable.

Claims

1. A method for human voice compensation based on laser-induced graphene, characterized in that: Includes the following steps: Step 1: Acquisition and frequency band analysis of patient vocal characteristics Step 1-1: Preset speech sample library construction: Select commonly used sentences covering different tones, speaking speeds and semantics to form a preset speech library containing several sets of samples. Each set of samples is a few seconds long, ensuring coverage of the core frequency band of human voice from 100Hz to 8kHz. Steps 1-2: Select a laser vibrometer with an accuracy of 0.01μm as a reference sensor, place it against the area of ​​strongest sound vibration on both sides of the thyroid cartilage in the patient's neck, fix the sensor position and mark it; guide the patient to read the sentences in the preset speech library in sequence, repeat each sentence multiple times, and the sensor collects the neck skin vibration signal at a sampling frequency of 20kHz, and simultaneously records the patient's breathing state and neck muscle tension during the pronunciation. Steps 1-3: Transmit the collected vibration signals to the MATLAB data analysis platform, perform spectrum analysis using Fast Fourier Transform, and plot the "frequency-vibration amplitude" curve; Steps 1-4: Extract the main vibration frequency bands and divide them into the low frequency band (100Hz-500Hz) corresponding to voiced sounds and low tones, the mid frequency band (500Hz-3kHz) corresponding to basic vowels and consonants, and the high frequency band (3kHz-8kHz) corresponding to voiceless consonants and tonal changes. Steps 1-5: Record the peak frequency, vibration amplitude range, and duration parameters for each frequency band to establish a personalized "frequency band-vibration characteristics" database for the patient; Step 2: LIG microstructure design and parameter optimization Step 2-1: Select microstructure scheme based on coverage frequency band Step 2-2: If Step 2-1 only needs to cover a single frequency band: for the high frequency band 3kHz-8kHz, use porous graphene microstructures to improve high frequency response sensitivity by utilizing their high specific surface area and elastic modulus; for the low frequency band 100Hz-500Hz, use graphene fiber microstructures to enhance low frequency capture capability by utilizing their large aspect ratio and flexibility. Step 2-3: If step 2-1 needs to cover the entire frequency band: design a gradient LIG microstructure array, draw N parallel LIG lines, and increase the laser scanning speed from the 1st line to the Nth line in a gradient of 20mm / s-500mm / s, so that each LIG line forms a microstructure with a gradient of inherent frequency, which together cover the entire frequency band from 100Hz to 8kHz. Step 2-4: If Step 2-1 requires coverage of mixed frequency bands: Divide the LIG pickup area according to the proportion of each frequency band, with the high frequency area accounting for 30%, the mid frequency area accounting for 50%, and the low frequency area accounting for 20%, corresponding to porous graphene microstructure, porous and fiber hybrid structure, and fiber microstructure, respectively. Steps 2-5: Using PI film, three types of samples are processed: porous graphene, graphene fiber, and gradient array. Steps 2-6: Use a vibration table to simulate vibrations of different frequencies and amplitudes within the range of 100Hz-8kHz, record the sample response amplitude using a laser vibrometer, and plot the "frequency-response sensitivity" curve; Steps 2-7: If the response sensitivity of a certain frequency band is lower than the preset value, adjust the laser parameters of the corresponding microstructure and repeat the test until the response sensitivity of each frequency band is not lower than the preset value. Establish a correspondence table of "frequency band-microstructure-laser parameters". Step 3: LIG pickup fabrication and packaging Step 3-1: Select a flexible PI film and cut it to fit the size of an adult's neck; Step 3-2: Wipe with alcohol, clean with ultrasound, and dry in sequence to ensure the surface of the PI film substrate is clean; Step 3-3: Transmittance Detection: Detect the transmittance of the PI film to a 1064nm laser; the transmittance should not be less than 90%. Steps 3-4: LIG microstructure fabrication: Using a 1064nm near-infrared laser, porous graphene, graphene fibers, or gradient arrays are fabricated according to the determined parameters. Steps 3-5: Electrode fabrication: Electrodes are formed at both ends of the LIG microstructure using silver paste printing; Steps 3-6: Wire connection: Attach the copper wire to the electrode and encapsulate it with epoxy resin; Steps 3-7: Continuity test: Use a multimeter to test the continuity between the wire and the LIG microstructure; Steps 3-8: Encapsulation: The entire microphone is encapsulated and cured using medical-grade silicone rubber; Steps 3-9: Fit Test: Fit the packaged microphone to the neck model and test its contact area under 5N pressure to ensure that the fit area is not less than 95%; Step 4: Training the AI ​​speech restoration model Step 4-1: Collect paired data of neck vibration signals and speech signals from several groups of normal people as the basic dataset; Step 4-2: Collect several sets of LIG microphone vibration signals and standard speech pairing data from patients to create personalized datasets; Step 4-3: Data preprocessing: Filter and normalize the vibration signal, extract MFCC features from the speech signal, and form training samples; Step 4-4: Using a CNN+LSTM hybrid model, extract the spatial and temporal features of the vibration signal. The fully connected layer outputs the speech MFCC features. Steps 4-5: Model Training and Validation: Train the model according to the set parameters and adjust the model structure based on the validation results; Steps 4-6: Personalized fine-tuning: Fine-tune the model using transfer learning until the accuracy of restoring speech semantics reaches over 98%; Step 5: Assembly and debugging of the voice compensation system Step 5-1: System Connection: Connect the LIG microphone to the data processing module and connect it to the speaker wirelessly; Step 5-2: Model Implantation: Deploy the trained AI speech restoration model to the data processing module; Step 5-3: Performance Testing: Test frequency band coverage, semantic accuracy, and wearing comfort; Step 5-4: System Optimization: If the detection indicators do not meet the standards, return to the corresponding step for adjustment; Step 5-5: Dynamic Updates: The system continuously records data during use and performs periodic optimizations based on changes in the patient's vocalization status.

2. The method for human voice compensation based on laser-induced graphene according to claim 1, characterized in that: During steps 1-2, when collecting vibration signals from the neck skin, keep the ambient noise below 30dB to avoid interfering with the vibration signal acquisition.

3. The method for human voice compensation based on laser-induced graphene according to claim 1, characterized in that: In steps 2-3, draw N≥30 parallel LIG lines with a line width of 10μm and a line spacing of 5μm.

4. The method for human voice compensation based on laser-induced graphene according to claim 1, characterized in that: In steps 2-5, the power density for processing porous graphene is 18000W / cm², and the scanning speed is 20-150mm / s; the power density for processing graphene fibers is 28000W / cm², and the scanning speed is 300-500mm / s; and the power density for processing gradient arrays is 20000W / cm², and the scanning speed is 150-300mm / s.

5. A method for human voice compensation based on laser-induced graphene according to claim 1, characterized in that: The preset value in steps 2-7 is 0.1V / g.

6. The method for human voice compensation based on laser-induced graphene according to claim 1, characterized in that: In steps 3-4, a 1064nm near-infrared laser is used to process porous graphene, graphene fibers, or gradient arrays according to determined parameters. The specific parameter settings are as follows: The parameters for processing porous graphene were set as follows: spot diameter 30μm, power density 18000-22000W / cm², scanning speed 20-150mm / s, and scanning overlap rate 6%. The parameters for processing graphene fibers were set as follows: spot diameter 30μm, power density 25000-30000W / cm², scanning speed 300-500mm / s, and scanning overlap rate 6%. The processing parameters for the gradient array region were set as follows: spot diameter 30μm, power density 20000W / cm², scanning speed 150-300mm / s, and scanning overlap rate 6%; 1064nm near-infrared laser was used for processing.

7. The method for human voice compensation based on laser-induced graphene according to claim 1, characterized in that: The data processing module in step 5-1 includes a signal amplification unit, a filtering unit, and an analog-to-digital conversion unit.

8. A method for human voice compensation based on laser-induced graphene according to claim 1, characterized in that: The testing indicators in step 5-4 include: frequency band coverage integrity (signal acquisition rate of each frequency band from 100Hz to 8kHz ≥98%), speech restoration accuracy (semantic accuracy ≥98%), and wearing comfort (no obvious pressure after continuous wearing for 2 hours).