A Smart Analysis and Management Method for Post-Throat Surgery Recovery Information
By collecting multi-source sensing data and constructing a muscle vibration coefficient matrix and a transfer learning model, the risk of postoperative recovery in pharyngeal surgery is predicted, and a personalized recovery path is generated. This solves the problems of data lack and abuse of intervention in the postoperative recovery process of pharyngeal surgery in existing technologies, and achieves precise monitoring and personalized intervention, thereby reducing the risk of postoperative complications.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-26
- Publication Date
- 2026-04-03
AI Technical Summary
Existing technologies are insufficient for precise monitoring and personalized intervention during the recovery process after pharyngeal surgery, especially during the recovery period after patients return home. Problems such as lack of data, delayed response, and overuse of interventions are prominent, leading to increased risks such as secondary damage to the local mucosa or postoperative vocal cord adhesions.
Multi-source sensor data are collected, including speech signals, pharyngeal surface micro-vibration signals, airway pressure disturbance signals during swallowing, and environmental aerosol particle concentration data. By constructing a muscle vibration coefficient matrix, asynchronous matching analysis, and transfer learning models, postoperative recovery risks are predicted, and personalized recovery paths are generated.
It enables multidimensional dynamic modeling of the recovery status after pharyngeal surgery, improves the sensitivity of abnormal state identification, reduces the incidence of postoperative complications, improves the efficiency of medical resource utilization, and realizes a personalized remote management mode.
Smart Images

Figure CN121416095B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of medical information analysis and health management technology, specifically to an intelligent analysis and management method for post-pharyngeal surgery recovery information. Background Technology
[0002] With the development of minimally invasive surgical techniques in otolaryngology, surgeries such as vocal cord polyp removal and laryngeal cancer resection are becoming increasingly common. Postoperatively, patients often face challenges during recovery, including difficulty speaking, recurring local tissue edema, unstable swallowing pain, and a high risk of postoperative infection. Current clinical management relies primarily on physician experience and patient feedback, making precise monitoring and personalized intervention difficult. This is especially true during the post-discharge recovery period, where data scarcity, delayed response, and misuse of interventions are particularly prominent.
[0003] Especially in the early stages after minimally invasive vocal cord surgery, the recovery of the pharyngeal region requires precise control of multiple parameters such as speech frequency, swallowing frequency, and environmental humidity. Improper control can lead to secondary damage to the local mucosa or chronic edema, and in severe cases, even postoperative vocal cord adhesion or permanent damage to voice quality. Currently, there is no intelligent management method for this postoperative recovery stage that can integrate multi-source data to assess recovery risks and intelligently identify potential deterioration trends to provide personalized feedback. Summary of the Invention
[0004] The purpose of this invention is to provide an intelligent analysis and management method for post-pharyngeal surgery recovery information to address the shortcomings of the prior art.
[0005] To achieve the above objectives, the present invention provides the following technical solution: an intelligent analysis and management method for post-pharyngeal surgery recovery information, comprising:
[0006] S100 collects multi-source sensory data of the target patient during the postoperative recovery period, including voice signals, pharyngeal surface micro-vibration signals, airway pressure disturbance signals during swallowing, and environmental aerosol particle concentration data.
[0007] S200, based on the cross-fitting of pharyngeal surface micro-vibration signals and airway pressure disturbance signals during swallowing, constructs a muscle vibration coefficient matrix P that reflects the tension fluctuation of the glottic muscle group;
[0008] S300, perform asynchronous matching analysis between the speech signal and the muscle vibration coefficient matrix P to extract the non-physiological vocal load factor F;
[0009] S400, based on the synergistic analysis of environmental aerosol particle concentration data and non-physiological vocalization load factor F, identifies the air quality interference intensity R that may induce chronic edema or local infection, and constructs the non-physiological vocalization load factor F and air quality interference intensity R into a postoperative glottic risk coefficient vector V.
[0010] S500, the risk coefficient vector V is input into the transfer learning model M trained by the postoperative recovery curve to obtain the abnormal risk prediction value W of the target patient in the future Δt time.
[0011] When the abnormal risk prediction value W exceeds the adaptive warning threshold T, the S600 generates a personalized recovery path that includes the intervention level, intervention method, and voice environment suggestions, and pushes the path to the patient's smart terminal and the attending physician's platform.
[0012] Preferably, wherein S200 includes:
[0013] S201, perform Hilbert transform on the denoised pharyngeal surface micro-vibration signal to extract the envelope feature curve and the trend of the main frequency variation of the signal;
[0014] S202, the airway pressure disturbance signal collected during swallowing is normalized in the time domain to obtain the standard pressure change profile.
[0015] S203 uses a dynamic time warping algorithm to perform nonlinear matching between the microseismic signal envelope curve and the airway pressure profile to establish an action alignment relationship.
[0016] S204. Based on the matching time points, construct the muscle vibration coefficient matrix P. Each element in the matrix is calculated from the micro-vibration amplitude and pressure disturbance gradient at the corresponding time.
[0017] Preferably, S300 includes:
[0018] S301 performs short-time energy extraction and spectral entropy analysis on postoperative speech signals to identify high-load vocal segments within the speech cycle.
[0019] S302, map the muscle tremor coefficient matrix P to the speech signal time axis, use a sliding window method to perform asynchronous feature association, and extract the muscle tension intensity sequence corresponding to the vocal segment;
[0020] S303, calculate the correlation index between the muscle tone intensity sequence and the speech energy curve in each vocal segment, and identify the mismatched segments where the muscle tone is abnormally higher than the speech output power;
[0021] S304, which combines the proportion of mismatched segments with the cumulative amplitude value, is defined as the non-physiological vocal load factor F.
[0022] Preferably, the S400 includes:
[0023] S401, extract the air disturbance intensity index D within the target monitoring period, and adjust the calculation according to the particulate matter size sensitivity factor and exposure time to obtain the air quality disturbance intensity R;
[0024] S402, the non-physiological vocalization load factor F and the air quality interference intensity R are bivariate orthogonally normalized to construct a set of joint influencing factors;
[0025] S403, based on the risk sensitization function fitted by the clinical retrospective sample, input F and R into the surface response model to obtain the nonlinear coupling effect value as the glottal risk weight;
[0026] S404, combine the glottal risk weight with the normalized F and R to construct the postoperative glottal risk coefficient vector V.
[0027] Preferably, the S500 includes:
[0028] S501, construct a time-series labeled sample set containing multi-stage postoperative recovery curves, using patient voice intensity fluctuations, muscle tone recovery trajectories, and historical environmental disturbance data as input features;
[0029] S502 uses a long short-term memory network as the basic model, completes the initial training in the source data domain, and achieves the transfer to the target patient data domain through parameter freezing and feature remapping mechanisms.
[0030] S503 uses the current patient's risk coefficient vector V as the latest input to the time series window, and combines it with historical feature trajectories to complete the recovery trend prediction.
[0031] S504, the model outputs the abnormal risk prediction value W within the time window Δt, which represents the probability that the patient will develop glottic edema or delayed recovery within the prediction period.
[0032] Preferably, the method for acquiring the patient's voice intensity fluctuations, muscle tone recovery trajectories, and historical environmental disturbance data includes:
[0033] S511 acquires continuous speech signals and extracts the speech intensity value for each cycle based on a sliding time window to construct a speech intensity time-series curve;
[0034] S512, principal component analysis was performed on the daily collected muscle vibration coefficient matrix P to extract representative tension features and construct muscle tension recovery trajectory curves in chronological order;
[0035] S513 synchronously records the changes in aerosol concentration in the air environment where the patient is located at different time periods, and constructs an environmental disturbance time series by combining location tags.
[0036] Preferably, the S600 includes:
[0037] S601, dynamically sets the warning threshold T based on the historical recovery database, and compares it with the current predicted value W to identify the intervention trigger state;
[0038] S602 determines the corresponding intervention level based on the patient risk level mapping table, and divides it into three categories: "suggestion of self-adjustment", "remote guidance intervention" and "suggestion of outpatient follow-up".
[0039] S603, combining patient speech intensity trends, muscle tone recovery curves and environmental exposure records, generates intervention suggestions and speech use environment adjustment instructions;
[0040] S604 pushes the generated recovery path in structured data format to both the patient's mobile terminal and the attending physician's platform simultaneously, enabling integrated rehabilitation management.
[0041] The technical effects and advantages provided by the present invention in the above technical solution are as follows:
[0042] 1. This invention achieves multidimensional dynamic modeling of the postoperative recovery status of the pharynx by integrating speech signal analysis, pharyngeal micro-vibration signal processing, airway pressure disturbance identification, and environmental particle concentration monitoring. It constructs a temporal feature set composed of speech intensity fluctuations, muscle tone recovery trajectories, and environmental disturbance history, and uses a transfer learning model to complete individualized recovery trend prediction and risk warning. Compared to existing methods that rely solely on single speech or behavioral data for evaluation, this invention introduces a heterogeneous data collaborative modeling mechanism of physiological signals and environmental factors, significantly improving the sensitivity of identifying abnormal postoperative recovery states (such as glottic edema and mucosal damage).
[0043] 2. The risk coefficient vector construction method and intervention level-driven path generation mechanism proposed in this invention can automatically generate a structured rehabilitation path containing intervention methods, intervention levels, and voice environment adjustment suggestions based on the patient's current status and historical trends. This path is then simultaneously pushed to the patient and attending physician's platforms in a linked manner, enabling a remote and personalized postoperative management model. This method not only improves the efficiency of medical resource utilization but also effectively reduces the incidence of postoperative complications, demonstrating strong clinical application value. Attached Figure Description
[0044] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this invention. For those skilled in the art, other drawings can be obtained based on these drawings.
[0045] Figure 1 This is a flowchart of the method of the present invention. Detailed Implementation
[0046] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0047] For examples, please refer to Figure 1 As shown in this embodiment, an intelligent analysis and management method for post-pharyngeal surgery recovery information includes:
[0048] S100 collects multi-source sensory data from the target patient during the postoperative recovery period, including voice signals, pharyngeal surface micro-vibration signals, airway pressure disturbance signals during swallowing, and environmental aerosol particle concentration data.
[0049] During the procedure, a high-sensitivity micro-vibration sensor array installed in the anterior region of the patient's neck is used to collect micro-vibration signals from the pharyngeal surface. This sensor array includes at least three miniature inertial measurement units distributed on both sides and in the center of the thyroid cartilage. The sensor sampling frequency is no less than 2000 Hz, used to capture high-frequency, low-amplitude vibration signals caused by micro-contractions of the pharyngeal muscles.
[0050] The acquired raw microseismic signals are first decomposed into multiple scales using a wavelet packet decomposition algorithm to distinguish between high-frequency noise and physiological vibration components. The noise threshold is determined based on the baseline vibration standard deviation of the patient at rest, typically set as the baseline mean plus twice the standard deviation. When the vibration amplitude exceeds this threshold, it is considered a valid muscle movement signal.
[0051] After denoising, the Hilbert transform is applied to calculate the envelope curve of the vibration signal, and the main energy concentration frequency band of the vibration is extracted. This frequency band is used to represent the micro-motion frequency characteristics of the pharyngeal muscle group and will subsequently serve as one of the input parameters for constructing the muscle vibration coefficient matrix.
[0052] In swallowing action recognition and airway disturbance signal acquisition, an integrated wearable air pressure sensing device is used. This device is attached to the area above the trachea in the patient's neck and records the instantaneous pressure changes inside the airway during swallowing through a piezoelectric pressure sensor. The sampling frequency is no less than 100 Hz.
[0053] To identify the timing of swallowing, the rate of change of airway pressure signal was continuously monitored. A first-order difference value greater than a set threshold (defined as 5 times the change in static mean pressure) was set as the trigger condition for the action. After triggering, the signal was sliced into segments within a 1-second interval before and after the trigger as candidate swallowing segments.
[0054] To distinguish between voluntary and involuntary swallowing behaviors, an improved convolutional neural network combined with a long short-term memory network was used to construct an action recognition model. The model was trained on labeled clinical samples, with pressure time series as input and category labels (voluntary / involuntary) as output. The model's accuracy was controlled to be above 95% through cross-validation.
[0055] After identification, only the disturbance signal corresponding to voluntary swallowing is retained for subsequent feature analysis to construct a feature set of muscle movement and airflow interference in swallowing behavior.
[0056] To obtain air pollution parameters in the patient's recovery environment that could potentially affect the quality of throat recovery, an air quality particle monitoring device was installed in the patient's daily activity space to collect real-time aerosol particle concentration data. This device, based on light scattering particle size analysis technology, can identify suspended particulate matter in the air with diameters between 0.3 and 10 micrometers.
[0057] During the data processing phase, the system performs weighted calculations on the number concentration and mass concentration of particles in different particle size ranges to generate particle size concentration distribution curves. To identify particulate components that are highly irritating to the pharyngeal mucosa, the particle size range of the highly sensitive region is set to 2.5 micrometers to 5 micrometers, with a weighting factor of 1.5.
[0058] Based on the particle size weighting result and the cumulative exposure time for the day, the air pollution intensity index D is generated using the following expression: Index D equals the particle size weighted average concentration multiplied by the exposure time multiplied by the sensitivity factor coefficient. The sensitivity factor coefficient is dynamically adjusted based on the patient's postoperative days and the physician's assessment level, ranging from 0.8 to 1.5.
[0059] S200 is based on the cross-fitting of pharyngeal surface micro-vibration signals and airway pressure disturbance signals during swallowing to construct a muscle vibration coefficient matrix P that reflects the tension fluctuation of the glottic muscle group.
[0060] After acquiring the denoised pharyngeal surface microvibration signal, a Hilbert transform is first performed on the signal to extract its envelope features and dominant frequency variation trend. The specific steps of the Hilbert transform include:
[0061] By applying analytic signal processing methods to the time-domain microseismic signal x(t), its complex form is obtained. Where H represents the Hilbert transform operation, g is the mathematical imaginary unit, and satisfies It is used to construct complex signals; and to extract the signal envelope curve from the complex signal z(t). The envelope curve is used to reflect the time-varying amplitude of the vibration energy of the pharyngeal tissues. A sliding window Fast Fourier Transform is used to extract the dominant frequency within each window, forming a dominant frequency variation curve f(t), which reflects the frequency distribution of the pharyngeal muscle group vibration over a continuous time period. The extracted envelope curve and dominant frequency curve serve as inputs for subsequent signal fitting and feature construction.
[0062] The airway pressure disturbance signals collected during swallowing are time-domain normalized to construct a standardized pressure change profile. The normalization process includes the following steps:
[0063] First, determine the time period from the start to the end of the swallowing action. Usually, the point from the point of maximum slope increase to the point of decrease in the signal is identified as the effective segment.
[0064] Let the signal segment be p(t). Calculate its maximum value pmax and minimum value pmin, and perform linear normalization: ; the normalized signal Equal-interval interpolation is performed at a fixed length (e.g., 100 sampling points) to ensure time alignment with the microseismic signal. The resulting standard pressure change profile can serve as a temporal representation of internal airflow disturbances during swallowing.
[0065] To achieve nonlinear temporal alignment between pharyngeal microseismic signals and airway pressure disturbance signals, a dynamic time warping algorithm is employed for signal matching. The implementation of the dynamic time warping algorithm includes: inputting two sets of time series data: the microseismic signal envelope curve A(t) and the normalized pressure profile pnorm(t), which may have different lengths; constructing a distance matrix D(i,j), where each element is the Euclidean distance between two time points, i.e.: In the formula, Let be the envelope value of the microseismic signal at the i-th time point. Let j represent the normalized pressure disturbance value at time point j. Using dynamic programming, search for the minimum-cost matching path in the distance matrix. k represents the number of matching point pairs in the path. The path cost is the minimum sum of distances between all points on the path;
[0066] The matching path W defines the optimal time alignment between the microseismic signal and the pressure disturbance signal, i.e., the pressure change time corresponding to each throat microseismic time point. This matching relationship provides the time synchronization basis for the subsequent construction of the muscle vibration coefficient matrix.
[0067] Based on the established signal matching path, at each matching time point, a muscle vibration coefficient matrix P is constructed to quantify the fluctuations in glottic muscle tension during the swallowing cycle. The specific process is as follows:
[0068] For each matching point (ik, jk), take the first derivative of the microseismic envelope value A(ik) with the corresponding point in the normalized pressure profile. This reflects the pressure disturbance gradient at that moment; the muscle vibration coefficient Pk at that point is defined as: Let P be a 1×N matrix of muscle kinetic coefficients, where N is the number of matching points, i.e., the number of samples within the swallowing cycle. The larger the muscle kinetic coefficient, the more intense the tension response of the pharyngeal muscle group at that moment, which is often related to excessive exertion or mucosal instability after surgery. This matrix P will serve as the key input feature vector for subsequently constructing a model for assessing glottic risk indicators and recovery status.
[0069] S300, perform asynchronous matching analysis between the speech signal and the muscle vibration coefficient matrix P to extract the non-physiological vocal load factor F.
[0070] During the postoperative speech recovery period, to identify abnormal vocal cord load in patients, short-time energy extraction and spectral entropy analysis of the speech signal are required. The specific processing procedure is as follows:
[0071] First, the original audio signal is divided into frames according to a time window, with a frame length of 25 milliseconds and a frame shift of 10 milliseconds.
[0072] For each frame of signal xn, calculate its short-time energy En, which is defined as the sum of the squares of the amplitudes of all sampling points in the current frame, and is used to reflect the sound intensity of the current frame.
[0073] Simultaneously, a Fast Fourier Transform is performed on each frame of the signal to obtain its spectral distribution S(f), the normalized probability P(f) of this distribution is calculated, and the spectral entropy is calculated accordingly. , used to measure the frequency complexity of vocal content;
[0074] By combining energy value and spectral entropy to set dual threshold judgment conditions, frames that meet the "high energy + low entropy" characteristics are identified as high-load vocal segments, representing vocal behaviors with heavy glottal burden.
[0075] To further analyze the glottic muscle tension corresponding to high-load vocal segments, it is necessary to map the previously constructed muscle vibration coefficient matrix P to the time axis of the speech signal and establish asynchronous feature correlation. The specific steps are as follows:
[0076] Since the time dimension of the muscle tremor coefficient matrix P is based on the swallowing cycle, it is inconsistent with the sampling axis of the speech signal, requiring nonlinear mapping. This mapping is achieved through a dynamic time warping algorithm, with the input being the tension change curve and the speech energy change curve in P.
[0077] The speech signal is segmented using a sliding window strategy, with each window having a length of 500 milliseconds and a sliding step of 250 milliseconds.
[0078] In each sliding window, the muscle vibration coefficient value of the corresponding time period is extracted according to the mapping path to form the muscle tension intensity sequence corresponding to the vocal segment, which serves as the basis for subsequent feature comparison.
[0079] To identify any significant mismatch between glottal muscle use and speech energy output in a vocal segment, it is necessary to analyze the correlation between muscle tone intensity sequences and speech energy curves. The implementation process is as follows:
[0080] For each vocal segment, two time series are constructed: one is the speech energy sequence E(t), and the other is the muscle tone intensity sequence P(t).
[0081] Calculate the Pearson correlation coefficient r for the two sequences. This coefficient is used to measure the degree of linear correlation between the two variables.
[0082] A mismatch identification threshold is set. When the correlation coefficient r < 0.3 and there is a continuous interval in the muscle tone intensity sequence where the value exceeds twice the average speech energy (i.e., abnormally high muscle tone and low vocal output), it is judged as a mismatch segment.
[0083] All mismatched segments will be included in the vocal load factor calculation in subsequent steps.
[0084] After identifying all vocal mismatch segments, the non-physiological vocal load factor F is calculated by combining their proportion of the total vocal time with the cumulative intensity of abnormal tension. The specific calculation method is as follows:
[0085] Let the total vocalization time be Ttotal, and the total duration of the mismatched segment be Tmismatch. Then the time scaling factor is: ;
[0086] The muscle tone intensity values within all mismatched segments are summed to obtain the cumulative value of abnormal tension. ,in This is the muscle tone value. Let N be the mean energy of the corresponding speech segment and N be the number of matching points. The non-physiological vocal load factor F is defined by linearly weighting the time scaling factor R and the cumulative tension abnormality value S: F = α⋅R + β⋅S; where the weighting coefficients α = 0.4 and β = 0.6, set based on clinical trial fitting results. A higher factor F indicates more significant non-physiological glottal overload behavior after surgery, suggesting the need for speech use intervention or medical follow-up.
[0087] S400, based on the synergistic analysis of environmental aerosol particle concentration data and non-physiological vocalization load factor F, identifies the air quality interference intensity R that may induce chronic edema or local infection, and constructs the non-physiological vocalization load factor F and air quality interference intensity R into a postoperative glottic risk coefficient vector V.
[0088] To identify the potential interference of air pollution on post-operative recovery after laryngeal surgery, an air pollution intensity index D needs to be extracted within a set monitoring period and adjusted according to individual sensitivity to construct an air quality interference intensity R. The air pollution intensity index D is calculated as follows: D equals the particle size-weighted average concentration multiplied by the exposure time multiplied by the sensitivity factor coefficient. The particle size-weighted average concentration is calculated based on the mass concentration of airborne particles in the range of 2.5 to 5 micrometers, with a weight of 1.5; the exposure time is measured in hours, based on the actual time the patient spends; the sensitivity factor coefficient is automatically adjusted based on the number of days after surgery and the doctor's score, with a value ranging from 0.8 to 1.5.
[0089] After obtaining D, to enhance its clinical applicability, the D value is normalized to the interval [0,1] and multiplied by a set of particle size-sensitive adjustment factor vectors S to enhance the response weight to high-sensitive particle size ranges, ultimately obtaining the air quality interference intensity R. This treatment enhances the ability to identify chronic mucosal irritants.
[0090] To unify the scale of indicators under different dimensions, the non-physiological vocalization load factor F and the air quality interference intensity R need to be subjected to bivariate orthogonal normalization to construct a set of joint influencing factors.
[0091] First, perform max-min normalization on F and R respectively to map them to the closed interval [0,1].
[0092] Considering the nonlinear coupling of postoperative risk in the joint space, an orthogonal transformation is further performed to construct a new, uncorrelated basis in the two-dimensional vector space. A normalized vector set [F′, R′] is defined as the set of joint influencing factors, used as input for subsequent nonlinear modeling. This normalization method eliminates potential linear coupling interference while maintaining the sensitivity of the original indicators, thus improving modeling stability.
[0093] To transform the nonlinear synergistic effect of the combined factors F' and R' into a quantifiable postoperative risk indicator, a risk sensitization function needs to be constructed based on clinical retrospective samples, and the glottic risk weight needs to be output through a surface response model.
[0094] The risk sensitization function is constructed using Gaussian radial basis functions. The expression is: the sensitization function value is equal to the deviation of the unit Gaussian function from F' and R', and the center point is the average value of F' and R' in the high-incidence area of glottic edema in the retrospective sample.
[0095] The surface response model is constructed using the support vector regression algorithm, with the probability of glottal abnormality in the backtracked samples as the target value and F' and R' as input variables to train the response surface.
[0096] The model output is defined as the glottic risk weight W, with a value range of 0 to 1. A higher value indicates that the patient is currently in a state of high risk of potential abnormal recovery.
[0097] After obtaining the glottal risk weight W, in order to construct a complete postoperative recovery assessment input vector, it needs to be combined with the normalized indices F' and R' to form the glottal risk coefficient vector V. The construction method is as follows: Define vector V as a three-dimensional real number vector: vector V is equal to the ordered arrangement of [F', R', W]; where F' and R' are the original indices after normalization, and W is the output result of the surface response model.
[0098] S500, the risk coefficient vector V is input into the transfer learning model M trained by the postoperative recovery curve to obtain the abnormal risk prediction value W of the target patient in the future time Δt.
[0099] To construct a supervised learning dataset for predicting postoperative recovery status, it is first necessary to build a time-labeled sample set of multi-stage postoperative recovery curves based on historical patient samples. This sample set takes physiological and environmental data from multiple time periods during the patient's recovery period as input and recovery outcome labels as output. Specifically, it includes the following steps:
[0100] The data recording period for each patient is from postoperative day 1 to day 30, with data recorded once a day to form a fixed-length time series;
[0101] Input features include speech intensity fluctuation curves, muscle tone recovery trajectories, and historical data on atmospheric environmental disturbances;
[0102] The output labels are manually labeled recovery status levels, divided into three categories: "normal recovery", "mild anomaly" and "significant delay", which are used as model supervision signals;
[0103] Each sample group is divided into days, ultimately forming an input feature sequence containing 30 time steps and corresponding recovery level labels, which are used for model training.
[0104] Postoperative speech intensity fluctuation characteristics of patients were extracted through continuous speech signal acquisition and time window analysis, as follows:
[0105] Wearable microphone sensors were used to collect raw speech signals from patients' daily natural speech, with a sampling frequency of no less than 16,000 Hz. The speech signals were divided into frame sequences with a frame length of 25 milliseconds and a frame shift of 10 milliseconds, and the short-time energy of each frame was calculated as a speech intensity index. A 1-hour sliding time window was used to extract the maximum, average, and standard deviation of speech intensity within each time window to represent the vocal state within that period. The above statistical values were spliced together in chronological order to form a speech intensity time-series curve, which was used to characterize postoperative vocal stability and recovery trend.
[0106] The muscle tone recovery trajectory is based on the previously constructed muscle vibration coefficient matrix P, and features are extracted and temporally modeled. The specific method is as follows:
[0107] Principal component analysis was used to extract the first three principal components of the daily generated muscle tremor coefficient matrix P, retaining no less than 90% of the original information. The first principal component was selected as the representative tension feature, and the change of this feature value over time was recorded to represent the postoperative tension recovery trend. The representative tension features extracted daily were spliced together according to the postoperative date to form the muscle tension recovery trajectory curve, which was used as one of the input features of the model.
[0108] Historical air disturbance data is obtained by correlating air particle concentration monitoring with location information to construct temporal characteristics of environmental factors. The specific steps are as follows:
[0109] The particle size monitoring range was set to 0.3 to 10 micrometers, and particulate matter concentration data was collected every 5 minutes using a laser scattering sensor. The location information of the patient's wearing device was collected simultaneously, and the air quality data was matched with the location information. The particulate concentration data of different locations and time periods each day were weighted and averaged to construct a daily aerosol disturbance index. The daily index was arranged in chronological order to form an environmental disturbance time series, which was used to analyze the impact of external pollution on the patient's recovery process.
[0110] To improve the model's generalization ability across different patients, a transfer learning strategy was used to construct a recovery trend model suitable for individualized prediction. The specific process is as follows:
[0111] The basic model adopts a two-layer long short-term memory network structure, with 128 hidden units in each layer and a time step input length of 30.
[0112] In the source data domain, i.e., historical patient data with relatively abundant labeled samples, preliminary training is completed, and the loss function is the cross-entropy function;
[0113] When migrating to the target patient data, the parameters of the previous layer are frozen, only the second layer is fine-tuned, and the input distribution is adjusted using a feature remapping mechanism;
[0114] The feature remapping method uses the maximum mean difference minimization algorithm to adjust the higher-order statistical properties of the source and target data to achieve transfer adaptation.
[0115] After model adaptation is complete, postoperative trend prediction is performed using the real-time input features of the current patient:
[0116] The risk coefficient vector V serves as the data input for the latest day in the time series, and together with the speech intensity, muscle tone trajectory, and environmental disturbance data from the historical 29 days, it forms a complete 30-day input sequence.
[0117] The input sequence is fed into a pre-trained transfer learning model, which then performs inference based on internal time dependencies and feature weights.
[0118] The final output is the predicted recovery trend within the future Δt time window, which can be set to a 3-day, 5-day, or 7-day window.
[0119] The model output is a continuous predicted probability value, defined as the abnormal risk prediction value W, which is used to quantify the likelihood of a patient developing glottic edema or delayed recovery within a time interval Δt.
[0120] When the abnormal risk prediction value W exceeds the adaptive warning threshold T, the S600 generates a personalized recovery path that includes the intervention level, intervention method, and voice environment suggestions, and pushes the path to the patient's smart terminal and the attending physician's platform.
[0121] To accurately determine whether a personalized intervention has been triggered, a warning threshold T needs to be dynamically set based on the patient's individual recovery pattern and compared with an abnormal risk prediction value W. This process includes the following technical steps:
[0122] A historical recovery database is constructed, containing the correspondence between predicted abnormal risk values W within 30 days post-surgery and actual recovery outcomes. Based on retrospective data from similar patient groups (matched by surgical type, age, post-operative days, etc.), the distribution of risk values leading to delayed recovery is extracted. A warning threshold T is set as the 75th percentile of this distribution, representing the boundary of the high-risk distribution, and can be dynamically adjusted according to changes in the patient's post-operative stage. When the current predicted value W is greater than or equal to T, it is determined to be an "intervention-triggered state," entering the personalized recovery path generation stage. This adaptive threshold strategy can balance individual differences and risk controllability, avoiding false alarms or missed alarms caused by static threshold settings.
[0123] After identifying the intervention trigger state, the corresponding intervention level needs to be determined based on the predicted risk level. This step is achieved through a mapping table:
[0124] The predicted value of abnormal risk W is divided into three level ranges:
[0125] Low risk: W≥T and W<T+0.1, corresponding intervention level is "self-adjustment recommended";
[0126] Medium risk: W≥T+0.1 and W<T+0.25, corresponding to "remote guidance and intervention";
[0127] High risk: W≥T+0.25, corresponding to "recommendation for outpatient follow-up";
[0128] The level of intervention determines the scope and delivery method of subsequent intervention plans. For example, self-adjustment suggestions are only sent to the patient's terminal, while high-risk suggestions are simultaneously pushed to the attending physician's platform.
[0129] This mapping table combines the continuity of model output values with the hierarchical logic of medical interventions, improving the accuracy of personalized responses.
[0130] After determining the intervention level, the system generates intervention suggestions and voice usage environment adjustment instructions based on the patient's recent recovery status data. The specific processing method is as follows:
[0131] The speech intensity trend is used to determine whether the patient has a persistent excessive vocalization behavior. If the average intensity has increased by more than 1.5 times over the past 3 days, a suggestion to "reduce vocalization frequency" is generated.
[0132] If the muscle tone recovery curve shows abnormal amplitude fluctuations, it suggests that there may be local tension instability. In this case, "short-term cold compress" and "avoiding high sound pressure output" are recommended.
[0133] If the particulate concentration in the environmental disturbance time series exceeds the set health standard limit (e.g., PM2.5 is higher than 75 micrograms per cubic meter), then add the suggestion to "upgrade the air filtration level of the residential environment";
[0134] All recommendations are generated by matching rules from a rule base built on a clinician knowledge base, which contains intervention templates for different risk combinations.
[0135] The final generated content constitutes a set of structured recovery path suggestions, which has highly individualized characteristics.
[0136] The generated recovery path information needs to be effectively communicated to patients and medical personnel. It should be pushed to multiple platforms using a structured data format to ensure data parsability and command consistency. The specific implementation is as follows:
[0137] The recovery path data is encapsulated in JavaScript object notation, with fields including: timestamp, intervention level, array of specific suggestion content, and push object identifier;
[0138] If the intervention level is "self-adjustment recommended", it will only be pushed to the patient's mobile terminal, with voice broadcast and message reminder functions;
[0139] If the intervention level is "remote guidance intervention" or "recommendation for outpatient follow-up", the recommendation will be pushed to the attending physician's workstation at the same time and uploaded to the medical data interface through an encrypted communication protocol;
[0140] All push notifications are uniquely identified, facilitating follow-up feedback tracking and intervention effectiveness evaluation. This coordinated push strategy ensures that patients and doctors maintain information synchronization throughout the recovery process, contributing to the establishment of a closed-loop rehabilitation management mechanism.
[0141] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.
Claims
1. A method for intelligent analysis and management of post-pharyngeal surgery recovery information, characterized in that: include: S100 collects multi-source sensory data of the target patient during the postoperative recovery period, including voice signals, pharyngeal surface micro-vibration signals, airway pressure disturbance signals during swallowing, and environmental aerosol particle concentration data. S200, based on cross-fitting of pharyngeal surface microvibration signals and airway pressure disturbance signals during swallowing, constructs a muscle vibration coefficient matrix P reflecting the tension fluctuations of the glottic muscle group, including: S201, perform Hilbert transform on the denoised pharyngeal surface micro-vibration signal to extract the envelope feature curve and the trend of the main frequency variation of the signal; S202, the airway pressure disturbance signal collected during swallowing is normalized in the time domain to obtain the standard pressure change profile. S203 uses a dynamic time warping algorithm to perform nonlinear matching between the microseismic signal envelope curve and the airway pressure profile to establish an action alignment relationship. S204. Based on the matching time points, construct the vibration coefficient matrix P. Each element in the matrix is obtained by calculating the micro-vibration amplitude and pressure disturbance gradient at the corresponding time. Specifically, for each matching point (ik, jk), take the first derivative of the micro-vibration envelope value A(ik) with the corresponding point in the normalized pressure profile. The muscle oscillation coefficient Pk is defined as: ;in, Let be the envelope value of the microseismic signal at the i-th time point; S300, perform asynchronous matching analysis on the speech signal and the muscle vibration coefficient matrix P to extract the non-physiological vocal load factor F, including: S301 performs short-time energy extraction and spectral entropy analysis on postoperative speech signals to identify high-load vocal segments within the speech cycle. S302, map the muscle tremor coefficient matrix P to the speech signal time axis, use a sliding window method to perform asynchronous feature association, and extract the muscle tension intensity sequence corresponding to the vocal segment; S303, calculate the correlation index between the muscle tone intensity sequence and the speech energy curve in each vocal segment, and identify the mismatched segments where the muscle tone is abnormally higher than the speech output power; S304, which combines the proportion of mismatched segments with the cumulative amplitude value, is defined as the non-physiological vocal load factor F, specifically including: The muscle tone intensity values within all mismatched segments are summed to obtain the cumulative value of abnormal tension. ,in This is the muscle tone value. Let N be the average speech energy corresponding to this segment, and N be the number of matching points. The time scaling factor R and the cumulative tension abnormality value S are linearly weighted to define the non-physiological vocal load factor F. The calculation expression of the time scaling factor is: R = Tmismatch / Ttotal; Ttotal is the total vocal time, and Tmismatch is the total duration of the mismatch segment. S400, based on the synergistic analysis of environmental aerosol particle concentration data and non-physiological vocalization load factor F, identifies the air quality interference intensity R that may induce chronic edema or local infection, and constructs a postoperative glottic risk coefficient vector V by combining the non-physiological vocalization load factor F and the air quality interference intensity R, including: S401, extract the air disturbance intensity index D within the target monitoring period, and adjust the calculation according to the particle size sensitivity factor and exposure time to obtain the air quality disturbance intensity R. Specifically, the calculation method of the air disturbance intensity index D is: D is equal to the particle size weighted average concentration multiplied by the exposure time multiplied by the sensitivity factor coefficient. After obtaining D, the D value is normalized to the interval [0,1], multiplied by a set of particle size sensitivity adjustment factor vectors S, and finally the air quality disturbance intensity R is obtained. S402, the non-physiological vocalization load factor F and the air quality interference intensity R are bivariate orthogonally normalized to construct a set of joint influencing factors; S403, based on the risk sensitization function fitted by the clinical retrospective sample, input F and R into the surface response model to obtain the nonlinear coupling effect value as the glottal risk weight; S404, combine the glottic risk weight with the normalized F and R to construct the postoperative glottic risk coefficient vector V; S500, the risk coefficient vector V is input into the transfer learning model M trained by the postoperative recovery curve to obtain the abnormal risk prediction value W of the target patient in the future Δt time. When the abnormal risk prediction value W exceeds the adaptive warning threshold T, the S600 generates a personalized recovery path that includes the intervention level, intervention method, and voice environment suggestions, and pushes the path to the patient's smart terminal and the attending physician's platform.
2. The intelligent analysis and management method for post-pharyngeal surgery recovery information according to claim 1, characterized in that: The S500 includes: S501, construct a time-series labeled sample set containing multi-stage postoperative recovery curves, using patient voice intensity fluctuations, muscle tone recovery trajectories, and historical environmental disturbance data as input features; S502 uses a long short-term memory network as the basic model, completes the initial training in the source data domain, and achieves the transfer to the target patient data domain through parameter freezing and feature remapping mechanisms. S503 uses the current patient's risk coefficient vector V as the latest input to the time series window, and combines it with historical feature trajectories to complete the recovery trend prediction. S504, the model outputs the abnormal risk prediction value W within the time window Δt, which represents the probability that the patient will develop glottic edema or delayed recovery within the prediction period.
3. The intelligent analysis and management method for post-pharyngeal surgery recovery information according to claim 2, characterized in that: The methods for obtaining the patient's voice intensity fluctuations, muscle tone recovery trajectory, and historical environmental disturbance data include: S511 acquires continuous speech signals and extracts the speech intensity value for each cycle based on a sliding time window to construct a speech intensity time-series curve; S512, principal component analysis was performed on the daily collected muscle vibration coefficient matrix P to extract representative tension features and construct muscle tension recovery trajectory curves in chronological order; S513 synchronously records the changes in aerosol concentration in the air environment where the patient is located at different time periods, and constructs an environmental disturbance time series by combining location tags.
4. The intelligent analysis and management method for post-pharyngeal surgery recovery information according to claim 1, characterized in that: The S600 includes: S601, dynamically sets the warning threshold T based on the historical recovery database, and compares it with the current predicted value W to identify the intervention trigger state; S602 determines the corresponding intervention level based on the patient risk level mapping table, and divides it into three categories: "suggest self-adjustment", "remote guidance intervention" and "suggest outpatient follow-up". S603, combining patient speech intensity trends, muscle tone recovery curves and environmental exposure records, generates intervention suggestions and speech use environment adjustment instructions; S604 pushes the generated recovery path in structured data format to both the patient's mobile terminal and the attending physician's platform simultaneously, enabling integrated rehabilitation management.
Citation Information
Patent Citations
Intelligent oral snore-ceasing equipment system based on multi-parameter monitoring and method thereof
CN120392405A
Auxiliary scheme generation method and system for dysphagia rehabilitation
CN120600228A