Dysphagia screening method based on non-negative orthogonal tensor decomposition
Through the non-negative orthogonal tensor decomposition method, the efficient non-invasiveness and high-precision problems of traditional swallowing disorder screening are solved, and stable screening is achieved in complex noise environments, which improves the sensitivity to weak pathological signals and meets the needs of real-time and large-scale processing.
Patent Information
- Application Number
- CN202510657949.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-08-08
AI Technical Summary
Traditional dysphagia screening methods rely on subjective assessment or high-cost equipment, making it difficult to efficiently and non-invasively screen dysphagia in the elderly population, and there are difficulties in modeling and redundancy in the extraction of speech signal characteristics.
The method based on non-negative orthogonal tensor decomposition is adopted to collect speech signals, construct speech spectrum tensors, and use non-negative orthogonal Tucker decomposition model to decompose energy spectrum tensors, extract feature sets, and combine them with neural networks or classifiers for screening.
It realizes high-precision, noise-resistant and robust screening of swallowing disorders, high clustering accuracy, can maintain stability in different noise environments, flexibly switch clustering modes, improve sensitivity to weak pathological signals, and meet real-time and large-scale processing needs.
Smart Images

Figure CN120452480A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of speech signal processing and dysphagia screening, and specifically relates to a dysphagia screening method based on non-negative orthogonal tensor decomposition. Background Art
[0002] Dysphagia is a significant clinical problem caused by a variety of neurological or structural diseases. Traditional screening methods rely on subjective assessments or high-cost imaging equipment (e.g., VFSS, FEES). Dysphagia is highly prevalent in the elderly over 65 years old, particularly those with neurodegenerative diseases. Dysphagia prevents the elderly from eating normally, potentially leading to severe malnutrition and, in severe cases, even life-threatening conditions. It can also easily lead to aspiration pneumonia. Therefore, early detection and intervention of dysphagia are crucial. Exploring an efficient, non-invasive, and easily accepted early screening method for dysphagia in the elderly is crucial. As a non-invasive biosignal, speech signals have broad application prospects in assisting the diagnosis of dysphagia. However, due to the time-frequency dynamics and high-dimensional structure of speech signals, feature extraction suffers from modeling difficulties and high redundancy. Summary of the Invention
[0003] In response to the above problems, the present invention proposes a dysphagia screening method based on non-negative orthogonal tensor decomposition.
[0004] The technical solution of the present invention is:
[0005] A swallowing disorder screening method based on non-negative orthogonal tensor decomposition, characterized by comprising the following steps:
[0006] S1. Collecting speech signals, specifically, having the subject speak in different tones according to the set corpus, thereby obtaining different speech signals;
[0007] S2. Preprocess the acquired voice signals and standardize different voice signals;
[0008] S3. Calculate the speech energy spectrum, including:
[0009] S31, voice framing, specifically, dividing the voice signal into frames starting from the first voice data point according to the set frame length, with an inter-frame overlap rate of 50%;
[0010] S32. Calculate the energy spectrum of each frame of the speech signal by short-time Fourier transform for the framed speech signal, retain the energy points of the first half of each frame of the speech signal and take their absolute values, and concatenate the energy points of each frame of the speech signal in the order of the framing to obtain a non-negative energy spectrum in the frequency dimension × time dimension.
[0011] S4. Using the methods of S1-S3, obtain the non-negative energy spectra of the subject speaking the same corpus in three different tones. Stack the three obtained non-negative energy spectra into a three-dimensional tensor, namely the speech spectrum tensor, whose three dimensions are frequency, time, and emotion respectively. Decompose the energy spectrum tensor using the non-negative orthogonal Tucker decomposition model. The decomposition model is as follows:
[0012]
[0013] in, Represents the speech spectrum tensor, C is the core tensor, d N is the speech spectrum tensor The dimension in the nth dimension, A (n) is the subspace matrix of the nth dimension, Q is the centralization constraint matrix, defined as The subspace matrix A is introduced (n) The orthogonality constraint, The length is r n A vector of all 1s, r n is the speech spectrum tensor The rank of the subspace in the nth dimension, The length is d n A vector of all 1s, d n Represents the length of the speech spectrum tensor in the nth dimension, and α is used to balance the error term and constraints The regularization parameter of , I is the identity matrix;
[0014] according to:
[0015]
[0016] in, C (n) The n-th matrix of the core tensor C, v and c are and vectorization of C; using the non-negative matrix solution method, combined with the accelerated proximal gradient algorithm, while implementing the orthogonality and non-negativity constraints on each factor matrix, after the model converges, the compressed representation feature set C of each speech spectrum tensor is obtained;
[0017] S5. Based on the obtained feature set C, a neural network model or a classifier is used to implement dysphagia screening.
[0018] The beneficial effects of the present invention are: (1) it realizes integrated high-precision screening on synthetic dysphagia speech tensor data. On the synthetic dysphagia speech tensor data with a dimension of 150×150×150, a block cluster size of 10×10×10, and a noise standard deviation of 3.5, the method of the present invention achieves a clustering accuracy of 100% and a fitting rate of 93.0% in a hard clustering scenario, and a fitting rate of 94.3% in a soft clustering scenario with an average signal-to-noise distortion ratio of 18.11dB; (2) it achieves excellent noise robustness. When the noise level increases from σ=0.5 to 8.0 , the clustering accuracy of the present invention under two different dimensions and clustering scales always remains above 85%, while the clustering accuracy of the existing method may drop to below 40%, fully ensuring the screening stability and reliability of swallowing signals in complex noise environments; (3) flexible clustering with adjustable cross-membership is realized. By adjusting the regularization parameter α, the present method can flexibly switch between hard clustering and soft clustering, which can not only achieve clear classification, but also retain the cross-membership information of weak features or overlapping features in swallowing disorder speech, thereby improving the sensitivity to weak pathological signals; (4) in typical scenarios, the rank r of each pattern factor is n Much smaller than d n , the storage and computing amount can be reduced by nearly 90%, combined with the parallel implementation of the accelerated proximal gradient algorithm, to meet the real-time and large-scale processing requirements of dysphagia screening. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 It is a schematic diagram of the overall process of the present invention.
[0020] Figure 2 Schematic diagram for constructing the speech energy spectrum tensor.
[0021] Figure 3 This is a diagram of Tucker's decomposition structure. DETAILED DESCRIPTION
[0022] The present invention will be described in detail below with reference to the accompanying drawings.
[0023] The method of the present invention is based on speech signals and the characteristics of speech energy spectrum, and has three main points: 1) the speech energy spectrum is non-negative, that is, all energy values are real numbers greater than 0; 2) the energy in the speech energy spectrum is mainly concentrated in the low-frequency area, and the energy spectrum as a whole presents a low-rank characteristic; 3) since the speech signal is dynamic, each energy spectrum contains two dimensions of time and frequency, and each dimension reflects different information. In order to better perform early screening of dysphagia, it is necessary to explore the characteristic subspace of each dimension. Based on the above characteristics of the energy spectrum, the present invention innovatively applies the non-negative orthogonal Tucker decomposition model to the entire method system for dysphagia screening. The main features of applying this method are: the non-negative constraint takes into account the non-negative characteristics of the speech signal energy spectrum, and can better capture the local feature differences between the energy spectra of speech signals of patients with dysphagia and those of normal people; the orthogonal constraint can independently explore the characteristic subspace of each dimension, which helps to extract the key information of each dimension of the energy spectrum; the combination of non-negative and orthogonal constraints also naturally introduces sparse modeling of the data, which helps the model extract key principal components. The decomposition-based model can also better fit the low-rank characteristics of the speech energy spectrum. Therefore, the non-negative orthogonal Tucker decomposition model can well meet the analysis needs of the speech energy spectrum.
[0024] like Figure 1 As shown in FIG. 1 , the overall process of the present invention mainly includes:
[0025] 1. Speech signal acquisition: To better reflect the vocalization of patients with dysphagia, this paper designs a new speech signal acquisition paradigm. That is, it selects corpus suitable for expressing various emotions to collect signals. For example, "The weather is so nice today." The subjects are asked to express it in a calm, depressed, or happy tone, thereby activating the vocal organs to different degrees, better reflecting the differences in acoustic characteristics between patients and normal people.
[0026] 2. Preprocessing: De-noise, amplitude normalize, and resample the speech signal according to actual needs. The purpose of resampling is to sample all speech signals to the same length. This length can be customized or the mean of all speech signal lengths can be calculated as a standard, so that subsequent speech signal processing can be performed in a unified framework.
[0027] 3. Calculate the Speech Energy Spectrum: First, perform speech framing. Framing means dividing the speech signal into segments of equal length, starting from the first speech data point, with overlapping frames. Considering the non-stationary nature of speech signals, the frame length is typically set to 10 to 25 ms, with a 50% overlap ratio. The energy spectrum is then calculated using a short-time Fourier transform (SFT). The energy spectrum is calculated for each frame of speech signal using the SFT, with the length of the energy spectrum set to the length of the frame. Because the energy spectrum obtained by the Fourier transform is symmetrical, only the energy points in the first half of each frame are retained and their absolute values are taken. The energy points of each frame are concatenated in the order of framing to obtain a non-negative energy spectrum in the frequency × time dimension.
[0028] 4. Non-negative orthogonal Tucker decomposition model decomposes the energy spectrum tensor: According to the above data collection paradigm, each participant reads the same speech signal with three different emotions. Since the content of the speech signals is consistent, but the emotional color is different, there is information redundancy between different speech signals. To solve this problem, the patent of this invention stacks the energy spectra of multiple speech signals of each participant into a three-dimensional tensor. A tensor is a high-order generalization form of a matrix. The three dimensions of the tensor are frequency, time, and emotion. The schematic diagram of the construction of the three-dimensional tensor is as follows: Figure 2 The main symbols are shown in Table 1.
[0029] Table 1 Main symbols
[0030]
[0031] Tucker decomposition is a classic tensor decomposition method. As shown in the figure below, Tucker decomposition can compress a three-dimensional tensor into the product of a set of low-rank factor matrices and a core tensor, where each low-rank factor matrix represents a subspace of that dimension, and the elements in the core tensor are the interaction weights of the principal components of each subspace, which can be used as a compressed representation of the original tensor. Tucker decomposition is suitable for extracting potential structures from high-dimensional, multimodal data. The structure of Tucker decomposition is as follows: Figure 3 shown.
[0032] However, the factor matrix in the standard Tucker decomposition is not constrained, the factors are uninterpretable (negative values, non-orthogonal), there is feature redundancy, and subspace overlap affects classification or clustering performance. The present invention combines the actual task requirements and applies the non-negative orthogonal Tucker decomposition model to perform high-dimensional feature extraction. The optimization model of this decomposition model is as follows:
[0033]
[0034] in, Represents the speech spectrum tensor, C is the core tensor, A (n) is the subspace matrix of the nth dimension. Each factor matrix A is introduced 9n) Orthogonality constraints. According to the equivalence between the following expressions:
[0035]
[0036] in C (n) The n-th matrix of the core tensor C, v and c are and C vectorization, while
[0037]
[0038] By utilizing a non-negative matrix solution method, combined with an accelerated proximal gradient algorithm, and simultaneously imposing orthogonal and non-negative constraints on each factor matrix, the model converges to obtain a compressed feature set C representing each high-dimensional speech energy spectrum tensor. Based on C, the backend can be trained using a deep learning model to achieve the final classification prediction. Alternatively, each C can be vectorized into a feature vector, and traditional classifiers such as Support Vector Machines (SVMs) can be used to screen for dysphagia.
Claims
1. A swallowing disorder screening method based on non-negative orthogonal tensor decomposition, characterized in that: The following steps are involved: S1. Collecting speech signals, specifically, having the subject speak in different tones according to the set corpus, thereby obtaining different speech signals; S2. Preprocess the acquired voice signals and standardize different voice signals; S3. Calculate the speech energy spectrum, including: S31, voice framing, specifically, dividing the voice signal into frames starting from the first voice data point according to the set frame length, with an inter-frame overlap rate of 50%; S32. Calculate the energy spectrum of each frame of the speech signal by short-time Fourier transform for the framed speech signal, retain the energy points of the first half of each frame of the speech signal and take their absolute values, and concatenate the energy points of each frame of the speech signal in the order of the framing to obtain a non-negative energy spectrum in the frequency dimension × time dimension. S4. Using the methods of S1-S3, obtain the non-negative energy spectra of the subject speaking the same corpus in three different tones. Stack the three obtained non-negative energy spectra into a three-dimensional tensor, namely the speech spectrum tensor, whose three dimensions are frequency, time, and emotion respectively. Decompose the energy spectrum tensor using the non-negative orthogonal Tucker decomposition model. The decomposition model is as follows: in, Represents the speech spectrum tensor, C is the core tensor, d N is the speech spectrum tensor The dimension in the nth dimension, A (n) is the subspace matrix of the nth dimension, Q is the centralization constraint matrix, defined as The subspace matrix A is introduced (n) The orthogonality constraint, The length is r n A vector of all 1s, The length is d n A vector of all 1s, α is used to balance the error term and constraints The regularization parameter of , I is the identity matrix; according to: in, C (n) The n-th matrix of the core tensor C, v and c are and vectorization of C; using the non-negative matrix solution method, combined with the accelerated proximal gradient algorithm, while implementing the orthogonality and non-negativity constraints on each factor matrix, after the model converges, the compressed representation feature set C of each speech spectrum tensor is obtained; S5. Based on the obtained feature set C, a neural network model or a classifier is used to implement dysphagia screening.
Citation Information
Cited By
Multi-modal sensing fusion swallowing rehabilitation evaluation system and method
CN120899193A