Alzheimer disease recognition method and system based on voice features
By extracting the sequence of structural parameters at the speech frame level, locating sparse segments with amplitude fluctuations and continuous jump breaks, and combining them with a set of abnormal time trend structural indicators, the problem of low accuracy in Alzheimer's disease speech recognition in existing technologies has been solved, and high-precision recognition of speech features of Alzheimer's disease has been achieved.
Patent Information
- Application Number
- CN202511187782.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-25
- Publication Date
- 2025-10-31
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing speech recognition technologies for Alzheimer's disease struggle to effectively capture subtle abnormalities in local speech segments, especially brief or non-continuous aberrations, leading to a decrease in recognition accuracy.
By extracting the sequence of speech frame-level structural parameters, locating sparse amplitude fluctuations and continuous jump breaks, and combining them with an abnormal time trend structure index set, a trend structure line is established for directional judgment, thereby achieving cross-scale feature coupling recognition of speech features in Alzheimer's disease.
It improves the discriminative power and accuracy of speech features in Alzheimer's disease, enhances the ability to detect weak abnormal information in speech, and significantly improves the stability and discrimination efficiency of recognition results.
Smart Images

Figure CN120877787A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of speech analysis technology, and in particular to a method and system for Alzheimer's disease identification based on speech features. Background Technology
[0002] The field of speech analysis technology involves technologies related to the acquisition, decomposition, feature extraction, and analysis of human speech signals. It mainly includes core aspects such as speech recognition, speech synthesis, speaker recognition, and emotion recognition. Typically, it involves modeling and matching the acoustic and semantic features of speech signals, quantifying and processing speech data to achieve the recognition and understanding of speech content or speaker states. Among these, Alzheimer's disease identification methods refer to a technical approach that extracts quantifiable indicators from the patient's speech, such as speech rate, pitch change rate, pause duration, jitter rate, and fundamental frequency fluctuations, and combines these with classification models or discrimination criteria for analysis and judgment. This typically involves constructing a speech feature dataset and using classifiers such as support vector machines or random forests for training and discrimination, thereby completing the identification of speech features related to Alzheimer's disease symptoms.
[0003] Current speech recognition methods for Alzheimer's disease focus on macroscopic speech parameters such as speech rate changes, fundamental frequency fluctuations, and pause times. They rely on overall speech features for modeling and training. However, in actual recognition, these methods are easily affected by changes in the speech context and individual speech differences. This leads to the model's inability to effectively capture subtle feature anomalies in local speech segments, especially when there are brief jumps or non-continuous abnormalities in the speech. The lack of sufficient parameter dimensions and dynamic structure references can easily cause the discrimination model to under-identify symptom features. For example, sudden breaks or sparse amplitude segments that frequently appear in the early speech manifestations of patients are often ignored because they are not continuous, resulting in a decrease in recognition accuracy. Summary of the Invention
[0004] The purpose of this invention is to address the shortcomings of existing technologies by proposing a method for Alzheimer's disease identification based on speech features.
[0005] To achieve the above objectives, the present invention adopts the following technical solution: an Alzheimer's disease identification method based on speech features, comprising the following steps:
[0006] S1: Obtain the continuous sampling point sequence of the patient's speech signal, perform frame division processing on the signal, extract the mean absolute value of amplitude, the difference between the maximum and minimum amplitude and the energy value in each frame, combine the parameters in sequence and arrange them in sequence according to the frame number to generate a speech frame-level structural parameter sequence.
[0007] S2: Based on the ratio of the maximum and minimum amplitude difference to the mean absolute value of amplitude in each frame of the speech frame-level structural parameter sequence, locate the frame segments in continuous frames whose ratio is less than the predefined dynamic range ratio threshold, and merge adjacent segments in chronological order to generate a list of sparse amplitude fluctuation segments.
[0008] S3: Based on the mean absolute value of amplitude in the speech frame-level structural parameter sequence, perform inter-frame difference operation on all frames and record the positions where there are two sets of positive and negative opposite relationships in the continuous difference, mark the start and end positions of abrupt breakage, and generate a set of continuous abrupt breakage segments.
[0009] S4: Based on the list of sparse amplitude fluctuation segments and the set of continuous jump fracture segments, extract the median time point of each overlapping segment and establish a trend structure line in chronological order. Perform directional judgment and record the number of points on the structure line to generate an abnormal time trend structure index set.
[0010] As a further aspect of the present invention, the speech frame-level structural parameter sequence includes a frame number index, a frame-level amplitude absolute value mean sequence, a frame-level maximum and minimum amplitude difference sequence, and a frame-level energy value sequence; the amplitude fluctuation sparse segment list includes the sparse segment start and end time, the sparse segment frame index, and the sparse segment ratio sequence; the continuous jump break segment set includes the break segment start and end position, the continuous falling frame segment number, and the amplitude mean difference sequence; and the abnormal time trend structure index set includes the trend structure point time median, trend direction change identifier, and the number of structure points.
[0011] As a further aspect of the present invention, the step of obtaining the speech frame-level structural parameter sequence specifically includes:
[0012] S111: Obtain the continuous sampling point sequence of the patient's speech signal, divide it into frames at equal intervals along the time axis, set the frame length and frame shift parameters, divide the entire speech signal into multiple adjacent speech frames, extract the absolute value sequence of amplitude in each frame, and generate the frame mean amplitude sequence.
[0013] S112: Based on the frame mean amplitude sequence, compare the maximum amplitude value and the minimum amplitude value in each frame, extract the amplitude range feature of each frame, perform inductive processing based on the original absolute amplitude value sequence of each frame, pair and associate it with the frame number, and obtain the frame-level structure index sequence.
[0014] S113: Based on the frame-level structure index sequence, arrange the frame numbers in ascending order, and combine the mean absolute value of amplitude, the amplitude range value and the relevant sequence content of each frame into a single frame structure parameter group. Summarize all frame-level structure parameter groups to obtain the speech frame-level structure parameter sequence.
[0015] As a further aspect of the present invention, the step of obtaining the list of sparse amplitude fluctuation segments specifically includes:
[0016] S211: Based on the speech frame-level structural parameter sequence, extract the maximum and minimum amplitude difference and the mean absolute amplitude value of each frame in sequence, combine them using the proportional relationship between the amplitude range and the mean, calculate the fluctuation ratio for the ratio corresponding to the consecutive frame number, and establish a frame index sequence structure in combination with the frame number order to obtain the consecutive frame fluctuation ratio index sequence.
[0017] S212: Based on the continuous frame fluctuation ratio index sequence, perform inter-frame screening operation according to the preset dynamic range ratio threshold, determine whether the intra-frame ratio is in the sparse segment determination interval, continuously record the frame number index that meets the condition, and obtain the continuous sparse frame sequence index.
[0018] S213: Based on the continuous sparse frame sequence index, perform screening and merging processing according to the continuity of frame number. If the frame gap is less than two frames, merge them into the same segment; otherwise, divide them into independent segments. Iterate through all sparse frame sequences and complete the merging processing. Record each merged segment as the segment boundary according to the frame number of the start and end frames, mark the time position, and obtain the list of sparse segments with amplitude fluctuation.
[0019] As a further aspect of the present invention, the step of obtaining the set of continuously abruptly fractured segments specifically includes:
[0020] S311: Based on the mean absolute value of amplitude in the speech frame-level structural parameter sequence, extract the mean items of two adjacent frames in the order of frame number and perform frame difference calculation to obtain the mean difference of consecutive frames. If the result is positive, it indicates that the amplitude is increasing, and negative value indicates that it is decreasing. Record the difference sign of each frame to form a difference change sign sequence. Find two consecutive sign reversal positions in the sequence, that is, consecutive reversal groups where the sign changes from positive to negative or from negative to positive. Establish a continuous change position index through the reversal frame number and calculate to obtain the continuous jump trend index value. If there are two difference reversals before the corresponding position, mark it as a jump change candidate segment.
[0021] S312: Based on the continuous jump trend index value, read the sequence and screen it in the order of frame number. For each item, determine whether the subsequent three frames constitute a continuous downward structure, determine whether the drop amplitude exceeds the set ratio value, set the continuous drop judgment threshold, and define the abnormal speech segment definition standard in Alzheimer's disease speech samples. Record the effective mutation area and obtain the set of continuous mean drop frame segments.
[0022] S313: Based on the set of continuously decreasing average frame segments, merge all candidate mutation segments continuously according to frame number. If the interval between two mutation segments is less than or equal to two frames, they are considered as the same segment. Construct a unified number for the broken segments. After the segments are merged, convert the frame number of each segment into a time interval. Record the start and end times of all segments in a unified manner to generate a set of continuously jumping broken segments.
[0023] As a further aspect of the present invention, the steps for obtaining the abnormal time trend structure indicator set are specifically as follows:
[0024] S411: Based on the list of sparse amplitude fluctuation segments and the set of continuous jump break segments, extract the start and end points of time for each frame segment, and construct all frame segments in the two sets as independent time intervals. Perform time axis matching operation segment by segment, determine whether there are overlapping frame segments between any two segments, check the proportion of the overlapping part in the two original segments, if the proportion is more than half of the frame length in both segments, it is determined to be a valid overlapping segment, and mark all segments that meet the conditions according to the corresponding relationship to obtain a set of bidirectional overlapping frame segments;
[0025] S412: Based on the bidirectional overlapping frame segment set, extract the start frame number and end frame number of all combined segments that meet the conditions, calculate the middle frame of the time interval as the representative point, convert the representative point into the corresponding time value, construct the median time point sequence, and arrange them in chronological order to obtain the complete time trend point trajectory and establish the median time point sequence.
[0026] S413: Based on the median time point sequence, determine the sequential relationship between adjacent time points in order, mark the case where the later time point is earlier than the previous time point as reverse, and mark the case where it is later as positive, mark the continuity of the overall trend direction and the number of direction changes, record the number and distribution characteristics of positive and reverse points, and combine the median time point density information to obtain the abnormal time trend structure index set.
[0027] As a further aspect of the present invention, the method further includes:
[0028] Based on the trend structure points and corresponding frame segment numbers marked in the abnormal time trend structure index set, the continuous jump trend speech frame segments are summarized, uniformly numbered and categorized with labels, and a frame order and time stamp index is established to generate Alzheimer's disease speech feature recognition results.
[0029] The Alzheimer's disease speech feature recognition results include an abnormal trend frame segment number index, trend label classification results, and time series markers.
[0030] As a further aspect of the present invention, the steps for obtaining the speech feature recognition results for Alzheimer's disease are specifically as follows:
[0031] S511: Based on the trend structure points marked in the abnormal time trend structure index set and the corresponding frame segment numbers, extract the start and end numbers of the frame segments covered by each trend segment, convert all frame numbers into the corresponding timestamp sequence in sequence, arrange the trend segments in chronological order, compare the number intervals of two adjacent trend segments, if the number interval value is within the allowable range, it is determined to be the same segment with continuous trend, and merge it into the same jump trend voice frame segment to establish a set of continuous jump trend frame segments;
[0032] S512: Based on the set of continuous jump trend frames, assign an independent number to each frame segment, and set a label based on the directional characteristics of the trend change in the frame segment. Summarize the number, frame start and end number, time start and end value, and trend label content to form an index list and establish a jump trend frame segment label index table.
[0033] S513: Based on the trend change direction, number of frames, time span, label type, and intra-frame trend switching attributes of each frame segment in the trend segment label index table, and according to the set frame segment distribution judgment conditions, the trend stability and trend switching characteristics of each frame segment are filtered and stratified, and the data of different feature labels are integrated to obtain the Alzheimer's disease speech feature recognition result.
[0034] An Alzheimer's disease identification system based on speech features includes:
[0035] The speech feature extraction module is used to perform S1: acquire a continuous sampling point sequence of the patient's speech signal, perform frame division processing on the signal, extract the mean absolute value of amplitude, the difference between the maximum and minimum amplitude and the energy value in each frame, combine the parameters in sequence and arrange them in sequence according to the frame number to generate a speech frame-level structural parameter sequence;
[0036] The low-amplitude region identification module is used to perform S2: based on the ratio of the maximum and minimum amplitude difference to the mean absolute amplitude value of each frame in the speech frame-level structural parameter sequence, locate the frame segments in continuous frames whose ratio is less than the predefined dynamic range ratio threshold, and merge adjacent segments in chronological order to generate a list of sparse amplitude fluctuation segments.
[0037] The module for marking broken segments is used to execute S3: based on the mean absolute value of amplitude in the speech frame-level structural parameter sequence, perform inter-frame difference operation on all frames and record the positions where there are two sets of positive and negative opposite relationships in the continuous difference, mark the start and end positions of abrupt breakage, and generate a set of continuous abrupt breakage segments;
[0038] The integrated trend anomaly module is used to execute S4: Based on the list of sparse amplitude fluctuation segments and the set of continuous jump break segments, the median time point of each overlapping segment is extracted and a trend structure line is established in chronological order. The direction of the structure line is judged and the number of points is recorded to generate an abnormal time trend structure index set.
[0039] The collection and recognition result module is used to execute S5: based on the trend structure points marked in the abnormal time trend structure index set and the corresponding frame segment number, it summarizes the continuous jump trend speech frame segments, performs unified numbering and label classification, establishes frame order and time stamp index, and generates Alzheimer's disease speech feature recognition results.
[0040] Compared with the prior art, the advantages and positive effects of the present invention are as follows:
[0041] In this invention, by serializing and integrating multiple parameters such as the mean absolute value of amplitude and the maximum and minimum difference, and combining multi-segment continuous ratio comparison and abrupt trend localization, a dual analysis mechanism for the sparsity and jumpiness of speech amplitude fluctuations is formed. By calculating the degree of overlap of trend indicators in adjacent frame segments and extracting the position in the trend structure accordingly, a trend structure line of the time series is established and its directional change and position density are extracted. This completes the quantitative classification of abnormal trends in the frame sequence, realizing cross-scale feature coupling recognition from speech micro-fluctuations to time series trends. This improves the discrimination and accuracy of speech features in Alzheimer's disease, enhances the detection capability of weak abnormal information in speech, and significantly improves the stability and discrimination efficiency of recognition results. Attached Figure Description
[0042] Figure 1 This is a flowchart of the main steps of the present invention;
[0043] Figure 2 This is a flowchart of the speech frame-level structural parameter sequence acquisition process of the present invention;
[0044] Figure 3 This is a flowchart of the process for obtaining the list of sparse amplitude fluctuations in this invention.
[0045] Figure 4 This is a flowchart of the process for obtaining the set of continuously jumping fracture sections in this invention;
[0046] Figure 5 This is a flowchart of the process for obtaining the abnormal time trend structure indicator set of the present invention;
[0047] Figure 6 This is a flowchart of the process for obtaining speech feature recognition results for Alzheimer's disease according to the present invention. Detailed Implementation
[0048] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0049] Please see Figure 1 A method for identifying Alzheimer's disease based on speech features includes the following steps:
[0050] S1: Obtain the continuous sampling point sequence of the patient's speech signal, perform frame division processing on the signal, extract the mean absolute value of amplitude, the difference between the maximum and minimum amplitude and the energy value in each frame, combine the three types of parameters in sequence into a set of frame-level feature records, and arrange them in sequence according to the frame number to generate a speech frame-level structure parameter sequence.
[0051] S2: Based on the ratio of the maximum and minimum amplitude difference to the mean absolute amplitude value in each frame of the speech frame-level structural parameter sequence, the ratio of adjacent frames is continuously screened to locate frames in which the ratio is less than a predefined dynamic range ratio threshold (a standard ratio value used in speech signal processing to detect sparse amplitude fluctuations (default 0.3). When the ratio of the maximum and minimum amplitude difference to the mean absolute amplitude value is lower than this threshold, it is determined to be a sparse speech segment. This threshold is determined based on statistical analysis of a healthy population speech database). Adjacent segments are then merged in chronological order to generate a list of sparse amplitude fluctuation segments.
[0052] S3: Based on the mean absolute amplitude value in the speech frame-level structural parameter sequence, perform inter-frame difference calculations on all frames and record the positions where two sets of positive and negative opposite relationships exist in the continuous differences. Screen whether the continuous decreasing trend of the mean amplitude value in the subsequent frames corresponding to the position satisfies a significant decreasing trend (the continuous amplitude mean decreasing feature defined in speech analysis (default decrease of more than 10% for 3 consecutive frames), used to identify speech breakpoints. This trend is verified through Alzheimer's disease speech feature research). Mark it as the start and end position of abrupt breakpoints, integrate the time identifiers of all breakpoint segments, and generate a set of continuously abrupt breakpoint segments.
[0053] S4: Based on the list of sparse amplitude fluctuation segments and the set of continuous jump fracture segments, find the combination segments in the two sets whose start and end times overlap by more than half the length of their respective segments, extract the median time point of each overlapping segment and establish a trend structure line in chronological order, determine the direction of the structure line and record the number of points, and generate an abnormal time trend structure index set.
[0054] S5: Based on the trend structure points and corresponding frame segment numbers marked in the abnormal time trend structure index set, summarize the speech frame segments with continuous jump trends, perform unified numbering and label classification, establish frame order and time stamp index, and generate Alzheimer's disease speech feature recognition results.
[0055] The speech frame-level structural parameter sequence includes frame number index, frame-level amplitude absolute mean sequence, frame-level maximum and minimum amplitude difference sequence, and frame-level energy value sequence. The amplitude fluctuation sparse segment list includes sparse segment start and end time, sparse segment frame index, and sparse segment ratio sequence. The continuous jump break segment set includes break segment start and end position, continuous falling frame segment number, and amplitude mean difference sequence. The abnormal time trend structure index set includes trend structure point time median, trend direction change indicator, and number of structure points. The Alzheimer's disease speech feature recognition results include abnormal trend frame segment number index, trend label classification results, and time series markers.
[0056] Please see Figure 2 Step S1 is as follows:
[0057] S111: Obtain the continuous sampling point sequence of the patient's speech signal, divide it into frames at equal intervals along the time axis, set the frame length and frame shift parameters, divide the entire speech signal into multiple adjacent speech frames, extract the absolute value sequence of amplitude in each frame, and generate the frame mean amplitude sequence.
[0058] After acquiring the continuous sampling point sequence of the patient's speech signal, a frame division operation is first performed on the entire speech signal. The frame length is set to 25 milliseconds, the frame shift to 10 milliseconds, and the signal is divided into multiple overlapping frame segments along the time axis. Each frame segment may contain 200 to 400 sampling points. Assuming the sampling frequency is 16kHz, each frame contains 400 sampling points. The sampled data are numbered sequentially as follows: After frame division, the absolute amplitude of each sample point in each frame is processed and converted into a positive number. This operation can be used for subsequent signal amplitude correlation statistics. The sequence of absolute sample point values is as follows: During the sampling process, a set of original audio signals with a length of 3 seconds was selected, and the sampling frequency was 16kHz. The total number of samples was 48,000 points, which could be divided into approximately 315 frames of data. Then, the absolute value mean of the amplitude was extracted from each frame sequentially. For example, the 10th frame of sampled data... After absolute value processing, it becomes The mean value is calculated to be 0.3, indicating that the overall amplitude level of the signal in this frame is 0.3. Based on this, the mean values of each frame are arranged sequentially to form a preliminary statistical sequence, namely the frame mean amplitude sequence, as shown below:
[0059] Table 1: Examples of Frame Mean Amplitude
[0060] Frame number Mean amplitude (unit: V) 1 0.29 2 0.31 3 0.26 4 0.28 … … 315 0.33
[0061] As shown in Table 1, the frame mean amplitude reflects the overall situation of signal amplitude fluctuation within each frame. Subsequent steps can use this sequence to further analyze the instantaneous dynamic change trend of the speech signal and finally obtain the frame mean amplitude sequence.
[0062] S112: Based on the frame mean amplitude sequence, compare the maximum and minimum amplitude values in each frame to extract the amplitude range feature of each frame, perform inductive processing based on the original absolute amplitude value sequence of each frame, pair and associate it with the frame number to obtain the frame-level structure index sequence.
[0063] After calling the frame mean amplitude sequence, it is necessary to extract the maximum and minimum values of the absolute amplitude sequence in each frame to reflect the range of signal fluctuations. The amplitude range is obtained by comparison. For example, the sampled absolute value sequence of frame 25 is... The maximum value is 0.2, the minimum value is 0.05, and the range is 0.15. After extracting the range values for each frame sequence, a range value sequence is formed for subsequent processing. Next, energy values are further collected based on each frame. These energy values can be obtained by performing an induction operation on the absolute value sequence. The specific computational logic is not involved here; the energy of each frame is represented through data encapsulation. For example, the energy of frame 25 is inductively derived from the total amplitude characteristics of the sampling points and defined as a representative value of 0.21. This energy value represents the relative amplitude of the signal within that frame. The activity level is then used to establish an index mapping between the amplitude range and the energy value and the frame number, forming a frame-level structured record sequence with a sequential index structure. In practice, the three indicators of the 25th frame are grouped into a triplet such as (0.3, 0.15, 0.21) and labeled with the frame number 25, indicating that its intra-frame amplitude mean is 0.3, amplitude range is 0.15, and energy level is 0.21. By combining the corresponding triplets of all frames with the frame number, a data structure set with position index characteristics is finally formed, which is represented as a frame-level structured index sequence.
[0064] S113: Based on the frame-level structure index sequence, arrange the frame numbers in ascending order, and combine the mean absolute value of amplitude, the amplitude range value, and the relevant sequence content of each frame into a single frame structure parameter group. Summarize all frame-level structure parameter groups to obtain the speech frame-level structure parameter sequence.
[0065] Based on the frame-level structure index sequence, the system traverses the frames in ascending order by frame number, extracting the three corresponding parameters for each frame: the mean absolute value of amplitude, the amplitude range, and the energy value. These three parameters are then merged into a frame structure parameter group according to a fixed arrangement. For example, taking frame 40 as an example, if its three parameters are 0.31, 0.16, and 0.22, its structure parameter group can be represented as (0.31, 0.16, 0.22). This parameter group is stored as an independent unit and appended to the end of the structure parameter sequence in numerical order. By traversing all frame numbers, the structure parameter groups of each frame are sequentially integrated into a unified data sequence, forming a complete sequence structure representation. This sequence contains all frame structure parameter units, totaling 315 groups, representing the trend of structure parameter changes from frame 1 to frame 315, while retaining the original frame number order. Within this structure sequence, parameters of any frame can be retrieved and compared subsequently. The final integrated complete data structure is the speech frame-level structure parameter sequence.
[0066] Please see Figure 3 Step S2 is as follows:
[0067] S211: Based on the speech frame-level structural parameter sequence, the maximum and minimum amplitude differences and the mean absolute amplitude values of each frame are extracted sequentially. The ratio between the amplitude range and the mean is used for combination. For the ratio corresponding to consecutive frame numbers, the formula is:
[0068] ;
[0069] The fluctuation ratio is calculated and obtained. Combined with the frame numbering order, a frame index sequence structure is established to obtain a continuous frame fluctuation ratio index sequence; where... Indicates the first Frame fluctuation ratio, This represents the difference between the maximum and minimum amplitudes in that frame. This represents the mean absolute value of the amplitude of the frame. Indicates the first The median of all absolute amplitude samples in the frame;
[0070] Based on the speech frame-level structural parameter sequence, it is necessary to extract the maximum and minimum amplitude difference and the mean absolute amplitude value of each frame. Each frame structure consists of sampling points at a sampling frequency of 16kHz and a frame length of 25ms, thus each frame contains 400 sampling points. For example, the sampling data in frame 12 is... Take the absolute value The maximum value is 0.36, the minimum value is 0.14, the amplitude range is 0.22, and the mean is 0.246. The ratio of the range to the mean is established for this frame to represent the amplitude fluctuation level within the frame. For all frames, the frame numbers are sequentially traversed to obtain the two parameters mentioned above for each frame. In this sub-step, the frame median is introduced as an auxiliary factor for symmetrical fluctuation. The offset between the median and the mean is further calculated and included in the ratio calculation. In the 12th frame, the median is 0.25, the mean is 0.246, and the offset is 0.004. Based on the above parameters, a ratio index is constructed to describe the amplitude fluctuation of the signal in the frame-level structure. When recording the results, it is necessary to maintain a synchronized structure with the frame number, forming a one-to-one correspondence between the number and the ratio. The example parameters are then used for calculation.
[0071] ;
[0072] ;
[0073] ;
[0074] ;
[0075] ;
[0076] Substitute into the original expression:
[0077] ;
[0078] This value indicates that the fluctuation ratio of frame 12 is 0.4663, which is above the 0.3 threshold and does not belong to the sparse segment candidate frame. The following are example calculation data for some frames:
[0079] Table 2. Amplitude Fluctuation Ratio Index
[0080] Frame number maximum amplitude minimum amplitude range mean median Ratio index 12 0.36 0.14 0.22 0.246 0.25 0.4663 13 0.40 0.19 0.21 0.265 0.28 0.4028 14 0.33 0.15 0.18 0.23 0.24 0.3879
[0081] As shown in Table 2, the ratio of each frame reflects the degree of amplitude fluctuation within the frame. This structure can be used to further screen the continuity between frames and finally obtain the sequence of continuous frame fluctuation ratio indicators.
[0082] The fluctuation ratio is used to measure the amplitude fluctuation of a speech signal within a single frame. It is an important numerical indicator reflecting the amplitude change state within a frame. Its core significance lies in expressing the relative proportional relationship between the difference between the maximum and minimum amplitudes and the mean amplitude, thus standardizing the amplitude fluctuation of the signal within a short time interval, thereby eliminating the interference of the overall energy level on the amplitude fluctuation judgment. The smaller the ratio, the more gentle and sparse the amplitude fluctuation in the frame signal, and vice versa, indicating that the signal has significant fluctuations. This indicator can capture speech segments with stable amplitudes, such as silent segments and soft-spoken segments, and can also assist in subsequent identification of structural segments such as speech boundaries and silent regions. It is an important intermediate parameter for realizing the structural segmentation of speech signals.
[0083] The formula's operational logic is based on a detailed characterization of the amplitude fluctuations within a speech frame, with the numerator using the range value. absolute difference between the median and the mean The addition operation aims to simultaneously measure the maximum fluctuation amplitude of the signal within a frame and the degree of deviation in the symmetry of the data distribution. The former reflects the difference in extreme values of the signal, while the latter characterizes the deviation of the amplitude distribution. The linear superposition of the two can comprehensively characterize the discreteness of the frame-level amplitude. The denominator uses the mean. The square root of the product of the median and the range The sum is used to balance the misjudgment of the ratio caused by the difference in signal strength between different frames. The mean serves as the baseline amplitude index, while the square root term of the product can adjust the downward trend of the ratio when the amplitude is high but the fluctuation is limited. The square root operation plays a convergence role and alleviates the nonlinear amplification caused by extreme amplitude values. By combining the structure between the numerator and denominator, a more discernible frame fluctuation ratio index is constructed, so as to extract stable and continuous sparse fluctuation feature segments in the low amplitude fluctuation region.
[0084] S212: Based on the continuous frame fluctuation ratio index sequence, perform inter-frame screening operation according to the preset dynamic range ratio threshold, determine whether the intra-frame ratio is in the sparse segment judgment interval, continuously record the frame number index that meets the condition, and obtain the continuous sparse frame sequence index.
[0085] The system calls the continuous frame fluctuation ratio index sequence and sets the dynamic range threshold for ratio judgment to 0.3. The threshold is derived from the statistical median value of the lower limit of amplitude ratio obtained from the analysis of a large number of samples in the healthy speech database. If the ratio of a frame is less than this value, it is considered that the amplitude fluctuation is sparse. The default lower limit is 0 and the upper limit is 0.3. The screening mechanism is set to continuous scanning. The frame numbers with a ratio lower than the threshold are registered one by one, and it is determined whether they have continuity. Taking frames 101 to 108 as an example, the ratios are 0.28, 0.25, 0.21, and 0, respectively. The ratios are 18, 0.15, 0.27, 0.29, and 0.31. Among them, the ratios of frames 101 to 107 are all less than 0.3, which meets the sparsity standard. The ratio of frame 108 is 0.31, which does not meet the threshold limit. Therefore, registration is stopped, forming a set of continuous sparse frame numbers. This process is repeated for the entire frame number sequence. All frame groups that meet the continuous low ratio are marked and form a set of frame numbers. Each item in the set can be regarded as a continuous frame. Finally, all sets are merged to form a unified index structure, namely the continuous sparse frame sequence index.
[0086] S213: Based on the continuous sparse frame sequence index, screening and merging are performed according to the continuity of frame number. If the frame gap is less than two frames, it is merged into the same segment; otherwise, it is divided into independent segments. All sparse frame sequences are traversed in turn and the merging process is completed. Each merged segment is recorded as the segment boundary according to the frame number of the start and end frames, and the time position is marked to obtain the list of sparse segments with amplitude fluctuation.
[0087] Based on the continuous sparse frame sequence index, frame segment merging is required. The number of frames between adjacent sparse frame segments is determined. If the gap between frames is no more than two frames, it can be considered a continuous segment. For example, the first group of sparse frames consists of frames 100 to 105, and the second group consists of frames 108 to 110, with frames 106 and 107 in between, totaling two frames. This meets the merging rule and is merged into a new segment of frames 100 to 110. If the gap between segments exceeds two frames, it is retained as two independent segments. The merging process... The system traverses the set of frame numbers in chronological order and records the start and end numbers of the merged segments in real time. When converting to actual time points, it can be multiplied by the frame shift time. For example, if each frame interval is 10ms, then frame numbers 100 to 110 represent the segment from 1.00 seconds to 1.10 seconds. After the sequence traversal is completed, the frame boundaries of all sparse segments are established. A recording structure with four dimensions, namely segment number, frame start point, frame end point, and time start and end, is established and summarized to form the final result list, namely the list of sparse segments with amplitude fluctuations.
[0088] Please see Figure 4 Step S3 is as follows:
[0089] S311: Based on the mean absolute value of amplitude in the speech frame-level structural parameter sequence, the mean values of adjacent frames are extracted sequentially according to frame number, and frame difference calculation is performed to obtain the mean difference of consecutive frames. If the result is positive, it indicates an increase in amplitude; if negative, it indicates a decrease. The sign of the difference in each frame is recorded to form a sequence of difference change signs. The positions where the sign reverses twice consecutively are found in the sequence, i.e., consecutive reversal groups where the sign changes from positive to negative or from negative to positive. An index of the continuous change positions is established by reversing the frame number, using the formula:
[0090] ;
[0091] The calculation obtains the continuous jump trend indicator value. If there are two difference reversals before the corresponding position, it is marked as a candidate segment for a jump change; where, , , , Representing frames respectively to frame The mean absolute value of the amplitude, This is a trend strength indicator within the continuous interval, with the symbol... This means taking the absolute value of the relative rate of change, ensuring that the judgment focuses only on the intensity and not on the directionality.
[0092] Based on the mean absolute amplitude value in the speech frame-level structural parameter sequence, the mean of each frame is extracted in order of frame number, and then the mean is subtracted from the next adjacent frame to obtain the direction of mean change between the two frames, constructing an inter-frame difference sequence. If the difference is positive, it indicates an increase in amplitude; if it is negative, it indicates a decrease in amplitude. The positive and negative sign sequences of all difference results are recorded, and by judging whether there is a sign reversal between two consecutive sets of positive and negative signs, it is identified whether the signal amplitude trend has undergone a continuous directional change, such as changing from positive to negative or from negative to positive twice. For example, the mean values of frames 25, 26, 27, and 28 are 0.33 and 0, respectively. 36, 0.32, 0.29, the differences are +0.03, -0.04, -0.03, corresponding to +, −, − signs, indicating one flip. If the difference in subsequent frame 29 is +0.02, then another flip occurs, satisfying the double flip criterion. The position index of this sequence is recorded as the candidate starting point of the mutation. Subsequently, the intensity evaluation of the continuous decreasing feature in this segment is introduced. Let the average amplitude of frames k to k+3 be 0.42, 0.37, 0.32, 0.28 respectively. Then the relative decreasing amplitudes of the three segments are 11.9%, 13.5%, and 12.5% respectively. The following three-segment ratio calculation method is used to substitute in the values:
[0093] Paragraph 1: ;
[0094] Paragraph 2: ;
[0095] Paragraph 3: ;
[0096] The three-segment overall trend indicator is: ;
[0097] If this indicator is greater than the continuous decline threshold of 0.30, it is considered a clear jump trend; the calculation data are summarized below:
[0098] Table 3. Mean Amplitude and Jump Trend Indicators of Continuous Frames
[0099] Frame number mean Adjacent difference Difference sign Ratio 1 Ratio 2 Ratio 3 Jump indicator 25 0.33 +0.03 + 26 0.36 −0.04 − 27 0.32 −0.03 − 28 0.29 +0.02 + 75 0.42 −0.05 − 0.119 0.135 0.125 0.379
[0100] As shown in Table 3, this jump index value is used to determine whether there is a significant fluctuation sequence between frames. The minimum segment span is 4 frames, and the final screening result is used to obtain the continuous jump trend index value.
[0101] The trend strength index is used to measure the change in the mean absolute value of the amplitude of multiple consecutive frames in a speech signal over time. Essentially, it quantifies whether there is a stable and obvious fluctuation trend in the signal within a local time window. By calculating and accumulating the relative rate of change of the mean amplitude between adjacent frames, this index can reveal whether there is a dynamic process of continuous decline or rise in the speech signal. It is especially valuable for detecting structural changes such as speech breaks and sudden changes in speech flow. The larger the trend strength index value, the stronger the amplitude fluctuation within that interval. If the index reaches a certain threshold and the direction is consistent, it can be judged as a jump segment or break segment with structural turning characteristics. It is an indispensable intermediate variable in the recognition of the structural boundary of speech signals.
[0102] The formula's operational logic aims to quantify the overall downward trend of the average amplitude across consecutive frames. It calculates the relative rate of change for three consecutive sets of adjacent frames and sums the absolute values of each rate of change to reflect the overall fluctuation of the average amplitude within the four-frame interval. , , They represent the first Frame to the Frame, First Frame to the Frame, First Frame to the The relative rate of change of frames uses absolute value operation to unify the dimension direction, focusing only on the intensity of change without considering its direction of increase or decrease. The three-term summation structure design can avoid misjudging the overall trend of change due to small fluctuations in a certain frame, and enhance the cumulative evaluation ability of the overall amplitude mean trend within the continuous change range. At the same time, the ratio form is used to normalize speech frames with different energy levels, which is suitable for speech break detection tasks under different speakers or different speech rates.
[0103] S312: Based on the continuous jump trend index value, read the sequence and screen it in the order of frame number. For each item, determine whether the next three frames constitute a continuous downward structure, determine whether the drop amplitude exceeds the set ratio value, set the continuous drop judgment threshold, correspond to the abnormal speech segment definition standard in Alzheimer's disease speech samples, record the effective mutation area, and obtain the continuous mean drop frame segment set.
[0104] Based on the continuous jump trend index value, all jump point positions that meet the preconditions are extracted. For each jump point, a four-frame sliding window is constructed by indexing three frames backward. A continuous comparison operation of the inter-frame mean is performed. If the mean of the next three frames is less than that of the previous frame and satisfies a decreasing relationship (e.g., frame m is 0.50, frame m+1 is 0.45, frame m+2 is 0.40, and frame m+3 is 0.35), and the decrease ratios between adjacent frames are 10%, 11.1%, and 12.5% respectively, then it is identified as a continuous downward trend segment. If the decrease ratio does not meet the 10% standard (e.g., 8% or 6%), then the candidate jump segment is removed. The continuous decrease threshold is set to 0.10, which is derived from the abnormal voice segment change threshold statistically obtained from Alzheimer's disease speech data. Based on the analysis of the mean decrease amplitude of 300 speech samples, the 10% setting value is determined as the dividing point to ensure data reproducibility. Finally, the frame segments that meet the decrease characteristic conditions are selected from all jump candidates to obtain the continuous mean decrease frame segment set.
[0105] S313: Based on the set of continuously decreasing average frame segments, merge all candidate mutation segments continuously according to frame number. If the interval between two mutation segments is less than or equal to two frames, they are considered as the same segment. Construct a unified number for the broken segments. After the segments are merged, convert the frame number of each segment into a time interval. Record the start and end times of all segments in a unified manner to generate a set of continuously jumping broken segments.
[0106] Based on a set of continuously decreasing average frame segments, the intervals between all segments are detected sequentially by frame number. If the interval between two segments is less than or equal to 2 frames, they can be merged into the same segment. After merging the segments, the start and end frame numbers of the segment are recorded, and the frame numbers are converted into time start and end positions according to a frame shift of 10ms. For example, if the frame numbers are 110 to 114, the corresponding time interval is 1.10 seconds to 1.14 seconds. All time intervals of the jump segments are recorded in this way, and the output result is in the format of four columns: segment number, start frame, end frame, and start and end time. For example, segment 1 is frame number 90 to 94, corresponding to a time interval of 0.90 to 0.94 seconds; segment 2 is frames 99 to 102, corresponding to 0.99 to 1.02 seconds. If the interval between the start frame of segment 2 and the end frame of segment 1 is less than or equal to 2 frames, they are merged into segment 1, starting frame 90 to ending frame 102, and the time interval is updated to 0.90 seconds to 1.02 seconds. Finally, all time series data of the jump segments are constructed and organized to obtain a set of continuous jump break segments.
[0107] Please see Figure 5Step S4 is as follows:
[0108] S411: Based on the list of sparse amplitude fluctuation segments and the set of continuous jump fracture segments, extract the start and end points of time for each frame segment, and construct all frame segments in the two sets into independent time intervals. Perform time axis matching operation segment by segment to determine whether there are overlapping frame segments between any two segments. Check the proportion of the overlapping part in the two original segments. If the proportion is more than half of the frame length in both segments, it is determined to be a valid overlapping segment. Mark all segments that meet the conditions according to the corresponding relationship to obtain a set of bidirectional overlapping frame segments.
[0109] Based on the frame time start and end point data from the list of sparse amplitude fluctuation segments and the set of continuous jump break segments, the start and end times of each segment in both types of segments are first extracted, and time intervals are constructed in milliseconds. For example, the start and end times of a sparse segment are 1.20 seconds to 1.40 seconds, corresponding to 20 frames, and the jump segment time is 1.30 seconds to 1.50 seconds, also corresponding to 20 frames. Then, the time ranges of these two segments are cross-compared, and the overlapping segment is found to be 1.30 seconds to 1.40 seconds, which is 10 frames. The overlapping frame count accounts for 50% of the total frame count of each segment, meeting the set bidirectional overlap standard. By comparing multiple segment pair combinations, the existence of similar conditions is screened one by one to obtain all combination segments that meet the frame overlap ratio requirement. These are then classified and organized by segment pair combination. To quantify the coverage, three parameters are extracted from each combination segment: the number of sparse segment frames, the number of jump segment frames, and the number of overlapping frames. The combination segment examples in the test set are as follows:
[0110] Table 4. Example of Overlapping Frame Coverage Ratio
[0111] Combination number sparse segment frame count Jump segment frame count Overlapping frames Sparse ratio Jump percentage 1 20 20 10 0.50 0.50 2 18 24 12 0.67 0.50 3 30 20 8 0.27 0.40
[0112] As shown in Table 4, the bidirectional ratios in Combination 1 and Combination 2 both reach or exceed the 50% threshold and are judged as valid overlapping segments. Combination 3 has an insufficient ratio and does not meet the conditions, so it is removed. Finally, a set of bidirectional overlapping frames is extracted.
[0113] S412: Based on the set of bidirectional overlapping frames, extract the start frame number and end frame number of all combined segments that meet the conditions, calculate the middle frame of the time interval as the representative point, convert the representative point into the corresponding time value, construct the median time point sequence, and arrange them in chronological order to obtain the complete time trend point trajectory and establish the median time point sequence.
[0114] Based on the set of bidirectional overlapping frames, the time interval information of each overlapping segment is extracted, that is, the start frame number and the end frame number are extracted. Based on the known frame shift of 10 milliseconds, the frame number is converted into a time value. For example, the start time of frame 140 is 1.40 seconds and the end time of frame 159 is 1.59 seconds. There are 20 frames in the segment. After sorting them in order of frame number, the middle frame number is frame 149, and the corresponding time is 1.49 seconds. This time is the median time point of the segment. The same operation is performed on all overlapping segments in this way to obtain all median time points. They are sorted and arranged in chronological order to construct a sequence. If there are different combinations of segments with the same median time point, the frame number priority method can be used to remove duplicates. Finally, the set of all single points representing the overlapping trend of the time period is obtained, and the median time point sequence is established.
[0115] S413: Based on the median time point sequence, determine the sequential relationship between adjacent time points, mark the case where the later time point is earlier than the previous time point as reverse, and the case where it is later as positive, mark the continuity of the overall trend direction and the number of direction changes, record the number and distribution characteristics of positive and reverse points, and combine the median time point density information to obtain the abnormal time trend structure index set.
[0116] Based on the median time point sequence, all time point pairs are traversed in chronological order. The magnitude relationship between two adjacent time points is compared. If the later time point is greater than the previous time point, it is recorded as a positive trend; otherwise, it is recorded as a negative trend. The trend direction is saved in a trend label list in the form of a tag. At the same time, the number of direction changes in the entire sequence is counted, that is, the total number of times the direction changes from positive to negative or vice versa. In addition, the total number of positive and negative segments and the number of points in each segment are also recorded to construct a positive and negative trend distribution structure. For example, if the median time point sequence is 1.20, 1.32, 1.25, 1.38, 1.31, 1.45 seconds, then the trend direction is upward, downward, upward, downward, upward, with a total of four direction changes, three positive segments and two negative segments, each segment with a length of 1 to 2 points. The trend labels, number of changes, and length of each segment are further summarized into a structured record to form a trend fluctuation cycle table. Combined with the overall density of the median points, the structure is classified, and finally, an abnormal time trend structure indicator set is obtained.
[0117] Please see Figure 6 The S5 steps are as follows:
[0118] S511: Based on the trend structure points marked in the abnormal time trend structure index set and the corresponding frame segment numbers, extract the start and end numbers of the frame segments covered by each trend segment, convert all frame numbers into the corresponding timestamp sequence in sequence, arrange the trend segments in chronological order, compare the number intervals of two adjacent trend segments, if the number interval value is within the allowable range, it is judged as the same segment with continuous trend, and merged into the same jump trend voice frame segment to establish a set of continuous jump trend frame segments.
[0119] Based on the trend structure points and frame segment numbers in the abnormal time trend structure indicator set, the frame segment range associated with each trend point is first analyzed to extract its corresponding start and end frame numbers. These are then converted into corresponding timestamps using a 10-millisecond frame shift standard. During this process, it's crucial to ensure the continuity of the mapping between frame numbers and time. Subsequently, all extracted frame segments are sorted in ascending order by their number, and the numbering interval between adjacent frame segments is determined. If the numbering interval between any two adjacent frame segments does not exceed 5 frames, it is considered to have temporal continuity. These frame segments are then aggregated into a single trend segment. For example, if frame segment A has start and end numbers of 2... From 05 to 215, and from 216 to 225, the frame numbering of segment B is 1 frame, which meets the continuity requirement. Therefore, they can be merged into trend segments numbered 205 to 225. In addition, in actual speech samples, if there are regions of continuous rising or falling pitch changes in a speech segment, the corresponding trend points will appear densely in the region, and then aggregate into jump trend segments. This process ultimately forms multiple jump trend segment sets, which constitute the basic data framework of continuous change structure. This type of operation is suitable for speech stream processing tasks that need to identify continuous voiceprint changes in speech detection scenarios, and the result is a set of continuous jump trend segments.
[0120] S512: Based on the set of continuous jump trend frames, assign an independent number to each frame segment, and set a label based on the directional characteristics of the trend change in the frame segment. Summarize the number, frame start and end number, time start and end value, and trend label content to form an index list and establish a jump trend frame segment label index table.
[0121] Based on a set of continuously changing trend frames, each changing segment is assigned a unique number label. A standardized numbering and naming convention is used for formatting, such as naming the first segment "TJ_001", the second segment "TJ_002", and so on. To further identify the directional characteristics of each segment, the directional characteristics of the trend structure points within each segment need to be identified and categorized. If all trend points within a segment are trending upwards, the segment is labeled "UP"; if they are trending downwards, it is labeled "DOWN"; if they alternate between upwards and downwards, it is labeled "MIX". Taking a specific trend frame segment, TJ_007, as an example, it contains 10 structure points, with 7 positive and 3 negative directions. Therefore, it is judged to be a mixed trend direction and labeled "MIX". Subsequently, all information such as the number, frame start and end numbers, corresponding timestamps, and trend direction labels are output in a unified format to construct a clearly structured label table for subsequent processing or retrieval. See the example in Table 5.
[0122] Table 5. Frame Tag Index Table for Jump Trends
[0123] serial number Frame start point Frame End Point Time start point (ms) End of time (ms) Trend tags TJ_001 120 136 1200 1360 UP TJ_002 137 158 1370 1580 DOWN TJ_003 159 180 1590 1800 MIX
[0124] As shown in Table 5, each segment is clearly marked according to its trend direction and frame position for subsequent recognition of Alzheimer's disease speech features. The final result is a jump trend frame segment label index table.
[0125] S513: Based on the trend change direction, number of frames, time span, label type and intra-frame trend switching status attribute information of each frame segment in the jump trend frame segment label index table, according to the set frame segment distribution judgment conditions, the trend stability and trend switching characteristics of each frame segment are screened and layered, and the data of different feature labels are integrated to obtain the Alzheimer's disease speech feature recognition results.
[0126] The system retrieves core fields from the jump trend frame segment label index table, including frame segment number, frame start and end, time interval, and trend label, to categorize and filter the status of each trend segment. First, based on the trend label, frames are categorized into three types: UP, DOWN, and MIX, and the number of frames in each type is counted. Then, numerical parameters such as the frame length distribution range, average number of frames, and time span are extracted for each type of trend segment. For example, the average frame length in the UP category is 18 frames, and the average time span is 180 milliseconds. Further, the switching frequency between frames within the same category is detected, such as TJ_ When the label between frame segments 003 and TJ_004 changes from DOWN to UP, a switching event is counted. When the number of consecutive switching events exceeds 3, the speech segment can be marked as a trend abnormal cluster segment. Such segments usually appear in the speech stream as areas of discontinuous speech or fluctuating intonation. A special identification field is set for such segments. Based on the integration of information such as segment numbers, labels, and time ranges, the marking results are summarized and output in a structured manner to establish an output dataset for Alzheimer's disease symptom determination. The final result is the speech feature recognition result for Alzheimer's disease.
[0127] An Alzheimer's disease identification system based on speech features includes:
[0128] The speech feature extraction module is used to perform S1: acquire a continuous sampling point sequence of the patient's speech signal, perform frame division processing on the signal, extract the mean absolute value of amplitude, the difference between the maximum and minimum amplitude and the energy value in each frame, combine the parameters in sequence and arrange them in sequence according to the frame number to generate a speech frame-level structural parameter sequence;
[0129] The low-amplitude region identification module is used to perform S2: based on the ratio of the maximum and minimum amplitude difference to the mean absolute amplitude value of each frame in the speech frame-level structural parameter sequence, it locates the frame segments in consecutive frames whose ratio is less than the predefined dynamic range ratio threshold, and merges adjacent segments in chronological order to generate a list of sparse amplitude fluctuation segments.
[0130] The module for marking broken segments is used to execute S3: based on the mean absolute value of amplitude in the speech frame-level structural parameter sequence, perform inter-frame difference operation on all frames and record the positions where there are two sets of positive and negative opposite relationships in the continuous difference, mark the start and end positions of abrupt breaks, and generate a set of continuous abrupt break segments;
[0131] The integrated trend anomaly module is used to execute S4: Based on the list of sparse amplitude fluctuation segments and the set of continuous jump fracture segments, the median time point of each overlapping segment is extracted and a trend structure line is established in chronological order. The direction of the structure line is judged and the number of points is recorded to generate an abnormal time trend structure index set.
[0132] The collection and recognition result module is used to execute S5: based on the trend structure points marked in the abnormal time trend structure index set and the corresponding frame segment number, it summarizes the continuous jump trend speech frame segments, performs unified numbering and label classification, establishes frame order and time stamp index, and generates Alzheimer's disease speech feature recognition results.
[0133] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention in any other way. Any person skilled in the art may make changes or modifications to the above-disclosed technical content to create equivalent embodiments that can be applied to other fields. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the protection scope of the present invention.
Claims
1. A method for Alzheimer's disease identification based on speech features, characterized in that, Includes the following steps: S1: Obtain the continuous sampling point sequence of the patient's speech signal, perform frame division processing on the signal, extract the mean absolute value of amplitude, the difference between the maximum and minimum amplitude and the energy value in each frame, combine the parameters in sequence and arrange them in sequence according to the frame number to generate a speech frame-level structural parameter sequence. S2: Based on the ratio of the maximum and minimum amplitude difference to the mean absolute value of amplitude in each frame of the speech frame-level structural parameter sequence, locate the frame segments in continuous frames whose ratio is less than the predefined dynamic range ratio threshold, and merge adjacent segments in chronological order to generate a list of sparse amplitude fluctuation segments. S3: Based on the mean absolute value of amplitude in the speech frame-level structural parameter sequence, perform inter-frame difference operation on all frames and record the positions where there are two sets of positive and negative opposite relationships in the continuous difference, mark the start and end positions of abrupt breakage, and generate a set of continuous abrupt breakage segments. S4: Based on the list of sparse amplitude fluctuation segments and the set of continuous jump fracture segments, extract the median time point of each overlapping segment and establish a trend structure line in chronological order. Perform directional judgment and record the number of points on the structure line to generate an abnormal time trend structure index set.
2. The Alzheimer's disease identification method based on speech features according to claim 1, characterized in that, The speech frame-level structural parameter sequence includes a frame number index, a frame-level amplitude absolute value mean sequence, a frame-level maximum and minimum amplitude difference sequence, and a frame-level energy value sequence. The amplitude fluctuation sparse segment list includes the start and end times of the sparse segments, the sparse segment frame index, and the sparse segment ratio sequence. The continuous jump break segment set includes the start and end positions of the break segments, the continuous decreasing frame segment numbers, and the amplitude mean difference sequence. The abnormal time trend structure index set includes the trend structure point time median, trend direction change identifier, and the number of structure points.
3. The Alzheimer's disease identification method based on speech features according to claim 1, characterized in that, The steps for obtaining the speech frame-level structure parameter sequence are as follows: S111: Obtain the continuous sampling point sequence of the patient's speech signal, divide it into frames at equal intervals along the time axis, set the frame length and frame shift parameters, divide the entire speech signal into multiple adjacent speech frames, extract the absolute value sequence of amplitude in each frame, and generate the frame mean amplitude sequence. S112: Based on the frame mean amplitude sequence, compare the maximum amplitude value and the minimum amplitude value in each frame, extract the amplitude range feature of each frame, perform inductive processing based on the original absolute amplitude value sequence of each frame, pair and associate it with the frame number, and obtain the frame-level structure index sequence. S113: Based on the frame-level structure index sequence, arrange the frame numbers in ascending order, and combine the mean absolute value of amplitude, the amplitude range value and the relevant sequence content of each frame into a single frame structure parameter group. Summarize all frame-level structure parameter groups to obtain the speech frame-level structure parameter sequence.
4. The Alzheimer's disease identification method based on speech features according to claim 1, characterized in that, The specific steps for obtaining the list of sparse amplitude fluctuation segments are as follows: S211: Based on the speech frame-level structural parameter sequence, extract the maximum and minimum amplitude difference and the mean absolute amplitude value of each frame in sequence, combine them using the proportional relationship between the amplitude range and the mean, calculate the fluctuation ratio for the ratio corresponding to the consecutive frame number, and establish a frame index sequence structure in combination with the frame number order to obtain the consecutive frame fluctuation ratio index sequence. S212: Based on the continuous frame fluctuation ratio index sequence, perform inter-frame screening operation according to the preset dynamic range ratio threshold, determine whether the intra-frame ratio is in the sparse segment determination interval, continuously record the frame number index that meets the condition, and obtain the continuous sparse frame sequence index. S213: Based on the continuous sparse frame sequence index, perform screening and merging processing according to the continuity of frame number. If the frame gap is less than two frames, merge them into the same segment; otherwise, divide them into independent segments. Iterate through all sparse frame sequences and complete the merging processing. Record each merged segment as the segment boundary according to the frame number of the start and end frames, mark the time position, and obtain the list of sparse segments with amplitude fluctuation.
5. The Alzheimer's disease identification method based on speech features according to claim 1, characterized in that, The specific steps for obtaining the set of continuously abruptly changing fracture segments are as follows: S311: Based on the mean absolute value of amplitude in the speech frame-level structural parameter sequence, extract the mean items of two adjacent frames in the order of frame number and perform frame difference calculation to obtain the mean difference of consecutive frames. If the result is positive, it indicates that the amplitude is increasing, and negative value indicates that it is decreasing. Record the difference sign of each frame to form a difference change sign sequence. Find two consecutive sign reversal positions in the sequence, that is, consecutive reversal groups where the sign changes from positive to negative or from negative to positive. Establish a continuous change position index through the reversal frame number and calculate to obtain the continuous jump trend index value. If there are two difference reversals before the corresponding position, mark it as a jump change candidate segment. S312: Based on the continuous jump trend index value, read the sequence and screen it in the order of frame number. For each item, determine whether the subsequent three frames constitute a continuous downward structure, determine whether the drop amplitude exceeds the set ratio value, set the continuous drop judgment threshold, and define the abnormal speech segment definition standard in Alzheimer's disease speech samples. Record the effective mutation area and obtain the set of continuous mean drop frame segments. S313: Based on the set of continuously decreasing average frame segments, merge all candidate mutation segments continuously according to frame number. If the interval between two mutation segments is less than or equal to two frames, they are considered as the same segment. Construct a unified number for the broken segments. After the segments are merged, convert the frame number of each segment into a time interval. Record the start and end times of all segments in a unified manner to generate a set of continuously jumping broken segments.
6. The Alzheimer's disease identification method based on speech features according to claim 1, characterized in that, The specific steps for obtaining the abnormal time trend structure indicator set are as follows: S411: Based on the list of sparse amplitude fluctuation segments and the set of continuous jump break segments, extract the start and end points of time for each frame segment, and construct all frame segments in the two sets as independent time intervals. Perform time axis matching operation segment by segment, determine whether there are overlapping frame segments between any two segments, check the proportion of the overlapping part in the two original segments, if the proportion is more than half of the frame length in both segments, it is determined to be a valid overlapping segment, and mark all segments that meet the conditions according to the corresponding relationship to obtain a set of bidirectional overlapping frame segments; S412: Based on the bidirectional overlapping frame segment set, extract the start frame number and end frame number of all combined segments that meet the conditions, calculate the middle frame of the time interval as the representative point, convert the representative point into the corresponding time value, construct the median time point sequence, and arrange them in chronological order to obtain the complete time trend point trajectory and establish the median time point sequence. S413: Based on the median time point sequence, determine the sequential relationship between adjacent time points in order, mark the case where the later time point is earlier than the previous time point as reverse, and mark the case where it is later as positive, mark the continuity of the overall trend direction and the number of direction changes, record the number and distribution characteristics of positive and reverse points, and combine the median time point density information to obtain the abnormal time trend structure index set.
7. The Alzheimer's disease identification method based on speech features according to claim 1, characterized in that, The method further includes: Based on the trend structure points and corresponding frame segment numbers marked in the abnormal time trend structure index set, the continuous jump trend speech frame segments are summarized, uniformly numbered and categorized with labels, and a frame order and time stamp index is established to generate Alzheimer's disease speech feature recognition results. The Alzheimer's disease speech feature recognition results include an abnormal trend frame segment number index, trend label classification results, and time series markers.
8. The Alzheimer's disease identification method based on speech features according to claim 7, characterized in that, The specific steps for obtaining the speech feature recognition results for Alzheimer's disease are as follows: S511: Based on the trend structure points marked in the abnormal time trend structure index set and the corresponding frame segment numbers, extract the start and end numbers of the frame segments covered by each trend segment, convert all frame numbers into the corresponding timestamp sequence in sequence, arrange the trend segments in chronological order, compare the number intervals of two adjacent trend segments, if the number interval value is within the allowable range, it is determined to be the same segment with continuous trend, and merge it into the same jump trend voice frame segment to establish a set of continuous jump trend frame segments; S512: Based on the set of continuous jump trend frames, assign an independent number to each frame segment, and set a label based on the directional characteristics of the trend change in the frame segment. Summarize the number, frame start and end number, time start and end value, and trend label content to form an index list and establish a jump trend frame segment label index table. S513: Based on the trend change direction, number of frames, time span, label type and intra-frame trend switching status attribute information of each frame segment in the jump trend frame segment label index table, according to the set frame segment distribution judgment conditions, the trend stability and trend switching characteristics of each frame segment are screened and layered, and the data of different feature labels are integrated to obtain the Alzheimer's disease speech feature recognition result.
9. An Alzheimer's disease identification system based on speech features, characterized in that, The system is used to implement the Alzheimer's disease identification method based on speech features as described in any one of claims 1-8, comprising: The speech feature extraction module is used to perform S1: acquire a continuous sampling point sequence of the patient's speech signal, perform frame division processing on the signal, extract the mean absolute value of amplitude, the difference between the maximum and minimum amplitude and the energy value in each frame, combine the parameters in sequence and arrange them in sequence according to the frame number to generate a speech frame-level structural parameter sequence; The low-amplitude region identification module is used to perform S2: based on the ratio of the maximum and minimum amplitude difference to the mean absolute amplitude value of each frame in the speech frame-level structural parameter sequence, locate the frame segments in continuous frames whose ratio is less than the predefined dynamic range ratio threshold, and merge adjacent segments in chronological order to generate a list of sparse amplitude fluctuation segments. The module for marking broken segments is used to execute S3: based on the mean absolute value of amplitude in the speech frame-level structural parameter sequence, perform inter-frame difference operation on all frames and record the positions where there are two sets of positive and negative opposite relationships in the continuous difference, mark the start and end positions of abrupt breakage, and generate a set of continuous abrupt breakage segments; The integrated trend anomaly module is used to execute S4: Based on the list of sparse amplitude fluctuation segments and the set of continuous jump break segments, the median time point of each overlapping segment is extracted and a trend structure line is established in chronological order. The direction of the structure line is judged and the number of points is recorded to generate an abnormal time trend structure index set. The collection and recognition result module is used to execute S5: based on the trend structure points marked in the abnormal time trend structure index set and the corresponding frame segment number, it summarizes the continuous jump trend speech frame segments, performs unified numbering and label classification, establishes frame order and time stamp index, and generates Alzheimer's disease speech feature recognition results.
Citation Information
Cited By
Narrowband satellite voice communication noise reduction system and method
CN121617408A
Multi-language adaptive identification method based on AI
CN121708902A