Method for predicting cognitive decline in the elderly with fusion of behavior and voice interaction

CN122551835APending Publication Date: 2026-08-11HAINAN PROVINCIAL GERIATRIC HOSPITAL
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-20
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0003]融合行为与语音交互的老年认知退化预测方法是指针对老年人在日常任务执行和问答交互过程中形成的行为序列与语音序列进行同步建模,并依据两类序列的耦合变化输出认知退化程度预测结果的一种处理方法,方法目的在于解决现有技术中过度依赖医学量表、访谈和单一数据分析导致的评估周期长、连续性不足、早期变化不易识别的问题,从而建立一种可连续执行、可量化输出、可用于早期预警的认知退化预测机制,方法希望达成在无需频繁线下检查条件下,对老年对象的认知状态进行周期性和连续性评估的效果,通过联合分析行为异常和语音异常,提高对轻度退化阶段和变化趋势阶段的识别能力,输出数值化退化评分、分级结果和趋势结果,为后续干预提供可直接使用的判定依据,进而实现认知退化风险的提前发现和管理

Benefits of technology

[0034] In this invention, the sequence cost value is calculated by calling the dynamic time warping algorithm. If the cost value exceeds the tolerance limit, a logarithmic penalty term is superimposed to fuse the correlation path, and a cross-modal asynchronous feature vector is established. This solves the problem of nonlinear temporal misalignment between behavioral pauses and speech delays, and improves the accuracy of cross-modal behavioral performance evolution correlation analysis and temporal alignment matching degree.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122551835A_ABST
    Figure CN122551835A_ABST
Patent Text Reader

Abstract

This invention relates to the field of behavior and speech recognition technology, specifically to a method for predicting cognitive decline in the elderly by integrating behavior and speech interaction. In this invention, a dynamic time warping algorithm is invoked to calculate the sequence cost value. If the cost value exceeds the tolerance limit, a logarithmic penalty term is superimposed to fuse the associated path, solving the problem of nonlinear temporal misalignment between behavioral pauses and speech delays. This improves the accuracy of cross-modal behavioral evolution correlation analysis and temporal alignment matching. A graph neural network algorithm is invoked to process cross-modal asynchronous feature vectors and behavioral intention topological connection tensors, deeply quantifying the spatial structural features of random wandering behavior and analyzing the deep interaction patterns between tactile tremors and thought blanks. By separating the expected time term of disordered residence within the logical judgment coefficient of cognitive state, a continuous evolution evaluation and grading result is output, providing a directly applicable numerical basis for degeneration judgment and avoiding the limitations of conventional mechanisms that overly rely on subjective interviews.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of behavior and speech recognition technology, and in particular to a method for predicting cognitive decline in the elderly by integrating behavior and speech interaction. Background Technology

[0002] Behavioral and speech recognition technologies primarily model the action sequences, operation order, pause duration, task completion paths, error trigger counts, and response delays exhibited by target objects within a continuous time range. They also identify behavioral change patterns through temporal analysis and analyze the speech signals generated during interaction, processing information such as speech rate, pronunciation intervals, pause distribution, pitch fluctuations, sentence completeness, repetition rate, and response delay. This technological field encompasses not only the recognition of single behavioral and speech information but also the joint modeling of the correlation between behavioral and speech information, thereby obtaining more stable state assessment results than single-modal analysis. In elderly cognitive assessment scenarios, cognitive decline is typically reflected simultaneously in two dimensions: task execution behavior and language expression ability. This includes issues such as disordered operation order, prolonged reaction time, increased sentence interruptions, increased repetitive expression, and decreased question-and-answer matching. The cross-disciplinary technology of behavioral and speech recognition has become an important implementation path for cognitive state analysis, abnormal trend judgment, and decline degree prediction.

[0003] The method for predicting cognitive decline in the elderly by integrating behavior and speech interaction refers to a processing method that synchronously models behavioral sequences and speech sequences formed by the elderly during daily task performance and question-and-answer interactions, and outputs a prediction result of the degree of cognitive decline based on the coupling changes of the two types of sequences. The method aims to solve the problems of long assessment cycles, insufficient continuity, and difficulty in identifying early changes caused by the over-reliance on medical scales, interviews, and single data analysis in existing technologies. It aims to establish a cognitive decline prediction mechanism that can be continuously executed, quantified, and used for early warning. The method hopes to achieve the effect of periodic and continuous assessment of the cognitive status of elderly subjects without the need for frequent offline examinations. By jointly analyzing behavioral and speech abnormalities, it can improve the ability to identify mild decline stages and trend stages, and output numerical decline scores, classification results, and trend results, providing a directly usable basis for subsequent intervention, thereby realizing the early detection and management of cognitive decline risks.

[0004] Existing technologies typically focus on separating and modeling single action sequences and single speech signals of the target object. They rely on linear threshold superposition to perform joint state determination. This conventional operating mechanism heavily depends on regular follow-up using medical scales and triggering with isolated single-modal data, severing the asynchronous coupling and evolutionary connections between cross-modal behaviors. In practice, isolated behavior monitoring often categorizes brief pauses in object searching as normal physiological slowness, and independent speech analysis also classifies slight slowing of speech rate as a reasonable phenomenon. This leads to multiple deep-seated abnormalities being filtered out by conventional error-tolerant mechanisms. The isolated data operation mode directly causes a severe discontinuity in the assessment cycle, making it difficult to capture continuous evolutionary trends. This, in turn, leads to the risk of missing early and subtle cognitive decline, significantly reducing the sensitivity of the early warning mechanism. Existing conventional separation modeling mechanisms are easily masked by the state of a single component not exceeding the limit, resulting in output judgment scores that remain within the healthy range. This delays early intervention and management, misses the early intervention window for mitigating cognitive decline, and increases the subsequent medical burden on families and society, as well as the risk of rapid deterioration of the condition. Summary of the Invention

[0005] The purpose of this invention is to address the shortcomings of existing technologies by proposing a method for predicting cognitive decline in the elderly by integrating behavioral and voice interaction.

[0006] To achieve the above objectives, the present invention adopts the following technical solution: a method for predicting cognitive decline in the elderly by integrating behavior and voice interaction, comprising the following steps:

[0007] Step 1: Extract the operation offset line spacing and determine the duration of the response vowel sounding by using the touch coordinate sequence and audio sampling signal. Calculate the absolute difference between the line spacing and the duration, compare the product values ​​within the set tolerance filtering limit, and generate the action pronunciation timing mapping matrix.

[0008] Step 2: Based on the action pronunciation timing mapping matrix, extract the object-finding pause step length and extract the extreme value of the word-ending trill, call dynamic time warping to calculate the cost value, determine whether the cost value exceeds the tolerance superimposed logarithmic penalty term fusion path, and establish a cross-modal asynchronous feature vector;

[0009] Step 3: Based on the cross-modal asynchronous feature vector, extract the wandering coordinate sequence and separate the fault nodes, calculate the straight-line distance between the coordinates and the nodes, compare the critical value to retain short-distance point pairs, take the inverse distance to allocate weights and aggregate the network matrix to obtain the behavior intention topology connection tensor;

[0010] Step 4: Based on the cross-modal asynchronous feature vector and the behavioral intention topology connection tensor, the graph neural network assigns node weights, calculates the weight blank product to construct the gain, superimposes the tremor attenuation operation to summarize the feature mean, and outputs the cognitive state logic judgment coefficient.

[0011] Step 5: Based on the cognitive state logic judgment coefficient, separate the expected time of disordered residence, extract the absolute value and calculate the reciprocal, multiply the reciprocal by the point and the judgment coefficient, summarize the sum of multiple values ​​and calculate the arithmetic mean to obtain the quantitative decay evolution evaluation constant.

[0012] As a further aspect of the present invention, the action pronunciation timing mapping matrix includes operation offset row spacing, response vowel pronunciation extension time scale, and multi-terminal product values; the cross-modal asynchronous feature vector includes object-finding pause step size, word-ending trill extreme value, and logarithmic penalty term fusion path; the behavioral intention topological connection tensor includes wandering coordinate sequence fault nodes, short-distance point pairs, and distance inverse weighted network matrix; the cognitive state logical judgment coefficient includes graph neural network node weights, weight blank product to construct gain, and summed feature mean; and the quantitative decay evolution evaluation constant includes the absolute value inverse of the disordered dwell time term, the result of the decision coefficient by point product, and the arithmetic mean of the sum of multiple values.

[0013] As a further aspect of the present invention, the specific steps for generating the action pronunciation timing mapping matrix are as follows:

[0014] Based on the touch coordinate sequence and audio sampling signal, the offset line spacing is extracted by segmenting the coordinate position, the extended time scale is extracted by truncating the audio waveform, the line spacing and time scale are spliced ​​to construct a combined node, the absolute difference of the data inside the node is calculated, the absolute difference elements are integrated to establish a line spacing and time scale difference sequence.

[0015] Based on the line spacing and time stamp difference sequence, a comparison threshold is defined, the size of the difference within the sequence is determined, elements exceeding the threshold are removed, the line spacing and time stamp of the remaining elements are extracted, the product of line spacing and time stamp is calculated, the product of line spacing and time stamp is aggregated, and an action pronunciation timing mapping matrix is ​​generated.

[0016] As a further aspect of the present invention, the specific steps for establishing the cross-modal asynchronous feature vector are as follows:

[0017] Based on the action pronunciation timing mapping matrix, the row and column data are split to extract the object-finding pause step length, high-frequency noise is filtered out to extract the word-ending trill extreme value, extreme value time nodes are matched, the object-finding pause step length and trill extreme value are merged to construct a combination item, and the extreme value step length combination item is aggregated to generate a modal extreme value step length set.

[0018] Based on the set of modal extreme step lengths, the internal pause step length and trill extreme value are analyzed, dynamic time warping is called to calculate the absolute term of the step length and extreme value deviation, weight multipliers are added to calculate the replacement value, the replacement value is summarized to form a continuous sequence, the sequence data is arranged with reference to the time sequence, and a path replacement value array is generated.

[0019] Based on the path cost array, the tolerance limit of the cost is defined, the specific numerical values ​​of the array elements are compared, the elements exceeding the limit are extracted and superimposed with logarithmic penalty multipliers, the penalty values ​​are fused with the initial path, multiple node data are merged, and a cross-modal asynchronous feature vector is established.

[0020] As a further aspect of the present invention, the dynamic time warping involves constructing a reference sequence by pausing step length, extracting trill extrema extremes to construct a comparison test sequence, establishing a two-dimensional cost grid by crossing the reference sequence and the comparison test sequence, locating the row and column intersection coordinate nodes within the two-dimensional cost grid, reading the pausing step length value and the corresponding trill extrema extreme value corresponding to the coordinate node, performing a subtraction operation between the corresponding pausing step length value and the corresponding trill extrema extreme value to obtain the node difference, extracting the absolute value of the node difference, and generating an absolute term for the step length and extreme value deviation.

[0021] As a further aspect of the present invention, the fusion penalty value and the initial path are: extracting the prior estimation matrix containing multiple time nodes within the initial path; arranging multiple over-limit elements in time sequence to construct a penalty value diagonal tensor corresponding to logarithmic penalty multipliers; calling the Kalman filter algorithm to perform the Hadamard product operation between the penalty value diagonal tensor and the multiple time node prior estimation matrices; calculating the posterior state covariance matrix with nonlinear penalty weights; extracting the main diagonal elements of the posterior state covariance matrix to cover the multiple original node scalar values ​​within the initial path; and then establishing the numerically updated associated fusion path tensor.

[0022] As a further aspect of the present invention, the specific steps for obtaining the behavioral intent topology connection tensor are as follows:

[0023] Based on the cross-modal asynchronous feature vector, axial values ​​are extracted to construct a wandering coordinate sequence, the interactive fault nodes are separated by the interruption time point, the difference between the coordinate sequence and the fault node position is calculated, the sum of the squared differences is accumulated to obtain the straight-line distance, and a set of wandering fault spacing is generated.

[0024] Based on the set of wandering fault spacings, a critical limit for distance comparison is defined, the size of the internal straight-line distance is determined, values ​​exceeding the critical limit are eliminated, the inverse of the straight-line distance within the limit is extracted, and the inverse of the distance is assigned to construct connection weights, thereby obtaining the behavioral intention topology connection tensor.

[0025] As a further aspect of the present invention, the specific steps for outputting the cognitive state logical judgment coefficient are as follows:

[0026] Based on the cross-modal asynchronous feature vector and the behavioral intention topology connection tensor, a graph neural network is invoked to extract vector interaction frequency values ​​and locate the tensor topology node coordinates. Mapping coefficients are assigned according to the frequency values, and the coefficients are aggregated to construct node weights, generating a frequency weight node distribution tensor.

[0027] Based on the frequency weight node distribution tensor, the word blank pause duration is stripped, the corresponding weight of the node is extracted, the product of duration and weight is calculated, the product is compared with the gain limit and the product value that does not exceed the limit is retained, the product values ​​within the limit are summarized, and a weight blank gain sequence is established.

[0028] Based on the weighted blank gain sequence, the touch tremor attenuation multiplier is extracted, the attenuation multiplier is calculated and the sequence element is multiplied bitwise, the bitwise product values ​​are accumulated to obtain the sum of the set, the sum is divided by the total number of elements to extract the truncated mean, and the cognitive state logic judgment coefficient is output.

[0029] As a further aspect of the present invention, the graph neural network loads a cross-modal asynchronous feature vector and a behavioral intention topology connection tensor, parses the internal mesh connection matrix structure of the behavioral intention topology connection tensor, extracts the intersecting nodes of rows and columns within the mesh connection matrix to establish a spatial index, locates the coordinates of the tensor topology nodes based on the spatial index, maps the corresponding data components within the cross-modal asynchronous feature vector according to the coordinates of the tensor topology nodes, reads the state record identifiers within the data components, parses the state record identifiers to obtain the occurrence parameters, and extracts the vector interaction frequency value.

[0030] As a further aspect of the present invention, the specific steps for obtaining the quantitative decay evolution evaluation constant are as follows:

[0031] Based on the cross-modal asynchronous feature vector and the behavioral intention topology connection tensor, the vector interaction frequency is extracted, the topology node coordinates are located and the frequency mapping node weights are matched, the blank pause duration is removed, the duration and weight are fused, and a weight blank gain sequence is established.

[0032] Based on the weighted blank gain sequence, the touch tremor attenuation factor is extracted and the amplitude of the attenuation factor modulation sequence element is introduced. The modulation elements are integrated to construct a global feature set. The global feature set is compressed to extract the truncated representation mean and output the cognitive state logic judgment coefficient.

[0033] Compared with the prior art, the advantages and positive effects of the present invention are as follows:

[0034] In this invention, the sequence cost value is calculated by calling the dynamic time warping algorithm. If the cost value exceeds the tolerance limit, a logarithmic penalty term is superimposed to fuse the correlation path, and a cross-modal asynchronous feature vector is established. This solves the problem of nonlinear temporal misalignment between behavioral pauses and speech delays, and improves the accuracy of cross-modal behavioral performance evolution correlation analysis and temporal alignment matching degree.

[0035] In this invention, a graph neural network algorithm is used to process cross-modal asynchronous feature vectors and behavioral intention topology connection tensors. Node weights are allocated according to the network topology, and multiple numerical averages are summed by superimposing tremor attenuation operations to output the cognitive state logical judgment coefficient. This deeply quantifies the spatial structural features of random wandering behavior and analyzes the deep interaction patterns between touch tremor and mental blankness.

[0036] In this invention, by separating the expected time term of disordered residence within the logical judgment coefficient of cognitive state, stripping the absolute value and performing the reciprocal calculation, a quantitative decay evolution evaluation constant is obtained, and a continuous evolution evaluation grading result is output, providing a numerical basis for directly applying the determination of degradation, thus avoiding the limitations of conventional mechanisms that rely too heavily on subjective interviews. Attached Figure Description

[0037] Figure 1 This is a schematic diagram of the main steps of the present invention. Detailed Implementation

[0038] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0039] Example 1

[0040] Please see Figure 1 This invention provides a technical solution: a method for predicting cognitive decline in the elderly by integrating behavior and voice interaction, comprising the following steps:

[0041] Step 1: Extract the operation offset line spacing and determine the duration of the response vowel sounding by using the touch coordinate sequence and audio sampling signal. Calculate the absolute difference between the line spacing and the duration, compare the product values ​​within the set tolerance filtering limit, and generate the action pronunciation timing mapping matrix.

[0042] Step 2: Based on the action-pronunciation temporal mapping matrix, extract the object-finding pause step length and extract the extreme value of the word-ending trill. Call dynamic time warping to calculate the cost value, determine the path fusion by exceeding the tolerance superimposed logarithmic penalty term, and establish a cross-modal asynchronous feature vector.

[0043] Step 3: Based on cross-modal asynchronous feature vectors, extract the wandering coordinate sequence and separate fault nodes, calculate the straight-line distance between coordinates and nodes, compare the critical value to retain short-distance point pairs, take the inverse distance to allocate weights and aggregate the network matrix to obtain the behavior intention topology connection tensor;

[0044] Step 4: Based on the cross-modal asynchronous feature vector and the behavioral intention topology connection tensor, the graph neural network assigns node weights, calculates the weight blank product to construct the gain, superimposes the tremor attenuation operation to summarize the feature mean, and outputs the cognitive state logical judgment coefficient.

[0045] Step 5: Based on the cognitive state logic judgment coefficient, separate the expected time of disordered residence, extract the absolute value and calculate the reciprocal, multiply the reciprocal by the point and the judgment coefficient, summarize the sum of multiple values ​​and calculate the arithmetic mean to obtain the quantitative decay evolution evaluation constant.

[0046] The action-pronunciation temporal mapping matrix includes operation offset row spacing, response vowel prolongation time scale, and multi-term product values ​​within the limit. The cross-modal asynchronous feature vector includes the object-finding pause step size, word-ending trill extrema, and logarithmic penalty term fusion path. The behavioral intention topological connection tensor includes wandering coordinate sequence fault nodes, short-distance point pairs, and distance inverse weighted network matrix. The cognitive state logical judgment coefficient includes graph neural network node weights, weight blank product to construct gain, and summed feature mean. The quantitative decay evolution evaluation constant includes the absolute value inverse of the expected dwell time term, the result of the judgment coefficient by point product, and the arithmetic mean of the sum of multiple values.

[0047] The specific steps for generating the action-pronunciation timing mapping matrix are as follows:

[0048] Based on the touch coordinate sequence and audio sampling signal, the offset line spacing is extracted by segmenting the coordinate position, the extended time scale is extracted by truncating the audio waveform, the line spacing and time scale are spliced ​​to construct a combined node, the absolute difference of the data inside the node is calculated, the absolute difference elements are integrated to establish a line spacing and time scale difference sequence.

[0049] Based on the line spacing and time stamp difference sequence, a comparison boundary threshold is defined, the size of the difference within the sequence is determined, elements exceeding the threshold are removed, the line spacing and time stamp of the remaining elements are extracted, the product of line spacing and time stamp is calculated, the product of line spacing and time stamp is aggregated, and an action pronunciation timing mapping matrix is ​​generated.

[0050] Based on the touch coordinate sequence and audio sampling signal, an adaptive Hamming window dual-mode alignment algorithm is used to segment the coordinate position, extract the offset line spacing, and truncate the audio waveform to extract extended time stamps. The audio sampling reference frequency is set to 44100 Hz, and a sliding Hamming window with a window length of 256 milliseconds and a superposition ratio of 50% is invoked. The start and end points of the audio waveform with a short-time zero crossing rate greater than 15 times are extracted from the sliding Hamming window, and the waveform is truncated to extract extended time stamps. The corresponding timestamps in the touch coordinate sequence are read, containing X-axis and Y-axis coordinate values. The corresponding timestamp's X-axis and Y-axis coordinate values ​​are then subtracted from the coordinate values ​​of the previous timestamp. The offset row spacing is extracted by bitwise subtraction. The offset row spacing and extended time scale are concatenated using vector concatenation instructions to construct a combined node. The concatenation dimension parameter is configured as the axial value of 1 to generate a floating-point array structure containing 3-dimensional elements. The L1 norm Manhattan distance statistical algorithm is used to calculate the absolute difference of the data inside the combined node. The first dimension offset row spacing value inside the array is retrieved and the second dimension extended time scale value is subtracted. The underlying absolute value conversion instruction is called to obtain the scalar absolute value and store it in the cache register. All the absolute difference elements recorded in the cache register are integrated to construct a one-dimensional floating-point array containing a continuously increasing time index to establish a row spacing and time scale difference sequence.

[0051] Based on the line-space timescale difference sequence, a comparison threshold is defined using a Gaussian mixture model-based three-sigma control limit setting mechanism. Historical difference samples from 1000 previous benchmark interaction records are extracted and input into the initial Gaussian distribution equation. The mean and standard deviation of the historical difference samples are calculated, and the mean plus three times the standard deviation is taken as the comparison threshold. A one-sided interval filtering algorithm is used to determine the magnitude of the differences within the line-space timescale difference sequence. Each element containing a difference within the sequence is read, and its value is input to the arithmetic logic unit along with the comparison threshold to execute floating-point scalar comparison instructions. Values ​​greater than the comparison threshold are discarded. For data exceeding the threshold, if the difference exceeds the threshold element, extract the remaining elements within the threshold that were not removed. These elements contain the original associated offset row distance and extended time scale. Calculate the product of the offset row distance and extended time scale using the matrix Hadamard dot product algorithm. The extracted offset row distance is used to form the first column vector, and the extracted extended time scale is used to form the second column vector. Perform the multiplication operation between the corresponding elements of the first and second column vectors to generate product terms. Call the tensor concatenation operation instruction to aggregate all product terms in chronological order. Allocate contiguous memory space to load the aggregated data block to generate the action-pronunciation time sequence mapping matrix.

[0052] The specific steps for establishing cross-modal asynchronous feature vectors are as follows:

[0053] Based on the action pronunciation time sequence mapping matrix, the row and column data are split to extract the object-finding pause step length, high-frequency noise is filtered out to extract the word-ending trill extreme value, extreme value time nodes are matched, the object-finding pause step length and trill extreme value are merged to construct a combination item, and the extreme value step length combination item is aggregated to generate a modal extreme value step length set.

[0054] Based on the modal extreme value step size set, the internal pause step size and trill extreme value are analyzed, dynamic time warping is called to calculate the absolute term of the step size and extreme value deviation, weight multiplier is added to calculate the replacement value, the replacement value is summarized to form a continuous sequence, the sequence data is arranged with reference to the time series, and the path replacement value array is generated.

[0055] Based on the path cost array, the tolerance limit of the cost is defined, the specific numerical values ​​of the array elements are compared, the elements exceeding the limit are extracted and superimposed with logarithmic penalty multipliers, the penalty values ​​are merged with the initial path, multiple node data are combined, and a cross-modal asynchronous feature vector is established.

[0056] Based on the action-pronunciation time-series mapping matrix, a tensor dimensionality reduction slicing algorithm is used to split row and column data to extract the target object pause step length. The slicing dimension parameters are configured such that index 0 in the first dimension corresponds to the time series, and index 1 in the second dimension corresponds to the step length sequence. A stripping operation along the second dimension is performed to read the non-zero continuous dwell time difference and calculate the target object pause step length. A Butterworth low-pass filter algorithm is used to remove high-frequency noise and extract word-ending trills. The low-pass filter cutoff frequency is set to 3000 Hz, and the filter order is configured to 4. Audio oscillation waveforms exceeding the cutoff frequency are filtered out. A local maximum search instruction is called to traverse the smoothed audio waveform after filtering and read the vibration... The highest amplitude value is taken as the extreme value of the word-ending trill. The extreme value time node is matched by a sliding time window cross-correlation algorithm. The sliding window length is preset to 50 milliseconds and the step size is configured to 10 milliseconds by reading the system initialization parameter table. The horizontal translation cross-correlation measurement is performed on the time axis. The timestamps corresponding to the cross-correlation coefficients are locked for matching. The multidimensional array splicing and stacking mechanism is used to merge the object-finding pause step size and the extreme value of the word-ending trill to construct a combination item. Double-precision floating-point storage space is allocated to generate a tuple structure containing dual feature elements. All extreme value step size combination items are aggregated in the order of timestamps and loaded into a contiguous memory block to generate a modal extreme value step size set.

[0057] Based on the modal extremum step size set, the internal object-finding pause step size and word-ending trill extrema are parsed using low-level data unpacking instructions. The dual-feature element tuple is decomposed into independent one-dimensional floating-point vectors. A dynamic time warping algorithm is called to calculate the absolute term of the step size and extremum deviation. The algorithm's global path constraint parameters are configured as equations with width constraints, and the constraint window width is set to 15 sampling frames. A two-dimensional cumulative cost matrix is ​​constructed to locate intersection coordinate points. The corresponding object-finding pause step size and word-ending trill extrema values ​​are read from the coordinate points and input into the arithmetic logic unit for subtraction. An absolute value conversion instruction is then called. Let the output scalar value be used as the absolute term of the deviation. Use the nonlinear activation mapping algorithm to add weight multipliers to calculate the substitution value. By extracting the constants of the system kernel mathematical library, configure the base of the activation equation to input the absolute term of the deviation. Set the temperature coefficient parameter to 0.5 to calculate and generate the corresponding weight multiplier. Perform floating-point scalar multiplication of the absolute term of the deviation and the weight multiplier to calculate the discrete substitution value. Retrieve memory and append write instructions to summarize all discrete substitution values ​​to form a continuous sequence. Sort the sequence data in ascending order with reference to the system timestamp. Allocate one-dimensional continuous addressing memory space to generate a path substitution value array.

[0058] Based on the path cost value array, an adaptive sliding quantile evaluation algorithm is used to define the cost value tolerance limit. The historical time observation window length is set to 500 discrete samples by loading system environment configuration items. All cost value samples within the historical time observation window are read, sorted in ascending order, and the value corresponding to the 95th percentile is extracted as the cost value tolerance limit. A scalar comparison filtering instruction is called to compare the specific values ​​of array elements. Each element value within the path cost value array is read and compared with the cost value tolerance limit, inputting the values ​​into a hardware comparator. If an element value is greater than the limit threshold, a truncation instruction is triggered, retaining out-of-bounds data as exceeding the limit. An interior-point logarithmic barrier function optimization algorithm is used to superimpose a logarithmic penalty multiplier on the exceeding elements, and the result is written into the initial register. The obstacle control parameter is initially set to 10, and a scalar decay operation with a step size of 0.1 is performed according to the calculation cycle. After decay, the parameter is multiplied with the over-limit element to generate a logarithmic penalty multiplier. The Kalman filter posterior state update algorithm is used to fuse the penalty value and the initial path. The prior estimated covariance matrix of the time nodes is extracted from the initial path. The logarithmic penalty multiplier is constructed into a penalty value diagonal matrix according to the time sequence. The penalty value diagonal matrix and the prior estimated covariance matrix are multiplied by Hadamard. The main diagonal elements of the covariance matrix after the operation are extracted to cover the original scalar of the initial path. Multi-node data are merged, and the updated multi-node data elements are written into the same tensor dimension in a column-wise append mode to establish a one-dimensional unaligned data sequence morphology cross-modal asynchronous feature vector.

[0059] Dynamic time warping is used to construct a reference sequence by pausing step length, extract trill extrema extremes to construct a comparison test sequence, establish a two-dimensional cost grid by crossing the reference sequence and the comparison test sequence, locate the row and column intersection coordinate nodes inside the two-dimensional cost grid, read the pausing step length value and the corresponding trill extrema extreme value corresponding to the coordinate node, perform the subtraction operation between the corresponding pausing step length value and the corresponding trill extrema extreme value to obtain the node difference, extract the absolute value of the node difference, and generate the absolute term of step length and extreme value deviation;

[0060] A one-dimensional floating-point array is continuously allocated and loaded using instructions to read the underlying raw pause step size. A 1024-byte contiguous address block is allocated from system memory, and the pause step size floating-point values ​​are sequentially pushed onto the memory stack in ascending order of timestamps to construct a reference baseline sequence. A data stream-limited queue arrangement algorithm is used to extract trill extreme values. By reading the system's factory-preset global environment variable configuration library and setting the maximum queue depth limit parameter to 256 items, extreme value data is pushed in sequentially, triggering a one-way auto-increment operation of the memory pointer to construct a comparison test sequence. Based on the reference baseline sequence and the comparison test sequence, a tensor orthogonal Cartesian algorithm is used. The Ernst product expansion algorithm intersects two sets of sequences and establishes a two-dimensional cost grid. It obtains the baseline sequence length constant from the system's underlying configuration file to define the grid's row dimension boundary and the test sequence length constant to define the grid's column dimension boundary. It calls a memory zeroing initialization instruction to generate a double-precision floating-point matrix with elements initially set to zero. The matrix's horizontal axis coordinate mapping is configured to reference the baseline sequence's integer index, and the matrix's vertical axis coordinate mapping is configured to compare with the test sequence's integer index to establish a complete two-dimensional cost grid. A two-pointer forward traversal addressing strategy is used to locate the row and column intersection coordinate nodes within the two-dimensional cost grid, and the row dimension pointer is initially biased. With the initial bias bit of the column dimension pointer set to 0 and the increment step size of the bidirectional pointer set to 1, a nested loop traverses the grid space matrix. The current pointer's two-dimensional cursor intersects at the physical address to locate the intersection coordinate node. A direct memory access instruction (DMI) is called to read the corresponding pause step value and the corresponding trill extreme value of the intersection coordinate node. The double-precision value temporarily stored at the underlying address is directly transferred from the static storage medium to the CPU's L1 cache register. An arithmetic logic unit (ALU) double-precision floating-point subtraction microprogram performs the subtraction operation between the corresponding pause step value and the corresponding trill extreme value. This is achieved through configuration... The minuend hardware register is loaded with the pause step value, the subtrahend hardware register is configured to load the trill extreme value, the underlying subtraction machine cycle trigger pulse is sent to obtain the calculated output node difference, the node difference is extracted by performing a bitwise mask clearing operation on the underlying sign bit, by loading the fixed Boolean mask constant of the whole system to perform a bitwise AND logic operation on the highest sign bit of the double-precision floating-point number, forcibly resetting the sign bit to zero to truncate the negative bit representing the data and retain the positive scalar data, and aggregating all the positive scalar data after the mask clearing operation is performed and storing them in the newly created tensor space in order to generate the step size and extreme value deviation absolute terms.

[0061] By fusing the penalty value with the initial path, the prior estimation matrix containing multiple time nodes in the initial path is extracted. The multiple over-limit elements are arranged in time order to construct a penalty value diagonal tensor corresponding to the logarithmic penalty multipliers. The Kalman filter algorithm is called to perform the Hadamard product operation between the penalty value diagonal tensor and the multiple time node prior estimation matrices. The posterior state covariance matrix with nonlinear penalty weight is calculated. The main diagonal elements of the posterior state covariance matrix are extracted to cover the scalar values ​​of multiple original nodes in the initial path, and then the numerically updated associated fused path tensor is established.

[0062] Based on the initial path and penalty value, a memory slicing read instruction is used to extract the prior estimation matrix containing multiple time nodes within the initial path. A preset sliding read step size constant parameter is obtained by reading the system hardware configuration file. This preset sliding read step size constant parameter is established by loading the initial programming environment parameter table of the motherboard's read-only memory. Based on the step size constant parameter of 16 bytes, data blocks are peeled layer by layer in an address-incrementing pattern to obtain the prior estimation matrix. A tensor diagonalization reconstruction algorithm is used to construct a penalty value diagonal tensor by arranging multiple over-limit elements in time sequence and corresponding logarithmic penalty multipliers. A 64-megabyte continuous video memory space is requested from the central processing unit to allocate a blank square. The matrix is ​​set to a 256x256 dimension. The system clock crystal is used to obtain nanosecond-level timestamps and ascending-order logarithmic penalty multipliers. A single-step offset write instruction is used to fill the ascending-order logarithmic penalty multipliers into the main diagonal physical address range of the blank matrix. The off-diagonal space is filled with double-precision floating-point zeros to generate a penalty numerical diagonal tensor. A Kalman filter posterior state update algorithm is used to estimate the penalty numerical diagonal tensor and multiple time-node prior matrices. The process noise covariance matrix constant is set to a scalar constant of 0.01 by writing to the Kalman gain control register. The measurement noise covariance matrix constant is also configured. The quantity is a scalar constant of 0.05. The underlying microinstructions of the tensor parallel computing framework are invoked to activate multiple arithmetic logic unit arrays. A parallel bitwise multiplication machine cycle control signal is issued, forcing the elements within the penalized diagonal tensor to perform floating-point multiplication with the elements of the prior estimation matrix at the same coordinate dimension to complete the Hadamard product operation. The state transition covariance iteration microprogram is invoked to calculate the posterior state covariance matrix with nonlinear penalty weights. The state transition matrix is ​​temporarily stored in the system state register of the previous machine cycle. The state transition matrix is ​​multiplied by the bitwise multiplication output intermediate result matrix. An additional microinstruction for subtracting the identity matrix is ​​added to output the posterior state covariance matrix. The variance matrix is ​​extracted using a diagonal element stripping and extraction algorithm. The diagonal elements of the posterior state covariance matrix cover the scalar values ​​of multiple original nodes within the initial path. By configuring the row and column addresses to synchronously increment by one floating-point word, the row and column index pointers are synchronously increased to traverse the matrix space and read the floating-point values ​​of the diagonal. The direct memory access overwrite instruction is called to directly write the read diagonal floating-point values ​​into the physical memory address block corresponding to the initial path, performing a forced replacement operation without residue. After the data is extracted and overwritten, the tensor dimension is continuously encapsulated with micro-instructions and loaded into the cache stack space to generate the associated fused path tensor after the value is updated.

[0063] The specific steps to obtain the behavioral intent topology connection tensor are as follows:

[0064] Based on cross-modal asynchronous feature vectors, axial values ​​are extracted to construct a wandering coordinate sequence. Interactive fault nodes are separated by removing the interruption time point. The difference between the coordinate sequence and the fault node position is calculated. The straight-line distance is obtained by accumulating the square of the difference and generating a set of wandering fault spacing.

[0065] Based on the set of distance intervals of wandering faults, the critical limit of distance comparison is defined, the size of the internal straight-line distance is determined, values ​​exceeding the critical limit are eliminated, the inverse of the straight-line distance within the limit is extracted, the inverse of the distance is assigned to construct the connection weight, and the behavioral intention topology connection tensor is obtained.

[0066] Based on cross-modal asynchronous feature vectors, tensor dimension mapping is used to extract axial values ​​from microinstructions. Address offset configuration parameters are sent to the motherboard memory controller to set the read offset to 12 bytes, locking the specific physical address range of the spatial coordinates within the vector memory block. A 32-bit single-precision floating-point bus is invoked to sequentially read the X-axis and Y-axis coordinate components, constructing a continuous coordinate point data array and building a wandering coordinate sequence. An isolated forest-based anomaly detection timing stripping algorithm is used to strip interrupt points. During system initialization, the preset tree structure quantity parameter is set to 100 trees, and the anomaly score judgment threshold is configured as a constant 0.75. The preset tree structure quantity parameter is written into the kernel global configuration table by the system bootloader. The timestamp data array within the wandering coordinate sequence is input into the isolated forest judgment model to calculate the isolation score at each time point, comparing the isolation score with the anomaly score. The system identifies and removes isolated timestamps with non-touch interaction events where the dwell time exceeds 2 seconds, truncates continuous trajectories, and separates interactive fault nodes. It uses the Euclidean distance squared sum to calculate the difference between the coordinate sequence and the fault node position using the underlying machine program. The CPU's internal minuend hardware register loads the current coordinate scalar within the wandering coordinate sequence, and the subtraction hardware register loads the corresponding coordinate scalar of the interactive fault node. The arithmetic logic unit activates, sending subtraction machine cycle control pulses to obtain the output deviation value and store it in the cache unit. For multiple output deviation values, the hardware multiplier array is called to perform self-multiplication to calculate the square term. The difference squared is accumulated through the memory accumulation register. The system's underlying square root operation microcode module is called to obtain positive floating-point values ​​to calculate the straight-line distance. All calculated straight-line distances are aggregated and allocated in chronological order to generate a set of wandering fault spacings.

[0067] Based on the set of wandering fault spacings, a dynamic thresholding algorithm based on local outlier factors is used to define the critical limits for distance comparison. Preset neighborhood space parameters are loaded by reading the system's factory-installed read-only memory. These parameters are directly written to a specified sector of the read-only memory during the device's factory programming stage, configuring values ​​for 20 discrete sample points. The local reachability density of each sample point within the wandering fault spacing set is calculated, and the corresponding outlier score is output. An outlier score greater than 1.5 corresponds to a physical distance scalar constant of 500 pixels as the comparison benchmark value, defining the critical limits for distance comparison. A single-instruction multiple-data-stream hardware comparator is used to truncate instructions to determine the magnitude of internal straight-line distances. The constant 500 pixels is written to the global comparison reference register through a scheduling vector processing unit. Multiple straight-line distances within the wandering fault spacing set are concurrently pushed into the hardware comparator array for batch floating-point scalar value comparison. For the operation, a hardware interrupt is triggered to eliminate discrete data exceeding the reference register's internal value and values ​​exceeding the critical limit. Data within the reserved data channel without a hardware interrupt is extracted. The reciprocal of the distance is calculated using a floating-point reciprocal hardware divider microinstruction. The divisor register of the floating-point division unit is constantly filled with a double-precision constant of 1.0. The divisor register is sequentially filled with the linear distance within the limit to trigger the hardware division operation. The machine cycle outputs the quotient set. The reciprocal of the distance is allocated, and a sparse matrix compression storage column-major order allocation algorithm is called to construct the connection weights. A 128x128 dimension blank square matrix space is requested from system memory. Based on the spatial topology mapping index corresponding to the interactive fault node, the reciprocal of the distance is inserted as a non-zero element into the specific row and column intersection physical address of the blank square matrix, and the column pointer array and row index array are updated synchronously. The assignment and conversion of all non-zero elements of the matrix are completed, and the aggregated network matrix structure is obtained to obtain the behavioral intention topology connection tensor.

[0068] The specific steps for outputting the logical judgment coefficients of the cognitive state are as follows:

[0069] Based on cross-modal asynchronous feature vectors and behavioral intention topology connection tensors, a graph neural network is called to extract vector interaction frequency values ​​and locate the coordinates of tensor topology nodes. Mapping coefficients are assigned according to the frequency values, and the coefficients are aggregated to construct node weights, generating a frequency weight node distribution tensor.

[0070] Based on the frequency weight node distribution tensor, the word blank pause duration is removed, the corresponding weight of the node is extracted, the product of duration and weight is calculated, the product is compared with the gain limit and the product value that does not exceed the limit is retained, the product values ​​within the limit are summarized, and a weight blank gain sequence is established.

[0071] Based on the weighted blank gain sequence, the touch tremor attenuation multiplier is extracted, the attenuation multiplier is calculated and the sequence element is multiplied bitwise, the bitwise product values ​​are accumulated to obtain the sum of the set, the sum is divided by the total number of elements to extract the truncated mean, and the cognitive state logic judgment coefficient is output.

[0072] Based on cross-modal asynchronous feature vectors and behavioral intent topology connection tensors, a graph attention network underlying tensor propagation algorithm is used to extract vector interaction frequency values. The behavioral intent topology connection tensor is loaded into the GPU shared memory by calling a parallel memory loading control instruction. The hidden layer neuron dimension parameters are configured as a 64-dimensional floating-point array. Microcode extraction is performed by calling memory pointer address offset to scan row-by-row along the first dimension of the cross-modal asynchronous feature vectors. Specific flag data is read to extract occurrence parameters, obtain vector interaction frequency values, and locate the tensor topology node coordinates. The tensor topology node coordinates are located by capturing the row and column indices temporarily stored in the spatial index register as double-precision scalars. The process is then performed according to vector... The interaction frequency values ​​are allocated mapping coefficients using hardware-level nonlinear activation function mapping instructions. The adaptive temperature scaling parameter is configured to be 0.25. The vector interaction frequency values ​​are pushed into the activation function hardware logic gate to perform floating-point exponentiation and output mapping coefficients. A general matrix multiplication hardware acceleration algorithm is used to aggregate coefficients to construct node weights. The matrix multiplication operation block length parameter is configured to be 16 x 16 units. The mapping coefficient matrix is ​​multiplied by the adjacent feature matrix. The output floating-point scalar accumulation result is written into a contiguous video memory block to construct node weights. All allocated and processed node weights are aggregated and written into a multidimensional array space according to the node arrangement order of the topology graph to generate a frequency weight node distribution tensor.

[0073] Based on the frequency-weighted node distribution tensor, a fixed-step memory slicing microinstruction is used to extract word blank pause durations. By sending address offset configuration parameters to the memory controller to set the read offset to 8 bytes, the contiguous storage data within the tensor's physical address block is locked. A 64-bit double-precision floating-point bus capture operation is invoked to separate scalar data representing residence time to extract word blank pause durations. A pointer-synchronous addressing capture operation is used to extract the node's corresponding weight. By shifting the address pointer 4 bits forward, the corresponding floating-point value stored in the physical space is read to extract the node's corresponding weight. A single-instruction multiple-data stream hardware multiplier microinstruction is invoked to calculate the product of the duration and weight. The pause duration data is loaded into the multiplicand register of the multiplication unit, and the corresponding weight of the node is loaded into the multiplier register. The system first sends out concurrent multiplication machine cycle control signals to obtain the double-precision floating-point product value calculation time and weighted product. It then uses a floating-point number out-of-bounds truncation statistical algorithm to compare the product with the gain limit. It loads the preset gain limit constant by reading the system read-only memory. The preset gain limit constant is written to the motherboard boot sector through system initialization to complete the preset setting and configure the constant value as 15.5. The product value and the preset gain limit constant are sent to the hardware comparator array to perform value comparison. The interrupt masking circuit is triggered to remove out-of-bounds data greater than the constant 15.5 and retain the product value that does not exceed the limit. It calls the continuous memory block to append and write the micro-instruction summary of the product value within the limit. It allocates a one-dimensional continuous addressing dynamic random access memory space block to fill in the data in sequence and establishes a weighted blank gain sequence.

[0074] Based on the weighted blank gain sequence, the touch tremor attenuation multiplier is extracted using kernel status word parsing instructions. Data from a specific register pushed onto the top of the internal status word stack in the previous system machine cycle is retrieved. A shift extraction operation mask is used to isolate the high-order data, and the lower 16-bit fixed-point components are read to extract the touch tremor attenuation multiplier. Vectorized bitwise multiplication concurrent instructions are used to calculate the bitwise product of the attenuation multiplier and the sequence elements. By configuring the vector processing unit to enable 128 parallel arithmetic logic cores, the multinomial floating-point elements within the weighted blank gain sequence are evenly distributed to the independent L1 cache of each arithmetic logic core. Single-precision floating-point multiplication opcodes are synchronously issued to execute the multiplication of the attenuation multiplier with the corresponding cached sequence elements, obtaining a parallel output array to calculate the bitwise product of the attenuation multiplier and the sequence elements. An accumulation register pipeline polling mechanism is used to accumulate the bitwise product value. The set sums are configured with an initial summation state of 0.0 as a scalar constant. The pipeline pointer traverses the array one by one, outputting all scalar elements in parallel and pushing them into the adder to perform continuous summation operations to obtain the set sum. Fixed-point signed division and bitmasking truncation algorithms are used to divide the sum by the total number of elements to extract the truncated mean. The hardware divider is configured to load the set sum scalar value into the dividend register, and write the original total number of elements participating in the accumulation operation into the divisor register. The divider operation clock cycle is activated to obtain the quotient value with a long decimal place. The underlying bitwise AND logic instruction is called to load the system-wide fixed constant mask to mask the last 4 decimal places of the quotient value, forcibly setting them to zero and discarding the remainder data to extract the truncated mean. The truncated mean is directly sent to the system output buffer bus and allocated a unique logical identifier to be loaded into the single-precision floating-point variable space. The cognitive state logic judgment coefficient is output.

[0075] A graph neural network is used to load cross-modal asynchronous feature vectors and behavioral intention topology connection tensors. The internal mesh connection matrix structure of the behavioral intention topology connection tensor is parsed, and the intersecting nodes of rows and columns inside the mesh connection matrix are extracted to establish spatial indices. The coordinates of the topology nodes of the tensor are located based on the spatial indices. The corresponding data components inside the cross-modal asynchronous feature vectors are mapped according to the coordinates of the topology nodes of the tensor. The state record identifiers inside the data components are read, the occurrence parameters are obtained by parsing the state record identifiers, and the vector interaction frequency values ​​are extracted.

[0076] Based on graph neural networks, a direct memory access parallel loading microinstruction is used to load cross-modal asynchronous feature vectors and behavioral intent topology connection tensors. By issuing source physical address and target graphics processing unit memory address block configuration instructions to the motherboard's direct memory access controller, the bus burst transmission length is set to 256 bytes. A sparse matrix compression row storage parsing algorithm is called to parse the internal mesh connection matrix structure of the behavioral intent topology connection tensor, and a preset tensor sparsity constant is read. This preset tensor sparsity constant is directly written to a specified sector of read-only memory through the system's underlying initialization boot program to complete the preset setting and configure it as a constant of 0.8. The underlying... The pointer offset traversal instruction scans the non-zero element value array and column index array to extract the intersecting nodes of the rows and columns within the mesh connection matrix structure. A 2-megabyte hash table contiguous memory block is allocated, and the row and column indices (double-precision floating-point values) are pushed sequentially into the mapping table as key-value pairs to establish a spatial index. Based on the spatial index, a binary search memory addressing microprogram is used to locate the coordinates of the tensor topology node. The algorithm is configured with an initial left boundary pointer constant of 0 and an initial right boundary pointer constant equal to the total number of elements in the hash table minus 1. The arithmetic logic unit is activated to perform an arithmetic right shift operation of 1 bit, calculates the physical address of the intermediate pointer, matches the target key-value pair, and returns the specific memory offset scalar to locate the tensor topology node. Point coordinates, based on tensor topology node coordinates, are mapped using multidimensional array tensor slicing mapping of underlying instructions to the corresponding data components within the cross-modal asynchronous feature vector. The first dimension of the slicing operation is set to start and end indices, mapping the tensor topology node coordinates to scalar values. The second dimension of the slicing operation is set to a length constant of 128 single-precision floating-point feature words. This triggers a hardware-level multiplexer channel to strip tuples at specific positions within contiguous storage blocks, mapping them to the corresponding data components within the cross-modal asynchronous feature vector. Microcode is used to capture specific bytes using a bitmask to read the internal state record identifier of the data component. A globally fixed decimal Boolean mask constant of 65280 is loaded to drive logic. The logic unit performs a bitwise AND operation on the first 16 bits of the data component header (two bytes) to isolate the payload feature, retains the high-order byte, reads the internal status record identifier of the data component, and uses a fixed-point shift decoding opcode to parse the status record identifier to obtain the occurrence parameters. The central processing unit's shift register is configured to load the original binary stream data of the status record identifier. An 8-bit right shift control pulse is issued to shift the high-order feature byte to the low-order data bus, strips redundant control bits, and obtains the occurrence parameters. The floating-point conversion hardware unit is called to directly convert the occurrence parameters into 32-bit unsigned integer data, loads it into an independent result register, and extracts the vector interaction frequency value.

[0077] The specific steps for obtaining the quantitative decay evolution evaluation constant are as follows:

[0078] Based on cross-modal asynchronous feature vectors and behavioral intention topology connection tensors, vector interaction frequencies are extracted, topology node coordinates are located and frequency mapping node weights are matched, blank pause durations are removed, durations and weights are fused, and a weight blank gain sequence is established.

[0079] Based on the weighted blank gain sequence, the touch tremor attenuation factor is extracted and the amplitude of the attenuation factor modulation sequence element is introduced. The modulation elements are integrated to construct a global feature set. The global feature set is compressed to extract the truncated representation mean and output the cognitive state logic judgment coefficient.

[0080] Based on cross-modal asynchronous feature vectors and behavioral intent topology connection tensors, a fixed-length scan microinstruction using tensor dimension indexing is employed to extract vector interaction frequencies. By writing the starting physical address to the memory controller register and configuring a single scan step size of 32 bytes, the data bus sequentially fetches single-precision integer values ​​from consecutive storage blocks and loads them into the central processing unit's general-purpose register array to extract vector interaction frequencies. A hash-mapping memory addressing algorithm is used to locate the topology node coordinates and match the frequency-mapped node weights. A preset addressing offset constant is loaded by reading system memory. This preset addressing offset constant is directly written to the underlying registers during the motherboard's power-on self-test (POST) phase and configured as a constant of 128. The arithmetic logic unit is called to append the preset addressing offset constant to the base address pointer and perform addition operations to obtain the absolute physical memory address, thus locating the topology node coordinates and driving the addressing process. The pointer retrieves the double-precision floating-point scalar value temporarily stored in the memory block to match the frequency mapping node weights. The hardware logic subtractor opcode with timestamp difference is used to remove the blank pause duration. The interrupt request corresponding to the system nanosecond-level timestamp is triggered by extracting two adjacent interactive actions. The subtractor is configured to receive the post-timestamp by the subtrahend port and the pre-timestamp by the subtrahend port. The subtraction execution clock pulse is sent to obtain the time difference to remove the blank pause duration. The duration and weight are merged by a single instruction multi-data stream floating-point concurrent multiplication microprogram. The blank pause duration is pushed into the first operand cache array of the vector processing unit and the frequency mapping node weights are pushed into the second operand cache array. The 16-way parallel hardware multiplier is activated to perform bitwise floating-point multiplication operations. The output product scalar is extracted and a continuous 4-megabyte dynamic random access memory space block is requested to fill the product scalar data in sequence to establish a weight blank gain sequence.

[0081] Based on the weighted blank gain sequence, the touch jitter attenuation factor is extracted using micro-instructions that read the offset from the underlying hardware state register. By locking the base address of a specific reserved block in the graphics processing unit's video memory, an address shift control signal is sent to shift the read pointer 16 bytes higher. The 32-bit data bus is then used to directly capture the solidified single-precision floating-point value within the physical space to extract the touch jitter attenuation factor. An adaptive scaling hardware pipeline multiplier algorithm is used to introduce the amplitude of the attenuation factor modulation sequence elements. The constant operand port of the hardware pipeline multiplier is configured to constantly load the scalar value of the touch jitter attenuation factor. The direct memory access controller drives the multiple floating-point elements within the weighted blank gain sequence to be sequentially pushed into the variable operand pipeline port, triggering the multiplier array to execute continuous scalar multiplication micro-instructions to calculate the modulation result and extract the modulation elements. A continuous physical memory space append allocation instruction is used to integrate the modulation elements and construct a global feature set. A continuous blank memory segment with a specification of 2048 floating-point words is requested from the operating system kernel, and the high-speed cache is called... The microcode directly maps all modulated elements into a blank memory segment according to the ascending time order to construct a global feature set. The arithmetic logic unit uses fixed-point truncation, accumulation, and division microcode to compress the global feature set and extract the truncated mean. The hardware adder's accumulation register is configured with an initial state of constant 0. All elements within the global feature set are traversed and pushed into the adder to perform accumulation operations and obtain the sum scalar. The total integer quantity constant of the elements participating in the accumulation operation is written to the divisor register of the divider hardware module. A hardware-level division operation is performed between the sum scalar and the quantity constant, outputting a double-precision floating-point quotient value within a machine cycle. A globally fixed hexadecimal Boolean mask constant 0xFFFFFF00 is loaded, driving the logic unit to perform a bitwise AND operation on the mantissa of the double-precision floating-point quotient value, forcibly masking the lower 8 bits of binary data to eliminate minor disturbances and extract the truncated mean. The truncated mean is pushed into the system output bus interface buffer and allocated a globally unique static memory identifier, loaded into a single-precision variable storage space, and outputs the cognitive state logic judgment coefficient.

[0082] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention in any other way. Any person skilled in the art may make changes or modifications to the above-disclosed technical content to create equivalent embodiments that can be applied to other fields. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the protection scope of the present invention.

Claims

1. A method for predicting cognitive decline in the elderly by integrating behavioral and voice interaction, characterized in that, Includes the following steps: Step 1: Extract the operation offset line spacing and determine the duration of the response vowel sounding by using the touch coordinate sequence and audio sampling signal. Calculate the absolute difference between the line spacing and the duration, compare the product values ​​within the set tolerance filtering limit, and generate the action pronunciation timing mapping matrix. Step 2: Based on the action pronunciation timing mapping matrix, extract the object-finding pause step length and extract the extreme value of the word-ending trill, call dynamic time warping to calculate the cost value, determine whether the cost value exceeds the tolerance superimposed logarithmic penalty term fusion path, and establish a cross-modal asynchronous feature vector; Step 3: Based on the cross-modal asynchronous feature vector, extract the wandering coordinate sequence and separate the fault nodes, calculate the straight-line distance between the coordinates and the nodes, compare the critical value to retain short-distance point pairs, take the inverse distance to allocate weights and aggregate the network matrix to obtain the behavior intention topology connection tensor; Step 4: Based on the cross-modal asynchronous feature vector and the behavioral intention topology connection tensor, the graph neural network assigns node weights, calculates the weight blank product to construct the gain, superimposes the tremor attenuation operation to summarize the feature mean, and outputs the cognitive state logic judgment coefficient. Step 5: Based on the cognitive state logic judgment coefficient, separate the expected time of disordered residence, extract the absolute value and calculate the reciprocal, multiply the reciprocal by the point and the judgment coefficient, summarize the sum of multiple values ​​and calculate the arithmetic mean to obtain the quantitative decay evolution evaluation constant.

2. The method for predicting cognitive decline in the elderly by integrating behavior and voice interaction according to claim 1, characterized in that, The action-pronunciation timing mapping matrix includes operation offset row spacing, response vowel prolongation time scale, and multiplicative values ​​of the limit product. The cross-modal asynchronous feature vector includes object-finding pause step size, word-ending trill extrema, and logarithmic penalty term fusion path. The behavioral intention topology connection tensor includes wandering coordinate sequence fault nodes, short-distance point pairs, and distance inverse weighted mesh matrix. The cognitive state logic judgment coefficient includes graph neural network node weights, weight blank product to construct gain, and summed feature mean. The quantitative decay evolution evaluation constant includes the absolute value of the disordered dwell time expectation term inverse, the result of the decision coefficient by point product, and the arithmetic mean of the sum of multiple values.

3. The method for predicting cognitive decline in the elderly by integrating behavior and voice interaction according to claim 1, characterized in that, The specific steps for generating the action pronunciation timing mapping matrix are as follows: Based on the touch coordinate sequence and audio sampling signal, the offset line spacing is extracted by segmenting the coordinate position, the extended time scale is extracted by truncating the audio waveform, the line spacing and time scale are spliced ​​to construct a combined node, the absolute difference of the data inside the node is calculated, the absolute difference elements are integrated to establish a line spacing and time scale difference sequence. Based on the line spacing and time stamp difference sequence, a comparison threshold is defined, the size of the difference within the sequence is determined, elements exceeding the threshold are removed, the line spacing and time stamp of the remaining elements are extracted, the product of line spacing and time stamp is calculated, the product of line spacing and time stamp is aggregated, and an action pronunciation timing mapping matrix is ​​generated.

4. The method for predicting age-related cognitive decline by integrating behavioral and voice interaction as described in claim 1, characterized in that, The specific steps for establishing the cross-modal asynchronous feature vector are as follows: Based on the action pronunciation timing mapping matrix, the row and column data are split to extract the object-finding pause step length, high-frequency noise is filtered out to extract the word-ending trill extreme value, extreme value time nodes are matched, the object-finding pause step length and trill extreme value are merged to construct a combination item, and the extreme value step length combination item is aggregated to generate a modal extreme value step length set. Based on the set of modal extreme step lengths, the internal pause step length and trill extreme value are analyzed, dynamic time warping is called to calculate the absolute term of the step length and extreme value deviation, weight multipliers are added to calculate the replacement value, the replacement value is summarized to form a continuous sequence, the sequence data is arranged with reference to the time sequence, and a path replacement value array is generated. Based on the path cost array, the tolerance limit of the cost is defined, the specific numerical values ​​of the array elements are compared, the elements exceeding the limit are extracted and superimposed with logarithmic penalty multipliers, the penalty values ​​are fused with the initial path, multiple node data are merged, and a cross-modal asynchronous feature vector is established.

5. The method for predicting cognitive decline in the elderly by integrating behavior and voice interaction according to claim 1, characterized in that, The dynamic time warping involves constructing a reference sequence using pause step lengths, extracting trill extrema extremes to construct a comparison test sequence, establishing a two-dimensional cost grid by crossing the reference sequence and the comparison test sequence, locating the row and column intersection coordinate nodes within the two-dimensional cost grid, reading the corresponding pause step length value and the corresponding trill extrema extreme value of the coordinate node, performing a subtraction operation between the corresponding pause step length value and the corresponding trill extrema extreme value to obtain the node difference, extracting the absolute value of the node difference, and generating an absolute term for the step length and extreme value deviation.

6. The method for predicting cognitive decline in the elderly by integrating behavior and voice interaction according to claim 4, characterized in that, The process involves fusing the penalty value with the initial path, extracting the prior estimation matrix containing multiple time nodes within the initial path, arranging the multiple over-limit elements in time sequence to construct a penalty value diagonal tensor, and then using the Kalman filter algorithm to perform the Hadamard product operation between the penalty value diagonal tensor and the multiple time node prior estimation matrices. This process calculates the posterior state covariance matrix with attached nonlinear penalty weights, extracts the main diagonal elements of the posterior state covariance matrix to cover the scalar values ​​of multiple original nodes within the initial path, and finally establishes the numerically updated associated fused path tensor.

7. The method for predicting cognitive decline in the elderly by integrating behavior and voice interaction according to claim 1, characterized in that, The specific steps to obtain the behavioral intent topology connection tensor are as follows: Based on the cross-modal asynchronous feature vector, axial values ​​are extracted to construct a wandering coordinate sequence, the interactive fault nodes are separated by the interruption time point, the difference between the coordinate sequence and the fault node position is calculated, the sum of the squared differences is accumulated to obtain the straight-line distance, and a set of wandering fault spacing is generated. Based on the set of wandering fault spacings, a critical limit for distance comparison is defined, the size of the internal straight-line distance is determined, values ​​exceeding the critical limit are eliminated, the inverse of the straight-line distance within the limit is extracted, and the inverse of the distance is assigned to construct connection weights, thereby obtaining the behavioral intention topology connection tensor.

8. The method for predicting cognitive decline in the elderly by integrating behavior and voice interaction according to claim 1, characterized in that, The specific steps for outputting the logical judgment coefficients of the cognitive state are as follows: Based on the cross-modal asynchronous feature vector and the behavioral intention topology connection tensor, a graph neural network is invoked to extract vector interaction frequency values ​​and locate the tensor topology node coordinates. Mapping coefficients are assigned according to the frequency values, and the coefficients are aggregated to construct node weights, generating a frequency weight node distribution tensor. Based on the frequency weight node distribution tensor, the word blank pause duration is stripped, the corresponding weight of the node is extracted, the product of duration and weight is calculated, the product is compared with the gain limit and the product value that does not exceed the limit is retained, the product values ​​within the limit are summarized, and a weight blank gain sequence is established. Based on the weighted blank gain sequence, the touch tremor attenuation multiplier is extracted, the attenuation multiplier is calculated and the sequence element is multiplied bitwise, the bitwise product values ​​are accumulated to obtain the sum of the set, the sum is divided by the total number of elements to extract the truncated mean, and the cognitive state logic judgment coefficient is output.

9. The method for predicting age-related cognitive decline by integrating behavior and voice interaction according to claim 1, characterized in that, The graph neural network loads cross-modal asynchronous feature vectors and behavioral intention topology connection tensors, parses the internal mesh connection matrix structure of the behavioral intention topology connection tensor, extracts the intersecting nodes of rows and columns within the mesh connection matrix to establish spatial indices, locates the coordinates of the tensor topology nodes based on the spatial indices, maps the corresponding data components within the cross-modal asynchronous feature vectors according to the coordinates of the tensor topology nodes, reads the state record identifiers within the data components, parses the state record identifiers to obtain occurrence parameters, and extracts the vector interaction frequency values.

10. The method for predicting age-related cognitive decline by integrating behavioral and voice interaction according to claim 1, characterized in that, The specific steps for obtaining the quantitative decay evolution evaluation constant are as follows: Based on the cross-modal asynchronous feature vector and the behavioral intention topology connection tensor, the vector interaction frequency is extracted, the topology node coordinates are located and the frequency mapping node weights are matched, the blank pause duration is removed, the duration and weight are fused, and a weight blank gain sequence is established. Based on the weighted blank gain sequence, the touch tremor attenuation factor is extracted and the amplitude of the attenuation factor modulation sequence element is introduced. The modulation elements are integrated to construct a global feature set. The global feature set is compressed to extract the truncated representation mean and output the cognitive state logic judgment coefficient.