Cognitive disorder diagnosis method combining super-dimensional calculation and large language model

Through hyperdimensional computing and attention fusion mechanism, multimodal sensor data is mapped to hyperdimensional space and classified. Combined with a large language model, it solves the problems of insufficient generalization ability and high computing resource consumption of traditional methods, and achieves efficient diagnosis of cognitive impairment.

CN120766992APending Publication Date: 2025-10-10SUN YAT SEN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510852326.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-10-10

AI Technical Summary

Technical Problem

The existing technologies are subject to the non-independent and identically distributed characteristics of sensor data across scenarios and the limited size of single-user training samples, resulting in insufficient generalization ability and overfitting of small samples in traditional machine learning methods. In addition, large language models consume high computing resources and have low modal fusion efficiency when processing long-term multimodal sensor data.

Method used

Hyperdimensional computing is used to map the feature vectors of multimodal samples into hyperdimensional space, an attention-inspired cross-modal feature fusion mechanism is introduced, and cognitive impairment diagnosis is performed through perceptron classification and action semantic mapping combined with a large language model.

Benefits of technology

By effectively utilizing the time series data analysis capabilities of large language models, we can solve the problems of insufficient generalization ability and high computing resource consumption of traditional methods, and achieve efficient diagnosis of cognitive impairments using multimodal sensor data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120766992A_ABST
    Figure CN120766992A_ABST
Patent Text Reader

Abstract

The invention provides a cognitive disorder diagnosis method combining super-dimensional calculation and a large language model. The cognitive disorder diagnosis method comprises the following steps: encoding a multi-modal sample feature vector into a multi-modal super-dimensional feature vector; aggregating the coded multi-modal super-dimensional feature vectors into a fused feature vector by using a cross-modal feature fusion mechanism inspired by attention; classifying through a perceptron, and mapping a classification result into an action semantic symbol; constructing a long-term activity sequence of the user through the sequence of the action semantic symbols; and constructing a structured prompt text based on a long-term activity sequence of an individual, and inputting the structured prompt text into a large language model to execute cognitive disorder diagnosis reasoning. According to the method, super-dimensional coding is introduced into an attention network mechanism, efficient and learnable dynamic weight distribution among modals is realized, and an original multi-modal sensor signal is converted into an action semantic symbol through action semantic mapping; the problems that the data dimension of a large language model, traditional machine learning depends on a large number of label samples, and the generalization ability of the model is insufficient are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of machine learning technology, and in particular to a method for diagnosing cognitive impairment by combining hyperdimensional computing with a large language model. Background Art

[0002] The development of intelligent sensing technology has provided a rich data foundation for research on human behavior perception, disease monitoring, and diagnosis. Multimodal sensor networks enable the continuous, multi-dimensional collection of human physiological information and spatial parameters through a variety of terminal devices, including smartphones, wearable wristbands, depth cameras, and radar.

[0003] Existing sensor-based motion recognition and diagnostic reasoning solutions mostly rely on machine learning methods, which can be divided into two paradigms based on model training architecture: centralized training and distributed training. However, existing research methods have the following technical bottlenecks:

[0004] First, due to uncontrollable variations in lighting, noise, and other aspects of different environments, as well as individual behavioral patterns, sensor data collected across scenarios often exhibit significant non-independent and identically distributed characteristics. Traditional machine learning methods often lack generalization capabilities due to differences in the distribution of local features. Second, the limited size of training samples for individual users can easily lead to overfitting in small samples. This poses significant challenges for disease diagnosis applications based on sensor systems.

[0005] To address the aforementioned technical issues, existing technologies have introduced large language models, which can effectively capture complex dependencies in long-term time series data. Through parameterized knowledge representation and contextual learning mechanisms, they can organically integrate domain expertise and prior information. This feature gives them unique advantages in application scenarios that require expertise, such as medical diagnosis. In addition, large language models acquire general modeling capabilities through pre-training, and combined with prompt engineering technology, they can significantly reduce dependence on labeled data in the target domain. However, when long-term multimodal sensor data is directly input, the existing large language model architecture faces significant challenges: First, the data analysis and reasoning process of large language models relies on a large number of complex calculations, and the increase in the length of the sample sequence will lead to an exponential increase in the consumption of computing resources; second, the spatial alignment and feature fusion problems of multimodal sensor data have not yet been effectively solved, and directly splicing raw data from different modalities will seriously reduce the efficiency of model reasoning. Summary of the Invention

[0006] In response to the shortcomings of the existing technology, the present invention provides a method for diagnosing cognitive impairment that combines hyperdimensional computing with a large language model. The present invention can realize the diagnosis of cognitive impairment based on long-term multimodal sensor data, and solves the limitations of traditional machine learning methods such as reliance on a large number of labeled samples and insufficient model generalization capabilities.

[0007] The technical solution of the present invention is: a method for diagnosing cognitive impairment combining hyperdimensional computing and a large language model, comprising the following steps:

[0008] S1), mapping the multimodal sample feature vector into the hyperdimensional space to obtain the encoded multimodal hyperdimensional feature vector;

[0009] S2) introduces an attention-inspired cross-modal feature fusion mechanism to perform weighted aggregation on the encoded multimodal hyperdimensional feature vectors to generate a fused feature vector;

[0010] S3) Classify the fused feature vector through the perceptron and map the classification result to the discrete action space to obtain the action semantic symbol;

[0011] S4), based on the temporal correlation characteristics of sample timestamps, construct the user's long-term activity sequence through the sequence of action semantic symbols;

[0012] S5) Construct structured prompt text based on the individual's long-term activity sequence and input it into a large language model to perform cognitive impairment diagnosis reasoning.

[0013] Preferably, in step S1), the multimodal sample feature vector is mapped into the hyperdimensional space by adopting a spatiotemporal hyperdimensional coding method that combines position-level coding with sequence-type data coding; the position-level coding decomposes the multimodal sample feature vector into a series of index-value pairs, and establishes a joint mapping of address index to position coding hyperdimensional vector and actual value to horizontal coding hyperdimensional vector, and the sequence-type data coding shifts the hyperdimensional vector through permutation operation.

[0014] Preferably, in step S1), mapping the multimodal sample feature vector into a hyperdimensional space to obtain a multimodal hyperdimensional feature vector specifically includes the following steps:

[0015] S11), preset a set of n D-dimensional super-dimensional vectors to quantize the numerical level of the feature vector X into n levels, and obtain the numerical level super-vectors {L1, L2, ..., L n}; For the feature vector X i-th feature x i , the numerical quantization process maps it to {L1,L2,...,L n}, and obtain the super-dimensional vector B representing its numerical level. i ;

[0016] S12) Initially, randomly generate m binary super-dimensional vectors of dimension D corresponding to the features {x1, x2, ..., x m} address coding, and obtain the address coding super vector {A1,A2,...,A m}; Among them, Ai Represents the i-th feature x of the feature vector X i Address code;

[0017] S13), for each feature x of the feature vector X i Perform hyperdimensional encoding, that is:

[0018]

[0019] Where S i is feature x i The result of hyperdimensional coding; A i 、B i are features x i Address encoding and hyperdimensional vectors; Represents the binding operation, that is, the bitwise XOR operation between vectors;

[0020] S14), encoding the temporal information between feature sequences through permutation and bundling operations, i.e.;

[0021]

[0022] Where H′ represents the result of encoding the temporal information between sequences; (i) represents the permutation operation, which encodes the sequence relationship by rearranging the elements of the hyperdimensional vector through circular right shift, (i) represents the number of elements to be shifted; represents the bundling operation, which is implemented by the addition operation between hyperdimensional vectors;

[0023] S15), perform dual polarization processing on H′, and the final coding result is expressed as:

[0024] H = δ(H′);

[0025] Where H represents the final encoded multimodal hyperdimensional feature vector, and δ represents the bipolarization function.

[0026] Preferably, in step S11), a D-dimensional binary L1 is first randomly generated to represent the minimum value of the feature value range, and then the D / n-bit elements of L1 are randomly flipped to generate L2. Similarly, each adjacent numerical level supervector of the next level is obtained by randomly selecting non-repeated D / n-bit elements from the numerical level supervector of the current level and flipping them. That is, each element is flipped only once throughout the process, and the final numerical level supervector L representing the maximum value is obtained. n The relationship to the minimum numerical level supervector L1 is bitwise inverted.

[0027] Preferably, in step S15), the dual polarization function is expressed as:

[0028]

[0029] Where x represents the element of the eigenvector X.

[0030] Preferably, in step S2), an attention-inspired cross-modal feature fusion mechanism is introduced to weightedly aggregate the multimodal hyperdimensional feature vectors to generate a fused feature vector, which specifically includes the following steps:

[0031] S21), the encoded multimodal hyperdimensional feature vector H j Generate the corresponding query vector Q through the embedding layer mapping with learnable parameters j , key vector K j Sum value vector V j ;

[0032] S22), query vector Q of each modality j The key vectors K of all modes are j Perform dot product operation to obtain the correlation coefficient Thus, the correlation matrix R between the modes is constructed;

[0033] S23), by introducing the learnable weight vector W r The correlation matrix is ​​further mapped and processed; and the attention weight α of each modality is obtained after normalization processing S23). j , where the calculation process of the attention mechanism is expressed as:

[0034]

[0035] Where Attention represents the attention mechanism; Q, K, and V are query vector, key vector, and value vector respectively; softmax represents the activation function; T represents the transpose operation; and E represents the embedding layer dimension.

[0036] S24), the value vector V of each mode j Perform weighted aggregation to obtain the fused feature vector, namely:

[0037] Fused({H j})=∑α j ·V j ;

[0038] Where Fused represents the fused feature vector.

[0039] Preferably, in step S3), the fused feature vector is classified by a multi-layer perceptron.

[0040] Preferably, in step S3), a mapping relationship exists between the symbol set predefined in the original action space, each action category corresponds to a corresponding semantic symbol tag, and the classification result is converted into an action semantic symbol through action semantic mapping.

[0041] Preferably, in step S5), the prompt text includes task objectives, background information, individual long-term activity sequences and individual long-term activity sequence descriptions; wherein the background information is used to describe the background and collection process description of multimodal data collection; the individual long-term activity sequence description provides a correspondence between symbolic tags in long-term activity sequence data and actual action semantics.

[0042] The beneficial effects of the present invention are:

[0043] 1. This invention effectively leverages the powerful time series data analysis and general modeling capabilities of large language models, overcoming the limitations of traditional machine learning methods, such as their reliance on large numbers of labeled samples and insufficient model generalization capabilities, in reasoning based on long-term multimodal sensor data.

[0044] 2. The present invention introduces hyperdimensional coding into the attention network mechanism to achieve efficient and learnable dynamic weight allocation between modalities, and converts the original multimodal sensor signals into action semantic symbols through action semantic mapping, solving the data dimensionality disaster difficulty of large language models. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 It is a schematic diagram of the process of the present invention;

[0046] Figure 2 Schematic diagram of the process of multimodal feature vector super-dimensional coding of the present invention;

[0047] Figure 3 Schematic diagram of the process of attention-inspired cross-modal feature fusion of the present invention. DETAILED DESCRIPTION

[0048] The specific embodiments of the present invention will be further described below with reference to the accompanying drawings:

[0049] like Figure 1 As shown, this embodiment provides a method for diagnosing cognitive impairment by combining hyperdimensional computing with a large language model, comprising the following steps:

[0050] S1), mapping the multimodal sample feature vector into the hyperdimensional space to obtain the encoded multimodal hyperdimensional feature vector;

[0051] This embodiment maps the multimodal sample feature vector into the hyperdimensional space by adopting a spatiotemporal hyperdimensional coding method that combines position-level coding with sequence-type data coding. The position-level coding decomposes the multimodal sample feature vector into a series of index-value pairs, and establishes a joint mapping of address index to position coding hyperdimensional vector and actual value to horizontal coding hyperdimensional vector. The sequence-type data coding shifts the hyperdimensional vector through permutation operation. Figure 2 As shown, the specific steps include:

[0052] S11), preset a set of n D-dimensional super-dimensional vectors to quantize the numerical level of the feature vector X into n levels, and obtain the numerical level super-vectors {L1, L2, ..., L n First, randomly generate a D-dimensional binary L1 to represent the minimum value of the feature value range, then randomly flip the D / n-bit elements of L1 to generate L2, and so on. Each adjacent numerical level supervector of the next level is obtained by randomly selecting non-repeated D / n-bit elements from the numerical level supervector of the current level and flipping them. That is, each element is flipped only once in the whole process, and finally represents the numerical level supervector L with the maximum value. n The relationship with the minimum value level super vector L1 is bitwise inverted. i , the numerical quantization process maps it to {L1,L2,...,L n}, and obtain the super-dimensional vector B representing its numerical level. i ;

[0053] S12) Initially, randomly generate m binary super-dimensional vectors of dimension D corresponding to the features {x1, x2, ..., x m} address coding, and obtain the address coding super vector {A1,A2,...,A m}; Among them, A i Represents the i-th feature x of the feature vector X i Address code;

[0054] S13), for each feature x of the feature vector X i Perform hyperdimensional encoding, that is:

[0055]

[0056] Where S i is feature x i The result of hyperdimensional coding; A i 、B i are features x i Address encoding and hyperdimensional vectors; Represents the binding operation, that is, the bitwise XOR operation between vectors;

[0057] S14), encoding the temporal information between feature sequences through permutation and bundling operations, i.e.;

[0058]

[0059] Where H′ represents the result of encoding the temporal information between sequences; (i) represents the permutation operation, which encodes the sequence relationship by rearranging the elements of the hyperdimensional vector through circular right shift, (i) represents the number of elements to be shifted; represents the bundling operation, which is implemented by the addition operation between hyperdimensional vectors;

[0060] S15), perform dual polarization processing on H′, and the final coding result is expressed as:

[0061] H = δ(H′);

[0062] Where H represents the final encoded multimodal hyperdimensional feature vector, and δ represents the bipolarization function.

[0063] The dual polarization function is expressed as:

[0064]

[0065] Where x represents the element of the eigenvector X.

[0066] S2) introduces an attention-inspired cross-modal feature fusion mechanism to perform weighted aggregation on the encoded multimodal hyperdimensional feature vectors to generate a fused feature vector; Figure 3 As shown, the specific steps include:

[0067] S21), the encoded multimodal hyperdimensional feature vector H j Generate the corresponding query vector Q through the embedding layer mapping with learnable parameters j , key vector K j Sum value vector V j ;

[0068] S22), query vector Q of each modality j The key vectors K of all modes are j Perform dot product operation to obtain the correlation coefficient Thus, the correlation matrix R between the modes is constructed;

[0069] S23), by introducing the learnable weight vector W r The correlation matrix is ​​further mapped and processed; and the attention weight α of each modality is obtained after normalization processing S23). j , where the calculation process of the attention mechanism is expressed as:

[0070]

[0071] Where Attention represents the attention mechanism; Q, K, and V are query vector, key vector, and value vector respectively; softmax represents the activation function; T represents the transpose operation; and E represents the embedding layer dimension.

[0072] S24), the value vector V of each mode j Perform weighted aggregation to obtain the fused feature vector, namely:

[0073] Fused({H j})=∑α j ·V j ;

[0074] Where Fused represents the fused feature vector.

[0075] S3) Classify the fused feature vector through the perceptron and map the classification result to the discrete action space to obtain the action semantic symbol;

[0076] The perceptron in this embodiment is a multi-layer perceptron. There is a mapping relationship between the predefined symbol sets in the original action space. Each action category corresponds to a corresponding semantic symbol label. The classification result is converted into an action semantic symbol through action semantic mapping.

[0077] For example, the category "walking" is associated with the symbolic tag "0," the category "using a mobile phone" is associated with the symbolic tag "1," and so on. Based on the above mapping relationship, the action perception results are converted into corresponding symbolic tags, realizing the mapping from feature space to action semantics.

[0078] S4), based on the temporal correlation characteristics of sample timestamps, construct the user's long-term activity sequence through the sequence of action semantic symbols;

[0079] For example, "walk (action semantic symbol is 0), use mobile phone (action semantic symbol is 1), use mobile phone (action semantic symbol is 1), walk (action semantic symbol is 0), sit down (action semantic symbol is 4)"; through the established symbol mapping rules, a standardized long-term activity sequence can be generated: "0,1,1,0,4,...", where the continuous repetition of symbols represents the continuous characteristics of the action.

[0080] S5) Construct structured prompt text based on the individual's long-term activity sequence and input it into a large language model to perform cognitive impairment diagnosis reasoning.

[0081] In this embodiment, the prompt text includes task objectives, background information, individual long-term activity sequences and individual long-term activity sequence descriptions; wherein the background information is used to describe the background and collection process of multimodal data collection; the individual long-term activity sequence description provides a correspondence between symbolic tags in long-term activity sequence data and actual action semantics.

[0082] The above embodiments and descriptions are only for explaining the principles and best embodiments of the present invention. Without departing from the spirit and scope of the present invention, the present invention may be subject to various changes and improvements, which shall fall within the scope of the invention to be protected.

Claims

1. A method for diagnosing cognitive impairment by combining hyperdimensional computing with a large language model, characterized in that: The steps include: S1), mapping the multimodal sample feature vector into the hyperdimensional space to obtain the encoded multimodal hyperdimensional feature vector; S2) introduces an attention-inspired cross-modal feature fusion mechanism to perform weighted aggregation on the encoded multimodal hyperdimensional feature vectors to generate a fused feature vector; S3) Classify the fused feature vector through the perceptron and map the classification result to the discrete action space to obtain the action semantic symbol; S4), based on the temporal correlation characteristics of sample timestamps, construct the user's long-term activity sequence through the sequence of action semantic symbols; S5) Construct structured prompt text based on the individual's long-term activity sequence and input it into a large language model to perform cognitive impairment diagnosis reasoning.

2. The method for diagnosing cognitive impairment combining hyperdimensional computing and a large language model according to claim 1, characterized in that: In step S1), the multimodal sample feature vector is mapped into the hyperdimensional space by adopting a spatiotemporal hyperdimensional coding method that combines position-level coding with sequence-type data coding; the position-level coding decomposes the multimodal sample feature vector into a series of index-value pairs, and establishes a joint mapping of address index to position coding hyperdimensional vector and actual value to horizontal coding hyperdimensional vector, and the sequence-type data coding shifts the hyperdimensional vector through permutation operation.

3. The method for diagnosing cognitive impairment combining hyperdimensional computing and a large language model according to claim 2, characterized in that: In step S1), the multimodal sample feature vector is mapped into a hyperdimensional space to obtain a multimodal hyperdimensional feature vector, which specifically includes the following steps: S11), preset a set of n D-dimensional super-dimensional vectors to quantize the numerical level of the feature vector X into n levels, and obtain the numerical level super-vectors {L1, L2, ..., L n }; For the feature vector X i-th feature x i , the numerical quantization process maps it to {L1,L2,...,L n }, and obtain the super-dimensional vector B representing its numerical level. i ; S12) Initially, randomly generate m binary super-dimensional vectors of dimension D corresponding to the features {x1, x2, ..., x m } address coding, and obtain the address coding super vector {A1,A2,...,A m }; Among them, A i Represents the i-th feature x of the feature vector X i Address code; S13), for each feature x of the feature vector X i Perform hyperdimensional encoding, that is: Where S i is feature x i The result of hyperdimensional coding; A i 、B i are features x i Address encoding and hyperdimensional vectors; Represents the binding operation, that is, the bitwise XOR operation between vectors; S14), encoding the temporal information between feature sequences through permutation and bundling operations, i.e.; Where H′ represents the result of encoding the temporal information between sequences; (i) represents the permutation operation, which encodes the sequence relationship by rearranging the elements of the hyperdimensional vector through circular right shift, (i) represents the number of elements to be shifted; represents the bundling operation, which is implemented by the addition operation between hyperdimensional vectors; S15), perform dual polarization processing on H′, and the final coding result is expressed as: H = δ(H′); Where H represents the final encoded multimodal hyperdimensional feature vector, and δ represents the bipolarization function.

4. The method for diagnosing cognitive impairment combining hyperdimensional computing and a large language model according to claim 3, characterized in that: In step S11), a D-dimensional binary L1 is first randomly generated to represent the minimum value of the feature value range, and then the D / n-bit elements of L1 are randomly flipped to generate L2. Similarly, each adjacent numerical level supervector of the next level is obtained by randomly selecting non-repeated D / n-bit elements from the numerical level supervector of the current level and flipping them. That is, each element is flipped only once throughout the process, and the final numerical level supervector L representing the maximum value is obtained. n The relationship to the minimum numerical level supervector L1 is bitwise inverted.

5. The method for diagnosing cognitive impairment combining hyperdimensional computing and a large language model according to claim 3, characterized in that: In step S15), the dual polarization function is expressed as: Where x represents the element of the eigenvector X.

6. The method for diagnosing cognitive impairment combining hyperdimensional computing and a large language model according to claim 1, characterized in that: In step S2), an attention-inspired cross-modal feature fusion mechanism is introduced to perform weighted aggregation on the multimodal hyperdimensional feature vectors to generate a fused feature vector, which specifically includes the following steps: S21), the encoded multimodal hyperdimensional feature vector H j Generate the corresponding query vector Q through the embedding layer mapping with learnable parameters j , key vector K j Sum value vector V j ; S22), query vector Q of each modality j The key vectors K of all modes are j Perform dot product operation to obtain the correlation coefficient Thus, the correlation matrix R between the modes is constructed; S23), by introducing the learnable weight vector W r Further mapping processing of the correlation matrix; And through normalization processing S23), the attention weight α of each modality is obtained j , where the calculation process of the attention mechanism is expressed as: Where Attention represents the attention mechanism; Q, K, and V are query vector, key vector, and value vector respectively; softmax represents the activation function; T represents the transpose operation; and E represents the embedding layer dimension. S24), the value vector V of each mode j Perform weighted aggregation to obtain the fused feature vector, namely: Fused({H j })=∑α j ·V j ; Where Fused represents the fused feature vector.

7. The method for diagnosing cognitive impairment combining hyperdimensional computing and a large language model according to claim 1, characterized in that: In step S3), there is a mapping relationship between the symbol sets predefined in the original action space, each action category corresponds to a corresponding semantic symbol tag, and the classification result is converted into an action semantic symbol through action semantic mapping.

8. The method for diagnosing cognitive impairment combining hyperdimensional computing and a large language model according to claim 1, characterized in that: In step S5), the prompt text includes task objectives, background information, individual long-term activity sequences and individual long-term activity sequence descriptions; wherein the background information is used to describe the background and collection process of multimodal data collection; the individual long-term activity sequence description provides a correspondence between symbolic tags in the long-term activity sequence data and actual action semantics.