Cognitive disorder intervention strategy generation method and device, equipment and storage medium
By collecting and fusing multimodal data to construct individualized cognitive maps, dynamically generating training tasks and analyzing cognitive trends in real time, this approach solves the problems of insufficient matching degree and lack of dynamic perception ability in existing cognitive impairment intervention programs, and achieves efficient and personalized cognitive intervention.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- PING AN TECH (SHENZHEN) CO LTD
- Filing Date
- 2026-01-06
- Publication Date
- 2026-04-17
AI Technical Summary
Existing technologies cannot achieve the fusion analysis and personalized modeling of multimodal data, resulting in insufficient matching between cognitive impairment intervention programs and users' actual cognitive needs. Furthermore, they lack the ability to dynamically perceive the real-time cognitive state of patients, affecting the effectiveness and sustainability of intervention outcomes.
It collects users' visual, speech, and text multimodal input data, generates a unified multimodal vector through cross-modal feature fusion, constructs an individualized cognitive map, dynamically generates training tasks and conducts multimodal interactions, analyzes cognitive states in real time, and generates intervention strategy suggestions.
It improved the accuracy of intervention strategies, reduced labor costs, enhanced service accessibility and user participation, and enabled real-time multidimensional perception and personalized training of cognitive states.
Smart Images

Figure CN121885218A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology and can be applied to the medical and financial technology fields. In particular, it relates to a method, device, equipment and storage medium for generating intervention strategies for cognitive impairment. Background Technology
[0002] Cognitive impairment diseases such as Alzheimer's disease and mild cognitive impairment are rapidly increasing globally, becoming a major health threat affecting the quality of life of the elderly. Traditional intervention methods have significant limitations: manual interventions heavily rely on professional therapists, which are not only costly in terms of manpower but also difficult to implement on a large scale; existing digital training systems generally adopt fixed process designs and lack the ability to dynamically perceive the real-time cognitive state of patients. In the field of fintech, standardized digital financial service interfaces fail to fully consider the special needs of elderly users and those with cognitive decline. Complex interaction processes and professional financial terminology often exceed their cognitive processing capabilities, leading to reduced service accessibility and increased financial risk. The core deficiency of the current technological system lies in its inability to achieve multimodal data fusion analysis and personalized modeling. It cannot accurately capture changes in cognitive state reflected by nonverbal cues such as facial expressions and tone of voice, and it lacks the ability to generate dynamic tasks based on medical knowledge graphs.
[0003] Existing solutions typically employ a single-modal, independent processing approach during feature extraction, making it difficult to establish cross-modal correlation analysis. Furthermore, the task generation stage relies excessively on pre-set templates, failing to adaptively adjust difficulty based on real-time user performance. These technical limitations directly result in insufficient alignment between the training scheme and the user's actual cognitive needs, impacting the effectiveness and sustainability of the intervention. Summary of the Invention
[0004] The purpose of this application is to provide a method, apparatus, device, and storage medium for generating intervention strategies for cognitive impairment, so as to improve the accuracy of intervention and reduce labor costs.
[0005] To address the aforementioned technical problems, embodiments of this application provide a method for generating intervention strategies for cognitive impairment, comprising: Collect multimodal input data from users, including visual, speech, and text. Features of each modality in the multimodal input data are extracted to obtain multimodal features. The different multimodal features are then fused to obtain a unified multimodal vector, which is used as user state description information. A personalized cognitive graph is constructed based on a predefined medical knowledge base and the user status description information; The user's material data is retrieved from the training knowledge base, and a training task is constructed based on the material data and the individualized cognitive graph; Perform the training task and interact with the user in a multimodal manner to generate a cognitive state record; The cognitive trends in the cognitive state records are analyzed and anomalies are detected. When a decline in the target cognitive dimension is detected, intervention strategy suggestions are generated.
[0006] To address the aforementioned technical problems, embodiments of this application provide an intervention strategy generation device for cognitive impairment, comprising: The data acquisition module is used to collect multimodal input data from users, including visual, speech, and text. The feature fusion module is used to extract features from each modality of the multimodal input data to obtain multimodal features, and to fuse the different multimodal features to obtain a unified multimodal vector, which is then used as user state description information. The graph construction module is used to construct an individualized cognitive graph based on a predefined medical knowledge base and the user status description information; The task construction module is used to retrieve the user's material data from the training knowledge base and construct training tasks based on the material data and the individualized cognitive graph. The task execution module is used to execute the training task and interact with the user in a multimodal manner to generate a cognitive state record; The suggestion generation module is used to analyze and detect anomalies in the cognitive state records. When a decline in the target cognitive dimension is detected, intervention strategy suggestion information is generated.
[0007] To solve the above-mentioned technical problems, one technical solution adopted by the present invention is to provide a computer device, including one or more processors; and a memory for storing one or more programs, such that the one or more processors implement the cognitive impairment intervention strategy generation method described in any one of the above-mentioned methods.
[0008] To solve the above-mentioned technical problems, one technical solution adopted by the present invention is: a computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the intervention strategy generation method for cognitive impairment as described in any one of the above-mentioned methods.
[0009] This invention provides a method, apparatus, device, and storage medium for generating intervention strategies for cognitive impairment. The method includes: collecting multimodal input data from a user (visual, speech, and text); extracting features from each modality of the multimodal input data to obtain multimodal features, fusing different multimodal features to obtain a unified multimodal vector, and using the unified multimodal vector as user state description information; constructing an individualized cognitive graph based on a predefined medical knowledge base and the user state description information; retrieving the user's material data from a training knowledge base, and constructing a training task based on the material data and the individualized cognitive graph; executing the training task and interacting with the user in a multimodal manner to generate a cognitive state record; analyzing and detecting anomalies in the cognitive trends in the cognitive state record, and generating intervention strategy suggestions when a decline in a target cognitive dimension is detected. This invention, by fusing multimodal data to construct a personalized cognitive graph, dynamically generating adaptive training tasks, and analyzing cognitive trends in real time, solves the problems of traditional intervention methods relying on manual intervention and lacking dynamic perception capabilities. It has the advantages of improving intervention accuracy, reducing labor costs, and enhancing service accessibility. Attached Figure Description
[0010] To more clearly illustrate the solutions in this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0011] Figure 1 This is a schematic diagram of an application environment for a method for generating intervention strategies for cognitive impairment according to an embodiment of the present invention; Figure 2 This is a flowchart illustrating the implementation of the cognitive impairment intervention strategy generation method provided in this application embodiment; Figure 3 yes Figure 2 A flowchart illustrating a specific implementation method of step S2; Figure 4 yes Figure 2 A flowchart illustrating a specific implementation method of step S3; Figure 5 yes Figure 2 A flowchart illustrating a specific implementation of step S4; Figure 6 yes Figure 5 A flowchart illustrating a specific implementation of step S44; Figure 7 yes Figure 2 A flowchart illustrating a specific implementation of step S5; Figure 8 yes Figure 2 A flowchart illustrating a specific implementation of step S6; Figure 9 This is a schematic diagram of the cognitive impairment intervention strategy generation device provided in the embodiments of this application; Figure 10 This is a schematic diagram of the computer device provided in the embodiments of this application. Detailed Implementation
[0012] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the terminology used herein in the specification of the application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application; the terms "comprising" and "having," and any variations thereof, in the specification, claims, and foregoing drawings of this application, are intended to cover non-exclusive inclusion. The terms "first," "second," etc., in the specification, claims, or foregoing drawings of this application are used to distinguish different objects, not to describe a particular order.
[0013] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0014] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.
[0015] The present invention will now be described in detail with reference to the accompanying drawings and embodiments.
[0016] It should be noted that the cognitive impairment intervention strategy generation method provided in this application embodiment is generally executed by a server, and correspondingly, the cognitive impairment intervention strategy generation device is generally configured in the server.
[0017] The method for generating intervention strategies for cognitive impairment provided in this invention can be applied to, for example... Figure 1In this application environment, the client communicates with the server via a network. The server can receive multimodal input data (visual, speech, and text) from the client and generate intervention strategy suggestions based on the multimodal input data. In this invention, the server sends the intervention strategy suggestions to the client. The client can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a standalone server or a server cluster consisting of multiple servers. The invention will be described in detail below through specific embodiments.
[0018] The cognitive impairment intervention strategy generation method provided in this application embodiment can be applied to mental health and emotion intervention scenarios in the medical field, or to personalized wealth management scenarios for high-net-worth clients and elderly clients in the fintech field.
[0019] In existing technologies, cognitive impairment rehabilitation training mainly relies on manual intervention or digital therapy using pre-set software. Manual intervention suffers from high costs and scarce resources, while digital therapies generally lack depth and are limited by rigid, one-dimensional tasks. Current systems cannot perceive patients' emotional fluctuations and cognitive load in real time, making it difficult to guarantee training effectiveness. In the financial services sector, standardized digital interfaces fail to consider the cognitive characteristics of elderly users, and complex operating procedures can easily lead to decision-making errors. The root cause of these problems lies in the lack of multi-dimensional, real-time perception capabilities for users' cognitive states in current technologies.
[0020] To address the aforementioned issues, the inventors discovered systemic deficiencies in existing systems regarding data acquisition dimensions, dynamic modeling capabilities, and feedback mechanisms. By analyzing the common needs of medical rehabilitation and financial services, they recognized that constructing a multimodal perception system is fundamental to achieving personalized services. Based on cross-modal feature fusion technology, they proposed a method to collaboratively analyze visual, speech, and text data to comprehensively capture user states. To solve the challenge of dynamic adjustment, they introduced a cognitive graph modeling method, combined with a medical knowledge base, to achieve individualized modeling. Ultimately, this resulted in a closed-loop system architecture of "perception-modeling-training-feedback." Therefore, this application proposes a method for generating intervention strategies for cognitive impairment, comprising: collecting multimodal input data of the user's vision, speech, and text; extracting features from each modality of the multimodal input data to obtain multimodal features, and fusing different multimodal features to obtain a unified multimodal vector, using the unified multimodal vector as user state description information; constructing an individualized cognitive map based on a predefined medical knowledge base and user state description information; retrieving the user's material data from the training knowledge base, and constructing a training task based on the material data and the individualized cognitive map; executing the training task and interacting with the user in a multimodal manner to generate a cognitive state record; analyzing and detecting anomalies in the cognitive trends in the cognitive state record, and generating intervention strategy suggestion information when a decline in the target cognitive dimension is detected.
[0021] Specifically, the system collects multi-dimensional data on user facial expressions, voice tone, and language expression in real time using multi-channel sensors. It employs deep convolutional networks to extract facial micro-expression features, analyzes speech emotion features using acoustic models, and parses text semantic features using natural language processing techniques. A cross-modal attention layer establishes the correlation between visual, speech, and text features, generating a feature vector containing the user's comprehensive cognitive state. Based on cognitive assessment standards in a medical knowledge base, the feature vector is mapped to quantitative indicators for dimensions such as memory, language, and executive function. A dynamically updated cognitive ability map is constructed by combining the temporal trends of the user's historical training data. Based on the weak dimensions identified in the map, corresponding training materials are matched from the knowledge base to generate personalized training programs including memory cards, language repetition, and logical reasoning. During execution, the system records the user's response accuracy, reaction time, and other indicators in real time, detecting abnormal fluctuations in cognitive ability through time series analysis. When a persistent decline in a specific dimension is detected, the system automatically adjusts the intensity of the training task or recommends intervention strategies.
[0022] It should be noted that the information generated in this application embodiment is an intervention strategy suggestion information, which is a suggestion report for the user, and not a specific treatment plan.
[0023] Please see Figure 2 , Figure 2 This paper illustrates one specific implementation of a method for generating intervention strategies for cognitive impairment.
[0024] It should be noted that if substantially the same result is obtained, the method of this invention is not based on... Figure 2 Limited to the order shown, this method includes the following steps: S1: Collect multimodal input data from users, including visual, speech, and text inputs.
[0025] Multimodal input data refers to visual images, voice signals, and text information collected synchronously through cameras, microphones, and text input devices. Specifically, it can be achieved by using the OpenCV library for video frame capture, ASR technology for speech-to-text conversion, and text input interfaces for data reception, in order to comprehensively capture the user's behavioral characteristics and cognitive state.
[0026] S2: Extract the features of each modality in the multimodal input data to obtain multimodal features, and fuse the different multimodal features to obtain a unified multimodal vector, and use the unified multimodal vector as user state description information.
[0027] Among them, the unified multimodal vector refers to the feature vector fused through a cross-modal attention mechanism. Specifically, it can be implemented using the multi-head attention mechanism in the Transformer architecture to eliminate the semantic gap between different modal data.
[0028] Please see Figure 3 , Figure 3 A specific implementation of step S2 is shown below: S21: Perform facial expression and human posture recognition on the visual input in the multimodal input data to obtain the visual modal features. S22: Perform acoustic emotion recognition on the speech input in the multimodal input data to generate the speech modal features. S23: Perform semantic parsing and fluency analysis on the text input in the multimodal input data to generate the text modal features. S24: Use a cross-modal attention mechanism to fuse the visual modal features, the speech modal features, and the text modal features to obtain the unified multimodal vector, which is then used as the user state description information.
[0029] Specifically, visual input obtains expression and pose feature vectors through facial keypoint detection and 3D pose estimation. For example, it uses 68 facial keypoints combined with human skeletal joint coordinates to generate features. Speech input generates sentiment probability distribution vectors through Mel-spectrum feature extraction and time-series modeling. For example, it uses 256-dimensional Mel-spectrum coefficients as input to a bidirectional LSTM network to output sentiment classification results. Text input generates semantic vectors through dependency parsing and semantic role labeling, and simultaneously generates a fluency score vector through word frequency statistics and inter-sentence correlation calculation. A cross-modal attention mechanism calculates the correlation weights between visual, speech, and text features. For example, it uses scaled dot product attention to calculate attention scores between different modalities, and then weights and fuses the feature vectors to generate a unified multimodal vector.
[0030] Among these, facial expression feature and human posture recognition refers to extracting features from a user's facial muscle movements and body posture using computer vision algorithms. Specifically, this can be implemented using the OpenFace toolkit combined with the OpenPose framework to capture non-verbal behavioral features. Acoustic emotion recognition involves extracting acoustic features such as prosody, pitch, and speech rate from speech signals for emotion classification. This can be implemented using the OpenSmile toolkit combined with an LSTM neural network to quantify a user's emotional state. Semantic parsing and fluency analysis involves performing syntactic analysis and semantic role labeling on text content, while simultaneously calculating sentence coherence metrics. This can be implemented using the BERT model combined with a text coherence evaluation algorithm to assess language logic and organizational ability. Cross-modal attention mechanisms involve establishing a correlation weight matrix between features from different modalities. This can be implemented using a multi-head attention mechanism to address the information redundancy problem among multimodal features.
[0031] Traditional multimodal feature fusion methods typically employ feature concatenation or weighted averaging, such as directly connecting features from different modalities or adding them according to a preset ratio, leading to feature redundancy and information conflicts. This solution, however, dynamically establishes feature associations through a cross-modal attention mechanism. For example, it automatically learns the interaction weights of different modal features during attention calculation, effectively eliminating interference from irrelevant features. Existing visual analysis technologies often focus on a single dimension, such as detecting only facial expressions or analyzing only body posture. This solution, however, enhances the completeness of nonverbal behavior representation through dual visual feature extraction, such as simultaneously capturing facial micro-expressions and body posture changes. This application can perceive the user's multi-dimensional cognitive state in real time. For example, by fusing facial expressions, posture, vocal emotion, and textual logic features, it accurately identifies the user's current cognitive load and emotional fluctuations. This solves the problem of misjudgment caused by traditional systems relying solely on single-modal data, such as avoiding the situation where user anxiety is ignored based solely on vocal content analysis. Furthermore, it provides a precise state description basis for the generation of personalized intervention strategies. For example, by using a unified multimodal vector, it accurately reflects the specific dimensions of the user's cognitive decline, providing a reliable basis for subsequent training task adjustments.
[0032] S3: Construct an individualized cognitive graph based on a predefined medical knowledge base and the user status description information.
[0033] Among them, the personalized cognitive graph refers to a dynamic knowledge graph that reflects the dimensions of a user's cognitive ability. Specifically, it can be implemented by combining graph neural networks with a medical ontology library to establish a quantitative model of a user's cognitive state.
[0034] Please see Figure 4 , Figure 4 A specific implementation of step S3 is shown below: S31: Generate the user's state score across multiple cognitive dimensions using a large language model based on the predefined medical knowledge base and the user's state description information. S32: Obtain the user's historical training data and perform temporal encoding on the historical training data using a Transformer temporal modeling layer to identify trend changes in the historical training data and generate target historical trend data. S33: Map the state score and the target historical trend data to the individualized cognitive map.
[0035] The large language model refers to a deep learning model with natural language understanding capabilities, specifically implemented using a pre-trained model based on the Transformer architecture. This model semantically associates unstructured medical knowledge base content with real-time user status data. The predefined medical knowledge base is a database containing medical standards and assessment indicators related to cognitive impairment. This can be implemented using a knowledge graph that integrates structured clinical guidelines and expert experience, providing medical evidence for status score generation. The Transformer temporal modeling layer is a neural network structure with a self-attention mechanism, specifically implemented using stacked multi-head attention modules. It extracts temporal features by capturing long-term dependencies in historical training data. The target historical trend data refers to a time-series feature vector reflecting the changing patterns of the user's cognitive abilities. This can be generated by segmenting and aggregating the temporal encoding results using a sliding window mechanism, representing the evolutionary trajectory of cognitive states.
[0036] Specifically, disease diagnostic criteria and cognitive assessment indicators from the medical knowledge base are transformed into vector representations and input together with user state descriptions into a large language model for joint reasoning. The large language model generates quantitative scores for cognitive dimensions such as memory ability and executive function through semantic matching. For example, it outputs a standardized score in the 0-1 range after comparing language fluency features with medical standards. Historical training data is segmented along the time dimension and input into a Transformer encoder. Its self-attention mechanism automatically identifies the correlation patterns between data at different time steps, such as discovering a decreasing trend in reaction time in three consecutive spatial training tasks. Finally, the state scores of each cognitive dimension at the current moment are spatially projected with the historical trend vectors to form a dynamic graph containing node attributes and connection weights, where nodes represent specific cognitive dimensions and edge weights reflect the strength of the correlation between dimensions.
[0037] Traditional methods use fixed templates to match medical knowledge base entries, which cannot adapt to the individual differences in cognitive characteristics. Existing time-series analysis mostly relies on statistical methods such as moving averages, which are difficult to capture the non-linear evolution of cognitive states. This application's embodiments achieve dynamic adaptation of medical knowledge through the semantic reasoning capabilities of a large language model, and utilize the Transformer's self-attention mechanism to establish a cross-time-step association model, so that the generated cognitive graph contains both a precise description of the current state and a deep representation of historical trends. This application achieves dynamic modeling and trend prediction of users' cognitive states, solving the technical deficiency of traditional systems in constructing personalized cognitive graphs. Specifically, the semantic processing of the medical knowledge base allows clinical standards to flexibly adapt to individual characteristics, overcoming evaluation biases caused by rigid rules; the time-series coding layer's in-depth mining of historical data can accurately identify gradual or abrupt patterns of cognitive ability, providing trend warnings for intervention strategies; the spatial mapping of multi-dimensional features forms a visualized cognitive state topology, providing a quantitative basis for generating personalized training tasks.
[0038] S4: Retrieve the user's material data from the training knowledge base, and construct a training task based on the material data and the individualized cognitive graph.
[0039] Among them, training task construction refers to generating targeted training schemes based on the cognitive graph defect dimensions. Specifically, it can be achieved by combining knowledge graph-based retrieval algorithms with reinforcement learning strategies to achieve accurate matching between training content and cognitive defects.
[0040] Please see Figure 5 , Figure 5 A specific implementation of step S4 is shown below: S41: Identify the cognitive dimensions to be strengthened based on the individualized cognitive graph. S42: Retrieve the user's material data from the training knowledge base. S43: Construct a task template including task objectives, scenarios, and task formats based on the cognitive dimensions to be strengthened and the material data. S44: Generate training instructions and interactive questions for the task template using a multimodal large model, and adjust the task difficulty coefficient in the task template to obtain the training task.
[0041] The training knowledge base refers to a database storing user-specific materials. Specifically, it can be constructed by collecting images, voice dialogue records, and text information from users' daily interactions, ensuring a high degree of relevance between training tasks and users' life scenarios. The task template refers to a framework structure containing task objectives, scene settings, and interaction methods. Specifically, it can utilize a predefined template library based on cognitive dimension classification, generating an initial task framework by matching the dimension to be strengthened with the material type. The multimodal large model refers to a generative model with multimodal processing capabilities for text, speech, and images. Specifically, it can adopt an architecture that integrates visual-language pre-trained models to automatically generate interactive instructions and questions that match the user's cognitive level. The task difficulty coefficient is an adjustable parameter reflecting the complexity of the training task. Specifically, it can be dynamically adjusted by analyzing users' historical error rates, response times, and task completion rates, forming a difficulty adaptive mechanism.
[0042] Specifically, firstly, by analyzing the cognitive dimension scores in the individualized cognitive graph, the target dimensions requiring focused training, such as memory or logical reasoning, are identified. Then, materials related to the user's life experiences, such as family photos and frequently used dialogue snippets, are retrieved from the training knowledge base. Based on the target dimensions and material types, a pre-defined task template framework is matched; for example, combining memory training objectives with an image sorting scenario to generate the basic task structure. Next, a multimodal large-scale model is invoked to automatically generate natural language operation instructions based on the scenario elements in the template, and interactive questions tailored to the user's cognitive level are designed, such as requiring the user to arrange event nodes in historical photos in chronological order. Finally, based on the user's error frequency and reaction speed in recent training, the complexity of elements in the task is dynamically increased or decreased, such as increasing the number of images or extending the problem-solving time limit, forming a training task sequence with adaptive difficulty.
[0043] Traditional cognitive training systems typically employ fixed question banks or preset task flows, failing to adjust content based on the user's real-time cognitive state. This application, however, achieves precise matching between training content and user capabilities by dynamically constructing task templates and adjusting difficulty in real-time. Existing technologies often use generic task materials lacking personalized relevance. This application, by utilizing a dedicated knowledge base to access materials from the user's everyday life, significantly enhances the immersion and acceptance of training tasks. Furthermore, existing systems rely on manually setting difficulty parameters. This application, through a multimodal large-scale model, automatically generates interactive content and dynamically adjusts it based on performance data, forming a closed-loop optimization mechanism. This application addresses the problems of existing cognitive training systems: monotonous tasks, rigid interactions, and an inability to dynamically adapt to the user's cognitive state. It achieves precise positioning of training objectives through individualized cognitive maps, avoiding wasted resources in ineffective training; enhances the relevance and attractiveness of training content through a user-specific material library, increasing user participation; achieves human-like natural interaction through multimodal generation technology, reducing resistance to mechanical operations; and maintains training tasks within the user's ability threshold through a dynamic difficulty adjustment mechanism, avoiding frustration while ensuring training effectiveness. Ultimately, this leads to a personalized intervention plan that can optimize training content in real time based on the user's cognitive state.
[0044] Please see Figure 6 , Figure 6 A specific implementation of step S44 is shown below: S441: Generate the training instructions and interactive questions for the task template using the multimodal large model. S442: Obtain the user's current performance information and historical error rate. S443: Adjust the task difficulty coefficient in the task template based on the current performance information and the historical error rate to obtain the training task.
[0045] The current performance information refers to the user's response speed, accuracy, and emotional feedback data during real-time interaction. This can be achieved by collecting multimodal interaction data from users in real time and extracting key indicators. Its purpose is to provide an immediate cognitive basis for difficulty adjustment, avoiding a mismatch between task difficulty and the user's current ability. The historical error rate refers to the frequency of errors in a specific cognitive dimension during historical training tasks. This can be achieved by statistically analyzing the error distribution data in the user's past task records. Its purpose is to identify the user's skill gaps through long-term data mining, ensuring that difficulty adjustments align with individual cognitive development patterns.
[0046] Specifically, when generating training tasks, the multimodal large model first generates interactive questions and guidance information containing text, images, and voice based on the task template. Subsequently, the system collects real-time data on the user's response speed, accuracy, and emotional fluctuations during task execution, forming current performance information. Simultaneously, it extracts the error frequency distribution of users across different cognitive dimensions from the historical database, calculating the historical error rate. Based on the immediate feedback from the current performance information, the system determines whether the user is experiencing cognitive overload or inattention; if an anomaly is detected, the difficulty coefficient is reduced. Combining this with long-term weaknesses identified in the historical error rate, the system specifically increases the task complexity in the relevant dimensions. Through the dual analysis of real-time and long-term data, the task difficulty coefficient is dynamically adjusted to a balance that avoids frustration while stimulating cognitive potential, ultimately generating a training task adapted to the user's current state.
[0047] Existing cognitive training systems typically employ fixed difficulty levels or adjustment rules based on single historical data, failing to detect real-time changes in user cognitive load. For example, traditional methods linearly increase difficulty based solely on historical accuracy, forcing users to continue high-difficulty training even when their performance temporarily declines due to emotional fluctuations, exacerbating frustration. This application, however, integrates real-time performance data with historical error patterns to construct a dynamic difficulty balancing mechanism. This mechanism automatically lowers task requirements when the user is in poor condition and continuously strengthens weak areas based on long-term skill deficiencies, achieving precise control of the personalized difficulty curve. This application solves the problem of poor training results caused by the inability of existing systems to dynamically adjust task difficulty. Specifically, it achieves dynamic matching between task difficulty and current cognitive state through real-time user performance data collection, avoiding decreased training efficiency due to tasks being too difficult or too easy; it establishes a difficulty adjustment strategy that aligns with individual cognitive development patterns by combining historical error rate analysis, ensuring that training content always targets the user's skill weaknesses; and it generates diverse interactive questions through a multimodal large-scale model, enhancing the fun and adaptability of training tasks, thereby improving user compliance and long-term training sustainability.
[0048] S5: Execute the training task and interact with the user in a multimodal manner to generate a cognitive state record.
[0049] Please see Figure 7 , Figure 7 A specific implementation of step S5 is shown below: S51: The training task is conducted through multiple rounds of interaction with the user via voice dialogue and visual display to collect the user's interaction responses in real time. S52: Semantic evaluation is performed based on the interaction responses using a large language model to obtain semantic evaluation results. S53: The cognitive state record is generated based on the semantic evaluation results.
[0050] The multi-turn interaction of voice dialogue and visual display refers to generating natural language questions using a speech synthesis engine and simultaneously displaying visual task scenarios using an image rendering engine. For example, speech synthesis is achieved using Tacotron 2, and a 3D dynamic scene is constructed using the Three.js framework. This interaction method can cover both auditory and visual perception channels, ensuring effective participation from users with different cognitive preferences. Semantic evaluation involves converting user voice responses into text and inputting it into a pre-trained large language model to extract logical coherence, keyword coverage, and sentiment indicators from the response content. For example, semantic vector encoding is performed using the GPT-4 model, and the semantic distance between the response and the standard answer is calculated using cosine similarity. This evaluation method overcomes the limitations of traditional rule matching, achieving human-like judgment of response quality. Cognitive state recording involves integrating semantic evaluation results with metadata such as interaction timestamps and response durations into structured time-series data. For example, JSON format is used to record the evaluation score, error type, and sentiment label for each interaction. This structured record provides a traceable, fine-grained data foundation for subsequent trend analysis.
[0051] Specifically, in financial customer risk assessment scenarios, when users answer investment preference questions via voice, the system simultaneously displays visual comparison charts of different financial products. The large language model not only determines whether the financial products chosen by the user match the risk level but also analyzes contradictions in their decision-making logic. For example, if a user claims to be a conservative investor but chooses a high-volatility product, the semantic assessment will mark this logical inconsistency. Each interaction generates a cognitive state record containing scores for decision accuracy, response speed, and logical consistency, forming a traceable cognitive ability change curve.
[0052] Traditional financial risk assessment systems only record the final choice result and cannot capture cognitive biases during the decision-making process. Fixed-choice question interaction restricts user freedom of expression, while this application's embodiment captures the natural decision-making process through multimodal interaction and identifies potential cognitive risks through semantic analysis. Existing technologies rely on manual review of abnormal transactions; this application's embodiment achieves real-time cognitive state monitoring through automated semantic assessment, significantly improving the timeliness of risk warnings. This application solves the problem that existing systems cannot perceive users' cognitive states in real time. In elderly financial service scenarios, when users continuously exhibit semantic logical confusion, the system can immediately adjust the task difficulty or trigger manual intervention to avoid investment errors caused by cognitive decline. In cognitive rehabilitation training, therapists can accurately locate the specific manifestations of memory decline or executive function impairment through semantic assessment results and develop targeted intervention plans.
[0053] S6: Analyze and detect anomalies in the cognitive state records. When a decline in the target cognitive dimension is detected, generate intervention strategy suggestions.
[0054] Please see Figure 8 , Figure 8 A specific implementation of step 6 is shown below: S61: Analyze the cognitive trends and detect anomalies in the cognitive state records using a time-series prediction model to obtain trend and anomaly detection results. S62: If the trend and anomaly detection results indicate a decline in the target cognitive dimension, generate intervention strategy recommendations based on the target cognitive dimension. S63: Generate a structured report based on the intervention strategy recommendations and provide the structured report to the user.
[0055] Among these, time-series prediction models refer to prediction algorithms based on time-series data modeling, specifically implemented using Long Short-Term Memory Networks or Autoregressive Integral Moving Average models, to capture the dynamic patterns of cognitive states changing over time. Anomaly detection algorithms are statistical methods for identifying deviations from normal patterns in data, specifically implemented using Isolation Forest algorithms or dynamic threshold setting methods, to distinguish between short-term fluctuations and persistent decline trends. Structured reports refer to the output format of intervention strategies organized according to a preset template, specifically implemented using JSON format or natural language generation technology, to transform professional assessment conclusions into actionable guidance recommendations.
[0056] Specifically, time-series data from cognitive state records are input into a trained time-series prediction model, which extracts feature change patterns along the time dimension using a sliding window mechanism. An anomaly detection algorithm performs residual analysis between the prediction results and the measured data. When the residuals of multiple consecutive time windows exceed a dynamic threshold, a decline determination for the target cognitive dimension is triggered. Intervention strategy recommendations are generated based on the type of decline dimension; for example, a memory reinforcement training plan is generated for memory decline, and a task decomposition strategy is generated for executive function decline. A structured report transforms the intervention strategy into a standardized operational guide containing task type, execution frequency, and difficulty gradient through predefined field mapping, and is then fed back to the user terminal via voice broadcast and a graphical interface.
[0057] Traditional methods rely on static interventions based on single assessment results, failing to distinguish between natural fluctuations in cognitive abilities and genuine decline. This solution, through time-series modeling and dynamic threshold detection, effectively filters out short-term interference factors and accurately identifies continuous decline trends. Existing systems typically use fixed thresholds to trigger warnings, easily leading to false alarms or missed alarms. This application's embodiment, based on residual analysis, uses dynamic threshold settings that adaptively adjust sensitivity according to individual user differences, improving detection reliability. Traditional intervention strategies rely solely on current state data; this application's embodiment combines historical trends to predict future decline risks, enabling preventative intervention. This application can capture the gradual change in cognitive abilities in real time, triggering an early warning mechanism in the early stages of decline, avoiding intervention delays caused by assessment lags in traditional methods. Differentiated intervention strategies are generated for the decline characteristics of different cognitive dimensions, addressing the mismatch between intervention measures and specific deficiencies in existing technologies. The standardized output format of structured reports lowers the barrier to understanding professional medical knowledge, enabling non-professionals to accurately implement intervention plans and improving the feasibility of rehabilitation training in home settings. In this embodiment, multimodal input data of the user's vision, speech, and text are collected; features of each modality in the multimodal input data are extracted to obtain multimodal features, and different multimodal features are fused to obtain a unified multimodal vector, which is used as user state description information; a personalized cognitive graph is constructed based on a predefined medical knowledge base and the user state description information; the user's material data is retrieved from the training knowledge base, and a training task is constructed based on the material data and the personalized cognitive graph; the training task is executed, and multimodal interaction is performed with the user to generate a cognitive state record; the cognitive trends in the cognitive state record are analyzed and anomalies are detected, and when the decline of the target cognitive dimension is detected, intervention strategy suggestion information is generated. This embodiment of the invention, by fusing multimodal data to construct a personalized cognitive graph, dynamically generates adaptive training tasks, and analyzes cognitive trends in real time, solves the problems of traditional intervention methods relying on manual intervention and lacking dynamic perception capabilities. It has the advantages of improving intervention accuracy, reducing labor costs, and enhancing service accessibility.
[0058] This application achieves real-time, multi-dimensional perception of users' cognitive states, overcoming the technical limitation of traditional systems' single-dimensional data collection. By dynamically constructing individualized cognitive maps, it overcomes the limitation of static evaluation models in failing to reflect dynamic changes in cognitive abilities. Based on a closed-loop training mechanism using multimodal interaction, it can adjust the training difficulty according to the user's real-time performance, effectively improving the targeting and timeliness of intervention strategies.
[0059] In mental health and emotion intervention scenarios, the system's accurate identification of emotional states makes it suitable for adjunctive treatment of anxiety and depression. For example, when the system detects a user's low mood, it can dynamically generate mindfulness meditation guidance, positive psychology Q&A, or pleasant memory recall tasks. Virtual companions can provide empathetic dialogue, becoming a readily accessible "digital psychological coach." In personalized wealth management scenarios for high-net-worth clients and senior clients, the application logic is that identifying client emotions can improve the service experience, and identifying agent status can ensure service quality. Specific applications: Intelligent Emotion Routing: When a call comes in, the system quickly determines the customer's emotions (anger, anxiety) through voice emotion analysis and prioritizes transferring the call to a more experienced customer service expert with stronger empathy.
[0060] Agent Status Monitoring and Assistance: Real-time monitoring of customer service staff's emotional state and fatigue levels. When the system detects that an agent is experiencing emotional exhaustion or decreased attention, it can automatically provide script suggestions, trigger rest reminders, or temporarily remove the agent from the high-stress queue, thereby ensuring service quality and caring for employee health.
[0061] In this embodiment, multimodal input data of the user's vision, speech, and text are collected; features of each modality in the multimodal input data are extracted to obtain multimodal features, and different multimodal features are fused to obtain a unified multimodal vector, which is used as user state description information; a personalized cognitive graph is constructed based on a predefined medical knowledge base and the user state description information; the user's material data is retrieved from the training knowledge base, and a training task is constructed based on the material data and the personalized cognitive graph; the training task is executed, and multimodal interaction is performed with the user to generate a cognitive state record; the cognitive trends in the cognitive state record are analyzed and anomalies are detected, and when the decline of the target cognitive dimension is detected, intervention strategy suggestion information is generated. This embodiment of the invention, by fusing multimodal data to construct a personalized cognitive graph, dynamically generates adaptive training tasks, and analyzes cognitive trends in real time, solves the problems of traditional intervention methods relying on manual intervention and lacking dynamic perception capabilities. It has the advantages of improving intervention accuracy, reducing labor costs, and enhancing service accessibility.
[0062] Please refer to Figure 9 As a response to the above Figure 2 The implementation of the method shown in this application provides an embodiment of an intervention strategy generation device for cognitive impairment, which is similar to... Figure 2 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.
[0063] like Figure 9As shown, the cognitive impairment intervention strategy generation device in this embodiment includes: a data acquisition module 71, a feature fusion module 72, a map construction module 73, a task construction module 74, a task execution module 75, and a suggestion generation module 76, wherein: Data acquisition module 71 is used to collect multimodal input data of the user, including visual, speech and text. The feature fusion module 72 is used to extract the features of each modality data in the multimodal input data to obtain multimodal features, and to fuse the different multimodal features to obtain a unified multimodal vector, and to use the unified multimodal vector as user state description information. The graph construction module 73 is used to construct an individualized cognitive graph based on a predefined medical knowledge base and the user status description information; The task construction module 74 is used to retrieve the user's material data from the training knowledge base and construct a training task based on the material data and the individualized cognitive graph. The task execution module 75 is used to execute the training task and interact with the user in a multimodal manner to generate a cognitive state record; The suggestion generation module 76 is used to analyze and detect anomalies in the cognitive state records. When a decline in the target cognitive dimension is detected, intervention strategy suggestion information is generated.
[0064] Furthermore, the feature fusion module 72 includes: A visual modality feature generation unit is used to perform facial expression feature and human posture recognition on the visual input in the multimodal input data to obtain the visual modality features; The speech modal feature generation unit is used to perform acoustic emotion recognition on the speech input in the multimodal input data and generate the speech modal features; The text modality feature generation unit is used to perform semantic parsing and fluency analysis on the text input in the multimodal input data to generate the text modality features; The user state description information generation unit is used to perform feature fusion on the visual modal features, the speech modal features and the text modal features using a cross-modal attention mechanism to obtain the unified multimodal vector, and use the unified multimodal vector as the user state description information.
[0065] Furthermore, the map construction module 73 includes: The status score generation unit is used to generate the user's status score in multiple cognitive dimensions based on the predefined medical knowledge base and the user status description information using a large language model. The temporal coding unit is used to acquire the user's historical training data and perform temporal coding on the historical training data through the Transformer temporal modeling layer to identify trend changes in the historical training data and generate target historical trend data. The mapping unit is used to map the state score and the target historical trend data into the individualized cognitive map.
[0066] Furthermore, the task construction module 74 includes: A cognitive dimension identification unit is used to identify the cognitive dimensions to be strengthened based on the individualized cognitive map; The material data retrieval unit is used to retrieve the user's material data from the training knowledge base; The task template construction unit is used to construct a task template including task objectives, scenarios, and task formats based on the cognitive dimension to be reinforced and the material data. The difficulty coefficient adjustment unit is used to generate training instructions and interactive questions for the task template through a multimodal large model, and adjust the task difficulty coefficient in the task template to obtain the training task.
[0067] Furthermore, the difficulty adjustment unit includes: The training instruction generation subunit is used to generate the training instructions and the interactive questions of the task template through the multimodal large model; The historical error rate acquisition subunit is used to acquire the user's current performance information and historical error rate; The coefficient adjustment subunit is used to adjust the task difficulty coefficient in the task template based on the current performance information and the historical error rate to obtain the training task.
[0068] Furthermore, the task execution module 75 includes: An interactive response acquisition unit is used to interact with the user in multiple rounds through voice dialogue and visual display of the training task, so as to acquire the user's interactive response in real time. A semantic evaluation unit is used to perform semantic evaluation based on the interaction response using a large language model to obtain a semantic evaluation result; A cognitive state record generation unit is used to generate the cognitive state record based on the semantic evaluation result.
[0069] Furthermore, it is suggested that module 76 include: An anomaly detection unit is used to analyze and detect anomalies in the cognitive state records using a time-series prediction model, and obtain trend and anomaly detection results. An intervention strategy suggestion information generation unit is used to generate intervention strategy suggestion information based on the target cognitive dimension if the trend and anomaly detection results show that the target cognitive dimension is declining. The structured report generation unit is used to generate a structured report based on the intervention strategy recommendation information and to feed the structured report back to the user.
[0070] To address the aforementioned technical problems, embodiments of this application also provide a computer device. Please refer to [link / reference needed]. Figure 10 , Figure 10 This is a basic structural block diagram of the computer device in this embodiment.
[0071] Computer device 8 includes a memory 81, a processor 82, and a network interface 83 that are interconnected via a system bus. It should be noted that... Figure 10 Only a computer device 8 with three components—memory 81, processor 82, and network interface 83—is shown. It should be understood that implementing all shown components is not required; more or fewer components may be implemented alternatively. Those skilled in the art will understand that this computer device is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions. Its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), and embedded devices.
[0072] Computer devices can include desktop computers, laptops, handheld computers, and cloud servers. These devices allow for human-computer interaction with users through keyboards, mice, remote controls, touchpads, or voice-activated devices.
[0073] The memory 81 includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 81 may be an internal storage unit of the computer device 8, such as the hard disk or memory of the computer device 8. In other embodiments, the memory 81 may also be an external storage device of the computer device 8, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the computer device 8. Of course, the memory 81 may include both internal storage units and external storage devices of the computer device 8. In this embodiment, the memory 81 is typically used to store the operating system and various application software installed on the computer device 8, such as program code for generating intervention strategies for cognitive impairment. In addition, the memory 81 may also be used to temporarily store various types of data that have been output or will be output.
[0074] In some embodiments, processor 82 may be a central processing unit (CPU), controller, microcontroller, microprocessor, or other data processing chip. This processor 82 is typically used to control the overall operation of the computer device 8. In this embodiment, processor 82 is used to run program code stored in memory 81 or process data, for example, to run the program code of the above-described method for generating intervention strategies for cognitive impairment, to implement various embodiments of the method for generating intervention strategies for cognitive impairment.
[0075] The network interface 83 may include a wireless network interface or a wired network interface, which is typically used to establish a communication connection between the computer device 8 and other electronic devices.
[0076] This application also provides another embodiment, namely, a computer-readable storage medium storing a computer program that can be executed by at least one processor to cause the at least one processor to perform the steps of the above-described method for generating an intervention strategy for cognitive impairment.
[0077] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods of the various embodiments of this application.
[0078] Obviously, the embodiments described above are merely some embodiments of this application, not all embodiments. The accompanying drawings show preferred embodiments of this application, but do not limit the scope of this application. This application can be implemented in many different forms; rather, these embodiments are provided to provide a more thorough and comprehensive understanding of the disclosure of this application. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing specific embodiments, or make equivalent substitutions for some of the technical features. Any equivalent structures made using the content of this application's specification and drawings, directly or indirectly applied to other related technical fields, are similarly within the scope of protection of this application.
Claims
1. A method of generating an intervention strategy for cognitive impairment, characterized by, include: Collect multimodal input data from users, including visual, speech, and text. Features of each modality in the multimodal input data are extracted to obtain multimodal features. The different multimodal features are then fused to obtain a unified multimodal vector, which is used as user state description information. A personalized cognitive graph is constructed based on a predefined medical knowledge base and the user status description information; The user's material data is retrieved from the training knowledge base, and a training task is constructed based on the material data and the individualized cognitive graph; Perform the training task and interact with the user in a multimodal manner to generate a cognitive state record; The cognitive trends in the cognitive state records are analyzed and anomalies are detected. When a decline in the target cognitive dimension is detected, intervention strategy suggestions are generated.
2. The cognitive impairment intervention strategy generation method according to claim 1, characterized by, The multimodal features include visual modal features, speech modal features, and text modal features; the process of extracting features from each modal data in the multimodal input data to obtain multimodal features, and fusing different multimodal features to obtain a unified multimodal vector, using the unified multimodal vector as user state description information, includes: The visual input in the multimodal input data is subjected to facial expression features and human posture recognition to obtain the visual modality features; Acoustic emotion recognition is performed on the speech input in the multimodal input data to generate the speech modal features; Semantic parsing and fluency analysis are performed on the text input in the multimodal input data to generate the text modal features; A cross-modal attention mechanism is used to fuse the visual modal features, the speech modal features, and the text modal features to obtain the unified multimodal vector, which is then used as the user state description information.
3. The cognitive impairment intervention strategy generation method according to claim 1, characterized by, The construction of a personalized cognitive graph based on a predefined medical knowledge base and the user status description information includes: The user's state score in multiple cognitive dimensions is generated by a large language model based on the predefined medical knowledge base and the user state description information. The user's historical training data is acquired, and the historical training data is time-series encoded through the Transformer time modeling layer to identify trend changes in the historical training data and generate target historical trend data. The state score and the target historical trend data are mapped to the individualized cognitive map.
4. The cognitive impairment intervention strategy generation method according to claim 1, characterized by, The step of retrieving the user's material data from the training knowledge base and constructing a training task based on the material data and the individualized cognitive graph includes: Based on the individualized cognitive map, the cognitive dimensions that need to be strengthened are identified; Retrieve the user's material data from the training knowledge base; Based on the cognitive dimension to be strengthened and the material data, a task template including task objectives, scenarios, and task formats is constructed. The training instructions and interactive questions for the task template are generated by a multimodal large model, and the task difficulty coefficient in the task template is adjusted to obtain the training task.
5. The cognitive impairment intervention strategy generation method according to claim 4, characterized by, The process of generating training instructions and interactive questions for the task template using a multimodal large model, and adjusting the task difficulty coefficient in the task template to obtain the training task, includes: The training instructions and interactive questions for generating the task template are generated through the multimodal large model; Obtain the user's current performance information and historical error rate; The training task is obtained by adjusting the task difficulty coefficient in the task template based on the current performance information and the historical error rate.
6. The cognitive impairment intervention strategy generation method according to any one of claims 1 to 5, characterized in that, The process of performing the training task and interacting with the user in a multimodal manner to generate a cognitive state record includes: The training task involves multiple rounds of interaction with the user through voice dialogue and visual display to collect the user's interaction response in real time; Semantic evaluation results are obtained by performing semantic evaluation based on the interaction response using a large language model. The cognitive state record is generated based on the semantic evaluation results.
7. The cognitive impairment intervention strategy generation method according to any one of claims 1 to 5, characterized by, The analysis and anomaly detection of cognitive trends in the cognitive state records, when detecting a decline in the target cognitive dimension, generates intervention strategy suggestions, including: The cognitive trends and anomaly detection in the cognitive state records are analyzed using a time-series prediction model to obtain trend and anomaly detection results. If the trend and anomaly detection results show that the target cognitive dimension is declining, then the intervention strategy recommendation information is generated based on the target cognitive dimension; A structured report is generated based on the intervention strategy recommendations, and the structured report is then fed back to the user.
8. An intervention strategy generation device for cognitive impairment, characterized by, include: The data acquisition module is used to collect multimodal input data from users, including visual, speech, and text. The feature fusion module is used to extract features from each modality of the multimodal input data to obtain multimodal features, and to fuse the different multimodal features to obtain a unified multimodal vector, which is then used as user state description information. The graph construction module is used to construct an individualized cognitive graph based on a predefined medical knowledge base and the user status description information; The task construction module is used to retrieve the user's material data from the training knowledge base and construct training tasks based on the material data and the individualized cognitive graph. The task execution module is used to execute the training task and interact with the user in a multimodal manner to generate a cognitive state record; The suggestion generation module is used to analyze and detect anomalies in the cognitive state records. When a decline in the target cognitive dimension is detected, intervention strategy suggestion information is generated.
9. A computer device, comprising: It includes a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the method for generating intervention strategies for cognitive impairment as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method for generating intervention strategies for cognitive impairment as described in any one of claims 1 to 7.