Intelligent assessment and intervention system for mental health of elderly population based on multi-modal perception

By using a multimodal perception-based intelligent assessment system that combines various assessment methods and micro-behavioral analysis, the system addresses the one-sidedness of mental health assessments for the elderly, enabling precise assessment and personalized intervention of their mental state.

CN122266760APending Publication Date: 2026-06-23SHANDONG XIEHE UNIV +2
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANDONG XIEHE UNIV
Filing Date
2026-03-18
Publication Date
2026-06-23

AI Technical Summary

Technical Problem

Existing technologies for assessing the mental health of the elderly are limited by single assessment methods, neglect micro-behavioral information, and lack in-depth integration of multimodal data, resulting in one-sided assessment results and a lack of targeted intervention programs.

Method used

A multimodal perception-based intelligent assessment system is adopted, which combines pure image, pure text, image-text combination and voice assessment methods to capture micro-behavioral information, construct a multi-dimensional assessment matrix, and perform nonlinear bias analysis through RBF SVM model and attention mechanism, dynamically assign weights, and generate personalized psychological intervention plans.

Benefits of technology

It enables a comprehensive quantitative assessment of the cognitive, emotional, and behavioral characteristics of the elderly, improving the accuracy of the assessment and the pertinence of the intervention plan, and generating personalized intervention plans that are more in line with the individual psychological characteristics of the elderly.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122266760A_ABST
    Figure CN122266760A_ABST
Patent Text Reader

Abstract

The application provides an old-age group mental health intelligent evaluation and intervention system based on multi-modal perception, and belongs to the technical field of mental evaluation and intervention. The system fuses four evaluation modes of pure images, pure texts and the like, captures micro behaviors in an evaluation process to construct a multi-dimensional evaluation matrix, extracts core representations and analyzes nonlinear cognitive biases, constructs and fuses a cognitive graph layer to realize dynamic weighting of evaluation modes and themes, inputs the weighting result into a pre-training model, and outputs an individualized mental intervention scheme. The application improves the accuracy and comprehensiveness of old-age mental evaluation, and makes the intervention scheme more targeted and in line with the cognitive physiological characteristics of the old-age group.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of psychological assessment and intervention technology, and in particular to an intelligent assessment and intervention system for the mental health of the elderly based on multimodal perception. Background Technology

[0002] With the accelerating aging of the population, the mental health of the elderly is receiving increasing social attention. Scientific and accurate mental health assessments are a prerequisite for effective interventions, but current mental health assessment techniques for the elderly still have significant limitations: 1. Current assessments often rely on single methods such as plain text questionnaires, pure image projective tests, or voice interviews. For example, plain text assessments are easily limited by the elderly's education level and comprehension ability, making it difficult to accurately reflect their true psychological state; pure image assessments (such as drawing tests), while able to capture subconscious expressions, lack linguistic support, easily leading to biased assessments; and voice assessments heavily depend on the assessor's subjective judgment, lacking objectivity. A single assessment method cannot comprehensively cover the cognitive, emotional, and behavioral characteristics of the elderly, thus reducing the accuracy of assessment results.

[0003] 2. The micro-behaviors of the elderly during the assessment process (such as pen pressure, keystroke frequency, facial micro-expressions, and fluctuations in physiological signs) contain rich information about their psychological activities. However, existing technologies often only focus on the assessment results themselves, neglecting the capture and analysis of dynamic micro-behaviors during the assessment process, thus missing key clues about their psychological state.

[0004] 3. While some existing technologies attempt to combine multiple assessment methods, they often remain at the level of simply summarizing results, failing to achieve deep integration and correlation analysis of information from different modalities. For example, image-text combined assessments merely present image and text results side by side without exploring the cognitive connections between them; multimodal data lacks a unified quantitative framework, making it difficult to form a holistic understanding of the mental health status of the elderly.

[0005] The above factors make the results of psychological assessments of the elderly too one-sided, which often makes intervention programs unable to address the core psychological problems of the elderly, resulting in poor intervention effects and failing to meet the diverse and personalized mental health needs of the elderly population.

[0006] Therefore, this invention proposes an intelligent assessment and intervention system for the mental health of the elderly based on multimodal perception. Summary of the Invention

[0007] This invention provides an intelligent assessment and intervention system for the mental health of the elderly based on multimodal perception, in order to solve the aforementioned technical problems.

[0008] This invention proposes an intelligent assessment and intervention system for the mental health of the elderly based on multimodal perception, comprising: The assessment module is used to assess the target elderly person sequentially based on assessment files related to the pure image assessment method, the pure text assessment method, the combined image and text assessment method, and the voice assessment method, and to obtain an assessment vector for each assessment method. The assessment vector contains a sub-vector of each assessment topic, and each assessment vector contains the same assessment topic. The micro-behavior analysis module is used to extract the cognitive behavior posture set, micro-expression posture set, and physiological posture set of the target elderly person for each assessment topic under each assessment method based on the captured micro-behavior information during the assessment process of each assessment method, and to construct the first multi-dimensional assessment matrix of the target elderly person based on the same assessment topic. The matrix construction module is used to extract core cognitive representations, core facial expression representations, and physiological auxiliary cognitive representations from the cognitive behavior posture set, micro-expression posture set, and physiological posture set of each assessment topic under the same assessment method, and construct the second multidimensional assessment matrix of the target elderly based on the same assessment method. The second multidimensional assessment matrix contains multiple representations of different assessment topics under the same assessment method, and each row vector corresponds to one representation. The layer construction module is used to analyze the first nonlinear cognitive bias of the target elderly person under each assessment method in the same first multidimensional assessment matrix and construct a first complete cognitive layer. At the same time, it analyzes the second nonlinear cognitive bias of the target elderly person under each assessment topic in the same second multidimensional assessment matrix and constructs a second complete cognitive layer. The weight assignment module is used to merge the first complete cognitive layer and the second complete cognitive layer to obtain a third complete cognitive layer, and to assign assessment weights to each assessment method and theme weights to each assessment topic based on the third complete cognitive layer, and to attach them to each assessment method and the corresponding assessment vector. The intervention module is used to input additional results into a pre-trained psychological assessment intervention model and output a psychological intervention plan adapted to the target elderly person.

[0009] Preferably, the evaluation module includes: The pure image assessment unit is used to acquire the drawn images of the target elderly based on each assessment theme in the pure image assessment method, and extract the brush stroke trajectory features, color distribution features, and composition complexity features of the drawn images to obtain the sub-vectors of the corresponding assessment theme under the pure image. The plain text assessment unit is used to obtain the input text content of the target elderly person based on each assessment topic in the plain text assessment method, and extract the semantic coherence features, emotional polarity features, and word frequency features of the input text content to obtain the sub-vector of the corresponding assessment topic under plain text. The image-text combined assessment unit is used to extract the image drawing features and text input features of the target elderly people based on each assessment topic in the image-text combined assessment method, and obtain the sub-vector of the corresponding assessment topic under the image-text combined method; The voice assessment unit is used to acquire the audio of the target elderly person's voice statement based on each assessment topic in the voice assessment method, and extract the prosodic features, speech rate features, and emotional intensity features of the voice statement audio to obtain the sub-vector of the corresponding assessment topic under the voice. The splicing unit is used to splice the sub-vectors of different assessment topics under the same assessment method to obtain the assessment vector of the corresponding assessment method.

[0010] Preferably, the micro-behavioral information includes all capture results for each assessment topic under different assessment methods; The columns of the first multidimensional evaluation matrix correspond to four evaluation methods: pure image, pure text, image and text combination, and voice. The rows correspond to cognitive behavior posture set, micro-expression posture set, and physiological posture set. The matrix elements are feature vectors of the corresponding posture set under the corresponding evaluation method.

[0011] Preferably, the matrix construction module includes: The sequence determination unit is used to perform temporal segmentation and feature encoding on the cognitive behavior posture set, micro-expression posture set and physiological posture set corresponding to each assessment topic under the same assessment method, so as to obtain the temporal feature sequence corresponding to each posture set. Density clustering unit is used to perform adaptive density clustering on each time-series feature sequence to generate several clusters, and to calculate the cluster center vector and intra-cluster feature dispersion of each cluster. The correction unit is used to retrieve the pre-stored benchmark posture feature library under the same evaluation method, perform multi-dimensional similarity matching between each cluster and the benchmark cluster, select the benchmark cluster with the highest matching degree as the correction reference, and correct the current cluster center vector and the feature dispersion within the cluster based on the statistical characteristics of the benchmark cluster. The representation determination unit is used to determine the core cognitive representation of each assessment topic by taking the cluster center vector of the corrected cognitive behavior posture set as the core cognitive representation and the cluster center vector of the corrected micro-expression posture set as the core facial expression representation. The physiological auxiliary determination unit is used to generate physiological auxiliary cognitive representations based on the cluster center vector of the corrected physiological posture set, and by using the ratio of the intra-cluster feature dispersion of the core cognitive representation to the mean intra-cluster feature dispersion of all representations as a weighting coefficient. The matrix construction unit is used to construct a multimodal representation matrix under this assessment method, with each assessment topic as the column and core cognitive representation, core facial expression representation, and physiological auxiliary cognitive representation as the row. It is regarded as the second multidimensional assessment matrix. The matrix elements include the feature vector of the corresponding representation and the corrected intra-cluster feature dispersion.

[0012] Preferably, the layer construction module includes: The normalization processing unit is used to normalize the feature vectors corresponding to the cognitive behavior posture set, micro-expression posture set, and physiological posture set in the same first multi-dimensional evaluation matrix. The dimensionality reduction processing unit is used to extract dimensionality from the normalized matrix and retain the core feature components related to cognitive biases in the elderly population. The model building unit is used to introduce temporal correlation weights, combine the assessment duration and response delay characteristics of the target elderly under each assessment method, and perform weighted correction on the core feature components corresponding to each assessment method to construct a nonlinear deviation quantification model for each assessment method. The nonlinear deviation quantification model is an SVM nonlinear deviation model based on radial basis function (RBF). The first deviation determination unit is used to calculate the nonlinear deviation value between each assessment method and the pre-stored benchmark matrix of healthy elderly people based on the nonlinear deviation quantification model, which is regarded as the first nonlinear cognitive deviation. The first layer construction unit is used to construct a spatial coordinate system with four assessment methods as the horizontal dimension and three types of posture sets as the vertical dimension. The first nonlinear cognitive deviation value of each assessment method is mapped to the heat value in the spatial coordinate system to generate a deviation heat distribution map. At the same time, the deviation correlation curve between assessment methods is embedded. The deviation heat distribution map and the deviation correlation curve are integrated to form a first complete cognitive layer containing the first deviation quantification value, the deviation distribution law and the correlation characteristics of assessment methods. The depth of the heat value corresponds to the magnitude of the deviation, and the slope of the correlation curve corresponds to the deviation change trend.

[0013] Preferably, the layer construction module further includes: The coefficient calculation unit is used to extract the feature vectors and corrected intra-cluster feature dispersion of the core cognitive representation, core facial expression representation, and physiological auxiliary cognitive representation in the same second multi-dimensional assessment matrix, calculate the dispersion variation coefficient of each representation under each assessment topic, and quantify the stability of the representation. The second deviation determination unit is used to input the multi-representation features corresponding to each assessment topic into a pre-constructed multi-representation fusion nonlinear deviation analysis model, compare it with the pre-stored benchmark representation data of the elderly healthy population under the same assessment method and assessment topic, calculate the nonlinear deviation quantification value of each assessment topic, and regard it as the second nonlinear cognitive deviation. The nonlinear deviation analysis model is a multi-representation fusion model based on the attention mechanism. Network building units are used to construct a topic-representation bias association network with assessment topics as nodes, second nonlinear cognitive bias values ​​as node weights, and representation stability as node association strength. The second layer construction unit is used to visualize and label the deviation levels of each assessment topic, integrate the correlation network and deviation level labeling to form a second complete cognitive layer containing the second deviation quantification value, the stability of the representation, and the deviation correlation between topics. In this layer, the node size corresponds to the deviation weight, and the thickness of the connection between nodes corresponds to the strength of the representation correlation.

[0014] Preferably, the weight assignment module includes: The feature extraction unit is used to extract the first deviation quantification features, the correlation features between assessment methods and the posture dimension distribution features corresponding to the assessment methods in the first complete cognitive layer, and to extract the second deviation quantification features, the representation stability features and the deviation correlation features between the assessment topics in the second complete cognitive layer. Combined with the adaptive calibration function constructed by the cognitive physiological characteristics of the elderly population, the heterogeneous interference in the features of the two types of layers is eliminated, and a calibrated feature set is generated. The vector generation unit is used to quantify the correlation strength of the features of the two types of layers based on the temporal response characteristics of the evaluation methods in the first complete cognitive layer and the coefficient of variation of the discreteness of the core representations in the second complete cognitive layer, and dynamically generate a fusion weight vector based on the correlation strength. The initial fusion map generation unit is used to integrate the calibrated feature set across dimensions according to the fusion weight vector, and optimize it based on the feature mutual information entropy to strengthen the intrinsic relationship between the evaluation method, evaluation theme and multimodal representation, and generate the initial fusion map. The graph correction unit is used to introduce the fusion benchmark dataset of elderly mental health, correct the initial fusion graph, and construct a three-dimensional cognitive coordinate system of assessment method, assessment topic, and multimodal representation based on the corrected initial fusion graph. The fused deviation quantification data, feature correlation strength, and representation stability parameters are mapped to the three-dimensional cognitive coordinate system to generate a three-dimensional dynamic fusion cognitive graph. Elderly cognitive deviation warning indicators and feature importance classification labels are embedded and integrated to form a third complete cognitive layer.

[0015] Preferably, the weight assignment module further includes: The parameter extraction unit is used to extract the feature correlation strength corresponding to each assessment method from the third complete cognitive layer. Feature credibility and the contribution of deviation Extract the representation stability corresponding to each assessment topic. Thematic relevance and deviation influence coefficient Where i=1,2,3,4 corresponds to four assessment methods, and j=1,2,...,n corresponds to various assessment topics. All extracted parameters have been normalized to the range of 0 to 1. The first calculation unit is used to calculate the evaluation weight for each evaluation method. ; The second calculation unit is used to calculate the topic weight of each assessment topic. .

[0016] Compared with the prior art, the beneficial effects of this application are as follows: 1. Overcoming the limitations of existing technologies with single assessment methods and the shortcomings of simply superimposing multimodal results, this technology integrates four assessment methods: pure image, pure text, combined image and text, and voice. At the same time, it captures micro-behavioral information during the assessment process to construct a multi-dimensional assessment matrix, achieving dual quantitative evaluation of the static characteristics of the assessment results and the dynamic characteristics of the micro-behavioral process. This fills the gap in existing technologies that neglect psychological cues in micro-behavior, comprehensively covering the cognitive, emotional, and behavioral characteristics of the elderly, and achieving a qualitative breakthrough in assessment dimensions.

[0017] 2. In response to the nonlinear characteristics of the psychological state of the elderly, an SVM model based on RBF and a multi-representation fusion model based on attention mechanism are used to analyze the nonlinear cognitive biases at the assessment method and assessment topic levels, respectively. In addition, a multi-dimensional dynamic weighting algorithm is designed in combination with the cognitive physiological characteristics of the elderly to replace the traditional fixed weighting method. This enables personalized dynamic weighting of assessment methods and assessment topics, making the assessment results more in line with the individual psychological characteristics of the elderly and significantly improving the accuracy of mental health assessment.

[0018] Other features and advantages of the invention will be set forth in the following description, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention may be realized and obtained by means of the structures particularly pointed out in the written description and the accompanying drawings.

[0019] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description

[0020] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1 This is a structural diagram of an intelligent assessment and intervention system for the mental health of the elderly based on multimodal perception, as described in an embodiment of the present invention. Detailed Implementation

[0021] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.

[0022] This invention proposes an intelligent assessment and intervention system for the mental health of the elderly based on multimodal perception, such as... Figure 1 As shown, it includes: The assessment module is used to assess the target elderly person sequentially based on assessment files related to the pure image assessment method, the pure text assessment method, the combined image and text assessment method, and the voice assessment method, and to obtain an assessment vector for each assessment method. The assessment vector contains a sub-vector of each assessment topic, and each assessment vector contains the same assessment topic. The micro-behavior analysis module is used to extract the cognitive behavior posture set, micro-expression posture set, and physiological posture set of the target elderly person for each assessment topic under each assessment method based on the captured micro-behavior information during the assessment process of each assessment method, and to construct the first multi-dimensional assessment matrix of the target elderly person based on the same assessment topic. The matrix construction module is used to extract core cognitive representations, core facial expression representations, and physiological auxiliary cognitive representations from the cognitive behavior posture set, micro-expression posture set, and physiological posture set of each assessment topic under the same assessment method, and construct the second multidimensional assessment matrix of the target elderly based on the same assessment method. The second multidimensional assessment matrix contains multiple representations of different assessment topics under the same assessment method, and each row vector corresponds to one representation. The layer construction module is used to analyze the first nonlinear cognitive bias of the target elderly person under each assessment method in the same first multidimensional assessment matrix and construct a first complete cognitive layer. At the same time, it analyzes the second nonlinear cognitive bias of the target elderly person under each assessment topic in the same second multidimensional assessment matrix and constructs a second complete cognitive layer. The weight assignment module is used to merge the first complete cognitive layer and the second complete cognitive layer to obtain a third complete cognitive layer, and to assign assessment weights to each assessment method and theme weights to each assessment topic based on the third complete cognitive layer, and to attach them to each assessment method and the corresponding assessment vector. The intervention module is used to input additional results into a pre-trained psychological assessment intervention model and output a psychological intervention plan adapted to the target elderly person.

[0023] In this embodiment, the image assessment method refers to a pure image assessment, which requires the target elderly person to draw an image of the assessment topic without text; the text assessment method refers to a pure text assessment, which requires the target elderly person to input text of the assessment topic without image drawing; the combined image and text assessment method refers to an assessment topic that requires both image drawing and text input; and the voice assessment method refers to an assessment topic that the target elderly person can verbally state the assessment topic. The assessment topics used in different assessment methods are all pre-set and can be used directly.

[0024] The assessment topics are divided into four core themes based on common psychological problems among the elderly: emotional state, cognitive function, social willingness, and life satisfaction.

[0025] The assessment document refers to the assessment task document adapted to various assessment methods based on the preset assessment theme. For pure image assessment methods, a drawing task guidance document is generated, such as simply showing "Draw your ideal old age"; for pure text assessment methods, a text question document is generated, such as "Please describe your mood in the past week"; for assessment methods combining images and text, a drawing + text description task document is generated, such as "Draw your daily activities and briefly describe them"; for voice assessment methods, a voice question and answer guidance document is generated, such as "Please talk about your relationship with your family".

[0026] An assessment vector is a structured and computable vector data obtained by extracting and quantifying the features of the responses of the target elderly to all assessment topics under a certain assessment method. It is a quantitative representation of the overall assessment results under that assessment method.

[0027] A subvector refers to a fixed-dimensional vector data obtained by extracting and quantifying the features of the responses of the target elderly person to a single assessment topic under a certain assessment method. It is a component unit of the assessment vector.

[0028] Microbehavioral information refers to various types of dynamic behavioral and physiological data that can be captured when the target elderly person completes assessment tasks for each assessment topic under various assessment methods.

[0029] In this embodiment, the cognitive behavior posture set is a 128-dimensional feature vector composed of operational behavior features during the evaluation process. It includes three subsets: operation frequency, operation duration, and operation stability, with each subset having 48, 40, and 40 dimensions respectively. The subsets are then concatenated after Z-Score standardization.

[0030] The micro-expression pose set consists of a 128-dimensional feature vector composed of the coordinate features of 68 facial key points and micro-expression features. The key point coordinate features are 96-dimensional, and the micro-expression features (type, duration, frequency of occurrence) are 32-dimensional, which are concatenated after feature encoding.

[0031] The physiological posture set is a 128-dimensional feature vector composed of heart rate fluctuation features, including four types of features: mean heart rate, standard deviation of heart rate, rate of change of heart rate, and correlation between heart rate and operational behavior. Each subset has 32 dimensions, which are then spliced ​​together after normalization.

[0032] In this embodiment, the feature data of each pose set is transformed into a 128-dimensional feature vector, which is then filled into the matrix according to the row and column correspondence to form a standardized first multidimensional evaluation matrix. Each evaluation topic corresponds to a first multidimensional evaluation matrix.

[0033] Core cognitive representation refers to the quantitative representation that reflects the core cognitive characteristics of the elderly, extracted from the cognitive behavioral posture set of each assessment topic under the same assessment method; core facial expression representation refers to the quantitative representation that reflects the core emotional characteristics of the elderly, extracted from the micro-facial expression posture set of each assessment topic under the same assessment method; physiological auxiliary cognitive representation refers to the quantitative representation that is used to help reflect the cognitive state of the elderly, based on the core features of the physiological posture set and weighted by the dispersion of the core cognitive representation.

[0034] In this embodiment, the feature vectors and discrete data of each core representation are filled into the matrix according to the row and column correspondence to form a standardized second multidimensional evaluation matrix. Each evaluation method corresponds to a second multidimensional evaluation matrix.

[0035] The first nonlinear cognitive bias refers to the nonlinear deviation between the micro-behavioral characteristics of the target elderly under the same first multidimensional assessment matrix and the baseline micro-behavioral characteristics of the elderly healthy population. It is a quantitative representation of psychological bias at the assessment method level.

[0036] The first complete cognitive layer refers to a layer that presents the nonlinear cognitive biases and related characteristics of the assessment methods for the target elderly in a visual form. That is, a spatial coordinate system is constructed and the bias values ​​are mapped to heat values. Combined with the bias correlation curves between assessment methods, a layer containing the quantitative values ​​of bias, distribution patterns and related characteristics is formed.

[0037] The second nonlinear cognitive bias refers to the nonlinear deviation between the multi-core representations of each assessment topic of the target elderly population and the baseline representations of healthy elderly populations under the same second multidimensional assessment matrix.

[0038] The second complete cognitive layer refers to a layer that presents the nonlinear cognitive biases and related characteristics of the target elderly at the assessment topic level in a visual form. That is, it constructs a topic-representation bias association network, combines the bias level labeling, and forms a layer that includes the deviation quantification value, representation stability, and topic association relationship.

[0039] The third complete cognitive layer is implemented based on the first and second complete cognitive layers.

[0040] The assessment weight refers to the weight value assigned to each assessment method based on the third complete cognitive layer, which reflects the contribution of the assessment method to the mental health assessment of the target elderly, and the sum of the weight values ​​of all assessment methods is 1.

[0041] The topic weight refers to the weight value assigned to each assessment topic based on the third complete cognitive layer, reflecting the contribution of the assessment topic to the mental health assessment of the target elderly, and the sum of the weight values ​​of all assessment topics is 1.

[0042] The psychological assessment and intervention model refers to a deep learning model trained on a multimodal dataset of elderly mental health, which outputs personalized psychological intervention plans based on weighted assessment vectors. It adopts a CNN+Transformer fusion model structure and is trained on the fusion model using multimodal assessment data of the elderly population and corresponding psychological intervention plans as training sets. The model input is the weighted assessment vector and multimodal representation features, and the output is a graded psychological intervention plan.

[0043] An intervention program refers to a personalized psychological intervention program that is adapted to the core psychological problems of the target elderly based on the results of the psychological health assessment. That is, according to the degree of deviation of the assessment results and the dimensions of the core problems, a corresponding graded intervention program is output, which includes three core types: emotional counseling, cognitive training, and professional medical referral.

[0044] In this embodiment, the additional result is a feature vector obtained by weighted fusion of the evaluation weight, the topic weight, and the evaluation vector of the corresponding evaluation method.

[0045] The beneficial effects of the above technical solution are as follows: By constructing a multi-module collaborative intelligent assessment and intervention system, it realizes the integrated application of four assessment methods: pure image, pure text, combined image and text, and voice. At the same time, by combining micro-behavioral analysis of the assessment process, it breaks through the limitations of existing single assessment methods and the single focus on assessment results. Through multi-dimensional matrix construction, nonlinear deviation analysis, and cognitive layer fusion, it achieves in-depth quantification and integration of multimodal information. Finally, based on the fusion results, it dynamically assigns weights and generates intervention plans through a dedicated model, which greatly improves the accuracy and comprehensiveness of mental health assessment for the elderly population. The generated intervention plans are more closely aligned with the core psychological problems of the elderly, effectively solving the problems of poor targeting and ineffective intervention in existing intervention plans.

[0046] This invention proposes an intelligent assessment and intervention system for the mental health of the elderly based on multimodal perception. The assessment module includes: The pure image assessment unit is used to acquire the drawn images of the target elderly based on each assessment theme in the pure image assessment method, and extract the brush stroke trajectory features, color distribution features, and composition complexity features of the drawn images to obtain the sub-vectors of the corresponding assessment theme under the pure image. The plain text assessment unit is used to obtain the input text content of the target elderly person based on each assessment topic in the plain text assessment method, and extract the semantic coherence features, emotional polarity features, and word frequency features of the input text content to obtain the sub-vector of the corresponding assessment topic under plain text. The image-text combined assessment unit is used to extract the image drawing features and text input features of the target elderly people based on each assessment topic in the image-text combined assessment method, and obtain the sub-vector of the corresponding assessment topic under the image-text combined method; The voice assessment unit is used to acquire the audio of the target elderly person's voice statement based on each assessment topic in the voice assessment method, and extract the prosodic features, speech rate features, and emotional intensity features of the voice statement audio to obtain the sub-vector of the corresponding assessment topic under the voice. The splicing unit is used to splice the sub-vectors of different assessment topics under the same assessment method to obtain the assessment vector of the corresponding assessment method.

[0047] The pen stroke trajectory feature refers to the movement trajectory features of the pen strokes when the target elderly person draws an image in a pure image assessment method. The Fourier transform algorithm is used to extract features such as the movement direction, trajectory length, curvature, and pen stroke pressure changes. For example, when an elderly person draws their later life, the pen stroke trajectory has small curvature and uniform pressure, reflecting their peaceful psychological state.

[0048] Color distribution characteristics refer to the types, distribution positions, and color saturation of colors used by the target elderly when drawing images in a pure image assessment method. The main color tone, number of colors, average saturation, and proportion of warm and cool colors are extracted through image color analysis algorithms. For example, if an elderly person draws an image with a warm color tone and moderate saturation, it reflects a high proportion of positive emotions.

[0049] Compositional complexity features refer to the characteristics of the number of elements, layout, and density of elements in an image when the target elderly person draws an image under a pure image assessment method. The number of elements, the uniformity of element distribution, and the proportion of white space in the image are extracted through image composition analysis algorithms. For example, if an elderly person draws an image with a moderate number of elements and a uniform layout, it reflects that their cognitive function is good.

[0050] Semantic coherence features refer to the logical connections and semantic coherence between sentences in the text content input by the target elderly in a pure text assessment method. The BERT model is used to extract features such as semantic similarity, sentence coherence logic, and topic consistency of the text. For example, when an elderly person answers about their mood in the past week, the text sentences are fluent and the topic is clear, and the semantic coherence feature value is high.

[0051] Emotional polarity features refer to the emotional tendencies expressed in the text content input by the target elderly in a pure text assessment method. Using the BERT model combined with an emotion dictionary, features such as the emotional tendency (positive / negative / neutral) and emotional intensity of the text are extracted. For example, if an elderly person's text content contains words such as happy and comfortable, the emotional polarity is positive and the intensity is high.

[0052] Lexical frequency features refer to the frequency of various words and the proportion of core words in the text content entered by the target elderly in the pure text assessment method. The core words, high-frequency word types, and the proportion of function words / content words in the text are extracted by word frequency statistics algorithm. For example, the high frequency of words such as family and companionship in the text of an elderly person reflects their psychological need for social companionship.

[0053] Image drawing features refer to various features of the target elderly when drawing images under the combined text and image assessment method. They are consistent with the image features under the pure image assessment method, including brush stroke trajectory features, color distribution features, and composition complexity features. The same feature extraction and noise correction algorithms as the pure image assessment unit are used to ensure the consistency of feature extraction.

[0054] Text input features refer to the various features of the target elderly when they input text in the combined text and image assessment method. These features are consistent with the text features in the pure text assessment method, including semantic coherence features, emotional polarity features, and word frequency features. The same feature extraction algorithm as the pure text assessment unit is used to ensure the consistency of feature extraction.

[0055] Prosodic features refer to the characteristics of speech such as pitch, volume, duration, and pause rhythm when the target elderly person makes a speech statement under the speech assessment method. The MFCC algorithm is used to extract features such as fundamental frequency, volume variation, pause interval, and stress position of speech. For example, if an elderly person's tone is steady and the volume is moderate when making a statement, it reflects that their emotional state is stable.

[0056] Speech rate characteristics refer to the number of words spoken per unit time and the frequency of pauses in sentences when the target elderly person makes a speech statement under the speech assessment method. The MFCC algorithm is used to extract features such as the average speech rate, the amplitude of speech rate fluctuation, and the duration of pauses between sentences. For example, if an elderly person's average speech rate is moderate and the amplitude of fluctuation is small, it reflects that their cognitive function is good.

[0057] Emotional intensity features refer to the intensity of emotions expressed by the target elderly person when making a speech in a speech assessment. The MFCC algorithm is used in combination with a speech emotion recognition model to extract the emotional intensity value of the speech and map it to the 0-1 range. The larger the value, the stronger the emotional expression.

[0058] The beneficial effects of the above technical solution are as follows: by subdividing the evaluation module into functional units, dedicated feature extraction and sub-vector generation units were designed for the four evaluation methods. Noise correction and speech enhancement measures were added to take into account the physiological characteristics of the elderly, ensuring the accuracy and adaptability of feature extraction. At the same time, the standardized splicing of sub-vectors was realized through the splicing unit, and a unified quantitative framework of evaluation vectors was constructed. This laid a structured and computable data foundation for subsequent multimodal information analysis and improved the quantitative accuracy of multimodal evaluation results.

[0059] This invention proposes an intelligent assessment and intervention system for the mental health of the elderly based on multimodal perception, wherein the micro-behavioral information includes all captured results for each assessment topic under different assessment methods; The columns of the first multidimensional evaluation matrix correspond to four evaluation methods: pure image, pure text, image and text combination, and voice. The rows correspond to cognitive behavior posture set, micro-expression posture set, and physiological posture set. The matrix elements are feature vectors of the corresponding posture set under the corresponding evaluation method.

[0060] In this embodiment, the capture results specifically include: Based on the pen pressure, drawing time and number of modifications captured during the image drawing process by the touch screen, a set of cognitive behavior postures under the pure image evaluation method is generated. Based on the keystroke frequency, deletion count, and input pause duration captured during the keyboard / touchscreen text input process, a set of cognitive behavioral postures is generated under the pure text evaluation method. It captures the switching time and operation sequence between image drawing and text input, and combines the captured micro-operation behaviors of image and text input to generate a cognitive behavior posture set under the image-text combined assessment method. It captures intonation changes, number of pauses, and stress distribution during speech presentations, and generates a set of cognitive behavioral postures under speech assessment methods. Simultaneously, based on the facial micro-expression frame sequence of the target elderly person captured by the camera for each assessment topic under each assessment method, the corresponding micro-expression posture set is obtained; Based on the heart rate fluctuations of the target elderly person captured by physiological sensors for each assessment topic under each assessment method, a corresponding set of physiological postures is obtained.

[0061] In this embodiment, the cognitive behavior pose set feature vector is generated based on temporal feature encoding, the micro-expression pose set is generated based on facial key point feature encoding, and all feature vectors are normalized using the Z-Score normalization algorithm.

[0062] In this embodiment, the feature vector refers to the fixed-dimensional vector data obtained by extracting, reducing, and normalizing the feature data of the cognitive behavior posture set, micro-expression posture set, and physiological posture set. It is a matrix element of the first multi-dimensional evaluation matrix. That is, the feature data of the three types of posture sets are converted into 128-dimensional feature vectors respectively, and the Z-Score normalization algorithm is used for normalization to eliminate the difference in dimensions and ensure the analyzability of the matrix elements.

[0063] The beneficial effects of the above technical solution are: it clearly defines the capture results of micro-behavioral information, and at the same time designs a fixed standardized structure for the first multi-dimensional evaluation matrix, transforming the scattered original micro-behavioral data into structured feature vectors and constructing a multi-dimensional matrix, which solves the problem of unorganized and unquantifiable micro-behavioral information in the existing technology. At the same time, through unified feature vector dimensions and normalization processing, it ensures the analyzability of matrix data.

[0064] This invention proposes an intelligent assessment and intervention system for the mental health of the elderly based on multimodal perception. The matrix construction module includes: The sequence determination unit is used to perform temporal segmentation and feature encoding on the cognitive behavior posture set, micro-expression posture set and physiological posture set corresponding to each assessment topic under the same assessment method, so as to obtain the temporal feature sequence corresponding to each posture set. Density clustering unit is used to perform adaptive density clustering on each time-series feature sequence to generate several clusters, and to calculate the cluster center vector and intra-cluster feature dispersion of each cluster. The correction unit is used to retrieve the pre-stored benchmark posture feature library under the same evaluation method, perform multi-dimensional similarity matching between each cluster and the benchmark cluster, select the benchmark cluster with the highest matching degree as the correction reference, and correct the current cluster center vector and the feature dispersion within the cluster based on the statistical characteristics of the benchmark cluster. The representation determination unit is used to determine the core cognitive representation of each assessment topic by taking the cluster center vector of the corrected cognitive behavior posture set as the core cognitive representation and the cluster center vector of the corrected micro-expression posture set as the core facial expression representation. The physiological auxiliary determination unit is used to generate physiological auxiliary cognitive representations based on the cluster center vector of the corrected physiological posture set, and by using the ratio of the intra-cluster feature dispersion of the core cognitive representation to the mean intra-cluster feature dispersion of all representations as a weighting coefficient. The matrix construction unit is used to construct a multimodal representation matrix under this assessment method, with each assessment topic as the column and core cognitive representation, core facial expression representation, and physiological auxiliary cognitive representation as the row. It is regarded as the second multidimensional assessment matrix. The matrix elements include the feature vector of the corresponding representation and the corrected intra-cluster feature dispersion.

[0065] In this embodiment, based on the assessment duration for each assessment topic for the target elderly person, the data is divided into segments of equal length, each with a duration of 5 seconds. If the last segment is less than 5 seconds, zeros are added to ensure that the duration of each segment is consistent. For example, if the assessment duration for an elderly person is 23 seconds, it is divided into 5 time segments, and the last segment is padded with 2 seconds of zero data.

[0066] Feature encoding refers to the process of quantizing and encoding the feature data of each time series segment, transforming unstructured time series data into structured feature-encoded data. It uses a combination of one-hot encoding and numerical encoding to encode the feature data of time series segments, generating segment feature vectors of fixed dimensions, with all segments having the same dimension of feature vectors.

[0067] A temporal feature sequence refers to the sequence data formed by arranging the feature encoding data of each temporal segment in the order of evaluation time. It concatenates the feature vectors of each temporal segment in chronological order to form a one-dimensional temporal feature sequence, and the sequence length is proportional to the number of temporal segments.

[0068] Adaptive density clustering refers to a density clustering method that adaptively adjusts clustering parameters based on the distribution characteristics of the data. Unlike density clustering with fixed parameters, it uses the DBSCAN algorithm and sets the neighborhood radius to a dynamic value based on the evaluation time. The longer the evaluation time, the larger the neighborhood radius. The minimum number of points is fixed at 10 to ensure the rationality and adaptability of the clustering results.

[0069] Clustering clusters refer to several data clusters into which a temporal feature sequence is divided by adaptive density clustering. The data within each cluster have a high degree of similarity. The DBSCAN algorithm divides the feature vectors of temporal segments with high similarity into the same cluster and removes noise points. For example, the temporal feature sequence of a certain cognitive behavior posture set is divided into 3 clusters and 2 noise points.

[0070] The cluster center vector is the mean vector of all feature vectors in each cluster. It is the core feature representation of the cluster. It is formed by calculating the mean of all feature vectors in each cluster to form a cluster center vector with the same dimension as the feature vectors in the cluster. For example, if a cluster contains 10 128-dimensional feature vectors, the 128-dimensional cluster center vector is obtained by calculating the mean of each dimension.

[0071] Intra-cluster feature dispersion refers to the degree of dispersion of all feature vectors within each cluster relative to the cluster center vector. It is an indicator reflecting the consistency of data within the cluster. The intra-cluster feature dispersion is calculated using the standard deviation method. The Euclidean distance between all feature vectors within the cluster and the cluster center vector is calculated, and the mean of the distances is taken as the intra-cluster feature dispersion. The smaller the value, the higher the consistency of data within the cluster.

[0072] The baseline posture feature library refers to a pre-built database that stores three types of posture feature data of healthy elderly people under the same assessment method. It serves as a reference benchmark for cluster correction. This database is built based on the micro-behavioral feature data of more than 5,000 healthy elderly people and is stored in three age groups: 60-70 years old, 70-80 years old, and 80 years old and above. The data includes cluster data under various assessment methods and assessment themes.

[0073] The benchmark cluster refers to the cluster in the benchmark posture feature library that has the highest similarity to the cluster of the target elderly person. It is a reference standard for cluster correction. The similarity between the cluster of the target elderly person and all clusters in the benchmark library is calculated by the cosine similarity algorithm, and the benchmark cluster with the highest similarity is selected as the correction reference.

[0074] Multi-dimensional similarity matching refers to a matching method that calculates the similarity between two clusters from multiple dimensions, including feature vector dimension, temporal distribution dimension, and dispersion dimension. It calculates the cosine similarity of the clusters in three dimensions: cluster center vector, temporal feature distribution, and intra-cluster feature dispersion, and takes the mean of the three similarity dimensions as the overall similarity to ensure the accuracy of the matching results.

[0075] The beneficial effects of the above technical solution are as follows: by subdividing the matrix construction module into functional units, a coherent feature processing flow of temporal partitioning, adaptive clustering, benchmark correction, and core representation extraction is designed. A benchmark posture feature library is introduced for correction based on individual differences in the elderly population. At the same time, physiological features are combined with core cognitive representations to generate physiological-assisted cognitive representations, which better suits the characteristics of the high correlation between physiology and cognition in the elderly. Finally, the constructed second multidimensional assessment matrix achieves accurate correlation between assessment topics and multiple core representations, improves the accuracy and stability of core representation extraction, and provides a high-quality quantitative data foundation for subsequent nonlinear deviation analysis at the assessment topic level.

[0076] This invention proposes an intelligent assessment and intervention system for the mental health of the elderly based on multimodal perception. The layer construction module includes: The normalization processing unit is used to normalize the feature vectors corresponding to the cognitive behavior posture set, micro-expression posture set, and physiological posture set in the same first multi-dimensional evaluation matrix. The dimensionality reduction processing unit is used to extract dimensionality from the normalized matrix and retain the core feature components related to cognitive biases in the elderly population. The model building unit is used to introduce temporal correlation weights, combine the assessment duration and response delay characteristics of the target elderly under each assessment method, and perform weighted correction on the core feature components corresponding to each assessment method to construct a nonlinear deviation quantification model for each assessment method. The nonlinear deviation quantification model is an SVM nonlinear deviation model based on radial basis function (RBF). The first deviation determination unit is used to calculate the nonlinear deviation value between each assessment method and the pre-stored benchmark matrix of healthy elderly people based on the nonlinear deviation quantification model, which is regarded as the first nonlinear cognitive deviation. The first layer construction unit is used to construct a spatial coordinate system with four assessment methods as the horizontal dimension and three types of posture sets as the vertical dimension. The first nonlinear cognitive deviation value of each assessment method is mapped to the heat value in the spatial coordinate system to generate a deviation heat distribution map. At the same time, the deviation correlation curve between assessment methods is embedded. The deviation heat distribution map and the deviation correlation curve are integrated to form a first complete cognitive layer containing the first deviation quantification value, the deviation distribution law and the correlation characteristics of assessment methods. The depth of the heat value corresponds to the magnitude of the deviation, and the slope of the correlation curve corresponds to the deviation change trend.

[0077] In this embodiment, the normalization processing unit is equipped with the Z-Score normalization algorithm to normalize all feature vectors in the matrix and eliminate dimensional differences; the dimensionality reduction processing unit is equipped with the PCA dimensionality reduction algorithm to extract core feature components related to cognitive biases in the elderly based on the variance contribution rate.

[0078] Core feature components refer to the main feature components that can reflect the cognitive biases of the elderly population, extracted through dimensionality reduction. That is, the principal components with a variance contribution rate of ≥90% are retained as core feature components to ensure that the core feature components can retain the main information of the original data. For example, the 128-dimensional feature vector is reduced to 20-dimensional core feature components with a variance contribution rate of 92%.

[0079] The model building unit is equipped with a temporal correlation weight calculation algorithm and an SVM model training interface. It calculates the temporal correlation weights and builds an SVM nonlinear bias quantization model based on the RBF kernel function.

[0080] The calculation process of temporal correlation weights: Initial correlation weight = 0.6 × (benchmark evaluation time / actual evaluation time) + 0.4 × (benchmark response delay / actual response delay); Temporal correlation weight = corresponding initial correlation weight / maximum value among all initial correlation weights.

[0081] In this embodiment, the start and end times of the assessment are recorded, and the difference between the two is the assessment duration; the time of the assessment task issuance and the time of the elderly person's first operation are recorded, and the difference between the two is the response delay. The average response delay of all assessment topics under a certain assessment method is taken as the response delay feature value of that assessment method.

[0082] In this embodiment, an SVM model based on radial basis function (RBF) is used. The core feature components are used as inputs and the deviation values ​​in the 0-1 interval are used as outputs. The model converges after being trained on data from elderly healthy people and abnormal people. The RBF kernel function parameter is 0.1, the penalty factor is 10, the loss function is the squared loss function, and the maximum number of iterations is 1000.

[0083] In this embodiment, the input of the SVM model based on radial basis function (RBF) is the core feature components of a certain assessment method after dimensionality reduction by PCA, and the output is the nonlinear deviation value between this assessment method and the benchmark of the elderly healthy population. The specific process is as follows: Weighted correction of the core feature vector of the elderly: ,in, These are the core feature components after PCA dimensionality reduction under a certain evaluation method; For time-series correlation weights; The mean vector of the core feature components of the baseline healthy elderly population of the same age group; RBF kernel function ,in, These are the weighted and corrected feature components; For health baseline characteristic components; This is the width parameter of the RBF kernel function, and it is adapted to the elderly scenario. , The standard deviation of the baseline characteristic component is 0.05 to 0.2. RBF-SVM Bias Quantization Decision Function ,in, The Lagrange multipliers for SVM are obtained after training with samples from elderly healthy / abnormal populations. ; For sample labeling, healthy sample Abnormal samples b represents the SVM bias term; This is the deviation quantification coefficient, with an empirical value of 1.0~2.0, to enhance the quantitative sensitivity of small deviations in the elderly; The final calculation formula for the first nonlinear cognitive bias ,in, The output value of the baseline decision function for healthy individuals is approximately 0.5. To The result was obtained after normalization.

[0084] like Then set it to 0, if Then set it to 1 to ensure .

[0085] The benchmark matrix for healthy elderly populations refers to a pre-constructed benchmark matrix that stores the feature data of the first multidimensional assessment matrix under various assessment methods for healthy elderly populations. This matrix is ​​constructed based on the micro-behavioral feature data of healthy elderly people, stored in layers according to age, and matched with the age layers of the target elderly population.

[0086] The deviation heat map is a distribution map generated after mapping the first nonlinear cognitive deviation value to a heat map value. It is a visual representation of the deviation value. In the spatial coordinate system, the intersection of each assessment method with the three types of posture sets is taken as the heat map point. The larger the deviation value, the darker the color of the heat map point. For example, the heat map point with a deviation value of 0.8 is dark red, and the heat map point with a deviation value of 0.2 is light pink.

[0087] The deviation correlation curve is a curve used to characterize the trend of deviation changes between different assessment methods. It is plotted with the assessment method on the x-axis and the deviation value on the y-axis, connecting the deviation values ​​of each assessment method. The slope of the curve reflects the trend of deviation changes. A positive slope indicates that the deviation value is increasing, and a negative slope indicates that the deviation value is decreasing.

[0088] The beneficial effects of the above technical solution are as follows: by subdividing the first layer construction function of the layer construction module into units, temporal correlation weights adapted to the cognitive characteristics of the elderly are introduced. PCA dimensionality reduction combined with SVM nonlinear model is used to achieve accurate quantification of the deviation at the assessment method level. At the same time, the first complete cognitive layer is constructed by visualizing the heat map and correlation curve. This not only realizes the quantitative presentation of the deviation value, but also explores the deviation correlation between assessment methods, making the deviation analysis results more comprehensive and intuitive, and in line with the cognitive physiological characteristics of the elderly population. This provides clear and quantitative psychological deviation data at the assessment method level for subsequent layer fusion.

[0089] This invention proposes an intelligent assessment and intervention system for the mental health of the elderly based on multimodal perception. The layer construction module further includes: The coefficient calculation unit is used to extract the feature vectors and corrected intra-cluster feature dispersion of the core cognitive representation, core facial expression representation, and physiological auxiliary cognitive representation in the same second multi-dimensional assessment matrix, calculate the dispersion variation coefficient of each representation under each assessment topic, and quantify the stability of the representation. The second deviation determination unit is used to input the multi-representation features corresponding to each assessment topic into a pre-constructed multi-representation fusion nonlinear deviation analysis model, compare it with the pre-stored benchmark representation data of the elderly healthy population under the same assessment method and assessment topic, calculate the nonlinear deviation quantification value of each assessment topic, and regard it as the second nonlinear cognitive deviation. The nonlinear deviation analysis model is a multi-representation fusion model based on the attention mechanism. Network building units are used to construct a topic-representation bias association network with assessment topics as nodes, second nonlinear cognitive bias values ​​as node weights, and representation stability as node association strength. The second layer construction unit is used to visualize and label the deviation levels of each assessment topic, integrate the correlation network and deviation level labeling to form a second complete cognitive layer containing the second deviation quantification value, the stability of the representation, and the deviation correlation between topics. In this layer, the node size corresponds to the deviation weight, and the thickness of the connection between nodes corresponds to the strength of the representation correlation.

[0090] In this embodiment, when the magnitude of the cluster center vector is not 0, the coefficient of variation of the dispersion = intra-cluster feature dispersion / magnitude of the cluster center vector; When the magnitude of the cluster center vector is 0, the coefficient of variation of the dispersion is 0.

[0091] The stability of representation refers to the degree of stability of core cognitive representation, core facial expression representation, and physiological auxiliary cognitive representation during the assessment process. It is judged according to the preset threshold of the coefficient of variation of dispersion. A coefficient of variation < 0.3 indicates that the representation is stable, 0.3 ≤ coefficient of variation ≤ 0.7 indicates that the representation is average, and a coefficient of variation > 0.7 indicates that the representation is unstable. For example, if the coefficient of variation of a certain core facial expression representation is 0.25, then its representation is judged to be stable.

[0092] In this embodiment, the input of the multi-representation fusion nonlinear deviation analysis model is three core representations (C / E / P) under a certain assessment topic, and the output is the nonlinear deviation value between the assessment topic and the benchmark of the healthy elderly population. Among them, the core cognitive representation C, the core facial expression representation E, and the physiological auxiliary cognitive representation P, the specific process is as follows: Dual-weight attention weight calculation: ,in, The coefficient of variation characterizes the dispersion of j; The maximum CV value among the three types of representations; This represents the inherent importance coefficient of the characteristics of aging. , , ; Attention weight normalization: , , To normalize the attention weights, and ; Multi-representation fusion features with correlation coefficients: , There are three types of core representations. The inter-modal correlation coefficient matrix is ​​constructed based on the Pearson correlation coefficient of the multimodal representation in the patent, with diagonal elements set to 1 and off-diagonal elements representing the inter-modal correlation coefficients. The resulting multi-representation feature vector; The final formula for calculating the second nonlinear cognitive bias: , It is the minimum value, and takes the value of ; This is a baseline fusion feature vector for healthy elderly individuals under the same assessment method / topic. The deviation quantification coefficient has an empirical value of 2.0 to 3.0. Validated by a sample of 1100 elderly individuals, this range can effectively improve the discrimination accuracy of moderate deviation in the elderly, with a discrimination accuracy of ≥88%. like Then set it to 0, if Then set it to 1.

[0093] The baseline representation data refers to the pre-constructed feature data of three core representations of healthy elderly people under the same assessment method and assessment theme. This data is stored in layers according to assessment method, assessment theme and age group, and is matched with the assessment conditions and age group of the target elderly people.

[0094] The topic-representation bias association network is a visual network constructed with assessment topics as nodes, core representations as association edges, second nonlinear cognitive bias values ​​as node weights, and the stability of representations as the node association strength. It uses the Gephi algorithm to construct the network and adopts a force-oriented layout to ensure the readability of the network.

[0095] The deviation level refers to the level reflecting the degree of psychological deviation at the assessment topic level, which is divided according to the second nonlinear cognitive deviation value. It maps the deviation value to the 0-1 range and divides it into four levels according to the size of the deviation value: no deviation (0-0.2), slight deviation (0.2-0.5), moderate deviation (0.5-0.8), and severe deviation (0.8-1.0). Different colors are used to mark each level.

[0096] The beneficial effects of the above technical solution are as follows: by subdividing the second layer construction function of the layer construction module into units, the stability of the representation is accurately quantified by using the coefficient of variation of discreteness. The multi-representation fusion model based on the attention mechanism improves the accuracy of the deviation analysis at the evaluation topic level. At the same time, the constructed topic-representation deviation association network effectively determines the deviation value, representation stability, and the deviation association relationship between topics, providing a reliable foundation for constructing the second complete cognitive layer.

[0097] This invention proposes an intelligent assessment and intervention system for the mental health of the elderly based on multimodal perception. The weight assignment module includes: The feature extraction unit is used to extract the first deviation quantification features, the correlation features between assessment methods and the posture dimension distribution features corresponding to the assessment methods in the first complete cognitive layer, and to extract the second deviation quantification features, the representation stability features and the deviation correlation features between the assessment topics in the second complete cognitive layer. Combined with the adaptive calibration function constructed by the cognitive physiological characteristics of the elderly population, the heterogeneous interference in the features of the two types of layers is eliminated, and a calibrated feature set is generated. The vector generation unit is used to quantify the correlation strength of the features of the two types of layers based on the temporal response characteristics of the evaluation methods in the first complete cognitive layer and the coefficient of variation of the discreteness of the core representations in the second complete cognitive layer, and dynamically generate a fusion weight vector based on the correlation strength. The initial fusion map generation unit is used to integrate the calibrated feature set across dimensions according to the fusion weight vector, and optimize it based on the feature mutual information entropy to strengthen the intrinsic relationship between the evaluation method, evaluation theme and multimodal representation, and generate the initial fusion map. The graph correction unit is used to introduce the fusion benchmark dataset of elderly mental health, correct the initial fusion graph, and construct a three-dimensional cognitive coordinate system of assessment method, assessment topic, and multimodal representation based on the corrected initial fusion graph. The fused deviation quantification data, feature correlation strength, and representation stability parameters are mapped to the three-dimensional cognitive coordinate system to generate a three-dimensional dynamic fusion cognitive graph. Elderly cognitive deviation warning indicators and feature importance classification labels are embedded and integrated to form a third complete cognitive layer.

[0098] In this embodiment, the elderly mental health integrated benchmark dataset contains multimodal assessment data of more than 1,000 elderly people, which are divided into three layers according to mental state: healthy, mild mood disorder, and moderate cognitive decline. The data dimensions cover three major categories: assessment methods, themes, and multimodal representations.

[0099] In this embodiment, the first deviation quantification feature refers to the quantified numerical feature of the first nonlinear cognitive deviation corresponding to each assessment method extracted from the first complete cognitive layer. It is the core feature reflecting the degree of deviation at the assessment method level. That is, the deviation heat value quantification data of each assessment method in the first complete cognitive layer is extracted and mapped to the 0-1 interval as the first deviation quantification feature. For example, the deviation heat value of the pure image assessment method is 0.65, and its first deviation quantification feature value is 0.65.

[0100] The correlation features between assessment methods refer to the correlation features of deviation changes between different assessment methods extracted from the first complete cognitive layer. These features reflect the correlation between the deviations of assessment methods. Specifically, quantitative data such as the slope and correlation coefficient of the deviation correlation curve in the first complete cognitive layer are extracted as correlation features between assessment methods. For example, the slope of the deviation correlation curve between pure image and pure text assessment methods is 0.2, and the correlation coefficient is 0.8, reflecting that the deviations of the two are positively correlated and the trend of change is gradual.

[0101] The posture dimension distribution feature refers to the deviation distribution features of the three types of posture sets—cognitive behavior, micro-expression, and physiological—extracted from the first complete cognitive layer under various assessment methods. It is a feature that reflects the deviation distribution law at the posture set level. That is, it extracts quantitative data such as the mean and variance of the heat values ​​corresponding to the three types of posture sets in the first complete cognitive layer as posture dimension distribution features. For example, the mean heat value of the physiological posture set under the four assessment methods is 0.4 and the variance is 0.08, which reflects that its deviation distribution is relatively uniform.

[0102] The second deviation quantification feature refers to the quantified numerical feature of the second nonlinear cognitive deviation corresponding to each assessment topic extracted from the second complete cognitive layer. It is the core feature reflecting the degree of deviation at the assessment topic level. That is, the weight value of each assessment topic node in the second complete cognitive layer is extracted as the second deviation quantification feature. For example, the node weight of the emotional state topic is 0.7, and its second deviation quantification feature value is 0.7.

[0103] The representation stability feature refers to the quantitative stability feature of each core representation extracted from the second complete cognitive layer under each assessment topic. It is a feature that reflects the reliability of the core representation. That is, the coarseness quantification data of the connection between each node in the second complete cognitive layer is extracted and mapped to the 0-1 interval as the representation stability feature. For example, the coarseness quantification value of the connection between the core cognitive representation under the social willingness topic is 0.85, and its representation stability feature value is 0.85.

[0104] Inter-topic deviation correlation features refer to the correlation features of deviations between different assessment topics extracted from the second complete cognitive layer. They are features that reflect the correlation between deviations between assessment topics. Specifically, the correlation coefficients between nodes in the topic-representation deviation correlation network in the second complete cognitive layer are extracted as inter-topic deviation correlation features. For example, the correlation coefficient between nodes of the topics of cognitive function and life satisfaction is 0.75, which reflects that the deviations of the two are strongly positively correlated.

[0105] The adaptive calibration function refers to a function constructed based on the cognitive and physiological characteristics of the elderly population to eliminate the heterogeneous interference of features between the first and second complete cognitive layers. It is the core function for achieving heterogeneous feature fusion. A nonlinear correction function based on the cognitive and physiological characteristics of the elderly is constructed, and its function expression is as follows: ,in These are the original eigenvalues. , , These are correction parameters set based on the age and cognitive function scores of the elderly population, where the parameters are... , , Based on age group and cognitive function scoring, the maximum score is 30 points, as shown in Table 1: Table 1. Comparison of Correction Parameters for Different Age Groups In this embodiment, heterogeneous interference refers to the feature fusion interference caused by the different dimensions, representation methods, and analysis levels of the features of the first and second complete cognitive layers. It is the main factor affecting the layer fusion effect. That is, by using the adaptability calibration function, features of different dimensions and representation methods are mapped to the same feature space to eliminate the interference caused by the differences in the dimensions and dimensions between heterogeneous features. For example, the two-dimensional spatial features of the first layer and the network node features of the second layer are uniformly mapped to a 64-dimensional feature space.

[0106] The calibrated feature set refers to the feature set with unified dimensions and unified representation method obtained after the features of the first and second complete cognitive layers are processed by the adaptive calibration function to eliminate heterogeneous interference. In other words, after all the extracted features are processed by the calibration function, they are classified and integrated according to the evaluation method, evaluation theme and multimodal representation to form a standardized calibrated feature set. The dimension of the feature set is unified to a preset fixed value.

[0107] Temporal response features refer to the temporal-related features such as assessment duration and response delay of the target elderly under each assessment method extracted from the first complete cognitive layer. They are features that reflect the cognitive response status of the elderly during the assessment process. That is, the temporal correlation weight values ​​of each assessment method in the construction process of the first complete cognitive layer are extracted as temporal response features. For example, the temporal correlation weight of the voice assessment method is 0.9, and its temporal response feature value is 0.9.

[0108] Feature association strength refers to the degree of correlation between features in the first and second complete cognitive layers. It is the core basis for dynamically generating the fusion weight vector. It is calculated by using the mutual information method to calculate the mutual information value between the features of the two types of layers. The mutual information value is normalized to the 0-1 range as the feature association strength. The larger the mutual information value, the higher the feature association strength. For example, the mutual information value between the temporal response feature of the first layer and the representation stability feature of the second layer is 0.8. After normalization, the feature association strength is 0.8.

[0109] The fusion weight vector refers to the vector data dynamically generated based on the feature association strength, used to perform weighted fusion of the calibrated feature set. It is the core basis for realizing feature weighted integration. That is, the analytic hierarchy process (AHP) is used to take the association strength of each feature as the judgment matrix. After hierarchical single sorting and consistency test, a fusion weight vector with the same dimension as the calibrated feature set is generated. The value of each element in the vector is the fusion weight of the corresponding feature, and the sum of all element values ​​is 1.

[0110] Cross-dimensional interweaving and integration refers to the cross-fusion and integration of features from different dimensions of the calibrated feature set according to the fusion weight vector. It is a process of achieving deep integration of assessment methods, assessment themes, and multimodal representation features. In other words, the features of assessment method dimension, assessment theme dimension, and multimodal representation dimension are weighted and superimposed according to the fusion weight vector to achieve the interweaving and fusion of features from different dimensions. For example, pure image assessment method features, emotional state theme features, and core expression representation features are weighted and superimposed according to their corresponding weights to form fused features.

[0111] Feature mutual information entropy optimization refers to the feature optimization process that strengthens the intrinsic relationship between the three elements—evaluation method, evaluation topic, and multimodal representation—by maximizing the mutual information entropy among them. This involves calculating the mutual information entropy among the three elements, adjusting the weights of the fused features using the gradient ascent method, maximizing the mutual information entropy, thereby strengthening the intrinsic relationship among them and improving the effectiveness of the fused features.

[0112] The initial fusion map refers to the initial map that reflects the relationship between the evaluation method, evaluation theme, and multimodal representation after the calibrated feature set is integrated through cross-dimensional interweaving and feature mutual information entropy optimization. In other words, the optimized fusion features are visualized and mapped according to the relationship between the three to form an initial fusion map containing feature quantification data and correlation strength. The map is a two-dimensional visualization map.

[0113] The Elderly Mental Health Integration Benchmark Dataset refers to a pre-constructed benchmark dataset for multimodal assessment of mental health in the elderly population, used to correct the initial integration map. This dataset contains multimodal assessment data from over 1,000 elderly individuals, categorized into three levels based on mental state: healthy, mild mood disorder, and moderate cognitive decline. The data dimensions cover three major categories: assessment methods, assessment topics, and multimodal representations, and are stored hierarchically by age group, allowing for the matching of corresponding benchmark data based on the age and mental state of the target elderly individual.

[0114] The three-dimensional cognitive coordinate system refers to a three-dimensional spatial coordinate system constructed with assessment method, assessment topic, and multimodal representation as the three coordinate axes. It is the foundation for realizing the three-dimensional visualization of fusion features. Specifically, a three-dimensional rectangular coordinate system is constructed with assessment method as the X-axis, assessment topic as the Y-axis, and multimodal representation as the Z-axis. The scales of the three coordinate axes are normalized according to the corresponding quantitative features and mapped to the 0-1 interval.

[0115] A three-dimensional dynamic fusion cognitive map refers to a three-dimensional dynamic visualization map generated by mapping the fused deviation quantification data, feature correlation strength, and representation stability parameters to a three-dimensional cognitive coordinate system. In other words, various quantification data are used as points, lines, and surfaces in the three-dimensional coordinate system to realize the three-dimensional dynamic presentation of the data. It also supports dynamic updates based on changes in feature parameters. For example, when the deviation value changes, the color and size of the corresponding point in the coordinate system will change synchronously.

[0116] The early warning sign for cognitive deviation in the elderly refers to a visual sign embedded in a three-dimensional dynamic fusion cognitive map to warn of significant cognitive deviations in the target elderly. Specifically, a threshold value for deviation quantification is set (e.g., deviation value ≥ 0.7). When the deviation value corresponding to a certain feature reaches the threshold, a red early warning sign is embedded at the corresponding position in the three-dimensional coordinate system. The sign is in the shape of a five-pointed star. At the same time, the assessment method, assessment theme and representation type corresponding to the deviation are marked. For example, if the deviation value of the plain text assessment method, cognitive function theme and core cognitive representation is 0.75, a red five-pointed star early warning sign is embedded at the corresponding position and relevant information is marked.

[0117] Feature importance grading refers to the visual annotation embedded in the 3D dynamic fusion cognitive map used to grade the importance of fused features. Specifically, feature importance is divided into three levels—Level 1, Level 2, and Level 3—based on feature association strength and representation stability, and is marked with yellow, blue, and green, respectively. Level 1 features are core features with association strength and stability ≥ 0.8; Level 2 features are important features with association strength and stability between 0.5 and 0.8; and Level 3 features are general features with association strength and stability ≤ 0.5. For example, a feature with an association strength of 0.85 and a representation stability of 0.9 is marked as Level 1 importance in yellow.

[0118] The third complete cognitive layer refers to the three-dimensional dynamic cognitive layer formed by integrating the corrected three-dimensional dynamic fusion cognitive map with early warning indicators of cognitive deviation in the elderly and the importance classification of features. It contains multi-dimensional information on assessment methods, assessment topics, and multimodal representations. It is a comprehensive quantitative and visual representation of the mental health status of the target elderly. That is, the three-dimensional dynamic fusion cognitive map is used as the basis, and early warning indicators and importance classifications are embedded in the map according to their corresponding positions to form a third complete cognitive layer that integrates data quantification, visualization, early warning prompts, and feature classification.

[0119] The beneficial effects of the above technical solution are as follows: By subdividing the layer fusion function of the weight assignment module into units, an adaptive calibration function that fits the cognitive physiological characteristics of the elderly population is constructed, effectively eliminating the heterogeneous feature interference between the first and second complete cognitive layers. The mutual information method and the hierarchical analysis method are used to realize the dynamic generation of fusion weights. The feature mutual information entropy optimization strengthens the intrinsic relationship between the assessment method, assessment topic, and multimodal representation. At the same time, the elderly mental health fusion benchmark dataset is introduced to correct the initial map. The final constructed third complete cognitive layer realizes the three-dimensional dynamic visualization of multi-dimensional features and also incorporates deviation warning and feature grading functions, providing a comprehensive, accurate, and visualized overall feature basis for the subsequent weight calculation of assessment methods and assessment topics.

[0120] This invention proposes an intelligent assessment and intervention system for the mental health of the elderly based on multimodal perception. The weight assignment module further includes: The parameter extraction unit is used to extract the feature correlation strength corresponding to each assessment method from the third complete cognitive layer. Feature credibility and the contribution of deviation Extract the representation stability corresponding to each assessment topic. Thematic relevance and deviation influence coefficient Where i=1,2,3,4 corresponds to four assessment methods, and j=1,2,...,n corresponds to various assessment topics. All extracted parameters have been normalized to the range of 0 to 1. The first calculation unit is used to calculate the evaluation weight for each evaluation method. ; The second calculation unit is used to calculate the topic weight of each assessment topic. .

[0121] In this embodiment, the feature correlation strength is calculated using the Pearson correlation coefficient, the stability of the representation is calculated using the coefficient of variation of dispersion, the deviation contribution is calculated as the deviation value of a single evaluation method / the sum of the deviation values ​​of all evaluation methods, the deviation influence coefficient is calculated as the deviation value of a single evaluation topic / the sum of the deviation values ​​of all evaluation topics, and the feature confidence is calculated as 1 - CV / 2.

[0122] In this embodiment, the n assessment topics are the four preset core assessment topics, j=1,2,3,4 corresponding to emotional state, cognitive function, social willingness, and life satisfaction, respectively.

[0123] The beneficial effects of the above technical solution are as follows: the calculation formulas for assessment weights and topic weights comprehensively consider multiple dimensions such as feature correlation strength, credibility, and bias contribution, realizing dynamic and personalized weight calculation. Unlike the fixed weight assignment method of existing technologies, the weight values ​​are more in line with the actual mental health characteristics of the target elderly. The additional weight assignment provides more accurate and targeted input data for subsequent psychological assessment and intervention models, improves the adaptability of the model's output intervention plan, and solves the problem of one-sided and lack of personalization in the weight assignment of existing technologies.

[0124] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.

Claims

1. A multimodal perception-based intelligent assessment and intervention system for the mental health of the elderly, characterized in that, include: The assessment module is used to assess the target elderly person sequentially based on assessment files related to the pure image assessment method, the pure text assessment method, the combined image and text assessment method, and the voice assessment method, and to obtain an assessment vector for each assessment method. The assessment vector contains a sub-vector of each assessment topic, and each assessment vector contains the same assessment topic. The micro-behavior analysis module is used to extract the cognitive behavior posture set, micro-expression posture set, and physiological posture set of the target elderly person for each assessment topic under each assessment method based on the captured micro-behavior information during the assessment process of each assessment method, and to construct the first multi-dimensional assessment matrix of the target elderly person based on the same assessment topic. The matrix construction module is used to extract core cognitive representations, core facial expression representations, and physiological auxiliary cognitive representations from the cognitive behavior posture set, micro-expression posture set, and physiological posture set of each assessment topic under the same assessment method, and construct the second multidimensional assessment matrix of the target elderly based on the same assessment method. The second multidimensional assessment matrix contains multiple representations of different assessment topics under the same assessment method, and each row vector corresponds to one representation. The layer construction module is used to analyze the first nonlinear cognitive bias of the target elderly person under each assessment method in the same first multidimensional assessment matrix and construct a first complete cognitive layer. At the same time, it analyzes the second nonlinear cognitive bias of the target elderly person under each assessment topic in the same second multidimensional assessment matrix and constructs a second complete cognitive layer. The weight assignment module is used to merge the first complete cognitive layer and the second complete cognitive layer to obtain a third complete cognitive layer, and to assign assessment weights to each assessment method and theme weights to each assessment topic based on the third complete cognitive layer, and to attach them to each assessment method and the corresponding assessment vector. The intervention module is used to input additional results into a pre-trained psychological assessment intervention model and output a psychological intervention plan adapted to the target elderly person.

2. The intelligent assessment and intervention system for the mental health of the elderly based on multimodal perception according to claim 1, characterized in that, The evaluation module includes: The pure image assessment unit is used to acquire the drawn images of the target elderly based on each assessment theme in the pure image assessment method, and extract the brush stroke trajectory features, color distribution features, and composition complexity features of the drawn images to obtain the sub-vectors of the corresponding assessment theme under the pure image. The plain text assessment unit is used to obtain the input text content of the target elderly person based on each assessment topic in the plain text assessment method, and extract the semantic coherence features, emotional polarity features, and word frequency features of the input text content to obtain the sub-vector of the corresponding assessment topic under plain text. The image-text combined assessment unit is used to extract the image drawing features and text input features of the target elderly people based on each assessment topic in the image-text combined assessment method, and obtain the sub-vector of the corresponding assessment topic under the image-text combined method; The voice assessment unit is used to acquire the audio of the target elderly person's voice statement based on each assessment topic in the voice assessment method, and extract the prosodic features, speech rate features, and emotional intensity features of the voice statement audio to obtain the sub-vector of the corresponding assessment topic under the voice. The splicing unit is used to splice the sub-vectors of different assessment topics under the same assessment method to obtain the assessment vector of the corresponding assessment method.

3. The intelligent assessment and intervention system for the mental health of the elderly based on multimodal perception according to claim 1, characterized in that, The micro-behavioral information includes all capture results for each assessment topic under different assessment methods; The columns of the first multidimensional evaluation matrix correspond to four evaluation methods: pure image, pure text, image and text combination, and voice. The rows correspond to cognitive behavior posture set, micro-expression posture set, and physiological posture set. The matrix elements are feature vectors of the corresponding posture set under the corresponding evaluation method.

4. The intelligent assessment and intervention system for the mental health of the elderly based on multimodal perception according to claim 1, characterized in that, The matrix construction module includes: The sequence determination unit is used to perform temporal segmentation and feature encoding on the cognitive behavior posture set, micro-expression posture set and physiological posture set corresponding to each assessment topic under the same assessment method, so as to obtain the temporal feature sequence corresponding to each posture set. Density clustering unit is used to perform adaptive density clustering on each time-series feature sequence to generate several clusters, and to calculate the cluster center vector and intra-cluster feature dispersion of each cluster. The correction unit is used to retrieve the pre-stored benchmark posture feature library under the same evaluation method, perform multi-dimensional similarity matching between each cluster and the benchmark cluster, select the benchmark cluster with the highest matching degree as the correction reference, and correct the current cluster center vector and the feature dispersion within the cluster based on the statistical characteristics of the benchmark cluster. The representation determination unit is used to determine the core cognitive representation of each assessment topic by taking the cluster center vector of the corrected cognitive behavior posture set as the core cognitive representation and the cluster center vector of the corrected micro-expression posture set as the core facial expression representation. The physiological auxiliary determination unit is used to generate physiological auxiliary cognitive representations based on the cluster center vector of the corrected physiological posture set, and by using the ratio of the intra-cluster feature dispersion of the core cognitive representation to the mean intra-cluster feature dispersion of all representations as a weighting coefficient. The matrix construction unit is used to construct a multimodal representation matrix under this assessment method, with each assessment topic as the column and core cognitive representation, core facial expression representation, and physiological auxiliary cognitive representation as the row. It is regarded as the second multidimensional assessment matrix. The matrix elements include the feature vector of the corresponding representation and the corrected intra-cluster feature dispersion.

5. The intelligent assessment and intervention system for the mental health of the elderly based on multimodal perception according to claim 3, characterized in that, The layer construction module includes: The normalization processing unit is used to normalize the feature vectors corresponding to the cognitive behavior posture set, micro-expression posture set, and physiological posture set in the same first multi-dimensional evaluation matrix. The dimensionality reduction processing unit is used to extract dimensionality from the normalized matrix and retain the core feature components related to cognitive biases in the elderly population. The model building unit is used to introduce temporal correlation weights, combine the assessment duration and response delay characteristics of the target elderly under each assessment method, and perform weighted correction on the core feature components corresponding to each assessment method to construct a nonlinear deviation quantification model for each assessment method. The nonlinear deviation quantification model is an SVM nonlinear deviation model based on radial basis function (RBF). The first deviation determination unit is used to calculate the nonlinear deviation value between each assessment method and the pre-stored benchmark matrix of healthy elderly people based on the nonlinear deviation quantification model, which is regarded as the first nonlinear cognitive deviation. The first layer construction unit is used to construct a spatial coordinate system with four assessment methods as the horizontal dimension and three types of posture sets as the vertical dimension. The first nonlinear cognitive deviation value of each assessment method is mapped to the heat value in the spatial coordinate system to generate a deviation heat distribution map. At the same time, the deviation correlation curve between assessment methods is embedded. The deviation heat distribution map and the deviation correlation curve are integrated to form a first complete cognitive layer containing the first deviation quantification value, the deviation distribution law and the correlation characteristics of assessment methods. The depth of the heat value corresponds to the magnitude of the deviation, and the slope of the correlation curve corresponds to the deviation change trend.

6. The intelligent assessment and intervention system for the mental health of the elderly based on multimodal perception according to claim 4, characterized in that, The layer construction module also includes: The coefficient calculation unit is used to extract the feature vectors and corrected intra-cluster feature dispersion of the core cognitive representation, core facial expression representation, and physiological auxiliary cognitive representation in the same second multi-dimensional assessment matrix, calculate the dispersion variation coefficient of each representation under each assessment topic, and quantify the stability of the representation. The second deviation determination unit is used to input the multi-representation features corresponding to each assessment topic into a pre-constructed multi-representation fusion nonlinear deviation analysis model, compare it with the pre-stored benchmark representation data of the elderly healthy population under the same assessment method and assessment topic, calculate the nonlinear deviation quantification value of each assessment topic, and regard it as the second nonlinear cognitive deviation. The nonlinear deviation analysis model is a multi-representation fusion model based on the attention mechanism. Network building units are used to construct a topic-representation bias association network with assessment topics as nodes, second nonlinear cognitive bias values ​​as node weights, and representation stability as node association strength. The second layer construction unit is used to visualize and label the deviation levels of each assessment topic, integrate the correlation network and deviation level labeling to form a second complete cognitive layer containing the second deviation quantification value, the stability of the representation, and the deviation correlation between topics. In this layer, the node size corresponds to the deviation weight, and the thickness of the connection between nodes corresponds to the strength of the representation correlation.

7. The intelligent assessment and intervention system for the mental health of the elderly based on multimodal perception according to claim 1, characterized in that, The weight assignment module includes: The feature extraction unit is used to extract the first deviation quantification features, the correlation features between assessment methods and the posture dimension distribution features corresponding to the assessment methods in the first complete cognitive layer, and to extract the second deviation quantification features, the representation stability features and the deviation correlation features between the assessment topics in the second complete cognitive layer. Combined with the adaptive calibration function constructed by the cognitive physiological characteristics of the elderly population, the heterogeneous interference in the features of the two types of layers is eliminated, and a calibrated feature set is generated. The vector generation unit is used to quantify the correlation strength of the features of the two types of layers based on the temporal response characteristics of the evaluation methods in the first complete cognitive layer and the coefficient of variation of the discreteness of the core representations in the second complete cognitive layer, and dynamically generate a fusion weight vector based on the correlation strength. The initial fusion map generation unit is used to integrate the calibrated feature set across dimensions according to the fusion weight vector, and optimize it based on the feature mutual information entropy to strengthen the intrinsic relationship between the evaluation method, evaluation theme and multimodal representation, and generate the initial fusion map. The graph correction unit is used to introduce the fusion benchmark dataset of elderly mental health, correct the initial fusion graph, and construct a three-dimensional cognitive coordinate system of assessment method, assessment topic, and multimodal representation based on the corrected initial fusion graph. The fused deviation quantification data, feature correlation strength, and representation stability parameters are mapped to the three-dimensional cognitive coordinate system to generate a three-dimensional dynamic fusion cognitive graph. Elderly cognitive deviation warning indicators and feature importance classification labels are embedded and integrated to form a third complete cognitive layer.

8. The intelligent assessment and intervention system for the mental health of the elderly based on multimodal perception according to claim 1, characterized in that, The weight assignment module also includes: The parameter extraction unit is used to extract the feature correlation strength corresponding to each assessment method from the third complete cognitive layer. Feature credibility and the contribution of deviation Extract the representation stability corresponding to each assessment topic. Thematic relevance and deviation influence coefficient Where i=1,2,3,4 corresponds to four assessment methods, and j=1,2,...,n corresponds to various assessment topics. All extracted parameters have been normalized to the range of 0 to 1. The first calculation unit is used to calculate the evaluation weight for each evaluation method. ; The second calculation unit is used to calculate the topic weight of each assessment topic. .