Task-aware cross-modal fusion-based health assessment method and system for stroke patients

CN122800261APending Publication Date: 2026-09-22THE SECOND AFFILIATED HOSPITAL OF ANHUI UNIVERSITY OF TRADITIONAL CHINESE MEDICINE (ACUPUNCTURE AND MOXIBUSTION HOSPITAL OF ANHUI PROVINCE)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611276791.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-21
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

[0004]然而,现有的脑卒中评估方式仍存在以下不足:一方面,量表评估依赖专业人员人工完成,评估效率较低、主观性较强,且多为单次评估,难以满足家庭康复场景下连续监测的需求,单纯依赖人工观察或单一模态数据也难以全面反映患者的健康状态;另一方面,现有基于人工智能的评估方法多针对单一任务建模,对多模态特征通常采用固定不变的融合方式,而NIHSS、FMA和mRS三类评估的关注重点并不相同:NIHSS更关注神经功能缺损,FMA更关注运动功能恢复,mRS更关注整体残疾程度和生活功能状态,采用固定的融合方式难以针对不同评估任务自适应地调整视觉信息与语音信息的贡献比例,从而影响评估结果的准确性

Benefits of technology

本申请提供了一种任务感知跨模态融合的脑卒中患者健康评估方法及系统,通过采集患者的视频数据和语音数据进行建模,并通过双向跨模态注意力机制分别获得视觉模态融合语音信息、语音模态融合视觉信息的跨模态表示,进而生成交互特征参与融合,能够同时利用患者的运动表现、语言表达、面部动作和口面部动作状态等多方面信息,克服了现有技术中单纯依赖人工观察或单一模态数据难以全面反映患者健康状态的不足;通过将目标临床评估任务对应的任务嵌入向量与视觉模态特征、语音模态特征、跨模态交互特征拼接后输入权重生成网络,生成随评估任务自适应变化的视觉模态权重、语音模态权重和交互权重,从而针对不同临床评估任务生成不同的任务融合特征;通过针对不同临床评估任务自适应地调整视觉模态特征、语音模态特征和跨模态交互特征的贡献比例,自动输出标准化的临床评估结果。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122800261A_ABST
    Figure CN122800261A_ABST
Patent Text Reader

Abstract

The application discloses a task-aware cross-modal fusion health assessment method and system for stroke patients, and relates to the technical field of artificial intelligence assisted medical assessment, which comprises the following steps: acquiring video and voice data of a stroke patient, and extracting modal features of the video and voice data respectively; mapping the visual and voice modal features to a unified feature space, performing bidirectional interaction through a cross-modal attention mechanism, splicing and linearly transforming two-way cross-modal representations to obtain cross-modal interaction features; acquiring a task embedding vector corresponding to a target clinical assessment task, generating a visual modal weight, a voice modal weight and an interaction weight according to the respective features and the task embedding vector; performing weighted fusion on the features according to the respective weights to obtain task fusion features; and outputting a clinical assessment result of the stroke patient according to the task fusion features. The application can adaptively adjust the contribution proportion of each feature for different clinical assessment tasks, and automatically output a standardized clinical assessment result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence-assisted medical assessment technology, and in particular to a task-aware cross-modal fusion method and system for health assessment of stroke patients. Background Technology

[0002] Stroke is characterized by high incidence and high disability rate. Patients often suffer from varying degrees of neurological deficits, motor dysfunction, and decreased ability to live independently after the onset of the disease. In clinical practice, standardized scales such as the National Institutes of Health Stroke Scale (NIHSS), the Fugl-Meyer Assessment (FMA), and the modified Rankin Scale (mRS) are commonly used. These scales are completed by doctors or rehabilitation therapists in hospitals or rehabilitation facilities through face-to-face observation and interviews.

[0003] In recent years, with the widespread use of video equipment, microphones, and telemedicine platforms, video and audio data of patients in hospitals, rehabilitation facilities, and homes can be easily collected and recorded. Meanwhile, the development of deep learning technology has made it possible to extract visual features such as limb movements, postures, and facial expressions from videos, and to extract speech features such as pronunciation clarity, speech rate, and pauses from audio, providing a technological foundation for automated health assessment of stroke patients.

[0004] However, existing stroke assessment methods still have the following shortcomings: On the one hand, scale assessments rely on professionals to complete them manually, resulting in low assessment efficiency, high subjectivity, and mostly single-assessment, making it difficult to meet the needs of continuous monitoring in home rehabilitation scenarios. Relying solely on manual observation or single-modal data is also insufficient to comprehensively reflect the patient's health status. On the other hand, existing AI-based assessment methods are mostly designed for single-task modeling, and typically use fixed fusion methods for multimodal features. However, the three types of assessments—NIHSS, FMA, and mRS—have different focuses: NIHSS focuses more on neurological deficits, FMA focuses more on motor function recovery, and mRS focuses more on the overall degree of disability and functional status. Using fixed fusion methods makes it difficult to adaptively adjust the contribution ratio of visual and speech information for different assessment tasks, thus affecting the accuracy of the assessment results. Summary of the Invention

[0005] The purpose of this application is to provide a task-aware cross-modal fusion method and system for health assessment of stroke patients, which can adaptively adjust the contribution ratio of visual modal features, speech modal features and cross-modal interaction features for different clinical assessment tasks, and automatically output standardized clinical assessment results.

[0006] To achieve the above objectives, this application provides the following solution: In a first aspect, this application provides a task-aware cross-modal fusion method for health assessment of stroke patients, including: acquiring video data and speech data of stroke patients, and extracting visual modal features and speech modal features respectively; The visual modal features and the speech modal features are mapped to a unified feature space. A cross-modal attention mechanism is used to perform bidirectional interaction. The two cross-modal representations obtained from the bidirectional interaction are concatenated and linearly transformed to obtain cross-modal interaction features. Obtain the task embedding vector corresponding to the target clinical assessment task. Then, concatenate the visual modal features, the speech modal features, the cross-modal interaction features, and the task embedding vector, and pass them through a weight generation network and normalization processing to generate visual modal weights, speech modal weights, and interaction weights. The visual modal features, the speech modal features, and the cross-modal interaction features are weighted and fused according to their respective weights to obtain the task fusion features corresponding to the target clinical assessment task. Based on the task fusion features, the clinical assessment results of the stroke patient under the target clinical assessment task are output.

[0007] In one embodiment, the method further includes: generating a health assessment report for the stroke patient based on the clinical assessment results, the health assessment report including patient identification, assessment time, and the clinical assessment results; The clinical assessment results of the stroke patients are stored in chronological order to form a record of changes in rehabilitation status, and the score change between two adjacent assessment results is calculated. The score change is used to assist in adjusting the rehabilitation treatment plan.

[0008] In one embodiment, the extraction of the visual modality features includes: The video data is subjected to video frame sampling, invalid frame removal, and image size normalization to obtain a video frame sequence; The video frame sequence is input into a video feature extraction network; The video frames in the video frame sequence are divided into multiple spatiotemporal image blocks. Each spatiotemporal image block is linearly mapped and positional encoding is added to obtain a video token sequence. The video token sequence is modeled using spatial and temporal self-attention, and the modeled video feature sequence is then pooled to obtain the visual modality features.

[0009] In one embodiment, the extraction of the speech modal features includes: The speech data is subjected to noise reduction, framing, endpoint detection, and silent segment removal to obtain effective speech segments; Perform a short-time Fourier transform on the effective speech segment and convert the transform result into a Mel spectrogram; The Mel spectrogram is input into the audio feature extraction network; The Mel spectrogram is divided into multiple spectrogram blocks, and each spectrogram block is linearly mapped and positional encoding is added to obtain an audio token sequence; The time-frequency variation relationship of the audio token sequence is modeled by a multi-layer self-attention network, and the modeled audio feature sequence is pooled to obtain the speech modal features.

[0010] In one embodiment, the mapping of the visual modal features and the speech modal features to a unified feature space is specifically calculated using the following formula: in, This represents the visual modal features. This represents the speech modal features. Represents the mapped visual modal features. This represents the mapped speech modal features. , , , These are learnable parameters; The bidirectional interaction via cross-modal attention mechanism is specifically calculated using the following formula: in, This represents the query matrix generated from the mapped visual modality features. , These represent the key matrix and value matrix generated from the mapped speech modal features, respectively. This represents the first cross-modal representation of visual modality fused with speech information; This represents the query matrix generated from the mapped speech modal features. , These represent the key matrix and value matrix generated from the mapped visual modal features, respectively. This represents the second cross-modal representation of speech modality fusion with visual information; The two-way cross-modal representation includes a first cross-modal representation of visual modality fused with speech information and a second cross-modal representation of speech modality fused with visual information; The cross-modal representations obtained from the bidirectional interaction are concatenated and linearly transformed to obtain the cross-modal interaction features, and the specific calculation formula is as follows: in, This represents the cross-modal interaction feature. Indicates feature splicing, , These are learnable parameters.

[0011] In one embodiment, the specific calculation formula for generating visual modality weights, speech modality weights, and interaction weights is as follows: in, , , These respectively represent the target clinical assessment tasks. The corresponding visual modality weights, the voice modality weights, and the interaction weights, This indicates the target clinical assessment task. The corresponding task embedding vector, , For learnable parameters, Indicates feature splicing; The formula for calculating the weighted fusion is as follows: in, This indicates the target clinical assessment task. The corresponding task fusion features.

[0012] In one embodiment, the target clinical assessment task includes an NIHSS scoring task, an FMA scoring task, and an mRS grading assessment task. For each of the NIHSS scoring task, the FMA scoring task, and the mRS grading assessment task, a corresponding task fusion feature is generated, and each task fusion feature is input into a multi-task health assessment model. The multi-task health assessment model includes a shared feature layer and multiple task prediction branches. The shared feature layer is used to extract a common health status representation from each task fusion feature. The multiple task prediction branches include an NIHSS score prediction branch, an FMA score prediction branch, and an mRS grading prediction branch. Each task prediction branch outputs a corresponding clinical assessment result based on the common health status representation.

[0013] In one embodiment, the multi-task health assessment model is trained using labeled training samples, which include video data and speech data of stroke patients, along with corresponding NIHSS-labeled scores, FMA-labeled scores, and mRS-labeled levels. The total loss function during training is a weighted sum of the prediction losses for each task. ,in, , , These represent the NIHSS score prediction loss, FMA score prediction loss, and mRS rank prediction loss, respectively. , , These are the preset weighting coefficients corresponding to the prediction loss of each task; the prediction loss of each task is calculated using the mean squared error between the labeled result and the prediction result: , , ,in, , , These represent the NIHSS annotation score, FMA annotation score, and mRS annotation level, respectively. These represent the NIHSS prediction score, FMA prediction score, and mRS prediction level, respectively.

[0014] Secondly, this application provides a task-aware cross-modal fusion health assessment system for stroke patients, comprising: The data acquisition module is used to acquire video and audio data of stroke patients; The feature extraction module is used to extract visual modal features from the video data and speech modal features from the speech data. The cross-modal interaction module is used to map the visual modal features and the speech modal features to a unified feature space, perform bidirectional interaction through a cross-modal attention mechanism, and concatenate and linearly transform the two cross-modal representations obtained from the bidirectional interaction to obtain cross-modal interaction features. The weight generation module is used to obtain the task embedding vector corresponding to the target clinical assessment task. The visual modal features, the speech modal features, the cross-modal interaction features and the task embedding vector are concatenated and then processed by the weight generation network and normalization to generate visual modal weights, speech modal weights and interaction weights. The feature fusion module is used to perform weighted fusion of the visual modal features, the speech modal features, and the cross-modal interaction features according to their respective weights to obtain the task fusion features corresponding to the target clinical assessment task. The health assessment module is used to output the clinical assessment results of the stroke patient under the target clinical assessment task based on the task fusion features.

[0015] Thirdly, this application provides an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor executes the computer program to implement the above-described task-aware cross-modal fusion method for health assessment of stroke patients.

[0016] According to the specific embodiments provided in this application, the following technical effects are disclosed: This application provides a task-aware cross-modal fusion method and system for health assessment of stroke patients. It models the patient's video and speech data and obtains cross-modal representations of visual modality fused with speech information and speech modality fused with visual information through a bidirectional cross-modal attention mechanism. This generates interactive features to participate in the fusion, simultaneously utilizing multiple aspects of the patient's motor performance, language expression, facial movements, and orofacial movements. This overcomes the shortcomings of existing technologies that rely solely on manual observation or single-modal data, which cannot comprehensively reflect the patient's health status. By concatenating the task embedding vector corresponding to the target clinical assessment task with visual modal features, speech modal features, and cross-modal interaction features, and inputting the result into a weight generation network, it generates visual modal weights, speech modal weights, and interaction weights that adaptively change with the assessment task, thus generating different task fusion features for different clinical assessment tasks. By adaptively adjusting the contribution ratios of visual modal features, speech modal features, and cross-modal interaction features for different clinical assessment tasks, it automatically outputs standardized clinical assessment results. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is an application environment diagram of a task-aware cross-modal fusion method for health assessment of stroke patients, as described in one embodiment of this application. Figure 2 A flowchart illustrating a task-aware cross-modal fusion method for health assessment of stroke patients, provided as an embodiment of this application; Figure 3 for Figure 2 A detailed flowchart illustrating the process of extracting visual modal features from video data in step S2. Figure 4 for Figure 2 A detailed flowchart illustrating the extraction of speech modal features from speech data in step S2. Figure 5A flowchart illustrating a task-aware cross-modal fusion method for health assessment of stroke patients, provided as another embodiment of this application; Figure 6 A schematic diagram of the functional modules of a task-aware cross-modal fusion health assessment system for stroke patients provided in an embodiment of this application; Figure 7 This is a schematic diagram of the functional modules of a task-aware cross-modal fusion health assessment system for stroke patients, provided as another embodiment of this application. Detailed Implementation

[0019] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0020] Example 1 This application provides a task-aware cross-modal fusion-based health assessment method for stroke patients, which can be applied to, for example... Figure 1 The application environment shown is illustrated. Terminal 101 communicates with server 102 via a network. A data storage system can store the data that server 102 needs to process. The data storage system can be set up independently, integrated into server 102, or placed in the cloud or on another server.

[0021] Terminal 101 can send the collected video and audio data of stroke patients performing preset assessment tasks to server 102. Upon receiving the video and audio data, server 102 extracts visual modal features from the video data and audio modal features from the audio data. It then interacts with the visual and audio modal features through a cross-modal attention mechanism to obtain cross-modal interaction features. It also obtains the task embedding vector corresponding to the target clinical assessment task. Based on the visual modal features, audio modal features, cross-modal interaction features, and task embedding vector, it generates visual modal weights, audio modal weights, and interaction weights. These weights are then weighted and fused to obtain task fusion features, which in turn output the clinical assessment results of the stroke patient under the target clinical assessment task. Server 102 can then feed back the obtained clinical assessment results to terminal 101, which displays the results for patients, families, doctors, or rehabilitation therapists to view.

[0022] Optionally, before sending data to the server 102, the terminal 101 may preprocess the collected video and audio data, such as performing human region detection on the video data, noise reduction on the audio data, and compression and desensitization processing on the data, and then send the preprocessed data to the server 102.

[0023] Furthermore, in some embodiments, the above-described evaluation method can also be implemented independently by either server 102 or terminal 101. For example, if the computing power of terminal 101 meets the requirements, terminal 101 can directly perform feature extraction, cross-modal fusion, and health assessment on the collected video and audio data; alternatively, server 102 can obtain the video and audio data to be evaluated from the data storage system and complete the above-described evaluation process independently.

[0024] The terminal 101 can be, but is not limited to, various desktop computers, laptops, smartphones, tablets, home rehabilitation terminals, IoT devices, and portable wearable devices; in some embodiments, the terminal 101 can also be a doctor's workstation. The server 102 can be implemented using a standalone server or a server cluster composed of multiple servers, or it can be a cloud server; in some embodiments, the server 102 can also be the server corresponding to a hospital information system or a telemedicine platform.

[0025] Example 2 This embodiment provides a task-aware cross-modal fusion method for health assessment of stroke patients, such as... Figure 2 As shown, this method is executed by a computer device, specifically by a terminal or server alone, or by both a terminal and a server. In this embodiment, the method is applied to... Figure 1 Taking server 102 as an example, the explanation includes the following steps: S1, acquire video and audio data of stroke patients.

[0026] S2 extracts visual modal features from video data and speech modal features from speech data.

[0027] S3 uses a cross-modal attention mechanism to interact visual modal features and speech modal features to obtain cross-modal interaction features.

[0028] S4. Obtain the task embedding vector corresponding to the target clinical assessment task. Based on the visual modal features, speech modal features, cross-modal interaction features, and task embedding vector, generate the visual modal weights, speech modal weights, and interaction weights corresponding to the target clinical assessment task.

[0029] S5. Based on visual modality weight, speech modality weight, and interaction weight, the visual modality features, speech modality features, and cross-modal interaction features are weighted and fused to obtain the task fusion features corresponding to the target clinical assessment task.

[0030] S6, based on task fusion characteristics, outputs the clinical assessment results of stroke patients under the target clinical assessment task.

[0031] By implementing steps S1 to S3, video and speech data of patients are collected for modeling. Cross-modal representations of visual modality fused with speech information and speech modality fused with visual information are obtained through a bidirectional cross-modal attention mechanism. Interactive features are then generated to participate in the fusion. This approach can simultaneously utilize multiple aspects of information, such as the patient's motor performance, language expression, facial movements, and orofacial movement state, overcoming the shortcomings of existing technologies that rely solely on manual observation or single-modal data and cannot comprehensively reflect the patient's health status. By implementing steps S4 to S5, the task embedding vector corresponding to the target clinical assessment task is introduced into the generation process of modal fusion weights. This allows the visual modal weights, speech modal weights, and interaction weights to adaptively change with the assessment task, thereby generating different task fusion features for different clinical assessment tasks. By implementing step S6, the contribution ratios of visual modal features, speech modal features, and cross-modal interaction features are adaptively adjusted for different clinical assessment tasks, automatically outputting standardized clinical assessment results.

[0032] The following detailed explanation will illustrate the implementation process of the above steps.

[0033] Regarding step S1 above, video data and audio data can be collected by terminal 101 and sent to server 102 when the stroke patient performs a preset assessment task, or they can be read by server 102 from the data storage system.

[0034] Let the multimodal data collected during the i-th evaluation be represented as: (Equation 1); In the above formula, Represents video data, Represents voice data, Indicates the type of assessment task. It includes auxiliary information such as patient identification, assessment time, and task number.

[0035] In one embodiment, the video data is RGB video data collected by the terminal 101 through a camera device, including video data of the patient performing upper limb raising, arm extension, finger movement, standing, sitting up, walking, turning, facial expression changes, mouth opening, mouth closing, tongue protrusion, or orofacial movements, which can be represented as: (Equation 2); In the above formula, Let represent the t-th frame of the image, where T represents the number of video frames contained in the current video segment. The speech data consists of speech signals collected by the terminal 101 through the microphone, including speech signals generated when the patient performs tasks such as reading aloud, repeating, naming, counting, questioning and answering, or free expression. It can be represented as: (Equation 3); In the above formula, This represents the nth audio sample point, where N represents the number of audio sample points.

[0036] In one implementation, the preset assessment tasks are set according to the assessment requirements of NIHSS, FMA, and mRS: for the NIHSS scoring task, data related to the patient's language expression, facial movements, limb movements, and consciousness response are collected; for the FMA scoring task, data related to the patient's limb movements, posture control, and balance ability are collected; for the mRS grading assessment task, data related to the patient's overall activity ability, language communication ability, and daily functional status are collected.

[0037] Optionally, before sending data, the terminal 101 can also perform data quality checks on the video and audio data. For example, it can use a human detection model to determine whether the video image contains a complete human body, the affected limb, the facial area, or the oral-facial area, and detect the signal-to-noise ratio, effective speech duration, silence ratio, and volume intensity of the audio signal. When the data quality does not meet the preset requirements, the patient is prompted to re-collect the data.

[0038] For step S2 above, visual modal features are extracted from the video data, such as... Figure 3 As shown, it includes the following sub-steps: S211, preprocess the video data to obtain a video frame sequence.

[0039] Specifically, firstly, a video frame sequence is obtained by sampling from the original video at a fixed frame rate or fixed interval: (Equation 4); In the above formula, This represents the sequence of sampled video frames, where L represents the number of sampled video frames.

[0040] Then, invalid frames are removed from the sampled video frames. Invalid frames include video frames that do not contain the patient's body or face area, or video frames with blurred or severely obscured images.

[0041] Finally, each retained frame is scaled to a preset size and pixel normalized to obtain the video tensor of the input video feature extraction network: (Equation 5); In the above formula, This represents the video tensor input to the video feature extraction network after preprocessing. Represents the video frame rate dimension. , These represent the height and width of the normalized image, respectively, and 3 represents the number of RGB color channels.

[0042] S212, input the video frame sequence into the video feature extraction network.

[0043] Specifically, the video feature extraction network is a TimeSformer network, which is used to model the spatial relationships within video frames and the temporal relationships between consecutive video frames through a self-attention mechanism.

[0044] S213, divide the video frames in the video frame sequence into multiple spatiotemporal image blocks, perform linear mapping on each spatiotemporal image block and add position encoding to obtain the video token sequence.

[0045] Specifically, firstly, the video input is divided into multiple spatiotemporal image blocks. Each image block is linearly mapped and then a positional code is added to obtain a video token representation. (Equation 6); In the above formula, This represents the initial token corresponding to the p-th spatiotemporal image block. This represents the p-th spatiotemporal image patch. Indicates vectorization operation, , For learnable parameters, This represents the position code corresponding to the p-th video image block.

[0046] Then, the initial tokens corresponding to all spatiotemporal image blocks are arranged into a video token sequence according to temporal and spatial order: (Equation 7); In the above formula, This represents the sequence of video tokens input to the TimeSformer network, where P represents the total number of video image blocks.

[0047] S214: Feature modeling of video token sequences is performed using spatial self-attention and temporal self-attention, and pooling is performed on the modeled video feature sequences to obtain visual modal features.

[0048] Specifically, after the video token sequence is input into the TimeSformer network, spatial and temporal self-attention are calculated layer by layer. The self-attention calculation is represented as follows: (Equation 8); In the above formula, Q represents the query matrix, K represents the key matrix, and V represents the value matrix. Representing feature dimension, Used to scale attention scores.

[0049] For the l The video feature update process is represented by a TimeSformer network with multiple layers as follows: (Equation 9); In the above formula, Indicates the first l The video feature sequence output by the layer, Indicates the first l The video feature sequence output from layer -1 This represents a TimeSformer feature extraction module.

[0050] go through After modeling with a TimeSformer network, the video feature sequence output from the last layer is processed. Perform pooling or extract classification tokens to obtain visual modality feature vectors: (Equation 10); In the above formula, This indicates a pooling operation or a categorized token extraction operation. This represents the visual modality feature vector.

[0051] Visual modal feature vector It is used to characterize a patient's limb range of motion, motor coordination, postural changes, gait performance, facial expression changes, and oral and facial movement status.

[0052] Regarding step S2 above, speech modal features are extracted from the speech data, such as... Figure 4 As shown, it includes the following sub-steps: S221, preprocess the speech data to obtain valid speech segments.

[0053] Specifically, the speech data is processed by noise reduction, framing, endpoint detection, silence removal, and volume normalization to obtain effective speech segments.

[0054] S222 performs a short-time Fourier transform on the effective speech segment and converts the transform result into a Mel spectrogram.

[0055] Specifically, first, a short-time Fourier transform is performed on the effective speech segments to obtain the time-frequency representation: (Equation 11); In the above formula, Time-frequency representation of speech signals. This represents the nth audio sample point. Represents the window function. Indicates the time frame index. Indicates frequency index, This indicates frame shift, and N represents the number of audio sample points. denoted by , where j represents the number of points in the Fourier transform and j represents the imaginary unit.

[0056] Then, the short-time Fourier transform result is converted into a Mel spectrogram: (Equation 12); Where M represents the Mel spectrum, This represents the Mel filter bank transform. This represents a constant set to prevent zero values ​​from being obtained when taking the logarithm.

[0057] The final speech feature input data is represented as follows: (Equation 13); in, This represents the time-frequency characteristic matrix corresponding to the Mel spectrogram. Represents the Mel frequency dimension. Indicates the number of audio time frames.

[0058] S223, input the Mel spectrogram into the audio feature extraction network.

[0059] Specifically, the audio feature extraction network can be the Spectrogram Transformer network.

[0060] S224, the Mel spectrogram is divided into multiple spectrogram blocks, each spectrogram block is linearly mapped and positional encoding is added to obtain the audio token sequence.

[0061] Specifically, firstly, the Mel spectrogram is divided into multiple spectrogram blocks. Each spectrogram block is vectorized, linearly mapped, and then audio position encoding is added to obtain the initial audio token. (Equation 14); In the above formula, This represents the initial audio token corresponding to the q-th spectrogram patch. This represents the q-th spectrum patch. Indicates vectorization operation, , For the learnable parameters of a linear mapping, This represents the position code corresponding to the q-th spectrogram block.

[0062] Then, all the initial audio tokens corresponding to the spectrum tiles are arranged into an audio token sequence according to time and frequency order: (Equation 15); In the above formula, This represents the audio token sequence before input to the Spectrogram Transformer network, and Q represents the total number of spectrogram tiles.

[0063] S225 uses a multi-layer self-attention network to model the time-frequency variation relationship of the audio token sequence, and performs pooling processing on the modeled audio feature sequence to obtain speech modal features.

[0064] Specifically, for the first l Layered audio features, whose update process is represented as follows: (Equation 16); In the above formula, Indicates the first l The audio feature sequence output by the layer Spectrogram Transformer network. Indicates the first l The audio feature sequence output by a -1 layer Spectrogram Transformer network. This represents a Transformer feature extraction module. l =1,2,…,L a L a This indicates the number of layers in the Spectrogram Transformer network.

[0065] After L a After layer modeling, the audio feature sequence output from the last layer is pooled to obtain the speech modality feature vector: (Equation 17); In the above formula, This indicates a pooling operation. This represents the audio feature sequence output by the last layer of the Spectrogram Transformer network. This represents the speech modal feature vector, used to characterize a patient's pronunciation clarity, speech rate, pauses, speech continuity, speech intensity variations, and fluency.

[0066] Regarding step S3 above, in one implementation, the visual modal features are first... and speech modal features Mapping to a unified feature space: (Equation 18); (Equation 19); In the above formula, Represents the mapped visual modal features. This represents the mapped speech modal features. , , , These are learnable parameters.

[0067] Then, a cross-modal attention mechanism is used to perform bidirectional interaction between the two modalities, resulting in two cross-modal representations. These two cross-modal representations include a first cross-modal representation of visual modality fused with speech information and a second cross-modal representation of speech modality fused with visual information. Specifically: Taking visual modality focusing on speech modality as an example: (Equation 20); In the above formula, This represents the query matrix generated from the mapped visual modality features. These represent the key matrix and value matrix generated from the mapped speech modal features, respectively. This represents the first cross-modal representation of visual modality fused with speech information.

[0068] Similarly, the speech modality focuses on the cross-modal representation of the visual modality as follows: (Equation 21); In the above formula, This represents the query matrix generated from the mapped speech modal features. , These represent the key matrix and value matrix generated from the mapped visual modal features, respectively. This represents the second cross-modal representation of speech modality fused with visual information.

[0069] Then, the cross-modal representations from the two directions are concatenated and subjected to a linear transformation to obtain the cross-modal interaction features: (Equation 22); In the above formula, Indicates cross-modal interaction features. Indicates feature splicing, , These are learnable parameters.

[0070] Regarding step S4 above, in one implementation, the set of clinical assessment tasks is represented as follows: (Equation 23); In the above formula, Represents a set of clinical assessment tasks. Corresponding to NIHSS scoring task, Corresponding to FMA scoring task, Corresponding to the mRS level assessment task.

[0071] For any evaluation task Generate the corresponding task embedding vector: (Equation 24); In the above formula, This indicates a task embedding function. Indicates the assessment task The task embedding vector.

[0072] In one implementation, the task embedding vector is a pre-set learnable vector. The task embedding vector corresponding to each assessment task is jointly optimized with other network parameters during model training to characterize the assessment focus of different clinical scale tasks.

[0073] Finally, the mapped visual modal features, speech modal features, cross-modal interaction features, and task embedding vectors are concatenated and input into the weight generation network, and the output is normalized using softmax. (Equation 25); In the above formula, Indicates the assessment task Weights of lower visual modality features Indicates the assessment task Weights of speech modal features Indicates the assessment task Weights of cross-modal interaction features. and For learnable parameters, This indicates feature splicing.

[0074] Regarding step S5 above, in one embodiment, the visual modal features, speech modal features, and cross-modal interaction features are weighted and fused according to visual modal weights, speech modal weights, and interaction weights to obtain the task fusion features corresponding to the target clinical assessment task. The weighted fusion is expressed as follows: (Equation 26); In the above formula, Indicates the target clinical assessment task The corresponding task fusion features.

[0075] Regarding step S6 above, in one embodiment, the fused features of each task are input into a multi-task health assessment model. The multi-task health assessment model includes a shared feature layer and multiple task prediction branches. The shared feature layer is used to extract common health status representations among different assessment tasks. For any task fused feature... The calculation process is expressed as follows: (Equation 27); In the above formula, This represents the output of the shared feature layer. This represents the task fusion features input to the shared feature layer. , To share feature layer parameters, It is a non-linear activation function.

[0076] Multiple task prediction branches include the NIHSS score prediction branch, the FMA score prediction branch, and the mRS rank prediction branch. For any score prediction branch, the output of the shared feature layer is used... The corresponding rating results are obtained: (Equation 28); In the above formula, For the predicted score, , These are the parameters for the corresponding prediction branch.

[0077] In one implementation, the multi-task health assessment model is trained using labeled training samples, including patient video data, voice data, and NIHSS-labeled scores, FMA-labeled scores, and mRS-labeled levels provided by doctors or rehabilitation therapists. The total loss function is: (Equation 29); In the above formula, , , These represent the NIHSS score prediction loss, FMA score prediction loss, and mRS rank prediction loss, respectively. , , These are the preset weighting coefficients corresponding to the prediction loss of each task.

[0078] In one implementation, NIHSS, FMA, and mRS are all output using regression methods, and the prediction loss for each task is calculated using the mean squared error. (Equation 30); (Equation 31); (Equation 32); In the above formula, , , These represent the NIHSS annotation score, FMA annotation score, and mRS annotation level, respectively. , , These represent the NIHSS prediction score, FMA prediction score, and mRS prediction level, respectively.

[0079] Alternatively, the mRS rank assessment task can also output the results in a classification manner, in which case the mRS rank prediction loss can be the cross-entropy loss.

[0080] By implementing step S6 above, the multi-task health assessment model extracts common health status representations among different assessment tasks through a shared feature layer, and then outputs the corresponding assessment results from the prediction branches of each task. Multiple standardized clinical assessment results can be obtained with a single data collection, avoiding the repetitive work of modeling and assessing different scales separately, and further improving assessment efficiency.

[0081] Optionally, after step S6, as Figure 5 As shown, it also includes: S7 generates a health assessment report for stroke patients based on clinical assessment results. The health assessment report includes patient identification, assessment time, and clinical assessment results.

[0082] Specifically, a health assessment report is generated based on the clinical assessment results: (Equation 33); In the above formula, This represents the health assessment report corresponding to the i-th assessment. Indicates patient identification. Indicates the assessment time. This indicates the data source, the type of evaluation task, and the model output information; , , These represent the NIHSS prediction score, FMA prediction score, and mRS prediction level, respectively.

[0083] S8 stores multiple assessment results of stroke patients in chronological order to form a record of changes in rehabilitation status, and calculates the score change between two adjacent assessment results. The score change is used to assist in adjusting the rehabilitation treatment plan.

[0084] Specifically, the results of multiple assessments are stored in chronological order to form a record of changes in rehabilitation status; (Equation 34); In the above formula, Records indicating changes in recovery status. These represent the health assessment reports corresponding to the 1st to the mth assessments, respectively. This indicates the number of assessments the patient has completed.

[0085] Then, calculate the change in score between two consecutive assessment results, for example: The change in NIHSS score is: (Equation 35); The change in FMA score is: (Equation 36); The change in mRS level is: , (Equation 37).

[0086] In equations 35, 36, and 37 above, , , These represent the changes in NIHSS score, FMA score, and mRS grade, respectively, between the i-th assessment and the (i-1)-th assessment. , , Let represent the NIHSS prediction score, FMA prediction score, and mRS prediction level output for the i-th evaluation, respectively. , , Let i represent the NIHSS prediction score, FMA prediction score, and mRS prediction level output for the (i-1)th evaluation, respectively, i=2,3,…,m.

[0087] By implementing steps S7 to S8 above, the patient's multiple assessment results are stored in chronological order, and the change in scores between two adjacent assessment results is calculated to form a record of changes in rehabilitation status. This allows doctors, rehabilitation therapists, or remote follow-up personnel to intuitively grasp the changing trends of the patient's neurological function, motor function, and degree of disability, providing an auxiliary basis for formulating or adjusting rehabilitation treatment plans.

[0088] In other embodiments, the video feature extraction network can also be a Video Swin Transformer network, a 3D-CNN network, or a 2D-CNN network; the audio feature extraction network can also be a wav2vec 2.0 network, a HuberT network, or a CNN-BiLSTM network; the cross-modal attention mechanism can also be replaced by gated fusion, feature concatenation fusion, or decision-level fusion, in which case the cross-modal interaction features are correspondingly gated fusion features, concatenation fusion features, or decision-level fusion features; in addition to outputting the NIHSS score, FMA score, and mRS level using a regression method, they can also be output using a classification method. For example, when the mRS level evaluation task uses a classification method, the corresponding prediction loss can be cross-entropy loss. The above alternative methods do not affect the generation of the task-aware fusion weights and the implementation of the weighted fusion process in this application.

[0089] Example 3 This application provides a task-aware cross-modal fusion health assessment system for stroke patients, which can be deployed in... Figure 1 In the terminal 101 and / or server 102 shown. For example... Figure 6 As shown, the system includes: a data acquisition module 301, a feature extraction module 302, a cross-modal interaction module 303, a weight generation module 304, a feature fusion module 305, and a health assessment module 306.

[0090] Data acquisition module 301 is used to acquire video and audio data of stroke patients; Feature extraction module 302 is used to extract visual modal features from the video data and extract speech modal features from the speech data; The cross-modal interaction module 303 is used to map the visual modal features and the speech modal features to a unified feature space, perform bidirectional interaction through a cross-modal attention mechanism, and concatenate and linearly transform the two cross-modal representations obtained from the bidirectional interaction to obtain cross-modal interaction features. The weight generation module 304 is used to obtain the task embedding vector corresponding to the target clinical assessment task, and to generate visual modal weights, speech modal weights, and interaction weights by concatenating the visual modal features, the speech modal features, the cross-modal interaction features, and the task embedding vector and then passing them through a weight generation network and normalization processing. The feature fusion module 305 is used to perform weighted fusion of the visual modal features, the speech modal features and the cross-modal interaction features according to each weight, so as to obtain the task fusion features corresponding to the target clinical assessment task; The health assessment module 306 is used to output the clinical assessment results of the stroke patient under the target clinical assessment task based on the task fusion features.

[0091] Optionally, such as Figure 7 As shown, the system also includes a report generation module 307, which is used to generate a health assessment report for the stroke patient based on the clinical assessment results, store the multiple assessment results of the stroke patient in chronological order to form a record of changes in rehabilitation status, and calculate the score change between two adjacent assessment results. The score change is used to assist in adjusting the rehabilitation treatment plan.

[0092] It should be noted that the specific implementation methods of the above modules can be found in the description of steps S1 to S8 in Embodiment 2, and will not be repeated here. The above modules can be implemented by software, hardware, or a combination of software and hardware; when implemented by software, each module can be a computer program unit stored in memory and executable by a processor.

[0093] Example 4 This application provides an electronic device including a processor and a memory. The memory stores a computer program, and when the processor executes the computer program, it implements the task-aware cross-modal fusion method for health assessment of stroke patients as described in Embodiment 2.

[0094] Optionally, the electronic device further includes a communication interface and a bus, wherein the processor, the memory, and the communication interface communicate with each other via the bus. The electronic device can be... Figure 1 The terminal 101 or server 102 shown.

[0095] This application also provides a computer-readable storage medium storing a computer program. When executed by a processor, the computer program implements the task-aware cross-modal fusion method for stroke patient health assessment as described in Embodiment 2. The computer-readable storage medium includes, but is not limited to, random access memory, read-only memory, erasable programmable read-only memory, optical disk, magnetic disk, or magneto-optical storage medium.

[0096] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.

[0097] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0098] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A task-aware cross-modal fusion method for health assessment of stroke patients, characterized in that, include: Video and audio data of stroke patients were acquired, and visual modal features and audio modal features were extracted respectively. The visual modal features and the speech modal features are mapped to a unified feature space. A cross-modal attention mechanism is used to perform bidirectional interaction. The two cross-modal representations obtained from the bidirectional interaction are concatenated and linearly transformed to obtain cross-modal interaction features. Obtain the task embedding vector corresponding to the target clinical assessment task. Then, concatenate the visual modal features, the speech modal features, the cross-modal interaction features, and the task embedding vector, and pass them through a weight generation network and normalization processing to generate visual modal weights, speech modal weights, and interaction weights. The visual modal features, the speech modal features, and the cross-modal interaction features are weighted and fused according to their respective weights to obtain the task fusion features corresponding to the target clinical assessment task. Based on the task fusion features, the clinical assessment results of the stroke patient under the target clinical assessment task are output.

2. The task-aware cross-modal fusion method for health assessment of stroke patients according to claim 1, characterized in that, Also includes: A health assessment report for the stroke patient is generated based on the clinical assessment results. The health assessment report includes the patient's identification, the assessment time, and the clinical assessment results. The clinical assessment results of the stroke patients are stored in chronological order to form a record of changes in rehabilitation status, and the score change between two adjacent assessment results is calculated. The score change is used to assist in adjusting the rehabilitation treatment plan.

3. The task-aware cross-modal fusion method for health assessment of stroke patients according to claim 1, characterized in that, The extraction of the visual modality features includes: The video data is subjected to video frame sampling, invalid frame removal, and image size normalization to obtain a video frame sequence; The video frame sequence is input into a video feature extraction network; The video frames in the video frame sequence are divided into multiple spatiotemporal image blocks. Each spatiotemporal image block is linearly mapped and positional encoding is added to obtain a video token sequence. The video token sequence is modeled using spatial and temporal self-attention, and the modeled video feature sequence is then pooled to obtain the visual modality features.

4. The task-aware cross-modal fusion method for health assessment of stroke patients according to claim 1, characterized in that, The extraction of the speech modal features includes: The speech data is subjected to noise reduction, framing, endpoint detection, and silent segment removal to obtain effective speech segments; Perform a short-time Fourier transform on the effective speech segment and convert the transform result into a Mel spectrogram; The Mel spectrogram is input into the audio feature extraction network; The Mel spectrogram is divided into multiple spectrogram blocks, and each spectrogram block is linearly mapped and positional encoding is added to obtain an audio token sequence; The time-frequency variation relationship of the audio token sequence is modeled by a multi-layer self-attention network, and the modeled audio feature sequence is pooled to obtain the speech modal features.

5. The method for health assessment of stroke patients based on task-aware cross-modal fusion according to claim 1, characterized in that, The specific calculation formula for mapping the visual modal features and the speech modal features to a unified feature space is as follows: in, This represents the visual modal features. This represents the speech modal features. Represents the mapped visual modal features. This represents the mapped speech modal features. , , , These are learnable parameters; The bidirectional interaction via cross-modal attention mechanism is specifically calculated using the following formula: in, This represents the query matrix generated from the mapped visual modality features. , These represent the key matrix and value matrix generated from the mapped speech modal features, respectively. This represents the first cross-modal representation of visual modality fused with speech information; This represents the query matrix generated from the mapped speech modal features. , These represent the key matrix and value matrix generated from the mapped visual modal features, respectively. This represents the second cross-modal representation of speech modality fusion with visual information; The two-way cross-modal representation includes a first cross-modal representation of visual modality fused with speech information and a second cross-modal representation of speech modality fused with visual information; The cross-modal representations obtained from the bidirectional interaction are concatenated and linearly transformed to obtain the cross-modal interaction features, and the specific calculation formula is as follows: in, This represents the cross-modal interaction feature. Indicates feature splicing, , These are learnable parameters.

6. The method for health assessment of stroke patients based on task-aware cross-modal fusion according to claim 5, characterized in that, The specific calculation formulas for the generated visual modality weights, speech modality weights, and interaction weights are as follows: in, , , Each represents a target clinical assessment task. The corresponding visual modality weights, the voice modality weights, and the interaction weights, This indicates the target clinical assessment task. The corresponding task embedding vector, , For learnable parameters, Indicates feature splicing; The formula for calculating the weighted fusion is as follows: in, This indicates the target clinical assessment task. The corresponding task fusion features.

7. The task-aware cross-modal fusion method for health assessment of stroke patients according to claim 1, characterized in that, The target clinical assessment tasks include the NIHSS scoring task, the FMA scoring task, and the mRS grading task; For each of the NIHSS scoring task, the FMA scoring task, and the mRS level assessment task, a corresponding task fusion feature is generated, and each task fusion feature is input into a multi-task health assessment model, which includes a shared feature layer and multiple task prediction branches. The shared feature layer is used to extract a common health status representation from the fusion features of each task. The multiple task prediction branches include an NIHSS score prediction branch, an FMA score prediction branch, and an mRS rank prediction branch. Each task prediction branch outputs a corresponding clinical assessment result based on the common health status representation.

8. The task-aware cross-modal fusion method for health assessment of stroke patients according to claim 7, characterized in that, The multi-task health assessment model is trained using labeled training samples, which include video data and voice data of stroke patients, as well as corresponding NIHSS labeled scores, FMA labeled scores, and mRS labeled levels. The total loss function during training is a weighted sum of the prediction losses for each task: in, , , These represent the NIHSS score prediction loss, FMA score prediction loss, and mRS rank prediction loss, respectively. , , These are the preset weighting coefficients corresponding to the prediction loss of each task; The prediction loss for each task is calculated using the mean squared error between the labeled and predicted results: in, , , These represent the NIHSS annotation score, FMA annotation score, and mRS annotation level, respectively. , , These represent the NIHSS prediction score, FMA prediction score, and mRS prediction level, respectively.

9. A task-aware cross-modal fusion health assessment system for stroke patients, characterized in that, include: The data acquisition module is used to acquire video and audio data of stroke patients; The feature extraction module is used to extract visual modal features from the video data and speech modal features from the speech data. The cross-modal interaction module is used to map the visual modal features and the speech modal features to a unified feature space, perform bidirectional interaction through a cross-modal attention mechanism, and concatenate and linearly transform the two cross-modal representations obtained from the bidirectional interaction to obtain cross-modal interaction features. The weight generation module is used to obtain the task embedding vector corresponding to the target clinical assessment task. The visual modal features, the speech modal features, the cross-modal interaction features and the task embedding vector are concatenated and then processed by the weight generation network and normalization to generate visual modal weights, speech modal weights and interaction weights. The feature fusion module is used to perform weighted fusion of the visual modal features, the speech modal features, and the cross-modal interaction features according to their respective weights to obtain the task fusion features corresponding to the target clinical assessment task. The health assessment module is used to output the clinical assessment results of the stroke patient under the target clinical assessment task based on the task fusion features.

10. An electronic device, comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor executes the computer program to implement the task-aware cross-modal fusion method for health assessment of stroke patients according to any one of claims 1 to 8.