Multi-modal dementia analysis auxiliary diagnosis method, system, device, medium and equipment

By extracting and fusing multimodal features from patient consultation videos, and combining large language models and clinical dementia rating scales, the shortcomings of existing dementia diagnostic methods have been addressed, achieving a more accurate and objective diagnosis of dementia symptoms.

CN121034590APending Publication Date: 2025-11-28BEIJING HUILONGGUAN HOSPITAL
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511105568.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-07
Publication Date
2025-11-28

AI Technical Summary

Technical Problem

Current methods for diagnosing dementia rely heavily on manpower and lack specificity. Imaging technologies and equipment are expensive and difficult to apply widely. Multimodal data is scattered and lacks unified standards. Data mining and analysis are not deep enough, making it difficult to fully understand the etiology and mechanisms.

Method used

By acquiring patient consultation videos, multimodal feature extraction is performed, including analysis of voice, text, and visual feature data. Time alignment and feature fusion are performed using a large language model and preset standards, and a comprehensive diagnosis is made in conjunction with the Clinical Dementia Rating Scale.

Benefits of technology

It enables more comprehensive and accurate diagnosis of dementia symptoms, improves the objectivity and accuracy of diagnosis, and reduces reliance on human resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121034590A_ABST
    Figure CN121034590A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of medical diagnosis, and provides a multi-modal dementia analysis auxiliary diagnosis method, system and device, a medium and equipment, and the method comprises the steps: obtaining an inquiry video of a patient, and carrying out the multi-modal feature extraction of the video, and obtaining a multi-modal feature data set; analyzing the voice feature data based on a large language model and a preset voice standard to obtain voice dimension information; analyzing the text feature data based on a large language model and a preset text standard to obtain text dimension information; analyzing the visual feature data based on a large language model and a preset visual standard to obtain visual dimension information; performing time alignment processing and feature fusion processing on the modal feature data to obtain a comprehensive multi-modal feature vector; and carrying out dementia symptom diagnosis analysis based on the large language model and a pre-constructed clinical dementia rating scale, and determining a diagnosis result based on an analysis result and the modal dimension information. The accuracy and efficiency of dementia diagnosis can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of medical diagnostic technology, and more specifically, to a multimodal dementia analysis-assisted diagnostic method, system, device, medium, and equipment. Background Technology

[0002] With the accelerating aging of the global population, the incidence of dementia is increasing year by year. Dementia not only seriously affects the quality of life of patients but also places a heavy care burden on their families, while also putting enormous pressure on social medical resources and economic development. Therefore, timely diagnosis of dementia is crucial for early intervention, slowing disease progression, and improving patient prognosis.

[0003] In current diagnostic methods, dementia diagnosis primarily relies on clinicians to collect medical history, conduct scale assessments, and perform clinical observations. However, this method is highly dependent on human intervention and suffers from significant limitations in specificity. With advancements in medical technology, imaging techniques such as MRI and CT scans are increasingly being used in the auxiliary diagnosis of dementia. These techniques provide doctors with direct evidence of changes in brain structure, but their sensitivity and specificity remain limited in the early diagnosis of dementia. Secondly, while molecular imaging techniques such as positron emission tomography (PET) can detect some pathological changes associated with dementia, their widespread application is constrained by the high cost of equipment and examinations. Furthermore, the diagnostic process of these technologies suffers from incomplete data integration, fragmented multimodal data, a lack of unified standards, and insufficient depth in data mining and analysis, making it difficult to fully understand the etiological mechanisms. Summary of the Invention

[0004] This disclosure provides at least one method, system, device, medium, and equipment for multimodal dementia analysis-assisted diagnosis. By comprehensively analyzing multimodal data, it can more comprehensively and accurately capture abnormal manifestations of patients in language, cognition, and vision, providing a more reliable basis for the diagnosis of dementia symptoms.

[0005] This disclosure provides a multimodal dementia analysis-assisted diagnostic method, including:

[0006] The patient's consultation video is acquired, and multimodal feature extraction processing is performed on the consultation video to obtain a multimodal feature dataset; wherein, the multimodal feature dataset includes speech feature data, text feature data, and visual feature data;

[0007] Speech analysis is performed on the speech feature data based on the large language model and the preset speech standard to obtain speech dimension information; and text analysis is performed on the text feature data based on the large language model and the preset text standard to obtain text dimension information; and visual analysis is performed on the visual feature data based on the large language model and the preset visual standard to obtain visual dimension information.

[0008] The speech feature data, text feature data, and visual feature data are subjected to time alignment and feature fusion processing respectively to obtain a comprehensive multimodal feature vector;

[0009] The comprehensive multimodal feature vectors are analyzed for dementia symptom diagnosis based on a large language model and a pre-constructed clinical dementia rating scale. Based on the analysis results, the speech dimension information, the text dimension information, and the visual dimension information, the diagnosis result for the patient is determined.

[0010] This disclosure provides a multimodal dementia analysis-assisted diagnostic system, including:

[0011] The video capture module is used to acquire videos of patients' consultations.

[0012] A multimodal feature extraction module, connected to the video acquisition module, is used to perform multimodal feature extraction processing on the consultation video to obtain a multimodal feature dataset; wherein, the multimodal feature dataset includes speech feature data, text feature data, and visual feature data;

[0013] A multi-dimensional analysis module is connected to the multimodal feature extraction module and the large language model, respectively, and is used to perform speech analysis on the speech feature data based on the large language model and a preset speech standard to obtain speech dimension information; and to perform text analysis on the text feature data based on the large language model and a preset text standard to obtain text dimension information; and to perform visual analysis on the visual feature data based on the large language model and a preset visual standard to obtain visual dimension information.

[0014] The feature processing module, connected to the multimodal feature extraction module, is used to perform time alignment processing and feature fusion processing on the speech feature data, text feature data and visual feature data respectively to obtain a comprehensive multimodal feature vector;

[0015] The diagnostic analysis module is connected to the feature processing module, the multi-dimensional analysis module, and the large language model, respectively. It is used to perform dementia symptom diagnostic analysis on the comprehensive multimodal feature vector based on the large language model and a pre-built clinical dementia rating scale, and to determine the diagnosis result of the patient based on the analysis results, the voice dimension information, the text dimension information, and the visual dimension information.

[0016] This disclosure provides a multimodal dementia analysis-assisted diagnostic device, comprising:

[0017] A multimodal feature extraction module is used to acquire the patient's consultation video and perform multimodal feature extraction processing on the consultation video to obtain a multimodal feature dataset; wherein, the multimodal feature dataset includes speech feature data, text feature data and visual feature data;

[0018] The multimodal dimension analysis module is used to perform speech analysis on the speech feature data based on a large language model and a preset speech standard to obtain speech dimension information; and to perform text analysis on the text feature data based on the large language model and a preset text standard to obtain text dimension information; and to perform visual analysis on the visual feature data based on the large language model and a preset visual standard to obtain visual dimension information.

[0019] The multimodal feature fusion module is used to perform time alignment processing and feature fusion processing on the speech feature data, text feature data and visual feature data respectively to obtain a comprehensive multimodal feature vector;

[0020] The diagnostic result determination module is used to perform dementia symptom diagnostic analysis on the comprehensive multimodal feature vector based on a large language model and a pre-built clinical dementia rating scale, and to determine the diagnostic result for the patient based on the analysis results, the speech dimension information, the text dimension information, and the visual dimension information.

[0021] This disclosure provides a computer device including a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the computer device is running, the processor communicates with the memory via the bus. When the machine-readable instructions are executed by the processor, they perform the multimodal dementia analysis-assisted diagnosis method as described in any of the above possible embodiments.

[0022] This disclosure provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the multimodal dementia analysis-assisted diagnosis method as described in any of the possible embodiments above.

[0023] The multimodal dementia analysis-assisted diagnosis method, system, device, medium, and equipment provided in this disclosure include: acquiring a patient's consultation video and performing multimodal feature extraction processing on the consultation video to obtain a multimodal feature dataset; wherein, the multimodal feature dataset includes speech feature data, text feature data, and visual feature data; performing speech analysis on the speech feature data based on a large language model and a preset speech standard to obtain speech dimension information; this speech dimension information covers key information reflecting the patient's language state, such as speech rate, intonation, and pronunciation quality. Additionally, performing text analysis on the text feature data based on a large language model and a preset text standard to obtain text dimension information; the text dimension information includes text features related to cognitive ability, such as semantic logic and content completeness. Finally, performing visual analysis on the visual feature data based on a large language model and a preset visual standard to obtain visual dimension information; the visual dimension information includes visual features such as facial expression changes, eye contact, and limb coordination. The speech, text, and visual feature data were subjected to time alignment and feature fusion processing respectively to obtain a comprehensive multimodal feature vector. Time alignment ensured that the feature data of different modalities remained consistent over time, while feature fusion effectively integrated the features of different modalities to comprehensively reflect the patient's overall condition. Based on a large language model and a pre-built clinical dementia rating scale, the comprehensive multimodal feature vector was used for dementia symptom diagnosis analysis. Based on the analysis results, speech dimension information, text dimension information, and visual dimension information, the diagnosis of the patient was determined.

[0024] In this way, by integrating speech, text and visual multimodal information and utilizing the analysis of comprehensive multimodal data, this disclosure can more comprehensively and accurately capture abnormal manifestations of patients in language, cognition and vision, and by using large language models and clinical dementia rating scales for comprehensive analysis, it can effectively improve the accuracy and objectivity of dementia diagnosis.

[0025] To make the above-mentioned objects, features and advantages of this disclosure more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0026] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the accompanying drawings referenced in the embodiments will be briefly described below. These drawings are incorporated in and constitute a part of this specification. They illustrate embodiments conforming to this disclosure and, together with the specification, serve to explain the technical solutions of this disclosure. It should be understood that the following drawings only show some embodiments of this disclosure and should not be considered as limiting the scope. Those skilled in the art can obtain other related drawings based on these drawings without creative effort.

[0027] Figure 1 A schematic diagram of the technical architecture of a multimodal dementia analysis-assisted diagnostic method provided in an embodiment of this disclosure is shown;

[0028] Figure 2 A flowchart of a multimodal dementia analysis-assisted diagnostic method provided by an embodiment of this disclosure is shown;

[0029] Figure 3 The flowchart illustrates a method for extracting multimodal features from a medical consultation video according to an embodiment of this disclosure.

[0030] Figure 4 A flowchart of a remote assisted diagnostic method provided by an embodiment of this disclosure is shown;

[0031] Figure 5 This diagram illustrates the structure of a multimodal dementia analysis-assisted diagnostic system provided in an embodiment of this disclosure.

[0032] Figure 6 This diagram illustrates the structure of a multimodal dementia analysis-assisted diagnostic device provided in an embodiment of this disclosure.

[0033] Figure 7 A schematic diagram of the structure of a computer device provided in an embodiment of this disclosure is shown. Detailed Implementation

[0034] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this disclosure, and not all of them. The components of the embodiments of this disclosure described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this disclosure provided in the accompanying drawings is not intended to limit the scope of the claimed disclosure, but merely represents selected embodiments of this disclosure. All other embodiments obtained by those skilled in the art based on the embodiments of this disclosure without inventive effort are within the scope of protection of this disclosure.

[0035] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.

[0036] In this document, the term "and / or" merely describes a relationship, indicating that three relationships can exist. For example, A and / or B can represent three cases: A alone, A and B simultaneously, and B alone. Furthermore, the term "at least one" in this document means any combination of at least two of any one or more elements. For example, including at least one of A, B, and C can mean including any one or more elements selected from the set consisting of A, B, and C.

[0037] To facilitate understanding of this embodiment, the executing entity of the multimodal dementia analysis-assisted diagnosis method provided in this disclosure will first be described in detail. The executing entity of the multimodal dementia analysis-assisted diagnosis method provided in this disclosure is a computer device. This computer device can be a terminal device or a server. The terminal device can also be a mobile device, user terminal, terminal, handheld device, computing device, vehicle-mounted device, wearable device, etc. The server can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud storage, big data, and artificial intelligence platforms. Optionally, this method can also be applied to an implementation environment composed of computer devices and servers.

[0038] Secondly, combining Figure 1 The schematic diagram of the technical architecture of this disclosure provides a detailed description of the execution logic of the multimodal dementia analysis-assisted diagnosis method provided in the embodiments of this disclosure. The execution process of the multimodal dementia analysis-assisted diagnosis method provided in the embodiments of this disclosure closely relies on the various levels and modules in this technical architecture. Supported by the service platform, patient management platform, medical service platform, appointment platform, medical record platform, and telemedicine platform provide a basic service framework and data interaction environment for the entire assisted diagnosis process. At the service business level, it covers diverse service content such as pre-diagnosis services, in-diagnosis and post-diagnosis services, as well as pre-diagnosis intelligent assistance and post-diagnosis intelligent assistance, providing patients with full-process medical assistance support. During the execution of the entire solution, the AI ​​model center, as the core intelligent processing unit, provides powerful algorithm support and model computing capabilities for each analysis step. In addition, various data resources such as sample data, medical record data, and chief complaint data in the data layer, as well as technical tools such as Hive, SQL, and MongoDB in the public components, and various operating systems at the system and environment level and physical machine servers and other hardware facilities at the foundation layer, together provide comprehensive support and guarantee for the execution of the entire multimodal dementia analysis-assisted diagnosis method, ensuring the accuracy, efficiency and stability of the diagnostic process.

[0039] The multimodal dementia analysis-assisted diagnosis method provided in this application will be described in detail below with reference to the accompanying drawings. See also Figure 2 The diagram shown is a flowchart of a multimodal dementia analysis-assisted diagnostic method provided in this embodiment of the present disclosure. The method includes the following steps S201 to S204:

[0040] S201, acquire the patient's consultation video, and perform multimodal feature extraction processing on the consultation video to obtain a multimodal feature dataset.

[0041] As we can understand, a consultation video refers to video footage recording the interaction between a patient and a doctor during a medical consultation. This footage may include the patient's voice, facial expressions, body language, and conversation with the doctor. By extracting feature data from different modalities from the consultation video, a multimodal feature dataset is obtained. Here, multimodality primarily refers to voice, text, and visual modalities. The multimodal feature dataset is a collection of voice feature data, text feature data, and visual feature data extracted from the consultation video.

[0042] Specifically, speech feature data is a set of extracted speech feature vectors that comprehensively covers various feature information of speech signals in the consultation video; text feature data is a set of numerical vectors obtained after text feature encoding, which can represent the key features of text information in the consultation video and can be used for dementia symptom analysis in the text dimension; visual feature data is comprehensive data containing visual feature information of patients in the consultation video, obtained after extracting keyframe features from the video and then aggregating time-series features, and can be used for dementia symptom analysis in the visual dimension.

[0043] For example, in order to more accurately obtain multimodal feature data from consultation videos that can be used for dementia analysis-assisted diagnosis, refer to Figure 3 As shown, the multimodal feature extraction process for medical consultation videos may include the following steps S301 to S303:

[0044] S301, extract the audio signal from the consultation video, extract the audio feature vector from the audio signal, and obtain the audio feature data.

[0045] Specifically, the audio signal in a medical consultation video is the sound signal recorded during the consultation process between the patient and the doctor. It contains various information such as the patient's speech content, tone, speech rate, and volume, and is an important carrier reflecting the patient's language state. A set of numerical vectors obtained by mathematically transforming and processing the audio signal can quantitatively represent various characteristics of the audio signal, such as spectral characteristics and temporal characteristics, for subsequent analysis and processing.

[0046] Here, speech signal extraction can be performed using audio processing libraries, such as the Librosa library in Python. Librosa provides rich audio reading and processing capabilities. By calling its load function, speech signals can be easily extracted from consultation video files and converted into digital audio formats, such as WAV format, for subsequent processing.

[0047] Furthermore, Mel-Frequency Cepstrum (MFCC) extraction technology can be used to extract speech features from the speech signal. MFCC is a widely used feature in speech recognition and speaker recognition, simulating the nonlinear perception of sound frequencies by the human ear. Specifically, the extracted speech signal is first pre-emphasized to enhance the high-frequency components, making the signal spectrum flatter for easier subsequent analysis. Then, framing is performed, dividing the continuous speech signal into multiple short frames, typically 20-30ms each. Next, each frame is windowed, such as using a Hamming window, to reduce spectral leakage. Then, a Fast Fourier Transform (FFT) is performed to convert the time-domain signal to the frequency-domain signal, obtaining the spectrum. The spectrum is then filtered using a Mel filter bank to simulate the nonlinear perception of sound frequencies by the human ear, obtaining the Mel spectrum. After taking the logarithm of the Mel spectrum, a Discrete Cosine Transform (DCT) is performed to finally obtain the MFCC feature vector. For example, for a speech segment containing 10 frames of speech signal, if 13-dimensional MFCC features are extracted from each frame, a 10×13 feature matrix can be obtained, which can be used as the speech feature vector of the speech segment.

[0048] S302, extract text information about the consultation video based on the voice signal, encode the text information into text features, and obtain the text feature data.

[0049] Specifically, speech recognition technology can convert the speech signals in a consultation video into text content, obtaining the text information of the consultation video. This text information can include the dialogue between the patient and the doctor, thus reflecting the patient's language expression and thought process. Text feature encoding is the process of converting text information into numerical vectors that computers can understand and process. Through specific encoding methods, semantic, syntactic, and other information in the text can be transformed into vector form.

[0050] For example, in the process of extracting text information, Automatic Speech Recognition (ASR) technology can be used, such as the open-source Kaldi toolkit. Kaldi is a powerful speech recognition framework that integrates various acoustic models and language model training methods. This can include: First, using Kaldi to extract acoustic features from the speech signal of the consultation video, such as the MFCC features mentioned earlier. Then, using a pre-trained acoustic model to decode the acoustic features, obtaining preliminary speech recognition results. Next, combining the language model to refine and optimize the preliminary results, improving recognition accuracy, and finally obtaining the text information of the consultation video.

[0051] Furthermore, when encoding textual features from the identified consultation videos, word embedding techniques from natural language processing, such as the Gemini Embedding model, can be used. In practical applications, Gemini Embedding demonstrates significant advantages in medical text classification tasks (such as ICD encoding prediction), as its vector representations directly reflect the hierarchical semantic relationships between medical terms (such as the implicit relationship between "type I diabetes" and "metabolic syndrome"), while traditional word embedding requires subsequent feature engineering to achieve similar results.

[0052] For example, taking the Gemini Embedding model as an example, this model can directly map complete sentences or paragraphs into high-dimensional semantic vectors, capturing global contextual information without word-by-word processing. Specifically, pre-trained models of Gemini Embedding, such as gemini-1.0-ultra-embedding, can be used to directly encode the entire text into a 3072-dimensional vector. This model also supports dynamic dimensionality reduction, reducing the vector to 128 / 256 / 512 dimensions to adapt to computational needs in different scenarios. The Gemini Embedding model can directly process long texts, supporting up to 2048 tokens, effectively avoiding the context fragmentation problem that may occur with traditional word-by-word embedding (such as Word2Vec). Furthermore, the Matryoshka Representation Learning (MRL) technology built into the Gemini Embedding model can flexibly adjust the output dimension while maintaining high accuracy.

[0053] Here, if the text length exceeds the model's processing limit, a sliding window mean pooling method can be used to segment the text, and then the processing results of each segment can be merged. At the same time, L2 normalization is performed on all embedding vectors to ensure the stability of similarity calculations (such as cosine similarity).

[0054] In some possible embodiments, after the text information is extracted, the extracted text can be cleaned to remove meaningless symbols, stop words, etc., and then the text can be segmented into words or phrases for subsequent encoding processing and text dimension analysis.

[0055] S303, extract multiple keyframe images from the consultation video according to a preset time interval, extract features from each keyframe image, and aggregate the time-series features of the feature extraction results of each keyframe image to obtain the visual feature data.

[0056] Here, keyframe images from a consultation video refer to representative image frames selected from the video according to certain rules (such as preset time intervals). These image frames reflect important visual information such as the patient's facial expressions and body movements during the consultation. The selected keyframe images are then analyzed and processed to extract features relevant to dementia diagnosis, such as facial feature points and body movement trajectories, and these features are converted into feature vectors. Simultaneously, since the keyframe images are selected in chronological order, time-series feature aggregation integrates and analyzes the feature extraction results from each keyframe image in chronological order to capture the changing patterns of the patient's visual performance over time, obtaining visual feature data that comprehensively reflects the patient's visual characteristics. The comprehensive data obtained after time-series feature aggregation, containing the patient's visual feature information from the consultation video, can be used for visual-based dementia symptom analysis.

[0057] For example, when extracting keyframe images, a content-based approach can be used, such as using the OpenCV library combined with image feature analysis. First, the histogram features of each frame in the consultation video are calculated; the histogram reflects the distribution of pixels in the image. Then, the difference between the histograms of adjacent frames is calculated. When the difference exceeds a certain threshold, the current frame is considered a keyframe. For example, setting the difference threshold to 0.2, when the calculated difference between the histograms of adjacent frames is greater than 0.2, the current frame is extracted as a keyframe. Keyframes are extracted at preset time intervals, such as every second, ensuring that the keyframes cover the important visual information of the consultation process.

[0058] Furthermore, when extracting features from keyframe images, adaptive selection of feature extraction methods can be employed. For example, for facial feature extraction, deep learning models such as OpenFace can be used. OpenFace is a facial behavior analysis toolkit based on convolutional neural networks, capable of detecting facial keypoints, estimating head pose, and recognizing facial action units. The extracted keyframe images are input into the OpenFace model, which outputs feature information such as the coordinates of facial keypoints and the activation level of facial action units. For limb movement feature extraction, the OpenPose model can be used. OpenPose can detect the positions of human joints, and by analyzing the motion trajectory of these joints, features of limb movements, such as speed and amplitude, can be obtained.

[0059] Here, n keyframe images can be extracted from the consultation video, with each image having a feature dimension of m, resulting in an n×m feature matrix. This matrix can then be further aggregated using a Long Short-Term Memory (LSTM) network, a special type of recurrent neural network capable of handling long-term dependencies in sequential data. The feature matrix is ​​input into the LSTM model in chronological order. The model processes the features at each time step and outputs a comprehensive feature vector as visual feature data. For example, setting the hidden layer dimension of the LSTM to 128, after training, the model outputs a 128-dimensional visual feature vector, which reflects the changes in the patient's visual behavior during the consultation process.

[0060] In some other embodiments, convolutional neural networks, such as VGG16 or ResNet, can be directly used to extract visual features; no specific limitation is made here. Taking VGG16 as an example, it consists of multiple convolutional layers, pooling layers, and fully connected layers. Each frame of the consultation video is input into the VGG16 model, and local features of the image are extracted through convolution and pooling operations. Finally, the feature vector of the image is obtained through the fully connected layer, which serves as the visual feature data.

[0061] S202, performing speech analysis on the speech feature data based on the large language model and preset speech standards to obtain speech dimension information; and performing text analysis on the text feature data based on the large language model and preset text standards to obtain text dimension information; and performing visual analysis on the visual feature data based on the large language model and preset visual standards to obtain visual dimension information.

[0062] Specifically, a large language model is a natural language processing model based on deep learning, possessing powerful language understanding and generation capabilities, and able to perform in-depth analysis and processing of input text or speech data. Different large language models can be selected for different modalities of data. For example, pre-trained language models such as BERT and GPT can be used to analyze text data; large language models such as Qwen2.5-VL can be used to analyze video data; and large language models such as Qwen2-Audio can be used to analyze audio data.

[0063] The preset speech standards are criteria used to measure whether speech feature data is normal. These can include speech rate and intonation standards (defining appropriate speech rate ranges and intonation patterns), sound quality standards (such as clarity of sound, presence of inappropriate silences or interruptions in responses), pronunciation accuracy standards (judging whether pronunciation conforms to standard pronunciation, whether there are mispronunciations, omissions, or substitutions), and speech fluency standards (assessing whether there are pauses or hesitations during speech). The preset text standards are criteria used to evaluate the quality of text feature data. These can include content completeness standards (checking whether the text conveys certain information, but the content is incomplete and incoherent, potentially omitting key information or details, or causing difficulty in discussing complex topics), content logic standards (judging whether the logic of the text content is reasonable), word choice standards (examining whether word choice is accurate and appropriate), and grammar standards (checking whether the text conforms to grammatical rules). The preset visual standards are criteria used to analyze visual feature data. These can include facial expression standards (judging whether facial expressions are natural and appropriate for the context) and body language standards (assessing whether body movements are coordinated and whether there are any abnormal behaviors).

[0064] Understandably, speech analysis involves inputting extracted speech feature data into a large language model and analyzing it in conjunction with predefined speech standards. For example, the large language model calculates the speech rate and compares it with the specified range in speech rate and intonation standards to determine if the speech rate is too fast or too slow; by analyzing the spectral characteristics of the speech, it assesses whether the sound quality meets the sound quality standards; speech recognition technology converts the speech into text and compares it with standard pronunciation to determine the accuracy of pronunciation; and it counts the number and duration of pauses in the speech to assess the fluency of speech.

[0065] Similarly, text analysis involves inputting text feature data into a large language model and analyzing it according to preset text standards. For example, it might check for the presence of key information to determine content completeness; use logical reasoning algorithms to analyze the logical relationships within the text and assess its logicality; and use lexicons and grammar rule checking tools to evaluate word choice and grammar. Visual analysis, on the other hand, involves inputting visual feature data into a large language model and analyzing it against preset visual standards. For instance, it might use facial expression recognition algorithms to match extracted facial features with preset facial expression standards to determine if facial expressions are normal; or analyze the trajectory and frequency of body movements and compare them with body language standards to assess whether body language is abnormal.

[0066] S203, the speech feature data, text feature data and visual feature data are subjected to time alignment processing and feature fusion processing respectively to obtain a comprehensive multimodal feature vector.

[0067] Specifically, since speech, text, and visual feature data may be asynchronous on the timeline, time alignment processing matches the feature data of these three modalities in chronological order, ensuring their temporal consistency. Furthermore, the time-aligned speech, text, and visual feature data can be integrated to form a comprehensive feature vector, providing a more complete reflection of the patient's condition. Here, the comprehensive multimodal feature vector is the feature vector containing information from speech, text, and vision modalities, obtained after time alignment and feature fusion processing.

[0068] The time alignment process can employ a timestamp-based approach, which involves recording a timestamp for each data point when extracting speech, text, and visual feature data. Then, based on these timestamps, the feature data from different modalities are matched to ensure temporal correspondence. For example, for a sentence in a medical consultation video, its start and end timestamps are recorded, and the speech, text, and visual features within the corresponding time period are identified and aligned.

[0069] For example, when fusing features from different modalities after time alignment, concatenation fusion or weighted fusion can be used. Concatenation fusion directly concatenates the speech, text, and visual feature vectors together to form a longer feature vector. For instance, assuming the speech feature vector has dimension n1, the text feature vector has dimension n2, and the visual feature vector has dimension n3, the combined multimodal feature vector after concatenation has dimension n1 + n2 + n3. Weighted fusion assigns different weights to each modal feature based on their importance, and then adds the weighted feature vectors together. For instance, if the speech feature weight is w1, the text feature weight is w2, and the visual feature weight is w3, and w1 + w2 + w3 = 1, then the combined multimodal feature vector V = w1V1 + w2V2 + w3V3, where V1, V2, and V3 are the speech, text, and visual feature vectors, respectively.

[0070] In some possible embodiments, after the modal features are time-aligned, they can be normalized to eliminate differences in the numerical range of different modal features, thereby improving model execution efficiency and stability.

[0071] S204, based on a large language model and a pre-constructed clinical dementia rating scale, perform dementia symptom diagnosis analysis on the comprehensive multimodal feature vector, and determine the diagnosis result for the patient based on the analysis results, the speech dimension information, the text dimension information, and the visual dimension information.

[0072] Here, the Clinical Dementia Rating Scale (as shown in Table 1) proposed in this disclosure is a targeted improvement based on the Clinical Dementia Rating (CDR) in related technologies. The CDR is a standardized tool widely used in clinical and research fields to assess the severity of dementia. It assesses the patient's cognitive function and daily living abilities from multiple dimensions, including memory, orientation, judgment and problem-solving abilities, social activities, family life and hobbies, and self-care abilities, through detailed interviews with the patient and their family. Professional clinicians determine the corresponding levels based on the assessment results. However, the Clinical Dementia Rating Scale of this disclosure is more aligned with the technology of the large language model itself, fully considering the importance of multimodal data (speech, text, and vision) in dementia diagnosis. It refines and optimizes the descriptions of the core features of each level, enabling more accurate matching and assessment with the multimodal feature information processed by the large language model.

[0073] Table 1

[0074]

[0075]

[0076] Specifically, after receiving a comprehensive multimodal feature vector, the large language model can perform detailed scoring based on the information in the feature vector and compare it with each item in the rating scale. Taking the memory assessment item in the rating scale as an example, the large language model will analyze it from multiple dimensions. In terms of speech dimension information, it focuses on the fluency and completeness of the patient's speech. For example, if the patient frequently pauses and repeats when answering questions, and is unable to narrate a matter completely, it may indicate a memory problem. The large language model can analyze the speech signal, count the number of pauses, the frequency of repetitions, and the coherence of the content, and compare these indicators with the standards of the memory-related levels in the rating scale to give a corresponding score. For example, if the number of pauses exceeds a certain threshold and the content coherence is poor, it may correspond to a CDR1 (mild dementia) or higher level memory score.

[0077] Furthermore, regarding textual information, large language models focus on examining the text's logical consistency. If the text exhibits increased grammatical errors, short sentences, or logical incoherence, it may reflect a decline in memory and cognitive abilities. Here, large language models can utilize pre-trained language models, such as DeepSeek (e.g., DeepSeek-R1), for semantic understanding and logical analysis of the text. DeepSeek-R1, based on the Retentive Network architecture, combines the advantages of Transformer and state-space models, enabling it to efficiently capture long-distance dependencies in text. Simultaneously, through dynamic gating mechanisms, this model further enhances its ability to model the text's logical structure, enabling more accurate parsing of semantic levels and logical connections within the text. DeepSeek-R1 can also assess the logicality of text through the Contextual Coherence Score (CCS). This metric comprehensively considers multiple dimensions, including semantic coherence, which uses contrastive learning-optimized embedding similarity to measure the naturalness and fluency of semantic connections between sentences; syntactic rationality, which uses a lightweight syntax tree parser to detect abnormal structures in the text to ensure sentences conform to grammatical rules; and information density, which applies information entropy-based local and global consistency analysis to assess the reasonableness of the distribution of information in the text, avoiding redundancy or omissions. Thus, based on the mapping relationship between CCS values ​​and clinical dementia rating scales, the large language model can generate more accurate memory and logicality scores, improving the accuracy of text quality assessment results.

[0078] Furthermore, regarding visual information, the large language model combines facial expressions and body language to assess a patient's memory status. For example, if a patient's expression is blank, their eyes are unfocused, and their body movements are uncoordinated when answering questions, this may be related to cognitive impairment caused by memory decline. Here, the large model can use a trained CNN model to recognize different facial expressions (such as smiling, frowning, blankness, etc.) and body movements (such as gestures, sitting posture, walking posture, etc.), and compare this information with different levels of visual feature performance on a rating scale to provide corresponding scores.

[0079] Understandably, after scoring each item on the rating scale, the large language model comprehensively considers all scoring results, as well as other relevant features from the speech, text, and visual dimensions, and uses its internal decision-making mechanism to determine the patient's diagnosis. For example, if the patient scores highly in memory, orientation, and calculation abilities, and their speech features show normal speaking speed, clear pronunciation, and fluent speech; their text features show accurate expression and strong logic; and their visual features show natural facial expressions and coordinated body movements, then the large language model will comprehensively determine the patient to be CDR0 (no dementia). Conversely, if the patient scores low on multiple assessment items, and their speech features show a significantly slower speaking speed and prominent recent memory loss; their text features show increased grammatical errors and weakened logic; and their visual features show difficulty with time orientation and apathy, then the large language model will determine the patient to be CDR1 (mild dementia) or a higher level.

[0080] In some possible embodiments, when performing dementia symptom diagnostic analysis on comprehensive multimodal feature vectors based on a large language model and a pre-built clinical dementia rating scale, the following (1) to (2) may also be included:

[0081] (1) Input the comprehensive multimodal feature vector into the large language model, perform semantic analysis on the comprehensive multimodal feature vector based on the large language model, and generate a dementia level probability distribution based on the analysis results and the pre-constructed clinical dementia rating scale;

[0082] (2) Determine the dementia severity classification label of the patient based on the probability distribution of the dementia level.

[0083] Specifically, after inputting the comprehensive multimodal feature vector into a large language model, the model performs semantic analysis. Large language models are typically based on deep learning architectures, such as the Transformer architecture, to analyze the relationships between elements in the input vector and uncover potential semantic information. After semantic analysis, the model matches and compares the results with features in the Clinical Dementia Rating Scale (CDR). For example, if the semantic analysis shows that the patient exhibits slower speech rate and increased repetitive speech in speech, increased grammatical errors and weakened logic in text, and apathy and poor coordination in visual aspects, the model will quantitatively compare these features with the features of each level in the CDR. Machine learning algorithms, such as logistic regression or neural networks, are used to calculate the probability of the patient being in each dementia level. Ultimately, a dementia level probability distribution is generated, containing numerical values ​​indicating the likelihood of the patient being in each dementia level. These values, presented as probabilities, more accurately reflect the uncertainty and diversity of the patient's dementia symptoms.

[0084] Furthermore, after obtaining the probability distribution of dementia grades, a classification label for the severity of dementia in patients can be determined based on this distribution. The probability distribution of dementia grades provides the probability values ​​for patients to fall into different dementia grades, and these probability values ​​reflect the degree to which the patient's symptoms fit each dementia grade. For example, the dementia grade with the highest probability is selected as the classification label. For instance, if the probability distribution of dementia grades shows that the probability of the patient being in CDR1 is 0.4, the probability of being in CDR0.5 is 0.3, the probability of being in CDR2 is 0.2, and the probabilities of other grades are even lower, then according to the principle of maximizing probability, the classification label for the patient's dementia severity would be determined as CDR1.

[0085] In some other embodiments, since the probability distribution of dementia grades may not always be clear, multiple grades may have similar probabilities. In such cases, a probability threshold can be set, and a grade is only considered a candidate classification label when its probability exceeds this threshold. If multiple grades have probabilities exceeding the threshold, a comprehensive judgment can be made by combining clinical experience and professional knowledge. For example, setting the probability threshold to 0.3, when the probability of a patient being in CDR1 is 0.35 and the probability of being in CDR0.5 is 0.32, although both probabilities exceed the threshold, considering the patient's medical history, family genetic factors, and other clinical information, the doctor may be more inclined to choose CDR1 as the final dementia severity classification label. In this way, the severity of the patient's dementia can be determined more accurately, providing a reliable basis for subsequent diagnosis and treatment.

[0086] Specifically, when determining the diagnosis of a patient based on the analysis results, voice dimension information, text dimension information, and visual dimension information, a diagnosis report can be generated based on the classification labels of dementia severity, combined with voice dimension information, text dimension information, and visual dimension information, to serve as the diagnosis of the patient.

[0087] The diagnostic report can list dementia severity classification labels, allowing doctors and patients to intuitively understand the patient's dementia level. It can also provide detailed descriptions of information across three dimensions: speech, text, and vision—a multimodal information description. Regarding the speech dimension, it records the patient's speech rate, such as the number of words spoken per minute; changes in pitch, including any abnormal increases or decreases in pitch; the frequency and duration of pauses; and the fluency and coherence of speech. For example, it might describe the patient as having "a 35% decrease in speech rate (within the normal reference range), decreased vocal clarity, frequent speech interruptions and prolongations, inaccurate pronunciation, unclear or incorrect pronunciation of some words, significant repetition of words, using mostly simple short sentences when answering questions, rarely initiating conversations, but generally able to understand others' speech."

[0088] Similarly, the text-level information description will cover vocabulary usage, such as the richness and specialization of the vocabulary; grammatical correctness, including the existence, type, and frequency of grammatical errors; semantic coherence and logic, whether the connection between sentences is natural, and whether the content is logical. For example, "The content is incoherent, the connection between sentences is poor, there are often contradictions or irrelevant situations, the word choice is inaccurate, vague words are often used to replace specific words, there are many grammatical errors, such as incomplete sentence structure and confused tenses, the frequency of first-person pronouns exceeds the standard by 1.8 times, and there are more typos and omissions in written expression."

[0089] Furthermore, descriptions of visual information include changes in facial expressions, such as whether smiles, frowns, or blank expressions frequently occur; eye focus, and whether the patient can maintain good eye contact with the doctor; and coordination of bodily movements, such as whether gestures are natural and gait is stable. For example, "The patient's facial expression is stiff, eye focus is low, they frequently look around, they make little eye contact with the doctor, their bodily coordination is poor, their gait is unsteady, and their gestures are unnatural."

[0090] Understandably, in addition to multimodal information descriptions, diagnostic reports can also provide quantitative indicators of abnormal features corresponding to each modality. These indicators are quantitative measurements of abnormal features in each modality, and can more accurately reflect the degree of abnormality. For example, for speech rate in the speech dimension, its deviation from normal speech rate can be calculated; for grammatical error frequency in the text dimension, the number of grammatical errors occurring per 100 words can be counted; for limb coordination in the visual dimension, motion analysis algorithms can be used to calculate the trajectory deviation of limb joints, etc.

[0091] Finally, the diagnostic report also includes a clinical feature checklist based on a pre-built Clinical Dementia Rating Scale (CDR). This checklist compares the patient's multimodal information with the features in the CDR, clearly showing the correspondence between the patient's performance in each feature and the dementia level. For example, in terms of speech features, a patient exhibits slower speech rate and increased repetitive speech, which corresponds to the speech features of CDR1. In terms of text features, the patient shows increased grammatical errors and weakened logical thinking, corresponding to the text features of CDR1. In terms of visual features, the patient exhibits apathy and uncoordinated movements, corresponding to the visual features of CDR1. Through the clinical feature checklist, doctors and patients can more intuitively understand the relationship between the patient's symptoms and the dementia level.

[0092] In some possible embodiments, the large language model can also provide corresponding suggestions and interventions in the diagnostic results. For example, for patients with suspected dementia (CDR0.5), further cognitive function examinations and regular follow-ups are recommended; for patients with mild dementia (CDR1), interventions such as cognitive training and drug treatment are recommended to slow disease progression. In this way, the multimodal dementia analysis-assisted diagnostic method of this disclosure can not only accurately diagnose the dementia status of patients, but also provide targeted suggestions for clinical treatment and rehabilitation.

[0093] Understandably, dementia is a chronic, progressive disease, meaning its progression is often slow and continuous, potentially exhibiting a gradual development over a long period. Furthermore, the rate of progression and manifestations vary significantly among different patients. Additionally, considering that some patients may have difficulty making frequent in-person follow-up appointments due to mobility issues or remote living conditions, patients can use mobile devices to provide information via voice or text to track their condition and facilitate follow-up. Voice messages can include the patient's self-reported emotional state, sleep, and eating habits; descriptions of recent physical discomfort such as dizziness, weakness, and numbness in the limbs; or feedback on changes in cognitive function, such as whether memory loss has worsened or the patient's ability to concentrate. Specifically, refer to... Figure 4 As shown, after determining the diagnosis of the patient, the following steps S401 to S404 may also be included:

[0094] S401, according to the preset follow-up visit cycle, receive follow-up visit information provided by the patient's terminal device.

[0095] Here, the preset follow-up visit cycle can be set comprehensively based on the characteristics of dementia and clinical treatment experience. For early-stage dementia patients, the follow-up visit cycle may be relatively long, such as once every three months; while for mid-to-late-stage patients, due to the potentially rapid changes in their condition, the follow-up visit cycle will be shortened accordingly, possibly once a month. The patient's terminal device can be a smartphone, tablet, or other device with data transmission capabilities. When the preset follow-up visit time arrives, a reminder message can be sent to the patient's terminal device. The patient can then provide text-based follow-up visit data through an application on their terminal device, covering recent physical condition, changes in lifestyle, etc. Alternatively, they can use the voice function to directly describe relevant information, generating voice-based follow-up visit data which is then uploaded to the system. After receiving this data, preliminary format verification and data integrity checks can be performed to ensure that the received follow-up visit information meets the requirements for subsequent analysis and processing.

[0096] S402, perform text analysis on the text follow-up data based on the large language model and preset speech standards to obtain text follow-up information; and perform speech analysis on the speech follow-up data based on the large language model and preset speech standards to obtain speech follow-up information.

[0097] Here, the text analysis process for text-based follow-up consultation data and the voice analysis process for voice-based follow-up consultation data are based on the same technical principles as steps S301 to S302 above. For details, please refer to steps S301 to S302, which will not be repeated here.

[0098] S403, perform time alignment processing and feature fusion processing on the text follow-up data and the voice follow-up data respectively to obtain a comprehensive multimodal follow-up feature vector.

[0099] The specific implementation process of multimodal feature fusion is the same as the technical principle of step S203 above. For details, please refer to step S203, and it will not be repeated here.

[0100] S404, based on a large language model and a pre-built clinical dementia rating scale, perform dementia symptom diagnosis analysis on the comprehensive multimodal follow-up feature vector, and determine the follow-up diagnosis result for the patient based on the analysis results, the voice follow-up information, the text follow-up information, and the diagnosis result.

[0101] Here, the comprehensive multimodal follow-up feature vector is input into a large language model. The large language model performs semantic analysis and feature interpretation, and, combined with a pre-built clinical dementia rating scale, assesses the severity of the patient's current dementia symptoms. Then, a comprehensive judgment is made by considering the analysis results, voice follow-up information, text follow-up information, and the initial diagnosis. For example, if the patient reports significant recent memory decline in the voice follow-up information, repeatedly mentions forgetting daily tasks in the text follow-up information, and the comprehensive multimodal follow-up feature vector analysis shows a high probability of mild dementia (CDR1) in the dementia grade probability distribution, while the initial diagnosis is suspected dementia (CDR0.5), the doctor can determine that the patient's condition may have progressed, confirm the follow-up diagnosis as mild dementia (CDR1), and formulate an appropriate treatment and care plan.

[0102] In some possible embodiments, the multimodal dementia analysis-assisted diagnosis method mentioned in this disclosure can also be applied to the diagnosis of different diseases, such as the auxiliary diagnosis of neuropsychiatric disorders like Parkinson's disease, autism spectrum disorder, and depression. By extracting and analyzing features from multimodal data such as video, audio, and text in different disease consultation scenarios, and based on the corresponding clinical assessment scales and diagnostic criteria for various diseases, a large language model is used for multidimensional information processing and comprehensive judgment to achieve accurate assessment and diagnostic analysis of disease symptoms. That is, the application scope of this method is not limited to dementia-related diseases; it has the ability to be extended to various disease diagnosis scenarios with multimodal data acquisition conditions and the ability to construct corresponding disease assessment systems, and is not specifically limited here.

[0103] The multimodal dementia analysis-assisted diagnostic method, system, device, medium, and equipment provided in this disclosure integrate speech, text, and visual multimodal information and utilize the analysis of comprehensive multimodal data to more comprehensively and accurately capture abnormal manifestations of patients in language, cognition, and vision. Furthermore, by using large language models and clinical dementia rating scales for comprehensive analysis, the accuracy and objectivity of dementia diagnosis are effectively improved.

[0104] Those skilled in the art will understand that, in the above-described method of the specific implementation, the order in which each step is written does not imply a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.

[0105] Based on the same inventive concept, this disclosure also provides a multimodal dementia analysis-assisted diagnosis system corresponding to the multimodal dementia analysis-assisted diagnosis method. Since the principle of the system in this disclosure for solving the problem is similar to the multimodal dementia analysis-assisted diagnosis method described above, the implementation of the system can refer to the implementation of the method, and the repeated parts will not be described again.

[0106] Reference Figure 5 The diagram shown is a schematic of a multimodal dementia analysis-assisted diagnostic system provided in an embodiment of this disclosure. The system includes:

[0107] The video capture module is used to acquire videos of patients' consultations.

[0108] A multimodal feature extraction module, connected to the video acquisition module, is used to perform multimodal feature extraction processing on the consultation video to obtain a multimodal feature dataset; wherein, the multimodal feature dataset includes speech feature data, text feature data, and visual feature data;

[0109] A multi-dimensional analysis module is connected to the multimodal feature extraction module and the large language model, respectively, and is used to perform speech analysis on the speech feature data based on the large language model and a preset speech standard to obtain speech dimension information; and to perform text analysis on the text feature data based on the large language model and a preset text standard to obtain text dimension information; and to perform visual analysis on the visual feature data based on the large language model and a preset visual standard to obtain visual dimension information.

[0110] The feature processing module, connected to the multimodal feature extraction module, is used to perform time alignment processing and feature fusion processing on the speech feature data, text feature data and visual feature data respectively to obtain a comprehensive multimodal feature vector;

[0111] The diagnostic analysis module is connected to the feature processing module, the multi-dimensional analysis module, and the large language model, respectively. It is used to perform dementia symptom diagnostic analysis on the comprehensive multimodal feature vector based on the large language model and a pre-built clinical dementia rating scale, and to determine the diagnosis result of the patient based on the analysis results, the voice dimension information, the text dimension information, and the visual dimension information.

[0112] Based on the same inventive concept, this disclosure also provides a multimodal dementia analysis-assisted diagnostic device corresponding to the multimodal dementia analysis-assisted diagnostic method. Since the principle of the device in this disclosure for solving the problem is similar to the multimodal dementia analysis-assisted diagnostic method described above, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be described again.

[0113] Reference Figure 6 The diagram shown is a schematic of a multimodal dementia analysis-assisted diagnostic device 600 provided in an embodiment of this disclosure. The device includes:

[0114] The multimodal feature extraction module 601 is used to acquire the patient's consultation video and perform multimodal feature extraction processing on the consultation video to obtain a multimodal feature dataset; wherein, the multimodal feature dataset includes speech feature data, text feature data and visual feature data;

[0115] The multimodal dimension analysis module 602 is used to perform speech analysis on the speech feature data based on a large language model and a preset speech standard to obtain speech dimension information; and to perform text analysis on the text feature data based on the large language model and a preset text standard to obtain text dimension information; and to perform visual analysis on the visual feature data based on the large language model and a preset visual standard to obtain visual dimension information.

[0116] The multimodal feature fusion module 603 is used to perform time alignment processing and feature fusion processing on the speech feature data, text feature data and visual feature data respectively to obtain a comprehensive multimodal feature vector;

[0117] The diagnostic result determination module 604 is used to perform dementia symptom diagnostic analysis on the comprehensive multimodal feature vector based on a large language model and a pre-built clinical dementia rating scale, and to determine the diagnostic result for the patient based on the analysis results, the speech dimension information, the text dimension information and the visual dimension information.

[0118] In some possible embodiments, the multimodal feature extraction module 601 is specifically used for:

[0119] Extract the audio signal from the consultation video, extract the audio feature vector from the audio signal, and obtain the audio feature data;

[0120] Based on the speech signal, extract text information about the consultation video, and encode the text information to obtain the text feature data;

[0121] Multiple keyframe images from the consultation video are extracted at preset time intervals, and features are extracted from each keyframe image. The feature extraction results of each keyframe image are then aggregated into time series features to obtain the visual feature data.

[0122] In some possible embodiments, the preset speech standard includes a speech rate and intonation standard, a sound quality standard, a pronunciation accuracy standard, and a speech fluency standard;

[0123] The preset text standards include content integrity standards, content logic standards, word choice standards, and grammar standards;

[0124] The preset visual standards include facial expression standards and body language standards.

[0125] In some possible embodiments, the diagnostic result determination module 604 is specifically used for:

[0126] The comprehensive multimodal feature vector is input into the large language model, semantic analysis is performed on the comprehensive multimodal feature vector based on the large language model, and a dementia level probability distribution is generated based on the analysis results and the pre-constructed clinical dementia rating scale.

[0127] The severity of dementia in the patients is classified and labeled based on the probability distribution of the dementia level.

[0128] In some possible embodiments, the diagnostic result includes a diagnostic report; the diagnostic result determination module 604 is specifically used for:

[0129] Based on the dementia severity classification labels, combined with the voice dimension information, the text dimension information, and the visual dimension information, a diagnostic report for the patient is generated; wherein, the diagnostic report includes dementia severity classification labels, multimodal information descriptions and abnormal feature quantification indicators corresponding to each modality, as well as a clinical feature comparison table based on a pre-constructed clinical dementia rating scale.

[0130] In some possible embodiments, the diagnostic result determination module 604 is further configured to:

[0131] According to a preset follow-up visit cycle, the system receives follow-up visit information from the patient's terminal device; wherein, the follow-up visit information includes text follow-up visit data and voice follow-up visit data.

[0132] Based on the large language model and the preset speech standard, the text follow-up data is analyzed to obtain text follow-up information; and based on the large language model and the preset speech standard, the speech follow-up data is analyzed to obtain speech follow-up information.

[0133] The text-based follow-up consultation data and the voice-based follow-up consultation data are subjected to time alignment processing and feature fusion processing respectively to obtain a comprehensive multimodal follow-up consultation feature vector;

[0134] Based on a large language model and a pre-built clinical dementia rating scale, the comprehensive multimodal follow-up feature vector is analyzed for dementia symptom diagnosis. Based on the analysis results, the voice follow-up information, the text follow-up information, and the diagnosis results, the follow-up diagnosis result for the patient is determined.

[0135] Based on the same technical concept, this disclosure also provides a computer device. (See also...) Figure 7 The diagram shows the structure of a computer device 700 provided in this embodiment of the present disclosure, including a processor 701, a memory 702, and a bus 703. The memory 702 stores execution instructions and includes a main memory 7021 and an external memory 7022. The main memory 7021, also called internal memory, is used to temporarily store computational data in the processor 701, as well as data exchanged with external memory 7022 such as a hard disk. The processor 701 exchanges data with the external memory 7022 through the main memory 7021.

[0136] In this embodiment, the memory 702 is specifically used to store application code that executes the solution of this application, and its execution is controlled by the processor 701. That is, when the computer device 700 is running, the processor 701 communicates with the memory 702 through the bus 703, so that the processor 701 executes the application code stored in the memory 702, and then executes the method described in any of the foregoing embodiments.

[0137] The memory 702 may be, but is not limited to, random access memory (RAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), etc.

[0138] Processor 701 may be an integrated circuit chip with signal processing capabilities. The aforementioned processor can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this invention. The general-purpose processor can be a microprocessor or any conventional processor.

[0139] It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the computer device 700. In other embodiments of this application, the computer device 700 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0140] This disclosure also provides a computer-readable storage medium storing a computer program that, when executed by a processor, performs the steps of the multimodal dementia analysis-assisted diagnosis method described in the above-described method embodiments. The storage medium can be a volatile or non-volatile computer-readable storage medium.

[0141] This disclosure also provides a computer program product carrying program code. The program code includes instructions that can be used to execute the steps of the multimodal dementia analysis-assisted diagnosis method described in the above method embodiments. For details, please refer to the above method embodiments, which will not be repeated here.

[0142] The aforementioned computer program product can be implemented through hardware, software, or a combination thereof. In one optional embodiment, the computer program product is specifically embodied in a computer storage medium; in another optional embodiment, the computer program product is specifically embodied in a software product, such as a software development kit (SDK), etc.

[0143] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems and devices described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. In the several embodiments provided in this disclosure, it should be understood that the disclosed systems and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division; in actual implementation, there may be other division methods. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection may be through some communication interfaces; the indirect coupling or communication connection of devices or units may be electrical, mechanical, or other forms.

[0144] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0145] In addition, the functional units in the various embodiments of this disclosure can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0146] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a processor-executable, non-volatile, computer-readable storage medium. Based on this understanding, the technical solution of this disclosure, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this disclosure. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0147] Finally, it should be noted that the above-described embodiments are merely specific implementations of this disclosure, used to illustrate the technical solutions of this disclosure, and not to limit it. The protection scope of this disclosure is not limited thereto. Although this disclosure has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features, within the scope of the technology disclosed in this disclosure. Such modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this disclosure, and should all be covered within the protection scope of this disclosure. Therefore, the protection scope of this disclosure should be determined by the protection scope of the claims.

Claims

1. A multimodal dementia analysis-assisted diagnostic method, characterized in that, include: The patient's consultation video is acquired, and multimodal feature extraction processing is performed on the consultation video to obtain a multimodal feature dataset; wherein, the multimodal feature dataset includes speech feature data, text feature data, and visual feature data; Speech analysis is performed on the speech feature data based on the large language model and the preset speech standard to obtain speech dimension information; and text analysis is performed on the text feature data based on the large language model and the preset text standard to obtain text dimension information; and visual analysis is performed on the visual feature data based on the large language model and the preset visual standard to obtain visual dimension information. The speech feature data, text feature data, and visual feature data are subjected to time alignment and feature fusion processing respectively to obtain a comprehensive multimodal feature vector; The comprehensive multimodal feature vectors are analyzed for dementia symptom diagnosis based on a large language model and a pre-constructed clinical dementia rating scale. Based on the analysis results, the speech dimension information, the text dimension information, and the visual dimension information, the diagnosis result for the patient is determined.

2. The method according to claim 1, characterized in that, The process of extracting multimodal features from the consultation video to obtain a multimodal feature dataset includes: Extract the audio signal from the consultation video, extract the audio feature vector from the audio signal, and obtain the audio feature data; Based on the speech signal, extract text information about the consultation video, and encode the text information to obtain the text feature data; Multiple keyframe images from the consultation video are extracted at preset time intervals, and features are extracted from each keyframe image. The feature extraction results of each keyframe image are then aggregated into time series features to obtain the visual feature data.

3. The method according to claim 1, characterized in that, The preset speech standards include speech rate and intonation standards, sound quality standards, pronunciation accuracy standards, and speech fluency standards; The preset text standards include content integrity standards, content logic standards, word choice standards, and grammar standards; The preset visual standards include facial expression standards and body language standards.

4. The method according to claim 1, characterized in that, The dementia symptom diagnostic analysis based on the comprehensive multimodal feature vector using a large language model and a pre-constructed clinical dementia rating scale includes: The comprehensive multimodal feature vector is input into the large language model, semantic analysis is performed on the comprehensive multimodal feature vector based on the large language model, and a dementia level probability distribution is generated based on the analysis results and the pre-constructed clinical dementia rating scale. The severity of dementia in the patients is classified and labeled based on the probability distribution of the dementia level.

5. The method according to claim 4, characterized in that, The diagnostic results include a diagnostic report; determining the diagnostic results for the patient based on the analysis results, the voice dimension information, the text dimension information, and the visual dimension information includes: Based on the dementia severity classification labels, combined with the voice dimension information, the text dimension information, and the visual dimension information, a diagnostic report for the patient is generated; wherein, the diagnostic report includes dementia severity classification labels, multimodal information descriptions and abnormal feature quantification indicators corresponding to each modality, as well as a clinical feature comparison table based on a pre-constructed clinical dementia rating scale.

6. The method according to any one of claims 1 to 5, characterized in that, After determining the diagnosis result for the patient, the process further includes: According to a preset follow-up visit cycle, the system receives follow-up visit information from the patient's terminal device; wherein, the follow-up visit information includes text follow-up visit data and voice follow-up visit data. Based on the large language model and the preset speech standard, the text follow-up data is analyzed to obtain text follow-up information; and based on the large language model and the preset speech standard, the speech follow-up data is analyzed to obtain speech follow-up information. The text-based follow-up consultation data and the voice-based follow-up consultation data are subjected to time alignment processing and feature fusion processing respectively to obtain a comprehensive multimodal follow-up consultation feature vector; Based on a large language model and a pre-built clinical dementia rating scale, the comprehensive multimodal follow-up feature vector is analyzed for dementia symptom diagnosis. Based on the analysis results, the voice follow-up information, the text follow-up information, and the diagnosis results, the follow-up diagnosis result for the patient is determined.

7. A multimodal dementia analysis-assisted diagnostic system, characterized in that, include: The video capture module is used to acquire videos of patients' consultations. A multimodal feature extraction module, connected to the video acquisition module, is used to perform multimodal feature extraction processing on the consultation video to obtain a multimodal feature dataset; wherein, the multimodal feature dataset includes speech feature data, text feature data, and visual feature data; A multi-dimensional analysis module is connected to the multimodal feature extraction module and the large language model, respectively, and is used to perform speech analysis on the speech feature data based on the large language model and a preset speech standard to obtain speech dimension information; and to perform text analysis on the text feature data based on the large language model and a preset text standard to obtain text dimension information; and to perform visual analysis on the visual feature data based on the large language model and a preset visual standard to obtain visual dimension information. The feature processing module, connected to the multimodal feature extraction module, is used to perform time alignment processing and feature fusion processing on the speech feature data, text feature data and visual feature data respectively to obtain a comprehensive multimodal feature vector; The diagnostic analysis module is connected to the feature processing module, the multi-dimensional analysis module, and the large language model, respectively. It is used to perform dementia symptom diagnostic analysis on the comprehensive multimodal feature vector based on the large language model and a pre-built clinical dementia rating scale, and to determine the diagnosis result of the patient based on the analysis results, the voice dimension information, the text dimension information, and the visual dimension information.

8. A multimodal dementia analysis-assisted diagnostic device, characterized in that, include: A multimodal feature extraction module is used to acquire the patient's consultation video and perform multimodal feature extraction processing on the consultation video to obtain a multimodal feature dataset; wherein, the multimodal feature dataset includes speech feature data, text feature data and visual feature data; The multimodal dimension analysis module is used to perform speech analysis on the speech feature data based on a large language model and a preset speech standard to obtain speech dimension information; and to perform text analysis on the text feature data based on the large language model and a preset text standard to obtain text dimension information; and to perform visual analysis on the visual feature data based on the large language model and a preset visual standard to obtain visual dimension information. The multimodal feature fusion module is used to perform time alignment processing and feature fusion processing on the speech feature data, text feature data and visual feature data respectively to obtain a comprehensive multimodal feature vector; The diagnostic result determination module is used to perform dementia symptom diagnostic analysis on the comprehensive multimodal feature vector based on a large language model and a pre-built clinical dementia rating scale, and to determine the diagnostic result for the patient based on the analysis results, the speech dimension information, the text dimension information, and the visual dimension information.

9. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 6.

10. A computer device, comprising a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method of any one of claims 1 to 6.

Citation Information

Cited By

  • AI-driven virtual doctor conference system and neurodevelopmental disorder assessment method

    CN121281844A