Device for generating mental health report based on multi-modal model and storage medium
Through the feature extraction and cross-modal attention mechanism of multimodal models, combined with fundus images, questionnaire and scale information, mental health reports are generated, which solves the subjectivity and efficiency of mental health assessment in the existing technology, and achieves more accurate and detailed mental health identification and report generation.
Patent Information
- Application Number
- CN202411982495.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-16
AI Technical Summary
The existing technology has problems such as high subjectivity, low evaluation efficiency, and difficulty in achieving large-scale screening and timely intervention in mental health assessments. The existing methods have failed to make full use of the potential connections between multimodal data and lack in-depth explanations and detailed support.
Using a multimodal model, fundus features and questionnaire features are extracted through multiple feature extraction modules, combined with scale information, a cross-modal fusion module is used to perform a cross-modal attention mechanism, obtain multimodal features, and generate a mental health report through a report generation module.
It improves the accuracy of mental health recognition and generates a complete mental health report, which can provide risk classification and specific scoring results for mental health problems such as anxiety and depression, helps patients and doctors better understand the identification results, and has important clinical application potential.
Smart Images

Figure CN120015304A_ABST
Abstract
Description
Technical Field
[0001] The present application generally relates to the field of multimodal data recognition technology. More specifically, the present application relates to a device for generating a mental health report using a multimodal model and a computer-readable storage medium. Background Art
[0002] In today's society, mental health problems such as depression and anxiety are becoming increasingly serious and have become an important public health issue worldwide. Early diagnosis and intervention of such mental illnesses are of great significance for improving the quality of life of patients. Traditional mental health assessments mainly rely on self-reports from patients and interviews and scale assessments by professional doctors. There is a certain degree of subjectivity, and the assessment efficiency is low, making it difficult to achieve large-scale screening and timely intervention. In recent years, with the development of deep learning and artificial intelligence technologies, especially breakthroughs in computer vision, natural language processing and other fields, intelligent recognition based on multimodal data has become an emerging solution.
[0003] As a non-invasive biomarker, fundus images can reflect a variety of health information of the human body. Studies have shown that fundus images can not only be used for the diagnosis of ophthalmic diseases, but also provide some information related to mental health. For example, changes in the microvascular structure of the retina may be associated with psychological diseases such as depression and anxiety. In addition, combined with the self-reported data of patients in the questionnaire scale, the prediction accuracy of mental health status can be further improved. However, existing methods are usually limited to the use of single modality data, or only combine multimodal data in a simple way, and fail to fully utilize the potential connections between different modalities. In addition, most of the current mental health risk prediction models are based on simple binary classification results, such as only predicting whether the patient has anxiety or depression symptoms. This prediction method lacks deeper explanations and detailed support, making it difficult to provide patients with personalized treatment recommendations or further diagnostic references.
[0004] In view of this, there is an urgent need to provide a solution for generating mental health reports using a multimodal model in order to improve the accuracy of mental health identification and obtain complete mental health reports. Summary of the invention
[0005] In order to at least solve one or more of the technical problems mentioned above, the present application proposes a solution for generating a mental health report using a multimodal model in multiple aspects.
[0006] In a first aspect, the present application provides a device for generating a mental health report based on a multimodal model, wherein the multimodal model includes multiple feature extraction modules, a cross-modal fusion module and a report generation module, and the device includes: a processor; and a memory on which computer instructions for generating a mental health report based on the multimodal model are stored, and when the computer instructions are executed by the processor, the device performs the following operations: collecting multimodal data and preprocessing the multimodal data, wherein the multimodal data at least includes a fundus image, scale information and questionnaire information related to mental health; based on the fundus image and the questionnaire information, using the multiple feature extraction modules to respectively extract fundus features and questionnaire features to obtain a fundus feature vector and a questionnaire feature vector, and representing the scale information as a scale feature vector; based on the fundus feature vector, the scale feature vector and the questionnaire feature vector, using the cross-modal fusion module to perform a cross-modal attention mechanism to obtain multimodal features; and inputting the multimodal features into the report generation module for a report generation operation to generate a mental health report.
[0007] In some embodiments, the preprocessing includes at least one or more of quality screening, image cropping, image size adjustment, or image color correction.
[0008] In some other embodiments, wherein the multiple feature extraction modules include a first feature module and a second feature module, the device further performs the following operations to obtain a fundus feature vector and a questionnaire feature vector: based on the fundus image, using the first feature module to extract fundus features to obtain the fundus feature vector; and based on the questionnaire information, using the second feature module to extract fundus features to obtain the questionnaire feature vector.
[0009] In some other embodiments, the device further performs the following operations to obtain the questionnaire feature vector: based on the questionnaire information, using the second feature module to extract the context semantic vector of each word in the answer text, the global vector of the question text and the word pair relationship matrix in the answer text in the questionnaire information; calculating the attention weight of each word in the answer text according to the context semantic vector and the global vector; and calculating the weighted sum according to the context semantic vector, the global vector, the word pair relationship matrix and the attention weight to obtain the questionnaire feature vector.
[0010] In some further embodiments, the first feature module comprises a self-supervised learning module, and the second feature module comprises a bidirectional recurrent network module.
[0011] In some other embodiments, the device further performs the following operations to represent the scale information as a scale feature vector: calculates the score of each question in the scale information; and normalizes the score of each question in the scale information to represent the scale information as the scale feature vector.
[0012] In some other embodiments, the device further performs the following operations to obtain multimodal features: merging the fundus feature vector, the scale feature vector and the questionnaire feature vector into a total feature vector for initialization; and based on the initialized total feature vector, using the cross-modal fusion module to execute a cross-modal attention mechanism and perform recursive updates to obtain the multimodal features.
[0013] In some other embodiments, the report generation module includes a first classification module and a natural language processing module, and the device further performs the following operations to generate a mental health report: based on the multimodal features, using the first classification module to perform a classification operation to obtain a target classification result; inputting the target classification result into the natural language processing module for parsing to obtain report text information; and combining the report text information with the medical knowledge base to generate the mental health report.
[0014] In some other embodiments, the device further performs the following operations: inputting the multimodal features into a second classification module for preliminary classification to obtain a preliminary classification result; comparing the preliminary classification result with the historical evaluation result to generate a feedback signal; optimizing the second classification module according to the feedback signal to obtain a final classification result based on the optimized second classification module; and combining the report text information, the medical knowledge base and the final classification result to generate a final mental health report.
[0015] In a second aspect, the present application provides a computer-readable storage medium, which includes computer program instructions for generating a mental health report based on a multimodal model. When the computer program instructions are executed by one or more processors, the operations performed by the device described in the first aspect are implemented.
[0016] Through a scheme for generating a mental health report based on a multimodal model as provided above, the embodiment of the present application extracts fundus features and questionnaire features through multiple feature extraction modules in the multimodal model to obtain fundus feature vectors and questionnaire feature vectors. Then, by combining the fundus feature vector, the questionnaire feature vector and the scale feature vector, a cross-modal fusion module is used to perform a cross-modal attention mechanism to capture the complex relationship between multimodal data and obtain multimodal features containing rich information. Furthermore, a report generation operation is performed by inputting the multimodal features into a report generation module to generate a mental health report. This can not only obtain risk classifications of mental health problems such as anxiety and depression, and improve the accuracy of mental health identification; it can also generate specific scoring results and give explanations to generate a completeness report to help patients and doctors better understand the identification results. It has important clinical application potential and is helpful for early screening and intervention of mental health problems. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] By reading the detailed description below with reference to the accompanying drawings, the above and other purposes, features and advantages of the exemplary embodiments of the present application will become easy to understand. In the accompanying drawings, several embodiments of the present application are shown in an exemplary and non-limiting manner, and the same or corresponding reference numerals represent the same or corresponding parts, wherein:
[0018] Figure 1 is an exemplary structural block diagram showing an apparatus for generating a mental health report based on a multimodal model according to an embodiment of the present application;
[0019] Figure 2 is an exemplary flowchart showing operations performed by an apparatus for generating a mental health report based on a multimodal model according to an embodiment of the present application;
[0020] Figure 3 is an exemplary schematic diagram showing a mental health report generated based on a multimodal model according to an embodiment of the present application;
[0021] Figure 4 It is an exemplary structural block diagram showing a device for generating a mental health report based on a multimodal model according to an embodiment of the present application. DETAILED DESCRIPTION
[0022] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present application.
[0023] It should be understood that the terms "include" and "comprising" used in the specification and claims of the present application indicate the presence of described features, wholes, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or collections thereof.
[0024] It should also be understood that the terms used in this application specification are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in this application specification and claims, unless the context clearly indicates otherwise, the singular forms of "a", "an" and "the" are intended to include plural forms. It should also be further understood that the term "and / or" used in this application specification and claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes these combinations.
[0025] As used in this specification and claims, the term "if" may be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" may be interpreted as meaning "upon determination" or "in response to determining" or "upon detection of [described condition or event]" or "in response to detecting [described condition or event]," depending on the context.
[0026] The specific implementation of the present application is described in detail below with reference to the accompanying drawings.
[0027] Figure 1 1 is an exemplary structural block diagram showing an apparatus 100 for generating a mental health report based on a multimodal model according to an embodiment of the present application. Figure 1 As shown in , the device 100 may include a processor 101 and a memory 102. The processor 101 may include, for example, a general-purpose processor ("CPU") or a dedicated graphics processor ("GPU"), and the memory 102 stores program instructions that can be executed on the processor. In some embodiments, the memory 102 may include, but is not limited to, a resistive random access memory RRAM (Resistive Random Access Memory), a dynamic random access memory DRAM (Dynamic Random Access Memory), a static random access memory SRAM (Static Random-Access Memory), and an enhanced dynamic random access memory EDRAM (Enhanced Dynamic Random Access Memory).
[0028] Furthermore, the memory 102 may store program instructions for generating a mental health report based on a multimodal model. In some embodiments, the multimodal model may include multiple feature extraction modules, a cross-modal fusion module, and a report generation module. When the program instructions are executed by the processor, the device 100 performs the following operations: collect multimodal data and pre-process the multimodal data, wherein the multimodal data at least includes fundus images, scale information and questionnaire information related to mental health; based on the fundus image and the questionnaire information, use multiple feature extraction modules to respectively extract fundus features and questionnaire features to obtain fundus feature vectors and questionnaire feature vectors, and represent the scale information as a scale feature vector; based on the fundus feature vector, the scale feature vector and the questionnaire feature vector, use a cross-modal fusion module to perform a cross-modal attention mechanism to obtain multimodal features; and input the multimodal features into the report generation module for report generation operation to generate a mental health report.
[0029] In some implementation scenarios, for example, a fundus acquisition device can be used to capture fundus images of multiple subjects, and scale information and questionnaire information related to the subjects' mental health (such as anxiety, depression, etc.) can be collected. Among them, the scale information contains relevant questions and corresponding scores; the questionnaire information contains question text and answer text. In some embodiments, the multimodal data can be preprocessed before being input into the multimodal model. In some embodiments, the preprocessing may include but is not limited to one or more of quality screening, image cropping, image resizing, or image color correction.
[0030] Specifically, images with poor quality (e.g., low resolution, high noise, etc.) can be screened out from the collected fundus images to achieve quality screening. For image cropping, a maximum square can be cut out from the center of the fundus image. As an example, the side length of the maximum square can be calculated based on side length = min (image width, image height), and then the starting point coordinates of the square (i.e., the coordinates of the upper left corner) are calculated, including and That is, the center coordinates of the fundus image are expanded outward so that the cropped square is the largest and completely located inside the original fundus image, so that important features in the fundus image are retained while possible interference elements in the periphery are removed.
[0031] For image resizing, it is possible to ensure that all fundus images have consistent shapes and sizes, which is convenient for subsequent processing and analysis. For example, the cropped fundus images can be adjusted to 224x224 pixels. In addition, image color correction can be performed on the fundus images to reduce the color deviation between images. For example, mapping the local average color to 40% grayscale helps to standardize the color of the fundus images and improve the learning effect and stability of the model.
[0032] Based on the fundus image, scale information and questionnaire information obtained as described above, when the program instructions are executed by the processor, the device 100 further executes based on the fundus image and questionnaire information, using multiple feature extraction modules to respectively extract fundus features and questionnaire features to obtain a fundus feature vector and a questionnaire feature vector, and represents the scale information as a scale feature vector.
[0033] In some embodiments, the plurality of feature extraction modules include a first feature module and a second feature module, wherein the device further performs the following operations to obtain a fundus feature vector and a questionnaire feature vector. That is, based on the fundus image, the first feature module is used to extract fundus features to obtain a fundus feature vector, and based on the questionnaire information, the second feature module is used to extract fundus features to obtain a questionnaire feature vector. In some embodiments, the aforementioned first feature module may include a self-supervised learning module, and the aforementioned second feature module may include a bidirectional recursive network module. Preferably, the aforementioned self-supervised learning module may be, for example, a RETFound model; the aforementioned bidirectional recursive network module may be, for example, a Span-level bidirectional retention model based on an attention mechanism (e.g., a BERT model).
[0034] In some implementation scenarios, for extracting fundus features using the first feature module to extract fundus feature vectors, the fundus image or the preprocessed fundus image x can be input into the first feature module (e.g., RETFound model) to extract high-level features from the fundus image. As an example, assuming that the feature vector extracted by the first feature module is F image , which is represented by the last layer of the first feature module as F image =f RETFound (x), that is, the extracted fundus feature vector is F image .
[0035] In some embodiments, for using the second feature module to extract fundus features to extract questionnaire feature vectors, first, based on the questionnaire information, the second feature module can be used to extract the context semantic vectors of each word in the answer text, the global vector of the question text, and the word pair relationship matrix in the answer text in the questionnaire information, and then the attention weights of each word in the answer text are calculated based on the context semantic vector and the global vector, and then the weighted sum is calculated based on the context semantic vector, the global vector, the word pair relationship matrix and the attention weight to obtain the questionnaire feature vector.
[0036] More specifically, we first perform word-level semantic analysis on each word in the question text and the answer text. i Perform word embedding representation and convert it into a high-dimensional vector h i Indicates the semantic features of a word: h i =Embed(ω i ), the entire answer text or question text can be converted into a corresponding word embedding matrix. In an implementation scenario, the answer text sequence X or the question text sequence Q can be input into the BERT model for word-level semantic parsing to convert each into a corresponding word embedding matrix
[0037] Next, the answer text is semantically parsed at the word pair level. For example, the word pair relationships are aggregated through a bidirectional recurrent network module to capture long-distance word relationships. In the implementation scenario, the bidirectional recurrent network module aggregates each word pair ω in the answer. i and ω j Perform analysis to generate a word pair relationship matrix between words Where l represents different relationship types. Each element M in the word pair relationship matrix ij Word ω i and ω j The semantic relationship between: in is the weight matrix to be learned, which represents the weight of the relationship between two words. In some implementation scenarios, after calculating the relationship matrix M between word pairs, it can also be semantically aggregated through a bidirectional recursive network module to further capture the relationship between distant words. That is, by semantically aggregating them through a multi-layer recursive network model, the semantic matrix M representing the output of the k-th layer recursive network can be obtained. (k) =Recurrent(M (k-1) ), thus, through this multi-layer semantic aggregation, the complex semantic relationships in the text can be better captured.
[0038] Furthermore, in order to better understand the semantics of each word in the entire sentence, the left-hand and right-hand context information of each word can be obtained through the bidirectional recurrent network module: By combining left- and right-facing contextual information and Get the contextual semantic vector of each word in the answer text: In addition, the word embedding matrix H of the question text can be q Processing is performed to generate a global vector of the question text through average pooling k represents the number of words.
[0039] Based on the contextual semantic vectors of each word in the answer text and the global vector of the question text obtained above, the attention weight of each word in the word pair can be calculated to measure the importance of each word relative to the question. For example, in an exemplary scenario, the word ω in the word pair i The corresponding attention weight α i The formula can be Calculated, where m represents the number of contextual semantic vectors of the word. Similarly, the word ω in the word pair can be obtained j The corresponding attention weight After obtaining the context semantic vector, global vector, word pair relationship matrix and attention weight, we can use the formula Calculate the questionnaire feature vector F text .
[0040] In some embodiments, the apparatus 100 further performs the following operations to represent the scale information as a scale feature vector. That is, the scores of each question in the scale information are calculated, and the scores of each question in the scale information are normalized to represent the scale information as a scale feature vector. Specifically, for each scale, the score of each question can be extracted as a part of the feature vector. As an example, assuming that the lth scale has N questions, and the score of each question is s i In order to maintain the comparability of scores between different scales, all scores are normalized, that is, s′ i =s i / s max , where s max is the maximum possible score for each question in the scale. Thus, the scale feature vector F of the scale can be obtained score =[s′ 1 ,s′ 2 ,…,s′ N Then, by concatenating the scale feature vectors of L scales, the scale feature vectors of all scale information can be obtained.
[0041] Further, when the program instructions are executed by the processor, the device 100 further executes a cross-modal attention mechanism based on the fundus feature vector, the scale feature vector, and the questionnaire feature vector using a cross-modal fusion module to obtain multimodal features. In some embodiments, the fundus feature vector, the scale feature vector, and the questionnaire feature vector are merged into a total feature vector for initialization, and then based on the initialized total feature vector, a cross-modal attention mechanism is executed using a cross-modal fusion module and recursively updated to obtain multimodal features.
[0042] In some implementation scenarios, the feature vector F is first image , scale feature vector F multi-score and the questionnaire feature vector F text Combined into a total feature vector by concatenation: X (0) =[F image ; F multi-score ; F text ] to initialize. Then, use the initialized total eigenvector X (0) Recursive updates are performed through the cross-modal attention mechanism to continuously optimize the feature representation: X (k+1) =σ(W k ·X (k) +b k ). Where σ represents the activation function ReLU, W k and b k Represent the weight and bias of the kth layer respectively. Weight W k The update relies on the back-propagation algorithm: Where η represents the learning rate, Represents the cross entropy loss function. In this scenario, when k reaches K, the recursion stops and the output is S = X (K) As an optimized multimodal feature representation.
[0043] Based on the multimodal features obtained above, when the program instructions are executed by the processor, the device 100 further performs the operation of inputting the multimodal features into the report generation module to perform a report generation operation to generate a mental health report. In some embodiments, the report generation module may include a first classification module and a natural language processing module, and the device further performs the following operations to generate a mental health report, namely, based on the multimodal features, using the first classification module to perform a classification operation to obtain a target classification result; inputting the target classification result into the natural language processing module for parsing to obtain report text information and combining the report text information with the medical knowledge base to generate a mental health report. In some implementation scenarios, the aforementioned first classification module may be, for example, a SHAP model, and the aforementioned natural language processing module may be, for example, a GPT model.
[0044] Specifically, when generating a mental health report, the input multimodal feature S can first be decomposed through, for example, the SHAP model, and the impact of each feature on the final classification result can be calculated. In some implementation scenarios, the contribution of each feature is quantified as a SHAP value (or Shapley value): Where xj represents the jth feature of the input, fF represents the model output containing the feature set F, V represents the set of all input features, and F represents a feature subset, which is F image , F multi-score or F text . Then, for example, GPT is used to determine the questionnaire answers that have the greatest impact on the prediction based on the size of the SHAP value, and these answers are highlighted in the interpretation of the questionnaire information. Similarly, the fundus image features and scale information scores can be interpreted and highlighted. Furthermore, GPT is combined with the standardized medical knowledge base DSM-V and / or NICE guidelines to automatically generate a mental health report.
[0045] In other embodiments, the device further performs the following operations: inputting the multimodal features into the second classification module for preliminary classification to obtain a preliminary classification result; comparing the preliminary classification result with the historical evaluation result to generate a feedback signal; optimizing the second classification module according to the feedback signal to obtain a final classification result based on the optimized second classification module and combining the report text information, the medical knowledge base and the final classification result to generate a final mental health report. In some embodiments, the aforementioned second classification module can be, for example, a multilayer perceptron model.
[0046] In some implementation scenarios, the input multimodal feature vector can be preliminarily classified by, for example, a multilayer perceptron model, and the probability distribution of different mental health risk categories can be output: P(y i |X) = softmax(W·X+b), where P(y i |X) represents the probability of each mental health risk category (e.g., normal, mild anxiety, moderate anxiety, severe anxiety, mild depression, moderate depression, and severe depression), and W and b represent the weight matrix and bias term of the multi-layer perceptron model, respectively. Next, the preliminary classification results are compared with the patient's historical assessment results (e.g., mental health assessment results ("history"), standard scales ("clinical_scores"), etc.) to generate a feedback signal to assist the multi-layer perceptron model in judging the rationality of the classification results. In some implementation scenarios, the aforementioned feedback signal can be expressed as F(y i )=α 1 ·Consistency(y i ,history)+α 2Matching(y i ,clinical_scores), where α 1 and α 2 They represent the weight parameters of historical consistency (“Consistency”) and scale score matching (“Matching”), which measure the consistency between the classification results and the historical data and scale scores, respectively.
[0047] Furthermore, through the feedback results, the weights of the multi-layer perceptron model are adjusted through back propagation until the classification results of the multi-layer perceptron model converge to stability. The final classification results obtained by the optimized multi-layer perceptron model can not only contain the classification information of each category, but also provide a specific score for each risk level, Risk_score(y i )=P(y i |X). For example, the Risk_score for moderate depression is 0.7, and the corresponding output moderate depression risk is 70%. In this scenario, the final classification result is combined with the above-mentioned report text information and medical knowledge base, and the GPT model is used for interpretation and analysis to generate the final mental health report. For example, for mild anxiety, the GPT model will recommend non-drug intervention, while for severe depression, it may recommend a combination of drug therapy and psychotherapy.
[0048] Combined with the above description, it can be seen that the embodiment of the present application extracts fundus features and questionnaire features through multiple feature extraction modules in the multimodal model to obtain fundus feature vectors and questionnaire feature vectors. Then, by combining the fundus feature vector, the questionnaire feature vector and the scale feature vector, the cross-modal fusion module is used to perform a cross-modal attention mechanism to capture the complex relationship between multimodal data and obtain multimodal features containing rich information. Furthermore, by inputting the multimodal features into the report generation module for report generation operation, a mental health report is generated. This can not only obtain the risk classification of mental health problems such as anxiety and depression, and improve the accuracy of mental health identification; it can also generate specific scoring results and give explanations to generate a completeness report to help patients and doctors better understand the identification results. It has important clinical application potential and is helpful for early screening and intervention of mental health problems.
[0049] Figure 2 2 is an exemplary flowchart showing the operation 200 performed by the apparatus for generating a mental health report based on a multimodal model according to an embodiment of the present application. Figure 2As shown in FIG. 1 , at step S201, multimodal data including fundus images, scale information related to mental health, and questionnaire information are collected, and at step S202, the multimodal data is preprocessed. In some embodiments, one or more preprocessing operations such as quality screening, image cropping, image size adjustment, or image color correction may be performed on the multimodal data. For more details about the preprocessing, please refer to the above Figure 1 The description made will not be repeated in this application.
[0050] Next, at step S203, the fundus feature vector and the questionnaire feature vector are extracted respectively through the first and second feature extraction modules. In some embodiments, the aforementioned first feature extraction module may be, for example, a RETFound model, and the second feature extraction module may be, for example, a BERT model. In extracting the questionnaire vector features, the second feature module may be used to extract the contextual semantic vectors of each word in the answer text, the global vector of the question text, and the word pair relationship matrix in the answer text, calculate the attention weights of each word in the answer text, and calculate the weighted sum to obtain the questionnaire feature vector. At step S204, the scale information is represented as a scale feature vector. For example, the scale feature vector is obtained by normalizing the scores of each question in the scale information. For more details about the aforementioned feature vector extraction, please refer to the above. Figure 1 The description made in will not be repeated in this application.
[0051] Further, at step S205, the fundus feature vector, the scale feature vector and the questionnaire feature vector are input into the cross-modal fusion module to execute the cross-modal attention mechanism to obtain multimodal features, and at step S206, the multimodal features are input into the report generation module to perform the report generation operation to obtain the report text information. In some embodiments, the report generation module may include a first classification module and a natural language processing module, and the aforementioned first classification module may be, for example, a SHAP model, and the aforementioned natural language processing module may be, for example, a GPT model. Specifically, the classification operation can be performed through the first classification module to obtain the target classification result, and then the target classification result can be input into the natural language processing module for parsing to obtain the report text information.
[0052] In addition, at step S207, the multimodal feature vector can also be input into the second classification module to obtain the scores of each mental health risk category and each risk level. In some embodiments, the second classification module can be, for example, a multi-layer perceptron model. The aforementioned mental health risk categories include, for example, normal, mild anxiety, moderate anxiety, severe anxiety, mild depression, moderate depression and severe depression. Further, by combining the scores of each mental health risk category and each risk level output at step S207 with the report text information at step S206 and combining the report text information with the medical knowledge base, a final mental health report is generated at step S208. For more details on generating the final mental health report, please refer to the above Figure 1 The description made in will not be repeated in this application.
[0053] Combined with the above description, it can be seen that the embodiment of the present application combines fundus images, scale information and questionnaire data, and uses a recursive cross-modal attention mechanism to optimize the feature expression of data in different modalities, thereby effectively improving the understanding and prediction accuracy of mental health status. Furthermore, the embodiment of the present application provides interpretability for the prediction results by using, for example, the SHAP model, so that the impact of each feature on the result can be quantified, thereby helping doctors and patients understand the decisions of the model. Furthermore, the embodiment of the present application also generates personalized explanatory reports and treatment recommendations by utilizing, for example, the GPT model, and combines it with a standardized medical knowledge base to provide personalized intervention plans suitable for patients.
[0054] The embodiment of the present application is not only applicable to personal mental health assessment, but also to large-scale screening and early intervention through multimodal data fusion and automated report generation methods. Its efficient and accurate automated processing flow can help medical institutions quickly identify potential high-risk groups, so as to intervene in time and improve the quality of life of patients. In addition, the embodiment of the present application generates comprehensive explanatory reports and personalized treatment recommendations by combining the patient's specific mental health risk situation, supports individualized psychological intervention and treatment, and can dynamically generate detailed treatment paths and lifestyle recommendations to help patients and doctors better understand the condition and take appropriate actions.
[0055] Figure 3 FIG. 1 is an exemplary schematic diagram showing a mental health report generated based on a multimodal model according to an embodiment of the present application. Figure 3 As shown in , the generated mental health report can include specific mental health assessment results, detailed analysis of multimodal data, and personalized treatment recommendations.
[0056] Figure 44 is an exemplary structural block diagram showing a device 400 for generating a mental health report based on a multimodal model according to an embodiment of the present application. It is understood that the device 400 may include an apparatus of an embodiment of the present application, and the device implementing the solution of the present application may be a single device (such as a computing device) or a multifunctional device including various peripheral devices.
[0057] like Figure 4 As shown in , the device of the present application may also include a central processing unit or central processing unit ("CPU") 411, which may be a general-purpose CPU, a dedicated CPU, or other information processing and program execution unit. Further, the device 400 may also include a large-capacity memory 412 and a read-only memory ("ROM") 413, wherein the large-capacity memory 412 may be configured to store various types of data, including various fundus images, scale information and questionnaire information related to mental health; based on fundus images and questionnaire information, algorithm data, intermediate results and various programs required to run the device 400. ROM 413 may be configured to store data and instructions required for power-on self-test of the device 400, initialization of various functional modules in the system, drivers for basic input / output of the system, and booting the operating system.
[0058] Optionally, the device 400 may also include other hardware platforms or components, such as the tensor processing unit ("TPU") 414, graphics processing unit ("GPU") 415, field programmable gate array ("FPGA") 416, and machine learning unit ("MLU") 417 shown. It is understood that although a variety of hardware platforms or components are shown in the device 400, they are merely exemplary and not restrictive, and those skilled in the art may add or remove corresponding hardware according to actual needs. For example, the device 400 may only include a CPU, related storage devices, and interface devices to implement the operations performed by the apparatus for generating a mental health report based on a multimodal model of the present application.
[0059] In some embodiments, in order to facilitate the transmission and interaction of data with an external network, the device 400 of the present application further includes a communication interface 418, so that it can be connected to a local area network / wireless local area network ("LAN / WLAN") 405 through the communication interface 418, and then connected to a local server 406 or to the Internet ("Internet") 407 through the LAN / WLAN. Alternatively or additionally, the device 400 of the present application can also be directly connected to the Internet or a cellular network through the communication interface 418 based on wireless communication technology, such as wireless communication technology based on the third generation ("3G"), the fourth generation ("4G") or the fifth generation ("5G") of the generation. In some application scenarios, the device 400 of the present application can also access the server 408 and the database 409 of the external network as needed to obtain various known algorithms, data and modules, and can remotely store various data, such as various types of data or instructions for presenting fundus images, scale information and questionnaire information related to mental health, etc.
[0060] The peripheral devices of the device 400 may include a display device 402, an input device 403, and a data transmission interface 404. In one embodiment, the display device 402 may include, for example, one or more speakers and / or one or more visual displays, which are configured to provide voice prompts and / or image video display for the generation of a mental health report based on a multimodal model of the present application. The input device 403 may include other input buttons or controls such as a keyboard, a mouse, a microphone, a gesture capture camera, etc., which are configured to receive input of audio data and / or user instructions. The data transmission interface 404 may include, for example, a serial interface, a parallel interface or a universal serial bus interface ("USB"), a small computer system interface ("SCSI"), a serial ATA, a FireWire ("FireWire"), a PCI Express, and a high-definition multimedia interface ("HDMI"), etc., which are configured for data transmission and interaction with other devices or systems. According to the solution of the present application, the data transmission interface 404 can receive the fundus image collected by the fundus acquisition device and the scale information and questionnaire information in the medical database, and transmit to the device 400 the fundus image, the scale information and questionnaire information related to mental health, or various other types of data or results.
[0061] The CPU 411, mass storage 412, ROM 413, TPU 414, GPU 415, FPGA 416, MLU 417 and communication interface 418 of the device 400 of the present application can be interconnected through a bus 419, and data interaction with peripheral devices can be achieved through the bus. In one embodiment, through the bus 419, the CPU 411 can control other hardware components in the device 400 and its peripheral devices.
[0062] Combination of the above Figure 4 The present invention describes a device that can be used to generate a mental health report based on a multimodal model for executing the present invention. It should be understood that the device structure or architecture here is only exemplary, and the implementation method and implementation entity of the present invention are not limited thereto, but can be changed without departing from the spirit of the present invention.
[0063] According to the above description in combination with the accompanying drawings, those skilled in the art can also understand that the embodiments of the present application can also be implemented by software programs. Therefore, the present application also provides a computer-readable storage medium, on which computer-readable instructions for generating a mental health report based on a multimodal model are stored. When the computer-readable instructions are executed by one or more processors, they can be used to implement the present application in combination with the accompanying drawings. Figure 1 The operations performed by the described apparatus for generating a mental health report based on a multimodal model.
[0064] It should be noted that although the operations of the method of the present application are described in a specific order in the accompanying drawings, this does not require or imply that the operations must be performed in this specific order, or that all the operations shown must be performed to achieve the desired results. On the contrary, the steps depicted in the flow chart can be performed in a different order. Additionally or alternatively, some steps can be omitted, multiple steps can be combined into one step, and / or one step can be decomposed into multiple steps.
[0065] It should be understood that when the terms "first", "second", "third" and "fourth" are used in the claims, the specification and the drawings of the present application, they are only used to distinguish different objects, rather than to describe a specific order. The terms "include" and "comprise" used in the specification and claims of the present application indicate the presence of the described features, wholes, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or their collections.
[0066] It should also be understood that the terms used in this application specification are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in this application specification and claims, unless the context clearly indicates otherwise, the singular forms of "a", "an" and "the" are intended to include plural forms. It should also be further understood that the term "and / or" used in this application specification and claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes these combinations.
[0067] Although the implementation methods of the present application are as above, the contents described are only examples adopted to facilitate the understanding of the present application, and are not intended to limit the scope and application scenarios of the present application. Any technician in the technical field described in the present application can make any modifications and changes in the form and details of the implementation without departing from the spirit and scope disclosed in the present application, but the scope of patent protection of the present application shall still be subject to the scope defined in the attached claims.
[0068] In addition, the collection and acquisition of various data in this application complies with relevant laws and regulations and is authorized by the data provider. Any organization or individual that needs to obtain external data must obtain authorization in accordance with the law and ensure data security. It is not allowed to illegally collect, use, process, or transmit unauthorized or unprotected data, or to illegally buy, sell, provide, or disclose unauthorized or unprotected data.
Claims
1. A device for generating a mental health report based on a multimodal model, wherein the multimodal model includes a plurality of feature extraction modules, a cross-modal fusion module and a report generation module, and the device includes: processor; as well as A memory having computer instructions for generating a mental health report based on a multimodal model stored thereon, wherein when the computer instructions are executed by a processor, the device performs the following operations: Collecting multimodal data and preprocessing the multimodal data, wherein the multimodal data at least includes fundus images, scale information and questionnaire information related to mental health; Based on the fundus image and the questionnaire information, using the multiple feature extraction modules to respectively extract fundus features and questionnaire features to obtain a fundus feature vector and a questionnaire feature vector, and representing the scale information as a scale feature vector; Based on the fundus feature vector, the scale feature vector and the questionnaire feature vector, using the cross-modal fusion module to perform a cross-modal attention mechanism to obtain multi-modal features; as well as The multimodal features are input into the report generation module to perform a report generation operation to generate a mental health report.
2. The apparatus according to claim 1, wherein the preprocessing comprises at least one or more of quality screening, image cropping, image size adjustment or image color correction.
3. The device according to claim 1, wherein the plurality of feature extraction modules include a first feature module and a second feature module, wherein the device further performs the following operations to obtain a fundus feature vector and a questionnaire feature vector: Based on the fundus image, extract fundus features using the first feature module to obtain the fundus feature vector; and Based on the questionnaire information, the second feature module is used to extract fundus features to obtain the questionnaire feature vector.
4. The apparatus according to claim 3, wherein the apparatus further performs the following operations to obtain the questionnaire feature vector: Based on the questionnaire information, using the second feature module to extract context semantic vectors of each word in the answer text, the global vector of the question text, and the word pair relationship matrix in the answer text in the questionnaire information; Calculate the attention weight of each word in the answer text according to the context semantic vector and the global vector; as well as A weighted sum is calculated based on the context semantic vector, the global vector, the word pair relationship matrix and the attention weight to obtain the questionnaire feature vector.
5. The apparatus of claim 3, wherein the first feature module comprises a self-supervised learning module, and the second feature module comprises a bidirectional recurrent network module.
6. The apparatus according to claim 1, wherein the apparatus further performs the following operations to represent the gauge information as a gauge feature vector: Calculating scores for each question in the scale information; and The scores of the various questions in the scale information are normalized to represent the scale information as the scale feature vector.
7. The apparatus according to claim 1, wherein the apparatus further performs the following operations to obtain multimodal features: Initialize by combining the fundus feature vector, the scale feature vector and the questionnaire feature vector into a total feature vector; and Based on the initialized total feature vector, the cross-modal fusion module is used to perform a cross-modal attention mechanism and perform recursive updates to obtain the multimodal features.
8. The apparatus according to claim 1, wherein the report generation module comprises a first classification module and a natural language processing module, and the apparatus further performs the following operations to generate a mental health report: Based on the multimodal features, using the first classification module to perform a classification operation to obtain a target classification result; Inputting the target classification result into the natural language processing module for parsing to obtain report text information; as well as The report text information is combined with the medical knowledge base to generate the mental health report.
9. The apparatus according to claim 8, wherein the apparatus further performs the following operations: Inputting the multimodal features into a second classification module for preliminary classification to obtain a preliminary classification result; comparing the preliminary classification result with the historical evaluation result to generate a feedback signal; Optimizing the second classification module according to the feedback signal to obtain a final classification result based on the optimized second classification module; as well as The report text information, the medical knowledge base and the final classification result are combined to generate a final mental health report.
10. A computer-readable storage medium, comprising computer program instructions for generating a mental health report based on a multimodal model, wherein when the computer program instructions are executed by one or more processors, the operations performed by the device according to any one of claims 1-9 are implemented.
Citation Information
Cited By
Identification device for isolated rapid eye movement sleep disorder typing transformation and storage medium
CN121191740A