Large model training method and system fusing psychological knowledge and multi-modal data
By constructing a multimodal psychological corpus and training a large model with a cross-modal encoder, the problem of insufficient multimodal data integration in existing technologies has been solved, thereby improving the professionalism of psychological assessment and intervention, and increasing the accuracy of empathy and service efficiency.
Patent Information
- Application Number
- CN202510953220.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-10
- Publication Date
- 2025-10-28
AI Technical Summary
Existing large-scale model training methods lack the ability to integrate multimodal data in psychological assessment and crisis intervention scenarios, and fail to incorporate clinical psychology theories, resulting in insufficient professionalism.
A multimodal psychological corpus was constructed, and a large model was pre-trained and trained on multimodal data using psychological data. Features were extracted from speech, text, and scales, and features were concatenated using a cross-modal encoder. The model parameters were then optimized based on evaluation metrics.
It significantly improved the professionalism of psychological assessment and intervention, increased the accuracy of empathy, reduced service costs, and enhanced the efficiency and quality of psychological counseling and treatment.
Smart Images

Figure CN120849893A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of model training, and in particular to a method and system for training large models that integrate psychological knowledge and multimodal data. Background Technology
[0002] Existing large-scale model training methods primarily focus on multimodal data processing in general domains (such as common languages and image recognition), without being specifically designed for psychological assessment scenarios. The training process typically relies on a single modality (such as text or image) and lacks the ability to integrate multimodal data such as speech and text. Furthermore, existing training methods do not incorporate clinical psychology theories, resulting in insufficient professionalism of the models in psychological assessment and crisis intervention scenarios.
[0003] The aforementioned shortcomings of existing technologies (such as limited data modalities and insufficient professional knowledge) directly result in their lack of professionalism in psychological assessment scenarios, and there is considerable room for improvement in the professional capabilities of large models. Summary of the Invention
[0004] The purpose of this invention is to disclose a method and system for training large models that integrates psychological knowledge and multimodal data, thereby solving the technical problems mentioned in the background art.
[0005] To achieve the above objectives, the present invention adopts the following technical solution:
[0006] This invention provides a method for training large models that integrates psychological knowledge and multimodal data, including:
[0007] S1, Construct a multimodal psycho-corpus;
[0008] S2, based on psychological data, pre-trains the original large model to obtain the first large model;
[0009] S3, based on a multimodal psycho corpus, trains the first major model to obtain the second major model;
[0010] S4 evaluates the second major model based on preset evaluation indicators, obtains evaluation results, and optimizes the parameters of the second major model based on the evaluation results.
[0011] Preferably, a multimodal psychocorpus is constructed, including:
[0012] S10, Obtain recordings of psychological counseling, CBT dialogue records, and scales;
[0013] S11, Extract speech features from the psychological counseling recording to obtain the first feature;
[0014] S12, extract text features from the CBT dialogue record to obtain the second feature;
[0015] S13, extract text features from the scale to obtain the third feature;
[0016] S14. Store the first feature, the second feature, and the third feature into the multimodal psychocorpus.
[0017] Preferably, psychological materials include psychology books.
[0018] Preferably, the original large model is pre-trained based on psychological data to obtain the first large model, including:
[0019] S20, perform text preprocessing on the psychological data to obtain the preprocessed text;
[0020] S21, using natural language processing techniques to obtain key information from preprocessed text;
[0021] S22, convert the key information into the first vector;
[0022] S23, use the first vector to pre-train the original large model to obtain the first large model.
[0023] Preferably, the psychological data undergoes text preprocessing to obtain preprocessed text, including:
[0024] The psychology data was processed sequentially through text cleaning, word segmentation, part-of-speech tagging, and part-of-speech restoration to obtain preprocessed text.
[0025] Preferably, the key information includes core conceptual entities, theoretical relational triples, clinical diagnostic criteria, intervention methodology, and scale assessment logic.
[0026] Preferably, the first major model is trained based on a multimodal psychological corpus to obtain the second major model, including:
[0027] The first, second, and third features in the multimodal corpus are vectorized to obtain the second vector;
[0028] The second vector is used to train the first large model to obtain the second large model.
[0029] Preferably, the first feature, second feature, and third feature in the multimodal corpus are vectorized to obtain a second vector, including:
[0030] Obtain the vectors corresponding to the first feature, the second feature, and the third feature respectively;
[0031] The vectors corresponding to the first feature, the second feature, and the third feature are concatenated to obtain the second vector.
[0032] Preferably, the second major model is evaluated based on preset evaluation indicators to obtain evaluation results, and the parameters of the second major model are optimized based on the evaluation results, including:
[0033] S41, Obtain the evaluation rules;
[0034] S42, Get the test set;
[0035] S43, input the test set into the second large model for calculation, and obtain the calculation results;
[0036] S44, the evaluation result is obtained based on the calculation results and evaluation rules;
[0037] S45, optimize the parameters of the second largest model based on the evaluation results.
[0038] This invention also provides a large model training system that integrates psychological knowledge and multimodal data, including a construction module, a first training module, a second training module, and an optimization module;
[0039] The building blocks are used to construct a multimodal psychocorpus;
[0040] The first training module is used to pre-train the original large model based on psychological data to obtain the first large model;
[0041] The second training module is used to train the first major model based on a multimodal psycho corpus to obtain the second major model;
[0042] The optimization module is used to evaluate the second major model based on preset evaluation indicators, obtain evaluation results, and optimize the parameters of the second major model based on the evaluation results.
[0043] Beneficial effects:
[0044] Compared to traditional multimodal large-scale model training methods, this invention significantly improves the professionalism of psychological assessment and intervention by integrating psychological theories with multimodal data to train a large-scale model. In terms of technical effectiveness, it effectively improves the model's empathy accuracy, enabling the provision of professional and accurate psychological assessment and intervention services to more people, thus contributing to improving the overall level of social mental health. In terms of economic benefits, it can be applied to fields such as psychological counseling and psychotherapy, improving service efficiency and quality while reducing service costs, demonstrating significant market potential and economic benefits. Attached Figure Description
[0045] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. It should be understood that the following drawings only illustrate certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without paying any creative work.
[0046] Figure 1 This is a schematic diagram of the large model training method that integrates psychological knowledge and multimodal data according to the present invention.
[0047] Figure 2 This is a schematic diagram of the large model training system that integrates psychological knowledge and multimodal data according to the present invention. Detailed Implementation
[0048] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.
[0049] like Figure 1 As shown in one embodiment, the present invention provides a method for training large models that integrates psychological knowledge and multimodal data, including...
[0050] S1, Construct a multimodal psychocorpus.
[0051] Preferably, a multimodal psychocorpus is constructed, including:
[0052] S10, obtain recordings of psychological counseling sessions, CBT conversation records, and scales.
[0053] By integrating psychological counseling recordings, CBT dialogue records, and scales, a multimodal psychological corpus covering voice and text is constructed, providing rich professional data support for model training.
[0054] All of the above data has been anonymized to prevent privacy leaks.
[0055] S11, extract speech features from the psychological counseling recording to obtain the first feature.
[0056] It can perform noise reduction, segmentation, and feature extraction (such as speech intensity, tone changes, and fluctuations) on speech data to obtain the first feature.
[0057] Furthermore, noise reduction is performed on the voice data, including:
[0058] Perform burst noise detection on the voice data and perform repair processing based on the detection results;
[0059] Noise reduction is performed on the audio data obtained after restoration.
[0060] Furthermore, burst noise detection is performed on the voice data, including:
[0061] The first step is to perform frame segmentation on the audio data to obtain multiple audio frames;
[0062] The second step is to obtain the energy spectral density for each frame.
[0063] The third step is to calculate the energy density comparison value for each frame.
[0064] The fourth step is to calculate the energy density comparison value between two adjacent speech frames to determine the judgment value.
[0065] The fifth step is to determine whether there is sudden noise based on the judgment value.
[0066] The noise detection method of this invention primarily detects abrupt changes in energy over time. Although background noise affects the overall energy level, this method can still effectively detect sudden noise events as long as the energy change of sudden noise is more significant (signal-to-noise ratio is sufficiently high). It is also robust to stable or slowly varying background noise.
[0067] Noise detection and repair can improve the quality of speech data, thereby improving the quality of the acquired features.
[0068] The energy spectral density can be obtained by first taking the Fourier transform result of the speech signal, and then taking the square of the modulus of the Fourier transform result.
[0069] Furthermore, calculate the energy density comparison value for each frame, including:
[0070] The energy density control value is calculated using the following formula:
[0071]
[0072] H t Let s(f) and s(t) be the energy density reference values for speech frame t, and let s(f) and s(t) represent the energy spectral densities of speech frames f and t, respectively.
[0073] This invention does not directly compare the energy spectral density of adjacent frames, but calculates the energy density comparison value. This upgrades the core detection target from "whether the energy changes abruptly" to "whether the rate of energy change changes abruptly" (second-order change), thereby significantly improving the sensitivity to real sudden noise events while suppressing interference from inherent changes in speech.
[0074] Furthermore, the formula for calculating the judgment value is as follows:
[0075]
[0076] ΔH is the judgment value of speech frame t, ΔT is the length of the speech frame, and H t-1 This is the energy density reference value for speech frame t-1.
[0077] Furthermore, based on the judgment value, the presence of sudden noise is determined, including:
[0078] If the judgment value is greater than the set judgment value threshold, it indicates that there is sudden noise.
[0079] For example, the threshold for the judgment value can be set to 3.2 bits / ms.
[0080] Furthermore, based on the detection results, remedial measures are implemented, including:
[0081] If the detection result indicates the presence of sudden noise, then the speech data is repaired using interpolation repair methods.
[0082] S12, extract text features from the CBT dialogue record to obtain the second feature.
[0083] The second feature can be obtained by performing calculations such as text data cleaning, word segmentation, part-of-speech tagging, and semantic parsing.
[0084] S13, extract text features from the scale to obtain the third feature.
[0085] Similarly, the scale can also be processed in the same way to obtain the third characteristic.
[0086] S14. Store the first feature, the second feature, and the third feature into the multimodal psychocorpus.
[0087] S2, based on psychological data, pre-trains the original large model to obtain the first large model.
[0088] Preferably, psychological materials include psychology books.
[0089] Psychology books contain basic knowledge in the field of psychology, enabling models to learn key information such as psychological theories, professional terminology, and clinical diagnostic criteria.
[0090] Preferably, the original large model is pre-trained based on psychological data to obtain the first large model, including:
[0091] S20, perform text preprocessing on the psychological data to obtain preprocessed text, including:
[0092] The psychology data was processed sequentially through text cleaning, word segmentation, part-of-speech tagging, and part-of-speech restoration to obtain preprocessed text.
[0093] When cleaning text, a domain-adaptive cleaning strategy can be used:
[0094] In the text cleaning stage, a domain-adaptive cleaning strategy is introduced to optimize for the specific characteristics of texts in the psychology field. For example, context-sensitive noise reduction is performed on technical terms (such as "cognitive behavioral therapy" and "psychological assessment scale") to avoid semantic loss due to terminology misjudgment. Specifically, domain dictionaries (such as a clinical psychology terminology database) can be used for terminology identification and filtering to ensure the accuracy of terminology.
[0095] Using multi-granularity word segmentation tools (such as jieba+HanLP) combined with named entity recognition (NER) technology, psychological concepts can be accurately extracted. For example, HanLP's named entity recognition module can be used to identify entity types such as "disease name," "treatment method," and "assessment scale" in the text.
[0096] By cleaning and segmenting words, we ensure the accuracy of technical terminology recognition, providing high-quality input for subsequent semantic analysis.
[0097] When tagging parts of speech, domain dictionaries (such as a clinical psychology terminology database) can be introduced to improve the accuracy of identifying specialized terms. For example, terms such as "anxiety," "depression," and "empathy" can be tagged to ensure their correctness in semantic parsing. A hybrid tagging method based on rules and statistics, combining dependency parsing and semantic role tagging techniques, is employed to analyze the semantic structure of the text and extract key concepts and relationships.
[0098] By semantic parsing, a richer semantic graph can be constructed, providing more refined semantic features for subsequent vector representation.
[0099] S21, using natural language processing techniques to obtain key information from preprocessed text.
[0100] Preferably, the key information includes core conceptual entities, theoretical relational triples, clinical diagnostic criteria, intervention methodology, and scale assessment logic. Examples of key information and corresponding acquisition techniques are shown in Table 1 below:
[0101] Table 1. Examples of Key Information and Comparison of Extraction Techniques
[0102]
[0103] S22, convert the key information into the first vector.
[0104] The extracted key information is transformed into vector representations using word embedding techniques (such as Word2Vec and BERT) to facilitate model learning.
[0105] S23, use the first vector to pre-train the original large model to obtain the first large model.
[0106] Self-supervised learning (such as masked language modeling, MLM) is used to enable the original large model to predict masked psychological terms.
[0107] Enable the model to acquire structured representations of knowledge in the field of psychology (e.g., understanding the association between "CBT" and "exposure therapy").
[0108] Training examples are shown in Table 2:
[0109] Table 2 Training Case Table
[0110]
[0111] Pre-training can significantly reduce the cost of data labeling.
[0112] S3, based on a multimodal psycho corpus, trains the first major model to obtain the second major model.
[0113] Preferably, the first major model is trained based on a multimodal psychological corpus to obtain the second major model, including:
[0114] Vectorize the first, second, and third features in the multimodal corpus to obtain the second vector, which includes:
[0115] Obtain the vectors corresponding to the first feature, the second feature, and the third feature respectively;
[0116] The vectors corresponding to the first feature, the second feature, and the third feature are concatenated to obtain the second vector.
[0117] Convert the first feature (such as pitch, pause frequency) into an acoustic vector;
[0118] Convert CBT dialogue text (such as patient descriptions) into semantic vectors;
[0119] Convert scale data (such as BDI scores) into structured vectors.
[0120] A second vector is obtained by concatenating the vectors corresponding to the first, second, and third features using a cross-modal encoder.
[0121] Feature fusion can improve a model’s ability to understand multimodal data.
[0122] Different modalities of data typically have different feature spaces. For example, text consists of discrete sequences of words, while images are represented by continuous matrices of pixel values. Cross-modal encoders can map these different modalities of data into a shared feature space. In this unified feature space, data from different modalities can correspond to and be compared with each other.
[0123] A unified feature representation provides the foundation for subsequent feature fusion operations. When data from different modalities are in the same feature space, their features can be easily concatenated, weighted, or summed for fusion operations.
[0124] The second vector is used to train the first large model to obtain the second large model.
[0125] Specifically, supervised learning can be used as a training method. Through cross-modal attention mechanisms, the model can learn the intermodal associations (e.g., voice tremor + text "I feel fear" → anxiety disorder) to achieve training.
[0126] S4 evaluates the second major model based on preset evaluation indicators, obtains evaluation results, and optimizes the parameters of the second major model based on the evaluation results.
[0127] Preferably, the second major model is evaluated based on preset evaluation indicators to obtain evaluation results, and the parameters of the second major model are optimized based on the evaluation results, including:
[0128] S41, Obtain the evaluation rules.
[0129] The evaluation rules include evaluation dimensions and specific indicators. For example, if the evaluation dimension is knowledge accuracy, the specific indicators are the accuracy rate of psychological concept recognition, the compliance rate of diagnostic criteria, etc.
[0130] For example, if the evaluation dimension is multimodal understanding, the specific indicator is the speech-text consistency score (such as the matching of speech emotion with text content).
[0131] S42, Get the test set.
[0132] The test set includes voice, CBT dialogue text and scale data, as well as corresponding standard information (such as real diagnostic labels (e.g., depression level)).
[0133] S43, input the test set into the second large model for calculation, and obtain the calculation results.
[0134] The calculation result can be the corresponding label.
[0135] S44, the evaluation result is obtained based on the calculation results and evaluation rules.
[0136] The difference between the calculation results and the standard information can be used as the evaluation result. For example, the distance between the vector of keywords in the calculation results and the vector of the standard information can be used as the evaluation result.
[0137] S45, optimize the parameters of the second largest model based on the evaluation results.
[0138] Optimization based on evaluation results is a typical iterative fine-tuning process:
[0139] If the model scores low on the "crisis intervention recommendation" task, locate the relevant parameter layer (such as the decision output layer).
[0140] Optimization methods:
[0141] Gradient descent: Calculate the evaluation loss (such as cross-entropy) → backpropagate to update parameters.
[0142] Targeted sampling: Weighting erroneous samples (such as misdiagnosed cases) during training.
[0143] Furthermore, the evaluation rules can reference international standards (such as the APA ethical guidelines and the EU AI Act) and industry practices to develop an "AI Psychological Assessment Standard Table," clearly defining the definitions, calculation methods, and evaluation standards for each assessment indicator. This standard table covers four core dimensions: human-centered interaction and collaborative experience, emotional understanding and value orientation, personalized understanding and intervention support, and ethical compliance. Each dimension includes multiple measurement directions (such as empathy accuracy, effectiveness of value guidance, and crisis referral response speed).
[0144] Based on the evaluation results, analyze the model's strengths and weaknesses in various aspects to determine the optimization direction. Employ optimization algorithms (such as stochastic gradient descent and Adam) to adjust and optimize the model's parameters to improve its performance across various evaluation metrics. Repeat the evaluation and optimization process until the model's performance meets the expected requirements.
[0145] like Figure 2 The present invention also provides a large model training system that integrates psychological knowledge and multimodal data, including a construction module, a first training module, a second training module and an optimization module;
[0146] The building blocks are used to construct a multimodal psychocorpus;
[0147] The first training module is used to pre-train the original large model based on psychological data to obtain the first large model;
[0148] The second training module is used to train the first major model based on a multimodal psycho corpus to obtain the second major model;
[0149] The optimization module is used to evaluate the second major model based on preset evaluation indicators, obtain evaluation results, and optimize the parameters of the second major model based on the evaluation results.
[0150] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A method for training large models that integrates psychological knowledge and multimodal data, characterized in that: include: S1, Construct a multimodal psycho-corpus; S2, based on psychological data, pre-trains the original large model to obtain the first large model; S3, based on a multimodal psycho corpus, trains the first major model to obtain the second major model; S4 evaluates the second major model based on preset evaluation indicators, obtains evaluation results, and optimizes the parameters of the second major model based on the evaluation results.
2. The method for training large models that integrates psychological knowledge and multimodal data according to claim 1, characterized in that, Constructing a multimodal psychocorpus, including: S10, Obtain recordings of psychological counseling, CBT dialogue records, and scales; S11, Extract speech features from the psychological counseling recording to obtain the first feature; S12, extract text features from the CBT dialogue record to obtain the second feature; S13, extract text features from the scale to obtain the third feature; S14. Store the first feature, the second feature, and the third feature into the multimodal psychocorpus.
3. The method for training large models that integrates psychological knowledge and multimodal data according to claim 2, characterized in that, Psychology materials include psychology books.
4. The method for training large models that integrates psychological knowledge and multimodal data according to claim 3, characterized in that, The original large model was pre-trained based on psychological data to obtain the first large model, which includes: S20, perform text preprocessing on the psychological data to obtain the preprocessed text; S21, using natural language processing techniques to obtain key information from preprocessed text; S22, convert the key information into the first vector; S23, use the first vector to pre-train the original large model to obtain the first large model.
5. The method for training large models that integrates psychological knowledge and multimodal data according to claim 4, characterized in that, The psychological data was preprocessed to obtain the preprocessed text, including: The psychology data was processed sequentially through text cleaning, word segmentation, part-of-speech tagging, and part-of-speech restoration to obtain preprocessed text.
6. The method for training large models that integrates psychological knowledge and multimodal data according to claim 4, characterized in that, Key information includes core conceptual entities, theoretical relational triples, clinical diagnostic criteria, intervention methodology, and scale assessment logic.
7. The method for training large models integrating psychological knowledge and multimodal data according to claim 1, characterized in that, The first major model was trained using a multimodal psychological corpus to obtain the second major model, which includes: The first, second, and third features in the multimodal corpus are vectorized to obtain the second vector; The second vector is used to train the first large model to obtain the second large model.
8. The method for training large models integrating psychological knowledge and multimodal data according to claim 1, characterized in that, Vectorize the first, second, and third features in the multimodal corpus to obtain the second vector, which includes: Obtain the vectors corresponding to the first feature, the second feature, and the third feature respectively; The vectors corresponding to the first feature, the second feature, and the third feature are concatenated to obtain the second vector.
9. The method for training large models that integrates psychological knowledge and multimodal data according to claim 1, characterized in that, The second major model is evaluated based on preset evaluation metrics to obtain evaluation results. Based on these results, the parameters of the second major model are optimized, including: S41, Obtain the evaluation rules; S42, Get the test set; S43, input the test set into the second large model for calculation, and obtain the calculation results; S44, the evaluation result is obtained based on the calculation results and evaluation rules; S45, optimize the parameters of the second largest model based on the evaluation results.
10. A large-scale model training system integrating psychological knowledge and multimodal data, characterized in that: It includes a construction module, a first training module, a second training module, and an optimization module; The building blocks are used to construct a multimodal psychocorpus; The first training module is used to pre-train the original large model based on psychological data to obtain the first large model; The second training module is used to train the first major model based on a multimodal psycho corpus to obtain the second major model; The optimization module is used to evaluate the second major model based on preset evaluation indicators, obtain evaluation results, and optimize the parameters of the second major model based on the evaluation results.
Citation Information
Patent Citations
Prediction and evaluation method and device for depression level of master and doctor-degree group
CN113436737A
Nutrition consultation method based on artificial intelligence
CN116913471A
Multi-modal personality analysis method, analysis model, readable storage medium and device
CN118916707A
Artificial intelligence psychological assessment method and system based on multi-modal input
CN119108112A