Tumor patient group psychotherapy method and system based on artificial intelligence
Through multimodal emotion recognition and adaptive intervention technology, group psychotherapy for cancer patients is evaluated and adjusted in real time, solving the problems of narrow coverage and slow response of traditional intervention, and achieving personalized precision intervention and resource optimization.
Patent Information
- Application Number
- CN202510957446.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-11
- Publication Date
- 2025-10-17
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional psychotherapy for cancer patients relies on manual intervention, has narrow coverage, slow response, and lacks data-driven iteration. It cannot meet the needs of large-scale patients and lacks systematic intervention of group dynamics, individual emotions, and physiological indicators.
Multimodal emotion recognition technology is used to collect patient data in real time. Through clustering and adaptive intervention, the treatment script is dynamically adjusted to achieve real-time accurate assessment and immediate intervention adjustment, forming a closed-loop optimization.
It has improved the individualized precision and resource utilization efficiency of group psychotherapy, enabling precise intervention and continuous optimization for a large number of patients.
Smart Images

Figure CN120809090A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of artificial intelligence, in particular to a tumor patient group psychological treatment method and system based on artificial intelligence. BACKGROUND
[0002] In the clinical practice of oncology, the incidence of psychological comorbidities such as anxiety and depression is as high as more than 40%. Traditional intervention mainly relies on one-on-one psychological therapists to implement cognitive-behavioral therapy (CBT), mindfulness-based stress reduction therapy (MBSR) and other modes, which requires a large amount of human, space and time resources, and it is difficult to cover the growing patient base.
[0003] Multimodal emotion recognition, real-time voice / emotion analysis and remote interaction technology of artificial intelligence (AI) are gradually embedded in medical scenarios for automatic assessment of patient status and provision of digital intervention. At the same time, the popularity of wearable physiological sensors makes large-scale and continuous monitoring possible, giving birth to a new paradigm of "AI-assisted group therapy" without therapist involvement.
[0004] Publications focus on single emotion recognition or single therapy push, lack of systematic intervention targeting "group dynamics-individual emotion-physiological indicators" three-dimensional coupling, and lack of automatic grouping of patients, group atmosphere regulation and individualized strategy linkage, still relying on artificial experience decision-making, resulting in limited timeliness and precision of intervention. SUMMARY
[0005] In order to overcome the shortcomings of the prior art, the present application aims to provide a tumor patient group psychological treatment method and system based on artificial intelligence, which realizes real-time accurate assessment, immediate intervention adjustment and closed-loop continuous optimization in a group context, overcoming the defects of traditional artificial psychological intervention in narrow coverage, slow response and lack of data-driven iteration.
[0006] To achieve the above-mentioned purpose, the present application provides the following scheme:
[0007] A tumor patient group psychological treatment method based on artificial intelligence, comprising:
[0008] Obtaining demographic information, disease course information and psychological scale scores of patients to obtain a first data set, using a wearable physiological sensor to collect physiological signals of heart rate, blood oxygen saturation and galvanic skin response of patients in real time to obtain a second data set, and synchronously collecting voice, text and facial image data of patients during group conversation to obtain a third data set;
[0009] Implementing multimodal emotion recognition and psychological demand assessment on the first data set, the second data set and the third data set to obtain emotion feature indicators and psychological demand indicators;
[0010] According to the emotional feature index, the psychological demand index, and the course stage, a clustering operation is performed on the patients to divide the patients into a target treatment group;
[0011] The conversation process of the target treatment group is analyzed in real time to obtain three interaction indexes, including speech balance, topic dominance, and cohesion within the group. When any of the interaction indexes or the emotional feature index exceeds a corresponding preset threshold, a guidance instruction is generated to adjust the discussion pace.
[0012] According to the emotional feature index and the interaction index, at least one intervention script is selected from a cognitive-behavioral therapy, an acceptance and commitment therapy, or a mindfulness stress reduction therapy. During the execution, the length, intensity, and content of the script are dynamically adjusted according to the real-time updated emotional feature index until the emotional feature index returns to a preset safe interval, and intervention effect data is obtained.
[0013] The emotional feature index, the psychological demand index, the interaction index, and the intervention effect data are summarized to generate a report containing emotional trends, interaction quality, and intervention effectiveness.
[0014] Preferably, multi-modal emotion recognition and psychological demand assessment are performed on the first data set, the second data set, and the third data set to obtain emotional feature indexes and psychological demand indexes, including:
[0015] The voice data in the third data set is subjected to noise reduction, framing, and Mel spectrum extraction, the text data in the third data set is subjected to word segmentation, stop word removal, and embedding vector mapping, and the facial image data in the third data set is subjected to face detection, key point positioning, and expression action unit coding;
[0016] The physiological signals in the second data set are used to calculate physiological emotional features; the physiological emotional features include heart rate variability and skin conductance response amplitude;
[0017] The processed voice data is input into a convolution-long short-term memory hybrid model to output a voice emotion vector, and the processed text data is input into a bidirectional gated recurrent network to output a text emotion vector;
[0018] The expression action unit coding is input into a residual convolutional network to output a facial emotion vector, and the physiological emotional features are input into a random forest classifier to output a physiological emotion vector;
[0019] The voice emotion vector, the text emotion vector, the facial emotion vector, and the physiological emotion vector are synchronized according to timestamps, and the synchronized four types of emotion vectors are input into an attention fusion network to calculate a comprehensive emotional feature index;
[0020] The comprehensive emotion feature index is combined with the psychological scale score in the first data set, and the combined coding is input into a supervised gradient boosting decision tree model to output a psychological demand index.
[0021] Preferably, the voice emotion vector, the text emotion vector, the facial emotion vector and the physiological emotion vector are synchronized by time stamp, and the synchronized four types of emotion vectors are input into an attention fusion network to calculate a comprehensive emotion feature index, including:
[0022] Based on the respective time stamps of the voice emotion vector, the text emotion vector, the facial emotion vector and the physiological emotion vector, a continuous time window W is divided according to a preset uniform sampling period Δt k , a uniform time reference is established;
[0023] In each time window W k , the missing emotion vectors are completed by linear interpolation, so that the voice emotion vector, the text emotion vector, the facial emotion vector and the physiological emotion vector have consistent sequence length and time index within the time window W k ;
[0024] The four types of emotion vectors aligned in the time window W k are stacked in a fixed modal order to form a multi-modal feature matrix M k ;
[0025] The multi-modal feature matrix M k is input into a multi-head self-attention fusion network to calculate modal weight coefficients and output a weighted fusion vector F k ;
[0026] The weighted fusion vectors F k of all time windows are subjected to pooling operation to obtain a comprehensive emotion feature index.
[0027] Preferably, the comprehensive emotion feature index is combined with the psychological scale score in the first data set, and the combined coding is input into a supervised gradient boosting decision tree model to output a psychological demand index, including:
[0028] Zero-mean normalization processing is performed on the comprehensive emotion feature index, and interval normalization processing is performed on the psychological scale score, so that the two types of features are in comparable dimension;
[0029] The normalized comprehensive emotion feature index and the psychological scale score are spliced into a fixed-length joint feature vector in a preset modal order;
[0030] inputting the fixed-length joint feature vector into a recursive feature elimination module to eliminate dimensions with a correlation with the target label lower than a preset threshold, to obtain a reduced feature vector;
[0031] inputting the reduced feature vector into a gradient boosting decision tree model that has completed supervised training, to output a corresponding psychological demand indicator;
[0032] performing temperature scaling calibration on the psychological demand indicator to improve the consistency of the confidence of the indicator in different patient groups.
[0033] Preferably, according to the emotional feature indicator, the psychological demand indicator, and the disease course stage, performing clustering operation on patients to divide the patients into a target treatment group, including:
[0034] For each patient, the emotional feature indicator, the psychological demand indicator, and the disease course stage code are spliced in a preset order to form a fixed-length patient feature vector;
[0035] standardizing each dimension in the fixed-length patient feature vector according to a zero-mean normalization rule to obtain a standardized patient feature vector set;
[0036] in a candidate clustering number interval , based on the standardized patient feature vector set, respectively calculating an average silhouette coefficient , a disease course stage purity , and a Davies-Bouldin index , and normalizing 0-1 to a standard index , and selecting the optimal clustering number with the maximum comprehensive clustering quality indicator according to the formula ; wherein , ; wherein and are the minimum and maximum values of the candidate clustering number interval, respectively;
[0037] Taking the standardized patient feature vector set as a sample space, using K-means++ to initialize centroids, and iteratively assigning each vector in the standardized patient feature vector set and updating the centroids under the Euclidean distance metric until the centroid displacement of the two consecutive times is less than a preset convergence threshold;
[0038] For boundary samples in the standardized patient feature vector set that are close to both adjacent centroids and have a distance difference lower than a set threshold, calling a K-nearest neighbor algorithm to reevaluate the neighborhood distribution and adjust the cluster attribution;
[0039] generating a target treatment group based on the corrected clusters.
[0040] Within the candidate number of clusters interval , the average silhouette coefficient , disease stage purity and Davies-Bouldin index are calculated respectively based on the set of standardized patient feature vectors , and the 0-1 normalization is standardized as the index .
[0041] Preferably, within the candidate number of clusters interval , the average silhouette coefficient , disease stage purity and Davies-Bouldin index are calculated respectively based on the set of standardized patient feature vectors , and the 0-1 normalization is standardized as the index , including:
[0042] For any patient sample in the cluster , the average Euclidean distance with other samples in the same cluster and the average Euclidean distance with the nearest neighbor sample in other clusters are calculated, and the single-sample silhouette coefficient is obtained according to ;
[0043] The arithmetic mean of sample is taken to obtain the average silhouette coefficient under the number of clusters ;
[0044] Let the th cluster contain patients, and the number of patients with the main disease stage be , then the disease stage purity of the th cluster is , and the weighted sum of the purity of each cluster according to the sample proportion is obtained ;
[0045] The original Davies-Bouldin index is calculated, and and are recorded within the range of all candidate numbers of clusters , and is executed to complete the 0-1 normalization; wherein is the average Euclidean distance from the th cluster sample to the cluster center, is the Euclidean distance between the th and the th cluster centers.
[0046] Preferably, the conversation process of the target treatment group is analyzed in real time for three types of interaction indicators: speech balance, topic dominance, and group cohesion. When any of the interaction indicators or the emotional characteristic indicators exceeds the corresponding preset threshold, guidance instructions are generated to adjust the discussion rhythm, including:
[0047] performing automatic speech recognition and speaker separation on the audio stream of the target treatment group to obtain a text sequence with timestamps and speaker labels;
[0048] In the sliding time window, count the speaking time of each member , calculate the total duration , and press Get speech balance ; The number of members in the target treatment group;
[0049] Perform topic distribution inference on the transcribed text and count the number of words each member spoke on the current dominant topic , calculate the total number of words , and press Get the topic dominance rate ;
[0050] Use sentence vector representation to calculate the average cosine similarity of all speech pairs in the same class , and press Gaining group cohesion ;
[0051] Balance the speech , Topic Dominance Rate , intra-group cohesion and emotional characteristic indicators With their respective preset thresholds 、 、 、 Compare, if If any of the indicators exceeds the corresponding threshold, a guidance instruction is generated and pushed to the conversation interface to adjust the discussion rhythm or invite low-participation members to speak.
[0052] Preferably, at least one intervention script from cognitive-behavioral therapy, acceptance-commitment therapy, or mindfulness-based stress reduction therapy is selected based on the emotional characteristic index and the interaction index, and during execution, the duration, intensity, and content of the script are dynamically adjusted based on the real-time updated emotional characteristic index until the emotional characteristic index returns to a preset safe range, thereby obtaining intervention effect data, including:
[0053] Current comprehensive emotional characteristic index and interaction metrics collection performing rule matching to generate at least one list of candidate intervention scripts, wherein the cognitive-behavioral therapy script is directed to a high anxiety-high cognitive dissonance state, the acceptance-commitment therapy script is directed to a high avoidance-low value focus state, and the mindfulness stress reduction therapy script is directed to a high physiological arousal-low emotional regulation state;
[0054] inputting the list of candidate intervention scripts, and determining a preferred function calculating a matching score for each script , selecting the top scripts and merging core exercise paragraphs to form a target intervention script;
[0055] setting an initial duration , intensity and content sequence for the target intervention script, and synchronously recording the center value of the emotional safety interval and the allowed deviation ;
[0056] during the execution of the script, continuously obtaining real-time emotional feature indicators at a preset sampling period , calculating emotional deviation
[0057] calculating a regulation coefficient according to and adjusting the duration , intensity
[0058] and content sequence of the script according to the regulation coefficient
[0059] , intensity and content sequence of the script are dynamically updated;
[0060] if the emotional feature indicators have been restored to the safety interval for consecutive sampling periods, it is determined that the emotional feature indicators have been restored to the safety interval, and a script ending process is executed and the regulation is stopped;
[0061] the start and end time, duration sequence , intensity sequence , content sequence and corresponding emotional feature indicator sequence of the script execution are summarized to generate the intervention effect data.
[0062] A tumor patient group psychological treatment system based on artificial intelligence, comprising:
[0063] The data acquisition unit is used for obtaining the demographic information, disease course information and psychological scale scores of the patient, obtaining a first data set, collecting the physiological signals of the heart rate, blood oxygen saturation and galvanic skin response of the patient in real time by using a wearable physiological sensor, obtaining a second data set, and synchronously collecting the voice, text and facial image data of the patient in the process of the group session, and obtaining a third data set;
[0064] The emotion recognition and demand assessment unit is used for implementing multi-modal emotion recognition and psychological demand assessment on the first data set, the second data set and the third data set, and obtaining emotion feature indexes and psychological demand indexes.
[0065] The clustering grouping unit is used for performing clustering operation on the patients according to the emotion feature indexes, the psychological demand indexes and the disease course stages, and dividing the patients into a target treatment group.
[0066] The interaction monitoring and guiding unit is used for analyzing three types of interaction indexes, namely, speech balance degree, topic dominance rate and intra-group cohesion, in the process of the group session of the target treatment group in real time, and generating a guiding instruction to adjust the discussion rhythm when any one of the interaction indexes or the emotion feature indexes exceeds a corresponding preset threshold.
[0067] The intervention execution unit is used for selecting at least one intervention script of cognitive-behavioral therapy, acceptance and commitment therapy or mindfulness stress reduction therapy according to the emotion feature indexes and the interaction indexes, and dynamically adjusting the length, intensity and content of the script according to the real-time updated emotion feature indexes in the execution process until the emotion feature indexes return to a preset safe interval, and obtaining intervention effect data.
[0068] The report generation unit is used for summarizing the emotion feature indexes, the psychological demand indexes, the interaction indexes and the intervention effect data, and generating a report containing emotion trend, interaction quality and intervention effectiveness.
[0069] According to the specific embodiments of the present application, the following technical effects are provided:
[0070] The present application integrates multi-modal data acquisition, emotion-demand joint assessment, automatic clustering grouping, real-time interaction monitoring and self-adaptive intervention closed loop in series through integrated design, which makes up for the problems of insufficient intervention precision and time lag in the prior art due to single modal recognition, artificial experience grouping and lack of immediate feedback; the intervention script is dynamically adjusted until the emotion indexes return to the safe interval, and the threshold and strategy weight are fed back in the form of a report, which realizes continuous optimization, thereby significantly improving the individualized precision, resource utilization efficiency and scalability of group therapy. BRIEF DESCRIPTION OF DRAWINGS
[0071] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed in the embodiments. Obviously, the drawings described below only illustrate some of the embodiments of the present application, and for those skilled in the art, other drawings can also be obtained from these drawings without creative labor.
[0072] Figure 1 The method flowchart provided for the embodiments of the present application is as shown in
[0073] Figure 2 The system structure schematic diagram provided for the embodiments of the present application is as shown in DETAILED DESCRIPTION
[0074] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the protection scope of the present application.
[0075] The purpose of the present application is to provide a tumor patient group psychological treatment method and system based on artificial intelligence, which realizes real-time accurate evaluation, immediate intervention adjustment and closed-loop continuous optimization in a group context, and overcomes the defects of traditional artificial psychological intervention, such as narrow coverage, slow response and lack of data-driven iteration.
[0076] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the present application will be further described in detail below with reference to the drawings and specific embodiments.
[0077] Figure 1 The method flowchart provided for the embodiments of the present application is as shown in Figure 1 The present application provides a tumor patient group psychological treatment method based on artificial intelligence, which comprises:
[0078] Step 100: obtaining demographic information, disease course information and psychological scale score of the patient to obtain a first data set, using a wearable physiological sensor to collect physiological signals of heart rate, blood oxygen saturation and galvanic skin response of the patient in real time to obtain a second data set, and synchronously collecting voice, text and face image data of the patient in the process of group conversation to obtain a third data set;
[0079] Step 200: implementing multi-modal emotion recognition and psychological demand evaluation on the first data set, the second data set and the third data set to obtain emotion feature indexes and psychological demand indexes;
[0080] Step 300: According to the emotional feature indicators, psychological demand indicators and disease course stages, the clustering operation is performed on the patients to divide the patients into a target treatment group;
[0081] Step 400: Real-time analysis of the speech balance, topic dominance and group cohesion of the conversation process of the target treatment group is performed, and when any interaction indicator or emotional feature indicator exceeds the corresponding preset threshold, a guide instruction is generated to adjust the discussion pace;
[0082] Step 500: According to the emotional feature indicators and interaction indicators, at least one intervention script of cognitive-behavioral therapy, acceptance and commitment therapy or mindfulness stress reduction therapy is selected, and the duration, intensity and content of the script are dynamically adjusted according to the real-time updated emotional feature indicators during the execution process until the emotional feature indicators return to the preset safe interval, and the intervention effect data is obtained;
[0083] Step 600: The emotional feature indicators, psychological demand indicators, interaction indicators and intervention effect data are summarized to generate a report containing emotional trends, interaction quality and intervention effectiveness.
[0084] In this embodiment, the data collection of step 100 adopts a three-stage process of "multi-modal synchronous collection → timing unified calibration → encrypted and secure upload". First, the structured questionnaire template is called in the patient initial login interface, guiding the patient to fill in the demographic fields such as gender, age range and education level, and reading the disease history, treatment plan and recurrence times through the electronic medical record interface. Then, the embedded psychological scale engine is started to complete the quantitative evaluation of GAD-7 and PHQ-9. The above information is encapsulated as a JSON-Schema object with version number and field integrity hash value in the terminal local, forming the first data set, providing a verifiable basis for subsequent incremental update and traceability.
[0085] When collecting the second data set, the wristband type photoelectric plethysmogram sensor and skin electrical response sensor combination conforming to the IEC80601-2-61 standard are selected, and synchronous sampling is completed through in-phase trigger pulse: the sampling frequency of heart rate and blood oxygen saturation is set to fifty hertz, and the sampling frequency of skin electrical response is set to ten hertz. In order to reduce motion artifacts, a three-axis accelerometer is integrated at the wristband end to output a motion energy level threshold in real time. When detecting violent action, this embodiment immediately inserts a marker frame and completes adaptive light intensity compensation within three hundred milliseconds. All physiological signals are corrected for baseline drift by fast discrete wavelet transform, and then the noise discrimination model trained based on the entropy weight method is used to remove segments with a signal-to-noise ratio of less than twelve decibels, outputting high-quality physiological feature streams with unified time stamps to form the second data set.
[0086] The third dataset is collected during the group therapy session. In this embodiment, an array microphone is arranged on the top of the session room, and the voice stream of each patient is mapped to an independent channel through beamforming technology. A local FPGA performs 16-kilohertz sampling, 256-point short-time Fourier transform, and adaptive echo cancellation. A synchronously arranged ultra-wide dynamic camera performs key point tracking on the patient's face at a rate of 30 frames per second and outputs an expression action unit vector in real time. To ensure cross-modal alignment of voice, text, and image data, this embodiment broadcasts a synchronization pulse at the end of each 200-millisecond sampling window. After the terminal captures the pulse, it encapsulates the current segment as a "content frame + start and end timestamp" structure. After the voice frame is converted to a text frame through end-side automatic speech recognition, it is written into a ring buffer together with the expression vector frame, forming the third dataset, which fundamentally avoids the multi-modal misalignment risk caused by traditional "recording first and then aligning".
[0087] To protect patient privacy, this embodiment performs field-level pseudo-anonymization on the edge side: all identity identifiers are encrypted into one-time tokens through elliptic curve algorithms, and the original identifiers are then deleted from the cache. Subsequently, the three types of data sets are pushed to the cloud Kafka stream through a TLS1.3 encrypted channel, along with a SHA-256 integrity check value.
[0088] In step 200 of this embodiment, uniform preprocessing is first performed on the multi-modal raw information of the third dataset: noise reduction, framing, and Mel spectrum extraction are sequentially performed on the voice data, tokenization, stop word removal, and embedding vector mapping are performed on the text data obtained through voice recognition, and expression action unit encoding is output after face detection and key point positioning on the facial image data. At the same time, the heart rate variability and skin conductance response amplitude of the physiological signals in the second dataset are calculated to form a single-modal feature set required for subsequent modeling. This embodiment calls a convolution-long short-term memory hybrid model to infer a voice emotion vector, a bidirectional gated recurrent network to infer a text emotion vector, a residual convolutional network to infer a facial emotion vector, and a random forest classifier to infer a physiological emotion vector according to the differentiated time sequence characteristics of each modal feature. After synchronously registering the four types of emotion vectors according to the time stamp, they are input into an attention fusion network to obtain a comprehensive emotion feature index that retains both inter-modal complementary information and time sequence weight information. Finally, this embodiment jointly encodes the comprehensive emotion feature index and the psychological scale scores in the first dataset to form a fixed-length feature vector, which is input into a supervised training gradient boosting decision tree model. The model output is a psychological demand index mapped from the patient's individual needs, providing quantitative evidence for subsequent clustering, grouping, and intervention script selection.
[0089] In the preferred embodiment, the embodiment first reads the original timestamps of the speech emotion vector, the text emotion vector, the facial emotion vector and the physiological emotion vector, and uniformly divides the entire conversation sequence into continuous time windows W with a preset uniform sampling period Δt as a step k Each time window thus corresponds to a unique start and end time, thereby forming a uniform time reference, so that the four types of emotion vectors can be aligned and compared on the same time axis.
[0090] Subsequently, for any time window W k , if it is found that there are missing frames in the window, the embodiment fills in the gaps by linear interpolation, so that the four types of emotion vectors of speech, text, face and physiology have consistent sequence length and time index in the window. After completion, the four types of aligned emotion vectors are stacked along the feature dimension in the fixed modal order of speech-text-face-physiology, to obtain a multi-modal feature matrix M k .
[0091] Next, the embodiment inputs the multi-modal feature matrix M k into the multi-head self-attention fusion network, and the network calculates the weight coefficients of the four types of modalities through parallel attention heads and generates a weighted fusion vector F k ; after all the time windows output the corresponding F k , the embodiment performs pooling operation (such as taking the average or maximum value) on these weighted fusion vectors, thereby obtaining a comprehensive emotion feature index with both time sequence representation and modality complementarity, providing a unified quantitative basis for subsequent psychological demand assessment and intervention decision-making.
[0092] The embodiment first performs zero-mean normalization on the comprehensive emotion feature index and interval normalization on the psychological scale score to correct the dimension difference and suppress the interference of abnormal values on the model weight. After normalization, the embodiment concatenates a joint feature vector with a length of two hundred and fifty-six dimensions in the fixed order of speech emotion vector, text emotion vector, facial emotion vector, physiological emotion vector and scale score, and writes a uniform timestamp and a hash check field at the end to ensure data integrity when tracing back. The vector serves as a fusion carrier of multi-modal emotion information and subjective scale information, laying a foundation for precise prediction of individual psychological needs.
[0093] To avoid redundant dimensions diluting the key signals, the embodiment calls a recursive feature elimination module to dynamically evaluate the contribution of each dimension to the label based on the cross-entropy loss reduction amplitude. When the contribution of a certain dimension is lower than 0.5%, it is marked as a candidate for deletion, and finally removed after the contribution has not recovered for three consecutive rounds. Through this process, the original 256-dimensional feature is compressed to a 48-dimensional core vector, effectively improving the subsequent reasoning speed and reducing the risk of overfitting. The embodiment then inputs the simplified feature vector into the gradient boosting decision tree model that has completed supervised training. The model uses five-fold cross-validation and early stopping strategy in the training stage, and the final tree depth is set to six, the sub-sample sampling rate is set to 0.8, and the learning rate is set to 0.05, ensuring the robustness in a small sample fluctuation environment.
[0094] After obtaining the original output of the model, the embodiment uses the temperature scaling method for confidence calibration. First, use the validation set as input, and solve the global temperature factor T by minimizing the log loss. When the validation set loss converges and the reduction amplitude is less than 0.2%, terminate the iteration. Then use the temperature factor to perform a logarithmic analysis transformation on all prediction scores, output the calibrated psychological demand indicators, and record them in the database as reliable basis for subsequent clustering grouping and intervention script selection.
[0095] The embodiment first concatenates the emotional feature indicators, psychological demand indicators and disease stage encodings of each patient in turn to generate a fixed-length feature vector; then performs zero-mean normalization on each dimension of the vector to obtain a set of standardized patient feature vectors with consistent dimensions. Within the preset clustering number search interval, the embodiment calculates the average silhouette coefficient, disease stage purity and Davies-Bouldin index for each candidate clustering number, then scales the last three to the standard interval of zero to one, and combines them into a comprehensive clustering quality indicator according to the weighting rule, so as to select the clustering number with the highest comprehensive indicator as the optimal value.
[0096] After determining the optimal clustering number, the embodiment takes all standardized patient feature vectors as the sample space, uses an improved centroid initialization strategy to make the initial centroid distribution more balanced; then iteratively assigns samples and updates centroid positions according to the Euclidean distance metric, and stops iteration when the centroid displacement of two consecutive iterations is less than the preset convergence threshold, obtaining the preliminary clustering result.
[0097] To improve the accuracy of boundary sample division, the embodiment detects samples that are close to two adjacent centroids at the same time and have a distance difference less than a set threshold, and calls a neighborhood algorithm based on the same distance metric to reevaluate their belonging clusters; if necessary, adjust the cluster attribution of the sample, and then mark the corrected clusters as the target treatment group. At the same time, the embodiment retains the calculation results of the above quality indicators within the entire candidate clustering number search interval, facilitating subsequent performance comparison and model updating.
[0098] As an optional implementation, the embodiment can divide patients into the following two categories:
[0099] Grouping by disease stage: classification based on the patient's diagnosis stage, treatment process to provide more matched psychological support. Specifically including:
[0100] (1) Initial diagnosis adaptation group: suitable for newly diagnosed patients, focusing on psychological impact and information support.
[0101] (2) Treatment process group: suitable for patients undergoing treatment, focusing on side effect coping and psychological adjustment.
[0102] (3) Rehabilitation and long-term management group: suitable for the recovery period after treatment, focusing on recurrence anxiety and quality of life.
[0103] (4) Palliative care group: suitable for the late stage or hospice care stage, focusing on life meaning and emotional support.
[0104] Grouping by psychological state, specifically including:
[0105] (1) Grouping based on emotional detection and psychological assessment scale to match suitable intervention strategies.
[0106] (2) High stress / anxiety group (suitable for patients with high anxiety scores, providing relaxation training, mindfulness intervention).
[0107] (3) Depression risk group (suitable for patients with high depression scores, focusing on cognitive adjustment and hope enhancement)
[0108] (4) Positive coping group (suitable for patients with good psychological regulation, encouraging peer support and experience sharing)
[0109] As shown in Table 1, after the comprehensive emotional indicators, psychological demand indicators and disease course stage coding are completed, the grouping inference module is called to classify individual patients using a decision tree. The module first retrieves the disease course node label. If it is newly diagnosed, it is directly marked as the initial diagnosis adaptation group. If it is in the chemotherapy or radiotherapy cycle, it enters the treatment process group. If it has completed the main treatment and has no signs of progression, it is marked as the rehabilitation management group. If the disease enters the irreversible stage, it is divided into the palliative care group. For samples not uniquely determined by the above disease course labels, the embodiment continues to compare the physiological arousal and anxiety index. If the index is higher than the threshold, it is classified into the high stress and anxiety group. Otherwise, calculate the depression prediction score. If it is higher than the threshold, it is classified into the depression risk group. The remaining samples are classified into the positive coping group. After grouping, a timestamped group label is generated for each patient and pushed to the script scheduler.
[0110] The script dispatcher generates intervention programs according to the proportion weight set in the table. Taking the treatment process group as an example, the dispatcher retrieves all cognitive behavioral therapy and mindfulness stress reduction fragments in the script repository, calls the random sampling function to select cognitive restructuring or behavioral experiment paragraphs with a probability of 70%, selects breathing meditation or body scan paragraphs with a probability of 30%, and arranges them in the order of cognitive restructuring first, breathing meditation second, and a round of execution unit. The initial diagnosis adaptation group and the palliative treatment group use acceptance and commitment therapy as the main method, and the dispatcher organizes the scripts in the order of value clarification exercise, emotion acceptance exercise and mindfulness awareness exercise, and ensures that the value clarification exercise accounts for no less than the set weight in a single round of script. Each time a round of execution unit is generated, the dispatcher writes it into the queue waiting for the real-time emotion regulator to call.
[0111] During the intervention execution, the real-time emotion regulator retrieves the latest emotional bias every five seconds and adjusts the script rhythm accordingly. If the bias exceeds the upper limit, the regulator preferentially prolongs the main therapy paragraphs with a higher proportion in the current script, shortens the secondary therapy paragraphs, and inserts additional short meditation modules if necessary; if the bias is below the lower limit, the regulator adjusts in the opposite direction to reduce the load. After the emotion stabilizes for five consecutive sampling periods, the regulator triggers the end process and combines the execution log, emotion trajectory and adjustment record into the intervention effect database, providing traceable data for subsequent efficacy evaluation and model retraining.
[0112] Table 1
[0113] Group type Primary therapy Secondary therapy Selection criteria Initial diagnosis group Acceptance and commitment therapy (60%) Cognitive behavioral therapy (40%) Newly diagnosed patients have a strong psychological impact and need to establish a sense of value and meaning Treatment process group Cognitive behavioral therapy (70%) Mindfulness stress reduction (30%) Negative thoughts are triggered by side effects, requiring cognitive restructuring Rehabilitation management group Mindfulness stress reduction (50%) Acceptance and commitment therapy (50%) Prevent relapse anxiety and improve present-moment awareness Palliative care group Acceptance and commitment therapy (80%) Mindfulness stress reduction (20%) Patients in the late stage need more exploration of life meaning and acceptance of suffering High stress / anxiety group Mindfulness stress reduction (70%) Cognitive behavioral therapy (30%) High physiological arousal requires reducing anxiety levels through mindfulness first Depression risk group Cognitive behavioral therapy (60%) Acceptance and commitment therapy (40%) Cognitive distortion is obvious and requires cognitive intervention Positive coping group Acceptance and commitment therapy (50%) Mindfulness stress reduction (50%) Psychological resilience is good, and value-oriented behavior can be deepened
[0114] In this embodiment, a real-time performance buffer is maintained for each round of script during intervention execution. When the cumulative emotional index decreases by less than 3% and the interaction balance improves significantly within three sampling periods, it is determined that the current main therapy paragraph is in the inefficient interval, and the dispatcher immediately reads the exercise unit matching the current emotional bias from the secondary therapy library and inserts it into the next time window to replace the last third of the original main therapy; if the emotional index returns to the safe interval and the cohesion index rises by more than 5% in the next two sampling periods, the time ratio of the main therapy to the secondary therapy is restored to the preset weight, otherwise the replacement structure is maintained until the end of the current script and the switching timestamp, replaced paragraph identifier and emotional change amount are recorded in the log to ensure that the therapy switching is based on quantitative basis and is fully traceable.
[0115] Optionally, the embodiment first calculates two types of average distances for each patient in the standardized patient feature vector set under each candidate cluster number: one is the average Euclidean distance between the patient and other members of the cluster, and the other is the average Euclidean distance between the patient and the nearest neighbor member of other clusters; then obtains the silhouette coefficient of a single sample by taking the ratio of the difference between the two types of average distances and the larger one, and obtains the average silhouette coefficient corresponding to the current cluster number by taking the arithmetic mean of the silhouette coefficients of all patients, to measure the compactness within the cluster and the separation between clusters.
[0116] Under the same candidate cluster number, the embodiment also calculates the distribution of disease stages of each cluster, calculates the proportion of patients in the main disease stage in the total number of patients in the cluster as the purity of the cluster, and obtains the overall disease stage purity by weighting the proportion of samples in the cluster; then calculates the Davies-Bouldin index, that is, calculates the average dispersion within each cluster, and then takes the average of the maximum ratio of the distance between the cluster centers and the adjacent clusters to obtain the separability index between clusters. Finally, the minimum and maximum values of the index are recorded in the range of all candidate cluster numbers, and the index corresponding to each cluster number is mapped to the interval of zero to one to form a comparable standardized index, which is used together with the average silhouette coefficient and the disease stage purity for comprehensive quality evaluation.
[0117] Specifically, the embodiment first performs speech activity detection at the audio acquisition end with a frame width of 25 milliseconds, and sends the continuous speech segment to the speech recognition engine in real time; the recognition engine is embedded with a speaker embedding module based on vector quantization, which completes speaker assignment at the frame level and generates speaker labels accurate to milliseconds. The audio stream is immediately written into a ring buffer after single-frame recognition, and the buffer exposes an interface for frame index retrieval, allowing the subsequent sliding time window to continuously pull the latest text segment in steps of ten frames. In order to ensure consistency between speech and emotion data, the embodiment appends a self-incrementing session sequence number and a sampling start timestamp to each text segment generated, thereby achieving cross-modal synchronization.
[0118] In the sliding window processing stage, the embodiment maintains a speaking duration accumulator for each member; at the end of the window, the accumulator values are read in sequence, divided by the total window duration, squared and summed, and then the result of subtracting the square sum is obtained to obtain the speaking balance. The text topic inference uses a fine-tuned sparse hidden vector model, which only retains the most probable topic as the dominant topic within the window, and then calculates the proportion of the number of words of each member under the topic, where the maximum proportion is the topic dominance rate. The intra-group cohesion index is linearly scaled from the average cosine similarity of sentence vectors, with a scaling factor of two to ensure that the index value interval is consistent with the speaking balance, facilitating subsequent unified threshold management.
[0119] The threshold detection adopts an event-driven manner; after the window result is written into the shared memory, the threshold comparison function is called immediately, and the speech balance, topic dominance, group cohesion, and comprehensive emotion score at the end of the sliding window are compared with the respective thresholds one by one. When any indicator exceeds the threshold, the embodiment generates a structured guidance instruction, which includes the target member identity, the prompt category, and the context script number to be executed; the instruction is delivered to the interactive interface rendering process through an asynchronous message queue, achieving one-second update to the end. The instruction is also written into the conversation log and triggers the adaptive threshold scheduler to record the information of the current triggering, providing a basis for subsequent threshold review and personalized model fine-tuning.
[0120] Optionally, in step 500 of the embodiment, a mapping table of three typical states and intervention scripts is first constructed, and a trigger threshold range is defined for each state. After the real-time comprehensive emotion indicators and sliding window interaction indicators enter the rule engine, according to the threshold, the current state is determined to belong to one of high anxiety with significant cognitive dissonance, high avoidance with insufficient value focus, or high physiological hyperactivity with insufficient emotional regulation, or a combination thereof. The rule engine extracts matching items from the script repository according to this and writes them into the candidate list, while labeling each script with expected emotion improvement amplitude, interaction recovery rate, and historical success rate, which are pre-statistical completed through an offline validation set to ensure that they do not rely on manual experience parameter tuning.
[0121] Then the embodiment performs script optimization on the candidate list. The optimization function multiplies the expected emotion improvement amplitude by 0.5, multiplies the interaction recovery rate by 0.3, and multiplies the historical success rate by 0.2, and the sum of the three results is the matching score. The top two scripts in the matching score ranking are selected, and their core exercise paragraphs are spliced into the target script in the order of cognitive restructuring, emotional acceptance, and breathing meditation. When the target script is generated, it is written into the initial execution duration of ten minutes, the intensity of level one, the content sequence index, and the emotion safety center value and the maximum allowed deviation of ten percent, providing boundary conditions for subsequent adaptive adjustment.
[0122] During script execution, the embodiment collects real-time emotion indicators every five seconds and calculates the deviation rate, which is the difference between the real-time emotion indicator and the safety center value divided by the safety center value. When the deviation rate is greater than ten percent, the script duration is automatically extended by twenty percent at the next exercise segment, and the intensity is increased by one level; when the deviation rate is less than negative ten percent, the script duration is shortened by twenty percent, and the intensity is reduced by one level; when the absolute value of the deviation rate is less than ten percent for five consecutive times, the end process is triggered, the meditation is ended, and the summary feedback is output. After execution, the embodiment records the complete time sequence, intensity sequence, content sequence, and emotion trajectory, and generates intervention effect data, which is used for subsequent threshold fine-tuning and script version iteration.
[0123] Further, in step 600 of the present embodiment, the present embodiment starts the asynchronous pipeline immediately after the intervention link ends, and uniformly pulls the archived emotional feature indicators, psychological demand indicators, interaction indicators, and intervention effect data. The pipeline first uses millisecond-level timestamps as anchor points to perform forward filling on the four types of time series to ensure consistent column length, and then uses the sliding mean method to suppress short-term spikes, extract emotional trend curves, interaction quality statistics, and intervention effectiveness curves, and write all original fields and derived fields to an encrypted cache, while generating a hash digest for subsequent integrity checking. The present embodiment then calls the report generation engine to arrange the emotional trend, interaction quality, and intervention effectiveness sections according to the combined template. The engine first visualizes the numerical results and key inflection points as line charts and bar charts, and then generates a natural language summary below the chart, pointing out the emotional fluctuation peak time, interaction balance improvement amplitude, and intervention achievement rate, etc. After completing the layout, the full text is output as a PDF-A to the target therapy group, and a machine-readable data package is attached, so that the medical information system can directly call and provide closed-loop data for subsequent algorithm retraining.
[0124] Corresponding to the above method, as shown in Figure 2 the present embodiment also provides a tumor patient group psychological treatment system based on artificial intelligence, comprising:
[0125] a data acquisition unit for acquiring demographic information, disease information and psychological scale scores of patients to obtain a first data set, using a wearable physiological sensor to collect physiological signals of heart rate, blood oxygen saturation and galvanic skin response of patients in real time to obtain a second data set, and synchronously collecting voice, text and facial image data of patients during group sessions to obtain a third data set;
[0126] an emotion recognition and demand assessment unit for implementing multi-modal emotion recognition and psychological demand assessment on the first data set, the second data set and the third data set to obtain emotional feature indicators and psychological demand indicators;
[0127] a clustering and grouping unit for performing clustering operation on patients according to the emotional feature indicators, the psychological demand indicators and the disease stage to divide the patients into a target therapy group;
[0128] an interaction monitoring and guiding unit for real-time analysis of speech balance, topic dominance and group cohesion of the target therapy group during the session process, and generating a guiding instruction to adjust the discussion pace when any of the interaction indicators or the emotional feature indicators exceeds the corresponding preset threshold;
[0129] An intervention execution unit is configured to select at least one intervention script of cognitive-behavioral therapy, acceptance and commitment therapy or mindfulness stress reduction therapy according to the emotional feature index and the interaction index, and dynamically adjust the length, intensity and content of the script during execution according to the real-time updated emotional feature index until the emotional feature index returns to a preset safe interval, so as to obtain intervention effect data.
[0130] A report generation unit is configured to aggregate the emotional feature index, the psychological demand index, the interaction index and the intervention effect data, and generate a report containing emotional trends, interaction quality and intervention effectiveness.
[0131] The present application has the following advantages:
[0132] (1) The present application realizes a multi-modal and full-process data closed loop, which simultaneously collects physiological signals, voice-text-face information and psychological scales, and aligns them with a unified time reference, thereby providing high-integrity and high-timeliness objective basis for emotional recognition, grouping decision and intervention evaluation.
[0133] (2) The present application optimizes the division of treatment groups by a comprehensive clustering quality evaluation function, so that the disease courses in the groups are more consistent and the emotional demands are closer, thereby significantly improving the pertinence of group intervention and resource utilization efficiency.
[0134] (3) The present application introduces real-time interaction monitoring and adaptive script adjustment mechanism, which can generate guidance instructions and dynamically adjust the intervention length, intensity and content within a second-level time scale when emotional or interaction abnormalities occur, thereby ensuring the safety, continuity and individualization of the treatment process.
[0135] (4) The present application visualizes the whole-process quantification results in the form of a three-dimensional report of emotional trends, interaction quality and intervention effectiveness, and encodes them through hash check and PDF-A three specifications, which not only meets the medical compliance requirements, but also provides a traceable evidence chain for subsequent model iteration and clinical decision-making.
[0136] The embodiments in the specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts of each embodiment can be referred to each other.
[0137] The principles and implementation modes of the present application are described by specific examples in this paper. The above description of the embodiments is only to help understand the method of the present application and its core idea; at the same time, for those skilled in the art, according to the idea of the present application, the specific implementation mode and application range will be changed. In summary, the content of the specification should not be understood as a limitation of the present application.
Claims
1. An artificial intelligence-based group psychotherapy method for cancer patients, characterized by: include: Obtaining the patient's demographic information, disease course information, and psychological scale scores to obtain the first data set, using wearable physiological sensors to collect the patient's heart rate, blood oxygen saturation, and skin galvanic response physiological signals in real time to obtain the second data set, and synchronously collecting the patient's voice, text, and facial image data during the group conversation to obtain the third data set; Performing multimodal emotion recognition and psychological needs assessment on the first data set, the second data set, and the third data set to obtain emotion feature indices and psychological needs indices; performing a clustering operation on the patients according to the emotional characteristic index, the psychological need index, and the disease stage, and dividing the patients into target treatment groups; Analyze the conversation process of the target treatment group in real time: three types of interaction indicators: speech balance, topic dominance, and group cohesion. When any of the interaction indicators or the emotional characteristic indicators exceeds the corresponding preset threshold, generate guidance instructions to adjust the discussion rhythm; According to the emotional characteristic index and the interaction index, at least one intervention script from cognitive-behavioral therapy, acceptance-commitment therapy, or mindfulness-based stress reduction therapy is selected, and during execution, the duration, intensity, and content of the script are dynamically adjusted according to the real-time updated emotional characteristic index until the emotional characteristic index returns to a preset safe range, thereby obtaining intervention effect data; The emotional characteristic indicators, the psychological needs indicators, the interaction indicators and the intervention effect data are summarized to generate a report including emotional trends, interaction quality and intervention effectiveness.
2. The method of group psychotherapy for cancer patients based on artificial intelligence according to claim 1, characterized in that: Performing multimodal emotion recognition and psychological needs assessment on the first data set, the second data set, and the third data set to obtain emotion feature indices and psychological needs indices, including: Performing noise reduction, framing, and Mel spectrum extraction on the speech data in the third data set, performing word segmentation, stop word removal, and embedding vector mapping on the text data in the third data set, and performing face detection, key point positioning, and expression action unit encoding on the facial image data in the third data set; Calculating physiological emotion features for the physiological signals in the second data set; the physiological emotion features include: heart rate variability and skin electrical response amplitude; The processed speech data is input into the convolution-long short-term memory hybrid model to output the speech emotion vector, and the processed text data is input into the bidirectional gated recurrent network to output the text emotion vector; Inputting the facial action unit encoding into the residual convolutional network to output a facial emotion vector, and inputting the physiological emotion feature into the random forest classifier to output a physiological emotion vector; Synchronizing the voice emotion vector, the text emotion vector, the facial emotion vector, and the physiological emotion vector according to timestamps, and inputting the synchronized four types of emotion vectors into an attention fusion network to calculate a comprehensive emotion feature index; The comprehensive emotional feature index and the psychological scale score in the first data set are jointly encoded, and the obtained joint encoding is input into a supervised trained gradient boosting decision tree model to output a psychological need index.
3. The method of group psychotherapy for cancer patients based on artificial intelligence according to claim 2, characterized in that: The speech emotion vector, the text emotion vector, the facial emotion vector, and the physiological emotion vector are synchronized according to timestamps, and the synchronized four types of emotion vectors are input into the attention fusion network to calculate the comprehensive emotion feature index, including: Based on the timestamps of the speech emotion vector, the text emotion vector, the facial emotion vector and the physiological emotion vector, a continuous time window W is divided according to a preset uniform sampling period Δt. k , establish a unified time base; In each time window W k The missing emotion vectors are supplemented by linear interpolation so that the speech emotion vector, the text emotion vector, the facial emotion vector and the physiological emotion vector are within the time window W. k Have consistent sequence length and time index; The time window W k The four aligned emotion vectors are stacked in a fixed modal order to construct a multimodal feature matrix M k ; The multimodal feature matrix M k Input the multi-head self-attention fusion network, calculate the weight coefficient of each modality, and output the weighted fusion vector F k ; The weighted fusion vector F for all time windows k Perform pooling operations to obtain comprehensive emotion feature indicators.
4. The method of group psychotherapy for cancer patients based on artificial intelligence according to claim 3, characterized in that: The comprehensive emotional feature index and the psychological scale score in the first data set are jointly encoded, and the obtained joint encoding is input into the supervised trained gradient boosting decision tree model to output the psychological need index, including: Performing zero-mean normalization processing on the comprehensive emotional characteristic index and performing interval normalization processing on the psychological scale score so that the two types of characteristics are in comparable dimensions; The normalized comprehensive emotional characteristic index and the psychological scale score are concatenated into a fixed-length joint feature vector according to the preset modal order; Input the fixed-length joint feature vector into a recursive feature elimination module to eliminate dimensions whose correlation with the target label is lower than a preset threshold to obtain a streamlined feature vector; Inputting the simplified feature vector into a gradient boosting decision tree model that has completed supervised training, and outputting the corresponding psychological need index; Temperature scaling calibration was performed on the psychological demand index to improve the confidence consistency of the index across different patient groups.
5. The method of group psychotherapy for cancer patients based on artificial intelligence according to claim 1, characterized in that: Based on the emotional characteristic index, the psychological need index, and the stage of the disease, a clustering operation is performed on the patients to divide them into target treatment groups, including: For each patient, the emotional characteristic index, the psychological demand index and the disease stage code are concatenated in a preset order to form a fixed-length patient feature vector; Normalizing each dimension in the fixed-length patient feature vector according to a zero-mean normalization rule to obtain a set of standardized patient feature vectors; In the candidate cluster number interval The average silhouette coefficient is calculated based on the standardized patient feature vector set. , Purity of disease course Davidson-Boulding Index , and 0-1 normalized to standard exponent , and then select the comprehensive clustering quality index according to the formula The maximum optimal number of clusters ;in, , ;in, and are the minimum and maximum values of the candidate cluster number interval respectively; The standardized patient feature vector set is used as the sample space and K-means++ is used to Initialize the centroids, iteratively distribute each vector in the standardized patient feature vector set under the Euclidean distance metric and update the centroids until two consecutive centroid displacements are less than a preset convergence threshold; For the boundary samples in the standardized patient feature vector set that are close to two adjacent centroids and whose distance difference is lower than a set threshold, the K nearest neighbor algorithm is called to re-evaluate the neighborhood distribution and adjust the cluster affiliation; After correction Based on the clusters, target treatment groups were generated; In the candidate cluster number interval The average silhouette coefficient is calculated based on the standardized patient feature vector set. , Purity of disease course Davidson-Boulding Index , and 0-1 normalized to standard exponent .
6. The method of group psychotherapy for cancer patients based on artificial intelligence according to claim 5, characterized in that: In the candidate cluster number interval The average silhouette coefficient is calculated based on the standardized patient feature vector set. , Purity of disease course Davidson-Boulding Index , and 0-1 normalized to standard exponent ,include: For any patient sample in the cluster , calculate the average Euclidean distance to other samples in the same cluster , and the average Euclidean distance to the nearest neighbor cluster samples , then press Obtain the single sample silhouette coefficient; All of samples Take the arithmetic mean to get the average silhouette coefficient under the number of clusters ; Record Clusters contain patients, of whom the number of patients in the main course of the disease was , then The purity of the disease course stage of each cluster is , sum the purity of each cluster according to its sample proportion, and get ; Find the original Davidson-Boulding index , and in the range of all candidate cluster numbers Internal Records and , and execute , complete 0-1 normalization; For the The average Euclidean distance from cluster samples to cluster centroid, For the With the The Euclidean distance between the centroids of clusters.
7. The method of group psychotherapy for cancer patients based on artificial intelligence according to claim 1, characterized in that: The conversation process of the target treatment group is analyzed in real time for three interaction indicators: speech balance, topic dominance, and group cohesion. When any of the interaction indicators or the emotional characteristic indicators exceeds the corresponding preset threshold, guidance instructions are generated to adjust the discussion rhythm, including: performing automatic speech recognition and speaker separation on the audio stream of the target treatment group to obtain a text sequence with timestamps and speaker labels; In the sliding time window, count the speaking time of each member , calculate the total duration , and press Get speech balance ; The number of members in the target treatment group; Perform topic distribution inference on the transcribed text and count the number of words each member spoke on the current dominant topic , calculate the total number of words , and press Get the topic dominance rate ; Use sentence vector representation to calculate the average cosine similarity of all speech pairs in the same class , and press Gaining cohesion within the group ; Balance the speech , Topic Dominance Rate , intra-group cohesion and emotional characteristic indicators With their respective preset thresholds 、 、 、 Compare, if If any of the indicators exceeds the corresponding threshold, a guidance instruction is generated and pushed to the conversation interface to adjust the discussion rhythm or invite low-participation members to speak.
8. The method of group psychotherapy for cancer patients based on artificial intelligence according to claim 1, characterized in that: Based on the emotional characteristic index and the interaction index, at least one intervention script from cognitive-behavioral therapy, acceptance-commitment therapy, or mindfulness-based stress reduction therapy is selected, and during execution, the duration, intensity, and content of the script are dynamically adjusted according to the real-time updated emotional characteristic index until the emotional characteristic index returns to a preset safe range, thereby obtaining intervention effect data, including: Current comprehensive emotional characteristic index and interaction metrics collection Conduct rule matching to generate at least one candidate intervention script list, where the cognitive-behavioral therapy script is targeted at the high anxiety-high cognitive dissonance state, the acceptance-commitment therapy script is targeted at the high avoidance-low value focus state, and the mindfulness-based stress reduction therapy script is targeted at the high physiological arousal-low emotion regulation state; Taking the candidate intervention script list as input, according to the optimization function Calculate the matching score for each script , select The largest front Scripts were copied and core practice paragraphs were combined to form a target intervention script; Set the initial duration for the target intervention script ,strength and content sequence , and simultaneously record the center value of the emotional safety zone With allowable deviation ; During script execution, the preset sampling period Continuously obtain real-time emotional feature indicators , calculate the sentiment bias ; according to Calculate the adjustment coefficient , and press ; Dynamic update script duration ,strength and content sequence ; If continuous Within a sampling period , it is determined that the emotional characteristic index has returned to a safe range, the script closing process is executed and the adjustment is stopped; Summarize the start and end time and duration of script execution , intensity sequence , content sequence And the corresponding emotional feature indicator sequence , generating the intervention effect data.
9. An artificial intelligence-based group psychotherapy system for cancer patients, characterized by: include: a data acquisition unit configured to obtain the patient's demographic information, disease course information, and psychological scale scores to obtain a first data set; to use wearable physiological sensors to collect the patient's heart rate, blood oxygen saturation, and skin galvanic response physiological signals in real time to obtain a second data set; and to synchronously collect the patient's voice, text, and facial image data during the group conversation to obtain a third data set; an emotion recognition and needs assessment unit, configured to perform multimodal emotion recognition and psychological needs assessment on the first data set, the second data set, and the third data set to obtain emotion feature indices and psychological needs indices; A clustering and grouping unit, configured to perform a clustering operation on the patients based on the emotional characteristic index, the psychological need index, and the disease stage, and divide the patients into target treatment groups; an interactive monitoring and guidance unit for treating the target group; The conversation process is analyzed in real time by three interaction indicators: speech balance, topic dominance, and group cohesion. When any of the interaction indicators or the emotional characteristic indicators exceeds the corresponding preset threshold, guidance instructions are generated to adjust the discussion rhythm. an intervention execution unit, configured to select at least one intervention script from cognitive-behavioral therapy, acceptance-commitment therapy, or mindfulness-based stress reduction therapy based on the emotional characteristic index and the interaction index, and dynamically adjust the duration, intensity, and content of the script according to the emotional characteristic index updated in real time during execution, until the emotional characteristic index returns to a preset safe range, thereby obtaining intervention effect data; A report generating unit is used to summarize the emotional characteristic indicators, the psychological need indicators, the interaction indicators and the intervention effect data to generate a report including emotional trends, interaction quality and intervention effectiveness.
Citation Information
Cited By
Alzheimer's disease data completion method based on multi-modal generation and fusion
CN121439269A