Rehabilitation state evaluation system for language psychological disorder patients based on multi-agent cooperation

By constructing a multi-agent collaborative assessment system for the rehabilitation status of patients with language-related psychological disorders, and utilizing the Transformer-XL and Wav2Vec2.0 architecture, the system achieves collaborative processing of multimodal data, solves the problems of poor dialect adaptability and single assessment dimensions, improves assessment accuracy and efficiency, provides interpretable assessment results, and is applicable to hospitals and community rehabilitation institutions at all levels.

CN121331389BActive Publication Date: 2026-04-21INSPUR SOFTWARE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
INSPUR SOFTWARE TECH CO LTD
Filing Date
2025-12-17
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing technologies for assessing patients with language-related psychological disorders suffer from poor dialect compatibility, fragmented multimodal data, limited assessment dimensions, and weak interpretability, resulting in low accuracy and efficiency, and failing to meet clinical needs.

Method used

A rehabilitation status assessment system for patients with language-related psychological disorders based on multi-agent collaboration was constructed. The system uses Transformer-XL and Wav2Vec2.0 as its basic architecture, integrates a multimodal attention mechanism, constructs a base model, and combines dialect-specific, text-specific, and graph-specific assessment agents. The collaborative processing and conflict resolution of multimodal data are achieved through a weighted fusion algorithm.

Benefits of technology

It enables a comprehensive assessment of language and psychological state, improves the accuracy and efficiency of assessment, adapts to different dialects and modalities, provides interpretable assessment results, and supports the development of personalized rehabilitation plans in clinical practice.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121331389B_ABST
    Figure CN121331389B_ABST
Patent Text Reader

Abstract

This invention discloses a rehabilitation status assessment system for patients with language-related psychological disorders based on multi-agent collaboration, belonging to the field of artificial intelligence technology. The technical problem it aims to solve is how to achieve a comprehensive assessment of language and psychology, improve the accuracy and efficiency of the assessment, and overcome the shortcomings of existing technologies such as poor dialect adaptation, multimodal fragmentation, single assessment dimension, and weak interpretability. It aims to meet the clinical needs for accurate and efficient assessment of patients with language-related psychological disorders. The technical solution is as follows: The system includes a base model construction unit, a specialized assessment agent construction unit, and a collaborative scheduling unit. The base model construction unit is used to construct a base model based on Transformer-XL and Wav2Vec2.0, integrating a multimodal attention mechanism to achieve cross-modal information interaction. The specialized assessment agent construction unit is used to construct dialect-specific assessment agents, text-specific assessment agents, and chart-specific assessment agents based on the base model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, specifically to a rehabilitation status assessment system for patients with language-related psychological disorders based on multi-agent collaboration. Background Technology

[0002] Rehabilitation assessments for patients with language-related psychological disorders need to simultaneously consider "language function" (such as pronunciation fluency and vocabulary ability) and "psychological state" (such as anxiety level and social willingness), and are significantly affected by "dialect diversity" and "data modality differences." Existing technologies have four core limitations:

[0003] ① Poor dialect adaptability: Existing assessment models are mostly trained based on Mandarin and cannot be adapted to dialects, resulting in low accuracy in identifying language defects in dialect patients and making them difficult to use in primary hospitals;

[0004] ② Fragmented processing of multimodal data: Clinical assessment data covers multiple modalities such as "voice (patient's pronunciation), text (patient's self-report / doctor's record), and charts (psychological assessment scales such as SAS anxiety scale and SDS depression scale), but existing technologies require manual processing of different modal data (such as using speech recognition models for voice and relying on manual reading of charts), without a collaborative mechanism, resulting in low assessment efficiency and easy to produce contradictory results;

[0005] ③ Single assessment dimension: Most systems only focus on language function assessment (such as pronunciation defect type, language fluency score), ignoring the correlation between psychological state and language rehabilitation. For example, anxious patients may have speech pauses due to tension. If only language ability is assessed, it is easy to misjudge as "moderate aphasia". It is necessary to combine psychological state for comprehensive judgment. Current technology lacks such cross-dimensional assessment ability.

[0006] ④ Weak clinical interpretability: The assessment results are mostly single scores without specific evidence. Doctors cannot trace the assessment logic, making it difficult to trust and use them for clinical decision-making.

[0007] Therefore, how to achieve a comprehensive assessment of language and psychology, improve the accuracy and efficiency of the assessment, overcome the shortcomings of existing technologies such as poor dialect adaptation, multimodal fragmentation, single assessment dimension, and weak interpretability, and meet the clinical needs for accurate and efficient assessment of patients with language and psychological disorders is a technical problem that urgently needs to be solved. Summary of the Invention

[0008] The technical objective of this invention is to provide a multi-agent collaborative assessment system for the rehabilitation status of patients with language and psychological disorders, in order to address how to achieve a comprehensive assessment of language and psychology, improve the accuracy and efficiency of the assessment, overcome the shortcomings of existing technologies such as poor dialect adaptation, multimodal fragmentation, single assessment dimension, and weak interpretability, and meet the clinical needs for accurate and efficient assessment of patients with language and psychological disorders.

[0009] The technical objective of this invention is achieved as follows: a multi-agent collaborative system for assessing the rehabilitation status of patients with language-related psychological disorders, comprising:

[0010] The base model building unit is used to construct a base model based on Transformer-XL (long text processing) and Wav2Vec2.0 (speech processing) architecture, integrating a multimodal attention mechanism to achieve cross-modal information interaction. The base model includes a speech branch layer, a text branch layer, a graph branch layer, and a multimodal fusion layer. The multimodal fusion layer is deployed on top of the speech, text, and graph branch layers and is connected to the cross-modal attention layer, allowing different modal representations, including speech, text, and graphs, to pay attention to each other, learn the correlation weights between each modality, and output a unified multimodal semantic representation.

[0011] The specialized evaluation agent construction unit is used to construct dialect-specific evaluation agents, text-specific evaluation agents, and chart-specific evaluation agents based on the base model. The dialect-specific evaluation agents, text-specific evaluation agents, and chart-specific evaluation agents are used to identify different modalities of input data in different dialects, texts, and charts, and to perform state evaluation based on the identified data, thereby generating dialect evaluation results, text evaluation results, and chart evaluation results.

[0012] The collaborative scheduling unit is used to integrate the dialect evaluation results, text evaluation results, and chart evaluation results generated by the dialect evaluation agent, text evaluation agent, and chart evaluation agent using a weighted fusion algorithm. This enables the fusion of results from different agents and the resolution of conflicts, ensuring the comprehensiveness and consistency of the evaluation.

[0013] Preferably, the speech branch layer adopts a Wav2Vec2.0 feature extractor and context encoder structure. The feature extractor is used to transform the original speech waveform into an acoustic feature sequence (such as 768-dimensional) through CNN stacking, and the context encoder (Transformer) is used to optimize the speech representation through contrastive learning.

[0014] The text branching layer is based on the Transformer-XL design and introduces relative position encoding and segment-level recurrence mechanism to handle dependency modeling of long texts (such as patient self-reports and scale descriptions).

[0015] The chart branch uses a CNN and Transformer structure. The CNN is used to extract the visual feature maps of the chart, and the Transformer is used to transform the feature maps into sequential visual representations.

[0016] As a preferred option, the base model building blocks include:

[0017] A multi-source data training set construction module is used to collect and preprocess dialect language data samples, multimodal psychological data samples, and clinical diagnosis and treatment data samples to construct a multi-source data training set. Dialect language data samples refer to dialect language disorder samples. Each dialect language data sample includes dialect speech (16kHz sampling rate, 1-3s per sample), dialect text annotation (corresponding dialect vocabulary and defect type), and Mandarin-mapped text. Multimodal psychological data samples include psychological assessment charts (such as paper / electronic charts and images of the SAS anxiety scale and SRS social response scale), text descriptions (patient self-reports of 'not wanting to speak, afraid of making mistakes,' and doctor's records of 'patients avoiding verbal communication'), and vocal emotional characteristics (such as speech segments with slow speech rate and low tone). Clinical diagnosis and treatment data samples include rehabilitation assessment records of patients with language and psychological disorders (language ability scores and psychological state levels annotated by doctors) and rehabilitation trajectory data (such as records of the correlation between changes in language fluency and changes in psychological state within one month).

[0018] The base model pre-training module is used to complete the dialect-Mandarin cross-linguistic alignment task, the language-psychological association prediction task, and the multimodal graph understanding task. The dialect-Mandarin cross-linguistic alignment task refers to learning the mapping relationship between dialect speech and text and Mandarin (such as mapping "falling rain" to "rain"), and optimizing the cross-linguistic representation ability of dialect features. The language-psychological association prediction task refers to predicting the psychological state level (such as "mild anxiety") based on the patient's speech (such as stuttering speech) and text (such as "nervous"), and modeling the correlation between language features and psychological state. The multimodal graph understanding task refers to extracting key information from psychological assessment graphs through CNN+Transformer to achieve automatic structuring of graph data.

[0019] The base model fine-tuning module is used to design loss functions including language assessment loss (cross-entropy), psychological assessment loss (MSE), and cross-modal consistency loss (cosine similarity). Based on the language-psychological integrated assessment data labeled by doctors, the base model is fine-tuned according to the loss functions to ensure that the assessment results output by the base model conform to the clinical diagnosis and treatment logic.

[0020] More effectively, the dialect-Mandarin cross-linguistic alignment task updates the parameters of the speech and text branches through backpropagation, thereby strengthening the semantic mapping relationship between dialects and Mandarin; specifically as follows:

[0021] After encoding the input speech and text data, cross-modal attention alignment is performed. Specifically: Speech encoding: Dialect and Mandarin speech are processed by the Wav2Vec2.0 feature extractor to obtain acoustic feature sequences respectively. and Text encoding: Semantic feature sequences are obtained from dialect and Standard Mandarin texts through the Transformer-XL text branch layer, respectively. and Cross-modal attention alignment: and , and Input the cross-modal attention layer and calculate the alignment weights between modalities using the following formula: ;in, Indicates the alignment weights between modalities; The feature dimension is represented by attention weights, which guide the base model to focus on semantically matched speech-text segments; vt stands for speech (video) and text (text) alignment, an abbreviation; T is the vector transpose symbol.

[0022] Contrast Alignment Loss: Construct positive and negative sample pairs, and optimize intermodal similarity using InfoNCE loss, the formula is as follows: ;in, Indicates the similarity between modalities; and τ represents the fusion representation of dialect and the fusion representation of Mandarin, respectively; N represents the total number of negative samples participating in the comparison; positive samples in the positive and negative samples are dialect-Mandarin pairs with the same semantics; negative samples in the positive and negative samples refer to randomly matched dialect-Mandarin pairs.

[0023] Speech-Text Matching (STM): This task uses a binary classification method to determine whether the input dialect speech and Mandarin text are semantically matched. Cross-entropy is used to measure the deviation between the binary classification prediction results of the base model and the actual matching labels (match / non-match).

[0024] The language-psychological association prediction task is as follows:

[0025] After encoding the input speech and text data, modality fusion is performed, specifically: Speech feature enhancement: Speech features are obtained through the Wav2Vec2.0 context encoder. Additional acoustic features, including fundamental frequency (F0), MFCC, and speech rate, are concatenated to enhance psychological cues in the speech. Text feature encoding involves processing patient text through a Transformer-XL text branch layer, employing a segmented loop mechanism for long texts to preserve cross-segment dependencies and extract text features. Multimodal fusion: and By learning the association weights between "voice anomaly features, text emotion words, and psychological states" through a cross-modal attention layer, a fused representation is obtained. ;

[0026] Mental cue matching task: Construct a "voice / text feature - mental cue word" matching task (e.g., voice pauses correspond to the "nervous" cue word), requiring the base model to predict the type of mental cue contained in the input features, and using cross-entropy to measure the deviation between the base model's prediction results and the true cue type labels;

[0027] Mental state level prediction: in fusion representation The system then connects to a classification head to predict psychological states from level 1 to 5. Weighted cross-entropy is used to measure the deviation between the prediction results and the true labels. The classification head includes a fully connected layer and a Softmax layer.

[0028] The multimodal graph comprehension task is as follows:

[0029] The chart image is processed by a CNN (ResNet-50) to extract shallow visual features (such as edges and textures), and then the feature map is serialized (divided into 16×16 patches) by a Transformer encoder (visual Transformer) to obtain the visual representation. ;

[0030] Text-assisted encoding: The text extracted from the chart (such as scale title and question description) is recognized by OCR and then processed through the Transformer-XL text branch layer to obtain the text representation. ;

[0031] Cross-modal attention fusion: integrating visual representations With text representation Cross-modal attention interaction is achieved through a cross-modal attention layer to obtain graph fusion representation. ;

[0032] Chart type classification: Determine whether the scale type is a binary or multi-class task, and use the cross-entropy function to measure the deviation between the output scale type prediction probability distribution and the true label, so as to help the base model focus on the layout features of different scales.

[0033] Entity recognition in charts: The structured extraction of key information from the scale is transformed into a sequence labeling task. The CRF layer is used to optimize the labeled sequence, and the negative log-likelihood loss function of CRF is used to measure the deviation between the labeled sequence predicted by the base model and the true labeled sequence.

[0034] Numerical prediction task: After fusing representations, the regression head is connected to predict the numerical information of the total score and individual item scores of the scale. The deviation between the predicted scores and the true labeled scores is measured by the MSE loss function.

[0035] More preferably, the loss function used in the base model fine-tuning module is as follows:

[0036] ;

[0037] in, Represents the loss function; Indicates language assessment loss; Indicates psychological assessment loss; Indicates cross-modal consistency loss; This represents the language evaluation weighting coefficient; This represents the weighting coefficient of the psychological assessment; Indicates the modal alignment weight coefficient; The learning rate uses cosine annealing (initial annealing). , minimum );

[0038] The language evaluation loss uses cross-entropy. For a multi-classification task targeting language defect types (such as naming aphasia, fluency defects, etc., totaling 6 categories), the formula is as follows:

[0039] ;

[0040] Where B (equal to 32) represents the training batch size; C (equal to 6) represents the number of language defect categories; This represents the one-hot encoding of sample b in category c (e.g., "named aphasia" corresponds to [1,0,0,0,0,0]). This represents the predicted probability of the base model, requiring training to converge. ;

[0041] The psychological assessment loss was calculated using the MSE (Mental State Examination) regression task for mental state scores (0-100 points, such as anxiety level), with the following formula:

[0042] ;

[0043] in, This indicates the doctor's actual psychological score (e.g., 52 points on the SAS scale). This represents the base model's predicted score, requiring convergence. (Corresponding scoring error ≤ 5 points);

[0044] The cross-modal consistency loss uses cosine similarity to ensure alignment of speech and text modal features, as shown in the following formula:

[0045] ;

[0046] in, This represents the speech feature vector of sample b; This represents the corresponding text feature vector; For L2 norm, it is required that (Corresponding cosine similarity ≥ 0.7).

[0047] As a preferred option, the dialect assessment agent is used to accurately identify the language and semantic content of the input dialect speech, and to quantitatively score the patient's language ability based on the fluency, vocabulary, and grammatical correctness of the speech, while simultaneously outputting a description of problems such as unclear pronunciation or disordered word order, as follows:

[0048] Dialect speech feature enhancement and recognition: The speech branch layer of the base model (Wav2Vec2.0) is invoked, combined with a customized dialect acoustic dictionary (including dialect phonemes and tone templates), and the dialect speech is classified into languages ​​and converted to text through MFCC feature extraction and dynamic time warping (DTW) algorithms. The dialect-Mandarin cross-language alignment task in the pre-training of the base model is also performed to correct semantic biases in the speech-to-text results. Specifically, MFCC feature extraction involves: constructing a Mel filter bank, increasing the filter density in the dialect feature-dense frequency band, and shifting the center frequency to the dialect-specific frequency band; extracting 12-16 dimensional static MFCC (Mel frequency cepstral coefficients), and combining the first-order difference (ΔMFCC), second-order difference (ΔΔMFCC), and frame energy to form a 36-48 dimensional dynamic feature vector, preserving the dialect tone variation trend.

[0049] Construction of a Language Proficiency Scoring Index System: A multi-dimensional language proficiency scoring index system is constructed, encompassing fluency, vocabulary size, and grammatical correctness. Based on the acoustic representation output from the speech branch layer of the base model, a scoring task head (composed of a fully connected layer and sigmoid activation) is integrated. Training is performed using clinically labeled dialect speech samples (including physician scoring labels). The final quantitative result is output through weighted fusion of scores from each dimension. The total language proficiency score (0-100 points) is calculated by weighting fluency, vocabulary size, and grammatical correctness, using the following formula: ;in, Indicates the smoothness rating; Indicates vocabulary size score; Indicates the score for grammatical correctness; The weighting of the fluency score; The weighting of the vocabulary score; Weights representing syntactic correctness; Smoothness rating ( (0-25 points) Based on the number of pronunciation pauses and speech rate stability, the formula is: ;in, This indicates the number of pauses per 10 seconds of speech (e.g., "eat...food" counts as 1 pause), k represents the pause penalty coefficient, k=2 means 2 points are deducted for each pause; v represents the patient's speech rate (words / minute). This represents the baseline value for normal speech speed in a dialect, taken as 150. Indicates the deviation in speaking speed; 1 point is deducted for every 10 words / minute of deviation, with a minimum of 0 points; vocabulary score ( (0-20 points) Based on the correct recognition rate of core dialect vocabulary, the formula is: ;in, This indicates the number of core vocabulary words tested; The number of words that are correctly pronounced and recognized; grammatical correctness score ( (0-15 points) Based on the grammatical error rate of dialect sentences, the formula is: ;in, Indicates the number of syntax errors; This indicates the number of words in the sentence; the minimum score is 0.

[0050] Autonomous reasoning and output: The reasoning engine has a built-in rule base (such as "speech rate fluctuation > 50% → mark 'speech rate unstable'"), and combines the scores of each dimension output by the language ability scoring index system to autonomously generate a language problem description in the format of "dialect type - recognized text - total language ability score - scores of each dimension - problem description", which is compatible with clinical electronic medical record systems.

[0051] As a preferred approach, the text assessment agent is used to extract psychologically relevant keywords (such as negative emotion words like "insomnia" and "irritability," and positive emotion words like "happy" and "calm") and linguistic feature words (such as absolute words like "always" and "never," and vague words like "seems" and "maybe") from text data such as psychological questionnaires, self-report texts, and medical records completed by patients. Based on the frequency and semantic association of keywords, it outputs a preliminary psychological state level (mild / moderate / severe anxiety / depression, or "no obvious abnormalities"), and simultaneously generates a psychological tendency analysis corresponding to the linguistic features (such as "frequent use of absolute words → poor emotional stability"). Specifically:

[0052] Long text feature extraction and keyword recognition: The text branch layer (Transformer-XL) of the base model is called, and long texts of more than 500 words (such as a patient's thousand-word self-report) are processed through segmented loop mechanism and relative position encoding to capture semantic dependencies across paragraphs; and a keyword recognition module is built using a BiLSTM-CRF model. The keyword recognition module is used to train based on psychological domain dictionary and clinical annotation data to achieve accurate location and classification of emotion words and feature words. The annotation types of clinical annotation data include negative emotion, positive emotion, absolute and fuzzy.

[0053] Psychological state level prediction: Keyword frequency and semantic similarity (based on BERT calculation and similarity with anxiety / depression benchmark words) are used as input features; on the basis of the text representation output by the Transformer-XL text branch layer, a hierarchical task head (fully connected layer + Softmax) is connected, and clinical text samples (including state levels labeled by psychologists) are used for training, combined with Focal Loss to solve the class imbalance problem;

[0054] Autonomous Reasoning and Output: The reasoning engine generates analysis conclusions based on the "feature word-psychological tendency" association rule. For example, if the frequency of 'insomnia' is >5 times and is accompanied by 'irritability', it indicates an anxiety tendency. The analysis conclusions include a list of keywords (including classification and frequency), a preliminary level of psychological state, and a description of the tendency analysis, which supports physicians in quickly locating core psychological clues.

[0055] As a preferred option, the chart assessment agent is used for clinically common psychological assessment charts such as the SAS (Self-Rating Anxiety Scale), SDS (Self-Rating Depression Scale), and SRS (Social Avoidance Scale). It automatically identifies the chart type, item selection status (e.g., "√" or "×"), and scoring scale (e.g., the "1-4" options for the SAS), calculates the total scale score and scores for each dimension, outputs the assessment results based on the score range, and associates targeted psychological intervention suggestions (e.g., "Mild anxiety → relaxation training recommended"). Specifically:

[0056] Visual feature extraction and information recognition of charts: The visual branch layer of the base model (ResNet-50+Transformer) is called to extract the visual features of the chart (such as the position of the option box, tick marks, and handwritten check marks). Combined with OCR technology, the chart title and title text are recognized to complete the chart type classification (accuracy ≥95%). The object detection model (YOLOv8) is used to locate the option box area, and the check status is identified by image segmentation technology (distinguishing between "checked" and "unchecked" based on the difference in pixel gray value).

[0057] Scale scoring calculation and correlation analysis: Construct a scale rule base to store the scoring criteria, total score calculation method and grading threshold of each scale; Combine the chart representation output by the base model to determine the scale type; Based on the identified check status and the corresponding rule base, automatically calculate the total score and scores of each dimension, and correlate the corresponding psychological state conclusions.

[0058] Autonomous reasoning and output: The reasoning engine connects to the clinical intervention knowledge base (including the mapping relationship between "score range and intervention recommendation") and outputs the assessment results. The assessment results include the chart type, the score of each question, the total score, the assessment results and intervention recommendations.

[0059] As a preferred approach, a weighted fusion algorithm is used to integrate the evaluation results of different agents to obtain the overall evaluation score. The formula is as follows:

[0060] ;

[0061] Where n represents the number of agents participating in the collaboration; For intelligent agents The evaluation score (the agent evaluation scores are all standardized to 0-100 points). This represents the weight of agent i in the dialect-specific evaluation, text-specific evaluation, or graph-specific evaluation. It is not a fixed value, but a comprehensive weighted average calculated dynamically based on accuracy, data reliability, and clinical importance, as shown in the following formula:

[0062] ;

[0063] in, The accuracy weight is represented by the formula: ;in, This represents the clinical validation accuracy of dialect-specific assessment agent i, text-specific assessment agent i, or graph-specific assessment agent i. The data reliability weight is based on the reliability of the input data source, and is applied to doctor-annotated texts or professional scales. Patient's self-spoken speech or self-reported text Family members described ; The clinical importance weighting is based on the impact of the assessment results on the treatment plan. Psychological state assessment (charts / text agents): Psychological intervention takes precedence over language training); Language proficiency assessment (dialect intelligence agent): Voice emotion assessment (voice agent): (To assist in verifying psychological state).

[0064] Furthermore, the collaborative scheduling unit also employs a conflict resolution mechanism to calculate the deviation of the agent's evaluation results. The formula is as follows:

[0065] ;

[0066] in, , This represents the evaluation scores of any two agents among the dialect-specific evaluation agent, text-specific evaluation agent, and graph-specific evaluation agent. If the situation is determined to be "conflict", a suggestion to "further evaluation is needed" will be output.

[0067] The multi-agent collaborative language and psychological disorder patient rehabilitation status assessment system of the present invention has the following advantages:

[0068] (I) This invention constructs a rehabilitation status assessment architecture for patients with language and psychological disorders based on multi-agent collaboration. Using a large rehabilitation model as the base model, it constructs assessment agents that are adapted to different dialects and different modal data. Each assessment agent is used to identify assessment data of different dialects, charts, texts, etc., and to perform status assessment based on the identified content, generate assessment conclusions, achieve high adaptation and efficient assessment, meet clinical requirements, and is applicable to rehabilitation departments of hospitals at all levels, community health service centers and children's rehabilitation institutions. It can achieve efficient adaptation assessment of multi-dialect and multi-modal data, and provide accurate basis for clinical development of personalized rehabilitation plans.

[0069] (II) This invention constructs a multi-agent collaborative architecture based on a large rehabilitation model, designs specialized assessment agents that are adapted to different dialects and modalities, and achieves comprehensive assessment of "language + psychology" through a collaborative mechanism, thereby improving the accuracy and efficiency of assessment, meeting the clinical needs for accurate and efficient assessment of patients with language and psychological disorders, and solving the problems of poor dialect adaptation, multimodal fragmentation, single assessment dimension and weak interpretability in existing technologies.

[0070] (III) The dialect assessment intelligent agent of the present invention is optimized for the acoustic and lexical features of different dialects, improves the accuracy of language defect identification for dialect patients, solves the pain point of "inaccurate assessment" in primary hospitals in various dialect areas, can cover more than 80% of language and psychological disorder patients nationwide, significantly improves dialect adaptability, and covers clinical needs in multiple regions.

[0071] (iv) The present invention enables multi-agent collaborative processing of speech, text and chart data to achieve simultaneous evaluation of "language function + psychological state" and shorten the evaluation time; at the same time, it avoids misjudgment of a single modality (such as misjudging the stuttering caused by anxiety as a language defect based solely on speech), and achieves efficient multi-modal collaboration with more comprehensive evaluation dimensions;

[0072] (v) The assessment report of this invention includes intelligent agent assessment basis and clinical correlation suggestions, allowing doctors to trace the scoring logic (such as "why mild anxiety is determined" and "what is the basis for the language defect type"), improving the trust in the existing "black box model", making it easier to use in the formulation of clinical rehabilitation plans, with strong clinical interpretability and high doctor trust. Attached Figure Description

[0073] The invention will be further described below with reference to the accompanying drawings.

[0074] Appendix Figure 1 This is a schematic diagram of the structure of a multi-agent collaborative language and psychological disorder patient rehabilitation status assessment system. Detailed Implementation

[0075] The following detailed description of the language and psychological disorder patient rehabilitation status assessment system based on multi-agent collaboration of the present invention is provided with reference to the accompanying drawings and specific embodiments.

[0076] Example 1: As shown in the attached document Figure 1 As shown, this embodiment provides a rehabilitation status assessment system for patients with language-related psychological disorders based on multi-agent collaboration. The system includes:

[0077] The base model building unit is used to construct a base model based on Transformer-XL (long text processing) and Wav2Vec2.0 (speech processing) architecture, integrating a multimodal attention mechanism to achieve cross-modal information interaction. The base model includes a speech branch layer, a text branch layer, a graph branch layer, and a multimodal fusion layer. The multimodal fusion layer is deployed on top of the speech, text, and graph branch layers and is connected to the cross-modal attention layer, allowing different modal representations, including speech, text, and graphs, to pay attention to each other, learn the correlation weights between each modality, and output a unified multimodal semantic representation.

[0078] The specialized evaluation agent construction unit is used to construct dialect-specific evaluation agents, text-specific evaluation agents, and chart-specific evaluation agents based on the base model. The dialect-specific evaluation agents, text-specific evaluation agents, and chart-specific evaluation agents are used to identify different modalities of input data in different dialects, texts, and charts, and to perform state evaluation based on the identified data, thereby generating dialect evaluation results, text evaluation results, and chart evaluation results.

[0079] The collaborative scheduling unit is used to integrate the dialect evaluation results, text evaluation results, and chart evaluation results generated by the dialect evaluation agent, text evaluation agent, and chart evaluation agent using a weighted fusion algorithm. This enables the fusion of results from different agents and the resolution of conflicts, ensuring the comprehensiveness and consistency of the evaluation.

[0080] In this embodiment, the speech branch layer adopts a Wav2Vec2.0 feature extractor and context encoder structure. The feature extractor is used to transform the original speech waveform into an acoustic feature sequence (such as 768-dimensional) through CNN stacking, and the context encoder (Transformer) is used to optimize the speech representation through contrastive learning.

[0081] The text branching layer is based on the Transformer-XL design and introduces relative position encoding and segment-level recurrence mechanism to handle dependency modeling of long texts (such as patient self-reports and scale descriptions).

[0082] The chart branch uses a CNN and Transformer structure. The CNN is used to extract the visual feature maps of the chart, and the Transformer is used to transform the feature maps into sequential visual representations.

[0083] The base model construction unit in this embodiment includes:

[0084] A multi-source data training set construction module is used to collect and preprocess dialect language data samples, multimodal psychological data samples, and clinical diagnosis and treatment data samples to construct a multi-source data training set. Dialect language data samples refer to dialect language disorder samples. Each dialect language data sample includes dialect speech (16kHz sampling rate, 1-3s per sample), dialect text annotation (corresponding dialect vocabulary and defect type), and Mandarin-mapped text. Multimodal psychological data samples include psychological assessment charts (such as paper / electronic charts and images of the SAS anxiety scale and SRS social response scale), text descriptions (patient self-reports of 'not wanting to speak, afraid of making mistakes,' and doctor's records of 'patients avoiding verbal communication'), and vocal emotional characteristics (such as speech segments with slow speech rate and low tone). Clinical diagnosis and treatment data samples include rehabilitation assessment records of patients with language and psychological disorders (language ability scores and psychological state levels annotated by doctors) and rehabilitation trajectory data (such as records of the correlation between changes in language fluency and changes in psychological state within one month).

[0085] The base model pre-training module is used to complete the dialect-Mandarin cross-linguistic alignment task, the language-psychological association prediction task, and the multimodal graph understanding task. The dialect-Mandarin cross-linguistic alignment task refers to learning the mapping relationship between dialect speech and text and Mandarin (such as mapping "falling rain" to "rain"), and optimizing the cross-linguistic representation ability of dialect features. The language-psychological association prediction task refers to predicting the psychological state level (such as "mild anxiety") based on the patient's speech (such as stuttering speech) and text (such as "nervous"), and modeling the correlation between language features and psychological state. The multimodal graph understanding task refers to extracting key information from psychological assessment graphs through CNN+Transformer to achieve automatic structuring of graph data.

[0086] The base model fine-tuning module is used to design loss functions including language assessment loss (cross-entropy), psychological assessment loss (MSE), and cross-modal consistency loss (cosine similarity). Based on the language-psychological integrated assessment data labeled by doctors, the base model is fine-tuned according to the loss functions to ensure that the assessment results output by the base model conform to the clinical diagnosis and treatment logic.

[0087] In this embodiment, the dialect-Mandarin cross-linguistic alignment task updates the parameters of the speech and text branches through backpropagation, thereby strengthening the semantic mapping relationship between dialects and Mandarin; specifically as follows:

[0088] (1) After encoding the input speech and text data, cross-modal attention alignment is performed. Specifically: Speech encoding: Dialect and Mandarin speech are processed by the Wav2Vec2.0 feature extractor to obtain acoustic feature sequences respectively. and Text encoding: Semantic feature sequences are obtained from dialect and Standard Mandarin texts through the Transformer-XL text branch layer, respectively. and Cross-modal attention alignment: and , and Input the cross-modal attention layer and calculate the alignment weights between modalities using the following formula: ;in, Indicates the alignment weights between modalities; The feature dimension is represented by attention weights, which guide the base model to focus on semantically matched speech-text segments; vt stands for speech (video) and text (text) alignment, an abbreviation; T is the vector transpose symbol.

[0089] (2) Contrast Alignment Loss: Construct positive and negative sample pairs, and optimize the intermodal similarity using InfoNCE loss. The formula is as follows: ;in, Indicates the similarity between modalities; and τ represents the fusion representation of dialect and the fusion representation of Mandarin, respectively; N represents the total number of negative samples participating in the comparison; positive samples in the positive and negative samples are dialect-Mandarin pairs with the same semantics; negative samples in the positive and negative samples refer to randomly matched dialect-Mandarin pairs.

[0090] (3) Speech-Text Matching (STM): This task uses a binary classification method to determine whether the input dialect speech and Mandarin text semantically match. Cross-entropy is used to measure the deviation between the model's binary classification prediction and the actual matching label (match / non-match). This task involves binary classification to determine whether the dialect speech and Mandarin text semantically match. The cross-entropy loss function is used to measure the deviation between the model's binary classification prediction and the actual matching label (match / non-match), guiding the model to optimize parameters to improve the accuracy of the matching judgment.

[0091] The training process for the dialect-Mandarin cross-language alignment task in this embodiment is as follows: ① Initialization: Wav2Vec2.0 uses pre-trained weights to initialize the speech branch, and Transformer-XL uses Chinese BERT to initialize the text branch; ② Training batches: Each batch uses a hybrid contrastive loss and STM task, with a learning rate set to 5e-5, and the AdamW optimizer is used; ③ Iterative optimization: The parameters of the speech branch and the text branch are updated through backpropagation to strengthen the semantic mapping relationship between dialect and Mandarin.

[0092] The language-psychological association prediction task in this embodiment is as follows:

[0093] (1) After encoding the input speech and text data, modality fusion is performed, specifically: speech feature enhancement: speech features are obtained through the Wav2Vec2.0 context encoder. Additional acoustic features, including fundamental frequency (F0), MFCC, and speech rate, are concatenated to enhance psychological cues in the speech. Text feature encoding involves processing patient text through a Transformer-XL text branch layer, employing a segmented loop mechanism for long texts to preserve cross-segment dependencies and extract text features. Multimodal fusion: and By learning the association weights between "voice anomaly features, text emotion words, and psychological states" through a cross-modal attention layer, a fused representation is obtained. ;

[0094] (2) Psychological cue matching task: Construct a “speech / text feature-psychological cue word” matching task (e.g., speech pauses correspond to the “nervous” cue word), requiring the base model to predict the type of psychological cue contained in the input features, and use cross-entropy to measure the deviation between the model’s prediction results and the actual cue type labels; the “speech / text feature-psychological cue word” matching task requires the model to predict the specific type of psychological cue contained in the input speech / text features, so binary cross-entropy is used as the loss function.

[0095] (3) Prediction of mental state levels: in fusion representation The system then uses a classification head to predict psychological states from level 1 to 5, and employs weighted cross-entropy to measure the deviation between the predicted results and the true labels. The classification head includes a fully connected layer and a Softmax layer. Specifically, in this task, a psychological state level prediction model is constructed based on the multimodal fusion representation H_fusion (the fusion representation is then fed into the classification head). To address the data imbalance problem (differences in the number of samples for different psychological state levels), a few categories are assigned higher weights, and then weighted cross-entropy is used to measure the deviation between the predicted results and the true labels.

[0096] The training process for the language-psychological association prediction task in this embodiment is as follows: ① Freeze some parameters and train only the multimodal fusion layer and the classification head (warm-up stage); ② Unfreeze all parameters and train using gradient accumulation (batch size=32, cumulative steps=4), with the learning rate decay strategy being cosine annealing; ③ Verify the model's accuracy in predicting psychological levels on the test set every 1000 steps and save the optimal model.

[0097] The multimodal graph understanding task in this embodiment is as follows:

[0098] (1) Extract shallow visual features (such as edges and textures) from the chart images using a CNN (ResNet-50), and then serialize the feature maps (segmenting them into 16×16 patches) using a Transformer encoder (visual Transformer) to obtain visual representations. ;

[0099] (2) Text-assisted encoding: The text extracted from the chart (such as scale title and question description) is recognized by OCR and then processed by the Transformer-XL text branch layer to obtain the text representation. ;

[0100] (3) Cross-modal attention fusion: integrating visual representations With text representation Cross-modal attention interaction is achieved through a cross-modal attention layer to obtain graph fusion representation. ;

[0101] (4) Chart type classification: Determine whether the scale type is a binary or multi-class task, use the cross-entropy function to measure the deviation between the predicted probability distribution of the output scale type and the true label, and assist the base model to focus on the layout features of different scales.

[0102] The "Chart Type Classification Task" uses cross-entropy as the loss function. Specifically, the model is based on the visual representation of the chart output by the visual branch of the base model and the OCR text representation output by the text branch. After cross-modal attention fusion, these representations are fed into the classification head. The core task is to determine the specific type of the input chart (such as SAS scale, SDS scale, SRS scale, etc.). The cross-entropy loss function is used to measure the deviation between the predicted probability distribution of the scale type output by the model and the true label, helping the model learn the differentiated features of different scales, such as layout and text.

[0103] (5) Entity recognition of charts: The structured extraction of key information from the scale is transformed into a sequence labeling task. The CRF layer is used to optimize the labeled sequence, and the negative log-likelihood loss function of CRF is used to measure the deviation between the labeled sequence predicted by the model and the real labeled sequence.

[0104] The model corresponding to the "chart entity recognition task" uses CRF negative log-likelihood as the loss function. Specifically, this model is designed to extract structured information from psychological assessment charts. It transforms the task into a sequence labeling task. To optimize the rationality of the labeled sequence, a CRF (Conditional Random Field) layer is connected to the model output layer. Therefore, the CRF negative log-likelihood loss function is used to measure the deviation between the labeled sequence predicted by the model and the actual labeled sequence.

[0105] (6) Numerical prediction task: After the fusion representation is connected to the regression head, the numerical information of the total score and individual item scores of the scale is predicted, and the deviation between the predicted score and the actual labeled score is measured by the MSE loss function.

[0106] The “Numerical Prediction Task” uses MSE as the loss function. This task focuses on extracting continuous numerical information from psychological assessment charts and is a regression task. The MSE loss function measures the deviation between the predicted score and the actual labeled score, thereby optimizing the model’s accuracy in extracting numerical information from the chart.

[0107] The training process for the multimodal graph understanding task in this embodiment is as follows: ① The visual branch is initialized with pre-trained ResNet-50 and ViT, and the text branch layer reuses the Transformer-XL parameters trained on the dialect-Mandarin cross-language alignment task; ② Multi-task joint training is adopted, with the learning rate set to 3e-5 and the weight decay to 0.01; ③ A Hard Negative Mining strategy is introduced to increase the weight of training samples for easily confused graph regions (such as similar rating scales).

[0108] The loss function used in the base model fine-tuning module in this embodiment is as follows:

[0109] ;

[0110] in, Represents the loss function; Indicates language assessment loss; Indicates psychological assessment loss; Indicates cross-modal consistency loss; This represents the language evaluation weighting coefficient; This represents the weighting coefficient of the psychological assessment; Indicates the modal alignment weight coefficient; (Language assessment has the highest priority) (Psychological assessment is secondary) (Modal alignment aid), and The learning rate uses cosine annealing (initial annealing). , minimum );

[0111] The language evaluation loss uses cross-entropy. For a multi-classification task targeting language defect types (such as naming aphasia, fluency defects, etc., totaling 6 categories), the formula is as follows:

[0112] ;

[0113] Where B (equal to 32) represents the training batch size; C (equal to 6) represents the number of language defect categories; This represents the one-hot encoding of sample b in category c (e.g., "named aphasia" corresponds to [1,0,0,0,0,0]). This represents the predicted probability of the base model, requiring training to converge. ;

[0114] The psychological assessment loss was calculated using the MSE (Mental State Examination) regression task for mental state scores (0-100 points, such as anxiety level), with the following formula:

[0115] ;

[0116] in, This indicates the doctor's actual psychological score (e.g., 52 points on the SAS scale). This represents the base model's predicted score, requiring convergence. (Corresponding scoring error ≤ 5 points);

[0117] The cross-modal consistency loss uses cosine similarity to ensure alignment of speech and text modal features, as shown in the following formula:

[0118] ;

[0119] in, This represents the speech feature vector of sample b; This represents the corresponding text feature vector; For L2 norm, it is required that (Corresponding cosine similarity ≥ 0.7).

[0120] In this embodiment, the dialect assessment agent is used to accurately identify the language and semantic content of the input dialect speech, and to quantitatively score the patient's language ability based on the fluency, vocabulary, and grammatical correctness of the speech. Simultaneously, it outputs a description of problems such as unclear pronunciation or disordered word order, as detailed below:

[0121] (1) Dialect speech feature enhancement and recognition: The speech branch layer (Wav2Vec2.0) of the base model is called, and a customized dialect acoustic dictionary (including dialect phonemes and tone templates) is combined. Through MFCC feature extraction and dynamic time warping (DTW) algorithm, the dialect speech is classified into languages ​​and converted into text. The dialect-Mandarin cross-language alignment task in the pre-training of the base model is executed to correct the semantic deviation in the speech-to-text result. Among them, MFCC feature extraction is specifically as follows: a Mel filter bank is constructed, the filter density is increased in the frequency band with dense dialect features, and the center frequency is shifted to the dialect-specific frequency band; 12-16 dimensional static MFCC (Mel frequency cepstral coefficients) are extracted, and the first-order difference (ΔMFCC), second-order difference (ΔΔMFCC) and frame energy are combined to form a 36-48 dimensional dynamic feature vector, which preserves the dialect tone change trend.

[0122] (2) Construction of a language proficiency scoring index system: A multi-dimensional language proficiency scoring index system including fluency, vocabulary size, and grammatical correctness is constructed. Based on the acoustic representation output by the speech branch layer of the base model, a scoring task head (composed of a fully connected layer + Sigmoid activation) is connected. Clinically labeled dialect speech samples (including physician scoring labels) are used for training. The final quantitative result is output by weighted fusion of scores from each dimension. Among them, the total language proficiency score (0-100 points) is calculated by weighting three items: fluency, vocabulary size, and grammatical correctness. The formula is as follows: ;in, Indicates the smoothness rating; Indicates vocabulary size score; Indicates the score for grammatical correctness; The weighting of the fluency score; The weighting of the vocabulary score; Weights representing syntactic correctness; (Fluency has the greatest impact on communication) (Vocabulary size is secondary) (Syllabic aid), and Smoothness rating ( (0-25 points) Based on the number of pronunciation pauses and speech rate stability, the formula is: ;in, This indicates the number of pauses per 10 seconds of speech (e.g., "eat...food" counts as 1 pause), k represents the pause penalty coefficient, k=2 means 2 points are deducted for each pause; v represents the patient's speech rate (words / minute). This represents the baseline value for normal speech speed in a dialect, taken as 150. Indicates the deviation in speaking speed; 1 point is deducted for every 10 words / minute of deviation, with a minimum of 0 points; vocabulary score ( (0-20 points) Based on the correct recognition rate of core dialect vocabulary, the formula is: ;in, This indicates the number of core vocabulary words tested; The number of words that are correctly pronounced and recognized; grammatical correctness score ( (0-15 points) Based on the grammatical error rate of dialect sentences, the formula is: ;in, Indicates the number of syntax errors; This indicates the number of words in the sentence; the minimum score is 0.

[0123] (3) Autonomous reasoning and result output: The reasoning engine has a built-in rule base (such as “speech rate fluctuation > 50% → mark 'speech rate unstable') and combines the scores of each dimension output by the language ability scoring index system to autonomously generate a language problem description in the format of “dialect type - recognized text - total language ability score - scores of each dimension - problem description”, which is compatible with the clinical electronic medical record system.

[0124] In this embodiment, the text assessment agent extracts psychologically relevant keywords (such as negative emotion words like "insomnia" and "irritability," and positive emotion words like "happy" and "calm") and linguistic feature words (such as absolute words like "always" and "never," and vague words like "seems" and "maybe") from the text data of the patient's psychological questionnaire, self-report text, and medical records. Based on the frequency and semantic association of keywords, it outputs a preliminary psychological state level (mild / moderate / severe anxiety / depression, or "no obvious abnormalities"), and simultaneously generates a psychological tendency analysis corresponding to the linguistic features (such as "frequent use of absolute words → poor emotional stability"). The details are as follows:

[0125] (1) Long text feature extraction and keyword recognition: The text branch layer (Transformer-XL) of the base model is called. Through the segmented loop mechanism and relative position encoding, long texts of more than 500 words (such as a patient's thousand-word self-report) are processed to capture the semantic dependencies across paragraphs. The BiLSTM-CRF model is used to build a keyword recognition module. The keyword recognition module is used to train based on the psychological domain dictionary and clinical annotation data to achieve accurate positioning and classification of emotional words and feature words. The annotation types of clinical annotation data include negative emotions, positive emotions, absolute and fuzzy.

[0126] (2) Psychological state level prediction: The keyword frequency and semantic similarity (based on BERT calculation and similarity with the anxiety / depression benchmark words) are used as input features; on the basis of the text representation output by the Transformer-XL text branch layer, the hierarchical task head (fully connected layer + Softmax) is connected, and clinical text samples (including the state level marked by psychologists) are used for training, and Focal Loss is combined to solve the class imbalance problem.

[0127] (3) Autonomous reasoning and result output: The reasoning engine generates analysis conclusions based on the "feature word-psychological tendency" association rule. For example, if the frequency of 'insomnia' is >5 times and is accompanied by 'irritability', it indicates an anxiety tendency. The analysis conclusions include a list of keywords (including classification and frequency), a preliminary level of psychological state and a description of the tendency analysis, which supports physicians in quickly locating core psychological clues.

[0128] The chart assessment agent in this embodiment is used for clinically common psychological assessment charts such as the SAS (Self-Rating Anxiety Scale), SDS (Self-Rating Depression Scale), and SRS (Social Avoidance Scale). It automatically identifies the chart type, item selection status (e.g., "√" or "×"), and scoring scale (e.g., the "1-4" options for the SAS), calculates the total scale score and scores for each dimension, outputs the assessment results based on the score range, and associates targeted psychological intervention suggestions (e.g., "Mild anxiety → relaxation training recommended"). Specifically:

[0129] (1) Visual feature extraction and information recognition of charts: The visual branch layer of the base model (ResNet-50+Transformer) is called to extract the visual features of the chart (such as the position of the option box, the tick marks, and the handwritten check marks). The chart title and title text are recognized by combining OCR technology to complete the chart type classification (accuracy ≥ 95%). The object detection model (YOLOv8) is used to locate the option box area, and the check status is identified by image segmentation technology (distinguishing between "checked" and "unchecked" based on the difference in pixel gray value).

[0130] (2) Scale scoring calculation and correlation analysis: Construct a scale rule base to store the scoring criteria, total score calculation method and grading threshold of each scale; Combine the chart representation output by the base model to determine the scale type; Based on the identified check status and the corresponding rule base, automatically calculate the total score and the score of each dimension, and associate the corresponding psychological state conclusion.

[0131] (3) Autonomous reasoning and result output: The reasoning engine is linked to the clinical intervention knowledge base (including the mapping relationship of “score range-intervention suggestion”) and outputs the evaluation results. The evaluation results include the chart type, the score of each question, the total score, the evaluation results and the intervention suggestions.

[0132] In this embodiment, a weighted fusion algorithm is used to integrate the evaluation results of different agents to obtain an overall evaluation score. The formula is as follows:

[0133] ;

[0134] Where n represents the number of agents participating in the collaboration; For intelligent agents The evaluation score (the agent evaluation scores are all standardized to 0-100 points). This represents the weight of agent i in the dialect-specific evaluation, text-specific evaluation, or graph-specific evaluation. It is not a fixed value, but a comprehensive weighted average calculated dynamically based on accuracy, data reliability, and clinical importance, as shown in the following formula:

[0135] ;

[0136] in, The accuracy weight is represented by the formula: ;in, This represents the clinical validation accuracy of dialect-specific assessment agent i, text-specific assessment agent i, or graph-specific assessment agent i. The data reliability weight is based on the reliability of the input data source, and is applied to doctor-annotated texts or professional scales. Patient's self-spoken speech or self-reported text Family members described ; The clinical importance weighting is based on the impact of the assessment results on the treatment plan. Psychological state assessment (charts / text agents): Psychological intervention takes precedence over language training); Language proficiency assessment (dialect intelligence agent): Voice emotion assessment (voice agent): (To assist in verifying psychological state).

[0137] In this embodiment, the collaborative scheduling unit also employs a conflict resolution mechanism to calculate the deviation of the agent's evaluation results. The formula is as follows:

[0138] ;

[0139] in, , This represents the evaluation scores of any two agents among the dialect-specific evaluation agent, text-specific evaluation agent, and graph-specific evaluation agent. If the situation is determined to be "conflict", a suggestion to "further evaluation is needed" will be output.

[0140] Example 2: The goal of this example is to achieve integrated detection of "language ability assessment (such as vocabulary size and fluency) + social anxiety state assessment (such as avoidance of communication and low mood)" through a multi-agent collaborative system, thereby addressing the pain points of institutions such as "difficulty in dialect assessment and lack of evidence for judging the psychological-language association". The specific process is as follows:

[0141] S1. Data preparation and preprocessing, as detailed below:

[0142] S101. Construction of the specialized dataset, as detailed below:

[0143] ① Children's dialect language data: Children's dialect speech, sampling rate 16kHz, single duration 1-2s, each marked with "pronunciation defect type";

[0144] ② Multimodal psychological data: including charts of the Social Anxiety Scale for Children (SASC), texts of observations by therapists, and audio emotional fragments;

[0145] S102. Data preprocessing, as detailed below:

[0146] ①Speech data: Background noise was removed using Audacity (samples with a signal-to-noise ratio ≥25dB ​​were retained), and MFCC features (40 dimensions) were extracted using torchaudio.

[0147] ②Chart data: Use OpenCV to crop the SASC scale area, mark the "check box positions", and generate training samples;

[0148] ③ Data partitioning: Divide the data into training set, validation set, and test set in a 7:2:1 ratio.

[0149] S2. Language and Psychological Rehabilitation Large Model Base Training: Follow the process of "multi-task pre-training → targeted fine-tuning", perform three pre-training tasks: dialect-Mandarin alignment task, language-psychological association prediction task, and chart comprehension task, and then perform targeted fine-tuning and optimization.

[0150] S3. Specific evaluation of the clinical application of intelligent agents, as detailed below:

[0151] S301. Dialect Assessment Agent: Language Ability Calculation. Specifically, by recording children's dialect pronunciations of 20 core words (15 seconds of audio duration), such as "māmā," "chīfàn," and "shuǎwánjù," the agent calculates a language ability score according to a formula, as shown in the example below:

[0152] Smoothness rating ( Number of pauses in a 15-second voice message: ("Mom...Mom", "Eat...dinner", "Play...toys" once each), speech rate v = 120 words / minute Substitute into the formula:

[0153] ;

[0154] Vocabulary score ( 14 out of 20 core vocabulary words were pronounced correctly. Substitute into the formula:

[0155] ;

[0156] Grammar accuracy score ( The three test sentences (e.g., "I play with toys," "Mom eats," "I want water") have an average length of 4 characters and one grammatical error ("I want water" should be "I want to drink water"). Substituting these into the formula:

[0157] ;

[0158] S302, Multimodal Psychological Assessment Agent: Social Anxiety Calculation, details as follows:

[0159] ①Chart-based assessment agent (SASC scale): Upload a screenshot of the SASC scale (20 questions), and the agent will identify and score each question. Substitute into the formula, for example:

[0160] → Mild social anxiety;

[0161] ② Text-based assessment agent: The therapist inputs the observation text "avoids eye contact, refuses to interact with other children, and only speaks to the mother". The agent extracts keywords such as "avoidance" and "refusal to interact" and outputs a psychological state score of 60 points (mild anxiety).

[0162] S4. Multi-agent cooperation and conflict resolution, as detailed below:

[0163] S401, Collaborative Weight Calculation and Comprehensive Scoring, details are as follows:

[0164] (1) Collaboration weight calculation: There are 3 agents participating in the collaboration (dialect agent, graph agent, and text agent), and the weights are calculated as follows:

[0165] ;

[0166] ①Accuracy weighting ( ): Based on the historical evaluation accuracy of the agent ( ):

[0167] ;

[0168] in, Clinical validation accuracy of agent i (e.g., graph agent) =0.95, then ), with a range of [0.5,1].

[0169] (2) Data reliability weight ( Based on the reliability of the input data source, for doctors' annotated text / professional scales Patient's self-reported speech / text Family members described .

[0170] (3) Clinical importance weight ( Impact of assessment results on treatment plan:

[0171] Psychological state assessment (chart / text agent): (Psychological intervention takes precedence over language training);

[0172] Language proficiency assessment (dialect agent): ;

[0173] Voice emotion assessment (voice agent): (To assist in verifying psychological state).

[0174] (4) Comprehensive evaluation score: The evaluation results of different agents are integrated using a "weighted fusion algorithm", and the formula is as follows:

[0175] ;

[0176] Where n is the number of agents participating in the collaboration. Let i be the weight of agent i. It is not a fixed value, but a comprehensive weighted average calculated dynamically based on accuracy, data reliability, and clinical importance.

[0177] S402, Conflict Simulation and Resolution: Assume that the emotional score output by the voice-based psychological agent is 6.1 (severe depression), which deviates from the score of 52.5 (mild anxiety) of the graph-based agent. The system triggered a conflict and prompted the therapist to "add 2 minutes of free activity audio and family description." The therapist then collected further audio (the patient was relatively relaxed and spoke at a rate of 80 words per minute). The family described the patient as "actively speaking at home, but silent in the unfamiliar environment of the facility." After reassessment, the emotional score rose to 45 points, and the deviation from the chart agent decreased to 15%, thus resolving the conflict.

[0178] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A rehabilitation status assessment system for patients with language-related psychological disorders based on multi-agent collaboration, characterized in that, The system includes: The base model building unit is used to construct a base model based on Transformer-XL and Wav2Vec2.0 architecture, integrating a multimodal attention mechanism to achieve cross-modal information interaction. The base model includes a speech branch layer, a text branch layer, a graph branch layer, and a multimodal fusion layer. The multimodal fusion layer is deployed on top of the speech, text, and graph branch layers and is connected to the cross-modal attention layer, allowing different modal representations, including speech, text, and graphs, to pay attention to each other, learn the correlation weights between each modality, and output a unified multimodal semantic representation. The specialized evaluation agent construction unit is used to construct dialect-specific evaluation agents, text-specific evaluation agents, and chart-specific evaluation agents based on the base model. The dialect-specific evaluation agents, text-specific evaluation agents, and chart-specific evaluation agents are used to identify different modalities of input data in different dialects, texts, and charts, and to perform state evaluation based on the identified data, thereby generating dialect evaluation results, text evaluation results, and chart evaluation results. The collaborative scheduling unit is used to integrate the dialect evaluation results, text evaluation results and graph evaluation results generated by the dialect-specific evaluation agent, the text-specific evaluation agent and the graph-specific evaluation agent using a weighted fusion algorithm, so as to realize the fusion of results from different agents and conflict resolution. The base model construction unit includes: The multi-source data training set construction module is used to collect and preprocess dialect language data samples, multimodal psychological data samples, and clinical diagnosis and treatment data samples to construct a multi-source data training set. Among them, the dialect language data samples refer to dialect language disorder samples, and each dialect language data sample includes dialect speech, dialect text annotation, and Mandarin mapped text; the multimodal psychological data samples include psychological assessment charts, text descriptions, and speech emotion features; the clinical diagnosis and treatment data samples include rehabilitation assessment records and rehabilitation trajectory data of patients with language and psychological disorders. The base model pre-training module is used to complete the dialect-Mandarin cross-linguistic alignment task, the language-psychological association prediction task, and the multimodal graph understanding task. The dialect-Mandarin cross-linguistic alignment task refers to learning the mapping relationship between dialect speech and text and Mandarin, and optimizing the cross-linguistic representation ability of dialect features. The language-psychological association prediction task refers to predicting the psychological state level based on the patient's speech and text, and modeling the correlation between language features and psychological state. The multimodal graph understanding task refers to extracting key information from psychological assessment graphs through CNN+Transformer to achieve automatic structuring of graph data. The base model fine-tuning module is used to design loss functions including language assessment loss, psychological assessment loss, and cross-modal consistency loss. Based on the language-psychological integrated assessment data labeled by doctors, the base model is fine-tuned according to the loss functions to ensure that the assessment results output by the base model conform to the clinical diagnosis and treatment logic.

2. The language and psychological disorder rehabilitation status assessment system based on multi-agent collaboration as described in claim 1, characterized in that, The speech branch layer adopts a Wav2Vec2.0 feature extractor and context encoder structure. The feature extractor is used to transform the original speech waveform into an acoustic feature sequence through CNN stacking, and the context encoder is used to optimize the speech representation through contrastive learning. The text branching layer is based on the Transformer-XL design, introducing relative position encoding and segmented loop mechanism to handle dependency modeling of long texts; The chart branch uses a CNN and Transformer structure. The CNN is used to extract the visual feature maps of the chart, and the Transformer is used to transform the feature maps into sequential visual representations.

3. The multi-agent collaborative language and psychological disorder patient rehabilitation status assessment system according to claim 1 or 2, characterized in that, The dialect-Mandarin cross-linguistic alignment task updates the parameters of the speech and text branches through backpropagation, thereby strengthening the semantic mapping relationship between dialects and Mandarin; specifically as follows: After encoding the input speech and text data, cross-modal attention alignment is performed. Specifically: Speech encoding: Dialect and Mandarin speech are processed by the Wav2Vec2.0 feature extractor to obtain acoustic feature sequences respectively. and Text encoding: Semantic feature sequences are obtained from dialect and Standard Mandarin texts through the Transformer-XL text branch layer, respectively. and ; Cross-modal attention alignment: and , and Input the cross-modal attention layer and calculate the alignment weights between modalities using the following formula: ;in, Indicates the alignment weights between modalities; The feature dimension is represented by attention weights, which guide the base model to focus on semantically matched speech-text segments. This indicates that the audio (video) and text (text) are aligned; the first letter is an abbreviation; T is the vector transpose symbol. Contrast Alignment Loss: Construct positive and negative sample pairs, and optimize intermodal similarity using InfoNCE loss, the formula is as follows: ;in, Indicates the similarity between modalities; and Let represent the fusion representation of dialect and the fusion representation of Mandarin, respectively; τ is the temperature coefficient; N represents the total number of negative samples participating in the comparison; positive samples in the positive and negative samples are set as dialect-Mandarin pairs with the same semantics; negative samples in the positive and negative samples refer to randomly matched dialect-Mandarin pairs. Speech-text matching: A binary classification task is used to determine whether the input dialect speech and Mandarin text semantically match. Cross-entropy is used to measure the deviation between the binary classification prediction results of the base model and the actual matching labels. The language-psychological association prediction task is as follows: After encoding the input speech and text data, modality fusion is performed, specifically: Speech feature enhancement: Speech features are obtained through the Wav2Vec2.0 context encoder. Additional acoustic features, including fundamental frequency (F0), MFCC, and speech rate, are concatenated to enhance psychological cues in the speech. Text feature encoding involves processing patient text through a Transformer-XL text branch layer, employing a segmented loop mechanism for long texts to preserve cross-segment dependencies and extract text features. Multimodal fusion: and By learning the association weights of "voice anomaly features - text emotion words - psychological state" through a cross-modal attention layer, a fusion representation is obtained. ; Mental cue matching task: Construct a "voice / text feature - mental cue word" matching task, requiring the base model to predict the type of mental cue contained in the input features. Cross-entropy is used to measure the deviation between the base model's prediction results and the true cue type labels, and weighted cross-entropy is used to measure the deviation between the prediction results and the true labels. Mental state level prediction: in fusion representation The system is then fed into a classification head to predict psychological states from 1 to 5, with the loss function being weighted cross-entropy; the classification head includes a fully connected layer and a Softmax layer. The multimodal graph comprehension task is as follows: The chart image is processed by a CNN to extract shallow visual features, and then the feature maps are serialized by a Transformer encoder to obtain the visual representation. ; Text-assisted encoding: After extracting the text from the chart and recognizing it using OCR, the text representation is obtained through the Transformer-XL text branch layer. ; Cross-modal attention fusion: integrating visual representations With text representation Cross-modal attention interaction is achieved through a cross-modal attention layer to obtain graph fusion representation. ; Chart type classification: Determine whether the scale type is a binary or multi-class task, and use the cross-entropy function to measure the deviation between the output scale type prediction probability distribution and the true label, so as to help the base model focus on the layout features of different scales. Entity recognition in charts: The structured extraction of key information from the scale is transformed into a sequence labeling task. The CRF layer is used to optimize the labeled sequence, and the negative log-likelihood loss function of CRF is used to measure the deviation between the labeled sequence predicted by the base model and the true labeled sequence. Numerical prediction task: After fusing representations, the regression head is connected to predict the numerical information of the total score and individual item scores of the scale. The deviation between the predicted scores and the true labeled scores is measured by the MSE loss function.

4. The language and psychological disorder rehabilitation status assessment system based on multi-agent collaboration as described in claim 1, characterized in that, The loss function used in the base model fine-tuning module is as follows: ; in, Represents the loss function; Indicates language assessment loss; Indicates psychological assessment loss; Indicates cross-modal consistency loss; This represents the language evaluation weighting coefficient; This represents the weighting coefficient of the psychological assessment; Indicates the modal alignment weight coefficient; The learning rate uses cosine annealing; The language evaluation loss uses cross-entropy. For multi-class classification tasks involving language defect types, the formula is as follows: ; Where B represents the training batch size; C represents the number of language defect categories; This represents the one-hot encoding of sample b in category c; This represents the predicted probability of the base model, requiring training to converge. ; The psychological assessment loss was calculated using MSE, a regression task for psychological state ratings, with the following formula: ; in, This indicates the doctor's actual psychological score; This represents the base model's predicted score, requiring convergence. ; The cross-modal consistency loss uses cosine similarity to ensure alignment of speech and text modal features, as shown in the following formula: ; in, This represents the speech feature vector of sample b; This represents the corresponding text feature vector; For L2 norm, it is required that .

5. The language and psychological disorder rehabilitation status assessment system based on multi-agent collaboration as described in claim 1, characterized in that, The dialect assessment agent is used to accurately identify the language and semantic content of the input dialect speech, and to quantitatively score the patient's language ability based on the fluency, vocabulary, and grammatical correctness of the speech. Simultaneously, it outputs a description of problems such as unclear pronunciation or disordered word order, as detailed below: Dialect Speech Feature Enhancement and Recognition: The speech branch layer of the base model is invoked, combined with a customized dialect acoustic dictionary, and MFCC feature extraction and dynamic time warping algorithms are used to complete dialect speech language classification and speech-to-text conversion. The dialect-Mandarin cross-language alignment task from the base model's pre-training is also performed to correct semantic biases in the speech-to-text results. Specifically, MFCC feature extraction involves: constructing a Mel filter bank, increasing filter density in dialect-feature-dense frequency bands, and shifting the center frequency towards dialect-specific frequency bands; extracting 12-16 dimensional static MFCCs, and combining first-order difference, second-order difference, and frame energy to form a 36-48 dimensional dynamic feature vector, preserving the dialect tone variation trend. Construction of a Language Proficiency Scoring Index System: A multi-dimensional language proficiency scoring index system is constructed, encompassing fluency, vocabulary size, and grammatical correctness. Based on the acoustic representation output from the speech branch layer of the base model, a scoring task head is integrated, and training is conducted using clinically labeled dialect speech samples. The final quantitative result is output through weighted fusion of scores from each dimension. The total language proficiency score is calculated by weighting three factors: fluency, vocabulary size, and grammatical correctness, using the following formula: ;in, Indicates the smoothness rating; Indicates vocabulary size score; Indicates the score for grammatical correctness; The weighting of the fluency score; The weighting of the vocabulary score; Weights representing grammatical correctness; Fluency scoring is based on the number of pronunciation pauses and speech rate stability, using the following formula: ;in, This represents the number of pauses per 10 seconds of speech; k represents the pause penalty coefficient, k=2 means 2 points are deducted for each pause; v represents the patient's speaking speed. This represents the baseline value for normal speech speed in a dialect, taken as 150. The score indicates the deviation in speaking speed; 1 point is deducted for every 10 words / minute of deviation, with a minimum of 0 points. Vocabulary score is based on the correct recognition rate of core dialect vocabulary, using the following formula: ;in, This indicates the number of core vocabulary words tested; The number of words correctly pronounced and recognized; the grammatical correctness score is based on the dialect sentence grammatical error rate, using the following formula: ;in, Indicates the number of syntax errors; This indicates the number of words in the sentence; the minimum score is 0. Autonomous Reasoning and Output: The reasoning engine has a built-in rule base and combines the scores of each dimension output by the language ability scoring index system to autonomously generate a language problem description in the format of "dialect type-recognized text-total language ability score-score of each dimension-problem description".

6. The language and psychological disorder rehabilitation status assessment system based on multi-agent collaboration as described in claim 1, characterized in that, The text assessment agent is used to extract psychologically relevant keywords and linguistic features from text data such as patient-completed psychological questionnaires, self-reported texts, and medical records. Based on keyword frequency and semantic association, it outputs a preliminary psychological state level and simultaneously generates a psychological tendency analysis corresponding to the linguistic features; specifically as follows: Long text feature extraction and keyword recognition: The text branch layer of the base model is called, and long texts of more than 500 characters are processed through segmented loop mechanism and relative position encoding to capture semantic dependencies across paragraphs; and a keyword recognition module is built using BiLSTM-CRF model. The keyword recognition module is used to train based on psychological domain dictionary and clinical annotation data to achieve accurate positioning and classification of emotion words and feature words. The annotation types of clinical annotation data include negative emotion, positive emotion, absolute and fuzzy. Psychological state level prediction: The structure uses keyword frequency and semantic similarity as input features; Based on the text representation output by the Transformer-XL text branch layer, a hierarchical task head is connected, and clinical text samples are used for training. FocalLoss is combined to solve the class imbalance problem. Autonomous reasoning and result output: The reasoning engine generates analysis conclusions based on the "feature word-psychological tendency" association rule; The analysis results include a list of keywords, a preliminary level of psychological state, and a description of the tendency analysis.

7. The language and psychological disorder rehabilitation status assessment system based on multi-agent collaboration as described in claim 1, characterized in that, The chart assessment agent is used for commonly used clinical psychological assessment charts such as SAS, SDS, and SRS. It automatically identifies the chart type, item selection status, and scoring scale, and calculates the total scale score and scores for each dimension. Based on the score range, it outputs the assessment results and associates targeted psychological intervention suggestions; as detailed below: Visual feature extraction and information recognition of charts: The visual branch layer of the base model is called to extract the visual features of the charts. The chart title and title text are recognized by combining OCR technology to complete the chart type classification. The object detection model is used to locate the option box area and the check status is identified by image segmentation technology. Scale scoring calculation and correlation analysis: Construct a scale rule base to store the scoring criteria, total score calculation method and grading threshold of each scale; Combine the chart representation output by the base model to determine the scale type; Based on the identified check status and the corresponding rule base, automatically calculate the total score and scores of each dimension, and correlate the corresponding psychological state conclusions. Autonomous reasoning and output: The reasoning engine connects to the clinical intervention knowledge base and outputs the assessment results. The assessment results include the type of chart, the score of each question, the total score, the assessment results and intervention recommendations.

8. The language and psychological disorder rehabilitation status assessment system based on multi-agent collaboration as described in claim 1, characterized in that, A weighted fusion algorithm is used to integrate the evaluation results of different agents to obtain the overall evaluation score. The formula is as follows: ; Where n represents the number of agents participating in the collaboration; Represents intelligent agents Evaluation score; This represents the weight of agent i in the dialect-specific evaluation, text-specific evaluation, or graph-specific evaluation. It is not a fixed value, but a comprehensive weighted average calculated dynamically based on accuracy, data reliability, and clinical importance, as shown in the following formula: ; in, The accuracy weight is represented by the formula: ;in, This represents the clinical validation accuracy of dialect-specific assessment agent i, text-specific assessment agent i, or graph-specific assessment agent i. Indicates the data reliability weight; Indicates the weight of clinical importance.

9. The rehabilitation status assessment system for patients with language and psychological disorders based on multi-agent collaboration according to claim 1 or 8, characterized in that, The collaborative scheduling unit also employs a conflict resolution mechanism to calculate the deviation of the agent's evaluation results. The formula is as follows: ; in, , This represents the evaluation scores of any two agents among the dialect-specific evaluation agent, the text-specific evaluation agent, and the graph-specific evaluation agent. If the condition is determined to be "conflict", a suggestion to "further evaluation is required" will be output.

Citation Information

Patent Citations

  • Speech interactive training system and speech interactive training method

    CN102063903A

  • Multi-modal language barrier screening system based on intelligent elderly assistant

    CN118588286A