Psychological risk dynamic monitoring method and system based on adolescent special large model

By using a large-scale model specifically designed for adolescents, combined with multimodal data and dynamic psychological baselines, this approach addresses the issues of insufficient scale fit and incomplete emotion recognition in existing psychological monitoring systems. It enables accurate identification and personalized monitoring of adolescent psychological risks, improves the authenticity of data collection and the scientific rigor of early warnings, and supports proactive prevention of adolescent psychological health issues.

CN122369956APending Publication Date: 2026-07-10CENTRAL SOUTH UNIVERSITY OF FORESTRY AND TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CENTRAL SOUTH UNIVERSITY OF FORESTRY AND TECHNOLOGY
Filing Date
2026-06-10
Publication Date
2026-07-10

Smart Images

  • Figure CN122369956A_ABST
    Figure CN122369956A_ABST
Patent Text Reader

Abstract

This invention discloses a method and system for dynamic monitoring of psychological risks based on a large-scale model specifically for adolescents. The method specifically includes: using multidimensional psychological feature vectors, multimodal emotion scores, and text emotion tendency scores as individual data, combined with normal emotional fluctuation data of adolescents of the same age as a reference benchmark, to construct an individual's dynamic psychological baseline through a variational autoencoder; determining whether there is a pathological psychological risk and triggering a graded warning by calculating the Euclidean distance between the current multidimensional psychological feature vector and the mean of the dynamic psychological baseline; and using a state machine model, taking response time, operational characteristics, and response patterns as inputs, outputting the probability of perfunctory responses, and dynamically adjusting questionnaire questions and guiding language using a Prompt process. This invention achieves three core objectives: accurate identification of adolescent psychological risks, effective differentiation between normal emotional fluctuations and pathological psychological risks, and personalized monitoring adaptation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the intersection of artificial intelligence and adolescent psychological intervention technology, and in particular to a method and system for dynamic monitoring of psychological risks based on a large-scale model specifically for adolescents. Background Technology

[0002] With the increasingly diversified growth environment of adolescents, multiple factors such as academic competition, peer conflicts, family environment influence, and the impact of the internet environment are intertwined, leading to a year-on-year increase in the incidence of adolescent mental health problems. The hidden and complex nature of psychological risks is also becoming increasingly prominent, posing a serious threat to the physical and mental health of adolescents. Against this backdrop, the need for dynamic monitoring of adolescent mental health risks is becoming increasingly urgent. However, existing psychological monitoring systems are mostly designed for general populations, and when applied to adolescents, they have significant technical defects and adaptation bottlenecks. Moreover, these defects are interconnected and mutually influential, forming a vicious cycle of "insufficient scale fit → distorted data collection → one-sided emotion recognition → incorrect risk assessment → failure of early warning," which seriously restricts the practicality and application effectiveness of the system.

[0003] The core flaw of existing psychological monitoring systems is essentially a serious disconnect between the "general design approach" and the "specific psychological needs of adolescents." The system development process has not fully considered the three core characteristics of adolescent psychological development: "developmental," "hidden," and "specific." These flaws are manifested in five aspects, which together prevent the system from meeting the precise and developmental monitoring needs of adolescent psychological risks and make it difficult to play an effective prevention and control role.

[0004] First, the generalization of assessment scales leads to insufficient monitoring accuracy: Most existing psychological monitoring systems use assessment scales developed based on adult psychological characteristics and needs, failing to specifically adapt them to the psychological characteristics, cognitive levels, and daily life scenarios of adolescents. Their core flaw is the lack of in-depth exploration and quantitative analysis of the unique psychological dimensions of adolescents. They simply use adult assessment indicators (such as workplace stress, social interaction difficulties, and family responsibility pressure) to determine the psychological state of adolescents, failing to accurately capture the core psychological challenges unique to this group, such as academic pressure, peer conflicts, confusion about self-identity, and difficulties adapting to school. The assessment dimensions of generalized scales are severely disconnected from the actual psychological needs of adolescents, resulting in the extracted psychological feature vectors failing to accurately and comprehensively reflect their psychological state, thus leading to serious risk misjudgments. For example, indicators such as "excessive work pressure" and "tense workplace relationships" in general questionnaires fail to address core concerns faced by adolescents, such as academic pressure and test anxiety. Furthermore, the psychological fluctuations experienced by adolescents due to exam failures or heavy workloads lack corresponding assessment dimensions for quantitative monitoring. Additionally, general questionnaires often use binary "yes / no" options, failing to capture subtle differences in adolescents' psychological states. This can easily lead to the misclassification of normal adolescent mood swings as psychological abnormalities or the omission of potential hidden psychological distress, further reducing monitoring accuracy. Moreover, the questions in general questionnaires are often abstract and obscure, not matching the cognitive level and comprehension abilities of adolescents. The large number of questions and lengthy questionnaires can easily trigger resistance and boredom among adolescents, leading to perfunctory answers, random selections, and quick skipping, further reducing the accuracy and effectiveness of the collected data.

[0005] Secondly, the subtlety of adolescents' emotional expression leads to incomplete identification: Adolescents are in puberty, their psychological development is not yet fully mature, their emotional regulation ability is relatively weak, and they are often reserved and introverted in nature, making their emotional expression quite subtle. Their true psychological state is often conveyed through obscure language, subtle behavioral changes, and hidden emotional signals, making it difficult to capture directly. However, existing psychological monitoring systems generally rely on single text or questionnaire data for emotion recognition, lacking the ability to integrate multimodal data and algorithms for mining adolescents' unique emotional expressions. This makes it impossible to effectively capture implicit emotional signals such as micro-expressions, voice frequency variations, and social language, resulting in the omission of a large number of potential psychological risks. The current system's emotion recognition scope is limited to explicit data such as questionnaire answers and straightforward text on social media platforms, without integrating multimodal data such as voice and video, and thus unable to capture adolescents' non-verbal emotional signals. For example, when teenagers answer questionnaires, they may outwardly state "normal mood," but their speech may contain negative emotional characteristics such as slowed speech, sudden drop in tone, and low voice. Their facial expressions may include frowning, shifty eyes, and downturned mouths—micro-expressions that current systems cannot recognize. Furthermore, teenagers often use internet slang like "emo" or "give up" to express negative emotions, and use vague expressions like "okay" or "not bad" to mask their psychological distress. General-purpose models cannot accurately understand these unique semantic expressions, leading to serious biases in emotion recognition and failing to reflect the teenagers' true psychological state. Moreover, existing emotion recognition algorithms are designed for a general purpose and are not specifically optimized for the emotional expression characteristics of teenagers. They cannot effectively distinguish between "temporary emotional fluctuations" and "persistent psychological distress," easily misjudging short-term negative emotions caused by temporary setbacks as long-term psychological problems, or ignoring long-term, hidden psychological distress. Many potential psychological risks fail to be identified in time, missing the optimal early intervention window and ultimately developing into serious psychological disorders.

[0006] Third, the lack of a dynamic baseline adaptation algorithm leads to distorted early warnings: Existing psychological monitoring systems generally use a "fixed threshold + single static baseline" model for early warning judgment. Its core flaw is the failure to incorporate the "physiological emotional fluctuations" of adolescents during puberty into the dynamic baseline construction process. The lack of a dynamic adaptation algorithm tailored to the psychological development characteristics of adolescents causes the system to confuse normal emotional changes during adolescence with pathological psychological problems, leading to two major problems: "distorted early warnings" and "alarm fatigue," severely impacting the efficiency and quality of psychological prevention and control work. On the one hand, adolescents are in puberty, experiencing an imbalance between physiological and psychological development, and their emotional fluctuations are more pronounced than in adults. These fluctuations are normal developmental changes, not pathological psychological problems. However, the fixed early warning thresholds of existing systems do not consider this characteristic, easily misinterpreting these normal emotional fluctuations as psychological risk signals and triggering numerous invalid early warnings. On the other hand, a single static baseline cannot capture the gradual deterioration of adolescents' psychological state, resulting in a significant problem of delayed early warnings. For example, the progression from mild mood swings to depressive tendencies in adolescents is gradual and slow. However, static baselines cannot be updated in real time to reflect these psychological changes. Warnings are only triggered when the adolescent's psychological problems become fully apparent and severe, leading to missed opportunities for early intervention and hindering the timely containment of further developments. Furthermore, existing systems lack a scientifically designed baseline update mechanism, failing to adjust baseline parameters in real time based on changes in the adolescent's developmental stage and life circumstances (such as exam season, summer / winter breaks, or family emergencies), further exacerbating the problem of inaccurate warnings.

[0007] Fourth, the rigid interactive data collection mechanism leads to insufficient data authenticity: Existing psychological monitoring systems primarily rely on fixed questionnaires and standardized interactive processes for data collection. Their core flaw is the lack of real-time monitoring and adaptive adjustment mechanisms for adolescents' responses. This makes it impossible to effectively identify invalid responses such as "perfunctory answers," "random filling," and "malicious selection," resulting in low data authenticity and validity. Data authenticity is the foundation for accurate psychological risk assessment; false or invalid data directly leads to incorrect risk assessments, rendering the system useless. Adolescents have limited patience and strong resistance. Faced with lengthy, obscure, and tedious text-based questionnaires, they are prone to perfunctory responses, exhibiting invalid behaviors such as rapid answering, random selection, repeatedly selecting the same option, and skipping subjective questions. However, existing systems only collect the final response results, failing to capture behavioral characteristics during the response process or adjust questionnaire content and guidance based on responses, resulting in poor data collection effectiveness. Furthermore, the existing system's interaction methods are relatively simple, mostly consisting of pure text questionnaires, which do not meet the interactive preferences of teenagers for fun and lightweight interaction, further exacerbating their resistance. At the same time, the system does not consider the individual differences of teenagers, adopting a uniform interaction process and questionnaire settings, which cannot adapt to the personalized needs of teenagers of different ages and personalities. As a result, some teenagers refuse to participate in the monitoring or give perfunctory answers due to poor interactive experience, further reducing the authenticity and effectiveness of data collection. Summary of the Invention

[0008] The purpose of this invention is to provide a method and system for dynamic monitoring of psychological risks based on a large-scale model specifically for adolescents, achieving three core objectives: accurate identification of adolescent psychological risks, effective differentiation between normal emotional fluctuations and pathological psychological risks, and personalized monitoring and adaptation. Ultimately, it aims to complete the entire process of building a dedicated assessment system, multimodal emotion recognition, and interactive adaptive adjustment, thereby solving at least one of the aforementioned problems in the prior art.

[0009] Firstly, the present invention provides a method for dynamic monitoring of psychological risk based on a large-scale model specifically designed for adolescents, the method specifically comprising: A general large language model was pre-trained using a general psychological corpus and then fine-tuned using a dedicated labeled dataset to form a large model specifically for adolescents. Based on the MHT scale and PHQ-9 scale, the indicators were standardized and modified, and multi-dimensional psychological feature vectors were extracted from the response data through a large model specifically for adolescents. Collect multimodal data streams, use a large model specifically designed for teenagers to integrate and identify micro-expressions and speech to output multimodal emotion scores, and perform semantic mining and high-risk keyword identification on text data to output text emotion tendency scores. Using multidimensional psychological feature vectors, multimodal emotion scores, and text emotion tendency scores as individual data, combined with normal emotion fluctuation data of adolescents of the same age as a reference benchmark, a dynamic psychological baseline for individuals is constructed through a variational autoencoder. By calculating the Euclidean distance between the current multidimensional psychological feature vector and the mean of the dynamic psychological baseline, it is determined whether there is a pathological psychological risk and a graded warning is triggered. Based on the state machine model, the system takes the response time, operational characteristics, and response patterns as inputs and outputs the probability of perfunctory responses. It also dynamically adjusts the questionnaire questions and guiding language in conjunction with the Prompt project.

[0010] Secondly, the present invention provides a dynamic monitoring system for psychological risk based on a large-scale model specifically designed for adolescents, the system specifically comprising: The model training module is used to pre-train a general large language model with a general psychological corpus, and then fine-tune it with a dedicated labeled dataset to form a large model specifically for teenagers. The feature extraction module is used to standardize the indicators based on the MHT scale and PHQ-9 scale, and extract multi-dimensional psychological feature vectors from the response data through a large model specifically for adolescents. The scoring module is used to collect multimodal data streams, and through a large model specifically for teenagers, it integrates the recognition of micro-expressions and speech to output multimodal emotion scores. It also performs semantic mining and high-risk keyword recognition on text data to output text emotion tendency scores. The baseline construction module is used to construct an individual's dynamic psychological baseline by using multidimensional psychological feature vectors, multimodal emotion scores, and text emotion tendency scores as individual data, combined with normal emotion fluctuation data of adolescents of the same age as a reference benchmark, through a variational autoencoder. The risk warning module is used to determine whether there is a pathological psychological risk and trigger a graded warning by calculating the Euclidean distance between the current multidimensional psychological feature vector and the mean of the dynamic psychological baseline. The perfunctory response assessment module is based on a state machine model. It takes response time, operational characteristics, and response patterns as inputs, outputs the probability of perfunctory responses, and dynamically adjusts questionnaire questions and guiding language in conjunction with the Prompt project.

[0011] Thirdly, the present invention provides a computer device, including: a memory and a processor, and a computer program stored in the memory, wherein when the computer program is executed on the processor, it implements the dynamic monitoring method for psychological risk based on a large model specifically for adolescents as described in any of the above methods.

[0012] Fourthly, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the method for dynamic monitoring of psychological risk based on a large model specifically for adolescents as described in any of the above methods.

[0013] Compared with the prior art, the present invention has at least one of the following technical effects: 1. This invention focuses on the psychological characteristics and emotional expression patterns of adolescents during their growth and development stages. It addresses the issues of insufficient monitoring accuracy caused by the generalization of existing universal scale systems and the concealment of emotional expression. It is applicable to scenarios where adolescents are concentrated, such as primary and secondary schools and universities. The core objectives are to accurately identify adolescent psychological risks, effectively distinguish between normal emotional fluctuations and pathological psychological risks, and personalize monitoring and adaptation. Ultimately, it completes the full-process technical implementation of the construction of a dedicated assessment system, multimodal emotion recognition, and interactive adaptive adjustment, providing intelligent support for adolescent psychological prevention and control. 2. This invention is an assessment system for adolescents that is optimized and constructed based on the MHT scale and the PHQ-9 scale. It extracts 128 core psychological feature vectors, which completely solves the problem of poor adaptability of general scales, making psychological monitoring more in line with the actual psychological needs of adolescents, and providing accurate and reliable feature support for the accurate judgment of psychological risks. 3. This invention integrates micro-expression, voice frequency conversion and text semantic mining into a multimodal dynamic emotion recognition algorithm, which breaks through the limitations of single data recognition. It can accurately capture the hidden emotional expression signals of adolescents, effectively solve the problem of "incomplete recognition due to the hidden emotional expression of adolescents", effectively avoid the omission of potential psychological risks, and provide timely basis for early intervention. 4. This invention incorporates a dynamic psychological baseline of physiological emotional fluctuations in adolescents, combined with a strict early warning triggering logic, which can accurately distinguish between normal emotional fluctuations and pathological psychological problems during adolescence. It effectively avoids the problems of early warning distortion and alarm fatigue, improves the scientificity and accuracy of early warning of psychological risks in adolescents, and can accurately identify pathological psychological risks, thus gaining valuable time for early intervention. 5. This invention effectively identifies and improves the perfunctory answering behavior of teenagers through real-time monitoring of answering behavior and adaptive dynamic adjustment mechanism of questions and words. It completely solves the core pain point of insufficient data collection authenticity in existing systems, provides high-quality and effective data support for psychological risk assessment, and significantly improves the value and reliability of monitoring data. 6. This invention uses transfer learning technology to perform secondary fine-tuning of a large language model specifically for teenagers, which significantly optimizes the model's logical reasoning ability when dealing with teenager-specific semantics and internet slang, solves the adaptation defects of general models, improves the model's accuracy in extracting and understanding teenagers' psychological information, and effectively solves the problem of general models' insufficient understanding of teenagers' implicit expressions. 7. The core technologies of this invention (dedicated scale optimization, multimodal emotion recognition, developmental baseline adaptation, and interactive adaptive adjustment) are all specifically designed for adolescents. They can be flexibly adapted to different stages of adolescent psychological risk monitoring scenarios, such as primary school, junior high school, senior high school, and university, without the need for large-scale modifications. They have broad application prospects and provide scientific and intelligent technical support for adolescent psychological risk prevention and control, effectively promoting the transformation of adolescent psychological intervention from "passive response" to "proactive prevention". Attached Figure Description

[0014] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0015] Figure 1 This is a flowchart illustrating a method for dynamic monitoring of psychological risk based on a large-scale model specifically for adolescents, provided by an embodiment of the present invention. Figure 2 This is a schematic diagram of the structure of a dynamic monitoring system for psychological risks based on a large-scale model specifically for adolescents, provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of the structure of a computer device provided in an embodiment of the present invention. Detailed Implementation

[0016] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.

[0017] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.

[0018] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0019] As used in this application specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if detected [the described condition or event]" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once detected [the described condition or event]," or "in response to detection [the described condition or event]."

[0020] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0021] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0022] In this application embodiment, the entity executing the process includes a terminal device. This terminal device includes, but is not limited to, devices capable of executing the methods disclosed in this application, such as servers, computers, smartphones, and tablets. Figure 1 A flowchart illustrating a dynamic monitoring method for psychological risk based on a large-scale model specifically for adolescents, as disclosed in an embodiment of the present invention, is shown below in detail: S101 uses a general psychological corpus to pre-train a general large language model, and then uses a dedicated labeled dataset for secondary fine-tuning to form a large model specifically for teenagers. S102, based on the MHT scale and PHQ-9 scale, standardizes the indicators and extracts multi-dimensional psychological feature vectors from the response data through a large model specifically for adolescents. S103 collects multimodal data streams, integrates micro-expression recognition and speech recognition through a large model specifically for teenagers to output multimodal emotion scores, and performs semantic mining and high-risk keyword recognition on text data to output text emotion tendency scores. S104 uses multidimensional psychological feature vectors, multimodal emotion scores, and text emotion tendency scores as individual data, combined with normal emotion fluctuation data of adolescents of the same age as a reference benchmark, and constructs the dynamic psychological baseline of individuals through variational autoencoders. S105, by calculating the Euclidean distance between the current multidimensional psychological feature vector and the mean of the dynamic psychological baseline, determines whether there is a pathological psychological risk and triggers a graded warning; S106, based on a state machine model, takes the response time, operational characteristics, and response patterns as inputs, outputs the probability of perfunctory responses, and dynamically adjusts questionnaire questions and guiding remarks in conjunction with the Prompt project.

[0023] In this embodiment, the application of current artificial intelligence technology in group psychological risk monitoring has been gradually promoted, and various general psychological monitoring systems have achieved certain application results in adult groups. However, there are obvious technical gaps and adaptation shortcomings in the customized optimization for adolescent groups. Existing general systems do not fully consider the special characteristics of adolescent psychological development during puberty. Problems such as generalized scales, single emotion recognition, and rigid early warning judgment are particularly prominent, which cannot meet the needs of precise and developmental monitoring of adolescent psychological risks.

[0024] Specifically, firstly, a general psychological corpus covering a wide range of psychological topics is collected. This corpus contains text data from various common psychological scenarios and is used to pre-train a general-purpose large-scale language model. Through large-scale input of this general-purpose psychological corpus, the general-purpose large-scale language model learns rich language patterns and basic psychological knowledge. Next, a dedicated labeled dataset is constructed. This dataset is specifically designed for the psychological characteristics, cognitive levels, and daily life scenarios of adolescents. The dataset contains a large number of real expression samples of adolescents facing core psychological distress such as academic pressure, peer conflict, confusion about self-identity, and difficulties adapting to school. These samples are detailed and labeled, clarifying their corresponding psychological dimensions and emotional states. This dedicated labeled dataset is used to fine-tune the pre-trained general-purpose large-scale language model, enabling the model to deeply understand the unique psychological dimensions of adolescents, accurately capture the psychological state of the adolescent group, and ultimately form a dedicated large-scale model suitable for monitoring adolescent psychological risks.

[0025] Based on the widely used MHT and PHQ-9 scales, their indicators were standardized and modified. The MHT scale is mainly used to assess anxiety in adolescents, while the PHQ-9 scale focuses on assessing depression. Standardization made the scale indicators more relevant to the actual situation of adolescents. When adolescents completed the responses, the data was input into a large-scale model specifically designed for them. Leveraging its powerful language understanding and analysis capabilities, this model deeply mines various psychological information related to adolescents from the response data, extracting multi-dimensional psychological feature vectors. These feature vectors comprehensively and accurately reflect the psychological state of adolescents in multiple aspects, including academic performance, peer relationships, and self-awareness, avoiding the insufficient monitoring accuracy caused by the lack of exploration of adolescent-specific psychological dimensions in general-purpose scales.

[0026] During monitoring in daily learning and life scenarios, online communication scenarios, and specialized psychological assessment scenarios (such as psychological interviews), multimodal data streams of adolescents are collected, including voice, video, and text data. The collected multimodal data is input into a large-scale model specifically designed for adolescents. Utilizing the model's multimodal fusion processing capabilities, micro-expressions in voice and video are fused and recognized. By analyzing features such as speech rate, tone, and intonation in voice, and micro-expressions such as facial expressions and eye movements in video, a comprehensive multimodal emotion score is output. Simultaneously, for text data, the large-scale model performs semantic mining to understand the true intentions and emotional tendencies expressed by adolescents in the text, and identifies high-risk keywords, such as words with serious psychological risk implications like "despair" and "life is meaningless," outputting a text emotion tendency score. This multimodal emotion recognition method can comprehensively capture the true psychological state of adolescents conveyed through subtle language, nuanced behaviors, and hidden emotional signals, effectively solving the problem of incomplete emotion recognition caused by existing systems relying on single text or questionnaire data.

[0027] Using multidimensional psychological feature vectors, multimodal emotion scores, and textual emotion tendency scores as individual data, this study comprehensively reflects the current psychological state of adolescents. Simultaneously, it collects emotional fluctuation data of adolescents of the same age during their normal development, using this as a reference benchmark. A variational autoencoder is then used to fuse the individual data with the normal emotional fluctuation data of adolescents of the same age. The variational autoencoder learns the latent distribution characteristics of the data and constructs a personalized dynamic psychological baseline for each adolescent based on individual differences and normal developmental patterns. This dynamic psychological baseline is updated in real time as the adolescent's developmental stage and life scenarios change (such as exam season, summer and winter vacations, family changes, etc.), accurately reflecting the range of normal psychological states of adolescents at different periods, avoiding the warning lag and distortion problems caused by existing systems using a single static baseline.

[0028] After establishing an individual's dynamic psychological baseline, the system acquires the adolescent's current multidimensional psychological feature vector in real time. By calculating the Euclidean distance between the current multidimensional psychological feature vector and the mean of the dynamic psychological baseline, the system assesses the degree of deviation between the adolescent's current psychological state and the normal state. If the Euclidean distance exceeds a pre-set threshold, it indicates a potential abnormality in the adolescent's psychological state, requiring further analysis to determine if it represents a pathological psychological risk. Based on the degree of deviation, the psychological risk is categorized into different levels, such as mild, moderate, and severe. When a pathological psychological risk is confirmed, the system triggers a tiered early warning mechanism, sending different levels of warning information to relevant personnel (such as parents, teachers, and psychological counselors) according to the risk level, enabling timely intervention to prevent further deterioration of the psychological problem.

[0029] Based on a state machine model, this system uses the duration of adolescents' responses, operational characteristics (such as click frequency and pause time), and answering patterns (such as repeatedly selecting the same option or quickly skipping questions) as input information. The state machine model analyzes and judges this input information to output the probability of adolescents giving perfunctory answers. When the probability of perfunctory answers exceeds a certain threshold, the system dynamically adjusts the questionnaire questions and guiding language using Prompt engineering techniques. For example, if it is found that an adolescent repeatedly and quickly selects the same option, the system can provide a prompt through Prompt engineering to guide the adolescent to think carefully before answering; or, based on the adolescents' responses, the difficulty and type of subsequent questions can be adjusted to better match the adolescents' interests and cognitive levels, thereby improving the adolescents' enthusiasm for answering and the authenticity and effectiveness of data collection. This solves the problem of insufficient data authenticity caused by the rigid interactive collection mechanism in existing systems.

[0030] This invention addresses the aforementioned industry pain points by optimizing a large-scale model specifically for adolescents, designing a multimodal dynamic emotion recognition algorithm, and constructing a "developmental maladaptive disorder" classification model. This effectively overcomes the accuracy bottlenecks and adaptability deficiencies of existing technologies, providing a new, feasible, and scalable technical solution and implementation path. This technology can be directly applied to the daily psychological prevention and control work in primary and secondary schools and universities, enabling routine monitoring of large-scale adolescent groups. It can also be flexibly adapted to various scenarios such as adolescent mental health service institutions and community mental health prevention and control centers, achieving an organic combination of large-scale group monitoring and personalized precision monitoring. This promotes the transformation of adolescent psychological intervention work from "passive response" to "proactive prevention," improving the efficiency and quality of psychological prevention and control work.

[0031] This invention addresses the four core deficiencies of existing technologies, focusing on the three core characteristics of adolescent psychological development: "developmental," "hidden," and "unique." It combines these with the practical needs of campus psychological prevention and control, constructing a dynamic monitoring system for psychological risks based on a large-scale model optimized specifically for adolescents. The core innovations revolve around four core technologies: "scale adaptation optimization and transfer learning mechanism," "multimodal dynamic emotion recognition algorithm," "developmental maladaptive classification model," and "interactive adaptive adjustment mechanism." Ultimately, this invention achieves three core objectives: accurate identification of adolescent psychological risks, effective differentiation between normal emotional fluctuations and pathological risks, and accurate collection of monitoring data.

[0032] In some embodiments, step S101 above, which involves pre-training the general large language model with a general psychological corpus and then fine-tuning it using a dedicated labeled dataset to form a large model specifically for adolescents, specifically includes: A general psychological corpus containing popular science knowledge, clinical psychological cases, and psychological questionnaire data was collected and constructed. Based on the general psychological corpus, a large language model using the Transformer architecture was pre-trained to obtain the pre-trained large language model. We collected multi-source data, including adolescent psychological questionnaire responses, adolescent social language data, adolescent emotional expression text data, and adolescent speech-to-text data. We standardized and labeled the multi-source data and ensured the consistency of the labeling to form a dedicated labeled dataset for adolescents covering different age groups. We used a youth-specific labeled dataset to fine-tune the pre-trained large language model. During the fine-tuning process, we adjusted the learning rate and introduced regularization constraints to optimize the mapping weights for youth-specific semantics, thus obtaining a youth-specific large model.

[0033] In this embodiment, the present invention breaks through the limitations of existing systems that are "generalized scales and lack dedicated model optimization", constructs a dedicated psychological assessment system that fits the psychological needs of adolescents, and at the same time, it uses transfer learning technology to fine-tune the general language model, so as to achieve accurate understanding of the semantic context of adolescents and effective extraction of psychological characteristics, completely solve the problems of insufficient scale fit and weak model reasoning ability, and significantly improve the accuracy and effectiveness of psychological monitoring.

[0034] To address the shortcomings of general-purpose language models in handling semantic and logical reasoning specific to adolescents, this invention employs transfer learning technology to adapt and optimize general-purpose language models for adolescent-specific psychological monitoring scenarios. Through parameter optimization in two stages—pre-training and secondary fine-tuning—the model significantly improves its accuracy in extracting adolescent psychological characteristics and emotional tendencies. The technology has a clear implementation process, well-defined parameter settings, and strong feasibility for practical application.

[0035] Pre-training foundational stage: Based on a general psychological corpus of over 500,000 data points, the large language model is pre-trained to ensure it possesses basic psychosemantic understanding and emotion recognition capabilities. The general psychological corpus covers various types of data, including popular science knowledge in psychology, clinical psychological cases, various psychological questionnaires, and emotional expression texts, ensuring the comprehensiveness and diversity of the data. Pre-training employs a Transformer architecture (such as BERT-base), with the model's hidden layer dimension set to 768, the number of attention heads to 12, the training batch size to 32, the learning rate to 0.0001, and the number of iterations to 1000. The cross-entropy loss function is used to optimize the model parameters, ensuring the model can accurately understand general psychosemantics and emotional expressions.

[0036] Secondary Fine-tuning Phase: 180,000 sets of adolescent-specific labeled data were collected to fine-tune the pre-trained large language model. The focus was on optimizing the model's processing parameters for adolescent-specific semantics to improve its adaptability. The adolescent-specific labeled dataset underwent rigorous screening and standardization, specifically including 80,000 sets of adolescent psychological questionnaire responses, 50,000 sets of adolescent social language data, 30,000 sets of adolescent emotional expression text data (label consistency Kappa ≥ 0.85, ensuring label accuracy), and 20,000 sets of adolescent speech-to-text and micro-expression annotation data. The dataset covers different age groups from elementary to high school, encompassing various situations such as normal psychological states, mild psychological distress, and serious psychological problems, ensuring the model's universality and relevance after fine-tuning.

[0037] The parameter optimization scheme for the second fine-tuning is as follows: the training batch size remains unchanged at 32, the learning rate is adjusted to 0.00005, the number of iterations is increased to 1500, and L2 regularization (regularization coefficient of 0.001) is added to further avoid model overfitting. The focus is on optimizing the mapping weights for adolescent-specific semantics, enabling the model to accurately interpret the true meaning of commonly used internet slang among teenagers such as "emo" and "badass," and accurately capture the subtle emotional expressions of teenagers. To ensure that the transfer learning process balances classification accuracy and parameter stability, the cross-entropy loss function is used in the pre-training stage, and the L2 regularization term is introduced for joint optimization in the second fine-tuning stage.

[0038] The loss functions for the pre-training and secondary fine-tuning stages are expressed as follows: ; ; in, This represents the cross-entropy loss during the pre-training phase. This represents the joint loss during the second fine-tuning phase. This represents the true label of sample i in category c. The model predicts the class probability, where N represents the number of samples and C represents the number of classes. This represents the model parameters to be optimized. This represents the L2 regularization coefficient.

[0039] In some embodiments, step S102 above, which involves standardizing the indicators based on the MHT scale and the PHQ-9 scale, and extracting multidimensional psychological feature vectors from the response data using a large-scale model specifically designed for adolescents, specifically includes: Based on the MHT scale and PHQ-9 scale, general indicators with low correlation to adolescent psychological distress were removed, and core dimensions specific to adolescents were added to form an adolescent-specific scale. These core dimensions include academic pressure, peer relationships, self-identity, and family environment. The response data generated by the adolescent-specific scale is input into the adolescent-specific big model. Through in-depth semantic analysis and psychological feature mapping, multi-dimensional psychological feature vectors are extracted. Based on the Z-score standardization formula, the multidimensional psychological feature vector is standardized.

[0040] In this embodiment, based on the MHT (Mental Health Assessment Scale for Primary and Secondary School Students) and PHQ-9 (Health Questionnaire), and combined with the core characteristics of adolescent psychological development and actual psychological needs, the existing scales are standardized and optimized. General assessment indicators that are not suitable for adolescents are removed, and four core dimensions specific to adolescents are added: academic pressure, peer relationships, self-identity, and family environment. At the same time, the question wording and option design are optimized to match the cognitive level and comprehension ability of adolescents, effectively improving the adolescents' answering experience and the authenticity of the data.

[0041] Specifically, the optimization and modification of the MHT and PHQ-9 scales follows three principles: "adaptability, relevance, and simplicity." First, it retains four subscales from the MHT scale that are highly relevant to the psychological needs of adolescents: learning anxiety, social anxiety, loneliness tendency, and impulsivity tendency. Subscales such as allergy tendency and physical symptoms, which are less correlated with the core psychological distress of adolescents, are removed. Second, the abstract expression of the PHQ-9 scale is simplified and optimized, transforming obscure professional terminology into easily understandable language for adolescents, reducing the difficulty of understanding the questions. The four newly added exclusive dimensions each have four sub-assessment indicators, forming a complete assessment system specifically for adolescents: the academic pressure dimension covers sub-indicators such as test anxiety, homework load, academic competition pressure, and pressure to enter higher education; the peer relationship dimension covers sub-indicators such as peer conflict, school bullying, social avoidance, and the need for peer approval; the self-identity dimension covers sub-indicators such as self-acceptance, low self-esteem, self-worth perception, and developmental confusion; and the family environment dimension covers sub-indicators such as parent-child communication, family atmosphere, parental expectations and pressure, and family support.

[0042] The optimized adolescent-specific scale comprises 8 subscales and 32 questions, reducing the number of questions by 40% compared to the general scale, effectively avoiding resistance from adolescents caused by lengthy questionnaires. The option design uses a five-category approach: "completely disagree - somewhat disagree - neutral - somewhat agree - completely agree," replacing the binary options of the general scale. This accurately captures subtle differences in adolescents' psychological states and more realistically reflects their psychological changes. Furthermore, the question wording has been differentiated to address the cognitive differences among adolescents of different age groups: For elementary school students, concrete and relatable wording is used, incorporating everyday learning and life scenarios; for middle and high school students, the wording aligns with their developmental stage, balancing scientific accuracy with ease of understanding, ensuring that adolescents of all ages can accurately understand the meaning of the questions and answer truthfully.

[0043] In terms of psychological feature extraction, a large-scale model specifically designed for adolescents was used to conduct in-depth analysis of the scale response data, extracting 128-dimensional psychological feature vectors, each corresponding to a subdivided psychological indicator. To eliminate differences in the dimensions of different indicators and improve the comparability of feature vectors, the Z-score standardization method was used to standardize each feature dimension, ensuring it meets the statistical characteristics of zero mean and unit variance, thereby providing accurate and reliable feature support for subsequent psychological risk assessment.

[0044] The Z-score standardization formula is as follows: ; in, Indicates the first The original feature values ​​of each psychological dimension This represents the mean of that dimension in the training samples. This represents the standard deviation of that dimension. This represents the standardized feature value.

[0045] In some embodiments, step S103 above, which involves collecting multimodal data streams, using a large-scale model specifically designed for adolescents, fusing and recognizing micro-expressions and speech to output multimodal emotion scores, and performing semantic mining and high-risk keyword identification on text data to output text emotion tendency scores, specifically includes: Simultaneously access text data streams, voice data streams, and video data streams, and perform data cleaning and format standardization to form a multimodal data stream; Based on multimodal data streams, micro-expression feature vectors and speech feature vectors are extracted and fused into a large model specifically for teenagers for processing, and multimodal emotion scores are output. Based on multimodal data streams, a large-scale model specifically designed for teenagers is used to semantically map popular online slang and vague expressions among teenagers in text data. Preset high-risk keywords are identified and assigned corresponding weights according to risk levels to obtain text sentiment scores.

[0046] In this embodiment, the present invention overcomes the limitations of existing systems that "recognize emotions in a single way and cannot capture hidden expressions," optimizes the multimodal data stream access and feature fusion capabilities of the risk observer agent, designs a dedicated multimodal dynamic emotion recognition algorithm, and integrates three major technologies: facial micro-expression, voice frequency conversion, and text semantic mining. This algorithm quantifies and scores the emotional state of adolescents, completely solving the problem of incomplete and inaccurate recognition caused by the hidden emotional expressions of adolescents, and improving the comprehensiveness and accuracy of emotion recognition.

[0047] Risk Observer has added two new data stream access channels: video and audio. Together with the existing text data stream, they form a multimodal data collection system of "text + audio + video" to comprehensively collect emotional signals from teenagers. All data streams are accessed through asynchronous interfaces and are encrypted and desensitized using SHA-256, strictly complying with the relevant requirements of the "Law on the Protection of Minors" regarding the protection of minors' personal information. At the same time, data cleaning rules have been optimized to effectively remove invalid data and ensure the quality and security of the collected data.

[0048] The access and processing flow of multimodal data streams has been specifically optimized to ensure the efficiency and security of data collection: First, a RESTful API multi-channel asynchronous interface has been built to support high-concurrency data access, with access latency controlled to ≤100ms, ensuring real-time transmission of data streams. Second, the format requirements for various data streams have been clearly defined. Text data supports multi-platform access (using UTF-8 encoding), ensuring that text data from different sources can be parsed normally; voice data uses WAV format with a sampling rate of 16kHz, collected every 5 minutes, ensuring the capture of emotional changes in speech; video data uses 720P MP4 format, collected every 10 minutes, focusing on capturing facial micro-expression changes. Third, a strict data anonymization and security management mechanism has been established. Data anonymization employs methods such as facial blurring, deletion of original voice data after voice feature extraction, and encrypted mapping of personal information. Simultaneously, a comprehensive access control and access tracking mechanism has been established to ensure that the personal information of minors is not leaked and to guarantee data security.

[0049] The data cleaning process is specifically optimized for the characteristics of adolescent data to ensure data validity: when cleaning text data, meaningless characters, garbled text, duplicate text and other abnormal text are removed, and valid answers and emotional expressions are retained; when cleaning voice data, noise data with a signal-to-noise ratio of <30dB is filtered out, and clear voice signals are retained to ensure the accuracy of voice feature extraction; when cleaning video data, invalid frames such as blurry frames and occluded frames are removed, and video frames that can clearly identify facial micro-expressions are retained.

[0050] Furthermore, based on multimodal data streams, micro-expression feature vectors and speech feature vectors are extracted and fused, then input into a large-scale model specifically designed for teenagers for processing, outputting a multimodal emotion score, specifically including: The speech data in the multimodal data stream is processed by frame segmentation. Mel frequency cepstral coefficients and their first-order difference coefficients are extracted from each frame of speech signal and combined to form a speech feature vector. The speech feature vector is used to characterize the acoustic properties of adolescent speech, such as pitch changes, speech rate fluctuations and tone fluctuations, which are related to emotions. Facial regions are precisely detected in each frame of video data in a multimodal data stream to locate areas with high incidence of micro-expressions. The amplitude and duration features of micro-expression movements are extracted and combined to form a micro-expression feature vector. The micro-expression feature vector is used to characterize the subtle changes in facial muscles of adolescents and their corresponding emotional tendencies. The areas with high incidence of micro-expressions include the eyes, eyebrows and mouth. The speech feature vector and the micro-expression feature vector are concatenated and fused to form a comprehensive feature vector that includes both speech attributes and facial expression attributes; The comprehensive feature vector is input into a large model specifically for adolescents. An attention mechanism is used to enhance the weights of the feature dimensions in the comprehensive feature vector that are highly correlated with negative emotions, and output a multimodal emotion score. The multimodal emotion score is used to represent the non-verbal comprehensive emotional state of adolescents at the current moment.

[0051] In this embodiment, the present invention extracts and quantifies non-verbal emotional signals of adolescents by fusing facial micro-expressions and voice frequency conversion features, effectively making up for the shortcomings of single text emotion recognition, accurately capturing the hidden emotional changes of adolescents, and improving the comprehensiveness and accuracy of emotion recognition.

[0052] Speech feature extraction: The MFCC (Mel-frequency cepstral coefficients) algorithm is used to preprocess the speech data (including pre-emphasis, framing, windowing, etc.) to eliminate noise interference in the speech signal. Then, the 12-dimensional Mel-frequency cepstral coefficients and their first-order difference coefficients are extracted to finally obtain a 24-dimensional speech feature vector. This vector can accurately capture the core features related to emotions in speech, such as pitch, speech rate, volume, and tone. For example, when there is a negative emotion, the speech rate slows down and the pitch drops sharply, while when there is a positive emotion, the speech rate is stable and the pitch is moderate.

[0053] Micro-expression feature extraction: Based on the CNN (Convolutional Neural Network) micro-expression recognition model, the MTCNN (Multi-Task Cascaded Convolutional Neural Network) algorithm is used to accurately detect the facial regions of adolescents. The focus is on extracting features from regions with obvious micro-expression changes, such as the eyes, eyebrows, and mouth, and finally obtaining a 64-dimensional micro-expression feature vector. At the same time, the amplitude of micro-expression movements (0-10 points, the higher the score, the larger the amplitude of movement) and duration (0-5 seconds) are quantified, and the corresponding emotional states are associated (e.g., frowning and avoiding eye contact correspond to anxiety, downturned corners of the mouth correspond to depression, and bright eyes correspond to positive emotions).

[0054] Feature fusion and quantization: The 24-dimensional speech feature vector and the 64-dimensional micro-expression feature vector are fused into an 88-dimensional comprehensive feature vector, which is then input into a large-scale model specifically designed for teenagers. Combined with speech and semantic information, a multimodal emotion score is quantified and output (the score ranges from 0 to 10, with higher scores indicating stronger negative emotions and lower scores indicating stronger positive emotions). The feature fusion process employs an attention mechanism, focusing on features highly correlated with negative emotions (such as slowed speech, frowning, and a low tone), improving the accuracy of negative emotion recognition and effectively capturing the hidden negative emotions of teenagers.

[0055] Furthermore, based on multimodal data streams, a large-scale model specifically designed for teenagers is used to semantically map popular online slang and vague expressions among teenagers in the text data. This identifies pre-defined high-risk keywords and assigns corresponding weights according to risk levels to obtain a text sentiment score. Specifically, this includes: Based on a large model specifically for teenagers, this study performs in-depth analysis of popular online slang and ambiguous expressions among teenagers in text data from multimodal data streams, combined with contextual information. It maps the obscure expressions commonly used by teenagers to their corresponding real emotional states and outputs a negative emotion intensity index, which is used to characterize the severity of negative emotions implied in the text. Based on a large model specifically for teenagers, the text data in the multimodal data stream is analyzed sentence by sentence. The system automatically identifies the keywords matched in the preset high-risk keyword database, assigns corresponding weight values ​​according to the risk level of the matched keywords, and dynamically adjusts the weight allocation based on the frequency of the keywords in the text data and the context, thus forming the high-risk keyword identification results. The negative emotion intensity index and the high-risk keyword identification results are weighted and fused together. The negative emotion intensity is used as the base score, and the weight of the high-risk keywords is used as the correction factor to output the text sentiment tendency score.

[0056] In this embodiment, a large model specifically designed for teenagers, after secondary fine-tuning, is used to conduct in-depth analysis of teenagers' text data, accurately uncover hidden negative emotions in the text, quantify the text emotion tendency score, and perform dual verification with multimodal emotion scores to further improve the accuracy and reliability of emotion recognition, ensuring that it can comprehensively and realistically reflect the emotional state of teenagers.

[0057] Dedicated Semantic Mining: Our dedicated large-scale model uses deep learning of adolescent social slang and internet buzzwords to accurately grasp the mapping relationships of semantics unique to teenagers. It can distinguish different emotional tendencies of the same expression by combining contextual information, avoiding semantic comprehension biases. For example, it can accurately distinguish the meaning of "playing dumb" in different contexts, recognizing both the negative "playing dumb" emotion arising from psychological distress in teenagers and the "playing dumb" expression used in everyday banter, ensuring the accuracy of emotion recognition.

[0058] High-risk keyword identification: The system analyzes the text data of teenagers sentence by sentence, automatically identifying high-risk keywords such as "despair," "don't want to live," and "life is meaningless." High-risk keywords are divided into three levels and assigned corresponding weights (Level 1 high-risk keywords have a weight of 0.8, such as "don't want to live" and "life is meaningless"; Level 2 high-risk keywords have a weight of 0.5, such as "despair" and "pain"; Level 3 high-risk keywords have a weight of 0.3, such as "sad" and "breakdown"). At the same time, the weight values ​​are dynamically adjusted based on the frequency of keyword occurrence and context to ensure the accuracy of high-risk emotion identification.

[0059] Sentiment Tendency Quantification: The text sentiment tendency score is calculated using a weighted method to simultaneously integrate the intensity of negative sentiment obtained from semantic mining and high-risk keyword signals, avoiding recognition bias caused by relying solely on a single textual clue. The calculation formula is as follows: ; in, The score indicates the sentiment tendency of the text. This represents the negative sentiment weight obtained from semantic mining. This represents the weight of the k-th high-risk keyword, where K represents the number of high-risk keywords hit.

[0060] To improve the robustness of the final sentiment score, both the text sentiment tendency score and the multimodal sentiment score are double-validated. If the difference between the two is no greater than 1 point, the average of the two is taken as the final sentiment score; otherwise, manual correction is performed considering the context. The formula for calculating the final sentiment score is as follows: ; ; in, Indicates multimodal sentiment score, This represents the final emotional score.

[0061] In some embodiments, step S104 above, which uses multidimensional psychological feature vectors, multimodal emotion scores, and text emotion tendency scores as individual data, and combines them with normal emotion fluctuation data of adolescents of the same age as a reference benchmark, to construct an individual's dynamic psychological baseline through a variational autoencoder, specifically includes: Using multidimensional psychological feature vectors, multimodal emotion scores, and text emotion tendency scores as individual data, we also collected normal emotion fluctuation data from multiple groups of adolescents of the same age to form a reference benchmark for normal fluctuations in the same age group. Individual data and a reference benchmark for normal fluctuations in the same age group are input into a variational autoencoder. The encoder part performs dimensionality reduction processing to extract latent vectors, which are used to distinguish between an individual's normal and abnormal psychological states. The latent vector is reconstructed by the decoder part of the variational autoencoder, and combined with the latent center of the normal fluctuation reference benchmark of the same age group, to determine the individual's dynamic psychological baseline and its normal fluctuation range and pathological risk range. The system slides daily with a fixed step size in a preset number of days, incorporating the latest multidimensional psychological data collected within the window into individual data, and iteratively updating and calculating the individual's dynamic psychological baseline and its normal fluctuation range and pathological risk range.

[0062] In this embodiment, the present invention overcomes the limitation of existing systems that "cannot distinguish between normal emotional fluctuations and pathological risks." Based on VAE (Variational Autoencoder), it constructs a psychological evolution trajectory model for adolescents and innovatively proposes a "developmental maladaptive disorder" classification model. It incorporates the physiological emotional fluctuations of adolescents into the construction process of dynamic psychological baselines and designs scientific and rigorous early warning trigger logic to effectively avoid the problems of excessive early warning and early warning distortion, thereby achieving accurate judgment of psychological risks for adolescents.

[0063] Using VAE (Variational Autoencoder) as the core algorithm and combining the stage characteristics of adolescent psychological development, a dynamic psychological baseline specifically for adolescents is constructed. This overcomes the shortcomings of the "single static" nature of general baselines and can accurately distinguish between normal emotional fluctuations and pathological psychological risks in adolescents, providing a reliable basis for psychological risk assessment.

[0064] Baseline data collection: The system collects core data such as 128-dimensional psychological feature vectors, multimodal emotion scores, and text emotion tendency scores of adolescents over the past 30 days. At the same time, it incorporates normal emotional fluctuation data of no less than 1,000 groups of adolescents of the same age as a reference benchmark. Individual data is collected once a day to ensure that the collected data is comprehensive, representative, and continuous, and can truly reflect the psychological change trajectory of adolescents.

[0065] Baseline Construction Logic: The VAE algorithm is used to perform dimensionality reduction and fusion processing on the collected multi-dimensional data to extract the core evolutionary features of adolescents' psychological states. Simultaneously, physiological emotional fluctuations in adolescents (such as normal emotional ups and downs during puberty, and brief pre-exam anxiety) are incorporated into the baseline calculation process to determine the normal fluctuation range and pathological risk range of the individual's dynamic psychological baseline. The VAE encoder adopts a multi-branch structure, capable of simultaneously processing multi-source inputs such as 128-dimensional psychological feature vectors, multimodal emotion scores, and text emotion tendency scores, outputting a 32-dimensional latent feature vector. The decoder uses an MLP (Multilayer Perceptron) architecture, with the reconstruction error (MSE) controlled to <0.01, ensuring the accuracy of baseline construction. To enable the VAE encoder to stably learn the latent distribution structure between "normal fluctuations" and "pathological deviations," a joint optimization method of "reconstruction loss + KL divergence loss + baseline consistency constraint" is used during the training phase. The baseline consistency constraint is used to narrow the distance between the latent vectors of normal samples and the reference center of the same age group, supporting the subsequent baseline construction logic.

[0066] The total training loss function for the VAE is as follows: ; ; ; in, This represents the total loss of the VAE. Indicates the reconstruction loss. Represents the latent distribution With prior distribution KL divergence between them Indicates baseline consistency loss. and These represent the weighting coefficients of the KL term and the baseline consistency term, respectively. This represents the i-th input sample used to train the VAE encoder. This represents the reconstructed sample corresponding to the i-th input sample. This represents the latent vector output by the encoder for the i-th reference sample with normal fluctuations in the same age group. This represents the potential center of a sample with normal fluctuations within the same age group. This indicates the number of training samples used to train the VAE encoder. This indicates the number of reference samples for normal fluctuations within the same age group.

[0067] During training, 50 rounds of unsupervised pre-training were first performed on all samples to stabilize the reconstruction capability. Then, 100 rounds of joint training were performed on reference samples with normal fluctuations and labeled risk samples. The optimizer used was Adam, with an initial learning rate of 0.001, a batch size of 64, and an Early Stopping strategy. Training was stopped when the validation set loss did not decrease for 10 consecutive rounds to avoid overfitting and ensure the generalizability of the latent space.

[0068] Baseline Dynamic Updates: A 7-day sliding window (1-day step) is used to update the dynamic psychological baseline daily, ensuring that the baseline can adapt to the psychological changes of adolescents in real time. At the same time, the baseline parameters are automatically adjusted according to the adolescent's developmental stage (such as from primary school to junior high school, from junior high school to senior high school) and changes in life scenarios (such as exam season, winter and summer vacations, family changes, etc.). An abnormal adjustment mechanism is established so that when adolescents encounter major life events (such as the death of a loved one, major setbacks, etc.), the normal fluctuation range of the baseline is temporarily adjusted to avoid misjudgment caused by special events and ensure the adaptability and accuracy of the baseline.

[0069] In some embodiments, step S105 above, which involves determining whether there is a pathological psychological risk and triggering a graded warning by calculating the Euclidean distance between the mean of the current multidimensional psychological feature vector and the dynamic psychological baseline, specifically includes: Calculate the Euclidean distance between the current multidimensional psychological feature vector and the mean vector of the dynamic psychological baseline in each dimension; Extract the standard deviation range corresponding to the dynamic psychological baseline, compare the Euclidean distance with twice the standard deviation range of the dynamic psychological baseline, and determine that the current psychological state deviates significantly from the normal fluctuation range when the Euclidean distance is greater than twice the standard deviation, and output the core condition satisfaction signal; The probability of psychological risk escalation within a preset number of days in the future, the trend of negative emotion scores between the current moment and the previous moment, and the abnormality score of behavior are obtained as indicators of collapse precursors. When the probability of psychological risk escalation reaches a preset probability threshold, the negative emotion score shows a continuous upward trend, and the abnormality score of behavior reaches a preset score threshold, a collapse precursor condition is met signal is output. When the core condition signal and the collapse precursor condition signal are met simultaneously, the current psychological state of the adolescents is comprehensively judged as a pathological psychological risk. The psychological risk is classified according to the degree of deviation of Euclidean distance and the severity of collapse precursor indicators, and a corresponding level of warning signal is output.

[0070] In this embodiment, based on an individual's dynamic psychological baseline, a scientific and rigorous early warning triggering logic is designed. The early warning is triggered only when the system determines that there is a pathological psychological risk, thus avoiding ineffective and excessive early warnings, ensuring the accuracy of the early warning, and providing a timely and reliable basis for early intervention of psychological risks in adolescents.

[0071] The core criterion is the Euclidean distance between the current adolescent's 128-dimensional psychological feature vector and the mean of the dynamic psychological baseline. This distance quantifies the degree of deviation of the current psychological state from the individual's normal range. A larger Euclidean distance indicates a greater deviation from the normal range and a higher likelihood of pathological psychological risk; a smaller distance indicates a closer proximity to the normal range and a higher likelihood of normal emotional fluctuations. The Euclidean distance calculation formula is as follows: ; in, This represents the Euclidean distance between the current psychological feature vector and the mean vector of the dynamic psychological baseline. Indicates the current time. Feature values ​​of each psychological dimension Indicates the dynamic psychological baseline at the 1st The mean of each psychological dimension.

[0072] Warning Triggering Conditions: The system sets strict warning triggering conditions. A pathological psychological risk is only identified and a warning is triggered when the Euclidean distance deviation exceeds twice the standard deviation, and simultaneously meets the "collapse precursor" thresholds (probability of psychological risk escalation ≥ 80% in the next 3 days, continuously rising negative emotion score, and behavioral abnormality score ≥ 7 points). To ensure the executability of the triggering logic, the following joint judgment rule is adopted: ; in, This indicates that an alert has been triggered. This indicates the range of standard deviations corresponding to the dynamic psychological baseline. This indicates the probability of an escalation of psychological risk over the next 3 days. and These represent the negative emotion scores at the current moment and the previous moment, respectively. This indicates the degree of behavioral abnormality. The system will only output a pathological psychological risk warning if all four of the above criteria are met simultaneously. The warnings are divided into three levels: mild, moderate, and severe, each corresponding to different intervention strategies: a mild warning corresponds to mild psychological distress, which is initially addressed by the homeroom teacher or the psychological counselor; a moderate warning corresponds to significant psychological distress, which is addressed by the school psychologist; and a severe warning corresponds to serious psychological problems, which requires intervention by a professional mental health institution.

[0073] Normal fluctuation judgment: If the deviation of the adolescent's Euclidean distance does not exceed 2 times the standard deviation, or does not meet the "collapse precursor" threshold, it is judged as normal physiological emotional fluctuation, and no warning is triggered. The system only continuously monitors the adolescent's psychological state and records the psychological evolution trajectory to ensure that neither pathological risks are missed nor invalid warnings are generated.

[0074] Furthermore, the acquisition of the probability of psychological risk escalation within a preset number of days in the future, the trend of negative emotion scores between the current moment and the previous moment, and the behavioral abnormality score as indicators of impending collapse specifically includes: The model inputs the multidimensional psychological feature vector, multimodal emotion score, and text emotion tendency score of the adolescent at the current moment into the risk evolution prediction model. Combined with the historical evolution trajectory of the dynamic psychological baseline, it predicts the probability that the psychological state will develop from the current deviation level to a more serious level within a preset number of days in the future, and outputs the probability of psychological risk escalation. Extract the final emotion score at the current moment and the final emotion score at the previous moment. When the emotion score at the current moment is higher than the emotion score at the previous moment, it is determined that the negative emotion is showing a continuous upward trend, and the negative emotion score change trend signal is output. The study collected data on adolescents' response behavior characteristics, interaction response duration, and operational regularity indicators within a preset time window. This data was then compared with the adolescents' own historical behavior patterns and referenced the behavioral baseline of adolescents in the same age group. A comprehensive quantitative calculation was then performed to obtain a behavioral abnormality score.

[0075] In this embodiment, a risk evolution prediction model is constructed. This model employs a sequential neural network architecture, using the adolescent's current multidimensional psychological feature vector, multimodal emotion score, and text sentiment tendency score as input layer data. The multidimensional psychological feature vector includes quantified values ​​for 12 dimensions such as academic pressure, peer relationships, and self-identity. The multimodal emotion score integrates micro-expression recognition results (e.g., frowning frequency, duration of drooping corners of the mouth) with speech features (e.g., speech rate fluctuations, pitch variation amplitude). The text sentiment tendency score is based on a semantic mining model's weighted calculation of high-risk keywords (e.g., "despair," "life is meaningless") in social texts. Secondly, the historical evolution trajectory of the dynamic psychological baseline is input into the model as time-series data. This trajectory records the average daily change in the adolescent's psychological feature vector over the past 30 days. By comparing the deviation patterns between the current data and the historical trajectory, and considering the psychological development patterns of adolescents of the same age, the model predicts the probability that the psychological state will escalate from the current level of deviation to a more severe level (e.g., from mild anxiety to depressive tendencies) within the next 7 days. Finally, it outputs a psychological risk escalation probability value within the range of 0-100%.

[0076] The system employs a two-tiered emotion scoring mechanism: The first tier is the instantaneous emotion score, generated by weighted fusion of multimodal emotion scores (60% weight) and text emotion tendency scores (40% weight). For example, if a teenager exhibits increased speaking speed in their speech (multimodal score +15) and the keyword "irritable" appears in their social text (text score +20), their instantaneous emotion score is 31. The second tier is the final emotion score, which is built upon the instantaneous score by adding a behavioral compensation factor. This factor is dynamically adjusted based on response time (2 points are deducted for every minute shorter) and operational regularity (5 points are deducted for selecting the same option more than 3 times consecutively). The system collects the final emotion score every 15 minutes, forming a time-series dataset. The trend determination module compares the current moment's final emotion score with the previous moment's: if the current score is higher than the previous moment's and the difference exceeds a preset threshold (e.g., 5 points), it indicates a continuous upward trend in negative emotion; if the upward condition is met in 3 consecutive comparisons, a trend reinforcement signal is output. For example, if a teenager's final emotional scores at 18:00, 18:15, and 18:30 are 42, 48, and 55 respectively, the system will determine that their negative emotions show a significant upward trend.

[0077] A three-dimensional behavioral assessment system is constructed: The first dimension is answering behavior characteristics, including the frequency of rapid answering (more than 20 questions per minute is considered abnormal), the number of times the same option is selected consecutively (more than 3 times triggers an alert), and the subjective question skipping rate (more than 50% is considered invalid answering); the second dimension is interaction response time, which assesses focus by calculating the deviation rate between the average answering time per question and the baseline value for adolescents of the same age (a deviation exceeding ±30% is considered abnormal); the third dimension is operational regularity, which uses a sliding window algorithm to analyze the standard deviation of operational data such as mouse movement trajectory and click interval time. When the standard deviation exceeds twice the historical mean, it is considered operational disorder. The system collects the above indicator data every 30 minutes and dynamically compares it with the adolescent's historical behavioral patterns over the past 7 days, while also referring to the behavioral baseline of adolescents of the same age (e.g., the average rapid answering frequency for the 14-16 age group is 5 times / hour). The abnormality calculation module uses a weighted scoring method: answering behavior characteristics account for 40%, interaction response time accounts for 35%, and operational regularity accounts for 25%, finally outputting a behavioral abnormality score within the range of 0-100 points. For example, if a teenager answers questions quickly 7 times during the monitoring period, skips 60% of subjective questions, and has a standard deviation of 2.5 times the historical mean for operational data, their behavioral abnormality score is 82.

[0078] In some embodiments, in step S106 above, the step of using a state machine model, taking response time, operational characteristics, and response patterns as inputs, outputting the probability of perfunctory responses, and dynamically adjusting questionnaire questions and guiding remarks in conjunction with the Prompt project, specifically includes: During the process of teenagers answering the questionnaire, the answering time for a single question, the overall answering time of the questionnaire, the frequency of answer deletion, the interval between option selection, the time spent on the page, the frequency of consecutively selecting the same option, and the skipping of subjective questions are collected in real time as answering behavior characteristic data. The answer behavior feature data is input into a state machine model, which includes a feature encoding layer, a temporal layer and a fully connected classification layer. The feature encoding layer is used to compress the answer behavior feature data into a behavior feature vector. The temporal layer is used to capture the evolution of the behavior feature vector in the continuous answering process. The fully connected classification layer is used to output the probability distribution of three states: normal answer, mild perfunctory answer and severe perfunctory answer. The probabilities of mild and severe perfunctory responses are weighted and fused to form the probability of perfunctory response. The probability of perfunctory response is then compared with multiple preset perfunctory probability thresholds to determine the current response status. Based on the current response status, the questionnaire questions and guiding language are dynamically adjusted through the Prompt project.

[0079] In one possible implementation, the probability of giving a perfunctory answer is compared with multiple preset perfunctory probability thresholds to determine the current answering state, including: maintaining a normal answering state when the probability of giving a perfunctory answer is lower than a first threshold; switching to a mild perfunctory state when the probability of giving a perfunctory answer is between the first threshold and a second threshold; and switching to a severe perfunctory state when the probability of giving a perfunctory answer is not lower than the second threshold. A buffer mechanism is set during the state switching process: when a single monitoring indicator is briefly abnormal but other indicators are normal and the duration of the abnormality does not exceed a preset duration, the state switching is not triggered; when the probability of giving a perfunctory answer briefly exceeds the threshold and the subsequent answering behavior returns to normal and remains so for more than a preset duration, the system automatically reverts to the normal answering state.

[0080] In one possible implementation, based on the current response status, the questionnaire questions and guiding remarks are dynamically adjusted through the Prompt process. This includes: maintaining the original questionnaire settings and using standardized guiding remarks when the response status is normal; automatically reducing the number of subsequent questions and simplifying the question descriptions when the response status is slightly perfunctory, while switching to gentle guiding remarks to alleviate the adolescent's resistance; and immediately pausing the current questionnaire through the Prompt process when the response status is severely perfunctory, switching to a fun or lightweight alternative interaction format, and using encouraging guiding remarks, and continuing the evaluation process after the adolescent's response status recovers.

[0081] In this embodiment, the present invention overcomes the limitations of existing systems such as "rigid interactive collection mechanism and insufficient data authenticity" by optimizing the state machine model of the intervention executor agent and designing a scientific interactive adaptive adjustment mechanism. By monitoring the adolescents' answering behavior in real time, the questionnaire questions and guiding words are dynamically adjusted to ensure the authenticity and effectiveness of the collected data, while improving the adolescents' answering experience and effectively solving the problem of data distortion caused by perfunctory answers.

[0082] Intervention implementers utilize an optimized state machine model to monitor adolescents' questionnaire-answering behavior in real time throughout the entire process, capturing passive feedback data related to perfunctory responses. The core monitoring indicators include three categories: response duration, operational characteristics, and response patterns. The probability of "perfunctory response" (ranging from 0-100%) is quantified and calculated, serving as the core basis for subsequent interactive adjustments. Simultaneously, the monitoring data is synchronized in real time to risk observers, providing supplementary support for psychological risk assessment. To balance temporal behavioral pattern recognition capabilities with online inference efficiency, the state machine model employs a lightweight network structure of "feature encoding layer + GRU temporal layer + fully connected classification layer." The input is a sequence of behavioral features within a continuous response window, and the output is the probability distribution of three states: normal response, mild perfunctory response, and severe perfunctory response.

[0083] The core calculation process of the state machine model is as follows: ; ; ; in, This represents the behavioral feature vector at time t. This represents a feature related to response time. Indicates operational characteristics, Indicate the pattern and characteristics of the answers. This represents the hidden state of the GRU at time t. This represents the hidden state of the GRU at time t-1. The output represents the probability of the three response states. and These represent the output layer weights and biases, respectively, and GRU stands for Gated Recurrent Unit. express An activation function is used to transform the output into a probability distribution.

[0084] The state machine model is divided into three states: normal response, mild perfunctory response, and severe perfunctory response. The state switches in real time based on the probability of a "perfunctory response." The quantification methods for monitoring indicators are clear and operable: Response time indicators include average response time per question and overall questionnaire response time; an average response time per question <3 seconds or an overall response time less than 50% of the standard time is considered abnormal. Operational characteristic indicators include answer deletion frequency, check interval, and page dwell time; excessively high answer deletion frequency or excessively short check intervals are considered abnormal. Response pattern indicators include consecutively checking the same option, skipping subjective questions, and contradictory answers; consecutively checking the same option ≥3 times or skipping all subjective questions are considered abnormal. Each indicator is assigned a corresponding weight according to the degree of abnormality (the higher the degree of abnormality, the greater the weight). During model training, manually labeled adolescent response logs are used as training samples. The ratio of normal responses, mild perfunctory responses, and severe perfunctory responses is controlled at 4:3:3. The training batch size is 32, the learning rate is 0.001, the number of iterations is 120, and the optimizer used is Adam. To improve state recognition accuracy, the cross-entropy loss function is used to optimize the model parameters during the training phase. The formula is as follows: ; in, This represents the cross-entropy loss of the state machine model during the training phase. This represents the number of training samples used to train the state machine model. Indicates the first The training sample at the th ... The actual label in the class state, This indicates that the state machine model predicts the first... The training sample belongs to the first... The probability of each class state. During training, a 20% validation set is used for model selection, and training is stopped early if the validation set loss does not decrease for 8 consecutive rounds.

[0085] The probability of "perfunctory response" is calculated as follows: To uniformly map multiple types of abnormal behavior into interpretable probability indicators, the abnormality weight normalization method is used to calculate the probability of perfunctory response, and the formula is as follows: ; in, This indicates the probability of giving a perfunctory answer. This indicates whether the m-th abnormal indicator is triggered or its normalized abnormal intensity, where M represents the number of abnormal indicators. This indicates the weight of the corresponding abnormal indicator. This represents the maximum anomaly weight value, which defaults to 2.0. State transitions use a threshold-triggered mechanism: when... When, it is judged as a normal answering state; when When, it is judged as a slightly perfunctory state; when At that time, it was judged as a state of severe perfunctory work.

[0086] Meanwhile, the model is equipped with a buffer mechanism. If a certain monitoring indicator is abnormal, but other indicators are normal and the abnormality lasts for less than 30 seconds, it will not be included in the calculation of abnormal weight. If the probability of perfunctory response rises to more than 30% for a short period of time, but the subsequent response behavior returns to normal and lasts for more than 1 minute, it will automatically revert to the normal response state to avoid misjudgment caused by operational errors (such as accidental touch or network lag).

[0087] Based on the quantitatively calculated probability of "perfunctory responses," the system automatically adjusts the subsequent evaluation process through the Prompt project, optimizes questionnaire questions and guiding language, and all adjustment strategies are designed based on the cognitive level and psychological characteristics of teenagers to avoid causing greater resistance due to improper adjustments. While improving the authenticity of the data, it significantly improves the teenagers' response experience.

[0088] Perfunctory response rate <30% (normal response status): Maintain the original questionnaire settings, using standardized guiding language that is concise, clear, and aligned with adolescents' cognitive abilities, without adding unnecessary or redundant expressions, such as "Please answer based on your true feelings, no need to overthink" and "Every choice you make helps us better understand you." Simultaneously, during the response process, the system displays a brief encouraging prompt every 8 questions completed (such as "Great job, keep it up!" or "Thank you for your honest response") to maintain adolescents' enthusiasm and avoid fatigue from prolonged responses. In this scenario, the questionnaire completion rate is ≥92%, and the data authenticity is ≥88%, effectively ensuring data quality.

[0089] 30% ≤ Perfunctory Probability < 60% (Mild Perfunctory State): Initiate the adjustment strategy of "reducing questions and burden, simplifying expression, and gentle guidance". The specific adjustments are as follows: First, automatically reduce the number of subsequent questions. Based on the number of questions already answered, reduce the number of subsequent questions by 30%-50%, prioritizing the retention of questions on core dimensions such as academic pressure and peer relationships, and eliminating questions with low relevance to the core psychological problems of adolescents, ensuring that the total questionnaire time is controlled within 10 minutes; Second, reduce the logical difficulty of the questions, simplify the expression of the questions, eliminate abstract and obscure expressions, and use more concrete language that is closer to the life scenarios of adolescents to ensure that adolescents can quickly understand the meaning of the questions; Third, switch to gentle guidance language, with a friendly and patient tone to alleviate the resistance of adolescents, such as "Take your time to answer, don't rush, just choose according to your true feelings" and "It's okay, even a simple choice can reflect your true state". At the same time, the frequency of encouraging prompts will be increased, with an encouraging message popping up every four questions completed, and simple interactive prompts will be added appropriately (such as "Your answer is very true, keep it up") to guide teenagers to answer questions seriously.

[0090] Perfunctory response rate ≥ 60% (severe perfunctory state): Immediately implement the adjustment strategy of "switching formats, lowering thresholds, and strengthening encouragement." Specific adjustments are as follows: First, immediately pause the current questionnaire to prevent teenagers from continuing to answer perfunctorily due to resistance. A gentle prompt will appear when pausing (e.g., "We've noticed you might be a little tired; let's continue in a more relaxed way"). Second, switch to a more engaging and lightweight interactive format to replace the traditional text-based questionnaire. Specific adjustments will be adapted to age groups: In elementary school, multiple-choice questions will be changed to situational judgment questions with simple cartoon illustrations, and text questions will be changed to voice responses, allowing teenagers to express their feelings orally. In middle and high school, multiple-choice questions will be changed to short sentence matching questions, and text questions will be changed to short sentence fill-in-the-blank questions, reducing the difficulty of answering. Third, use encouraging guiding language with a warm and positive tone to give teenagers ample affirmation and encouragement, eliminating their resistance. For example, "This question is very simple; tell us your thoughts, we're listening carefully," or "Don't be afraid to express yourself; your feelings are important." Once the teenagers' response status returns to normal (the probability of perfunctory responses drops below 30%, and they consistently answer ≥3 questions), the system will continue to evaluate the process. If severe perfunctory responses reappear, the above adjustment strategy will be repeated until the questionnaire is completed or the teenagers explicitly indicate that they are giving up.

[0091] To ensure the scientific validity and rationality of the adaptive adjustment strategy, the system regularly collects feedback from teenagers (such as "Do you find the current questionnaire format easy?" and "Do you feel comfortable with the guiding language?"). Combining this feedback with monitoring data, the system continuously optimizes the parameters of the adjustment strategy (such as the percentage of questions reduced, the wording, and the suitability of the interaction format). Simultaneously, an adaptability assessment mechanism for the adjustment strategy is established, with pilot tests conducted quarterly. Based on the test results, the adjustment thresholds and strategy details for different age groups are adjusted to ensure that the adjustment strategy always aligns with the cognitive level and psychological characteristics of teenagers.

[0092] Reference Figure 2 An embodiment of the present invention provides a dynamic monitoring system 2 for psychological risk based on a large-scale model specifically for adolescents. The system 2 specifically includes: Model training module 201 is used to pre-train a general large language model with a general psychological corpus and then fine-tune it with a dedicated labeled dataset to form a large model specifically for adolescents. Feature extraction module 202 is used to standardize indicators based on the MHT scale and PHQ-9 scale, and extract multi-dimensional psychological feature vectors from the response data through a large model specifically for adolescents. The scoring calculation module 203 is used to collect multimodal data streams, and through a large model specifically for teenagers, it integrates the recognition of micro-expressions and speech to output multimodal emotion scores. It also performs semantic mining and high-risk keyword recognition on text data to output text emotion tendency scores. The baseline construction module 204 is used to construct the dynamic psychological baseline of an individual by using multidimensional psychological feature vectors, multimodal emotion scores and text emotion tendency scores as individual data, combined with normal emotion fluctuation data of adolescents of the same age as a reference benchmark, through variational autoencoder. The risk warning module 205 is used to determine whether there is a pathological psychological risk and trigger a graded warning by calculating the Euclidean distance between the current multidimensional psychological feature vector and the mean of the dynamic psychological baseline. The perfunctory response assessment module 206 is used to output the probability of perfunctory responses based on the state machine model, taking the response time, operational characteristics and response patterns as inputs, and dynamically adjusting the questionnaire questions and guiding words in conjunction with the Prompt project.

[0093] It is understandable that, such as Figure 1 The content of the embodiments of the dynamic monitoring method for psychological risk based on a large-scale model for adolescents shown are all applicable to the embodiments of the dynamic monitoring system for psychological risk based on a large-scale model for adolescents. The specific functions implemented by the embodiments of the dynamic monitoring system for psychological risk based on a large-scale model for adolescents are as follows: Figure 1 The illustrated embodiment of the dynamic monitoring method for psychological risk based on a large-scale model specifically for adolescents is the same, and the beneficial effects achieved are the same as those shown. Figure 1 The beneficial effects achieved by the embodiment of the dynamic monitoring method for psychological risk based on a large model specifically for adolescents are also the same.

[0094] It should be noted that the information interaction and execution process between the above systems are based on the same concept as the method embodiments of the present invention. For details on their specific functions and technical effects, please refer to the method embodiments section, which will not be repeated here.

[0095] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the system can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0096] Reference Figure 3The present invention also provides a computer device 3, including: a memory 302 and a processor 301, and a computer program 303 stored in the memory 302. When the computer program 303 is executed on the processor 301, it implements the dynamic monitoring method for psychological risk based on a large model specifically for adolescents as described in any of the above methods.

[0097] The computer device 3 may be a desktop computer, laptop, handheld computer, or cloud server, etc. The computer device 3 may include, but is not limited to, a processor 301 and a memory 302. Those skilled in the art will understand that... Figure 3 The computer device 3 is merely an example and does not constitute a limitation on the computer device 3. It may include more or fewer components than shown in the figure, or combine certain components, or different components, such as input / output devices, network access devices, etc.

[0098] The processor 301 can be a Central Processing Unit (CPU), but it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.

[0099] In some embodiments, the memory 302 may be an internal storage unit of the computer device 3, such as a hard disk or memory of the computer device 3. In other embodiments, the memory 302 may be an external storage device of the computer device 3, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the computer device 3. Furthermore, the memory 302 may include both internal and external storage units of the computer device 3. The memory 302 is used to store the operating system, applications, boot loader, data, and other programs, such as the program code of the computer program. The memory 302 can also be used to temporarily store data that has been output or will be output.

[0100] This invention also provides a computer-readable storage medium storing a computer program thereon. When the computer program is run by a processor, it implements the method for dynamic monitoring of psychological risk based on a large model specifically for adolescents, as described in any of the above methods.

[0101] In this embodiment, if the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include at least: any entity or device capable of carrying computer program code to a photographing device / terminal device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electrical carrier signals or telecommunication signals.

[0102] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0103] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0104] In the embodiments disclosed in this application, it should be understood that the disclosed devices / terminal equipment and methods can be implemented in other ways. For example, the device / terminal equipment embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling or direct coupling or communication connection may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

Claims

1. A method for dynamic monitoring of psychological risk based on a large-scale model specifically for adolescents, characterized in that, The method specifically includes: A general large language model was pre-trained using a general psychological corpus and then fine-tuned using a dedicated labeled dataset to form a large model specifically for adolescents. Based on the MHT scale and PHQ-9 scale, the indicators were standardized and modified, and multi-dimensional psychological feature vectors were extracted from the response data through a large model specifically for adolescents. Collect multimodal data streams, use a large model specifically designed for teenagers to integrate and identify micro-expressions and speech to output multimodal emotion scores, and perform semantic mining and high-risk keyword identification on text data to output text emotion tendency scores. Using multidimensional psychological feature vectors, multimodal emotion scores, and text emotion tendency scores as individual data, combined with normal emotion fluctuation data of adolescents of the same age as a reference benchmark, a dynamic psychological baseline for individuals is constructed through a variational autoencoder. By calculating the Euclidean distance between the current multidimensional psychological feature vector and the mean of the dynamic psychological baseline, it is determined whether there is a pathological psychological risk and a graded warning is triggered. Based on the state machine model, the system takes the response time, operational characteristics, and response patterns as inputs and outputs the probability of perfunctory responses. It also dynamically adjusts the questionnaire questions and guiding language in conjunction with the Prompt project.

2. The method according to claim 1, characterized in that, The process of pre-training a general-purpose large language model using a general-purpose psychological corpus and then fine-tuning it using a dedicated labeled dataset to form a large-scale model specifically for adolescents includes: A general psychological corpus containing popular science knowledge, clinical psychological cases, and psychological questionnaire data was collected and constructed. Based on the general psychological corpus, a large language model using the Transformer architecture was pre-trained to obtain the pre-trained large language model. We collected multi-source data, including adolescent psychological questionnaire responses, adolescent social language data, adolescent emotional expression text data, and adolescent speech-to-text data. We standardized and labeled the multi-source data and ensured the consistency of the labeling to form a dedicated labeled dataset for adolescents covering different age groups. We used a youth-specific labeled dataset to fine-tune the pre-trained large language model. During the fine-tuning process, we adjusted the learning rate and introduced regularization constraints to optimize the mapping weights for youth-specific semantics, thus obtaining a youth-specific large model.

3. The method according to claim 1, characterized in that, The standardization of indicators based on the MHT and PHQ-9 scales, and the extraction of multidimensional psychological feature vectors from the response data using a large-scale model specifically designed for adolescents, specifically include: Based on the MHT scale and PHQ-9 scale, general indicators with low correlation to adolescent psychological distress were removed, and core dimensions specific to adolescents were added to form an adolescent-specific scale. These core dimensions include academic pressure, peer relationships, self-identity, and family environment. The response data generated by the adolescent-specific scale is input into the adolescent-specific big model. Through in-depth semantic analysis and psychological feature mapping, multi-dimensional psychological feature vectors are extracted. Based on the Z-score standardization formula, the multidimensional psychological feature vector is standardized.

4. The method according to claim 1, characterized in that, The process involves collecting multimodal data streams, using a large-scale model specifically designed for teenagers, fusing micro-expression and speech data to output multimodal emotion scores, and performing semantic mining and high-risk keyword identification on text data to output text sentiment tendency scores. Specifically, this includes: Simultaneously access text data streams, voice data streams, and video data streams, and perform data cleaning and format standardization to form a multimodal data stream; Based on multimodal data streams, micro-expression feature vectors and speech feature vectors are extracted and fused into a large model specifically for teenagers for processing, and multimodal emotion scores are output. Based on multimodal data streams, a large-scale model specifically designed for teenagers is used to semantically map popular online slang and vague expressions among teenagers in text data. Preset high-risk keywords are identified and assigned corresponding weights according to risk levels to obtain text sentiment scores.

5. The method according to claim 4, characterized in that, The process involves extracting micro-expression feature vectors and speech feature vectors based on multimodal data streams, fusing them, and inputting them into a large-scale model specifically designed for teenagers for processing, outputting a multimodal emotion score, specifically including: The speech data in the multimodal data stream is processed by frame segmentation. Mel frequency cepstral coefficients and their first-order difference coefficients are extracted from each frame of speech signal and combined to form a speech feature vector. The speech feature vector is used to characterize the acoustic properties of adolescent speech, such as pitch changes, speech rate fluctuations and tone fluctuations, which are related to emotions. Facial regions are precisely detected in each frame of video data in a multimodal data stream to locate areas with high incidence of micro-expressions. The amplitude and duration features of micro-expression movements are extracted and combined to form a micro-expression feature vector. The micro-expression feature vector is used to characterize the subtle changes in facial muscles of adolescents and their corresponding emotional tendencies. The areas with high incidence of micro-expressions include the eyes, eyebrows and mouth. The speech feature vector and the micro-expression feature vector are concatenated and fused to form a comprehensive feature vector that includes both speech attributes and facial expression attributes; The comprehensive feature vector is input into a large model specifically for adolescents. An attention mechanism is used to enhance the weights of the feature dimensions in the comprehensive feature vector that are highly correlated with negative emotions, and output a multimodal emotion score. The multimodal emotion score is used to represent the non-verbal comprehensive emotional state of adolescents at the current moment.

6. The method according to claim 4, characterized in that, The method, based on multimodal data streams, uses a large-scale model specifically designed for teenagers to perform semantic mapping on popular online slang and vague expressions among teenagers in the text data. It identifies pre-defined high-risk keywords and assigns corresponding weights according to risk levels to obtain a text sentiment score, specifically including: Based on a large model specifically for teenagers, this study performs in-depth analysis of popular online slang and ambiguous expressions among teenagers in text data from multimodal data streams, combined with contextual information. It maps the obscure expressions commonly used by teenagers to their corresponding real emotional states and outputs a negative emotion intensity index, which is used to characterize the severity of negative emotions implied in the text. Based on a large model specifically for teenagers, the text data in the multimodal data stream is analyzed sentence by sentence. The system automatically identifies the keywords matched in the preset high-risk keyword database, assigns corresponding weight values ​​according to the risk level of the matched keywords, and dynamically adjusts the weight allocation based on the frequency of the keywords in the text data and the context, thus forming the high-risk keyword identification results. The negative emotion intensity index and the high-risk keyword identification results are weighted and fused together. The negative emotion intensity is used as the base score, and the weight of the high-risk keywords is used as the correction factor to output the text sentiment tendency score.

7. The method according to claim 1, characterized in that, The method uses multidimensional psychological feature vectors, multimodal emotion scores, and text emotion tendency scores as individual data, combined with normal emotion fluctuation data of adolescents of the same age as a reference benchmark, to construct an individual's dynamic psychological baseline through a variational autoencoder, specifically including: Using multidimensional psychological feature vectors, multimodal emotion scores, and text emotion tendency scores as individual data, we also collected normal emotion fluctuation data from multiple groups of adolescents of the same age to form a reference benchmark for normal fluctuations in the same age group. Individual data and a reference benchmark for normal fluctuations in the same age group are input into a variational autoencoder. The encoder part performs dimensionality reduction processing to extract latent vectors, which are used to distinguish between an individual's normal and abnormal psychological states. The latent vector is reconstructed by the decoder part of the variational autoencoder, and combined with the latent center of the normal fluctuation reference benchmark of the same age group, to determine the individual's dynamic psychological baseline and its normal fluctuation range and pathological risk range. The system slides daily with a fixed step size in a preset number of days, incorporating the latest multidimensional psychological data collected within the window into individual data, and iteratively updating and calculating the individual's dynamic psychological baseline and its normal fluctuation range and pathological risk range.

8. The method according to claim 1, characterized in that, The process of determining whether there is a pathological psychological risk and triggering a graded warning by calculating the Euclidean distance between the current multidimensional psychological feature vector and the mean of the dynamic psychological baseline specifically includes: Calculate the Euclidean distance between the current multidimensional psychological feature vector and the mean vector of the dynamic psychological baseline in each dimension; Extract the standard deviation range corresponding to the dynamic psychological baseline, compare the Euclidean distance with twice the standard deviation range of the dynamic psychological baseline, and determine that the current psychological state deviates significantly from the normal fluctuation range when the Euclidean distance is greater than twice the standard deviation, and output the core condition satisfaction signal; The probability of psychological risk escalation within a preset number of days in the future, the trend of negative emotion scores between the current moment and the previous moment, and the abnormality score of behavior are obtained as indicators of collapse precursors. When the probability of psychological risk escalation reaches a preset probability threshold, the negative emotion score shows a continuous upward trend, and the abnormality score of behavior reaches a preset score threshold, a collapse precursor condition is met signal is output. When the core condition signal and the collapse precursor condition signal are met simultaneously, the current psychological state of the adolescents is comprehensively judged as a pathological psychological risk. The psychological risk is classified according to the degree of deviation of Euclidean distance and the severity of collapse precursor indicators, and a corresponding level of warning signal is output.

9. The method according to claim 1, characterized in that, The state machine model, taking response time, operational characteristics, and response patterns as inputs, outputs the probability of perfunctory responses. It also dynamically adjusts questionnaire questions and guiding statements using the Prompt process, specifically including: During the process of teenagers answering the questionnaire, the answering time for a single question, the overall answering time of the questionnaire, the frequency of answer deletion, the interval between option selection, the time spent on the page, the frequency of consecutively selecting the same option, and the skipping of subjective questions are collected in real time as answering behavior characteristic data. The answer behavior feature data is input into a state machine model, which includes a feature encoding layer, a temporal layer and a fully connected classification layer. The feature encoding layer is used to compress the answer behavior feature data into a behavior feature vector. The temporal layer is used to capture the evolution of the behavior feature vector in the continuous answering process. The fully connected classification layer is used to output the probability distribution of three states: normal answer, mild perfunctory answer and severe perfunctory answer. The probabilities of mild and severe perfunctory responses are weighted and fused to form the probability of perfunctory response. The probability of perfunctory response is then compared with multiple preset perfunctory probability thresholds to determine the current response status. Based on the current response status, the questionnaire questions and guiding language are dynamically adjusted through the Prompt project.

10. A dynamic monitoring system for psychological risk based on a large-scale model specifically for adolescents, characterized in that, The system specifically includes: The model training module is used to pre-train a general large language model with a general psychological corpus, and then fine-tune it with a dedicated labeled dataset to form a large model specifically for teenagers. The feature extraction module is used to standardize the indicators based on the MHT scale and PHQ-9 scale, and extract multi-dimensional psychological feature vectors from the response data through a large model specifically for adolescents. The scoring module is used to collect multimodal data streams, and through a large model specifically for teenagers, it integrates the recognition of micro-expressions and speech to output multimodal emotion scores. It also performs semantic mining and high-risk keyword recognition on text data to output text emotion tendency scores. The baseline construction module is used to construct an individual's dynamic psychological baseline by using multidimensional psychological feature vectors, multimodal emotion scores, and text emotion tendency scores as individual data, combined with normal emotion fluctuation data of adolescents of the same age as a reference benchmark, through a variational autoencoder. The risk warning module is used to determine whether there is a pathological psychological risk and trigger a graded warning by calculating the Euclidean distance between the current multidimensional psychological feature vector and the mean of the dynamic psychological baseline. The perfunctory response assessment module is based on a state machine model. It takes response time, operational characteristics, and response patterns as inputs, outputs the probability of perfunctory responses, and dynamically adjusts questionnaire questions and guiding language in conjunction with the Prompt project.