Intelligent psychological portrait diagnosis and treatment method and system based on large language model driving

The intelligent psychological profiling and diagnosis system driven by a large language model, combined with multimodal data fusion and personalized visual intervention, solves the problem of the single method of psychological health assessment and intervention, realizes real-time and safe psychological state assessment and intervention, and promotes the inclusive development of psychological services.

CN121964069APending Publication Date: 2026-05-01LANZHOU JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
LANZHOU JIAOTONG UNIV
Filing Date
2025-12-19
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing mental health assessment and intervention methods are limited in scope and lack integrated analysis capabilities, failing to achieve real-time capture of mental states and effective means from assessment to intervention.

Method used

The intelligent psychological profiling and diagnosis system driven by a large language model constructs a fully localized, multimodal, real-time closed-loop intelligent psychological profiling and visual intervention system through multimodal data fusion perception, psychological profile generation, and personalized visual intervention. It includes four stages: multimodal data perception and preprocessing, cognitive reasoning and generation of psychological profiles, decision-making and generation of personalized visual intervention, and human-computer interaction and system evolution.

Benefits of technology

It enables contactless and immediate psychological state assessment and intervention, lowers the threshold for psychological assessment, promotes the universalization of psychological services, provides immediate and immersive psychological support, enhances user experience, and ensures user data privacy and security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121964069A_ABST
    Figure CN121964069A_ABST
Patent Text Reader

Abstract

The invention provides an intelligent psychological portrait diagnosis and treatment method and system based on large language model driving, relates to the technical field of psychological portrait diagnosis and treatment, and realizes real-time evaluation and intervention on the psychological state of a user through four core stages of multi-modal data perception, psychological portrait generation, personalized visual intervention and man-machine interaction evolution. According to the system, a high-performance edge computing platform and a deep learning model are adopted, a structured psychological portrait is generated in combination with a psychological knowledge base, and emotion adjustment support is provided through visual content. The main innovation points of the method comprise cross-modal psychological assessment, a real-time closed-loop intervention mechanism, low-cost expert experience replication, a high-performance integrated solution, a sustainable-evolution man-machine cooperation mode and privacy security guarantee, and aims to promote the popularity and high efficiency of psychological health services.
Need to check novelty before this filing date? Find Prior Art

Description

Intelligent Psychological Profiling Diagnosis and Treatment Method and System Driven by Large Language Model Technical Field

[0001] This invention relates to the field of psychological profiling and diagnosis technology, specifically to an intelligent psychological profiling and diagnosis method and system driven by a large language model. Background Technology

[0002] Current mental health assessment and intervention mainly rely on traditional psychological scales and interviews, physiological signal monitoring, computer vision-based facial expression recognition, and large language model dialogue systems. The main problems with existing technologies are that the assessment methods are singular, the ability to integrate and analyze is poor, the ability to capture mental states in real time is not possible, and there is a lack of effective means from assessment to intervention.

[0003] 1.1 The Application Potential of Large Language Model (LLM) in the Field of Psychological Services

[0004] Large Language Models (LLMs) have demonstrated immense application potential in the field of mental health services. Their core value lies in their ability to transform vast amounts of psychological knowledge, treatment cases, and communication skills into scalable, personalized, and readily accessible interactive capabilities. By utilizing structured resources such as authoritative journals, treatment guidelines, and cognitive behavioral therapy (CBT), LLMs can act as a patient and tireless primary care partner, providing 24 / 7 initial emotional assessments, generation of psychoeducational content, and guidance on cognitive restructuring. They can not only collect user status data through multi-turn dialogues to create preliminary psychological profiles and identify core emotional characteristics such as anxiety and depression, but also generate reassuring dialogues, mindfulness practice guidance, or crisis intervention resource recommendations based on evidence-based medicine principles. However, realizing their potential requires a rigorous ethical framework—the model must clearly define its boundaries, avoid providing clinical diagnoses beyond safe limits, and seamlessly connect users to human experts when high-risk signals are identified. In the future, LLM is expected to become a "power multiplier" in the mental health service ecosystem, empowering rather than replacing counselors. By automating document processing, conversation summary generation, and progress tracking, it can significantly improve the efficiency of professionals, thereby enabling high-quality mental health support to reach a wider population.

[0005] 1.2 Psychological Image Generation Technology

[0006] Psychological image generation technology refers to the use of LLM models to automatically generate visual content with specific emotional valence and symbolic meaning based on psychological state assessment results or treatment goals. Its core principle lies in embedding prior psychological knowledge through text prompts—for example, generating images such as a "tranquil forest" or a "stabilizing foundation" for anxiety, and using color psychology and compositional metaphors to guide the user's cognitive emotional regulation. This technology has been proven to build dynamic art therapy tools, helping users externalize and reconstruct their inner experiences through personalized images, and also enabling "visual empathy" within digital mental health platforms, providing exposure contexts or mindfulness focuses for cognitive behavioral therapy.

[0007] 1.3 Multimodal Data Fusion

[0008] This system can construct a three-dimensional data acquisition matrix encompassing visual, auditory, and environmental perception. High-definition cameras capture visual signals such as facial micro-expressions, pupil changes, and clothing details; temperature and humidity sensors and a light monitoring module perceive the environmental atmosphere. Multi-source data undergoes spatiotemporal alignment and feature complementation processing to form a complete emotional data map. In psychological service scenarios, this integration is crucial: the system's powerful data acquisition and analysis capabilities enable it to mimic the comprehensive judgment of a human counselor, converging discrete behavioral signals (such as avoidant eye contact and somber clothing colors) into a coherent "psychological profile." This significantly reduces the risk of misjudgment and provides highly reliable decision-making basis for subsequent personalized generative feedback, ultimately constructing a contextualized AI system that more closely resembles human perception.

[0009] 1.4 Development Trends of Intelligent Psychological Profiling Therapy Technology

[0010] Intelligent psychological profiling technology is becoming more perceptive, considerate, professional, and responsible. Its ultimate goal is to build a more inclusive, efficient, and precise mental health service ecosystem, allowing more people to benefit from the "heart" well-being brought about by technological advancements. Future intelligent psychological profiling will no longer be limited to a single data source. By collaboratively analyzing multimodal data such as text, voice, micro-expressions, gait, physiological signals, and even EEG, the system can more comprehensively capture the complex relationship between an individual's outward behavior and inner psychological state, thereby constructing a more three-dimensional and accurate psychological profile. The core role of large language models in psychological profiling is becoming increasingly prominent. LLM not only processes and analyzes massive amounts of psychological data, but more importantly, its capabilities in affective computing and theory of mind are rapidly evolving. This means that AI is not only "computing" data, but also attempting to "understand" human emotions, beliefs, and intentions, thereby enabling deeper cognitive modeling and more empathetic interaction. As technology deeply penetrates the most private psychological domains, data security, privacy protection, and ethical norms have become core prerequisites for development. Future development must be accompanied by a strict ethical framework to ensure the legality of data collection, protect user privacy using technologies such as federated learning, differential privacy, and data encryption, and clearly define the auxiliary boundaries of AI to prevent the abuse of technology. This is both a constraint on technological development and a guarantee for it to gain social trust and be applied healthily in the long term.

[0011] With the continuous development of artificial intelligence technology, especially the exponential progress in computing power and algorithms, computing capabilities have experienced explosive growth, making it possible to process multimodal data captured by cameras in real time. Simultaneously, breakthroughs in Large Language Model (LLM) cognitive reasoning technology form the technological backbone of this project. This invention constructs a fully localized, multimodal, real-time closed-loop intelligent psychological profiling and visual intervention system by using technologies such as multimodal data fusion perception, psychological profile generation, and personalized visual intervention. Summary of the Invention

[0012] The purpose of this invention is to provide an intelligent psychological profiling and diagnosis system driven by a large language model, so as to solve the problems mentioned in the background art.

[0013] To achieve the above objectives, this invention provides the following technical solution: an intelligent psychological profiling and diagnosis system driven by a large language model. The operation mechanism of the entire system can be divided into four core stages: multimodal data perception and preprocessing stage, cognitive reasoning and generation stage of psychological profiling, decision-making and generation stage of personalized visual intervention, and human-computer interaction and system evolution stage. The specific process is as follows:

[0014] Phase 1: Multimodal Data Sensing and Preprocessing

[0015] The system collects user status information through an intelligent sensing network and sends the raw video stream to the edge computing module. The YOLO-V5n face detection model determines facial features. Then, the collected images are sent to the scene understanding model to classify and extract features of the user's clothing and posture. The above features are then packaged into a structured, time-aligned feature vector sequence. The system parses the heterogeneous data into prompts that can be understood by the LLM and sends them to the QWen3-32B model for deep analysis.

[0016] Phase Two: Cognitive Reasoning and Generation of Mental Profiling

[0017] The QWen3-32B model first uses a visual-language alignment understanding system to match these visual and contextual cues according to an emotion classification system and assess the severity of the state. Then, it infers based on clinical intervention strategies and cases. After the inference is completed, the model generates a structured psychological profile according to the output format set during fine-tuning.

[0018] Phase 3: Decision-making and Generation of Personalized Visual Interventions

[0019] The edge computing module receives the initial suggested direction from the mental profile, and generates prompt words by selecting the most matching visual elements from the rule base based on the specific content of the profile through the dynamic prompt word engineering algorithm. Then, it uses the lightweight StableDiffusion image generation model.

[0020] Phase Four: Human-Computer Interaction and System Evolution

[0021] The system continuously monitors the user's behavior after an interaction, and subsequent interaction data will be recorded as new data points for a new round of fine-tuning of the visual generation model.

[0022] Furthermore, the structured psychological profile typically contains several fixed fields:

[0023] Emotional states: dominant emotion and secondary emotion;

[0024] Significant behavioral features: Identify key visual evidence supporting this judgment;

[0025] Potential causal inference: probability analysis combined with context;

[0026] Preliminary recommendations: Provide supportive suggestions based on a psychology knowledge base.

[0027] Furthermore, the cue words are subject to double constraints before the Stable Diffusion image generation model:

[0028] Safety constraints: The generation of any images that may cause discomfort is strictly prohibited through negative warning words;

[0029] Aesthetic and style constraints: By integrating a specific LoRA model, the style of the output image is controlled to ensure that it conforms to the intervention target and is aesthetically pleasing.

[0030] Furthermore, it comprises the following components:

[0031] The edge computing module is used to build a high-performance edge computing platform with end-to-end localization.

[0032] The intelligent sensing network consists of 4K high-definition cameras and an 8-channel microphone array;

[0033] Human-computer interaction interface, used to display generated images;

[0034] QWen3-32B includes a coordination and control module, a visual analysis module, a psychological query module, a knowledge base, a model service layer, an evaluation and optimization module, and a deployment and maintenance module.

[0035] The coordination and control module is responsible for intent recognition, task scheduling, and result fusion.

[0036] The visual analysis unit is specifically designed to process image and video content, enabling object detection, scene understanding, and visual question answering.

[0037] The Psychological Inquiry Platform focuses on the field of mental health, providing emotion analysis, psychological assessment, and counseling dialogue.

[0038] The knowledge base serves as the system's memory center, integrating vector retrieval and knowledge graphs to provide accurate information retrieval and fact verification.

[0039] The model service layer provides unified management of model inference;

[0040] The evaluation and optimization module continuously monitors system performance;

[0041] The deployment and maintenance module ensures the stable operation of the entire system.

[0042] Compared with the prior art, the beneficial effects of the present invention are:

[0043] 1. Precise psychological profiling and cross-modal assessment

[0044] It enables non-contact psychological state assessment, providing a comprehensive and immediate psychological profile, breaking through the limitations of traditional methods, and eliminating the need for wearing devices or filling out questionnaires.

[0045] 2. Real-time closed-loop intervention

[0046] It achieves a closed loop from psychological state perception to intervention within seconds, providing immediate and immersive psychological support, regulating emotions through visual content, and enhancing user experience.

[0047] 3. Low-cost replication of expert experience

[0048] By embedding expert knowledge into large-scale models, the threshold for psychological assessment can be lowered, promoting the universalization of psychological services and breaking down geographical and resource limitations.

[0049] 4. High-performance integrated solution

[0050] It adopts a high-performance, low-power hardware platform to achieve flexible deployment and long-term operation, making it suitable for a variety of application scenarios.

[0051] 5. Sustainable Evolution of Human-Machine Collaboration

[0052] By building a collaborative ecosystem of "AI screening + expert in-depth treatment," the system is continuously optimized based on user feedback, forming a virtuous cycle of evolution and improving industry efficiency.

[0053] 6. Privacy, security, and privacy

[0054] The system operates entirely locally in a closed loop, ensuring user data privacy and security and eliminating the risk of data leakage. Attached Figure Description

[0055] Figure 1 is a schematic diagram of the system flow of the present invention. Detailed Implementation

[0056] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0057] Please refer to Figure 1. This invention, focusing on the innovative construction and practical application of an "intelligent psychological profiling and visual intervention system based on a locally deployed multimodal large language model," proposes several core technical paths and structural designs. Its innovation lies not only in the significant upgrade and replacement of existing mental health assessment and intervention methods, but also in the system-level technological breakthroughs brought about by cross-domain integration (artificial intelligence + embedded computing + clinical psychology).

[0058] One of the core technologies of this invention is to deeply integrate visually enhanced large language models with psychological reasoning capabilities, and to achieve efficient and secure localized deployment and operation on the resource-constrained edge computing platform RK3588.

[0059] One of the technical challenges lies in enabling complex large-scale vision models to run rapidly on small devices. Large models supporting image recognition and complex inference typically require enormous computing resources and power consumption, making them difficult to run directly on the RK3588 development board (about the size of a mobile phone), let alone meet real-time interactive requirements. By simplifying the model and removing unnecessary parameters, the model becomes smaller and lighter while maintaining its core capabilities. Simultaneously, through numerical compression, the model's parameters are converted from high-precision (FP32) to low-precision (INT4 / INT8), reducing model size and computational load, and improving running speed.

[0060] The second technical challenge was enabling the general-purpose model to possess expert-level psychological reasoning abilities. While general visual-language models can recognize that "a person is crying," they lack the depth to understand the potential causes, severity, and appropriate professional intervention. We integrated comprehensive psychological resources to form a dedicated knowledge base, including classic emotion theories and psychological scale norms. Using this professional knowledge base, we conducted targeted retraining of the model. After training, the entire system achieved a perfect closed loop from "perception" to "understanding" to "action."

[0061] Structured Expression and Visual Generation Mechanism of Mental Profiling

[0062] One of the core technologies of this invention is the seamless integration of the deep reasoning capabilities of a large language model with an image generation system. The model's output is not only a natural language description, but also, through a meticulously designed dynamic mapping system of "psychological state-visual elements," transforms structured profiles into detailed descriptions rich in emotional imagery, thereby driving the generative model to create visual content with precise emotion regulation capabilities. The technical challenge of this mapping process lies in bridging the semantic gap between the abstract nature of psychological concepts and the concreteness of visual elements. For example, how to transform the suggestion "needing warmth and hope" into a concrete image prompt: "At dawn, sunlight penetrates the clouds, shining on newly sprouted green shoots."

[0063] 1. Visual Decoding of Psychological Valence: This invention constructs a "visual element-psychological valence" mapping rule base. This rule base associates psychological concepts with a series of quantifiable visual elements.

[0064] 2. Dynamic Prompt Algorithm: The core of the system includes a prompt generation algorithm. This algorithm can parse keywords (such as emotion type and suggestion direction) in the psychological profile, retrieve matching visual elements from the mapping rule base, and dynamically combine them into a highly detailed, machine-readable natural language prompt to accurately guide the image generation model.

[0065] 3. Constraints on the Safety and Effectiveness of Generated Content: To ensure the positive effects of the intervention, a dual constraint mechanism is introduced during the image generation process. First, aesthetic constraints: technologies such as LoRA or ControlNet are used to control the style, color, and composition of the output image to ensure it aligns with the target emotion. Second, safety constraints: built-in filters strictly prevent the generation of any image content that may evoke discomfort or negative associations, guaranteeing the safety and positive nature of the intervention.

[0066] A framework for assessing mental state through multimodal data perception and fusion

[0067] Psychological state assessment does not rely solely on a single facial expression image, but rather on a comprehensive evaluation through the fusion of multi-channel and multi-dimensional data. Specifically, this invention combines visual data (facial micro-expressions, clothing neatness, color scheme, posture), contextual data (current time, user's historical emotional baseline, recent interaction records), and a psychological knowledge base (emotion theories, population norms, cultural differences) to provide the model with a three-dimensional perceptual input.

[0068] These heterogeneous data, input in a unified format combining visual features, numerical / categorical parameters, and natural language context, achieve a comprehensive fusion of information about the user's psychological state. The key technical challenge lies in organizing this data from different modalities into a structure that the model can understand.

[0069] To address this, this invention employs a multimodal Prompt Engineering approach, designing a structured prompt template. This template organizes visual features, user metadata, and the current context into a text sequence, forming a "context" that the model can understand, thereby guiding the model to mimic expert-level comprehensive reasoning. For example, by inputting information such as "droopy corners of the mouth, loose clothing, and the time being a weekday morning," the model can more accurately infer "fatigue and low mood on a weekday morning" rather than simply "sadness."

[0070] Fully localized closed-loop and highly adaptable privacy and security architecture design

[0071] This invention pays special attention to the system's closed nature and privacy security, ensuring its high reliability and credibility in sensitive mental health scenarios.

[0072] In terms of security, after system deployment, the entire closed loop—from signal perception, data analysis, decision-making reasoning to intervention feedback—is completed automatically on the local device without any manual intervention. All of the user's sensitive data (original images, psychological profiles) begins and ends with the device, completely eliminating the risk of data leakage and misuse, and establishing a cornerstone of user trust.

[0073] In terms of highly adaptable architecture, this invention adopts a modular and loosely coupled design principle. The core Psychological Inference Engine (VLLM), Visual Generation Engine, and Hardware Driver Layer are independent of each other. This makes the system highly portable: on the one hand, the optimized model and algorithm can be deployed on edge platforms of different performance levels (such as Jetson series and other ARM architecture chips); on the other hand, the system supports flexible functional expansion, such as integrating a microphone for voice emotion analysis or connecting to a heart rate wristband to obtain physiological data in the future, without having to reconstruct the entire system. This design enables the technical solution to be applied not only to auxiliary tools in professional counseling rooms, but also to a wide range of scenarios such as smart homes, in-vehicle systems, and mobile devices, promoting the widespread development of intelligent psychological support technology.

[0074] The system of this invention consists of the following core components:

[0075] ① Edge computing module (RK3588)

[0076] ② Intelligent sensing network (4K high-definition camera + 8-channel microphone array)

[0077] ③ Human-computer interaction interface (4K OLED display)

[0078] ④ Inference algorithm (QWen3-32B model conversion)

[0079] The overall architecture of this invention is based on the Rockchip RK3588 high-performance edge computing platform, constructing a fully localized end-to-end intelligent psychological profiling system. Its core uses a quantized version of the QWen3-32B multimodal large model accelerated by an NPU as the "brain," responsible for processing visual data captured by the camera: first, a lightweight YOLO-V5n model driven by the NPU completes face detection and feature extraction; then, the QWen3-32B fuses visual and textual information to generate a structured psychological profile analysis; subsequently, the system automatically constructs emotional prompts based on the analysis results and calls a lightweight Stable Diffusion model, also deployed on the NPU, to generate image feedback in real time that conforms to psychological intervention strategies. Finally, the results are rendered and displayed through a local GUI interface. This architecture fully leverages the heterogeneous computing capabilities of the RK3588 (CPU+NPU+GPU), achieving a low-latency, high-efficiency closed loop from visual perception and cognitive reasoning to image generation while ensuring all sensitive data is processed completely offline and protecting user privacy. The QWen3-32B comprises a coordination and control module, a visual analysis module, a psychological query module, a knowledge base, a model service layer, an evaluation and optimization module, and a deployment and maintenance module. The coordination and control module acts as the system's brain, responsible for intent recognition, task scheduling, and result fusion. The visual analysis module specifically processes image and video content, enabling object detection, scene understanding, and visual question answering. The psychological query module focuses on the mental health field, providing sentiment analysis, psychological assessment, and counseling dialogue. The knowledge base serves as the system's memory center, integrating vector retrieval and knowledge graphs to provide accurate information retrieval and fact-checking. The model service layer uniformly manages model inference. The evaluation and optimization module continuously monitors system performance, and the deployment and maintenance module ensures the stable operation of the entire system.

[0080] Local Deployment and Training of Large Language Models

[0081] To achieve precise perception and personalized intervention in the field of mental health, the core intelligent module of this system is deployed on the high-performance edge computing platform RK3588, employing a simplified and optimized visual version of the QW3 Vision-Enhanced Language Model (VLLM). This visual language model possesses powerful image understanding capabilities, directly analyzing visual information such as facial expressions, micro-expressions, and clothing details to accurately identify and interpret users' psychological states and emotional characteristics. Mental health is influenced by a variety of complex factors, including external environmental stress, internal cognitive patterns, and immediate emotional fluctuations. Therefore, relying on big data, artificial intelligence, and multimodal fusion technologies, especially breakthroughs in facial expression recognition and affective computing, can provide strong support for mental health intervention. By combining visual recognition with the large language model, the system can not only analyze users' real-time emotional states but also perform deep reasoning using a psychological knowledge base to generate personalized psychological profiles and intervention suggestions. Such a system can automatically complete emotion recognition, state analysis, and psychological support without human intervention, generating images with emotional compensation effects for visual intervention, thereby improving the accessibility, timeliness, and effectiveness of mental health support.

[0082] To meet the professional needs of the mental health field, this invention has undergone refined training based on psychological knowledge, building upon the original general language model. This process integrates various professional resources to ensure that the model can deeply understand human emotional and cognitive mechanisms and provide scientifically effective support strategies.

[0083] First, this system integrates a professional psychology knowledge base, including emotion classification systems, psychological scale norms, clinical intervention strategies, and psychotherapy cases. This data provides a solid theoretical foundation for the model, organizing the case experiences and intervention strategies of senior psychological counselors and therapists into structured text. This enables the model to deeply understand human emotional and cognitive mechanisms and generate scientific and effective support strategies and content based on evidence-based principles.

[0084] Secondly, domain-adaptive transfer learning serves as the framework for fine-tuning techniques. Based on a pre-trained general model (QW1.5B), transfer learning is performed using data from the aforementioned psychology domain, enabling the model to smoothly transition from general language capabilities to a professional system with psychological and emotional understanding and support capabilities, thus achieving domain adaptation.

[0085] In this way, the model can not only process users' facial images but also generate warm and professional psychological profiles by combining semantic context, and then generate visual content with emotion regulation functions based on these profiles. The model's performance on mental health tasks is continuously optimized, ultimately achieving an automatic closed loop from emotion recognition to intervention suggestions, providing users with immediate, reliable, and personalized psychological support. This domain-adaptive fine-tuning approach ensures that the model smoothly transforms from a general language model into a professional intelligent system with the ability to understand and support psychological emotions.

[0086] System operation mechanism

[0087] The entire system's operation mechanism can be divided into four core stages: multimodal data perception and preprocessing stage, cognitive reasoning and generation stage of psychological profiling stage, decision-making and generation stage of personalized visual intervention stage, and human-computer interaction and system evolution stage.

[0088] Data acquisition utilizes a 4K high-definition camera, an 8-channel microphone array, temperature and humidity sensors, and light sensors. The YOLO-V5n model (existing) is used for face detection and alignment, and facial movements, clothing colors, and pose features are extracted. Combined with environmental data, feature vectors are constructed.

[0089] The feature vectors are input into the QWen3-32B model, which is fine-tuned using psychological knowledge, for alignment and causal inference. The model outputs emotional state, significant behavioral characteristics, potential causal inferences, and preliminary suggested directions. The entire process is completed by the RK3588 edge computing module. (Innovative Technology)

[0090] Based on the "suggested direction," matching visual elements are retrieved from the mental state-visual element mapping rule base. Detailed image cue words are dynamically generated and input into an NPU-accelerated lightweight Stable Diffusion model.

[0091] The generated images are displayed on a 4K OLED screen. At the same time, the system records the user's interaction behavior, providing data support for subsequent model optimization and personalized adaptation.

[0092] The specific process is as follows:

[0093] Phase 1: Multimodal Data Sensing and Preprocessing

[0094] This stage marks the beginning of interaction with the real world. The goal is to collect user state information with high fidelity and low latency, preparing standardized "feed" for subsequent AI models. When the user enters the device's sensing range (standing in front of the camera), the system automatically triggers the data acquisition process through background subtraction or motion detection algorithms, requiring no active user input. This achieves a truly "seamless" experience, avoiding psychological interference caused by device operation. The 4K high-definition camera begins capturing the user's facial video stream at a rate of multiple frames per second. Simultaneously, environmental sensors (such as temperature, humidity, and light sensors integrated into the device) continuously record current environmental data, providing contextual background for psychological state analysis.

[0095] The acquired raw video stream is fed into the NPU of the RK3588 chip. After the face detection model (YOLO-V5n) determines facial features, it first quickly locates the face region in each frame and performs pose correction and alignment. Next, the facial feature point detection model is activated, accurately identifying dozens of feature points of key areas such as eyebrows, eyes, nose, and lips (e.g., identifying the 15th action unit AU15 indicating a downturned mouth, or AU4 indicating furrowed brows). Then, the acquired image is fed into the scene understanding model, which classifies and extracts features from the user's clothing (color saturation, formality of style), posture (straight / hunched), etc. All these features—from facial action unit values ​​to clothing color HSV values—are packaged into a structured, time-aligned feature vector sequence.

[0096] After that, the system parses the heterogeneous data into prompts that the LLM can understand and sends them into the large model for in-depth analysis.

[0097] Phase Two: Cognitive Reasoning and Generation of Mental Profiling

[0098] The structured cues generated in the previous stage are fed into the fine-tuned QWen3-32B model. The model first performs visual-linguistic alignment understanding. At this point, the knowledge injected into the model during the fine-tuning in the field of psychology begins to play a role. The system matches these visual and contextual cues with its learned "emotion classification system" and refers to "psychological scale norms" to assess the severity of the state (whether it is mild depression or a possible moderate tendency towards depression). Then, causal inference is made based on knowledge of "clinical intervention strategies and cases." For example, it might generate a conclusion like: "A tired and depressed expression on a weekday morning, combined with unkempt clothing, may suggest poor sleep quality or recent high work stress, indicating a certain risk of lack of motivation." After inference, the model generates a structured psychological profile according to the output format set during fine-tuning. This profile typically contains several fixed fields:

[0099] Emotional state: dominant emotion and secondary emotion.

[0100] Significant behavioral features: Identify key visual evidence that supports this judgment.

[0101] Potential causal inference: Probability analysis in conjunction with context

[0102] Preliminary recommendations: Provide supportive suggestions based on a psychology knowledge base.

[0103] This stage involves multi-step chain reasoning based on prior psychological knowledge, and organizes the results into a professional, structured format before passing them on to the next stage.

[0104] Phase 3: Decision-making and Generation of Personalized Visual Interventions

[0105] The system receives the "preliminary suggested direction" from the mental profile and activates the built-in dynamic prompting algorithm. This algorithm accesses a predefined "mental state-visual element" mapping rule base. This rule base integrates color psychology, art therapy theory, and symbolism. Based on the specific content of the profile, the algorithm selects the most matching visual elements from the rule base and combines them into a highly detailed, high-quality prompt. The generated prompt is then fed into the lightweight Stable Diffusion image generation model, also deployed on the RK3588. Before generation, the system imposes two constraints:

[0106] Safety constraints: Strictly prohibit the generation of any images (violence, horror, sadness) that may cause discomfort through negative prompts.

[0107] Aesthetic and stylistic constraints: By integrating specific LoRA models, the style of the output images (such as digital art, watercolor, impressionism) is controlled to ensure that they conform to the intervention objectives and are aesthetically pleasing.

[0108] The generated high-definition images are rendered and displayed on the device's 4K OLED screen. The screen's mirror-like properties allow users to see their own image presented alongside the generated intervention image in a creative way (overlay, juxtaposition, or gradual blending), creating an immersive "dialogue" experience. Simultaneously, the system may provide simple on-screen text guidance, such as, "Please take a deep breath and feel this tranquility," combining visual intervention with simple mindfulness exercises to enhance the intervention's effectiveness.

[0109] Phase Four: Human-Computer Interaction and System Evolution

[0110] The system continuously monitors user behavior after each interaction. For example, after viewing the generated "Tranquil Forest" image, does the user linger and stare, or quickly leave? On their next visit, are their micro-expressions and postures more relaxed than before? This implicit interaction data is recorded as new data points. In the cloud, the system aggregates anonymized data from a large number of anonymous users to analyze which types of visual interventions are most effective for which psychological profile characteristics. These findings can be used to optimize and update the "psychological state-visual element" mapping rule base, and even to fine-tune the visual generation model. Throughout the entire operation, the system consistently adheres to its role as an "assistant." If, during the inference phase, the model identifies extremely high-risk signals (such as an expression of extreme despair accompanied by specific behavioral patterns), its built-in ethical safety safeguards are triggered. Instead of attempting interventions beyond its capabilities, a friendly prompt is displayed on the screen, suggesting the user contact a professional psychologist or call a mental health assistance hotline, thus ensuring a safe and responsible referral.

[0111] This four-stage operating mechanism is interconnected, together forming a privacy-secure, real-time, personalized intelligent psychological support system implemented at the edge.

[0112] 4.2 The Structured Expression and Visual Generation Mechanism of Psychological Profiling

[0113] One of the core technologies of this invention is the seamless integration of the deep reasoning capabilities of a large language model with an image generation system. The output of the large language model is not only a natural language description, but also, through a meticulously designed dynamic mapping system of "psychological state-visual elements," transforms structured profiles into detailed descriptions rich in emotional imagery, thereby driving the generative model to create visual content with precise emotion regulation capabilities. The technical challenge of this mapping process lies in bridging the semantic gap between the abstract nature of psychological concepts and the concreteness of visual elements. For example, how to transform the suggestion "needing warmth and hope" into a concrete image prompt: "At dawn, sunlight penetrates the clouds, shining on newly sprouted green shoots."

[0114] The entire process is a rule-driven, dynamically generated, and multi-constrained automated system, with the specific process as follows:

[0115] 1. Structured Psychological Profile Input and Parsing. The system receives a psychological profile generated by the QWwn3-32B model. This profile is in structured JOSN format or natural language text and contains four key fields: emotional state, salient behavioral features, inference of potential causes, and preliminary suggested directions. By parsing the key fields, the system extracts the core intervention goals (such as "alleviating anxiety") and affective valence (such as "needing warmth").

[0116] 2. Psychological State - Visual Element Mapping Retrieval. The system accesses a pre-built "visual element - psychological valence" mapping rule base. This rule base is built based on psychological theories (such as color psychology, art therapy, and symbolic metaphor), mapping abstract psychological concepts to quantifiable visual elements. Based on the parsed keywords, the system retrieves matching combinations of visual elements from the rule base to form a visual element set. For example, "relieve anxiety + need warmth" is mapped to the visual elements "warm yellow tone, soft lighting + open natural landscape, stable objects + symmetry, low contrast, no sharp objects".

[0117] 3. Dynamic Cue Word Generation. The cue word generation algorithm converts a set of visual elements into a highly detailed, natural language cue word readable by the image generation model. The algorithm ranks the visual elements according to their sentiment valence intensity, then converts the elements into fluent descriptive language, and finally applies style constraints based on system configuration or user history preferences.

[0118] 4. Dual Constraint Filtering Mechanism. Before the prompt words are fed into the image generation model, the system performs safety and aesthetic filtering through a dual constraint module. The system calls a negative prompt word library to prohibit the generation of content that may cause discomfort, such as violence, horror, sadness, and gore. At the same time, by loading a pre-trained LoRA model or ControlNet control network, the style (such as watercolor, impressionism, minimalism), color distribution, and composition structure of the output image are forcibly controlled to ensure that the image meets the intervention target and has visual appeal.

[0119] 5. Lightweight Image Generation Model Inference. The processed prompts are input into a lightweight Stable Diffusion model deployed on an edge device (RK3588 NPU). The model generates high-resolution images based on the prompts, progressively generating images through a diffusion process. Each step is guided by the prompts and constraints, and is accelerated using dedicated NPU operators, ensuring that image generation can be completed within seconds even in resource-constrained environments.

[0120] This invention employs a multimodal Prompt Engineering approach, designing a structured prompt template. This template organizes visual features, user metadata, and the current context into a text sequence, forming a "context" that the model can understand, thereby guiding the model to mimic expert-level comprehensive reasoning. For example, by inputting information such as "droopy corners of the mouth, loose clothing, and the time being a weekday morning," the model can more accurately infer "fatigue and low mood on a weekday morning" rather than simply "sadness."

[0121] This invention employs a fully localized closed-loop and highly adaptable privacy and security architecture design to ensure high reliability and trustworthiness in sensitive mental health scenarios.

[0122] In terms of security, after system deployment, the entire closed loop—from signal perception, data analysis, decision-making reasoning to intervention feedback—is completed automatically on the local device without any manual intervention. All of the user's sensitive data (original images, psychological profiles) begins and ends with the device, completely eliminating the risk of data leakage and misuse, and establishing a cornerstone of user trust.

[0123] In terms of highly adaptable architecture, this invention adopts a modular and loosely coupled design principle. The core Psychological Inference Engine (VLLM), Visual Generation Engine, and Hardware Driver Layer are independent of each other. This makes the system highly portable: on the one hand, the optimized model and algorithm can be deployed on edge platforms of different performance levels (such as Jetson series and other ARM architecture chips); on the other hand, the system supports flexible functional expansion, such as integrating a microphone for voice emotion analysis or connecting to a heart rate wristband to obtain physiological data in the future, without having to reconstruct the entire system. This design enables the technical solution to be applied not only to auxiliary tools in professional counseling rooms, but also to a wide range of scenarios such as smart homes, in-vehicle systems, and mobile devices, promoting the widespread development of intelligent psychological support technology.

[0124] In the specific implementation, a high-performance embedded edge computing platform, RK3588, was first selected as the core computing node of the system. This platform has an 8-core CPU, a 6T NPU, and powerful image processing and local AI model inference capabilities. Based on the RK3588 platform, a multimodal large language model, QWen3-32B, fine-tuned with knowledge from the field of psychology, was deployed. The model uses efficient parameter fine-tuning techniques such as LoRA (Low-Rank Adaptation), and can complete psychological state analysis, emotion recognition, and personalized intervention strategy generation with a response speed of less than 800ms after receiving user facial images and contextual data.

[0125] This large model supports multimodal input, with the input format being a user's facial image plus a structured description of the current situation (e.g., "Time period: weekday morning, recent emotional baseline: stable, interaction history: none"). To enrich the evaluation dimensions, the system can be expanded to integrate several biosignal sensing modules, including a heart rate detection module (e.g., MAX30102) and a skin conductance response sensor (GSR). All sensor data is acquired and preprocessed through the STM32 main control module and uploaded to the RK3588 simultaneously with the images in the form of structured text.

[0126] To meet the needs of personalized visual intervention, this system features a high-precision image generation and display subsystem. Based on the lightweight Stable Diffusion model also deployed on the RK3588 NPU, this subsystem can generate emotionally modulating visual content in real time according to the emotional prompts output by the large language model, and then render and output it through a high-definition display screen.

[0127] The specific workflow is as follows: After the RK3588 platform completes a comprehensive analysis of the user's facial image and contextual data, it outputs a structured psychological profile and suggestions, such as: "Emotional state: Anxiety; Significant features: Frowning, stiff posture; Suggestion: Provide a warm and open environment to alleviate tension." This text is then parsed by a prompt word engineering algorithm, transformed into a detailed image description, and drives an image generation model to create corresponding visual content. Once generated, the system automatically switches the display content to present the user with a customized intervention image.

[0128] Furthermore, to enhance the system's adaptability, various visual intervention modes can be designed, such as a "tranquil natural landscape" mode for anxiety and a "warm and positive imagery" mode for low mood. The entire process can be automatically selected and executed by the model based on real-time psychological profiles.

[0129] In terms of communication architecture, to ensure efficient and stable data interaction between modules within the system, this invention defines a lightweight communication protocol based on local sockets or shared memory to ensure a low-latency closed loop from data acquisition and model inference to image rendering.

[0130] In practical applications, this system can be deployed in psychological counseling rooms, school psychological counseling centers, corporate lounges, or smart home environments as a private psychological support tool. The system's local interactive interface allows users to view historical emotion curves, feedback on intervention effects, and enables professionals to calibrate the model's recommendations.

[0131] In terms of software control logic, the system is divided into 5 main state machine processes:

[0132] 1. Data acquisition status: Periodic or multimodal triggered reading of image and sensor data.

[0133] 2. Psychological profile generation status: Input the data into the VLLM model to complete the psychological state assessment and structured output.

[0134] 3. Cue word construction and image generation status: Dynamically generate image cue words based on the psychological profile and drive the generation model to create visual content.

[0135] 4. Content rendering and display status: Output the generated image to the display screen to complete user feedback.

[0136] 5. Interaction Records and Learning Status: Encrypted and anonymized interaction data is recorded for subsequent model optimization and personalized adaptation.

[0137] The entire system's operation cycle can be triggered in real time or started on a schedule based on user behavior, and intervention strategies and content styles can be dynamically adjusted based on feedback.

[0138] In a typical experiment, the system was deployed in a corporate health management area and ran continuously for 30 days, providing employees with immediate emotional support. The model achieved an accuracy rate of over 88% in recognizing user emotions, and the generated intervention images received positive feedback from users, such as "soothing effect" and "relevant to their current mood," demonstrating good practicality, user acceptance, and privacy security.

[0139] This invention breaks through the static limitations of traditional art installations by employing three innovative pillars: multi-sensor fusion, edge AI computing, and mirrored interactive design. It constructs a complete emotional closed loop of "perception-analysis-generation-interaction." By capturing users' physiological signals in real time, it performs data denoising, feature extraction, and emotional model matching, mapping physiological data to seven basic emotions such as anger, pleasure, and calmness. Finally, it visualizes the emotional state through dynamic artistic brushstrokes, lighting effects, or color gradients. This technical solution has significant advantages in practicality, innovation, stability, and scalability, which can be summarized into the following beneficial effects:

[0140] 1. It achieves deep integration and precise profiling of cross-modal psychological assessment, breaking through the limitations of traditional methods.

[0141] Through the innovative application of the multimodal large language model (QWen3-32B), a deep understanding and semantic interpretation of unstructured visual information (facial expressions, micro-expressions, clothing details) has been achieved, thereby constructing an unprecedented "psychological profiling" capability that is close to the level of human experts. Traditional monitoring of physiological signals such as electrocardiograms and electroencephalograms, while objective, requires physical contact and is difficult to interpret; while traditional questionnaires and scales heavily rely on users' subjective expressions, leading to concealment, misinterpretation, and retrospective bias. This technology opens up a new path for "non-contact objective behavioral analysis."

[0142] The key innovation of this invention lies in its use of cross-modal alignment and reasoning based on "visual-language." The model does not simply classify facial expressions (e.g., identifying them as "sad"), but rather integrates a vast amount of detailed information, including dozens of facial action units (AUs), clothing color and style, and posture, and places it within a rich psychological knowledge context for comprehensive reasoning. For example, the system can not only identify "drooping corners of the mouth" (AU15) and "frowning brows" (AU4), but also, combined with the visual cue of "loose and casual clothing," infer that the user may be in a state of "depressed mood accompanied by a tendency towards self-neglect." This deep association far surpasses traditional algorithms. Furthermore, by employing generative profiling rather than classification-based judgment, the system outputs not cold labels or probability values, but a rich, nuanced, and structured natural language description, much like a counselor verbally recounting their observations and inferences, greatly enhancing the richness and usability of the information.

[0143] This technology provides a 360-degree panoramic view of mental state. This is revolutionary for mental health screening, daily emotion tracking, and even creative industries (such as user experience testing). Users do not need to wear any devices or laboriously fill out questionnaires; they can obtain an instant, comprehensive, and in-depth mental state report simply in a natural state. This makes large-scale, high-frequency mental health assessments possible, providing an extremely efficient tool for the early detection of mental health problems.

[0144] 2. It created a real-time closed-loop intervention mechanism of "assessment-suggestion-generation," enhancing the immediacy and immersive experience of psychological support.

[0145] This system is not a passive analysis tool, but an active intervention system. It creatively combines "psychological profiling" with the reasoning ability of large language models and the generative ability of diffusion models, realizing a closed loop from "perception" to "intervention" in seconds, and providing a brand-new digital psychological support experience.

[0146] Based on the "suggestion" field in the psychological profile, the system automatically transforms it into a detailed description rich in emotion and imagery through prompt word engineering, thereby driving the image generation model to create visual content with specific emotional valence. For example, for anxiety, the generated image might be "a tranquil forest shrouded in morning mist, with soft light and a sense of peace." This personalized content generation based on semantic understanding surpasses traditional methods such as random recommendations or fixed image libraries. Second, visual therapy healing intervention. This solution digitizes and automates principles from psychology such as "visual therapy," "art therapy," and "mindfulness meditation," using carefully generated images to guide the user's emotions and cognition, providing non-verbal, heartfelt emotional support and comfort.

[0147] It provides 24 / 7 instant visual empathy. When a user experiences emotional fluctuations, the system can provide a complete analysis-feedback intervention within seconds, without appointment or waiting. This immediacy is crucial for alleviating acute emotional distress (such as sudden anxiety or stress). The generated images, as a non-verbal form of communication, are more readily accepted by users in an emotional state, offering a more immersive and therapeutic experience than simple textual advice, truly achieving "silence speaks louder than words."

[0148] 3. It has achieved deep integration of domain knowledge and low-cost replication of expert experience, promoting the universalization of psychological services.

[0149] By performing a fine-tuning (LoRA) of the general large model in the field of psychology, this system solidifies expensive and scarce expert knowledge and technology into a software model that can be replicated on a large scale, greatly reducing the threshold for obtaining high-quality psychological assessments.

[0150] Constructing a "digital twin" of expert knowledge. The fine-tuning process not only utilizes a structured psychology knowledge base, but more importantly, incorporates the case experience and intervention strategies of senior consultants. This allows the model to not simply retrieve knowledge, but to mimic the expert's thinking patterns and decision-making processes, making "experience-based" judgments and suggestions, resulting in more practical and humane outputs; simultaneously, the application of efficient parameter fine-tuning techniques. Employing PEFT techniques such as LoRA allows us to quickly equip the model with professional domain capabilities with relatively low computational costs and data volume, avoiding the huge overhead of training a large model from scratch, and providing technical feasibility for rapid iteration and customization (such as customization for different groups such as children and working professionals).

[0151] This technology breaks down geographical and resource limitations. Excellent mental health counselors are concentrated in first- and second-tier cities and are expensive. This system can deploy some of the assessment and initial support capabilities of top experts at extremely low marginal cost to any corner with RK3588 devices, such as schools in remote areas, community centers, break rooms in ordinary companies, and even homes. It acts as a tireless, ubiquitous "primary mental health consultant," capable of large-scale screening, daily emotional care, and immediate initial intervention, thereby benefiting more people and truly promoting the "universalization" and "equalization" of mental health services.

[0152] 4. It provides a highly integrated and low-power all-in-one solution, laying the hardware foundation for large-scale applications.

[0153] This solution selects the RK3588 as the hardware platform and conducts in-depth software and hardware co-optimization, ultimately forming a high-performance, low-power, and highly integrated terminal solution, laying a solid physical foundation for the productization and large-scale promotion of the technology.

[0154] A breakthrough in engineering large-scale model inference has been achieved at the edge. Heterogeneous computing units such as CPUs, NPUs, and GPUs have been coordinated on a single chip to successfully deploy and efficiently run a complex multimodal pipeline, including a visual encoder, a large language model, and an image generation model. This involves a series of complex engineering technologies such as model quantization, operator optimization, memory scheduling, and power consumption and thermal management, demonstrating the feasibility of building complex AI systems on edge devices and providing a valuable practical example for the industry.

[0155] To facilitate deployment and maintenance and suit various scenarios, this terminal device only requires power and internet access (for initial model updates, not data transmission) to operate independently, functioning like a plug-and-play "psychological CT scanner." Its low power consumption allows for continuous operation without the need for expensive server room environments. This integrated form factor enables highly flexible deployment in school counseling rooms, corporate HR departments, hospital waiting areas, nursing home rooms, and even smart homes, significantly expanding the boundaries of application scenarios and making it possible to move from "single projects" to "large-scale deployment."

[0156] 5. A sustainable and evolving human-machine collaboration model has been established, opening a new chapter in the ecosystem of mental health.

[0157] This system is not intended to replace human experts, but rather to build a new human-machine collaborative ecosystem where "AI conducts initial screening and intervention, while human experts focus on in-depth treatment," optimizing resource allocation and enabling the entire ecosystem to continuously evolve.

[0158] The system defines itself as both an "facilitator" and an "amplifier," handling large-scale, routine screening and initial support work. This frees human experts from the burden of tedious initial screening and simple question-and-answer sessions, allowing them to focus more on complex therapeutic processes requiring deep empathy and creativity. Simultaneously, a continuously optimizing flywheel effect is designed. After obtaining user authorization and undergoing rigorous anonymization and aggregation, the system's interaction data (such as which images are most effective at enhancing a particular emotion) can be used for continuous model iteration and optimization. This makes the system increasingly "intelligent" with use, and its interventions increasingly "precise," forming a virtuous cycle of self-evolution.

[0159] This technology enhances the efficiency and quality of the entire mental health industry. For institutions such as hospitals, schools, and businesses, it's like having a tireless "first responder," significantly improving service coverage and response efficiency while reducing overall labor costs. For counselors, it provides powerful tools, improving their work efficiency and the scientific basis of their decisions. For users, it provides more immediate, convenient, and continuous psychological support. Ultimately, this solution has the potential to become a core infrastructure redefining mental health service processes, ushering in a new era of readily accessible high-quality mental health support.

[0160] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from its spirit or essential characteristics. Therefore, the embodiments should be considered illustrative and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be included within the present invention, and no reference numerals in the claims should be construed as limiting the scope of the claims.

[0161] Furthermore, it should be understood that although this specification describes embodiments, not every embodiment contains only one independent technical solution. This narrative style is merely for clarity. Those skilled in the art should consider the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.

Claims

1. A method for intelligent psychological profiling and diagnosis based on a large language model, characterized in that: The entire system's operation mechanism can be divided into four core stages: multimodal data perception and preprocessing, cognitive reasoning and generation of psychological profiles, decision-making and generation of personalized visual interventions, and human-computer interaction and system evolution. The specific process is as follows: Stage 1: Multimodal Data Perception and Preprocessing. The system collects user status information through an intelligent sensing network and sends the raw video stream to an edge computing module. The YOLO-V5n face detection model determines facial features. Then, the collected images are sent to a scene understanding model to classify and extract features from the user's clothing and posture. These features are then packaged into a structured, time-aligned feature vector sequence. The system parses the heterogeneous data into prompts that the LLM can understand and sends them to the QWen3-32B model for deep analysis. Stage 2: Cognitive Reasoning and Generation of Psychological Profiles (QW...) The en3-32B model first uses a visual-language alignment understanding system to match visual and contextual cues according to an emotion classification system and assess the severity of the state. Then, it infers based on clinical intervention strategies and cases. After inference, the model generates a structured psychological profile according to the output format set during fine-tuning. The third stage is the decision-making and generation of personalized visual intervention. The edge computing module receives the initial suggested direction from the psychological profile and generates prompts by selecting the most matching visual elements from the rule base based on the specific content of the profile using a dynamic prompt word engineering algorithm. Then, it generates prompt words through a lightweight StableDiffusion image generation model. The fourth stage is human-computer interaction and system evolution. The system continuously monitors the user's behavioral response after each interaction. Subsequent interaction data is recorded as new data points for a new round of fine-tuning of the visual generation model.

2. The psychological profiling and diagnostic method according to claim 1, characterized in that: The structured psychological profile typically includes several fixed fields: emotional state: dominant and secondary emotions; salient behavioral features: key visual evidence supporting the judgment; potential causal inference: probability analysis combined with context; preliminary suggested directions: supporting suggestions based on a psychological knowledge base.

3. The psychological profiling and diagnostic method according to claim 1, characterized in that: Before the Stable Diffusion image generation model, the cue words are subject to dual constraints: safety constraints: negative cue words are used to strictly prohibit the generation of any images that may cause discomfort; aesthetic and style constraints: by integrating a specific LoRA model, the style of the output image is controlled to ensure that it conforms to the intervention goal and is aesthetically pleasing.

4. A system for performing the psychological profiling and diagnostic method according to claim 1, characterized in that: The system comprises the following components: an edge computing module, which constructs a high-performance, end-to-end localized edge computing platform; an intelligent sensing network consisting of 4K high-definition cameras and an 8-channel microphone array; a human-computer interaction interface for displaying generated images; and the QWen3-32B, which includes a coordination and control module, a visual analysis entity, a psychological query entity, a knowledge base, a model service layer, an evaluation and optimization module, and a deployment and maintenance module. The coordination and control module handles intent recognition, task scheduling, and result fusion. The visual analysis entity specifically processes image and video content, enabling object detection, scene understanding, and visual question answering. The psychological query entity focuses on the field of mental health, providing sentiment analysis, psychological assessment, and counseling dialogue. The knowledge base serves as the system's memory center, integrating vector retrieval and knowledge graphs to provide accurate information retrieval and fact-checking. The model service layer uniformly manages model inference. The evaluation and optimization module continuously monitors system performance. The deployment and maintenance module ensures the stable operation of the entire system.