Artificial intelligence multi-mode technology-based psychological evaluation system and evaluation method for house tree person

Through the House-Tree-Person psychological assessment system based on artificial intelligence multimodal technology, combined with dynamic behavior and static image data, an in-depth fusion analysis of the House-Tree-Person test is achieved, which solves the subjectivity and efficiency problems of traditional assessments and provides a more accurate and efficient psychological state assessment.

CN120600293APending Publication Date: 2025-09-05NINGBO BAOXING INTELLIGENT ENG
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510478646.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

Traditional house-tree-person psychological assessment technology has problems such as strong subjectivity in interpretation, lack of standardization, results easily affected by the subjects' defensive psychology, single information dimension, and low analysis efficiency. Existing artificial intelligence psychological assessment research mostly stays in single modality analysis and lacks in-depth analysis solutions that effectively integrate drawing results and processes.

Method used

The House-Tree-Person psychological assessment system based on artificial intelligence multimodal technology is used. The dynamic behavior data and static image data during the painting process are synchronously captured through the data acquisition module, and the visual features and behavioral features are extracted respectively using convolutional neural networks and long short-term memory networks. A comprehensive psychological state assessment report is generated through the multimodal fusion analysis module.

Benefits of technology

It achieves an objective and in-depth evaluation of the House-Tree-Person test, significantly improves the accuracy, objectivity and efficiency of the assessment, overcomes the limitations of traditional methods, and provides a more comprehensive psychological state assessment tool.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120600293A_ABST
    Figure CN120600293A_ABST
Patent Text Reader

Abstract

The invention discloses an artificial intelligence multi-modal technology-based psychological evaluation system and evaluation method for a house tree person, and aims to solve the technical problems of high subjectivity, low efficiency and single information dimension of traditional HTP evaluation. The system comprises a data acquisition module which is used for synchronously acquiring room tree person drawing image data of a user and a dynamic behavior data flow in a drawing process; the image analysis module is used for analyzing the image data by adopting a first artificial intelligence model to extract static visual features; the behavior analysis module is used for analyzing the behavior data flow by adopting a second artificial intelligence model to extract dynamic behavior characteristics; the multi-modal fusion analysis module is used for carrying out fusion processing on the static and dynamic characteristics to generate multi-modal characteristics; and the psychological assessment report generation module is used for generating an assessment report based on the multi-modal features. According to the method, the objectivity, the accuracy, the efficiency and the assessment depth of psychological assessment are remarkably improved by fusing and analyzing the static content and the dynamic process information of the drawing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of artificial intelligence and psychological assessment technology. Specifically, it relates to a psychological state assessment system and method using multimodal information processing technology. More specifically, it relates to a House-Tree-Person (HTP) psychological assessment system based on artificial intelligence multimodal technology and its implementation method. The present invention can be applied to various scenarios requiring psychological state assessment, such as psychological counseling, educational assessment, career development assessment, and personal self-exploration. Background Art

[0002] The House-Tree-Person (HTP) test, a classic projective technique, is widely used in clinical psychology and counseling practice. The fundamental principle of this technique is to have individuals draw images of specific themes—houses, trees, and people—in order to indirectly explore their underlying personality structure, emotional landscape, underlying conflicts, and how they experience their relationships with their family, environment, and interpersonal networks. The long-standing popularity of the HTP test is largely due to its nonverbal format, relative ease of use, and its perceived ability to access and elicit information at a deep, subconscious level.

[0003] However, a closer examination reveals that traditional HTP tests do face some inherent challenges in practical application. Primarily, the subjectivity of interpretation is a major concern. Assessment results are largely determined by the analyst's professional expertise, clinical experience, and even personal judgment. This often makes it difficult to achieve uniformity and standardization in scoring criteria and reference norms, posing a challenge to the objectivity of the assessment results. Secondly, upon realizing they are in an assessment context, participants may activate psychological defense mechanisms, intentionally or unintentionally embellishing their drawings to project a specific image. While HTPs are designed to mitigate direct verbal defenses, indirect embellishments, such as those manifested in the act of drawing, can still occur, jeopardizing the authenticity of the results. Furthermore, from a scientific and empirical perspective, the reliability and validity of traditional HTP tests are often limited by the lack of large-scale, standardized sample databases and efficient quantitative analysis methods. This makes it difficult to fully guarantee their universality and cross-cultural validity.

[0004] Another notable technical limitation is that traditional analysis methods often focus excessively on the static outcome of the finished painting—the information conveyed by the image itself. In contrast, there is often a lack of systematic attention and effective means to capture the dynamic behavior of subjects during the painting process—such as hesitation, repeated revisions, drawing order, and time distribution—which can contain rich psychological information. This results in a relatively single dimension of information, limiting the depth and breadth of the assessment. Finally, relying solely on manual evaluation and interpretation of HTPs is typically time-consuming and inefficient, making it difficult to meet the modern demand for large-scale psychological screening or rapid assessment services.

[0005] In recent years, the rapid development of artificial intelligence (AI), particularly breakthroughs in computer vision, natural language processing, and machine learning, has opened up new possibilities for the intelligent innovation of psychological assessment tools. Multimodal technology, a key research area in AI, focuses on integrating and processing information from diverse sources or data types (e.g., images, text, speech, physiological signals, behavioral trajectory data), aiming to achieve more comprehensive and reliable understanding and analysis than a single source alone. While some studies have begun to explore the application of AI in psychological assessment, such as identifying emotions or inferring personality traits through analyzing text content, facial microexpressions, or speech prosody, the systematic application of advanced multimodal AI technologies, particularly complex multimodal methods that deeply integrate information about the content and process of a painting, to the classic projective test of house-tree-person (HTPP) to overcome the limitations of these traditional methods, remains in its early stages of exploration. Existing AI-based psychological assessment research often remains limited to single-modality analysis. Even when multimodal data is involved, few specifically address the specific challenge of integrating content and process analysis in HTP tests. In particular, there is still a lack of technical solutions that can simultaneously capture static image features and dynamic behavior sequence features and effectively integrate them to improve evaluation accuracy, objectivity and efficiency.

[0006] Therefore, how to effectively use artificial intelligence and multimodal fusion technology to build a psychological assessment system and corresponding methods that can achieve objective, efficient and in-depth analysis of house-tree-person paintings, especially to effectively integrate information on both the final result of the painting and the painting production process, thereby significantly improving the accuracy, objectivity and application efficiency of the assessment, constitutes a key technical issue that needs to be overcome urgently in the current intersection of mental health technology and artificial intelligence. Summary of the Invention

[0007] The present invention aims to overcome the technical problems existing in existing House-Tree-Person (HTP) psychological assessment technologies, such as strong subjectivity in interpretation, lack of standardization, susceptibility of results to the defensive psychology of subjects, single information dimension, and low analysis efficiency, and to provide a more objective, accurate, efficient and in-depth House-Tree-Person psychological assessment system and assessment method based on artificial intelligence multimodal technology.

[0008] To solve the above technical problems, the present invention provides a house-tree-person psychological assessment system based on artificial intelligence multimodal technology, the design of which includes:

[0009] A data collection module, whose function is to not only receive the image data of the house-tree-person drawing completed by the user through the interactive interface, that is, the static drawing result, but also synchronously and in real time collect the dynamic behavior data stream of the user throughout the entire drawing process to capture detailed information about the drawing process;

[0010] an image analysis module configured to analyze the house-tree-person painting image data using a first artificial intelligence model to extract static visual features that reflect the content of the painting, such as the structure, layout, lines, and colors of elements in the image;

[0011] a behavior analysis module configured to analyze the dynamic behavior data stream using a second artificial intelligence model, specifically for processing the behavior data generated during the drawing process to extract dynamic behavior features that can reflect the state of the user's drawing process, such as the rhythm, strength, hesitation or modification of the drawing;

[0012] A key multimodal fusion analysis module, which is configured to effectively fuse the aforementioned static visual features with dynamic behavioral features to generate a multimodal feature that can more comprehensively and three-dimensionally represent the user's comprehensive psychological state. This feature also contains information about the "result" and "process" of the painting;

[0013] and a psychological assessment report generation module, which is configured to automatically generate a structured and easy-to-understand house-tree-person psychological assessment report based on the multimodal features and further combined with a preset psychological knowledge base stored in the system.

[0014] In a specific implementation, the first artificial intelligence model is preferably a convolutional neural network or a variant thereof. Leveraging its advantages in image recognition, the model is suitable for identifying and quantitatively analyzing static visual features in HTP paintings.

[0015] In another specific implementation, the second artificial intelligence model is preferably a long short-term memory network or a variant thereof. This model is good at processing the time series characteristics of the dynamic behavior data stream and can effectively extract dynamic behavior features reflecting changes in psychological state.

[0016] In order to accurately capture the painting process, the dynamic behavior data stream may specifically include at least one data item selected from the following group: pen tip coordinate sequence, stroke timestamp, pen pressure value, pen tip tilt, painting pause time and position, number and range of erasing actions, and the drawing order and duration of each element.

[0017] According to a further embodiment of the present invention, the multimodal fusion analysis module may adopt a feature layer fusion strategy or a decision layer fusion strategy to perform the fusion processing, aiming to integrate the information advantages of the two modalities and generate a more reliable basis for psychological state assessment.

[0018] In another embodiment, in order to make the assessment report more readable and practical, the psychological assessment report generation module can use natural language processing technology to convert the internal analysis results into fluent and natural text, and generate a comprehensive house-tree-person psychological assessment report that includes an overview of the psychological state, risk warnings and suggestions.

[0019] In order to improve the convenience and experience of users, in some embodiments, the data acquisition module may also include additional canvas management functions, such as supporting multiple canvas creation, layer management, etc.

[0020] At the same time, in order to assist users to better complete the painting task, in some embodiments, the data acquisition module may also be configured with a painting assistance function, such as providing guidance or prompt information.

[0021] Accordingly, the present invention also provides a House-Tree-Person psychological assessment method based on artificial intelligence multimodal technology. This method is characterized by the following core steps: first, receiving user-drawn House-Tree-Person image data through an interactive interface and simultaneously capturing dynamic behavioral data streams during the drawing process; then, analyzing the image data using a first artificial intelligence model to extract static visual features, and using a second artificial intelligence model to analyze the behavioral data stream to extract dynamic behavioral features; then, fusing the extracted static visual features with the dynamic behavioral features to generate a comprehensive multimodal feature set; and finally, generating a House-Tree-Person psychological assessment report based on this multimodal feature set and a pre-set psychological knowledge base.

[0022] In a preferred embodiment, the method employing the first artificial intelligence model for analysis may employ a convolutional neural network or a variant thereof; and employing the second artificial intelligence model for analysis may employ a long short-term memory network or a variant thereof. Furthermore, the dynamic behavior data stream underlying the analysis should include at least one key piece of information: brush pressure, drawing speed, pause patterns, and erasing behavior.

[0023] Through the aforementioned technical solution, this invention leverages multimodal artificial intelligence technology to innovatively integrate and analyze the static image content of a painting with the dynamic painting process, achieving an objective and in-depth assessment of the House-Tree-Person test. This significantly improves the accuracy, objectivity, efficiency, and information dimensionality of the assessment, effectively overcoming the limitations of existing technologies and providing a valuable intelligent assessment tool for the field of mental health services. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 This is a schematic diagram of the system of Example 1. DETAILED DESCRIPTION

[0025] In order to more clearly illustrate the purpose, technical solutions and advantages of the present invention, the following will be described in detail through specific embodiments. It should be understood that these embodiments are only used to illustrate several implementations of the present invention and are not intended to limit the entire scope of the present invention. At the same time, if there is no conflict, the technical features of each embodiment can be combined with each other.

[0026] Example 1.

[0027] The core of the embodiments of the present invention is to provide a house-tree-person psychological assessment system and assessment method based on artificial intelligence multimodal technology.

[0028] The system combines the analysis of the final image content of the user's house-tree-person (HTP) drawing with the capture of dynamic behavior during the drawing process, aiming to achieve a more objective, accurate, efficient and in-depth psychological state assessment experience.

[0029] like Figure 1 As shown, in this embodiment, the system is deployed on a computing device, typically equipped with a processor, memory, and a user interface for interaction, such as a tablet computer equipped with a stylus or a touch-enabled computer. The system's operation relies on the close collaboration of several key internal modules, including: data acquisition, image analysis, behavior analysis, multimodal fusion analysis, and finally, the psychological assessment report generation module that presents the final results.

[0030] The data collection module forms the front end of user interaction with the system. It not only provides a virtual canvas for users to draw images of houses, trees, and people, but also plays a key role in comprehensively recording relevant data. Users can create multiple independent canvas spaces as needed, draw houses, trees, and people separately, or choose to draw them all on the same canvas.

[0031] The system allows users to name and categorize canvases, which is very helpful for subsequent long-term personal status tracking or comparative analysis between different assessments. The settings of the painting tools strive to simulate the real experience, providing options that imitate the brushstroke effects of different physical media such as pencils and markers, as well as eraser tools for modification. It also has an intuitive and easy-to-use color palette, and the brush thickness can also be adjusted according to the user's needs. For the smoothness of the interaction, the system often supports common gestures, such as pinching to zoom and rotate the canvas, or quickly performing undo or redo through specific gestures (such as three-finger sliding). The layer management function gives users greater flexibility, allowing different elements to be placed on independent layers, which can be easily reviewed, modified, or adjusted to appear and disappear separately without interfering with other parts of the picture.

[0032] Another core function of the data acquisition module is the synchronous and precise capture of the dynamic behavioral data stream during the drawing process. As the user's fingertip or stylus moves across the screen, leaving a trace, the system records a series of detailed behavioral parameters in the background. This generates a rich data stream, including the precise trajectory of the stylus tip's movement, namely the stylus tip coordinate sequence; along with high-precision timestamp information, recording the start, end, and duration of each stroke, forming the stroke timestamp. On pressure-sensitive hardware devices, subtle changes in the user's applied pressure are also captured, reflecting the amount of force applied and even the intensity of emotion. Even for more advanced devices, such as the Apple Pencil, changes in the stylus tip's tilt relative to the screen can be recorded as indirect behavioral clues. Crucially, any hesitation, deliberation, or pauses during the creative process are reflected as data on the duration and location of drawing pauses; and any modifications made are quantified by the number and extent of erasing actions. Furthermore, the user's chosen drawing order (houses first, trees second, and people last) and the time spent on each element are also fully recorded. This combination of parameters paints a dynamic picture of the user's drawing behavior, providing insights beyond the static image for subsequent in-depth analysis.

[0033] In order to better serve different users, the data acquisition module can also be configured with a drawing assistance function in some implementation forms. Users can choose to turn on auxiliary lines such as grid lines, horizontal lines, etc., or call some basic shape templates, such as the basic outline of a house or human body proportion lines for reference. This will be very helpful for users who have little drawing experience or need structural guidance. In some embodiments, the system can give some gentle prompts in a timely manner based on preset rules during the drawing process, such as providing guiding suggestions for common structural defects, such as the lack of roots in drawn trees or the imbalance of character proportions. Of course, these aids and prompts are optional, and their original intention is to assist rather than force. The primary purpose of the entire system and method is to ensure that the user's spontaneous completion is true expression, and the system will not interfere too much with the user's spontaneous expression.

[0034] After the drawing task is completed, this module is responsible for securely storing the final house-tree-person drawing image data and the complete dynamic behavior data stream. This typically involves data encryption and supports cloud synchronization and local backup mechanisms to prevent data loss. Furthermore, to facilitate user retention and interaction with other applications, the system supports exporting the drawing image to common image file formats such as PNG or SVG.

[0035] The image analysis module focuses on processing the finalized image data of the house-tree-person painting transmitted by the data acquisition module. Its core is the first artificial intelligence model, which is used to deeply analyze these static images and extract the static visual features contained therein, including element structure, layout, line quality, and color application. In a typical implementation, the model uses convolutional neural network (CNN) technology. It should be noted that the convolutional neural network here does not refer to a specific fixed structure, but refers to a type of deep learning network that uses convolution operations to extract spatial hierarchical features of images. Its actual application may include the classic CNN architecture or more advanced variants, such as residual networks (ResNet) and densely connected networks (DenseNet), which have been proven to perform well in image analysis tasks.

[0036] Trained on a large amount of psychologically annotated HTP image data, this first AI model can automatically identify and quantify key visual elements in an image. For example, it can analyze the structural features of a house and link them to the user's psychological state. For example, the shape of the roof may be related to an individual's openness or defensiveness, while the size and proportion of doors and windows may reflect their sociability. It can also interpret the form of trees. For example, the density of the crown may be associated with emotional fullness or repression, the thickness of the trunk may symbolize the strength of vitality, and the presence of roots may touch upon subconscious feelings of security. Similarly, the posture and details of a person's body—whether the arms are extended or retracted—may reflect initiative, the depiction of facial expressions relates to self-perception, and the harmony of body proportions is related to self-esteem. Furthermore, the first AI model analyzes the overall layout of elements in the image, the quality of line work (smooth and confident, or hesitant and fragmented), and the choice and application of color. All of these extracted visual features are ultimately converted into structured data, laying the foundation for subsequent fusion analysis. This automated and quantitative feature extraction is undoubtedly a big step forward in objectivity compared to traditional interpretation that relies on subjective experience.

[0037] The behavior analysis module operates in parallel with the image analysis module. It is specifically responsible for interpreting the dynamic behavior data stream recorded by the data acquisition module. Given the inherent time series properties of this data stream, some specific system implementations utilize a second AI model to process it. This model is typically a long short-term memory (LSTM) network or its functionally similar variants, such as the gated recurrent unit (GRU). The core advantage of these models lies in their ability to effectively process sequential information and capture long-term dependencies in the data, which is precisely suited to analyzing the dynamic changes in the painting process. By learning from a large number of painting behavior sequences and their associated psychological state labels, the second AI model is able to deeply understand the time series characteristics of the dynamic behavior data stream and extract meaningful dynamic behavior features from it. These features often reveal the user's immediate psychological state fluctuations during the creation process. The second artificial intelligence model can identify significant changes in painting speed and associate these changes with emotional fluctuations or the fluency of thinking; it can quantify the continuous pattern of brushstroke pressure, which can be used to analyze the intensity of emotions or the degree of inner tension; it can interpret pause patterns, such as repeated long pauses in a specific area of ​​the picture. This dynamic behavioral feature may imply inner conflict or uncertainty; it can also analyze the pattern of erasing behavior. Frequent and large-scale erasing may point to anxiety, insecurity or excessive perfectionism. Furthermore, the order in which the user draws each element can also be interpreted by the model as a reflection of the inner priority of different psychological themes. It is the extraction of these dynamic behavioral features that enables this technical solution to obtain more data on observing psychological processes compared to the evaluation system or method of the existing technology, and provides information that the static picture itself cannot fully display.

[0038] The multimodal fusion analysis module is the most important innovation in this system's design. Its significance lies in the fact that the current psychological state is not determined solely by the content of the image or the act of drawing, but rather by the complex interaction of the two. Therefore, this module receives static visual features from the image analysis module and dynamic behavioral features from the behavior analysis module, and applies a fusion processing strategy to organically combine these two different sources and properties of information. Its goal is to generate a more comprehensive, three-dimensional, and accurate multimodal feature representation of the user's overall psychological state.

[0039] There are various possible designs for specific fusion strategies. Some implementations employ feature-level fusion, such as concatenating a vector representing visual features with a vector representing behavioral features, and then feeding this fused, longer vector into a subsequent classification or regression model. Alternatively, attention-based fusion can be employed, allowing the model to dynamically learn how much weight to assign to visual and behavioral cues when assessing different psychological indicators. Other implementations may employ decision-level fusion, first generating preliminary assessments based on static and dynamic features, and then integrating these preliminary assessments using ensemble learning methods (such as weighted voting or model stacking) to arrive at a final conclusion. Regardless of the strategy employed, the core idea is to leverage the synergistic effects of multimodal information. For example, if a static image shows a region with chaotic lines (possibly indicating anxiety), while behavioral data shows frequent erasures and unusual pauses when drawing that region, multimodal fusion can significantly enhance confidence in the anxiety assessment. This fusion analysis not only brings about the superposition of information, but also significantly improves the robustness and depth of the evaluation results through cross-validation and complementarity.

[0040] The psychological assessment report generation module, the final step in the system and assessment methodology, transforms complex analysis results into information that users or professionals can understand and use. This module receives the comprehensive multimodal features carefully constructed by the multimodal fusion analysis module and combines them with a psychological knowledge base stored within the system. This knowledge base is typically quite rich, containing, for example, a large number of validated association rules between HTP image features and psychological indicators. It also incorporates the application framework of classic psychological theories (such as Jungian archetypes and Maslow's hierarchy of needs) in HTP interpretation, and even uses machine learning to uncover deep patterns from a large number of real-world case studies. Based on this information, the module ultimately generates a detailed and structured house-tree-person psychological assessment report.

[0041] To make the report flow naturally and be easily understood, in many implementations, this module utilizes natural language processing (NLP) technology. Using large-scale language models such as GPT or BERT, fine-tuned to the style of psychological assessment reports, it can "translate" the numerical and symbolic analysis results into human language. The generated report typically consists of several sections: first, an overview of the user's current psychological state (such as emotional stability, stress level, and interpersonal communication patterns), supplemented by a concise rating or score in some implementations. Second, potential risk warnings are identified. For example, if the system finds that certain feature combinations strongly indicate depressive tendencies, high anxiety, potential aggression, or unprocessed trauma, these warnings will be provided. Next, the report delve into details, analyzing the symbolic meaning of houses, trees, and people, as well as their interactions within the image. Finally, and most importantly, based on the assessment results, the report provides personalized improvement recommendations, which may include recommendations for seeking further professional counseling, specific and feasible self-adjustment methods, or identifying patterns that may need attention in interpersonal relationships. The goal of this report is to be both insightful and compassionate, and to provide guidance.

[0042] This embodiment also provides an implementation of a house-tree-person psychological assessment method based on artificial intelligence multimodal technology. The method includes the following steps:

[0043] S1: Multimodal Data Acquisition

[0044] This method uses a system interface to perform house-tree-person (HTP) drawing. During this phase, the system performs data acquisition tasks: receiving and storing the user's completed HTP image data (static recording); simultaneously, during the user's drawing process, it captures behavioral details in real time and synchronously, forming a dynamic behavioral data stream. This dynamic behavioral data stream records the pen tip coordinate sequence, stroke timestamp, pen pressure value (obtained when hardware supports it), pen tip tilt, drawing pause time and position data, and the number and range of erase actions. Furthermore, the drawing order and duration of each element are also recorded in this dynamic behavioral data stream. The final image is recorded synchronously with the creative process behavior, providing a data foundation for subsequent multimodal analysis.

[0045] S2: Static Visual Feature Extraction

[0046] After acquiring the painting data, the method enters the analysis phase. One step involves analyzing the image data of the house-tree-person painting. The system invokes a first artificial intelligence model (in some implementations, a convolutional neural network (CNN) or its variants) to process the static image. This model identifies and quantifies the visual elements in the image and their organization, extracting static visual features. These features encompass the structural form, spatial layout, use of lines, and color selection of the houses, trees, and people in the painting, forming a description of the image's content.

[0047] S3: Dynamic Behavior Feature Extraction

[0048] The method analyzes the collected dynamic behavior data stream. This step uses a second artificial intelligence model (in some implementations, a long short-term memory (LSTM) network or its variants) to analyze the time series characteristics of the dynamic behavior data stream. This model is used to analyze temporal patterns in the behavior data and extract dynamic behavior features. The model can quantify and analyze drawing speed, brush pressure patterns, the frequency and duration of pause patterns, and the pattern and location of erasing behaviors. These dynamic features are used to characterize the user's immediate psychological state during the creation process, such as emotional intensity, confidence, or anxiety level.

[0049] S4: Multimodal Feature Fusion

[0050] The subsequent step is fusion processing. This step integrates the static visual features generated by S2 with the dynamic behavioral features generated by S3. The system adopts a fusion processing strategy, which can include feature layer fusion (such as splicing two feature vectors or using an attention mechanism for weighted combination) or decision layer fusion (such as first making a preliminary judgment based on the features of a single modality and then integrating these judgments) to generate a multimodal feature representation. The fused feature set contains information from both the content of the picture and the creative process, and is used to provide a more comprehensive representation of the psychological state than a single modality analysis. For example, the line features (static features) of a certain area of ​​the picture are combined with the erasure and pause features (dynamic features) when drawing the area, which can be used to enhance the confidence in the judgment of a specific psychological state (such as anxiety).

[0051] S5: Psychological Assessment Report Generation

[0052] This step generates an assessment result based on the multimodal features generated by S4 and combined with the system's built-in psychological knowledge base. This knowledge base stores relevant psychological theoretical models, HTP interpretation rules, and case studies. Based on this information, the system generates a structured house-tree-person psychological assessment report. Some implementations use natural language processing (NLP) technology to convert the analysis results into text. The report content includes an overview of the user's overall psychological state, potential risk warnings, detailed interpretation of the painting, and personalized recommendations.

[0053] By executing steps S1 to S5 above, the method of the present invention realizes a multi-dimensional analysis of the house-tree-person test, thereby improving the objectivity, efficiency and information richness of the evaluation.

[0054] It's important to note that user data security and privacy are paramount throughout the design and implementation of this system. All collected data, whether images or behavioral streams, is encrypted during storage and transmission. During model training and large-scale data analysis, the system should support and prioritize anonymization technologies to ensure the proper protection of personally identifiable information and adhere strictly to relevant data protection regulations, such as GDPR, and professional ethical standards in the mental health field.

[0055] The system and method disclosed by the present invention have broad application prospects. In the field of psychological counseling, it can be a powerful assistant for counselors, quickly forming a preliminary impression of the visitor's status during the first interview, and assisting in formulating a counseling plan. In the educational assessment scenario, it can be used to conduct regular screening of the mental health status of student groups, and to detect individuals who may have emotional distress or developmental disorders as early as possible so as to intervene in time. In organizational and human resource management (workplace assessment), it can be used to assess employees' stress adaptability, teamwork style, and even assist in judging job matching. Of course, it can also be used as a convenient personal self-exploration tool to help ordinary users, in a private environment, enhance their understanding and awareness of their own inner world through this classic projection method.

[0056] Those skilled in the art will appreciate that the foregoing descriptions are merely preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention. The ultimate scope of protection of the present invention is defined by the appended claims.

Claims

1. A house-tree-person psychological assessment system based on artificial intelligence multimodal technology, characterized by: include: A data acquisition module is used to receive the house-tree-person drawing image data drawn by the user through the interactive interface, and synchronously collect the dynamic behavior data stream during the drawing process; an image analysis module configured to analyze the house-tree-person painting image data using a first artificial intelligence model to extract static visual features reflecting the content of the painting; a behavior analysis module configured to analyze the dynamic behavior data stream using a second artificial intelligence model to extract dynamic behavior features reflecting the painting process; a multimodal fusion analysis module configured to fuse the static visual features with the dynamic behavioral features to generate a multimodal feature representing the user's comprehensive psychological state; The psychological assessment report generation module is configured to generate a house-tree-person psychological assessment report based on the multimodal features and in combination with a preset psychology knowledge base.

2. The system according to claim 1, wherein: The first artificial intelligence model is a convolutional neural network, which is used to identify static visual features in painting images, including element structure, layout, line quality and color application.

3. The system according to claim 1 or 2, characterized in that The second artificial intelligence model is a long short-term memory network, which is used to process the dynamic behavior characteristics of the dynamic behavior data stream, including time series characteristics, and extract painting speed, brush pressure changes, pause patterns, erasing behavior and drawing order.

4. The system according to claim 1, wherein: The dynamic behavior data stream includes at least one data selected from the following group: pen tip coordinate sequence, pen stroke timestamp, pen stroke pressure value, pen tip tilt, painting pause time and position, number and range of erasing actions, and drawing order and duration of each element.

5. The system according to claim 1, wherein: The multimodal fusion analysis module adopts a feature layer fusion strategy or a decision layer fusion strategy to integrate the static visual features and dynamic behavioral features into a more comprehensive psychological state representation.

6. The system according to claim 1, wherein: The psychological assessment report generation module uses natural language processing technology to convert the multimodal features into a text report containing an overview of the psychological state, potential risk warnings and personalized suggestions.

7. The system according to claim 1, wherein: The data acquisition module also includes a canvas management function that supports multiple canvas creation, layer management, and export of painting content.

8. The system according to claim 1, wherein: The data acquisition module is also equipped with a drawing assistance function, which provides element guidance prompts or basic shape templates.

9. A house-tree-person psychological assessment method based on artificial intelligence multimodal technology, characterized by: The following steps are involved: Receiving image data of a house-tree-person drawing drawn by a user through an interactive interface, and synchronously collecting dynamic behavior data streams during the drawing process; Analyzing the house-tree-person painting image data using a first artificial intelligence model to extract static visual features reflecting the content of the painting; Using a second artificial intelligence model to analyze the dynamic behavior data stream to extract dynamic behavior features reflecting the painting process; fusing the static visual features with the dynamic behavioral features to generate a multimodal feature representing the user's comprehensive psychological state; Based on the multimodal features and combined with a preset psychology knowledge base, a house-tree-person psychological assessment report is generated.

10. The method according to claim 9, characterized in that The step of using the first artificial intelligence model for analysis specifically includes using a convolutional neural network or a variant thereof for analysis; the step of using the second artificial intelligence model for analysis specifically includes using a long short-term memory network or a variant thereof for analysis; and the dynamic behavior data stream includes at least one of brush pressure, painting speed, pause mode and erasing behavior.

Citation Information

Cited By

  • Artificial intelligence psychological assessment method and system based on multi-modal input

    CN121287146A

  • Interactive multi-modal artificial intelligence psychological behavior assessment method and system based on drawing analysis

    CN121891007A