House-tree-person test analysis method and device based on multimodal time series data
Through multimodal time series data collection and feature fusion, the problems of labor expenditure and analysis limitations of traditional house-stories tests are solved, and more efficient and objective psychological assessment is achieved.
Patent Information
- Application Number
- CN202510714676.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2045-05-30
AI Technical Summary
The traditional house tree test relies on manual observation and static image analysis, which consumes a lot of manpower and time, lacks the collection and analysis of multimodal information such as dynamic behavior, expressions, and voice, making it difficult to meet the needs of large-scale psychological assessment.
Multimodal time series data acquisition, including video, audio, handwriting and physiological information, features are extracted through CNN-LSTM and Transformer models, cross-modal consistency scores are calculated and the characteristics are fused to generate comprehensive analysis results.
It improves the comprehensiveness and objectivity of test analysis, reduces the dependence on manual observation, and provides more comprehensive and reliable psychological assessment results.
Smart Images

Figure CN120277395B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of data processing technology, and in particular to a house-tree-person test analysis method and device based on multimodal time series data. Background Art
[0002] The House-Tree-Person (HTP) psychological projective test reveals a subject's psychological state by analyzing their drawings of houses, trees, and people. In traditional HTP testing, psychological evaluators primarily collect data from the drawing phase and subsequently interpret the data.
[0003] This process primarily relies on psychological evaluators observing the subject's drawing process and psychological behavior in real time, which consumes significant manpower and time, making it difficult to meet the needs of large-scale psychological assessments. Furthermore, the analysis process primarily relies on the evaluator's observation notes and the final drawing results. The data dimension is limited to static image content, lacking the collection and analysis of multimodal information such as dynamic behavior, facial expressions, voice, and intonation during the drawing process. Summary of the Invention
[0004] The embodiments of the present disclosure provide a house-tree-person test analysis method and device based on multimodal time series data.
[0005] In a first aspect, an embodiment of the present disclosure proposes a House-Tree-Person Test analysis method based on multimodal time series data, comprising: obtaining multimodal data generated by a target subject during a House-Tree-Person Test; performing feature extraction on the multimodal data based on the time series to obtain multimodal features; splicing the multimodal features to obtain a splicing vector; calculating a cross-modal consistency score based on the splicing vector, and fusing it with the splicing vector to obtain feature fusion data; and analyzing the feature fusion data to obtain an analysis result.
[0006] In a second aspect, an embodiment of the present disclosure proposes a house-tree-person test analysis device based on multimodal time series data, comprising: a multimodal data acquisition module, configured to acquire multimodal data generated by a target subject during a house-tree-person test; a multimodal feature extraction module, configured to extract features from the multimodal data based on the time series to obtain multimodal features; a splicing vector generation module, configured to splice the multimodal features to obtain a splicing vector; a feature fusion data generation module, configured to calculate a cross-modal consistency score based on the splicing vector, and fuse it with the splicing vector to obtain feature fusion data; and an analysis result generation module, configured to analyze the feature fusion data to obtain an analysis result.
[0007] In a third aspect, an embodiment of the present disclosure provides an electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that when the at least one processor executes the instructions, it is possible to implement the house-tree-person test analysis method based on multimodal time series data as described in any implementation method of the first aspect.
[0008] In a fourth aspect, an embodiment of the present disclosure provides a non-transitory computer-readable storage medium storing computer instructions, which are used to enable a computer to implement the house-tree-person test analysis method based on multimodal time series data as described in any implementation method of the first aspect when executed.
[0009] In a fifth aspect, an embodiment of the present disclosure provides a computer program product comprising a computer program, which, when executed by a processor, can implement the house-tree-person test analysis method based on multimodal time series data as described in any implementation manner in the first aspect.
[0010] The House-Tree-Person Test analysis method and device based on multimodal time series data provided by the embodiments of the present disclosure collects time series data of multimodal data during the House-Tree-Person drawing process. Based on the structure and analysis points of the House-Tree-Person Test, the time series data of different elements in the drawing process are identified and encoded. This allows for a comprehensive analysis of the subject's drawing sequence, layout, symbolic features, etc., cross-modally verifying the test results and improving the comprehensiveness, efficiency, and objectivity of the test analysis.
[0011] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Other features, objects and advantages of the present disclosure will become more apparent from a reading of the detailed description of non-limiting embodiments made with reference to the following drawings:
[0013] Figure 1 is an exemplary system architecture of the present disclosure;
[0014] Figure 2 A flowchart of a first house-tree-person test analysis method based on multimodal time series data provided by an embodiment of the present disclosure;
[0015] Figure 3 A flowchart of a second house-tree-person test analysis method based on multimodal time series data provided by an embodiment of the present disclosure;
[0016] Figure 4A flowchart of a third House-Tree-Person test analysis method based on multimodal time series data provided by an embodiment of the present disclosure;
[0017] Figure 5A A flowchart of a fourth House-Tree-Person test analysis method based on multimodal time series data provided by an embodiment of the present disclosure;
[0018] Figure 5B A flowchart of a fifth house-tree-person test analysis method based on multimodal time series data provided by an embodiment of the present disclosure;
[0019] Figure 6 A flowchart of a sixth house-tree-person test analysis method based on multimodal time series data provided by an embodiment of the present disclosure;
[0020] Figure 7 A structural block diagram of a house-tree-person test analysis device based on multimodal time series data provided by an embodiment of the present disclosure;
[0021] Figure 8 A schematic structural diagram of an electronic device suitable for executing a house-tree-person test analysis method based on multimodal time series data, provided by an embodiment of the present disclosure. DETAILED DESCRIPTION
[0022] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those skilled in the art that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description. It should be noted that the embodiments in the present disclosure and the features in the embodiments can be combined with each other unless there is a conflict.
[0023] In the technical solutions disclosed herein, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0024] Figure 1 An exemplary system architecture 100 is shown to which embodiments of the disclosed multimodal time series data-based house-tree-person test analysis method, apparatus, electronic device, and computer-readable storage medium can be applied.
[0025] like Figure 1As shown, system architecture 100 may include terminal devices 101, 102, 103, a network 104, and a server 105. Network 104 is a medium for providing communication links between terminal devices 101, 102, 103 and server 105. Network 104 may include various connection types, such as wired or wireless communication links or fiber optic cables.
[0026] Users can use terminal devices 101, 102, 103 to interact with server 105 via network 104 to receive or send messages, etc. Terminal devices 101, 102, 103 and server 105 may be installed with various applications for enabling information communication between them, such as instant messaging applications.
[0027] Terminal devices 101, 102, 103 and server 105 can be either hardware or software. When terminal devices 101, 102, 103 are hardware, they can be various electronic devices with display screens, including but not limited to smartphones, tablet computers, laptop computers, and desktop computers. When terminal devices 101, 102, 103 are software, they can be installed in the electronic devices listed above. They can be implemented as multiple software or software modules, or as a single software or software module, and are not specifically limited here. When server 105 is hardware, it can be implemented as a distributed server cluster consisting of multiple servers, or as a single server. When the server is software, it can be implemented as multiple software or software modules, or as a single software or software module, and are not specifically limited here.
[0028] Server 105 can provide various services through various built-in applications. The data or information required by these applications to provide these services can be obtained from terminal devices 101, 102, and 103 via network 104, or can be pre-stored locally on server 105 in various ways. Therefore, when server 105 detects that this data is already stored locally, it can choose to directly obtain it from the local location. In this case, exemplary system architecture 100 may also exclude terminal devices 101, 102, and 103 and network 104.
[0029] Because acquiring various types of data or information and analyzing and processing them may require a significant amount of computing resources and significant computing power, the House-Tree-Person Test analysis methods based on multimodal time series data provided in the subsequent embodiments of this disclosure are generally executed by a server 105 possessing significant computing power and resources. Accordingly, the House-Tree-Person Test analysis apparatus based on multimodal time series data is also generally located in the server 105. However, it should also be noted that, if the terminal devices 101, 102, and 103 also possess sufficient computing power and resources, the terminal devices 101, 102, and 103 may also perform the various operations previously assigned to the server 105 through the corresponding applications installed thereon, thereby outputting the same results as those of the server 105. In particular, when there are multiple terminal devices with different computing capabilities, but the terminal device where the corresponding application is determined to have stronger computing capabilities and more remaining computing resources, the terminal device can be allowed to perform the aforementioned operations, thereby appropriately reducing the computing pressure on server 105. Accordingly, the house-tree-person test analysis device based on multimodal time series data can also be installed in terminal devices 101, 102, and 103. In this case, exemplary system architecture 100 may also not include server 105 and network 104.
[0030] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.
[0031] Please refer to Figure 2 , Figure 2 This is a flowchart of a House-Tree-Person test analysis method based on multimodal time series data provided by an embodiment of the present disclosure, wherein process 200 includes the following steps:
[0032] Step 201: Acquire multimodal data generated by a target subject during the House-Tree-Person Test.
[0033] This step is intended to be performed by the subject of the House-Tree-Person test analysis method based on multimodal time series data (e.g. Figure 1 The server 105 shown in the figure obtains data generated by a target subject during the HTP test. The House-Tree-Person (HTP) psychological projective test analyzes the subject's drawings of houses, trees, people, and their accessories to reveal their psychological state. The target subject primarily refers to the subject taking the HTP test. Multimodal data refers to data generated by the target subject during the HTP test, containing information in multiple modalities, such as images, text, audio, video, and sensor data.
[0034] In this embodiment, the multimodal data mainly includes:
[0035] 1. Video information: While the subject is drawing, use a high-definition camera or mobile device to capture the subject's facial expressions, gaze direction, body posture, hand movements, and the corresponding drawing content in real time.
[0036] 2. Audio Information: Using a microphone, the subject's speech content, voice characteristics (volume, pitch, speaking speed, pauses, etc.), and ambient sound are collected. Audio content collected includes the subject's own speech and conversations with the psychological evaluator, such as descriptions of the painting and its meaning, and explanations of specific details in the painting.
[0037] 3. Handwriting information: If drawing is done with pen and paper, a camera or high-definition camera can be used to capture the order and movements of the subject's strokes during the drawing process, as well as the corresponding drawing elements and moments. For example, the path and speed of the pen tip movement when drawing a certain element, the line strength, pauses, erasures, corrections, modifications, etc., and the times at which they occur. If drawing is done with a pressure-sensitive digital tablet or tablet, the pressure, pen speed, trajectory coordinates, pauses, erasures, corrections, etc., and the times at which they occur, can be recorded in real time when drawing a certain element.
[0038] 4. Physiological information: Use a multichannel biosensor or wearable device such as a smartwatch or ring to collect physiological parameters such as heart rate, respiration, electrodermal conduction, electroencephalogram (EEG), and electromyography (EMG).
[0039] In this embodiment, by collecting multimodal time series data such as video, audio, handwriting, and physiological signal lights in real time during the painting process, the subject's expression, movement, language, and physiological projection can be captured more comprehensively, greatly improving the ability to capture key nodes of the painting, allowing evaluators or automated systems to upgrade from "single static picture" to "full dynamic observation", and to explore richer psychological clues in multi-dimensional information.
[0040] In the actual process of conducting the House-Tree-Person test, preparation and guidance for the test need to be carried out in advance:
[0041] 1. Pre-test information collection: Collect basic demographic information of the test-taker (gender, age, education level, occupation, medical history, reasons for participating in the House-Tree-Person Test, etc.), and mark the specific purpose of the "House-Tree-Person Test" (such as diagnostic assistance, educational assessment, or clinical research) during registration.
[0042] (1) Prepare the tools for the standard House-Tree-Person drawing test, including A4 paper, pencil, eraser, etc.; if electronic drawing is used, prepare a digital tablet / tablet and adjust the pen strokes to pencil strokes;
[0043] (2) If multimodal data acquisition equipment is required, preliminary debugging is performed (camera, microphone, handwriting device, physiological equipment, etc.).
[0044] 2. Professional guidance and process tips:
[0045] (1) Inform the subject of the requirements for the House-Tree-Person Drawing Test. For example, "Please use a pencil to draw a picture containing a house, a tree, or a person on the A4 white paper in front of you. As long as it contains these three elements, the specific shape of the picture is up to you. I hope you can draw what you think in your heart."
[0046] (2) Inform the subject of things to avoid, such as “do not sketch or copy.”
[0047] (3) If the subject has recently taken the House-Tree-Person Test according to standardized instructions, alternative instructions should be given to avoid a decrease in reliability and validity due to repeated testing.
[0048] Step 202: Extract features from the multimodal data based on the time series to obtain multimodal features.
[0049] This step aims to extract features from multimodal data based on time series by the above-mentioned execution subject to obtain multimodal features.
[0050] In some optional implementations of this embodiment, such as Figure 3 As shown, the process mainly includes:
[0051] Step 301: Time series segmentation is performed on the multimodal data according to a unified time window to obtain time series segmentation data.
[0052] In this step, the execution entity divides all modal data into a unified time window (such as 1 second, 2 seconds or 5 seconds), or uses a variable length window according to specific needs to ensure that the segmentation boundaries of all modalities are aligned on the same time axis.
[0053] Step 302: Extract emotion features based on the video information and audio information in the time-series segmented data to obtain emotion feature information.
[0054] Step 303: Match the emotional feature information with the preset segments in the House-Tree-Person test process.
[0055] During this process, the execution entity uses CNN-LSTM (a deep learning model architecture that combines a convolutional neural network (CNN) and a long short-term memory (LSTM) network) to extract emotional time series data, including facial expressions, voice intonation, and speech rate. The CNN layer extracts local time-frequency features of facial expressions or audio within a short time window. The LSTM layer captures the temporal dependencies of emotional changes over a longer time frame. The data is then segmented into pre-set segments based on the time when the house, tree, person, and accessory objects were drawn, and the emotional state labels are matched to the segments. Pre-set segments can include, for example, the pre-test phase, the house phase, the tree phase, the person phase, and the accessory phase.
[0056] Step 304: Calculate the target subject's cooperation feature information in the preset segments during the House-Tree-Person Test based on the emotion feature information, video information, and audio information.
[0057] During this process, the executor can use Transformer (a groundbreaking architecture in the field of deep learning that mainly relies on the self-attention mechanism to process sequence data and is mainly used in multiple fields such as natural language processing and computer vision) or self-attention mechanism to encode the subject's emotions, posture, movements, language and other interaction patterns during the entire drawing test and the interaction with the psychological assessor, calculate the subject's cooperation characteristics in different time segments, form time series data and perform segmented matching.
[0058] Step 305: Perform static feature extraction and dynamic feature extraction based on the video information and handwriting information to obtain painting feature information.
[0059] Step 306: Match the drawing feature information with the preset segments in the House-Tree-Person test process.
[0060] In this step, the processes implemented by the above-mentioned execution entities mainly include:
[0061] (1) Use CNN (convolutional neural network) to extract static features of images: static features include the content and features of the picture after the house-tree-person drawing test is completed, including the recognition of picture content elements and their features (house, tree, person, accessory elements and their features, such as windowless house, broken tree trunk, etc.), picture perspective, size and proportion, position layout, picture cut-off, shadow, perspective, symmetry, omission, richness and other features.
[0062] (2) Use LSTM (Long Short-Term Memory Network) to process the handwriting trajectory or the time series records of drawing speed, pressure, etc., and extract the time-dependent dynamic features of the drawing action: the dynamic features include but are not limited to the following features: the overall drawing time of the subject during the drawing process, the drawing time of each element of the house, tree, person and accessories, etc.; the overall drawing order of the house, tree, person and accessories and the specific drawing order of each element, such as drawing the house first, then the tree, and finally the person; drawing the face first, then the torso, etc.; the starting and ending time of the strokes during the drawing process, the speed change curve, the pause time change, etc.; the distribution of handwriting pressure, the area of uneven force, the thickness / density of the lines during the drawing process; the traces of repeated corrections, erasures or modifications during the drawing process.
[0063] Step 307: Extract features from the physiological data based on the preset segments in the House-Tree-Person test process to obtain physiological feature information.
[0064] In this step, the execution entity can use a 1D-CNN (a variant of a convolutional neural network (CNN) primarily used to process one-dimensional sequence data) to process continuous physiological signals. The preprocessed physiological signals are segmented and convolutional computations are performed to capture local temporal patterns and extract changes in physiological features. These include, but are not limited to, fluctuations in heart rate variability (HRV), electrodermal (GSR) activity, and electroencephalogram (EEG) activity over time during the drawing process. These fluctuations are saved as time series data and segmented for matching.
[0065] Step 203: Splice the multimodal features to obtain a spliced vector.
[0066] This step aims to combine the extracted multimodal features into a vector at the same time step by the execution subject, such as [emotional features, cooperation, handwriting speed, handwriting pressure, HRV, GSR, stage...].
[0067] Step 204: Calculate the cross-modal consistency score based on the splicing vector, and fuse it with the splicing vector to obtain feature fusion data.
[0068] In this embodiment, whether the modalities such as emotion, physiology, and handwriting are consistent or abnormal at the same time is quantified, a cross-modal consistency score is calculated, and the data after feature fusion is embedded. The specific calculation method of the cross-modal consistency score includes: for multiple key modes at the same time step, the mode is defined as a positive mode or a negative mode, and each mode outputs an activation value. The larger the value of the positive mode, the more active and abnormal it is; the opposite is true for the negative mode. Set a threshold, if it exceeds the threshold, it is defined as a high activation state, and if it is below the threshold, it is defined as a low activation state. If at the same moment, the proportion of the number of the same activated mode to the total number is used as the consistency score, then when the proportion is greater than 50%, it is recorded as high cross-modal consistency.
[0069] Based on this, the team uses cross-attention to fuse features, using painting features as the dominant modality. This involves leveraging a key modality to proactively query time-series segments from other modalities to identify the most relevant behavioral or emotional changes. This cross-attention mechanism generates a new set of features (such as physiological fluctuations associated with painting behavior), which are then embedded into the fused multimodal data for subsequent analysis and decision-making.
[0070] In this embodiment, the dominant modality can be selected as the painting feature, and emotional and physiological features can be associated with the painting feature through a cross-attention mechanism. By calculating the similarity between the painting feature and the emotional and physiological features, attention weights are assigned to the emotional and physiological features. Based on these assigned attention weights, the painting feature information, emotional and physiological feature information, and cross-modal consistency scores are fused to generate a feature fusion data output that comprehensively considers all information. For example, when a drawing of a tree trunk shows repeated erasing, the model will look for related fluctuations in the emotional and physiological signals, helping to cross-infer whether the subject may be experiencing psychological conflict or anxiety.
[0071] Exemplarily, the output may include: outputting the fused multimodal time series features at each moment, where each time step includes:
[0072] (1) Multimodal fusion vector: contains emotions (such as anxiety and happiness), handwriting (speed and pressure), physiology (HRV and GSR), and cooperation at each time step;
[0073] (2) Cross-modal consistency score: As an independent dimension, it indicates whether multimodal high synchronization occurs at this moment;
[0074] (3) Stage labels: such as “house stage”, “tree stage”, “person stage”, etc. (text or discrete encoding), which help to distinguish different painting sub-processes;
[0075] (4) Timestamp / timing alignment: Ensure that the features of each modality are aligned and output on the same time axis.
[0076] Step 205: Analyze the feature fusion data to obtain analysis results.
[0077] This step aims to enable the execution subject to analyze the feature fusion data from multiple dimensions and obtain corresponding analysis results to determine the various states of the target object during the House-Tree-Person Test.
[0078] The House-Tree-Person Test analysis method based on multimodal time series data provided by the disclosed embodiments collects multimodal time series data from the House-Tree-Person drawing process. Based on the structure and analysis key points of the House-Tree-Person Test, it identifies and encodes the time series data of different elements in the drawing process. This allows for a comprehensive analysis of the test subject's drawing sequence, layout, and symbolic features, cross-modally validating the test results and improving the comprehensiveness, efficiency, and objectivity of the test analysis. Furthermore, through multimodal data collection, a unified timeline, and automated emotion recognition, the method reduces over-reliance on manual observation and provides objective cross-modal data verification and support for key points such as "whether the house has no windows" and "peak anxiety when drawing people," making the analysis results more objective, consistent, and repeatable.
[0079] In this embodiment, in order to ensure the accuracy of the acquired multimodal data, the acquired data may also be preprocessed. Figure 4 , Figure 4 A flowchart of a house-tree-person test analysis method based on multimodal time series data provided by an embodiment of the present disclosure, namely, Figure 2 Step 201 in the process 200 shown provides a specific implementation method. The other steps in the process 200 are not adjusted. The specific implementation method provided in this embodiment is replaced by step 201 to obtain a new complete embodiment. The process 400 includes the following steps:
[0080] Step 401: Obtain the original data generated by the target subject during the House-Tree-Person Test.
[0081] Step 402: normalize and reduce noise on the original data to obtain first preprocessed data.
[0082] In this step, the above-mentioned execution entity preprocesses the multiple collected raw data, such as frame rate normalization, denoising, lighting correction and face detection of video information; facial key point annotation of faces; segmentation of speech and silent audio segments, noise reduction processing, and extraction of speech-related features (timbre, volume, speaking speed, pauses, etc.); difference or downsampling processing of handwriting information; cleaning abnormal strokes (such as stroke delay and excessive jitter) and converting them to a unified coordinate system; filtering, denoising and baseline correction of physiological signals, and removing motion artifacts or environmental noise.
[0083] Step 403: Perform time synchronization processing on the first pre-processed data to obtain time synchronization data.
[0084] In this step, the execution entity aligns heterogeneous data sources, such as video, audio, handwriting, and physiological data, using a unified time base (e.g., the system clock) to form a multivariate time series. For example, alignment can be performed based on specific events (e.g., the start and end of a sound, the moment a pen touches down, etc.).
[0085] Step 404: Use timestamps and tags to annotate the time-synchronized data to obtain multimodal data.
[0086] In this step, the execution entity annotates key events during the test subject's drawing process and stores them in the form of timestamps and tags. For example, the start and end times of the test subject's drawing of different elements such as the house, tree, person, and accessories are annotated; key moments such as the subject's pen placement / lifting, corrections, and significant emotional or verbal fragments are automatically identified and annotated; and moments when the psychological assessor provides prompt guidance, such as "Can you add some more details?" or when the test subject proactively says, "I want to add more windows to the house."
[0087] The multimodal data collected are preprocessed through the above process, which can eliminate interference from some noise signals and synchronize the multimodal data in time, facilitate subsequent processing based on the data, and improve the accuracy of the data.
[0088] In this embodiment, the above step 205, the process of analyzing the feature fusion data and obtaining the analysis results, can be analyzed from multiple angles and dimensions. Figure 5A , Figure 5A A flowchart of a house-tree-person test analysis method based on multimodal time series data provided by an embodiment of the present disclosure, namely, Figure 2 Step 205 in the process 200 shown provides a specific implementation method. The other steps in the process 200 are not adjusted. The specific implementation method provided in this embodiment is replaced by step 205 to obtain a new complete embodiment. The process 500a includes the following steps:
[0089] Step 501a: Analyze the time period during which at least one of the painting feature information, the emotion feature information, and the physiological feature information exceeds a preset data threshold based on the painting feature information, the emotion feature information, and the physiological feature information.
[0090] During this process, the execution entity analyzes whether at least one of the drawing feature information, emotional feature information, and physiological feature information exceeds a preset data threshold and records the time period during which the threshold is exceeded. For example, when a physiological signal (such as heart rate or galvanic skin response) exceeds a preset normal range (fixed threshold / ratio) for a certain period of time, the execution entity can mark this period. The time series features of this period exceeding the physiological data threshold are then extracted to determine the drawing stage, the elements drawn, the detailed features of the elements, and the dynamic characteristics of the drawing behavior for subsequent analysis.
[0091] Step 502a: Based on the time period, the painting feature information, the emotional feature information, and the physiological feature information are compared and analyzed to obtain analysis results.
[0092] After determining that a threshold has been exceeded, the system combines the drawing, emotional, and physiological characteristics of that period for comparative analysis to generate results. For example, using a preset threshold, the system automatically identifies "unusual passages" or "significant emotional peaks" to match relevant drawing content. For example, if a subject experiences 15 consecutive seconds of high anxiety and GSR activity (T = 100-150s), indicating they are drawing a "human face," the analysis can link the psychological intention of face drawing to anxiety.
[0093] Please refer to Figure 5B , Figure 5B A flowchart of a house-tree-person test analysis method based on multimodal time series data provided by an embodiment of the present disclosure, namely, Figure 2 Step 205 in the process 200 shown provides another specific implementation. The other steps in the process 200 are not adjusted. The specific implementation provided in this embodiment is replaced by step 205 to obtain a new complete embodiment. The process 50b includes the following steps:
[0094] Step 501b: Based on the cross-modal consistency score, analyze the distribution of the painting feature information, the emotional feature information, and the physiological feature information when the cross-modal consistency is high.
[0095] During this process, the execution entity analyzes the distribution of consistency across multiple modalities (physiological data, emotional fluctuations, drawing behavior, etc.) to extract relevant drawing features (for example, focusing on periods of high cross-modal consistency). Within these periods, features related to the drawing behavior, such as handwriting speed, pressure, and pause duration, as well as characteristics of the drawing phase (such as drawing a "person" face or drawing "house" details), can be extracted for in-depth analysis. For example, high cross-modal consistency between repeated erasures and heart rate fluctuations can further confirm that the subject is experiencing anxiety. If this occurs while drawing a house, it can be inferred that the subject may be anxious about family-related issues.
[0096] Step 502b: Analyze based on the distribution and the preset House-Tree-Person Projective Test knowledge base to obtain analysis results.
[0097] In this embodiment, the pre-set House-Tree-Person projective test knowledge base primarily includes: ① Theoretical knowledge of the House-Tree-Person test: including the static and dynamic characteristics and cooperation level of each element (house, tree, person, and other components), and their corresponding psychological interpretations; ② The subject's emotional and physiological state during the House-Tree-Person test, and their corresponding interpretations, as shown in Tables 1 and 2.
[0098] Table 1
[0099]
[0100] Table 2
[0101]
[0102] By analyzing the obtained distribution and the corresponding drawing features, the analysis conclusion corresponding to the drawing features can be searched in the preset House-Tree-Person Projective Test knowledge base to obtain the analysis results.
[0103] In some optional implementations of this embodiment, step 205, analyzing the feature fusion data to obtain analysis results, may also include performing cluster analysis on the feature fusion data using unsupervised learning methods to obtain cluster analysis results. In this embodiment, the execution entity clusters multimodal data using unsupervised learning methods (such as K-means and DBSCAN), which can reveal similarities and differences between different groups of subjects. For example, different group patterns may exist between subjects' drawing behaviors (such as handwriting pressure and speed) and emotional fluctuations (such as anxiety and joy), revealing the psychological characteristic patterns of specific groups in the House-Tree-Person Test, thereby expanding the understanding of traditional theories.
[0104] In some optional implementations of this embodiment, step 205, analyzing the feature fusion data to obtain analysis results, may also include: using a pre-trained autoencoder to perform anomaly detection on the feature fusion data based on preset regular behavior pattern data to obtain analysis results. In this embodiment, the execution entity may use autoencoders and other technologies to model regular behaviors based on the temporal characteristics of multimodal data, thereby automatically detecting abnormal emotional patterns or painting behaviors. These abnormal patterns may be new connections between emotional reactions and painting performance. This process mainly includes:
[0105] Step 1: Train the autoencoder: The autoencoder can be trained to learn the correlation patterns between normal emotional fluctuations and drawing behavior. The autoencoder learns how to compress input data (such as emotional fluctuations, handwriting characteristics, heart rate, GSR, etc.) and map it into a latent space so that it can reconstruct the original input. If the model can reconstruct the data well, it indicates that the input data is normal.
[0106] Step 2: Anomaly Detection: When new input data differs significantly from the autoencoder model's reconstruction, it indicates that the data contains an unusual pattern. For example, if a painting session includes an extremely asymmetrical house drawing accompanied by unusual emotional fluctuations (such as extreme anxiety or joy), the autoencoder may not be able to reconstruct this data well, indicating that this data segment has an unusual pattern.
[0107] Step 3: Identifying Potential Relationships Between Emotions and Behaviors: Anomaly detection can uncover new relationships between emotional fluctuations and drawing behaviors, such as the connection between certain emotional fluctuations (e.g., depression and anxiety) and repeatedly modified or extremely asymmetrical elements in drawings (e.g., houses or trees). For example, certain emotional fluctuations (e.g., depression or excessive joy) may be associated with unusual drawing behaviors (e.g., extremely asymmetrical houses or abstract trees). Anomaly detection can uncover these potential connections between abnormal behaviors and emotional fluctuations.
[0108] Please refer to Figure 6 , Figure 6 This is a flow chart of another voice interaction method provided by an embodiment of the present disclosure, wherein process 600 includes the following steps:
[0109] Step 601: Acquire multimodal data generated by the target subject during the House-Tree-Person Test.
[0110] Step 602: Extract features from the multimodal data based on the time series to obtain multimodal features.
[0111] Step 603: Splice the multimodal features to obtain a spliced vector.
[0112] Step 604: Calculate the cross-modal consistency score based on the splicing vector, and fuse it with the splicing vector to obtain feature fusion data.
[0113] Step 605: Analyze the feature fusion data to obtain analysis results.
[0114] The above steps 601-605 are similar to the following Figure 2 Steps 201-205 shown are consistent. For the same content, please refer to the corresponding part of the previous embodiment and will not be repeated here.
[0115] Step 606: Generate an analysis report based on the analysis results.
[0116] In this step, the execution entity can generate an analysis report based on the analysis results obtained in any of the above-mentioned embodiments. In specific implementations, the execution entity can further adjust the report content, language style, etc. using a large language model prompting project to output an expert report (for clinicians, other psychological assessment and counseling staff) interpreting the House-Tree-Person Projective Test and a subject report (for the subject, their family members, teachers, etc.).
[0117] For example, the expert report could focus on: Based on the model's output (including video and handwriting playback), combined with the House-Tree-Person test's semiotics and clinical knowledge base, the model's analysis could be tailored to the individual's age, background, and past medical history, highlighting key points for further evaluation or interviews. Initial clinical intervention recommendations could be provided, combined with input from psychological questionnaires or interviews.
[0118] The test-taker's report can focus on: Using language accessible to non-professionals, explaining the overall characteristics of the House-Tree-Person structure and its comprehensive interpretation; providing appropriate mental health advice if the report indicates significant emotional distress or decreased cooperation. It should be noted that this test is not intended to be diagnostic, but rather a supplementary assessment.
[0119] Further references Figure 7 As an implementation of the methods shown in the above figures, the present disclosure provides an embodiment of a house-tree-person test analysis device based on multimodal time series data. Figure 2 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.
[0120] like Figure 7As shown, the House-Tree-Person Test analysis device 700 based on multimodal time series data of this embodiment may include: a multimodal data acquisition module 701, a multimodal feature extraction module 702, a splicing vector generation module 703, a feature fusion data generation module 704, and an analysis result generation module 705. The multimodal data acquisition module 701 is configured to acquire multimodal data generated by a target subject during the House-Tree-Person Test; the multimodal feature extraction module 702 is configured to perform feature extraction on the multimodal data based on the time series to obtain multimodal features; the splicing vector generation module 703 is configured to splice the multimodal features to obtain a splicing vector; the feature fusion data generation module 704 is configured to calculate a cross-modal consistency score based on the splicing vector and fuse it with the splicing vector to obtain feature fusion data; and the analysis result generation module 705 is configured to analyze the feature fusion data to obtain an analysis result.
[0121] In this embodiment, in the House-Tree-Person Test Analysis Device 700 based on multimodal time series data, the specific processing of the multimodal data acquisition module 701, the multimodal feature extraction module 702, the splicing vector generation module 703, the feature fusion data generation module 704 and the analysis result generation module 705 and the technical effects thereof can be referred to respectively. Figure 2 The relevant descriptions of steps 201-205 in the corresponding embodiment are not repeated here.
[0122] This embodiment exists as an apparatus embodiment corresponding to the above-mentioned method embodiment. This embodiment provides a house-tree-person test analysis device based on multimodal time series data. By collecting the time series of multimodal data during the house-tree-person drawing process, it identifies and encodes the time series data of different elements in the drawing process based on the structure and analysis points of the house-tree-person test, so as to comprehensively analyze the subject's drawing order, layout, symbol features, etc., verify the test results cross-modally, and improve the comprehensiveness, efficiency, and objectivity of the test analysis.
[0123] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can implement the house-tree-person test analysis method based on multimodal time series data described in any of the above embodiments when executing.
[0124] According to an embodiment of the present disclosure, the present disclosure further provides a readable storage medium storing computer instructions, which are used to enable a computer to implement the house-tree-person test analysis method based on multimodal time series data described in any of the above embodiments when executed.
[0125] According to an embodiment of the present disclosure, the present disclosure further provides a computer program product, which, when executed by a processor, can implement the house-tree-person test analysis method based on multimodal time series data described in any of the above embodiments.
[0126] Figure 8 A schematic block diagram of an example electronic device 800 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are provided as examples only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0127] like Figure 8 As shown, device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. RAM 803 may also store various programs and data required for the operation of device 800. Computing unit 801, ROM 802, and RAM 803 are interconnected via a bus 804. An input / output (I / O) interface 805 is also connected to bus 804.
[0128] Various components in device 800 are connected to I / O interface 805, including an input unit 806, such as a keyboard, mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a magnetic disk, optical disk, etc.; and a communication unit 809, such as a network card, modem, wireless communication transceiver, etc. The communication unit 809 allows device 800 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0129] The computing unit 801 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 performs the various methods and processes described above, such as the House-Tree-Person Test analysis method based on multimodal time series data. For example, in some embodiments, the House-Tree-Person Test analysis method based on multimodal time series data can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the House-Tree-Person Test analysis method based on multimodal time series data described above can be performed. Alternatively, in other embodiments, the computing unit 801 may be configured in any other appropriate manner (eg, by means of firmware) to execute the house-tree-person test analysis method based on multimodal time series data.
[0130] Various embodiments of the systems and techniques described above can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0131] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0132] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fibers, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0133] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0134] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0135] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact via a communication network. This client-server relationship arises through computer programs running on the respective computers and establishing a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or cloud host. This is a hosting product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosting and virtual private server (VPS) services.
[0136] According to the technical solution of the embodiments of the present disclosure, by collecting time series of multimodal data during the house-tree-person drawing process, the structure and analysis points of the house-tree-person test are targeted, and the time series data of different elements in the drawing process are identified and encoded to comprehensively analyze the subject's drawing order, layout, symbol features, etc., to verify the test results cross-modally and improve the comprehensiveness, efficiency, and objectivity of the test analysis.
[0137] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not a limitation herein.
[0138] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.
Claims
1. A House-Tree-Person Test analysis method based on multimodal time series data, characterized in that: The method comprises: Obtain multimodal data generated by the target subject during the House-Tree-Person Test; Extracting features from the multimodal data based on a time series to obtain multimodal features; Splicing the multimodal features to obtain a splicing vector; Calculating a cross-modal consistency score based on the splicing vector and fusing it with the splicing vector to obtain feature fusion data; Analyzing the feature fusion data to obtain analysis results; The multimodal data includes: video information, audio information, handwriting information and physiological data; The extracting features of the multimodal data based on the time series to obtain multimodal features includes: Performing time series segmentation on the multimodal data according to a unified time window to obtain time series segmentation data; Extracting emotion features based on the video information and audio information in the time series segmentation data to obtain emotion feature information; Matching the emotional characteristic information with the preset segments in the house-tree-person test; Calculating cooperation feature information of the target object in a preset segment during the House-Tree-Person test based on the emotion feature information, video information, and audio information; Performing static feature extraction and dynamic feature extraction based on the video information and handwriting information to obtain painting feature information; Matching the drawing feature information with the preset segments in the house-tree-person test; The physiological data is subjected to feature extraction based on the preset segments in the house-tree-person test process to obtain physiological feature information.
2. The House-Tree-Person Test analysis method based on multimodal time series data according to claim 1, characterized in that: The step of obtaining multimodal data generated by the target subject during the House-Tree-Person Test includes: Obtain the original data generated by the target subject during the House-Tree-Person Test; Performing normalization and noise reduction processing on the original data to obtain first preprocessed data; Performing time synchronization processing on the first preprocessed data to obtain time synchronization data; The time-synchronized data is annotated using a timestamp and a tag to obtain the multimodal data.
3. The House-Tree-Person Test analysis method based on multimodal time series data according to claim 1, characterized in that: The static feature extraction and dynamic feature extraction based on the video information and handwriting information to obtain the painting feature information includes: Using a convolutional neural network to extract image static features from the video information to obtain static features; Using a long short-term memory network to extract dynamic features from the handwriting information to obtain temporal dynamic features; The painting feature information is formed based on the static features and the temporal dynamic features.
4. The House-Tree-Person Test analysis method based on multimodal time series data according to claim 1, characterized in that: The calculating the cross-modal consistency score based on the splicing vector and fusing it with the splicing vector to obtain feature fusion data includes: Taking the painting feature information as the dominant modality, the emotional feature information, the physiological feature information and the painting feature information are associated through a cross-attention mechanism; Calculating the similarity between the painting feature information and the emotional feature information and the physiological feature information, and assigning attention weights to the emotional feature information and the physiological feature information; Based on the assigned attention weights, the painting feature information, the emotional feature information, the physiological feature information and the cross-modal consistency score are fused to obtain the feature fusion data.
5. The House-Tree-Person Test analysis method based on multimodal time series data according to claim 1, characterized in that: The analyzing the feature fusion data to obtain analysis results includes: Analyzing, based on the painting feature information, the emotional feature information, and the physiological feature information, a time period in which at least one of the painting feature information, the emotional feature information, and the physiological feature information exceeds a preset data threshold; Based on the time period, the painting feature information, the emotion feature information, and the physiological feature information are compared and analyzed to obtain the analysis result.
6. The House-Tree-Person Test analysis method based on multimodal time series data according to claim 1, characterized in that: The analyzing the feature fusion data to obtain analysis results includes: Based on the cross-modal consistency score, analyzing the distribution of the painting feature information, the emotional feature information, and the physiological feature information when the cross-modal consistency is high; An analysis is performed based on the distribution and a preset House-Tree-Person Projective Test knowledge base to obtain the analysis result.
7. The House-Tree-Person Test analysis method based on multimodal time series data according to claim 1, characterized in that: The analyzing the feature fusion data to obtain analysis results includes: Cluster analysis is performed on the feature fusion data using an unsupervised learning method to obtain a cluster analysis result.
8. The House-Tree-Person Test analysis method based on multimodal time series data according to claim 1, characterized in that: The analyzing the feature fusion data to obtain analysis results includes: The analysis result is obtained by performing anomaly detection on the feature fusion data based on preset regular behavior pattern data through a pre-trained autoencoder.
9. The House-Tree-Person Test analysis method based on multimodal time series data according to claim 1, characterized in that: The calculating the cross-modal consistency score based on the splicing vector includes: Classifying the multimodal data in the concatenated vector into a positive mode or a negative mode; Obtaining activation values of the positive mode and the negative mode respectively; Determine the number of the positive mode and the negative mode in the same activation state based on the activation value and the preset threshold; The cross-modal consistency score is calculated based on the quantity.
10. The House-Tree-Person Test analysis method based on multimodal time series data according to any one of claims 1 to 9, characterized in that: The method further comprises: An analysis report is generated based on the analysis results.
11. A house-tree-person test analysis device based on multimodal time series data, characterized in that: The device comprises: a multimodal data acquisition module configured to acquire multimodal data generated by a target subject during a House-Tree-Person test; a multimodal feature extraction module, configured to extract features from the multimodal data based on a time series to obtain multimodal features; a splicing vector generating module, configured to splice the multimodal features to obtain a splicing vector; a feature fusion data generating module, configured to calculate a cross-modal consistency score based on the splicing vector, and fuse it with the splicing vector to obtain feature fusion data; an analysis result generating module, configured to analyze the feature fusion data to obtain an analysis result; The multimodal data includes: video information, audio information, handwriting information and physiological data; The multimodal feature extraction module is further configured to: Performing time series segmentation on the multimodal data according to a unified time window to obtain time series segmentation data; Extracting emotion features based on the video information and audio information in the time series segmentation data to obtain emotion feature information; Matching the emotional characteristic information with the preset segments in the house-tree-person test; Calculating cooperation feature information of the target object in a preset segment during the House-Tree-Person test based on the emotion feature information, video information, and audio information; Performing static feature extraction and dynamic feature extraction based on the video information and handwriting information to obtain painting feature information; Matching the drawing feature information with the preset segments in the house-tree-person test; The physiological data is subjected to feature extraction based on the preset segments in the house-tree-person test process to obtain physiological feature information.
12. An electronic device, characterized in that: include: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the house-tree-person test analysis method based on multimodal time series data according to any one of claims 1 to 10.
13. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to enable the computer to execute the house-tree-person test analysis method based on multimodal time series data according to any one of claims 1 to 10.
14. A computer program product, characterized in that The invention comprises a computer program, which, when executed by a processor, implements the steps of the house-tree-person test analysis method based on multimodal time series data according to any one of claims 1 to 10.
Citation Information
Patent Citations
Multi-mode depression detection method, system, medium and equipment
CN119580951A
Campus green space ownership perception evaluation method and system based on multi-modal learning
CN119862400A