Psychological assessment method and system based on line analysis and multi-modal visual language model

CN122842937APending Publication Date: 2026-09-29JINHUA YUNQI PSYCHOLOGICAL CONSULTING CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611005425.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-07
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

但是,若仅将原始线图输入通用多模态视觉语言模型进行分析,模型容易依据图像整体观感给出泛化判断,缺少线型检测、纸面校正、九宫格映射、笔画速度、笔画方向、起笔钩和回笔钩等可追溯的结构化事实依据

Benefits of technology

[0019]本发明的有益效果是:第一,本发明针对线解析测评建立了专用的结构化特征融合方式,将纸张摆放方向、纸面校正结果、九宫格位置、线型类别、检测框尺寸面积、笔画速度、笔画方向、起笔钩和回笔钩等线解析要素与原始图像绑定,形成多模态测评记录,使传统依赖人工经验的线解析观察过程具备可量化、可追溯和可复核的技术基础。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122842937A_ABST
    Figure CN122842937A_ABST
Patent Text Reader

Abstract

This invention discloses a psychological assessment method and system based on line analysis and a multimodal visual language model. The method includes: acquiring an image of a line drawing created by the test subject on A4 paper, and performing paper surface detection, perspective correction, and horizontal, vertical, or tilted placement recognition; inputting the corrected paper surface image into a line visual detection model to identify line type, detection box position, size, area, and confidence level, and mapping it to a nine-square grid on the paper to form a structured detection summary; when real-time video mode is enabled, capturing pen tip or fingertip trajectories and analyzing dynamic stroke features such as stroke speed, direction, pauses, starting hooks, and ending hooks. This invention improves the consistency and traceability of feature extraction, auxiliary analysis, and manual review in line analysis assessment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to an assessment method, specifically a psychological assessment auxiliary analysis method based on line analysis and a multimodal visual language model. Background Technology

[0002] Line drawing analysis is an assessment method that uses observation of a test-taker's line drawing on paper to supplement the analysis of their emotional state, behavioral tendencies, cognitive styles, and psychological development characteristics. In actual assessments, test-takers typically draw lines freely on ordinary A4 paper using a black pen. The assessor then combines this information with factors such as the paper's orientation, the line's position on the paper, line size, line type, line direction, starting stroke, pen movement, ending stroke, pen speed, pen pressure, pauses, and sense of control to make a comprehensive judgment. This type of assessment can preserve the natural expression during the actual paper drawing process, but its analysis results often rely on the assessor's professional experience and on-site observation skills.

[0003] Existing techniques for analyzing the psychology of drawing include methods that collect drawing data through online drawing platforms, identify drawing content through image features, predict psychological states through machine learning models, or collect drawing process data using sensors. These methods typically analyze broadly defined artworks, color distributions, image content, online handwriting trajectories, or sensor data, focusing primarily on drawing image feature extraction, training psychological state prediction models, or automatic report generation. However, line analysis assessment does not primarily focus on the complete drawing content or color expression, but rather on specific elements such as the spatial relationship between paper and lines, the combination of lines, the nine-square grid area where lines are located, line type variations, the stroke process, and the starting and ending points of strokes. Current technologies lack a dedicated structured feature system for line analysis assessment scenarios, making it difficult to reliably transform the manual observation rules in line analysis into machine-recognizable, calculable, and verifiable data.

[0004] Meanwhile, in real-world paper-and-pen testing scenarios, the A4 line drawings created by test subjects are typically obtained through camera captures or image uploads. The image acquisition process is easily affected by factors such as shooting angle, paper orientation (horizontal or vertical), paper tilt, perspective distortion, and differences in line thickness and scale. If the paper surface cannot be pre-detected, positioned, and perspective corrected, the actual position, size, and grid area of ​​the lines on the paper are difficult to determine accurately, hindering unified analysis across different testing images. While traditional manual line analysis relies on the experience of the evaluator, in batch testing, remote testing, or scenarios using ordinary evaluators, it is prone to problems such as unstable line type recognition, omissions of starting and ending strokes, inconsistent line position judgments, and incomplete recording of process information.

[0005] With the development of visual recognition models and multimodal visual language models, it has become possible to directly use artificial intelligence models to assist in the analysis of line drawings on paper. However, if the original line drawing is simply input into a general multimodal visual language model for analysis, the model tends to make generalized judgments based on the overall visual impression of the image, lacking traceable structured factual evidence such as line type detection, paper correction, nine-square grid mapping, stroke speed, stroke direction, starting hook, and ending hook. At the same time, line analysis evaluation has strong professional rule attributes. If a general model cannot combine it with a line analysis knowledge base for reasoning, it is also easy to ignore the special judgment logic in the line analysis system, resulting in insufficient stability of the analysis results and making it difficult for evaluators to manually review them.

[0006] Furthermore, psychological assessment processes are inherently sensitive and involve a degree of privacy. If test-takers directly see the system-generated psychological analysis conclusions or risk warnings during the assessment, it may influence their subsequent statements and affect the objectivity of the assessment process. Existing dual-screen interaction or display control solutions are mostly used for general information display and lack a role isolation display mechanism for the psychological assessment process. This makes it difficult to ensure that assessors can view the test charts, structured summaries, auxiliary analysis results, and psychological profiles, while test-takers only see the introductory text, collection status, waiting prompts, or neutral feedback. Moreover, the results of a single line graph assessment are usually insufficient to reflect the long-term changes of the test-taker, and existing technologies do not adequately disclose the continuous updating relationship between the results of previous line graph analyses, dialogue records, manual review information, and the individual's psychological profile.

[0007] Therefore, there is a need to provide a psychological assessment auxiliary analysis method and system based on line analysis and a multimodal visual language model. This method can perform paper surface detection, perspective correction, line target detection, nine-square grid positioning, dynamic stroke analysis, and structured feature fusion on line drawings drawn on ordinary A4 paper and with a black pen. The original image, structured detection summary, and line analysis knowledge base are jointly input into the multimodal visual language model to generate non-diagnostic psychological assessment auxiliary analysis results for assessors. At the same time, the stability, traceability, and security of the assessment auxiliary analysis are improved through role isolation display and psychological profile update mechanism. Summary of the Invention

[0008] To address the aforementioned problems, this invention provides a psychological assessment auxiliary analysis method based on line analysis and multimodal visual language models, which can effectively overcome the shortcomings of existing technologies.

[0009] This invention is achieved through the following technical solution: a psychological assessment method based on line analysis and a multimodal visual language model, comprising the following steps: Step 1: Obtain the line drawing image on paper created by the test subject, and save the original image corresponding to the line drawing image on paper; Step 2: Perform paper surface detection on the paper line drawing image to obtain paper surface contour information, and perform perspective correction on the paper surface based on the paper surface contour information to generate a corrected paper surface image of uniform scale. At the same time, output the paper placement direction information, which includes at least one of horizontal placement, vertical placement, or tilted placement. Step 3: Input the corrected paper image into the line visual detection model, identify the line detection objects in the corrected paper image through the line visual detection model, and output the line type, detection confidence, detection box position, detection box size and detection box area of ​​each line detection object; Step 4: Based on the relative position of the detection box in the corrected paper image, map each line detection object to the paper grid position, and fuse the paper placement direction information, perspective correction result, line type, detection confidence, detection box position, detection box size, detection box area and paper grid position to form a structured detection summary. Step 5: Bind the original image and the structured detection summary as a multimodal assessment record of the same assessment round, and input the multimodal assessment record into the multimodal visual language model service. The multimodal visual language model service combines the line parsing knowledge base to perform comprehensive reasoning on the original image and the structured detection summary to generate non-diagnostic psychological assessment auxiliary analysis results for assessors. Step Six: Through role isolation display control, display the non-diagnostic psychological assessment auxiliary analysis results, the structured test summary and / or line annotation diagram on the assessor's side display interface, and ensure that the test subject's side display interface only displays assessment guidance information, collection status information, waiting prompt information or neutral feedback information during the actual assessment process, without displaying the non-diagnostic psychological assessment auxiliary analysis results.

[0010] As a preferred technical solution, the paper line drawing image is an image collected after the test subject draws lines freely on multiple A4 sheets of paper with a black pen. The system generates image records for the corresponding paper according to the drawing order and configures an evaluation round index for each image record. The multimodal visual language model service performs a comprehensive analysis based on the changing trends of line type, line position, line size, starting position, and / or drawing control across different evaluation rounds.

[0011] As a preferred technical solution, when performing paper surface detection on the paper line drawing image, the paper line drawing image is first subjected to grayscale processing, blurring processing, binarization processing, morphological closing operation, and contour filtering to obtain a quadrilateral paper surface contour with an area conforming to a preset range. Then, a perspective transformation matrix is ​​calculated based on the four corner points of the quadrilateral paper surface contour, and the paper surface is corrected to a corrected paper surface image with a preset horizontal or vertical size using the perspective transformation matrix. When paper surface detection fails, the original image or auxiliary markers are used for paper orientation estimation.

[0012] As a preferred technical solution, the line visual detection model is a target detection model trained based on real tester's paper line drawing data. The real tester's paper line drawing data is labeled according to a line analysis annotation system, which includes multiple labels such as paper placement direction label, line type label, detection box position label, nine-square grid position label, line size label, intersection relationship label, broken line label, and swirl label. The line type includes multiple labels such as straight line, intersecting line, chaotic line, chaotic swirl line, circle line, soft line, swirl line, hard zigzag line, and broken line. When the detection confidence of any line detection object is lower than the preset confidence threshold, the line detection object will not be included in the structured detection summary, or will be marked as an object to be reviewed by the evaluator.

[0013] As a preferred technical solution, when mapping each line detection object to the nine-square grid position on the paper, the width and height directions of the corrected paper image are divided into three equal parts to form nine position regions, and the position region to which each line detection object belongs is determined according to the center coordinates of the detection frame of each line detection object; the detection frame area is the normalized area of ​​the detection frame area relative to the paper area, and the detection frame size includes the detection frame width, the detection frame height, or the line scale information calculated from the detection frame width and the detection frame height.

[0014] As a preferred technical solution, when enabling the real-time video detection mode, a dynamic stroke tracking step is also included: the visual acquisition module acquires video images of the subject drawing lines, identifies key points of the hand, and uses the fingertip of the index finger or the tip of the pen as trajectory proxy points. When the trajectory proxy point is located within the paper boundary and meets the writing posture conditions, record the coordinates and timestamp of the trajectory proxy point; When the trajectory proxy point leaves the paper boundary, the trajectory speed is lower than the preset static threshold and remains so for a preset time, or the trajectory duration exceeds the preset maximum time, the current stroke recording ends.

[0015] As a preferred technical solution, it also includes a dynamic stroke summary generation step: calculating the stroke duration, average speed, peak speed, speed variance, overall direction, starting direction and ending direction based on the trajectory point sequence of each stroke, and determining the starting hook shape and ending hook shape based on the directional changes of the trajectory of the first segment of the stroke and the trajectory of the second segment of the stroke, respectively. The stroke duration, average speed, peak speed, speed variance, overall direction, starting direction, returning direction, starting hook shape, returning hook shape, and number of strokes are used to form a dynamic stroke summary, which is then input into the multimodal visual language model service along with the structured detection summary.

[0016] As a preferred technical solution, the structured detection summary includes the structured line parsing feature vector and reliability information corresponding to the evaluation round; The structured line parsing feature vector includes multiple features such as paper placement direction encoding, line type category counting, paper nine-square grid position, detection box normalized area, overall stroke direction, stroke average speed, stroke speed variance, starting hook shape, and returning hook shape. The reliability information includes one or more of the following: paper surface detection reliability, line target detection reliability, and dynamic stroke acquisition reliability. The reliability of paper surface detection is determined based on whether the paper surface contour detection is successful, the stability of paper surface corner points, and the quality of perspective correction. The reliability of line target detection is determined based on the mean or median confidence level of the line detection object. The reliability of dynamic stroke acquisition is determined based on the number of effective strokes, the completeness of trajectory acquisition, and the continuity of trajectory. The system weights and fuses the paper detection reliability, line target detection reliability, and dynamic stroke acquisition reliability to generate a comprehensive reliability. Each fusion weight is non-negative and the sum of the weights is 1. The comprehensive reliability is used as part of the priority for manual review by the evaluator or as input prompts for the multimodal visual language model.

[0017] As a preferred technical solution, the multimodal visual language model service is a locally deployed or privately deployed multimodal visual language model service; When generating the non-diagnostic psychological assessment auxiliary analysis results, the knowledge rules related to paper placement, paper line relationship, line type, line position, line size, pen speed, pen pressure, starting and ending strokes in the line parsing knowledge base are first retrieved according to the structured detection summary. Then, the retrieved knowledge rules, original image, structured detection summary and role-limiting prompt words are input into the multimodal visual language model service. The role-restricted prompts limit the output results to non-diagnostic terms such as "may suggest", "tends to", or "needs further understanding"; when the visual appearance of the original image is inconsistent with the structured detection summary, the inconsistency is marked as requiring manual review by the evaluator.

[0018] The present invention provides a psychological assessment auxiliary analysis system based on line parsing and multimodal visual language model, comprising a client and a server; The client includes a visual acquisition module, an evaluator-side display interface, a test subject-side display interface, an audio module, a network communication module, and a role isolation display control module. The visual acquisition module is used to acquire images of the test subject's line drawing on paper and / or videos of the drawing process. The assessor's side display interface is used to display the test chart, structured test summary, non-diagnostic psychological test auxiliary analysis results and psychological profile information, while the test subject's side display interface is used to display test guidance information, collection status information, waiting prompt information or neutral feedback information. The server includes a paper correction module, a line visual detection model, a stroke process analysis interface, a feature fusion module, a multimodal visual language model service, a line parsing knowledge base module, and a psychological profile management module. The paper correction module is used to perform paper surface detection, perspective correction, and paper placement orientation recognition on the paper line drawing image. The line visual detection model is used to identify line detection objects and output line type, detection confidence, detection box position, detection box size, and detection box area. The stroke process analysis interface is used to generate dynamic stroke summaries based on the trajectory point sequence corresponding to the drawing process video. The feature fusion module is used to fuse the original image, paper orientation information, perspective correction results, line type, detection confidence, detection box position, detection box size, detection box area, paper grid position, and dynamic stroke summary into a multimodal assessment record. The multimodal visual language model service is used to generate non-diagnostic psychological assessment auxiliary analysis results in conjunction with the line parsing knowledge base module. The psychological profile management module is used to aggregate the results of previous paper-based line chart analysis, dialogue records, and manual review information according to the same workspace and the same subject thread, and to generate or update the structured psychological profile of the corresponding subject. The role isolation display control module is used to control the non-diagnostic psychological assessment auxiliary analysis results to be displayed only on the assessor's side display interface, and to restrict the display interface on the test subject's side display of psychological analysis conclusions during the actual assessment process.

[0019] The beneficial effects of this invention are as follows: First, this invention establishes a dedicated structured feature fusion method for line analysis evaluation, which binds line analysis elements such as paper placement direction, paper surface correction results, nine-square grid position, line type category, detection frame size and area, stroke speed, stroke direction, starting hook and ending hook to the original image, forming a multimodal evaluation record, thus enabling the traditional line analysis observation process that relies on manual experience to have a quantifiable, traceable and verifiable technical foundation.

[0020] Secondly, this invention uses A4 paper detection, perspective correction, and recognition of horizontal, vertical, or tilted placement to enable the processing of line positions, line sizes, and nine-square grid areas under a unified paper coordinate system, thereby reducing the impact of image acquisition errors on subsequent line analysis results.

[0021] Third, this invention inputs the structured detection summary, dynamic stroke process summary, original image, and line analysis knowledge base output by the line visual detection model into the multimodal visual language model. This allows the model analysis to no longer rely solely on the overall impression of the original image, but to combine detection facts such as line type, position, size, speed, direction, and stroke start and end points with line analysis expertise for comprehensive reasoning. This improves the stability of the auxiliary analysis results and the convenience of review for evaluators.

[0022] Fourth, this invention uses role-isolation display control to allow the assessor to display the test chart, structured test summary, non-diagnostic psychological test auxiliary analysis results, and psychological profile information, while the test subject only displays the test guidance, data collection status, waiting prompts, or neutral feedback during the actual test. This can reduce the risk of test subjects being psychologically influenced after seeing the analysis content in advance.

[0023] Fifth, this invention can aggregate the results of previous line graph analyses, dialogue records, and manual review information based on the same subject thread, and generate or update a structured psychological profile, so that the results of a single line graph assessment can be included in continuous assessment and consultation records, providing a reference for subsequent retesting, manual review, and longitudinal observation. Attached Figure Description

[0024] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0025] Figure 1 This is a system block diagram of the present invention; Figure 2 This is a flowchart of the present invention. Detailed Implementation

[0026] All features disclosed in this specification, or all steps in all disclosed methods or processes, may be combined in any way, except for mutually exclusive features and / or steps.

[0027] Any feature disclosed in this specification (including any appended claims, abstract, and drawings) may be replaced by other equivalent or similar features for a similar purpose, unless specifically stated otherwise. That is, unless specifically stated otherwise, each feature is merely one example of a series of equivalent or similar features.

[0028] like Figure 1 and Figure 2 As shown, this embodiment provides a psychological assessment auxiliary analysis method and system based on line parsing and a multimodal visual language model. It is used to collect, correct, identify, fuse features, and perform auxiliary analysis on line drawings created by test subjects under ordinary paper-and-pen drawing conditions. The system as a whole includes a client and a server. The client can be located in a small host with a double-sided screen, or it can be composed of an assessment terminal with a visual acquisition module, a display module, an audio module, a network communication module, and a main control computing unit. The client is used to complete test subject guidance, static image acquisition, dynamic stroke acquisition, assessor operation control, role isolation display, and data uploading. The server is used to complete A4 paper inspection, perspective correction, line visual inspection, dynamic stroke analysis, line parsing feature fusion, multimodal visual language model reasoning, line parsing knowledge base retrieval, and psychological profile management. The multimodal visual language model service can be a locally deployed or privately deployed VLM visual language model service, so as to generate non-diagnostic psychological assessment auxiliary analysis results for assessors while protecting the privacy of assessment data, in conjunction with the line parsing knowledge base.

[0029] Before the assessment begins, the assessor creates or selects a workspace and dialogue thread for the test taker via the client. The system establishes a session identifier for this assessment. This session identifier is used to associate the original image, corrected paper image, line detection results, dynamic stroke summary, structured detection summary, VLM visual language model output results, assessor manual review information, and psychological profile update records for this assessment. The assessor-side display on the client displays the workspace, dialogue thread, acquisition control, image detection results, line annotation diagram, structured detection summary, non-diagnostic psychological assessment auxiliary analysis results, and the psychological profile update entry. The test taker-side display on the client displays the assessment guidance, acquisition status, waiting prompts, and neutral feedback. In the actual assessment mode, psychological analysis conclusions, risk assessment results, and psychological profile content are not displayed to avoid the test taker being influenced by the analysis conclusions during the assessment process.

[0030] During the assessment, the client prompts the test-taker to freely draw lines on multiple A4 sheets of paper using a black pen, either through the test-taker's side display interface or audio module. Preferably, the test-taker completes line drawings on three A4 sheets sequentially, with each sheet corresponding to one assessment round. The system records these line drawings in the order they were drawn as the first line drawing, the second line drawing, and the third line drawing. After each sheet is completed, the assessor can trigger a photo detection through the assessor's side display interface. The client captures the image of the line drawing on the paper through the top visual acquisition module and saves the original image corresponding to that line drawing. If the client supports image uploads, the completed line drawing image can also be obtained through image upload. Each original image is bound to the test-taker's thread, session identifier, and assessment round index to facilitate subsequent differentiation of line positions, line type changes, and drawing control trends across different sheets.

[0031] After obtaining the original image, the client sends it to the A4 paper correction module on the server. The A4 paper correction module first converts the original image to grayscale to reduce the impact of ambient and background colors on paper edge recognition. Then, it blurs and binarizes the grayscale image to make the paper edges and lines more prominent. Next, it connects any broken parts in the paper edges using morphological closing operations and uses contour filtering to find quadrilateral paper contours with areas within a preset range. Once the system obtains the quadrilateral paper contour, it calculates the perspective transformation matrix based on the four corner points of the contour and uses this matrix to correct tilted or perspective-distorted paper into a uniformly sized horizontal or vertical corrected paper image. If the corrected paper width is greater than its height, the system determines the paper is horizontal; if the corrected paper height is greater than its width, the system determines the paper is vertical; if the paper edge is deflected relative to a reference direction, the system outputs the tilt angle. If the paper outline detection fails, the system can downgrade to using the original image for subsequent detection, or it can use auxiliary markers, image edge directions, or manually selected corner points to estimate the paper orientation and locate the paper.

[0032] After paper calibration, the server inputs the calibrated paper image into the line visual detection model. The line visual detection model is trained using A4 paper line drawings created by real evaluators, with the training data labeled according to a line analysis annotation system. This system includes labels for paper orientation, line type, detection box position, grid position, line size, intersection, broken line, and swirl. Through this annotation system, paper orientation, paper-line relationships, line-line relationships, and line type, which originally relied on human experience for judgment in line analysis, are transformed into structured labels that the visual target detection model can learn and output. After inferring from the calibrated paper image, the line visual detection model outputs the line type, detection confidence, detection box center coordinates, detection box width, detection box height, and detection box area for each detected line. Line type categories can include straight lines, intersecting lines, chaotic lines, chaotic swirling lines, circular lines, soft lines, swirling lines, hard folds, and broken lines. For detection objects with a confidence level lower than the preset confidence threshold, the system may not include them in the structured detection summary, or may mark them as objects to be reviewed by the evaluator, in order to avoid the low confidence detection results directly affecting subsequent auxiliary analysis.

[0033] To ensure that the spatial position of lines on the paper corresponds to the line analysis rules, the server, after obtaining the detection box, divides the width and height directions of the corrected paper image into three equal parts, forming nine position regions. The system determines the position of each line detection object within this nine-grid based on the center coordinates of its detection box and divides the detection box area by the paper area of ​​the corrected paper image to obtain the normalized area of ​​the detection box. The system can also determine line scale features based on the detection box width, detection box height, and the ratio between them, used to characterize the size relationship of lines on the paper. For multiple line detection objects on the same sheet of paper, the system statistically analyzes the number of each line type, the line distribution within each nine-grid region, the intersection relationships between different line detection objects, and the occurrence of special line types such as broken lines and swirls, thereby forming paper-line relationship and line-line relationship data corresponding to the line analysis knowledge system.

[0034] When using only static image acquisition mode, the system generates a structured detection summary for the evaluation round based on the original image, the corrected paper image, the paper orientation, line detection results, the detection box size and area, and the nine-square grid mapping results. When real-time video detection mode is enabled, the client also acquires video images during the subject's line drawing process and identifies key hand points or pen tip positions through a dynamic stroke tracking module. The system can select the index fingertip or pen tip position as a trajectory proxy point. When the trajectory proxy point is within the A4 paper boundary and meets the writing posture conditions, the client records the coordinates and timestamp of the trajectory proxy point. The distance and time difference between adjacent trajectory points are used to calculate the trajectory speed. When the trajectory proxy point leaves the paper boundary, the trajectory speed is below a preset static threshold for a preset time, or the duration of a single trajectory exceeds a preset maximum time, the system determines that the current stroke has ended and enters the waiting recording state for the next stroke.

[0035] After dynamic stroke tracking is completed, the client sends the trajectory point sequence, duration, average speed, peak speed, and speed variance of each stroke to the stroke process analysis interface on the server. The stroke process analysis interface calculates the overall direction, starting direction, and ending direction of the stroke based on the trajectory point sequence, and determines the starting hook and ending hook shapes based on the directional changes of the initial and subsequent trajectories of the stroke. If the direction of the initial trajectory has a significant angular deviation relative to the overall direction of the stroke, the system determines that a starting hook exists; if the direction of the subsequent trajectory has a significant angular deviation relative to the overall direction of the stroke, the system determines that a ending hook exists. The system can also statistically analyze the number of effective strokes, the number of pauses, the pause time, speed fluctuations, and trajectory continuity to reflect the rhythm changes and control characteristics during the subject's drawing process. The dynamic stroke summary and the static image detection summary are then combined in the subsequent feature fusion process, enabling the model to obtain not only the final shape of the line but also the speed, direction, pauses, and starting and ending stroke information during the line formation process.

[0036] The server-side feature fusion module binds information from each evaluation round, including the original image, corrected paper image, paper orientation, perspective correction result, line type, line type count, detection confidence, detection box size, normalized detection box area, grid position, stroke speed, stroke direction, starting hook shape, returning hook shape, and number of strokes, into a single multimodal evaluation record. For evaluation rounds without dynamic stroke acquisition, the system can mark dynamic fields such as stroke speed, speed variance, starting hook shape, and returning hook shape as default values ​​and indicate in the prompt that this round only includes static visual features. For evaluation rounds with dynamic stroke acquisition, the system stores the dynamic stroke summary and the static detection summary in correspondence and uses them together as input for subsequent VLM visual language model analysis.

[0037] To enable structured detection summaries to be expressed in a uniform manner, the structured line analytical feature vector corresponding to a single sheet of paper can be represented as: Where r represents the line drawing round or paper number, O represents the paper placement direction, θ represents the angle of the paper relative to the reference direction, n represents the number of lines, C represents the line type category set, p represents the position of the nine-square grid, a represents the normalized area of ​​the detection box, d represents the overall direction of the stroke, v represents the average speed, σ represents the speed variance, h represents the starting hook shape or the ending hook shape, and K represents the number of valid detection objects. For evaluation rounds where dynamic stroke acquisition is not enabled, the corresponding dynamic fields can be marked as default values, and the prompt should state that this round only has static visual features.

[0038] To indicate the reliability of the tests across different testing rounds to the evaluators, the system can also calculate the overall reliability of the r-th sheet. The overall reliability of the r-th sheet can be expressed as: The reliability of paper surface detection can be determined based on the success of paper contour detection, the stability of paper corner points, and the quality of perspective correction. The reliability of line target detection can be determined based on the mean or median confidence level of the detected line objects. The reliability of dynamic stroke acquisition can be determined based on the number of effective strokes, the completeness of acquisition, and the continuity of the trajectory. The overall reliability can be used as a priority for manual review by evaluators, or as one of the prompts input into the VLM visual language model. When the paper surface correction quality is low, the overall confidence level of line detection is low, or the dynamic stroke acquisition is incomplete in a certain evaluation round, the system will display a prompt on the evaluator's side indicating that this round needs priority review, and will explain the uncertainty of the relevant detection information in the model prompt.

[0039] The set of multimodal evaluation records input to the VLM visual language model can be represented as: Where I represents the original or corrected image of A4 paper, F represents the structured line analytical feature vector, ρ represents the overall reliability, R represents the number of sheets of paper used in a single evaluation, and D represents the evaluation dialogue and human review information. This multimodal evaluation record set unifies the original image, structured line analytical features, reliability information, and evaluator dialogue and review information, enabling the subsequent VLM visual language model to simultaneously obtain image facts, detection facts, reliability hints, and human supplementary information within the same input context.

[0040] The VLM visual language model, combined with the line parsing knowledge base and role-restricted cue word output for auxiliary analysis, can be represented as: Where M represents the multimodal visual language model service deployed locally or privately, Z represents the set of multimodal assessment records input to the VLM visual language model, K represents the line parsing knowledge base, P represents role-limiting and non-diagnostic output prompts, and Y represents the auxiliary analysis results output by the VLM visual language model. Before performing auxiliary analysis, the system retrieves relevant knowledge rules from the line parsing knowledge base based on keywords such as paper placement, line relationship, line type, grid position, line size, pen speed, pen pressure, and starting and ending strokes in the structured detection summary. The retrieved knowledge rules, original image, structured detection summary, dynamic stroke summary, and role-limiting prompts are then input into the VLM visual language model. The role-limiting prompts define the model's role as a line parsing psychological assessment counseling auxiliary analyst and require the output to use non-diagnostic terms such as "may suggest," "tends to," and "needs further understanding" to avoid generating definitive diagnostic conclusions. The model output may include overall observation, line feature description, non-diagnostic psychological assessment auxiliary analysis, and companionship suggestions.

[0041] Once the VLM visual language model obtains the original image and structured detection summary, its analysis process does not solely rely on the overall appearance of the original image. Instead, it simultaneously references the structured facts output by the line visual detection model and the professional rules in the line parsing knowledge base. For example, when the detection summary shows that lines on a piece of paper are concentrated in a specific nine-square grid area, the normalized area of ​​the detection box is small, and there are broken line features, the model needs to provide an explanatory statement based on these structured facts in its output. When the dynamic stroke summary shows that a stroke has a low average speed, a long pause time, or a distinct return stroke hook, the model can incorporate this as procedural observation information into its analysis. If there is a significant conflict between the visual appearance of the original image and the structured detection summary—for example, lines are visible in the original image but the detection model fails to recognize them, or the detection model identifies non-line traces as line objects—the system marks this conflict as requiring manual review by an evaluator and prompts the evaluator in the output results to prioritize the original image and the manual review.

[0042] The role-isolation display control module determines the scope of displayed content based on the current display object and assessment mode. In the real assessment mode, the assessor's side display interface can show the original image, corrected paper image, line annotation diagram, structured test summary, dynamic stroke summary, comprehensive reliability, non-diagnostic psychological assessment auxiliary analysis results, manual review entry, and psychological profile update entry. The test subject's side display interface only shows neutral content such as "Please draw lines freely on the paper," "Image acquisition in progress," "Please wait," and "This round of assessment is complete," without displaying model analysis conclusions, risk assessments, psychological profiles, or assessor notes. If the system is in debugging mode or a non-real assessment scenario, the role-isolation display control module can allow both display interfaces to simultaneously display partial content for developers or assessors to debug the equipment, but this mode is set differently from the real assessment mode.

[0043] On the assessor's side, the assessor can view the line annotation diagrams and detection summaries for each assessment round and manually review the system's recognition results. Assessors can confirm whether the line type of a specific line detection object is correct, whether the nine-square grid position is accurate, whether the detection box covers valid lines, whether the dynamic stroke trajectory is complete, and whether the overall reliability is reasonable. Assessors can also add observations or consultation records in the dialogue thread. The system saves the assessor's review results and the model output results together in the corresponding subject's thread, serving as the basis for subsequent psychological profile updates and longitudinal comparisons.

[0044] After the same test subject completes one or more line graph assessments, the psychological profile management module generates or updates a structured psychological profile based on the analysis results of previous line graphs, dialogue records, and manual review information within the same workspace and the same test subject's thread. The psychological profile can be updated incrementally according to the assessment rounds; the update process can be represented as follows: In this system, U represents the psychological profile after the assessment, Y represents the VLM (Visual Language Model) auxiliary analysis results corresponding to the assessment, D represents the dialogue and manual review information of the assessment, and G represents the psychological profile generation or update function. The structured psychological profile can include fields such as basic information, main problems, emotional state, personality traits, cognitive patterns, coping styles, support systems, consultation suggestions, risk assessment, and progress assessment. When generating or updating the profile, the system requires that it only organize information explicitly appearing in the current thread and use non-diagnostic, auxiliary expression methods. If a field lacks sufficient evidence, the system can mark the field as "requires further understanding" or leave it empty to avoid the model artificially supplementing information that has not been collected. The updated psychological profile is saved as a traceable historical record, allowing assessors to view changes at different points in time during subsequent consultations or retests.

[0045] In a specific application scenario, after the evaluator launches the client and selects a test subject thread, the client displays drawing instructions on the test subject's side. After the test subject completes the first A4 paper line drawing, the evaluator clicks to take a picture for inspection. The system captures the original image and completes paper surface inspection, perspective correction, line target detection, and nine-square grid positioning, then generates a structured inspection summary for the first round of evaluation. If real-time video inspection is not enabled in this round, the system marks the dynamic stroke field as the default value and inputs the original image and structured inspection summary into the VLM visual language model service. The VLM visual language model service combines the line parsing knowledge base to generate non-diagnostic auxiliary analysis. The evaluator's side displays the analysis results and labeled diagram, while the test subject's side only displays a waiting or completion prompt. The above process is repeated for the second and third A4 papers. The system saves the inspection results corresponding to each paper according to the evaluation round, and after the third paper is completed, it comprehensively analyzes the line distribution, line type changes, and drawing control trends between the three rounds.

[0046] In another specific application scenario, the evaluator activates real-time video detection mode. The client records key hand points or pen tip trajectories during the subject's drawing process, and generates dynamic stroke data by combining the trajectory point sequence, coordinates, and timestamps. After the subject completes the drawing, the system acquires both a static image of the line drawing on paper and analyzes the dynamic stroke data for stroke quantity, speed, direction, pauses, starting and ending strokes. The server merges the static detection summary and the dynamic stroke summary and inputs them into the VLM visual language model service, enabling the model to simultaneously obtain both the "final line shape" and the "process characteristics of drawing the line." Therefore, the system's output auxiliary analysis results include not only the spatial distribution and line type of the lines on the paper, but also process information such as drawing speed, trajectory continuity, pauses, and starting and ending strokes, more closely resembling the information obtained by a traditional evaluator during on-site observation.

[0047] In this embodiment, the training data for the line visual detection model can come from A4 paper line drawings created by real evaluators. During data annotation, annotators label the paper orientation, line type, detection box position, grid position, line size, intersection relationships, broken lines, and swirls according to the line analysis annotation system. The annotated data is used to train the line visual detection model, enabling it to identify straight lines, intersecting lines, chaotic lines, soft lines, hard zigzag lines, swirling lines, and broken lines in actual evaluation images, and output structured detection results related to line analysis auxiliary analysis. As the amount of training data increases, the system can expand the line type category set or optimize the detection confidence threshold, but this expansion does not change the basic processing flow of "paper correction, line detection, grid mapping, structured summarization, knowledge base enhanced VLM analysis, and role isolation display" in this embodiment.

[0048] Through the above implementation methods, this invention can transform information such as paper placement, paper-line relationships, line type, line spatial position, stroke speed, stroke direction, starting and ending strokes, and reliability information from line analysis assessments into machine-processable multimodal assessment records under conditions of ordinary A4 paper and black pen drawing. These records, along with the original image and the line analysis knowledge base, are then input into the VLM visual language model service to generate non-diagnostic psychological assessment auxiliary analysis results for assessors. This implementation method retains the naturalness of real paper-and-pen assessments while improving the consistency of line detection, feature recording, model analysis, and manual review. Simultaneously, through role isolation display and psychological profile update mechanisms, it reduces the risk of test subjects being influenced by analytical conclusions and provides a traceable longitudinal record for subsequent consultations and retesting.

[0049] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions conceived without inventive effort should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope defined in the claims.

Claims

1. A psychological assessment method based on line analysis and a multimodal visual language model, characterized in that, Includes the following steps: Step 1: Obtain the line drawing image on paper created by the test subject, and save the original image corresponding to the line drawing image on paper; Step 2: Perform paper surface detection on the paper line drawing image to obtain paper surface contour information, and perform perspective correction on the paper surface based on the paper surface contour information to generate a corrected paper surface image of uniform scale. At the same time, output the paper placement direction information, which includes at least one of horizontal placement, vertical placement, or tilted placement. Step 3: Input the corrected paper image into the line visual detection model, identify the line detection objects in the corrected paper image through the line visual detection model, and output the line type, detection confidence, detection box position, detection box size and detection box area of ​​each line detection object; Step 4: Based on the relative position of the detection box in the corrected paper image, map each line detection object to the paper grid position, and fuse the paper placement direction information, perspective correction result, line type, detection confidence, detection box position, detection box size, detection box area and paper grid position to form a structured detection summary. Step 5: Bind the original image and the structured detection summary as a multimodal assessment record of the same assessment round, and input the multimodal assessment record into the multimodal visual language model service. The multimodal visual language model service combines the line parsing knowledge base to perform comprehensive reasoning on the original image and the structured detection summary to generate non-diagnostic psychological assessment auxiliary analysis results for assessors. Step Six: Through role isolation display control, display the non-diagnostic psychological assessment auxiliary analysis results, the structured test summary and / or line annotation diagram on the assessor's side display interface, and ensure that the test subject's side display interface only displays assessment guidance information, collection status information, waiting prompt information or neutral feedback information during the actual assessment process, without displaying the non-diagnostic psychological assessment auxiliary analysis results.

2. The psychological assessment method based on line parsing and multimodal visual language model according to claim 1, characterized in that, The paper line drawing image is an image collected after the test subject draws lines freely on multiple A4 sheets of paper with a black pen. The system generates image records for the corresponding paper according to the drawing order and configures an evaluation round index for each image record. The multimodal visual language model service performs a comprehensive analysis based on the changing trends of line type, line position, line size, starting position, and / or drawing control across different evaluation rounds.

3. The psychological assessment method based on line parsing and multimodal visual language model according to claim 1, characterized in that, When performing paper surface detection on the paper line drawing image, the paper line drawing image is first processed by grayscale conversion, blurring, binarization, morphological closing operation and contour filtering to obtain a quadrilateral paper surface contour with an area that meets the preset range. Then, the perspective transformation matrix is ​​calculated based on the four corner points of the quadrilateral paper surface contour, and the paper surface is corrected to a corrected paper surface image with a preset horizontal or vertical size using the perspective transformation matrix. When paper surface detection fails, the paper orientation is estimated by using the original image or auxiliary markers.

4. The psychological assessment method based on line parsing and multimodal visual language model according to claim 1, characterized in that, The line visual detection model is a target detection model trained based on real tester paper line drawing data. The real tester paper line drawing data is labeled according to a line analysis annotation system, which includes multiple labels such as paper placement direction label, line type label, detection box position label, nine-square grid position label, line size label, intersection label, broken line label, and swirl label. The line type includes multiple types such as straight line, intersecting line, chaotic line, chaotic swirl line, circle line, soft line, swirl line, hard zigzag line, and broken line. When the detection confidence of any line detection object is lower than the preset confidence threshold, the line detection object will not be included in the structured detection summary, or will be marked as an object to be reviewed by the evaluator.

5. The psychological assessment method based on line parsing and multimodal visual language model according to claim 1, characterized in that, When mapping each line detection object to the nine-grid position on the paper, the width and height directions of the corrected paper image are divided into three equal parts to form nine position regions. The position region to which each line detection object belongs is determined according to the center coordinates of the detection frame of each line detection object. The detection frame area is the normalized area of ​​the detection frame area relative to the paper area. The detection frame size includes the detection frame width, detection frame height, or line scale information calculated from the detection frame width and detection frame height.

6. The psychological assessment method based on line parsing and multimodal visual language model according to claim 1, characterized in that, When the real-time video detection mode is enabled, a dynamic stroke tracking step is also included: the visual acquisition module acquires video images of the subject drawing lines, identifies key points of the hand, and uses the position of the index finger tip or the position of the pen tip as trajectory proxy points. When the trajectory proxy point is located within the paper boundary and meets the writing posture conditions, record the coordinates and timestamp of the trajectory proxy point; When the trajectory proxy point leaves the paper boundary, the trajectory speed is lower than the preset static threshold and remains so for a preset time, or the trajectory duration exceeds the preset maximum time, the current stroke recording ends.

7. The psychological assessment method based on line parsing and multimodal visual language model according to claim 6, characterized in that, It also includes a dynamic stroke summary generation step: calculating the stroke duration, average speed, peak speed, speed variance, overall direction, starting direction and ending direction based on the trajectory point sequence of each stroke, and determining the starting hook shape and ending hook shape based on the directional changes of the trajectory of the first segment of the stroke and the trajectory of the second segment of the stroke, respectively. The stroke duration, average speed, peak speed, speed variance, overall direction, starting direction, returning direction, starting hook shape, returning hook shape, and number of strokes are used to form a dynamic stroke summary, which is then input into the multimodal visual language model service along with the structured detection summary.

8. The psychological assessment method based on line parsing and multimodal visual language model according to claim 1, characterized in that, The structured detection summary includes the structured line parsing feature vector and reliability information corresponding to the evaluation round; The structured line parsing feature vector includes multiple features such as paper placement direction encoding, line type category counting, paper nine-square grid position, detection box normalized area, overall stroke direction, stroke average speed, stroke speed variance, starting hook shape, and returning hook shape. The reliability information includes one or more of the following: paper surface detection reliability, line target detection reliability, and dynamic stroke acquisition reliability. The reliability of paper surface detection is determined based on whether the paper surface contour detection is successful, the stability of paper surface corner points, and the quality of perspective correction. The reliability of line target detection is determined based on the mean or median confidence level of the line detection object. The reliability of dynamic stroke acquisition is determined based on the number of effective strokes, the completeness of trajectory acquisition, and the continuity of trajectory. The system weights and fuses the paper detection reliability, line target detection reliability, and dynamic stroke acquisition reliability to generate a comprehensive reliability. Each fusion weight is non-negative and the sum of the weights is 1. The comprehensive reliability is used as part of the priority for manual review by the evaluator or as input prompts for the multimodal visual language model.

9. The psychological assessment method based on line parsing and multimodal visual language model according to claim 1, characterized in that, The multimodal visual language model service is a locally deployed or privately deployed multimodal visual language model service; When generating the non-diagnostic psychological assessment auxiliary analysis results, the knowledge rules related to paper placement, paper line relationship, line type, line position, line size, pen speed, pen pressure, starting and ending strokes in the line parsing knowledge base are first retrieved according to the structured detection summary. Then, the retrieved knowledge rules, original image, structured detection summary and role-limiting prompt words are input into the multimodal visual language model service. The role-limiting prompts limit the output results to non-diagnostic terms such as "may prompt", "tend to", or "needs further understanding"; when the visual appearance of the original image is inconsistent with the structured detection summary, the inconsistency is marked as requiring manual review by the evaluator.

10. A psychological assessment auxiliary analysis system based on line parsing and multimodal visual language model, characterized in that, Including both client and server sides; The client includes a visual acquisition module, an evaluator-side display interface, a test subject-side display interface, an audio module, a network communication module, and a role isolation display control module. The visual acquisition module is used to acquire images of the test subject's line drawing on paper and / or videos of the drawing process. The assessor's side display interface is used to display the test chart, structured test summary, non-diagnostic psychological test auxiliary analysis results and psychological profile information, while the test subject's side display interface is used to display test guidance information, collection status information, waiting prompt information or neutral feedback information. The server includes a paper correction module, a line visual detection model, a stroke process analysis interface, a feature fusion module, a multimodal visual language model service, a line parsing knowledge base module, and a psychological profile management module. The paper correction module is used to perform paper surface detection, perspective correction, and paper placement orientation recognition on the paper line drawing image. The line visual detection model is used to identify line detection objects and output line type, detection confidence, detection box position, detection box size, and detection box area. The stroke process analysis interface is used to generate dynamic stroke summaries based on the trajectory point sequence corresponding to the drawing process video. The feature fusion module is used to fuse the original image, paper orientation information, perspective correction results, line type, detection confidence, detection box position, detection box size, detection box area, paper grid position, and dynamic stroke summary into a multimodal assessment record. The multimodal visual language model service is used to generate non-diagnostic psychological assessment auxiliary analysis results in conjunction with the line parsing knowledge base module. The psychological profile management module is used to aggregate the results of previous paper-based line chart analysis, dialogue records, and manual review information according to the same workspace and the same subject thread, and to generate or update the structured psychological profile of the corresponding subject. The role isolation display control module is used to control the non-diagnostic psychological assessment auxiliary analysis results to be displayed only on the assessor's side display interface, and to restrict the display interface on the test subject's side display of psychological analysis conclusions during the actual assessment process.