A Method and System for Screening and Classifying Mental States Based on Shared Concern Ability

By acquiring eye-tracking images and portrait videos, generating eye-tracking information, and using a neural network model for psychological state classification, this technology solves the problems of relying on subjective judgment of personnel and limitations of two-dimensional context in existing technologies, and achieves more efficient and accurate psychological state screening.

CN116687407BActive Publication Date: 2026-01-30BEIJING INST OF TECH +3
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202310646532.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-02
Publication Date
2026-01-30
Estimated Expiration
2043-06-02

AI Technical Summary

Technical Problem

Existing psychological state screening methods rely heavily on the subjective judgment of the personnel performing the screening, and eye-tracking-based screening methods are limited to two-dimensional scenarios, resulting in high rates of missed and false detections, and low screening efficiency and accuracy.

Method used

By acquiring eye-tracking images and portrait videos, eye-tracking information is generated, eye-tracking features are extracted, and a trained neural network model is used to classify psychological states. Combined with multi-context test tasks, the user's psychological state is assessed.

Benefits of technology

It enables a more objective and comprehensive screening of psychological states, reduces the rate of missed and false detections, and improves the accuracy and efficiency of screening.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116687407B_ABST
    Figure CN116687407B_ABST
Patent Text Reader

Abstract

This application provides a method and system for screening and classifying psychological states based on shared attention ability. The method, after acquiring eye-tracking images and portrait videos, generates eye-tracking information based on the images and videos, extracts eye-tracking features from the information, and inputs these features into a trained screening model to obtain classification results. The eye-tracking images and portrait videos are images and videos captured when a user performs a test action according to interactive instructions, which guide the user to view a target area. This method screens users' psychological states at the level of eye-tracking behavior, providing a more objective and comprehensive assessment, reducing false negatives and false positives, and improving screening accuracy and efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a method and system for screening and classifying psychological states based on shared attention ability. Background Technology

[0002] A mental state refers to the complete characteristics of mental activity within a certain period of time. Based on the characteristics of mental states, various mental states can be distinguished, such as attention, fatigue, tension, relaxation, sadness, and joy. Different mental states can produce different manifestations; therefore, mental state screening and classification can be performed based on user behavior to detect their mental state. For example, users with autism exhibit specific manifestations of mental states, including eye avoidance, unusual repetitive movements, preferences, and behavioral stereotypes. Therefore, mental state screening and classification methods can help assess the risk of autism.

[0003] Psychological state screening and classification methods can employ psychological scales, where personnel provide professional and highly accurate judgments based on authoritative standards. Alternatively, screening and classification equipment can be used to collect specific data, which is then compared with a control group through data statistics, analysis, and visualization to obtain effective information on the psychological state of the screened users.

[0004] However, using psychological scales for screening and classifying mental states places high demands on the personnel performing the tasks. They need extensive experience in psychological state screening to arrive at professional and accurate judgments. Therefore, unlike data-driven approaches, scale-based results often heavily rely on the personnel's subjective opinions; different subjective views and interpretations of the scales can lead to different results. Screening and classifying mental states using equipment primarily focuses on acquiring brain imaging data (neuroimaging), posture control patterns, and eye-tracking data. Furthermore, eye-tracking-based methods are limited to two-dimensional scenarios and lack authenticity in their social aspects. Summary of the Invention

[0005] This application provides a psychological state screening and classification method and system based on shared attention ability, in order to solve the problem of missed detections and false detections caused by the high dependence on the personnel and equipment used in psychological state screening and the limitation of features extracted to a two-dimensional context.

[0006] Firstly, this application provides a method for screening and classifying mental states based on shared attention capacity, including:

[0007] Acquire eye-tracking images and portrait videos, which are images and videos captured when the user performs test actions according to interactive instructions, and the interactive instructions are used to guide the user to view the target area;

[0008] An eye-tracking image set is obtained based on the portrait video and the eye-tracking images, and eye-tracking information is generated based on the eye-tracking image set and the portrait video, wherein the eye-tracking information includes reaction time, fixation time and response result;

[0009] Eye movement features are extracted from the eye movement information. The eye movement features include eye fixation features, overall fixation features, and eye behavior features. The eye fixation features are used to characterize reaction time, the overall fixation features are used to characterize fixation rate, and the eye behavior features are used to characterize correct response rate.

[0010] The eye movement features are input into a trained screening model to obtain the classification information output by the screening model. The screening model is a neural network model trained based on sample eye movement data. The sample eye movement data includes eye fixation features, overall fixation features, and eye behavior feature data with classification labels.

[0011] Secondly, this application provides a psychological state screening and classification system based on shared attention capacity, including:

[0012] The acquisition module is used to acquire eye-tracking images and portrait videos. The eye-tracking images and portrait videos are images and videos captured when the user performs test actions according to interactive instructions. The interactive instructions are used to guide the user to view the target area.

[0013] The preprocessing module is used to obtain an eye-tracking image set based on the portrait video and the eye-tracking images, and to generate eye-tracking information based on the eye-tracking image set and the portrait video, wherein the eye-tracking information includes reaction time, fixation time and response result;

[0014] The feature extraction module is used to extract eye movement features from the eye movement information. The eye movement features include eye fixation features, overall fixation features, and eye behavior features. The eye fixation features are used to characterize reaction time, the overall fixation features are used to characterize fixation rate, and the eye behavior features are used to characterize correct response rate.

[0015] The screening module is used to input the eye movement features into a trained screening model to obtain the classification information output by the screening model. The screening model is a neural network model trained based on sample eye movement data. The sample eye movement data includes eye fixation features, overall fixation features, and eye behavior feature data with classification labels.

[0016] As can be seen from the above technical solutions, this application provides a psychological state screening and classification method and system based on shared attention ability. The method, after acquiring eye-tracking images and portrait videos, generates eye-tracking information based on the images and videos, extracts eye-tracking features from the information, and inputs these features into a trained screening model to obtain classification results. The eye-tracking images and portrait videos are images and videos captured when the user performs test actions according to interactive instructions, which guide the user to view the target area. This method screens the user's psychological state at the level of eye-tracking behavior, providing a more objective and comprehensive assessment of the user's psychological state, reducing the false negative and false positive rates, and improving screening accuracy and efficiency. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is a schematic diagram of the structure of the psychological state screening and classification system in the embodiments of this application;

[0019] Figure 2 This is a schematic diagram of the psychological state screening and classification system from the user's perspective in an embodiment of this application;

[0020] Figure 3 This is a schematic diagram of the target area in an embodiment of this application;

[0021] Figure 4 This is a flowchart illustrating the psychological state screening and classification method in the embodiments of this application;

[0022] Figure 5 This is a schematic diagram of the process for obtaining the line-of-sight direction in an embodiment of this application;

[0023] Figure 6 This is a schematic diagram of the target plane in an embodiment of this application;

[0024] Figure 7 This is a schematic diagram of the psychological state screening and classification system in the embodiments of this application. Detailed Implementation

[0025] To make the objectives and implementation methods of this application clearer, the exemplary implementation methods of this application will be clearly and completely described below with reference to the accompanying drawings of the exemplary embodiments of this application. Obviously, the exemplary embodiments described are only some embodiments of this application, and not all embodiments.

[0026] Psychological state screening and classification methods can detect users' psychological state, but these methods have drawbacks such as high requirements for the personnel performing the screening, the screening results being highly dependent on the personnel's subjective opinions, and eye-tracking-based psychological state screening and classification methods being limited to two-dimensional scenarios and unable to objectively and comprehensively assess users' psychological state, resulting in low efficiency and low accuracy in psychological state screening.

[0027] To improve the accuracy and efficiency of screening, some embodiments of this application provide a psychological state screening and classification method based on shared attention ability. This method can screen users' psychological states at the level of eye movement behavior, providing a more objective and comprehensive assessment of their psychological states, reducing the false negative and false positive rates, and improving the efficiency and accuracy of psychological state screening. It should be noted that the psychological state screening and classification method provided in this application can be applied to psychological state classification based on eye movement data to improve screening accuracy and efficiency; for example, it can assist in assessing the risk of autism.

[0028] The aforementioned psychological state screening and classification method can be applied to psychological state screening and classification systems, such as... Figure 1 , Figure 2 As shown, the psychological state screening and classification system includes a support 11, a tabletop 12, and a data acquisition device 13. The support 11 is adjustable, and the tabletop 12 is detachably connected to the fixed support 11 to adjust the height of the tabletop 12 and the support 11 according to the user's height. The data acquisition device 13 is mounted on the tabletop 12 and is used to acquire the user's facial video and eye-tracking images. To facilitate more natural interaction with the user, a housing can be provided on the edge of the tabletop 12 to house the data acquisition device 13, thus concealing it and preventing it from attracting the user's attention.

[0029] The acquisition device 13 includes an eye acquisition device and a human image acquisition device. The eye acquisition device can be a depth camera used to acquire eye movement images of the user's eyes, and the human image acquisition device is used to capture human image videos of the user during the testing process.

[0030] To assess a user's psychological state, multiple test paradigms can be set up. These paradigms are pre-defined scenario-based test tasks designed to assess the user's reactions when performing test actions. Therefore, the psychological state screening and classification system also includes an instruction device 15, which can generate interactive instructions to guide the user in performing corresponding test actions.

[0031] In this embodiment, the psychological state screening and classification method is used for psychological state classification. For example, for autistic patients, their joint attention ability is lacking, making it difficult for them to look at objects in the corresponding direction according to instructions. Therefore, in order to analyze the performance differences of autistic patients, the testing paradigm can be a test task that requires the user to look at a target area. The instruction device 15 and the user are located on opposite sides of the table 12, and the instruction device 15 guides the user to look at the target area through language or action guidance.

[0032] The target area can be a pre-defined location, such as... Figure 3 As shown, the test task scenario is divided into multiple Regions of Interest (ROIs), including four target regions and a background region. The four target regions are ROI_1, ROI_2, ROI_3 and ROI_4, and the background region is ROI_5 (Other region).

[0033] For ease of observation by users, such as Figure 1 , Figure 2 , Figure 3 As shown, a marker 14 can be placed at a preset target area to attract the user's attention. For example, for autistic individuals, since the majority of autistic individuals are in the early childhood stage, the marker 14 can be set according to the preferences of early childhood users, for example... Figure 3 The doll shown.

[0034] It is understandable that the instruction device 15 can be a device with voice interaction capabilities, and can generate interactive instructions through preset program steps to guide the user to perform corresponding test actions.

[0035] To improve efficiency, the mental state screening and classification system includes a controller configured to execute a mental state screening and classification method. The controller can be connected to the acquisition device 13 and the command device 15. During testing, the controller controls the command device 15 to generate interactive commands and controls the acquisition device 13 to execute a data capture program.

[0036] like Figure 4 As shown, Figure 4 This is a flowchart illustrating the psychological state screening and classification method based on shared attention ability in the embodiments of this application, specifically including the following:

[0037] S100, acquires eye-tracking images and portrait videos.

[0038] Among them, eye-tracking images and portrait videos are images and videos captured when the user performs test actions according to interactive instructions. The interactive instructions are used to guide the user to view the target area. For example, such as Figure 3 As shown, Figure 3This is a schematic diagram of a psychological state screening and classification system as seen from the user's perspective. The instruction device 15 and the user are located on opposite sides of a table 12. The data acquisition device 13 is positioned on the table 12, with four pre-set target areas (ROI_1, ROI_2, ROI_3, and ROI_4) located on either side of the instruction device 15. A marker 14 is placed in each target area. Before the test begins, the height of the table 12 and the support 11 is adjusted according to the user's height, and the data acquisition device 13 is calibrated to ensure accurate positioning of the user's eyes, avoiding errors in eye movement data caused by differences in height and eye size. At the start of the test, the data acquisition device 13 is activated, controlling the instruction device 15 to generate interactive commands to guide the user in performing test actions, i.e., guiding the user to view the target areas. The data acquisition device 13 then captures eye movement images and facial videos of the user performing the test actions.

[0039] To improve screening accuracy, users can be guided to view multiple preset target areas sequentially. Therefore, in some embodiments, multiple test paradigms can be generated sequentially based on the number of target areas. Each test paradigm is associated with an instruction to perform a test action of viewing the target area. The test paradigms are then encapsulated as interactive instructions and sent via instruction device 15. For example, as... Figure 3 As shown, four target areas are set, and four test paradigms can be generated based on the four target areas. Each test paradigm is associated with a test action that executes the viewing of the target area.

[0040] To enhance the interactive experience, the instruction device 15 can be a humanoid robot, for example, such as... Figure 3 As shown, the humanoid robot can point to one of the markers 14 with its finger and issue the voice command "Look at this," thus sending an interactive command to guide the user to look at marker 14, i.e., the target area. Similarly, interactive commands guiding the user to look at the remaining three markers 14 can be sent sequentially. Each interactive command can be set with a preset time; for example, after sending one interactive command, wait 3-5 seconds before sending the next interactive command. The above process can be achieved by capturing eye-tracking images and facial videos of the user as they perform the test actions according to the interactive commands using the acquisition device 13.

[0041] In this embodiment, since users in different psychological states exhibit different eye movement data when performing test actions—for example, autistic patients may show eye avoidance—using eye movement changes from multi-scenario test tasks as analysis and prediction information allows for a more objective and comprehensive assessment of users' psychological states.

[0042] S200, obtains a set of eye-tracking images based on portrait videos and eye-tracking images.

[0043] Since portrait videos are dynamic image sequences and eye-tracking images are individual eye-tracking entries, each representing eye-tracking data for a single frame, preprocessing is required after acquiring the portrait video and eye-tracking images. This involves reading frames from the portrait video to obtain the eye-tracking images corresponding to the same time point. In some embodiments, image frames from the portrait video can be read, and the frame number positions corresponding to the start and end nodes of the portrait video can be marked to obtain the frame number range. The eye-tracking images are then traversed, and those falling within the frame number range are filtered to generate an eye-tracking image set, i.e., selecting the valid eye-tracking images corresponding to the time points in the portrait video.

[0044] S300 generates eye-tracking information based on a set of eye-tracking images and a portrait video.

[0045] After acquiring the set of eye-tracking images and the portrait video, it is necessary to extract eye-tracking information to analyze the user's psychological state. This eye-tracking information includes reaction time, fixation duration, and response result. Reaction time is the duration between the time the interaction command is sent and the start time of the user's effective fixation; fixation duration is the effective fixation duration of the user on the target area; and response result is the correctness of the user's execution of the test action according to the interaction command.

[0046] For reaction time, the starting time of effective gaze when the user performs the test action can be obtained from the eye-tracking image set, that is, the time when the user's gaze falls on the target area. Then, the sending time of the interaction command can be obtained, and the interval between the sending time and the starting time can be calculated to obtain the reaction time.

[0047] For the response, image frames from a portrait video can be acquired, and facial recognition can be performed on these frames to identify the eye positions, thus determining the user's gaze origin and direction. Next, a target plane is acquired, and the gaze point is calculated based on the gaze origin, gaze direction, and the target plane. By comparing the gaze point with the target area, it is determined whether the user's gaze falls on the target area, thereby generating the response. Here, the target plane is the plane containing the target area, and the gaze point is the intersection of the user's gaze and the target plane.

[0048] To improve the accuracy of eye-tracking data, a head posture recognition step can be added when determining the gaze direction. The gaze direction is determined based on the eye position and head posture. Therefore, in some embodiments, image frames from a portrait video can be acquired, and face recognition can be performed on these frames to identify facial features, including the coordinates of the corners of the eyes, the tip of the nose, and the corners of the mouth. Head posture estimation is then performed based on these facial features to obtain the user's head posture, and gaze direction estimation is performed based on the eye features and head posture to obtain the user's gaze direction.

[0049] like Figure 5 As shown, an image frame containing a face can be input into a preset face recognition model for face recognition, and the face image of the image frame can be cropped. The face image can then be input into a preset face feature point detection model for face feature point recognition, identifying the position of the eyes. Next, the face image can be input into a preset head pose estimation model, which extracts the coordinates of facial features, namely the corners of the eyes, the tip of the nose, and the corners of the mouth. Based on the position of these feature points, head pose estimation is performed to obtain the head pose. Finally, the head pose and eye position are input into a preset gaze estimation model, which performs gaze estimation to obtain the gaze direction.

[0050] The head pose estimation model described above can output three Euler angles: raw (rotation around the y-axis), pitch (rotation around the x-axis), and roll (rotation around the z-axis). The gaze estimation model outputs the direction vector of the gaze direction, such as... Figure 6 As shown, its coordinate system is a right-handed coordinate system with the midpoint between the user's eyes as the origin, the direction of the face pointing towards the acquisition device 13 as the Z-axis, and the Y-axis pointing vertically upwards. Figure 6 As shown, the coordinate system (spatial coordinate system) provided by the acquisition device 13 is centered on the camera of the acquisition device 13. The Z-axis points in the direction the camera is facing, the Y-axis points downwards, and the X-axis points to the right when viewed from the direction the Z-axis points. This coordinate system is a right-handed coordinate system. The target plane is the XOY plane of the spatial coordinate system, and the point of view is the two-dimensional coordinate of the intersection of the user's line of sight and the target plane under the target plane.

[0051] After acquiring the starting point, direction, and target plane of the gaze, normalization is performed on these elements to obtain the starting coordinates of the gaze point and the direction vector of the gaze direction. This involves normalizing the spatial coordinate system provided by the acquisition device 13, the spatial coordinates of the gaze direction output by the gaze estimation model, and the spatial coordinates of the head posture output by the head posture estimation model. Then, based on the starting coordinates and direction vector, the coordinates of the gaze landing point are calculated using the principle of vector parallelism. Finally, based on the coordinates of the gaze landing point and the target area, it can be determined whether the user's gaze falls on the target area.

[0052] like Figure 6 As shown, point K(x1, y1, z1) represents the coordinates of the midpoint of the user's eyes in the spatial coordinate system, i.e., the coordinates of the starting point of the gaze. Simultaneously, point K is also the origin of the user's head pose coordinate system. Regions S1, S2, S3, S4, and S5 represent the areas occupied by the projection of the target region onto the target plane. The direction vector of the gaze direction output by the gaze estimation model is... Based on the principle of vector parallelism, the coordinates G(x2, y2, z2) of the point where the line of sight falls are calculated according to the following formula:

[0053]

[0054]

[0055]

[0056] After obtaining the coordinates of the gaze point, the range of the target area can be determined. This range is the area within the target plane that the target area occupies, calculated beforehand based on the distance between the target area and the acquisition device 13. By comparing the coordinates of the gaze point with the range, if the coordinates of the gaze point are within the range, it is determined that the user's gaze is on the target area, and the response result is marked as correct. If the coordinates of the gaze point are not within the range, it is determined that the user's gaze is not on the target area, and the response result is marked as incorrect.

[0057] S400 extracts eye movement features from eye movement information.

[0058] A user's attention to different target areas can be measured by reaction time, fixation rate, and correct response rate. When a user performs a test action of viewing certain areas of interest, the shorter the reaction time, the higher the fixation rate, and the higher the correct response rate, the better the user's co-attention ability. Therefore, after obtaining eye movement information, eye movement features can be extracted from the eye movement information to assess the user's co-attention ability and thus screen the user's psychological state.

[0059] In some embodiments, eye movement information can be input into a trained eye movement modality classification model to obtain eye movement features. The training of the eye movement modality classification model involves acquiring multiple sample eye movement information and their corresponding eye movement features as training sample data to train the classifier. The sample eye movement information includes sample eye information labeled with different psychological states. For example, in a psychological state screening classification method used to assist in assessing the risk of autism, the sample eye information includes sample eye information labeled with "having autism" and sample eye information labeled with "not having autism." Therefore, the trained eye movement modality classification model can be applied to autism screening.

[0060] The eye movement features include eye fixation features, overall fixation features, and eye behavior features, as shown in the table below:

[0061]

[0062] Among these, eye gaze features are used to characterize reaction time and assess different core user abilities and characteristics. Reaction time can be calculated based on eye movement image sequences, specifically the duration between the end of the interaction command and the user's first effective gaze point. Overall gaze features characterize gaze rate, used to comprehensively assess a user's gaze behavior in different sub-scenarios under various contexts; this can be obtained by calculating the ratio of the user's effective gaze time within the test period. Eye behavior features characterize correct response rate, used to assess a user's core abilities; this can be obtained by statistically analyzing the ratio of the number of correct responses to the number of interaction commands issued.

[0063] In some embodiments, a logistic regression (LR) model can be used to train the classifier. During training, eye movement features, overall eye movement features, and eye behavior features corresponding to multiple samples of eye movement information are calculated, resulting in a total of 9 dimensions. To simplify the training and prediction of the machine learning model, the features can be dimensionality reduced, retaining the main features, reducing the amount of data, and improving the efficiency of training and prediction. For example, a recursive feature elimination algorithm with cross-validation (RFECV) can be used for feature dimensionality reduction, selecting the optimal 75% of features (7 dimensions) as the feature combination for the eye movement modality. This feature combination is then used to train the classifier for the eye movement modality, resulting in an eye movement modality classification model.

[0064] It is understandable that this application proposes a feature extraction scheme based on multiple feature groups, which is not limited to eye gaze-related features, but also extracts important evaluation indicators such as overall gaze features and behavioral features in addition to eye gaze features, so as to more comprehensively evaluate the different psychological states of users in multiple contexts.

[0065] S500 inputs eye-tracking features into the trained screening model to obtain the classification information output by the screening model.

[0066] The screening model is a neural network model trained on sample eye-tracking data, which includes eye fixation features, overall eye fixation features, and eye behavior features with classification labels. This can be achieved by obtaining a sample eye feature set, which includes multiple sample eye features labeled with different mental states. These sample eye feature sets are then fused into sample eye-tracking data. A neural network model is trained based on this sample eye-tracking data to obtain the screening model. For example, a mental state screening classification method is used to assist in assessing the risk of autism. The sample eye feature set includes eye features labeled with "having autism" and eye features labeled with "not having autism." These sample eye feature sets are fused into sample eye-tracking data, and a logistic regression (LR) model is trained to obtain the screening model.

[0067] Eye-tracking features are input into a trained screening model, which analyzes these features and outputs classification information. This classification information includes a screening value, which represents the probability of assigning an eye-tracking feature to a psychological state label. The screening value can be used to determine the user's psychological state, thereby generating a screening result.

[0068] In some embodiments, a screening threshold can be preset. After obtaining the classification information output by the screening model, the screening value in the classification information is read, and the screening threshold is obtained. The screening value and the screening threshold are compared. If the screening value is greater than or equal to the screening threshold, the eye movement feature is marked as the first psychological state. If the screening value is less than the screening threshold, the eye movement feature is marked as the second psychological state.

[0069] For example, in autism screening, the first mental state is "having autism," and the second mental state is "not having autism." If the screening value output by the model is less than the screening threshold, the user's mental state is determined to be the second mental state, generating a screening result of "not having autism." Conversely, if the value is greater than the threshold, a screening result of "having autism" is generated. The screening result can be represented using different symbols; for example, the number "0" represents "not having autism," and the number "1" represents "having autism."

[0070] Based on the above-mentioned psychological state screening and classification methods, such as Figure 7 As shown, some embodiments of this application also provide a psychological state screening and classification system based on shared attention ability, including:

[0071] The acquisition module is used to acquire eye-tracking images and portrait videos.

[0072] Among them, eye-tracking images and portrait videos are images and videos captured when users perform test actions according to interactive instructions. The interactive instructions are used to guide users to view the target area.

[0073] The preprocessing module is used to obtain a set of eye-tracking images from the portrait video and eye-tracking images, and to generate eye-tracking information from the set of eye-tracking images and the portrait video.

[0074] Eye movement information includes reaction time, fixation time, and response result.

[0075] The feature extraction module is used to extract eye movement features from eye movement information.

[0076] Among them, eye movement features include eye fixation features, overall fixation features, and eye behavior features. Eye fixation features are used to characterize reaction time, overall fixation features are used to characterize fixation rate, and eye behavior features are used to characterize correct response rate.

[0077] The screening module is used to input eye-tracking features into the trained screening model to obtain the classification information output by the screening model.

[0078] The screening model is a neural network model trained based on sample eye movement data, which includes eye gaze features, overall gaze features, and eye behavior features with classification labels.

[0079] In some embodiments, the psychological state screening and classification system further includes a monitoring module for recording the user's testing process. The monitoring module may include a global camera and a monitor. The global camera records the user's overall testing status in real time for subsequent evaluation and filtering of the user's condition and data availability. The monitor observes the footage recorded by the global camera in real time.

[0080] As can be seen from the above technical solutions, this application provides a psychological state screening and classification method and system based on shared attention ability. The method, after acquiring eye-tracking images and portrait videos, generates eye-tracking information based on the images and videos, extracts eye-tracking features from the information, and inputs these features into a trained screening model to obtain classification results. The eye-tracking images and portrait videos are images and videos captured when the user performs test actions according to interactive instructions, which guide the user to view the target area. This method screens the user's psychological state from the perspective of eye-tracking behavior, providing a more objective and comprehensive assessment of the user's psychological state, reducing the false negative and false positive rates, and improving screening accuracy and efficiency.

[0081] For ease of explanation, the above description has been provided in conjunction with specific embodiments. However, the exemplary discussion above is not intended to be exhaustive or to limit the embodiments to the specific forms disclosed above. Various modifications and variations can be obtained based on the above teachings. The selection and description of the above embodiments are for the purpose of better explaining the principles and practical applications, thereby enabling those skilled in the art to better utilize the embodiments and various different variations of embodiments suitable for specific application considerations.

Claims

1. A mental state screening classification method based on a common interest ability, characterized in that, The method comprises the following steps: obtaining eye movement images and portrait videos, the eye movement images and the portrait videos being images and videos obtained by shooting when a user performs a test action according to an interactive instruction, the interactive instruction being used to guide the user to watch a target region; obtaining an eye movement image set according to the portrait videos and the eye movement images, and generating eye movement information according to the eye movement image set and the portrait videos, the eye movement information comprising a reaction time length, a gaze time length and a response result; extracting eye movement features from the eye movement information, the eye movement features comprising eye gaze features, overall gaze features and eye behavior features, the eye gaze features being used to represent the reaction time, the overall gaze features being used to represent a gaze rate, and the eye behavior features being used to represent a correct response rate; inputting the eye movement features into a trained screening model to obtain classification information output by the screening model, the screening model being a neural network model trained according to sample eye movement data, the sample eye movement data comprising eye gaze feature data, overall gaze feature data and eye behavior feature data with classification labels; the number of the target regions is multiple, and the method further comprises the following steps: generating multiple test paradigms in sequence according to the number of the target regions, each of the test paradigms being associated with an instruction to perform a test action of watching a target region; encapsulating the test paradigms as the interactive instruction, and sending the interactive instruction; the step of generating the eye movement information according to the eye movement image set and the portrait videos comprises the following steps: obtaining a sending time point of the interactive instruction; obtaining a starting time point of effective gaze of the user when performing the test action according to the eye movement image set; calculating an interval time length between the sending time point and the starting time point to obtain the reaction time length; the step of generating the eye movement information according to the eye movement image set and the portrait videos comprises the following steps: obtaining an image frame of the portrait video; identifying an eye position of the image frame to obtain a visual line starting point and a visual line direction of the user; obtaining a target plane, the target plane being a plane where the target region is located; calculating a visual line landing point according to the visual line starting point, the visual line direction and the target plane, the visual line landing point being an intersection point of the visual line of the user and the target plane; comparing the visual line landing point and the target region to generate the response result.

2. The psychological state screening classification method according to claim 1, characterized by, the step of obtaining the eye movement image set according to the portrait videos and the eye movement images comprises the following steps: reading an image frame of the portrait video; marking image frame number positions of a start node and an end node of the portrait video to obtain a frame number range; traversing the eye movement images; screening the eye movement images located in the frame number range to generate the eye movement image set.

3. The psychological state screening classification method according to claim 1, wherein after the step of obtaining the image frame of the portrait video, the method further comprises the following steps: identifying a face feature of the image frame, the face feature comprising an eye corner coordinate, a nose tip coordinate and a mouth corner coordinate; performing head posture estimation according to the face feature to obtain a head posture of the user; performing visual line estimation according to the eye position and the head posture to obtain a visual line direction of the user.

4. The psychological state screening classification method according to claim 1, wherein According to the line-of-sight starting point, the line-of-sight direction and the target plane, the step of calculating the line-of-sight landing point comprises: Performing normalization processing on the line-of-sight starting point, the line-of-sight direction and the target plane to obtain a starting point coordinate of the line-of-sight starting point and a direction vector of the line-of-sight direction; According to the starting point coordinate and the direction vector, the coordinate of the line-of-sight landing point is calculated.

5. The mental state screening classification method of claim 1, wherein, The method further comprises: Obtaining a sample eye feature set, the sample eye feature set comprising a plurality of sample eye features labeled with different psychological state labels; Fusing the sample eye feature set into sample eye movement data; Training the neural network model based on the sample eye movement data to obtain a screening model.

6. The mental state screening classification method of claim 1, wherein, The step of inputting the eye movement feature into the trained screening model to obtain classification information output by the screening model further comprises: Reading a screening value in the classification information, the screening value being a classification probability of the eye movement feature belonging to a psychological state label; Obtaining a screening threshold value; If the screening value is greater than or equal to the screening threshold value, marking the eye movement feature as a first psychological state; If the screening value is less than the screening threshold value, marking the eye movement feature as a second psychological state.

7. A mental state screening classification system based on common interest ability, characterized in that, The system is configured with the psychological state screening and classification method of claim 1, and the system comprises: A collection module configured to obtain eye movement images and portrait videos, the eye movement images and the portrait videos being images and videos obtained by shooting when a user performs a test action according to an interactive instruction, the interactive instruction being used to guide the user to watch a target area; A preprocessing module configured to obtain an eye movement image set from the portrait videos and the eye movement images, and to generate eye movement information from the eye movement image set and the portrait videos, the eye movement information comprising reaction time, fixation time and response result; A feature extraction module configured to extract eye movement features from the eye movement information, the eye movement features comprising eye fixation features, overall fixation features and eye behavior features, the eye fixation features being used to represent reaction time, the overall fixation features being used to represent fixation rate, and the eye behavior features being used to represent correct response rate; A screening module configured to input the eye movement features into a trained screening model to obtain classification information output by the screening model, the screening model being a neural network model trained according to sample eye movement data, the sample eye movement data comprising eye fixation features, overall fixation features and eye behavior feature data with classification labels.

Citation Information

Patent Citations

  • Autism spectrum disorder screening system and method based on eye movement and facial expression

    CN115429271A