Eye movement calibration recognition method, device and system based on artificial intelligence

A dual-branch eye movement calibration method using deep neural networks with cross-modal feature fusion addresses the inaccuracies in existing methods, improving the identification of complex anomalies and enhancing clinical assessment reliability.

CN120318896APending Publication Date: 2025-07-15SHANGHAI ZEHNIT MEDICAL TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510399913.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

Existing eye movement calibration methods are difficult to accurately identify complex anomalies, such as undershoot, overshoot or response delays, and relying on manual interpretation or simple rule algorithms to lead to misleading and inefficiency.

Method used

A deep neural network model with a two-branch structure is adopted, combined with a cross-modal feature fusion mechanism, and the eye movement data is encoded and feature extraction is used to identify eye movement calibration states, including the cross-fusion of eye movement trajectory and spatial distribution characteristics, to achieve multi-objective joint discrimination.

Benefits of technology

It improves the evaluation accuracy and processing efficiency of eye movement calibration, enhances the ability to identify abnormal behaviors, and improves the practicality and prediction accuracy of the model in clinical evaluation scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318896A_ABST
    Figure CN120318896A_ABST
Patent Text Reader

Abstract

The invention discloses an artificial intelligence-based eye movement calibration identification method, device and system, and the method comprises the steps: obtaining the eye movement data of a subject in an eye movement calibration process; based on the eye movement data, constructing input data of at least two different dimensions, including a first type of data used for representing eyeball movement track changes; the second type of data is used for representing spatial distribution characteristics of eyeball positions; the first type of data and the second type of data are input into a trained intelligent recognition model, and the intelligent recognition model is constructed based on a deep neural network structure, has at least two branch coding structures and is used for extracting features of input data of different dimensions; the recognition model further comprises a cross-modal feature fusion mechanism which is used for carrying out fusion processing on feature vectors output by the branch coding structures. And recognizing and outputting an eye movement calibration result through the intelligent recognition model. According to the invention, intelligent identification of eye movement calibration is realized, and evaluation accuracy and processing efficiency are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the fields of artificial intelligence and biomedical engineering, and particularly relates to an eye movement calibration recognition method, device, and system based on artificial intelligence. Background Art

[0002] Existing eye movement calibration methods mainly rely on manual interpretation or simple rule algorithms to judge whether the eye movement trajectory is consistent with the calibration target, and it is difficult to accurately identify complex abnormal conditions such as undershoot, overshoot, or reaction delay. For example, in the saccade test used clinically, the calibration factor is mostly compared with a set threshold to determine whether the calibration passes. In actual clinical tests, even if the subject does not cooperate well, the calibration factor may fall within the normal range, and for some special subjects, even if they cooperate normally, the calibration factor may fall outside the normal range. Another example is that during the clinical eye nystagmus view examination, the operator usually analyzes through software, that is, judges whether the calibration passes by the numerical threshold of the calibration factor, which easily leads to incorrect calibration factors, and all subsequent eye movement examinations output incorrect data regarding the rotation angle and rotation speed of the eyeball, bringing great trouble and misguidance to the clinical interpretation of the test report.

[0003] In addition, for abnormal conditions such as undershoot and overshoot commonly seen clinically, clinicians also need to judge based on experience, with low efficiency and poor consistency. Summary of the Invention

[0004] To solve at least one of the above technical problems, the present application provides an eye movement calibration recognition method, computer program product, device, and system based on artificial intelligence. The present application uses a dual-branch structure to encode different types of eye movement data, and combines a cross-modal feature fusion mechanism to achieve intelligent recognition of eye movement calibration, improving the evaluation accuracy and processing efficiency.

[0005] In a first aspect, the present application discloses an eye movement calibration recognition method based on artificial intelligence, including:

[0006] Obtain the eye movement data of a subject during the eye movement calibration process;

[0007] Based on the eye movement data, construct at least two types of input data with different dimensions, including: the first type of data, used to represent the change in the eye movement trajectory; the second type of data, used to represent the spatial distribution characteristics of the eye position;

[0008] Input the first type of data and the second type of data into a trained intelligent recognition model. The intelligent recognition model is constructed based on a deep neural network structure and has at least two branch encoding structures for respectively extracting the features of different dimensions of input data; the recognition model also includes a cross-modal feature fusion mechanism for fusing and processing the feature vectors output by each branch encoding structure;

[0009] The result of eye movement calibration is recognized and output through an intelligent recognition model; wherein, the result of eye movement calibration includes whether the eye movement calibration is successful, and / or whether there is an undershoot, and / or whether there is an overshoot.

[0010] In some embodiments, the first type of data includes any one of an image, a sequence, and vector data for describing the change of an eye movement trajectory; the second type of data includes any one of a density map, a heat map, a position histogram, and a coordinate statistical vector for describing the spatial distribution of an eye position or a fixation point.

[0011] In some embodiments, the input data is processed by an intelligent recognition model through the following steps to recognize and output the result of eye movement calibration: extracting dynamic trajectory behavior features from the first type of data to obtain a global behavior feature vector; extracting static spatial features from the second type of data to obtain a spatial distribution feature vector; using a cross-attention mechanism to cross-fuse the global behavior feature vector and the spatial distribution feature vector to generate a fused feature representation; and performing multi-object joint discrimination on the fused feature through a shared feature processing layer and a plurality of task-specific branch structures to output a multi-dimensional recognition result.

[0012] In some embodiments, extracting dynamic trajectory behavior features from the first type of data to obtain a global behavior feature vector specifically includes: dividing the input first type of data into multiple data blocks; converting each data block into an embedding vector; adding a learnable position encoding and a classification marker vector representing overall feature information to each embedding vector sequence. Using a multi-layer Transformer encoder, including a multi-head attention mechanism and a feed-forward network structure, to perform global feature extraction on the embedding vector sequence containing the position encoding and the classification marker; and extracting the classification marker vector from the feature vectors output by the multi-layer encoder as the global behavior feature vector of the trajectory map.

[0013] In some embodiments, extracting static spatial features from the second type of data to obtain a spatial distribution feature vector specifically includes: enhancing the features of the input second type of data through convolution; performing multi-stage downsampling through a depthwise separable convolution structure and a channel expansion structure to extract multi-scale spatial features; converting the multi-scale spatial features into feature vectors of a unified dimension; and performing a linear transformation on the pooled feature vectors to output a spatial distribution feature vector of a unified dimension.

[0014] In some embodiments, a cross-attention mechanism is adopted to cross-fuse the global behavior feature vector and the spatial distribution feature vector to generate a fused feature representation, which specifically includes: mapping the trajectory behavior feature vector and the spatial distribution feature vector to a feature space of a unified dimension respectively; interacting and fusing the two types of features after the unified dimension through a multi-layer encoder; based on the output of the fusion unit, extracting the global fused feature through an attention pooling layer and outputting the fused feature representation.

[0015] In some embodiments, through a shared feature processing layer and multiple task-specific branch structures, multi-object joint discrimination is performed on the fused feature, and a multi-dimensional recognition result is output, which specifically includes: uniformly regularizing the fused feature representation to extract shared discriminant information; based on the shared discriminant information, performing different recognition tasks by adopting multiple task-specific branch structures; wherein, the multiple task-specific branch structures include multiple classification task branch structures and / or at least one regression task branch structure; wherein: each classification branch structure includes a group of bilinear layers for performing binary classification recognition operations for different classification tasks respectively; the classification tasks include calibration state recognition, overshoot recognition, and undershoot recognition; the regression task branch structure includes at least one fully connected layer and an output layer for regression prediction, which is used to output continuous numerical results of the fixation position or calibration error based on the fused feature representation; splicing the prediction results of each task branch to generate a structured recognition result, including whether the calibration is successful, whether there is an undershoot, and / or whether there is an overshoot.

[0016] In some embodiments, both the first type of data and the second type of data are image data, and at least two branch encoding structures of the intelligent recognition model are respectively a Vision Transformer network structure and an improved ConvNext convolutional network.

[0017] In some embodiments, the input data input to the intelligent recognition model further includes visual target auxiliary data; wherein, the acquisition of the visual target auxiliary data includes: acquiring data related to the visual target display in the calibration experiment; based on the data related to the visual target display, constructing the visual target auxiliary data, including: a third type of data for representing the position change data of the visual target display; and a fourth type of data for representing the spatial distribution feature of the visual target display position.

[0018] In a second aspect, the present application discloses a computer program product, including: computer-executable instructions or a computer program, and when the computer-executable instructions or the computer program are executed by a processor, the artificial intelligence-based eye movement calibration recognition method described in any one of the above is implemented.

[0019] In a third aspect, the present application discloses an artificial intelligence-based eye movement calibration recognition device, including:

[0020] An information interface module, configured to receive eye movement input data, where the eye movement input data includes: first type of data for representing changes in the eye movement trajectory; second type of data for representing the spatial distribution characteristics of the eye position.

[0021] A data processor, coupled to the information interface module, where the data processor is provided with a trained intelligent recognition model. The intelligent recognition model is constructed based on a deep neural network structure and has at least two branch coding structures for respectively extracting features of different types of input data; the recognition model further includes a cross-modal feature fusion mechanism for performing fusion processing on the feature vectors output by each branch coding structure. The data processor is configured to execute the following instructions:

[0022] Input the first type of data and the second type of data into the intelligent recognition model.

[0023] Output the result of eye movement calibration through the intelligent recognition model; where the result includes whether the eye movement calibration is successful, and / or whether there is an undershoot, and / or whether there is an overshoot.

[0024] In a fourth aspect, the present application discloses an eye movement calibration recognition system based on artificial intelligence, including:

[0025] An eye movement acquisition module, configured to obtain eye movement data of a subject during the eye movement calibration process.

[0026] A data construction module, configured to construct at least two different-dimensional eye movement input data based on the eye movement data, including: first type of data for representing changes in the eye movement trajectory; second type of data for representing the spatial distribution characteristics of the eye position.

[0027] An intelligent recognition module, configured to, according to the received first type of data and the second type of data, recognize and output the result of eye movement calibration through a trained intelligent recognition model; where the intelligent recognition model includes:

[0028] A first coding module, configured to extract dynamic trajectory behavior features from the input first type of data to obtain a global behavior feature vector.

[0029] A second coding module, configured to extract static spatial features from the input second type of data to obtain a spatial distribution feature vector.

[0030] A feature fusion module, which uses a two-stream cross-attention mechanism to cross-fuse the global behavior feature vector and the spatial distribution feature vector to generate a fused feature representation.

[0031] A multi-task decision module, including a shared feature processing layer and multiple task branch structures, configured to perform multi-object joint discrimination on the fused features and output a multi-dimensional recognition result.

[0032] Compared with the prior art, the present invention has at least one of the following beneficial effects:

[0033] 1. The present application adopts at least two branch coding structures to adapt to eye movement data of different dimensions, extracts features from the first type of data and the second type of data respectively, and improves the coding accuracy; and adopts a cross-modal feature fusion mechanism to interactively model different modal feature vectors to realize the associated representation of eye movement trajectory behavior features and spatial distribution, and improves the comprehensive discrimination ability of the model for the eye movement calibration state.

[0034] 2. When the present application extracts features from the first type of data (eye movement trajectory data), it globally and dynamically models the first type of data through data block partitioning, positional encoding and multi-layer Transformer encoders, enhancing the model's understanding ability of trajectory continuity, directionality and offset patterns, thereby improving the recognition accuracy of abnormal behaviors such as under-shoot and over-shoot. When extracting features from the second type of data (eye position spatial distribution data), a multi-stage downsampling method is adopted to perform multi-scale feature extraction on the eye position spatial distribution data, improving the extraction ability of features such as fixation focus and distribution deviation, and contributing to improving the accuracy of eye movement calibration test recognition.

[0035] 3. The present application introduces a two-stream cross-attention mechanism in the feature fusion process, and generates a cross-modal fusion feature representation through an attention pooling strategy, effectively realizing the deep fusion expression of dynamic trajectories and static distributions, and enhancing the model's understanding ability of semantic alignment and complementary information between modalities. The multi-task decision module adopts a combination of shared features and task-specific structures to support the simultaneous output of structured recognition results such as whether the calibration is successful, over-shoot / under-shoot state, and fixation deviation angle, realizes multi-objective joint optimization, and improves the practicality and prediction accuracy of the model in clinical evaluation scenarios.

[0036] 4. The present application further introduces a visual target display as auxiliary data, constructs a third type of data (visual target position change) and a fourth type of data (visual target spatial distribution), and incorporates them into the model processing flow through a new coding module or input fusion method, enabling the model to have "target-guided" semantic information, strengthening the logical pairing ability between eye movement behaviors and visual target presentations, thereby improving the discrimination accuracy of the model for the cooperation degree of eye movement calibration tests and the sensitivity to abnormal behaviors. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] The above characteristics, technical features, advantages and their implementation manners of the present invention will be further described below in a clear and understandable manner in combination with the drawings in the preferred embodiments.

[0038] Figure 1 is a flowchart of an embodiment of the eye movement calibration recognition method based on artificial intelligence of the present application;

[0039] Figure 2 It is a structural block diagram of the artificial intelligence recognition model in the embodiments of the present application;

[0040] Figure 3 It is an example of an eye movement trajectory map with normal calibration in an embodiment of the present application;

[0041] Figure 4 It is an example of a spatial distribution map with normal calibration in an embodiment of the present application;

[0042] Figure 5 It is an example of an eye movement trajectory map with typical under - shoot in an embodiment of the present application;

[0043] Figure 6 It is an example of an eye movement trajectory map with typical over - shoot in an embodiment of the present application;

[0044] Figure 7 It is a structural diagram of the artificial intelligence recognition model in an embodiment of the present application;

[0045] Figure 8 It is a schematic diagram of the ROC curve for model evaluation in an embodiment of the present application. Detailed implementation manners

[0046] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will describe the specific implementation manners of the present invention with reference to the accompanying drawings. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings, and other implementation manners can also be obtained.

[0047] To make the drawings concise, only the parts related to the invention are schematically shown in each drawing, and they do not represent the actual structure of the product. Additionally, to make the drawings concise and easy to understand, for components with the same structure or function in some drawings, only one of them is schematically shown, or only one of them is marked. In this document, "one" not only means "only one", but also means "more than one" situation.

[0048] It should also be further understood that the term "and / or" used in the specification and claims of the present application refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0049] In this text, it should be noted that unless otherwise clearly specified and defined, the terms "install", "connect", and "link" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium, and it can be the communication inside two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.

[0050] In addition, in the description of the present application, the terms "first", "second", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.

[0051] In one embodiment, the present application discloses an eye movement calibration and recognition method based on artificial intelligence, as Figure 1 shown, including:

[0052] S100, obtaining the eye movement data of the subject during the eye movement calibration process;

[0053] Specifically, the acquisition of the eye movement data in this embodiment can be actively collecting the eye movement data of the subject during the eye movement calibration process, such as collecting the eye movement data of the subject in the eye movement calibration test through a video oculography; it can also be receiving the eye movement data of the subject during the eye movement calibration process. This embodiment does not limit the acquisition method of the eye movement data, nor does it limit the data type of the eye movement data. For example, the eye movement data includes the change sequence of the eye position over time.

[0054] S200, based on the eye movement data, constructing at least two different-dimensional input data, including:

[0055] The first type of data, used to represent the change of the eye movement trajectory;

[0056] The second type of data, used to represent the spatial distribution characteristics of the eye position;

[0057] Specifically, after obtaining the eye movement data, it is necessary to further process the eye movement data to construct the model input data of different dimensions. The first type of data includes any one of the image, sequence, and vector data used to describe the change of the eye movement trajectory; the second type of data includes any one of the density map, heat map, position histogram, and coordinate statistical vector used to describe the spatial distribution of the eye position or fixation point. For example, based on the eye movement data, constructing an eye movement trajectory image or an eye movement trajectory time series as the first type of data; calculating the position distribution of the eye position (substantially equivalent to the fixation point) during the entire calibration process and constructing it as a fixation heat map or a coordinate distribution histogram as the second type of data.

[0058] Further, synchronously obtain the visual target display information of the calibration test, and construct the visual target presentation data (the third type of data) and the spatial layout diagram (the fourth type of data).

[0059] S300. Input the first type of data and the second type of data into the trained intelligent recognition model. The intelligent recognition model is constructed based on a deep neural network structure and has at least two branch coding structures for respectively extracting the features of input data of different dimensions. The recognition model also includes a cross-modal feature fusion mechanism for fusing the feature vectors output by each branch coding structure.

[0060] Specifically, the intelligent recognition model in this embodiment has at least two branch coding structures and a cross-modal feature fusion mechanism. The at least two branch coding structures can be the same network structure or different network structures. Since at least two different dimensions of input data are input, corresponding branch coding structures are used to adaptively process the input data of different dimensions for feature extraction. This application does not limit the specific number of coding structures used or the number of branches, which can cover structure reuse and expansion. Of course, the number of branches can also be determined according to the dimension or type of the input data. For example, in this embodiment, if the input data only includes the first type of data and the second type of data, a double-branch coding structure (preferably, a double-branch heterogeneous coding structure) can be used to respectively extract the features of the first type of data and the second type of data, and a cross-modal fusion mechanism is used to jointly extract and fuse the feature vectors output by each branch coding structure to achieve cross-modal recognition, improve the model's ability to model the association between trajectory behavior and eye position spatial distribution, and enhance the recognition accuracy and stability of abnormal calibration results.

[0061] S400. Output the result of eye movement calibration through the intelligent recognition model. Among them, the result of eye movement calibration includes whether the eye movement calibration is successful, and / or whether there is an undershoot, and / or whether there is an overshoot.

[0062] In the above step S400, the input data is processed by the intelligent recognition model in the following steps to identify and output the eye movement calibration result:

[0063] S410. Extract the dynamic trajectory behavior features from the first type of data to obtain the global behavior feature vector.

[0064] S420. Extract the static spatial features from the second type of data to obtain the spatial distribution feature vector.

[0065] S430. Use the cross-attention mechanism to cross-fuse the global behavior feature vector and the spatial distribution feature vector to generate a fused feature representation.

[0066] S440 performs multi-object joint discrimination on the fused features through a shared feature processing layer and multiple task-specific branch structures, and outputs multi-dimensional recognition results.

[0067] Obviously, the processing of this recognition output is carried out through specific network structures (each module) in the intelligent recognition model. Next, we will focus on elaborating the model architecture of the intelligent recognition model:

[0068] The intelligent recognition model adopted in this application, as Figure 2 shown, specifically includes:

[0069] (1) The first encoding module 100 is used to extract dynamic trajectory behavior features from the input first type of data and obtain a global behavior feature vector.

[0070] Preferably, the first encoding module 100 includes:

[0071] The data partitioning unit 110 is used to partition the input first type of data into multiple data blocks.

[0072] The linear embedding unit 120 is used to convert each data block into an embedding vector.

[0073] The position encoding unit 130 is used to add learnable position encoding and a classification marker vector representing overall feature information to each embedding vector sequence.

[0074] The encoder unit 140 uses multiple layers of encoders, including a multi-head attention mechanism and a feed-forward network structure, and is used to perform global feature extraction on the embedding vector sequence containing position encoding and classification markers.

[0075] The feature extraction output unit 150 is used to extract the classification marker vector from the output of the encoder unit as the global behavior feature vector of the trajectory graph.

[0076] Taking the data form of the first type of data as sequence data as an example, assuming that the first type of data includes a two-dimensional coordinate sequence of the eye position recorded by an eye tracker changing over time. To process the above sequence input, the first encoding module in this embodiment can be constructed based on a time series modeling neural network structure, such as a time series encoding network based on a bidirectional gated recurrent unit (Bi-GRU) or a long short-term memory network (Bi-LSTM); or based on a time series Transformer structure (Time Transformer), which can more effectively capture long-term dependencies and trajectory turning features. The specific processing flow is as follows:

[0077] S111, first partition the input first type of data into multiple data blocks.

[0078] For the eye movement trajectory sequence data, it can be divided into several time windows at a fixed step size;

[0079] S112. Convert each data block into an embedding vector;

[0080] For each data block, it can be transformed into a fixed-length vector representation through MLP, 1D convolution, or a small RNN, which is equivalent to constructing a local trajectory segment feature representation;

[0081] S113. Add a learnable position encoding and a classification token vector representing the overall information to each embedding vector;

[0082] In this step, the temporal structure is retained by adding the position encoding, and the classification token [CLS] is used as the first token of the Transformer sequence for global information aggregation.

[0083] S114. Use a multi-layer encoder for global modeling;

[0084] The multi-layer Transformer Block can be directly used. Each layer contains multi-head attention and a feed-forward network, and the embedding vector with the added position encoding and classification token information is used for feature extraction sequence.

[0085] S115. Extract the [CLS] vector as the global behavior feature of the trajectory;

[0086] Specifically, extract the [CLS] output from the output of the multi-layer encoder as the aggregated semantic vector of the sequence as a whole, which is used as the global behavior feature vector of the eye movement trajectory sequence data.

[0087] Through the above example, the first type of data is the eye movement trajectory data in the form of a time series. In the present invention, the original trajectory sequence can still be divided into multiple time segments (data blocks), converted into embedding vectors respectively, and combined with the learnable position encoding and classification tokens to construct an embedding sequence. This sequence can be input into a multi-layer attention-based Transformer encoder for global feature extraction, so as to obtain the aggregated semantic representation of the complete trajectory behavior, which is used as the basic feature for subsequent multi-modal fusion and calibration recognition. Of course, the first type of data can also be image data, and the first encoding module is also processed according to the above processing flow. The subsequent embodiments will elaborate in detail on the processing situation when the first type of data is an eye movement trajectory image.

[0088] (2) The second encoding module 200 is used to extract static spatial features from the input second type of data to obtain a spatial distribution feature vector;

[0089] The second encoding module 200 includes:

[0090] A feature extraction unit 210, configured to perform feature enhancement on the input second-type data through convolution;

[0091] A multi-stage downsampling unit 220, including a depthwise separable convolution structure and a channel expansion structure, for performing multi-stage downsampling and extracting multi-scale spatial features;

[0092] A global pooling unit 230, for converting the multi-scale spatial features into a feature vector of a unified dimension;

[0093] A feature mapping unit 240, for performing a linear transformation on the pooled feature vector and outputting a spatial distribution feature vector of a unified dimension.

[0094] Similarly, the present application does not limit the presentation form of the second-type data, which may be a set of fixation point coordinates reflecting the spatial distribution of eye positions (fixation points), a spatial distribution histogram, statistical features of spatial heat zone coordinates, a probability density vector of spatial distribution, etc. The second encoding module can select an adapted MLP network or a structure-aware encoder for feature modeling according to the input type, to effectively abstractly represent the spatial characteristics of fixation behavior and further improve the discriminative ability of the model for abnormal distribution features. For example, when the input second-type data does not have an image structure, the second encoding module can adopt a structure composed of a multi-layer perceptron, normalization processing, and linear mapping to achieve abstract modeling and semantic enhancement of the distribution features, so as to output a spatial feature representation with the same dimension as the trajectory behavior features, which is suitable for cross-modal feature fusion tasks. At this time, the second encoding module includes:

[0095] A feature preprocessing unit: configured to perform preprocessing according to the presentation form of the second-type data, including preset format conversion, embedding construction, or vector standardization; for example, if the input is a structured feature vector or statistical data, normalization or positional embedding encoding (if it is a point set) can be performed first;

[0096] A feature enhancement unit: for modeling the preprocessed spatial distribution vector of the input and extracting spatial distribution features. If the input spatial distribution vector is a structured feature vector, a non-linear transformation can be adopted to enhance the feature expression ability; specifically, a multi-layer perceptron network structure (such as MLP) or a lightweight fully connected network can be used to implement;

[0097] A feature compression unit: for performing dimensionality reduction or principal component extraction on the enhanced vector (such as Dropout + Linear + Activation);

[0098] A feature mapping unit: mapping the feature vector to a unified dimension consistent with the output of the first encoding module for subsequent feature fusion processing.

[0099] (3) Feature fusion module, which uses the cross-attention mechanism to cross-fuse the global behavior feature vector and the spatial distribution feature vector to generate a fused feature representation;

[0100] The feature fusion module 300 includes:

[0101] A dimension unification unit 310, used to map the trajectory behavior feature vector and the spatial distribution feature vector to a feature space of unified dimension respectively;

[0102] A feature fusion unit 320, including a multi-layer Transformer encoder, is used to interactively fuse the above two types of features after unifying the dimensions;

[0103] A fusion feature output unit 330, including an attention pooling layer, is used to extract a global fusion feature according to the output of the feature fusion unit and output a fusion feature representation;

[0104] (4) A multi-task decision module 400, including a shared feature processing layer and multiple task-specific branch structures, is used to perform multi-target joint discrimination on the fused feature representation and output a multi-dimensional recognition result.

[0105] The multi-task decision module 400 includes:

[0106] A feature sharing unit 410 is used to unify the fused feature representation and extract shared discriminant information;

[0107] The prediction unit 420 performs different recognition tasks based on the shared discrimination information by using multiple task-specific branch structures. The prediction unit uses multiple task-specific branch structures, including multiple classification task branch structures and / or at least one regression task branch structure; wherein:

[0108] Each classification branch includes a set of bilinear layers for performing binary classification recognition operations for different classification tasks; the classification tasks include calibration state recognition, overshoot recognition, and undershoot recognition; the regression task branch structure includes at least one fully connected layer and an output layer for regression prediction, which is used to output the continuous numerical results of the gaze position or calibration error based on the fused feature representation;

[0109] The output splicing unit 430 is used to splice the prediction results of each task branch to generate a structured recognition result, including whether the calibration is successful and the abnormality type determination (whether there is undershoot and / or overshoot, etc.).

[0110] In another embodiment of the present application, we take the eye movement trajectory diagram and eye position spatial distribution diagram in the eye movement calibration process as an example of the input model data, and perform eye movement calibration recognition through the intelligent recognition model. The following is the specific implementation process:

[0111] 1. Collection of training data: Collect the eye movement sample data of each subject during the calibration test, and label the collected eye movement sample data according to the pre-established annotation standard for calibration data (this standard is generally formulated by senior clinicians). To ensure the accuracy of annotation, the doctor can finally check the annotation quality. Data annotation includes the degree of cooperation when looking at three positions in the calibration data (classified into four levels according to the degree of cooperation: excellent, good, poor, very poor. Only when the first two levels are reached at all three positions can it be considered that the calibration is passed). The three positions correspond to left calibration, middle calibration, and right calibration, as well as the classification of whether there is an undershoot waveform.

[0112] 2. Construction of the input data of the model: The input of the model in this example uses the eye movement trajectory map and the eye position spatial distribution map. The size of the picture is the set size, such as a picture of 220*1220. Both the trajectory map and the spatial distribution map take the center of the picture as the 0 position of the ordinate and the leftmost side as the 0 position of the abscissa, and the coordinate axes are hidden when drawing. As Figures 3 - 6 shown (the coordinate axes are hidden in all the figures. Among them, for the eye movement trajectory map, the hidden coordinate axes are: the horizontal axis is time, and the vertical axis is the change in the fixation angle; for the eye position spatial distribution map, the hidden coordinate axes are: the horizontal axis is the eye position or the fixation position, and the vertical axis is the probability distribution). Figure 3 is the eye movement trajectory map of normal calibration (calibration successful); Figure 4 is the eye position spatial distribution map of normal calibration; Figure 5 is an example of a typical undershoot eye movement trajectory map. Figure 5 The place indicated by the arrow in it shows an obvious undershoot observed during the right saccade, and it takes multiple saccades to reach the target position. This indicates that the initial saccade amplitude of the patient is insufficient and one or more compensatory saccades are required for correction. Figure 6 is an example of a typical overshoot eye movement trajectory map. The Figure 6 arrow in it indicates the overshoot. This figure shows the eye movement trajectory during the reverse saccade following the left and right saccades. It can be seen from the figure that the initial saccade amplitude of the patient is too large. This failure to fixate on the visual target requires correcting the reverse eye movement to reposition the fixation.

[0113] In this embodiment, the input data of the constructed model can be an eye movement trajectory map and an eye position spatial distribution map containing coordinate axes; of course, an eye movement trajectory map and an eye position spatial distribution map with hidden coordinate axes can also be used. Since the intelligent recognition model processes the pixel distribution of images rather than the "coordinate axis annotations" in the images, what the model focuses on is the trajectory path form and hotspot area distribution of the images rather than text annotations or coordinate scales. Therefore, whether to display the coordinate axes has no impact on model feature extraction. Instead, hiding the coordinate axes can reduce ineffective visual interference. Secondly, after hiding the coordinate axes, the content in the image can represent the eye movement behavior characteristics more "purely", and the training is more focused. Thirdly, from the perspective of representational integrity, since the eye movement trajectory map and the eye position spatial distribution map are usually generated on a unified standard canvas (such as a 256*256 image), and the pixel positions themselves carry spatial position meanings, the model will implicitly capture the spatial layout during convolution or self-attention operations. Therefore, even if the coordinate axes are hidden, as long as the image pixels are generated according to rules, the spatial relationship is still effectively modeled.

[0114] In summary, when the eye movement trajectory map and the fixation spatial distribution map are used as model inputs, a standard image format without coordinate axes is adopted to reduce the interference of non-semantic elements on model training. Since the spatial coordinates are embedded in the pixel distribution during image generation, hiding the coordinate axes will not affect the spatial structure modeling. Instead, it is beneficial to improve the feature focusing ability and generalization performance of the model.

[0115] 3. Construction of the intelligent recognition model: The intelligent recognition model in this embodiment adopts a design of dual-branch encoding + feature fusion + multi-task output. The model architecture is as Figure 7 shown. This model adopts a dual-branch architecture. Preferably, for these two types of data, namely the eye movement trajectory map and the spatial distribution map, a dual-branch heterogeneous encoding structure can be adopted, that is: one branch (the first encoding module) uses Vision Transformer (ViT) to extract global behavior features from the eye movement trajectory map; the other branch (the second encoding module) uses ConvNext Adapter (ConvNext convolutional network) to extract multi-scale spatial features from the eye position spatial distribution map. The two types of modal features achieve feature interaction and integration through the self-attention mechanism in the fusion module, and finally, the recognition of the calibration state and deviation type is completed through the multi-task decision module.

[0116] (1) The first encoding module and its working process:

[0117] The first encoding module adopts the Vision Transformer architecture. The input image is first divided into multiple image patches of equal size, and then mapped into a sequence of patch vectors through linear embedding; a learnable classification token [CLS] is added to the head of the sequence, and a position encoding vector is added to each patch to preserve the spatial order information. The processed sequence is input into a multi-layer Transformer encoder to extract the global dynamic features of the trajectory map. Preferably, this module includes 8 layers of Transformer Blocks, and each layer includes a multi-head attention substructure and a feed-forward network to fully explore the dynamic behavior features of the trajectory map. The following specifically elaborates the process with examples as follows:

[0118] ① Input and segmentation of the trajectory map:

[0119] In this embodiment, the first encoding module is mainly used to extract the features of the eye movement trajectory map; this first encoding module is based on the VIT architecture and uses the trajectory map as the input. In the first encoding module, the input image is first divided into multiple image patches (Embedded Patches) of equal size and converted into embedding vectors. Taking the input image with 3 channels (RGB), a height of 220, and a width of 1220 as an example, the input image R 3×220×1220 is divided into 20*20 patches (image blocks), and the size of each patch is 20*20; calculate the number of patches N patches as:

[0120]

[0121] ② Construction of the embedding vector of the image patch:

[0122] Each image patch obtains an embedding vector P through the linear projection matrix We i :

[0123] Among them:

[0124] The i-th patch (image block), in the embodiment of the present application, i = 1, 2,... 671;

[0125] Convert the i-th patch into a one-dimensional vector;

[0126] W e : Linear projection matrix, which maps it to the embedding space. For example, in this embodiment, W e ∈R(20*20*3)*768, which means mapping the patch to a 768-dimensional feature space;

[0127] Pi ∈R 768 : represents the embedding vector of the patch.

[0128] ③ Add positional encoding and classification token:

[0129] The positional encoding adds a learnable [CLS] token at the beginning of the embedding vector sequence and combines it with the standard ViT (Vision Transformer) positional encoding. Among them, the constructed embedding sequence E:

[0130] Where:

[0131] e CLS represents the classification token vector (learnable), placed at the front;

[0132] E pos represents the positional encoding vector of the standard ViT;

[0133] P represents the embedding vector, respectively representing the 1st to the Nth patches embedding vectors; in this embodiment, N patches = 671, and the final sequence length is 672 (= 671 + 1)

[0134] ④ Transformer encoder:

[0135] Adopt multiple layers of Transformer encoders (8 layers of Transformer Block are adopted in this embodiment), and each layer of Transformer Block performs multi-head attention MHA and feed-forward network FFN operations:

[0136] Z′ j = MHA(LayerNorm(Z i-1 )) + Z i-1 ;

[0137] Z j = FFN(LayerNorm(Z′ j )) + Z′ j ;

[0138] Where: MHA: multi-head attention mechanism;

[0139] FFN: feed-forward neural network;

[0140] Z j : the output of the jth layer; in this embodiment, j = 1, 2,..., 8;

[0141] LayerNorm represents the normalization operation.

[0142] The computational decomposition of multi-head attention (e.g., 12 heads), where the attention of a single head is expressed as:

[0143]

[0144] Z: The input sequence;

[0145] respectively represent the query, key, and value weights of the t-th attention head;

[0146] d k : The dimension of the key vector (scaling factor);

[0147] softmax ensures the normalization of attention weights;

[0148] This is the Scaled Dot-Product Attention mechanism.

[0149] After concatenating all the attention heads (taking 12 heads as an example):

[0150] MSA(Z) = Concat(head1,..., head 12 )W out

[0151] where: W out represents the linear transformation matrix for the multi-head output, which is used to map the vector after concatenating multiple attention heads back to the unified output dimension of the Transformer layer. In this embodiment, if 12 attention heads are used and the output dimension of each attention head is 64, then the dimension after concatenation is 12 * 64 = 768, and the output linear transformation matrix W out ∈R 768*768 ; the output dimension of the transformed vector is also 768.

[0152] ⑤ Extract the final trajectory global feature:

[0153] Finally, extract the vector corresponding to the first position (i.e., the [CLS] token) from the Transformer output sequence as the global behavior feature F of the entire eye movement trajectory map track :

[0154] F track = Z L [CLS] ∈ R 768 ; where:

[0155] Z L : Represents the output sequence of the last layer of the Transformer;

[0156] [CLS] represents the position of the first classification token in the output sequence of the last layer of the Transformer; it is used to extract the global feature representation of the entire input in the Vision Transformer;

[0157] R 768 represents a real vector space of dimension 768, F track is an embedding vector in this space.

[0158] (2) The second encoding module and its working process:

[0159] In this embodiment, the second encoding module adopts an improved convolutional network structure, such as the ConvNext architecture + a custom adapter, such as ConvNextAdapter. This module adopts a progressive spatio-temporal feature extraction strategy. First, it extracts high-dimensional representative features through convolution (preferably an asymmetric convolution kernel), and successively undergoes a multi-stage downsampling process (preferably four stages) to achieve spatial dimension compression and channel expansion. Each downsampling stage contains multiple convolutional modules and normalization structures for extracting multi-scale spatial distribution features. The extracted multi-channel feature maps are globally average pooled to form a fixed-dimensional vector representation, which is further linearly projected onto a unified feature space for alignment and fusion with the trajectory features.

[0160] The second encoding module in this embodiment can adopt a histogram encoding module. By designing a depthwise separable convolution adapter ConvNextAdapter, it uses the above progressive spatio-temporal feature extraction strategy for feature extraction; specifically, the working process is as follows:

[0161] ① Feature enhancement is performed on the input second type of data (such as a spatial eye position distribution map) through convolution to extract deeper features. As Figure 7 shown in ConvNextAdapter, initial convolution extraction is now performed through Conv2D. By performing an initial shallow convolution operation on the original eye position spatial distribution map X hist ∈ R B×C×H×W the feature map of the first stage is obtained:

[0162] X stem = Stem(X hist ) ∈ R B×64×74×1220 ; where,

[0163] X hist represents the input of the original spatial distribution map (histogram);

[0164] Stem(X hist ): Initial convolution module;

[0165] X stem: The output shallow feature map,

[0166] R B×C×H×W represents a real vector space of dimension B * 64 * 74 * 1220; where B represents the batch size (e.g., B = 16 or 32); 64 is the number of channels after convolution, and the feature size of the spatial distribution map is 74×1220.

[0167] In this step, the input layer captures high-dimensional features through convolution, preferably through an asymmetric convolution kernel, such as an asymmetric convolution kernel (7, 3), and converts the input image into a feature map with a higher number of channels.

[0168] ② Multi-stage feature abstraction.

[0169] In this embodiment, a multi-stage downsampling structure is adopted to achieve feature abstraction through spatial compression and channel expansion, where the depthwise separable convolution block optimizes the parameter efficiency. Preferably, it includes a downsampling structure with four stages, and each stage contains at least one ConvNeXt Block and a downsampling operation. For the convenience of understanding, the following only expands and describes the feature transformation formula with the first two stages as an example:

[0170] First-stage downsampling: Further compress the spatial dimension:

[0171] X1 = Downsample1(X stem ) ∈ R 128*74*610

[0172] Second-stage downsampling:

[0173] X2 = Downsample2(X1) ∈ R 256*19*1220

[0174] where: X1 and X2 respectively represent the intermediate feature maps obtained after the downsampling stage. It can be seen that the number of channels gradually increases to 256 after these two-stage samplings to enhance the semantic capacity.

[0175] ③ Through global pooling, convert the feature map into a vector.

[0176] Perform global average pooling (Global Average Pooling) on the feature map X2 to convert the spatial dimension into a vector:

[0177] F hist-raw = GlobalAvgPool(X2) ∈ R B×256 ;

[0178] where F hist-raw represents the original spatial distribution feature vector.

[0179] ④Align the dimensions of the linear layer through feature projection.

[0180] Project the vector with a dimension of 256 to the dimension (768) consistent with the trajectory feature:

[0181] F hist = hist_proj(F hist-raw ) ∈ R B×768 ; where:

[0182] hist_proj(·): represents the linear feature mapping function (Linear Projection), which is used to map the low-dimensional feature vector to the high-dimensional feature space.

[0183] F hist-raw represents the original spatial distribution map feature vector (unprojected) obtained after global average pooling.

[0184] F hist : represents the output vector after projecting the spatial distribution map feature.

[0185] (3) Feature fusion module and its working process:

[0186] The feature fusion module is used to fuse the two-modal feature vectors extracted and output by the first encoding module and the second encoding module, realize modal interaction, and provide a unified feature representation for the final discrimination task. This module first projects the features of the two modalities to a unified high-dimensional space through linear mapping, preferably 1536 dimensions, to meet the input requirements of the subsequent multi-head attention mechanism. Subsequently, the two types of projected vectors are concatenated to form a sequence feature, and fusion processing is performed through a two-stream cross-attention mechanism including multiple layers of Transformer encoders. Finally, an attention pooling layer is used to perform weighted aggregation on the fusion sequence to obtain the final fusion feature representation, which is used as the input for subsequent multi-task discrimination. The specific implementation process is as follows:

[0187] ①Feature mapping, unified dimension:

[0188] In this embodiment, the feature fusion module projects the features output by the first encoding module and the second encoding module to a higher-dimensional space (such as a 1536-dimensional extended space) respectively:

[0189] T = track_proj(F track ) ∈ R B×1536 ;

[0190] H = hist_proj(F hist ) ∈ R B×1536 ;

[0191] where track_proj and hist_proj are both linear layers;

[0192] Ftrack is the encoded vector of the eye movement trajectory map, with a dimension of B * 768;

[0193] Fhist is the encoded vector of the eye position spatial distribution map, with a dimension of B * 768;

[0194] T and H are the projected feature vectors respectively, with a dimension of B * 1536.

[0195] In this step, the features are mapped to a richer expression space, and in order to unify the feature dimensions of the two modalities, it is convenient to perform concatenation (Concat) and fusion (such as Multi-Head Attention) subsequently.

[0196] ② Feature fusion:

[0197] First, the two-modal features are concatenated into a token sequence for multi-layer Transformer encoding;

[0198] C = concat([T, H], dim = 1) ∈ R B×2×1536 ; where:

[0199] C represents the token sequence after bimodal concatenation, with a dimension of B * 2 * 1536;

[0200] In this way, a token sequence with a length of 2 is formed, and each token is a modality; as the input of the multi-modal Transformer encoder.

[0201] Then, a multi-layer Transformer encoder is used for feature fusion:

[0202] In this embodiment, feature mixing is performed through 4 layers of Transformer encoders (8-head attention), and the attention mechanism is used to establish dependencies and interactions between modalities. Specifically, an attention pooling layer is introduced, with the cross-modal feature mean as the query vector to achieve adaptive feature aggregation and output a 1536-dimensional fused feature.

[0203] C′ = TransformerEncoder(C) ∈ R B×2×1536 ;

[0204] The above C′ represents the token sequence after modality fusion.

[0205] ③ Finally, the sequence after modality fusion is aggregated to form the final fused feature vector F fused :

[0206] F fused = MultiHeadAttention(C′, C′, C′) ∈ R B×1536 ;

[0207] In this embodiment, a self-attention mechanism is adopted, using the sequence itself as the query (Q), key (K), and value (V), so that the two modal tokens can learn information from each other.

[0208] In this embodiment, the fusion module uses the self-attention mechanism to interactively model the modal sequence. Specifically, the trajectory features and the gaze distribution features are concatenated to form a fused token sequence C′, which is used as the query, key, and value input of the multi-head attention mechanism. By calculating the attention weights between different modalities, the adaptive fusion of semantic features is achieved. The final output is the fused unified representation F fused , used for subsequent multi-task discrimination.

[0209] (4) Multi-task decision-making module and its workflow:

[0210] Through the shared feature processing layer and multiple task-specific branch structures, the fused features are subjected to multi-objective joint discrimination and multi-dimensional recognition results are output. The shared feature processing layer in this module is used to unify the fused feature representation and generate shared discrimination information through normalization and activation functions; multiple task branches correspond to different discrimination targets, including but not limited to: calibration success judgment, undershoot / overshoot classification, gaze error regression, etc. Each task branch contains several fully connected layers and output layers, supporting binary classification, regression or multi-classification output. All task output results will be spliced to form a structured recognition vector, which is used to comprehensively evaluate the quality and accuracy of the eye movement calibration test. In terms of specific structural implementation, the shared feature processing layer can realize feature reduction through LayerNorm and 512-dimensional GELU activation; multiple task-specific branch structures can use two parallel branches-4 binary classifiers using a bilinear layer structure (256-dimensional latent space). The final output layer splices 4×2=8-dimensional prediction values to support end-to-end multi-objective joint optimization.

[0211] ① Reduce features through feature sharing layer:

[0212] S=Shared(F fused )∈R B×512 ;

[0213] Usually includes LayerNorm, GELU activation, and linear layer; used to extract shared discriminant information.

[0214] ② Handle different tasks through multi-task specific branch structure:

[0215] Specifically, an independent branch structure can be set for each task, including classification tasks and / or regression tasks, such as:

[0216] Three types of binary classification tasks (each output is 2 dimensions)

[0217] Output calculation process for the head of the first task branch:

[0218] Where:

[0219] is the output result of the first binary classification task (e.g., whether the calibration is successful);

[0220] binary_class_heads[1] represents the first task branch among multiple binary classifiers (index starting from 1);

[0221] S is the shared feature representation, coming from the previous layer feature reduction module, with dimension B×512;

[0222] ∈R B×2 , indicating that the output result is a 2D vector for each sample (used for softmax to determine two classes);

[0223] B is the batch size, representing the number of samples.

[0224] Each binary classification task in the multi-task decision module is implemented by an independent task branch structure. As shown in the output calculation formula for the head of the first task branch above, the first task branch sends the shared feature vector S into the binary classification head structure binary_class_heads[1] exclusive to the first task for processing, and outputs the classification scores of each sample under two classes for judging the binary classification result of this task (e.g., whether the calibration is successful). The calculation of the remaining binary classification tasks is similar and will not be elaborated here.

[0225] For the regression task, a regression task branch is adopted. Taking the task of predicting the calibration angle offset as an example, the same shared feature S is input, and one or more continuous real values are output. We can design it as:

[0226] O reg = regression_head(S) ∈R B×D ; Where:

[0227] O reg represents the regression prediction output;

[0228] regression_head represents the regression head module, generally an MLP composed of linear layers;

[0229] D is the output dimension. For example, D = 1 represents a single value (such as the offset angle), and D>1 represents multi-dimensional regression (such as predicting for the left / right eyes separately);

[0230] B represents the batch size (number of samples).

[0231] This model realizes the alignment of heterogeneous feature spaces through parametric projection, and uses the global attention mechanism of Transformer to overcome the modality deviation problem in traditional fusion methods. A specially designed asymmetric convolutional downsampling strategy effectively handles the spatial heterogeneity of distribution map data while maintaining the integrity of temporal features. The multi-task weight sharing mechanism significantly reduces the number of model parameters while ensuring the prediction accuracy.

[0232] ③ Concatenate the prediction results of each task branch to generate a structured recognition result.

[0233] Finally, a structured result is output, such as including 3 binary classifications (whether the calibration is successful, whether there is an underrun and / or whether there is an overshoot) + 1 regression value (calibration eye position offset).

[0234] 4. Evaluation Metrics:

[0235] During the training iteration process, the accuracy (accuracy, %), sensitivity (sensitivity, %), specificity (specificity, %), F1 score, and the area under the receiver operating characteristic curve (Area Under the Receiver Operating Characteristic Curve, AUC) are used to evaluate the prediction results.

[0236] 4.1 Accuracy is the proportion of the number of correctly predicted samples to the total number of samples:

[0237]

[0238] TP: True Positive, TN: True Negative, FP: False Positive, FN: False Negative

[0239] 4.2 Sensitivity is the proportion of correctly identified positive samples:

[0240]

[0241] 4.3 Specificity is the proportion of correctly identified negative samples:

[0242]

[0243] 4.4 The calculation formula for the F1 score (F1 Score) is as follows:

[0244]

[0245] 4. Model Training and Results

[0246] (1) Experimental Configuration: This embodiment is implemented using the Pytorch library, and the model is trained using the GPU of NVIDIA RTX3090. The training uses a batch size of 8, calculates the loss, sets the initial learning rate to 6e-5, linearly increases the learning rate in the first 5 epochs until it reaches the maximum value of 3e-4 at the 5th epoch, and then adopts the cosine annealing strategy to decay periodically according to the cosine function to avoid gradient oscillation caused by too large a learning rate.

[0247] (2) Evaluation Results

[0248] Results of Annotated Data: After data annotation, in this embodiment, 1280 cases of data are divided into a training set, a tuning set, and a test set in a ratio of 8:1:1. A vestibular function calibration recognition model with a two-stream Transformer + ConvNext architecture is used for model training, and the manually annotated results are used as the gold standard for comparison. There are 1024 cases in the training set, 56.25% of the data is calibrated successfully, and 45.13% of the data has an underflow.

[0249] (3) Model Evaluation Results:

[0250] To verify the effectiveness of the intelligent recognition model described in the present invention in the determination of the results of the eye movement calibration test, the calibration cooperation determination ability of the model at different fixation positions (middle, left, right) was systematically evaluated, mainly including indicators such as accuracy, sensitivity, specificity, F1 score, MCC (Matthews correlation coefficient), and AUC (area under the ROC curve). Some results are as Figure 8 shown.

[0251] From the evaluation results, the model shows extremely high discrimination ability in the calibration judgments at the three fixation positions:

[0252] The accuracy of the calibration data in the middle is the highest, reaching 97.7% (95% confidence interval is 93.3% - 99.2%);

[0253] The accuracy of the leftward calibration is slightly lower, at 94.5% (95% confidence interval 89.1% - 97.3%);

[0254] Among the overall judgment performance indicators, the indicators with the best performance of the model are as follows:

[0255] The sensitivity is as high as 99.0% (95% CI: 94.4% - 99.8%)

[0256] The specificity is the highest at 96.9% (95% CI: 84.3% - 99.4%)

[0257] The highest F1 score was 96.5% (95% CI: 93.8% - 99.0%).

[0258] The highest MCC was 84.4% (95% CI: 72.3% - 94.5%).

[0259] In addition, as can be observed from the ROC curve shown in Figure 8 , the AUC of the intelligent recognition model in the calibration recognition tasks at each position was close to 0.99 (95% confidence interval was 0.98 - 1.00), showing almost perfect classification and discrimination ability.

[0260] Further analyzing the model's recognition ability for abnormal calibration behaviors - especially the under - shoot waveform, the results showed that although it was slightly lower than the cooperation judgment, it still had high practical value:

[0261] The accuracy rate was 87.5% (95% CI: 80.7% - 92.2%).

[0262] The sensitivity was 89.7% (95% CI: 79.2% - 95.2%).

[0263] The specificity was 85.7% (95% CI: 75.7% - 92.1%).

[0264] The F1 score was 86.7% (95% CI: 79.6% - 92.6%).

[0265] The MCC was 75.1% (95% CI: 63.3% - 86.2%).

[0266] The AUC was 0.93 (95% CI: 0.88 - 0.97).

[0267] The above results indicate that the model described in the present invention shows excellent recognition performance in various task scenarios, especially outstanding in the recognition of calibration cooperation status, and also has strong robustness in the recognition of under - shoot anomalies, verifying the feasibility and effectiveness of the proposed multi - modal fusion structure and multi - task output strategy in the actual eye movement calibration test scenario.

[0268] In another embodiment of the present application, based on any of the above - mentioned embodiments, the input data input to the intelligent recognition model further includes visual target auxiliary data; wherein, the acquisition of the visual target auxiliary data includes:

[0269] Acquiring data related to the visual target display in the calibration test;

[0270] Based on the data related to the visual target display, constructing the visual target auxiliary data, including:

[0271] The third type of data, which is used to represent the position change data of the visual target display; for example, the presentation coordinates of the visual target at different time points, the path sequence, etc.; and

[0272] The fourth type of data, which is used to represent the spatial distribution characteristics of the visual target display position; for example, the density heat map of the target points, the occurrence frequency map, the statistical histogram, etc.

[0273] In this embodiment, in view of the fact that the input data of the model has added visual target auxiliary data to improve the accuracy of eye movement calibration recognition, the intelligent recognition model in this embodiment can complete the structural integration process of the above visual target auxiliary data in one of the following two ways:

[0274] (1) Adopt the multi-branch structure expansion method

[0275] The intelligent recognition model further includes: a third encoding module and a fourth encoding module, where:

[0276] The third encoding module is used to encode and model the third type of data, and extract the behavioral features representing the dynamic changes of the visual target;

[0277] The fourth encoding module is used to extract spatial features from the fourth type of data, and extract the spatial distribution characteristics presented by the visual target.

[0278] The feature fusion module performs unified mapping and fusion processing on the feature vectors output by the first to fourth encoding modules to obtain the final multi-modal integration - fusion feature representation, which is then used for the recognition output of the subsequent multi-task decision module.

[0279] (2) Adopt the input-level fusion method

[0280] In another implementation, the intelligent recognition model still adopts the original double-branch encoding structure, that is, it only includes the first encoding module and the second encoding module. However, in this implementation, the following processing needs to be performed before the input data is input into the model:

[0281] Pre-fuse the third type of data with the first type of data to construct combined trajectory-visual target input data, and input it into the first encoding module for feature extraction;

[0282] Pre-fuse the fourth type of data with the second type of data to construct combined eye position-target spatial distribution input data, and input it into the second encoding module for feature extraction.

[0283] The above fusion can be achieved through methods such as channel superposition, image stitching, position alignment, etc., which are used to enrich the input semantics of each encoding module and improve the context integrity of the encoding representation.

[0284] Similarly, the subsequent feature fusion module further fuses the feature vectors output by the first encoding module and the second encoding module to generate a fused feature representation, and performs multi-object joint discrimination on the fused features through the task decision module to output multi-dimensional recognition results.

[0285] In this embodiment, the visual target display information is introduced as an auxiliary input, providing a target reference for the calibration task of the model, which helps the model to more accurately judge whether the eye movement behavior is consistent with the visual target guidance, thereby effectively improving the discrimination accuracy of various abnormalities such as calibration success, overshoot, and undershoot.

[0286] Another embodiment of the present application discloses a computer program product, including: computer-executable instructions or a computer program, which, when executed by a processor, implement the method for eye movement calibration recognition based on artificial intelligence in any of the above embodiments.

[0287] Another embodiment of the present application discloses an eye movement calibration recognition device based on artificial intelligence, which includes:

[0288] An information interface module, configured to receive eye movement input data, where the eye movement input data includes:

[0289] The first type of data, used to represent the change in the eye movement trajectory;

[0290] The second type of data, used to represent the spatial distribution characteristics of the eye position;

[0291] A data processor, coupled to the information interface module, where the data processor is provided with a trained intelligent recognition model, and the intelligent recognition model is constructed based on a deep neural network structure, and its model structure is the same as the intelligent recognition model in any of the above embodiments;

[0292] The data processor is configured to execute the following instructions:

[0293] Input the first type of data and the second type of data into the intelligent recognition model;

[0294] Output the result of eye movement calibration through the intelligent recognition model; where the result includes whether the eye movement calibration is successful, and / or whether there is an undershoot, and / or whether there is an overshoot.

[0295] Another embodiment of the present application discloses an eye movement calibration recognition system based on artificial intelligence, including:

[0296] An eye movement acquisition module, used to acquire the eye movement data of the subject during the eye movement calibration process;

[0297] A data construction module, used to construct at least two different dimensions of eye movement input data based on the eye movement data, including:

[0298] The first type of data is used to represent the change in the eye movement trajectory;

[0299] The second type of data is used to represent the spatial distribution characteristics of the eye position;

[0300] The intelligent recognition module is used to recognize and output the result of eye movement calibration according to the received first type of data and the second type of data through a trained intelligent recognition model; where:

[0301] The intelligent recognition model includes:

[0302] The first encoding module is used to extract the dynamic trajectory behavior characteristics from the input first type of data to obtain a global behavior feature vector;

[0303] The second encoding module is used to extract the static spatial characteristics from the input second type of data to obtain a spatial distribution feature vector;

[0304] The feature fusion module uses a two-stream cross-attention mechanism to cross-fuse the global behavior feature vector and the spatial distribution feature vector to generate a fused feature representation;

[0305] The multi-task decision module includes a shared feature processing layer and multiple task branch structures, and is used to perform multi-object joint discrimination on the fused features and output multi-dimensional recognition results.

[0306] For the model architecture of the intelligent recognition model in this embodiment, reference may also be made to the model architecture of any of the foregoing embodiments. To avoid repetition, it will not be elaborated here.

[0307] It should be noted that the above embodiments can be freely combined as needed. The above are only the preferred embodiments of the present invention. It should be pointed out that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can still be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. An eye movement calibration and recognition method based on artificial intelligence, characterized in that, Comprising: Obtaining eye movement data of a subject during eye movement calibration; Based on the eye movement data, constructing input data of at least two different dimensions, including: The first type of data, used to represent the change in the eye movement trajectory; The second type of data, used to represent the spatial distribution characteristics of the eye position; Inputting the first type of data and the second type of data into a trained intelligent recognition model, the intelligent recognition model is constructed based on a deep neural network structure and has at least two branch encoding structures for respectively extracting the features of input data of different dimensions; the recognition model also includes a cross-modal feature fusion mechanism for performing fusion processing on the feature vectors output by each branch encoding structure; Identifying and outputting the result of eye movement calibration through the intelligent recognition model; wherein, the result of the eye movement calibration includes whether the eye movement calibration is successful, and / or whether there is an undershoot, and / or whether there is an overshoot.

2. The artificial intelligence-based eye movement calibration recognition method according to claim 1, wherein The first type of data includes any one of images, sequences, and vector data used to describe the change in the eye movement trajectory; The second type of data includes any one of density maps, heat maps, position histograms, and coordinate statistical vectors used to describe the spatial distribution of eye positions or fixation points.

3. The eye movement calibration and recognition method based on artificial intelligence according to claim 1, characterized in that Processing the input data through the intelligent recognition model to identify and output the result of eye movement calibration through the following steps: Extracting dynamic trajectory behavior features from the first type of data to obtain a global behavior feature vector; Extracting static spatial features from the second type of data to obtain a spatial distribution feature vector; Using a cross-attention mechanism to cross-fuse the global behavior feature vector and the spatial distribution feature vector to generate a fused feature representation; Performing multi-objective joint discrimination on the fused feature through a shared feature processing layer and multiple task-specific branch structures to output a multi-dimensional recognition result.

4. The eye movement calibration and recognition method based on artificial intelligence according to claim 3, characterized in that The extracting dynamic trajectory behavior features from the first type of data to obtain a global behavior feature vector specifically includes; Dividing the input first type of data into multiple data blocks; Converting each data block into an embedding vector; Adding learnable position encoding and a classification marker vector representing overall feature information to each embedding vector sequence. Using a multi-layer encoder, including a multi-head attention mechanism and a feed-forward network structure, to perform global feature extraction on the embedding vector sequence containing position encoding and classification markers; Extracting the classification marker vector from the feature vectors output by the multi-layer encoder as the global behavior feature vector of the trajectory map.

5. The eye movement calibration and recognition method based on artificial intelligence according to claim 3, characterized in that The extracting static spatial features from the second type of data to obtain a spatial distribution feature vector specifically includes: Enhancing the features of the input second type of data through convolution; Performing multi-stage downsampling through a depthwise separable convolution structure and a channel expansion structure to extract multi-scale spatial features; Converting the multi-scale spatial features into feature vectors of a unified dimension; Performing a linear transformation on the pooled feature vectors to output a spatial distribution feature vector of a unified dimension.

6. The eye movement calibration and recognition method based on artificial intelligence according to claim 3, characterized in that, The cross-attention mechanism is adopted to cross-fuse the global behavior feature vector and the spatial distribution feature vector to generate a fused feature representation, which specifically includes: Mapping the trajectory behavior feature vector and the spatial distribution feature vector into a feature space of a unified dimension respectively; Interactively fusing the two types of features after unifying the dimension through a multi-layer Transformer encoder; Based on the output of the fusion unit, extracting the global fusion feature through an attention pooling layer and outputting the fused feature representation.

7. The eye movement calibration and recognition method based on artificial intelligence according to claim 3, wherein, The fused feature representation is subjected to multi-object joint discrimination through a shared feature processing layer and multiple task-specific branch structures to output a multi-dimensional recognition result, which specifically includes: Performing unified regularization on the fused feature representation to extract shared discriminant information; Based on the shared discriminant information, different recognition tasks are executed by adopting multiple task-specific branch structures; wherein, the multiple task-specific branch structures include multiple classification task branch structures and / or at least one regression task branch structure; wherein: each classification branch structure includes a group of bilinear layers for performing binary classification recognition operations for different classification tasks respectively; the classification tasks include calibration state recognition, overshoot recognition, and undershoot recognition; the regression task branch structure includes at least one fully connected layer and an output layer for regression prediction, and is used to output continuous numerical results of the fixation position or calibration error based on the fused feature representation; The prediction results of each task branch are concatenated to generate a structured recognition result, including whether the calibration is successful, whether there is an undershoot, and / or whether there is an overshoot.

8. The eye movement calibration and recognition method based on artificial intelligence according to any one of claims 1-7, characterized in that, Both the first type of data and the second type of data are image data, and at least two branch encoding structures of the intelligent recognition model are a Vision Transformer network structure and an improved ConvNext convolutional network respectively.

9. The eye movement calibration and recognition method based on artificial intelligence according to claim 1, wherein The input data input to the intelligent recognition model further includes visual target auxiliary data; wherein, the acquisition of the visual target auxiliary data includes: Acquiring data related to the visual target display in the calibration experiment; Based on the data related to the visual target display, constructing visual target auxiliary data, including: The third type of data, which is used to represent the position change data of the visual target display; and The fourth type of data, which is used to represent the spatial distribution feature of the visual target display position.

10. A computer program product, comprising: A computer-executable instruction or a computer program, characterized in that when the computer-executable instruction or the computer program is executed by a processor, the eye movement calibration recognition method based on artificial intelligence according to any one of claims 1 to 9 is implemented.

11. An eye movement calibration and recognition device based on artificial intelligence, characterized in that, Including: An information interface module configured to receive eye movement input data, where the eye movement input data includes: The first type of data, which is used to represent the change of the eye movement trajectory; The second type of data, which is used to represent the spatial distribution feature of the eye position; A data processor, coupled to the information interface module, is provided with a trained intelligent recognition model. The intelligent recognition model is constructed based on a deep neural network structure and has at least two branch coding structures for respectively extracting features of different types of input data. The recognition model further includes a cross-modal feature fusion mechanism for fusing the feature vectors output by each branch coding structure. The data processor is configured to execute the following instructions: Input the first type of data and the second type of data into the intelligent recognition model; Output the result of eye movement calibration through the intelligent recognition model; wherein the result includes whether the eye movement calibration is successful, and / or whether there is an undershoot, and / or whether there is an overshoot.

12. An eye movement calibration and recognition system based on artificial intelligence, characterized in that, Comprising: An eye movement acquisition module for acquiring eye movement data of a subject during the eye movement calibration process; A data construction module for constructing at least two different-dimensional eye movement input data based on the eye movement data, including: The first type of data for representing the change in the eye movement trajectory; The second type of data for representing the spatial distribution characteristics of the eye position; An intelligent recognition module for identifying and outputting the result of eye movement calibration according to the received first type of data and the second type of data through the trained intelligent recognition model; wherein: The intelligent recognition model includes: A first coding module for extracting dynamic trajectory behavior features from the input first type of data to obtain a global behavior feature vector; A second coding module for extracting static spatial features from the input second type of data to obtain a spatial distribution feature vector; A feature fusion module that uses a two-stream cross-attention mechanism to cross-fuse the global behavior feature vector and the spatial distribution feature vector to generate a fused feature representation; A multi-task decision module, including a shared feature processing layer and multiple task branch structures, for performing multi-objective joint discrimination on the fused features and outputting multi-dimensional recognition results.

Citation Information

Cited By

  • Eye movement detection method and system based on multimode data fusion and dynamic visual target adjustment

    CN122074962A