Bionic visual thinking chain cross-modal device, visualization system and method based on eye tracking
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-24
- Publication Date
- 2026-08-11
AI Technical Summary
[0013]本发明的目的是解决现有视觉思维链模型缺乏真实视觉思维过程模拟、难以建立符合真实人类视觉思维模式的视觉思维链的技术问题,而提供一种基于眼动追踪的仿生视觉思维链跨模态模型、可视化系统及方法
[0077]1、本发明提供的基于眼动追踪的仿生视觉思维链跨模态模型,采用待测图像眼动注视区域图像块及其注视顺序和注视时长,其来源于反映人类真实视觉认知的眼动追踪数据,可真实反映人类视觉认知的动态特性和时序模式,使大语言模型按照人类眼动轨迹的顺序逐步处理视觉信息,模拟人类真实视觉思维过程,从而建立符合真实人类视觉思维模式的视觉思维链;
Smart Images

Figure CN121523540B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to visual thought chains, specifically to a biomimetic visual thought chain cross-modal model, visualization system, and method based on eye-tracking. Background Technology
[0002] With the rapid development of artificial intelligence technology, large language models have achieved remarkable results in tasks such as visual understanding, question answering, and reasoning. By integrating multiple input modalities such as text and images, they can perform complex cognitive tasks including logical reasoning, causal inference, and analogy mapping.
[0003] Chain-of-Thought (CoT) is a significant breakthrough in the field of large language models in recent years. It enhances the reasoning ability of models by generating explicit intermediate reasoning steps. With the extension of cross-modal technology to the image domain, the concept of Visual Chain-of-Thought has emerged, aiming to demonstrate the reasoning path of large language models when processing visual information.
[0004] The Multimodal-CoT (Multimodal Chain-of-Thought Reasoning in Language Models) proposed by Zhuosheng Zhang et al. in Transactions on Machine Learning Research, 2024, became a foundational work for cross-modal thought chain models. It incorporates visual information into the thought chain model framework through a two-stage training strategy (reason generation → answer reasoning). Significant progress was achieved in benchmark tests such as ScienceQA (Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering, a science question answering dataset from NeurIPS 2022 (The 36th Conference on Neural Information Processing Systems)), demonstrating the feasibility of cross-modal thought chain models. Subsequently, various methods emerged, each exploring different technical paths, such as:
[0005] The VoT model (Vision-of-Thought, from arXiv:2404.03622, 2024, Microsoft Research) simulates the cognitive process of the human mind, generating mental images during reasoning and tracking changes in visual states using a visual-spatial canvas, generating a corresponding visual representation after each reasoning step. However, its drawback lies in the fact that the generated mental images are entirely created by the algorithm, lacking support from objective human cognitive data, and therefore cannot be verified as conforming to real human visual thinking patterns.
[0006] The IPVR model (Interactive Prompting for Visual Reasoning, from See, Think, Confirm: Interactive Prompting Between Vision and Language Models for Knowledge-based Visual Reasoning / arXiv:2301.05226, 2023) designs a three-stage structured framework: See-Think-Confirm. The See stage involves overall observation and key element identification of the input image; the Think stage performs logical reasoning and relational analysis based on the observation results; and the Confirm stage verifies the reasoning results and makes a final decision. While this progressive approach enhances the structure of reasoning, it primarily relies on text prompts to guide the reasoning process and lacks a deep understanding mechanism for visual content.
[0007] The G-CoT model (Graph-guided Chain-of-Thought, from Graph Chain-of-Thought: Augmenting Large Language Models by Reasoning on Graphs / ACL 2024, arXiv:2404.07103) models the reasoning process as a graph structure, where nodes represent reasoning steps and edges represent logical relationships. This method integrates visual, historical, and linguistic signals, generating decisions through graph information propagation. While it has advantages in scenarios requiring multi-moment information fusion, such as autonomous driving, the graph construction process is complex, computationally expensive, and struggles to handle large-scale, dynamically changing reasoning tasks.
[0008] While the aforementioned models have made progress in cross-modal understanding and reasoning, they still fall short of the fundamental understanding of human visual cognition:
[0009] (1) Lack of realistic simulation of visual thinking process: When humans understand complex visual information, they form a series of orderly visual thinking patterns, from overall perception to local analysis, from main elements to secondary details. This thinking process has a clear temporal sequence and hierarchy. However, the above model cannot truly reflect the dynamic characteristics and temporal patterns of human visual cognition. In essence, it is still a product of computer algorithms, rather than a biomimetic simulation based on real human cognitive data.
[0010] (2) Lack of genuine visual reasoning ability: The models mentioned above all lack step-by-step reasoning analysis of image content and cannot perform deep reasoning through the dynamic shift of visual attention like humans. More importantly, these models have not actually established a visual thought chain that conforms to the real human visual thinking pattern. Their main goal is to improve the model's performance on cross-modal tasks by imitating a certain reasoning process, rather than truly building a reasoning mechanism based on visual cognition. Human visual thinking is not entirely based on language and textual thinking, such as the process by which humans acquire visual information about faces, dynamic movements, etc.
[0011] (3) Fundamental limitations of bounding box annotation: Currently widely used bounding box annotation methods have multiple limitations. Traditional bounding boxes can only annotate entities with clearly defined boundaries, and they perform extremely poorly for uncountable nouns and non-fixed-shape objects, such as wind, flowing coffee, rising steam, drifting smoke, and flowing water. These phenomena are all abstract information, but they are extremely important in daily visual understanding and cannot be accurately described by bounding boxes. At the same time, bounding boxes cannot accurately describe irregular shapes, gradient areas, or scattered visual elements. In addition, bounding boxes can only identify locations and cannot express the degree of human cognitive investment in different areas, nor can they effectively express the depth and importance of different thought processes.
[0012] (4) Lack of visual thinking temporal modeling: Human visual thinking processes have a clear temporal order and logical hierarchy. When understanding complex images, humans perform visual analysis according to specific cognitive strategies, forming a thought sequence with clear logic. However, although the above model introduces a multi-step processing mechanism, it fails to capture the true visual thinking temporal sequence. Summary of the Invention
[0013] The purpose of this invention is to solve the technical problems of existing visual thought chain models lacking simulation of real visual thinking processes and making it difficult to establish visual thought chains that conform to real human visual thinking patterns, and to provide a biomimetic visual thought chain cross-modal model, visualization system and method based on eye tracking.
[0014] To achieve the above objectives, the technical solution adopted by the present invention is as follows:
[0015] A biomimetic visual thought chain cross-modal model based on eye tracking is characterized by including an input module, a first image encoder, a second image encoder, a first multilayer perceptron, a second multilayer perceptron, a third multilayer perceptron, a fourth multilayer perceptron, a fifth multilayer perceptron, a large language model, an image-text alignment module, and an output module.
[0016] The output of the input module is connected to the input of the first image encoder, the second image encoder, and the first input of the large language model, respectively. The input module is used to receive the image to be tested, the image block of the eye movement fixation region of the image to be tested and its fixation sequence and fixation duration, and the text question of the image to be tested, and transmit them to the first image encoder, the second image encoder, and the large language model, respectively.
[0017] The output of the first image encoder is connected to the input of the first multilayer perceptron and the second multilayer perceptron, respectively, for extracting features of the image to be tested; the output of the first multilayer perceptron is connected to the first input of the image-text alignment module, for mapping the features of the image to be tested to a feature space compatible with the image-text alignment module; the output of the second multilayer perceptron is connected to the second input of the large language model, for mapping the features of the image to be tested to a feature space compatible with the large language model.
[0018] The output of the second image encoder is connected to the input of the third multilayer perceptron, which is used to extract the features of the eye-tracking gaze region image patch of the image under test, and fuse the features according to the gaze order and gaze duration to obtain the fused features of the eye-tracking gaze region image patch of the image under test; the output of the third multilayer perceptron is connected to the third input of the large language model, which is used to map the fused features of the eye-tracking gaze region image patch of the image under test to a feature space compatible with the large language model;
[0019] The first output of the large language model is connected to the input of the fourth multilayer perceptron. The output of the fourth multilayer perceptron is connected to the second input of the image-text alignment module. The output of the image-text alignment module is connected to the input of the fifth multilayer perceptron. The output of the fifth multilayer perceptron is connected to the fourth input of the large language model. The second and third outputs of the large language model are connected to the first and second inputs of the output module, respectively. The large language model is used to obtain the predicted text response of the image under test, as well as the eye-tracking vector prediction matrix and the eye-tracking predicted text response. The output module is used to output the eye-tracking vector prediction matrix and the eye-tracking predicted text response of the image under test, thereby obtaining the biomimetic visual thought chain and the text thought chain.
[0020] Furthermore, both the first image encoder and the second image encoder are image feature extraction networks;
[0021] The first, second, third, fourth, and fifth multilayer perceptrons all employ residual connections and layer normalization structures.
[0022] Furthermore, the first image encoder and the second image encoder employ CLIP or ViT;
[0023] The large language model is GPT or LLaMA.
[0024] The present invention also provides a bionic visual thought chain visualization system based on eye tracking, which is characterized by including an eye tracker for acquiring eye tracking data of the image to be tested, a preprocessing module, the above-mentioned bionic visual thought chain cross-modal model based on eye tracking, and a visual thought chain visualization module.
[0025] The output of the eye tracker is connected to the first input of the preprocessing module. The second input of the preprocessing module is used to receive the image to be tested and the text question of the image to be tested. The output is connected to the input of the input module in the bionic visual thinking chain cross-modal model based on eye tracking. The preprocessing module is used to obtain the eye-tracking gaze region image block, its gaze order and gaze duration, and the eye-tracking vector matrix from the image to be tested based on the eye-tracking data of the image to be tested. It then transmits these, along with the image to be tested and the text question of the image to be tested, to the input of the input module in the bionic visual thinking chain cross-modal model based on eye tracking.
[0026] In the eye-tracking-based bionic visual thinking chain cross-modal model, the output end of the output module is connected to one input end of the visual thinking chain visualization module. The other input end of the visual thinking chain visualization module is used to receive the image to be tested, and the visual thinking chain visualization module is used to visualize the bionic visual thinking chain on the image to be tested.
[0027] Furthermore, the sampling frequency of the eye tracker is ≥60Hz.
[0028] This invention also provides a biomimetic visual thought chain visualization method based on eye tracking. The method, based on the aforementioned biomimetic visual thought chain visualization system based on eye tracking, is characterized by including the following steps:
[0029] Step 1: Select a visual reasoning dataset including original images, text questions, and text answers. Then, use an eye tracker to acquire eye-tracking data of the original images in the visual reasoning dataset. Transfer the data along with the original images to the preprocessing module for preprocessing to obtain the eye-tracking gaze region image patches, their gaze order, gaze duration, and eye-tracking vector matrix corresponding to the original images. Annotate the eye-tracking vector matrix and the text answer on the corresponding original images. Then, use the original images, the corresponding text questions, the eye-tracking gaze region image patches, their gaze order, and gaze duration as the eye-tracking dataset and input it into the biomimetic visual thinking chain cross-modal model based on eye-tracking.
[0030] Step 2: Based on eye tracking, the bionic visual thinking chain cross-modal model learns the features of an original image, the corresponding eye-tracking gaze region image patch, gaze order, and gaze duration from the eye-tracking dataset, and uses these features to reason about text questions to obtain eye-tracking predicted text answers and eye-tracking vector prediction matrices.
[0031] Step 3: Calculate the loss function based on the eye-tracking predicted text answer, the eye-tracking vector prediction matrix and the corresponding text answer, and the eye-tracking vector matrix. Then, adjust the parameters of the first multilayer perceptron, the second multilayer perceptron, the third multilayer perceptron, the fourth multilayer perceptron, the fifth multilayer perceptron and the image-text alignment module in the eye-tracking-based bionic visual thought chain cross-modal model according to the loss function. Then, return to Step 2 and train the eye-tracking-based bionic visual thought chain cross-modal model using the next original image and the corresponding text question, the eye-tracking gaze region image patch and its gaze order and gaze duration in the eye-tracking dataset until the loss function converges.
[0032] Step 4: Use an eye tracker to collect eye tracking data of the image to be tested, and then transmit it along with the image to be tested to the preprocessing module; The preprocessing module extracts the eye-tracking region image blocks, their gaze order and gaze duration from the image to be tested based on the eye tracking data of the image to be tested, and transmits them along with the image to be tested and the text question of the image to be tested to the input module of the trained bionic visual thought chain cross-modal model based on eye tracking;
[0033] Step 5: The trained eye-tracking-based bionic visual thinking chain cross-modal model learns the features of the test image, the eye-tracking gaze region image block of the test image, and its gaze sequence and gaze duration. Based on these features, it infers the text question of the test image to obtain the eye-tracking predicted text answer and eye-tracking vector prediction matrix of the test image. Then, the eye-tracking vector prediction matrix of the test image is transmitted to the visual thinking chain visualization module.
[0034] Step 6: Transmit the image to be tested to the visual thinking chain visualization module. The visual thinking chain visualization module draws the gaze region on the image to be tested according to the eye movement vector prediction matrix of the image to be tested, and marks the row number of the gaze region in the eye movement vector prediction matrix at the center of the gaze region, thus obtaining the image to be tested containing the visual thinking chain and completing the visualization of the bionic visual thinking chain.
[0035] Furthermore, step 2 specifically involves:
[0036] Step 2.1: The input module transmits an original image from the eye-tracking dataset to the first image encoder, transmits the eye-tracking gaze region image block corresponding to the original image, along with its gaze order and gaze duration, to the second image encoder, and transmits the text question corresponding to the original image to the large language model;
[0037] Step 2.2: The first image encoder extracts the features of the original image and transmits them to the first multilayer perceptron and the second multilayer perceptron respectively; the first multilayer perceptron maps the features of the original image to a feature space compatible with the image-text alignment module to obtain the original image alignment mapping features and transmits them to the image-text alignment module; the second multilayer perceptron maps the features of the original image to a feature space compatible with the large language model to obtain the original image model mapping features and transmits them to the large language model.
[0038] The second image encoder extracts features from the eye-tracking fixation region image patch, and then performs feature fusion based on the fixation order and fixation duration of the eye-tracking fixation region image patch to obtain the eye-tracking fixation region image patch fused features, which are then transmitted to the third multilayer perceptron; the third multilayer perceptron maps the eye-tracking fixation region image patch fused features to a feature space compatible with the large language model to obtain the eye-tracking fixation region image patch mapped features, which are then transmitted to the large language model.
[0039] Step 2.3: The large language model infers the text question based on the mapping features of the original image model, obtains the text answer, and transmits it to the fourth multilayer perceptron; the fourth multilayer perceptron maps the text answer to a feature space compatible with the image-text alignment module, obtains the text answer mapping features, and transmits them to the image-text alignment module.
[0040] Step 2.4: The image-text alignment module aligns and fuses the original image alignment mapping features and text response mapping features to obtain image-text alignment features and transmits them to the fifth multilayer perceptron; the fifth multilayer perceptron maps the image-text alignment features to a feature space compatible with the large language model to obtain image-text alignment mapping features and transmits them to the large language model.
[0041] Step 2.5: Based on the original image model mapping features and eye-tracking fixation region image patch mapping features obtained in Step 2.2, and the image-text alignment mapping features obtained in Step 2.4, the large language model performs reasoning on the text question to obtain the eye-tracking predicted text answer and eye-tracking vector prediction matrix, and transmits them to the output module for output.
[0042] Further, in step 3, the loss function is calculated using the following formula:
[0043]
[0044] in, For loss function, The visual question-answering loss is calculated based on eye-tracking predicted text responses and text answers. The eye-tracking prediction loss is calculated based on the eye-tracking vector prediction matrix and the eye-tracking vector matrix. The thought chain constraint loss is calculated based on the eye-tracking vector prediction matrix and the eye-tracking vector matrix; α, β, and γ are the weights of the question-answering loss, eye-tracking prediction loss, and thought chain constraint loss, respectively.
[0045] Furthermore, step 1 specifically includes:
[0046] Step 1.1: Select a visual reasoning dataset from the public dataset that includes the original images, text questions, and text answers. Then, select multiple original images from the visual reasoning dataset in an equalized manner and randomly and uniformly group them.
[0047] Step 1.2: Set the eye-tracking calibration accuracy threshold and the gaze angle speed threshold;
[0048] Step 1.3: Invite testers and perform eye-tracking calibration on each of them. If the eye-tracking calibration accuracy is less than or equal to the eye-tracking calibration accuracy threshold, have the testers read the text question corresponding to one set of original images and watch that set of original images. At the same time, use an eye tracker to record eye-tracking data. Then have the testers answer the text question to obtain the eye-tracking data and the testers' answers. Otherwise, repeat the eye-tracking calibration.
[0049] Step 1.4: Check whether the tester's answer is correct based on the text answer. If it is correct, retain the corresponding eye-tracking data and the corresponding original image, text question and text answer. If not, remove the corresponding eye-tracking data and the corresponding original image, text question and text answer to obtain the initial eye-tracking dataset.
[0050] Step 1.5: Calculate the gaze velocity of all sampling points in each group of eye tracking data in the initial eye tracking dataset. If the gaze velocity of a sampling point is lower than the gaze velocity threshold, the sampling point is determined as a fixation point; otherwise, it is determined as a saccade. This way, the fixation points of each group of eye tracking data in the initial eye tracking dataset are obtained.
[0051] Step 1.6: Based on the fixation points of each group of eye-tracking data in the initial eye-tracking dataset, obtain M fixation regions, their fixation order, and fixation duration. Based on the M fixation regions and their fixation order, obtain the eye movement vector matrix of each group of eye-tracking data. Then, calculate the minimum bounding rectangle region of the M fixation regions in each group of eye-tracking data, where M is an integer and M≥1.
[0052] Step 1.7: Based on the minimum bounding rectangle of the M gaze regions in each set of eye-tracking data, extract M eye-tracking gaze region image patches from the corresponding original images, and label the eye-tracking vector matrix and text answer on the corresponding original images. Then, use the original images, the corresponding text questions, the eye-tracking gaze region image patches and their gaze order and gaze duration as the eye-tracking dataset, and input them into the bionic visual thought chain cross-modal model based on eye-tracking.
[0053] Furthermore, step 1.6 specifically includes:
[0054] Step 1.6.1: Determine whether adjacent fixation points in one set of eye-tracking data in the initial eye-tracking dataset meet the following conditions: time interval < 200ms and gaze angle < 5°. If so, merge them into one fixation point to obtain the merged fixation point; then obtain M fixation regions and their fixation order based on the merged fixation point.
[0055] Step 1.6.2: Count the number of fixation points in each of the M fixation regions, and then calculate the fixation duration of each of the M fixation regions based on the sampling frequency of the eye tracker.
[0056] Step 1.6.3: Assume that all M fixation regions are ellipses, and then solve for the elliptical equation parameters of the M fixation regions using the following formulas to obtain the elliptical equations of the M fixation regions:
[0057]
[0058] in, , Let x and y be the x and y coordinates of the i-th fixation point in the k-th fixation region, respectively, where k and i are integers, and 1 ≤ k ≤ M, 1 ≤ i ≤ N. k , Let A be the number of fixation points in the k-th fixation region. k Bk C k D k E k F k All of these are the ellipse equation parameters for the k-th gaze region;
[0059] Step 1.6.4: Calculate the geometric parameters of the center, semi-major axis, semi-minor axis, and rotation angle of the M gaze regions using the following formulas:
[0060]
[0061]
[0062]
[0063]
[0064]
[0065] in, , Let x and y be the geometric center coordinates of the k-th gaze region, respectively. , These are the major and minor semi-axis of the k-th fixation region, respectively. Let be the rotation angle of the k-th gaze region;
[0066] Step 1.6.5: Based on the geometric parameters of the M fixation regions—center, semi-major axis, semi-minor axis, and rotation angle—obtain the eye movement vectors for each of the M fixation regions. Then, the eye movement vectors are arranged from top to bottom according to the fixation order of the fixation areas, resulting in an eye movement vector matrix of size M×5. The data in the k-th row corresponds to the k-th gaze region.
[0067] Step 1.6.6: Calculate the minimum bounding rectangle region [x] for each of the M gaze regions using the following formula. mink ,y mink ,w k ,h k ]:
[0068]
[0069]
[0070]
[0071]
[0072]
[0073]
[0074] in, , Let x and y be the x-coordinates of the left and right boundaries of the minimum bounding rectangle of the k-th gaze region, respectively. , These are the ordinates of the lower and upper boundaries of the minimum bounding rectangle of the k-th gaze region, respectively. , These are the width and height of the minimum bounding rectangle of the k-th gaze region, respectively;
[0075] Step 1.6.7: Repeat steps 1.6.1-1.6.6 until the minimum bounding rectangle of the M gaze regions in each group of eye-tracking data in the initial eye-tracking dataset is obtained.
[0076] Compared with the prior art, the present invention has the following beneficial effects:
[0077] 1. The biomimetic visual thinking chain cross-modal model based on eye tracking provided by this invention uses image blocks of the eye-tracking gaze region of the image under test, as well as their gaze order and gaze duration. These are derived from eye-tracking data that reflects real human visual cognition. This model can truly reflect the dynamic characteristics and temporal patterns of human visual cognition, enabling the large language model to process visual information step by step according to the sequence of human eye movement trajectories, simulating the real human visual thinking process, thereby establishing a visual thinking chain that conforms to the real human visual thinking pattern.
[0078] 2. The bionic visual thought chain cross-modal model based on eye tracking provided by this invention can obtain the eye movement vector prediction matrix and the eye movement prediction text answer, thereby simultaneously outputting the bionic visual thought chain and the text thought chain, improving the interpretability of the thought process after the fusion of text and image information.
[0079] 3. The biomimetic visual thought chain cross-modal model based on eye tracking provided by this invention uses a multilayer perceptron with residual connections and layer normalization structure, which can improve the stability and expressive power of feature mapping.
[0080] 4. The biomimetic visual thinking chain visualization system based on eye tracking provided by this invention visualizes the visual thinking process on the original image to be tested through the visual thinking chain visualization module.
[0081] 5. The bionic visual thought chain visualization system based on eye tracking provided by this invention has an eye tracker sampling frequency of ≥60Hz, which can ensure the capture of detailed features of human visual cognition.
[0082] 6. The bionic visual thought chain visualization method based on eye tracking provided by this invention is based on real eye tracking data. It uses an eye tracking dataset containing image blocks of eye-tracking gaze regions and their gaze order and gaze duration to train a cross-modal model of the bionic visual thought chain. The eye tracking vector matrix is integrated into the visual thought chain. Through real human eye tracking data, it guides the construction of the visual thinking process of the cross-modal model of the bionic visual thought chain based on eye tracking.
[0083] 7. The bionic visual thinking chain visualization method based on eye tracking provided by this invention uses the eye tracking vector matrix to preserve the gaze order, which can reflect the human visual analysis strategy from the whole to the part and from the primary to the secondary. It also uses the gaze duration to reflect the degree of human cognitive investment in different gaze areas, enabling the cross-modal model of the bionic visual thinking chain based on eye tracking to learn the real path and thinking logic of human visual analysis, thereby improving the interpretability and accuracy of cross-modal reasoning.
[0084] 8. The biomimetic visual thought chain visualization method based on eye tracking provided by this invention uses ellipses to describe the gaze area, which can handle uncountable nouns and non-fixed shape objects, accurately describe dynamic phenomena such as wind, flowing coffee, and rising hot air, solves the fundamental problem that traditional bounding boxes cannot accurately describe irregular shapes and dynamic phenomena, and significantly improves the ability to express visual information. Attached Figure Description
[0085] Figure 1 This is a schematic diagram of the structure of an embodiment of the biomimetic visual thought chain cross-modal model based on eye tracking of the present invention;
[0086] Figure 2 This is a schematic diagram of an original image from the GQA dataset selected in step 1.1 of the embodiment of the biomimetic visual thought chain visualization method based on eye tracking of the present invention;
[0087] Figure 3 This is a flowchart of step 1.3 in an embodiment of the biomimetic visual thought chain visualization method based on eye tracking of the present invention;
[0088] Figure 4 This is a schematic diagram of step 1.3, obtaining eye-tracking data, in an embodiment of the biomimetic visual thought chain visualization method based on eye tracking of the present invention.
[0089] Figure 5 This is a schematic diagram of the gaze region obtained in step 1.6.1 of the embodiment of the bionic visual thought chain visualization method based on eye tracking of the present invention. In the figure, 1-6 are the serial numbers of the 6 gaze regions, indicating the gaze order.
[0090] Figure 6This is the eye-tracking gaze region image block obtained in step 1.7 of the embodiment of the biomimetic visual thought chain visualization method based on eye tracking of the present invention, wherein (a)-(f) respectively correspond to Figure 5 The fixation regions numbered 1-6 in the middle;
[0091] Figure 7 This is a schematic diagram illustrating the training of the cross-modal model of the bionic visual thought chain based on eye tracking in steps 2-3 of the embodiment of the eye-tracking-based bionic visual thought chain visualization method of the present invention. Detailed Implementation
[0092] The following detailed description, in conjunction with the accompanying drawings and specific embodiments, provides a further detailed account of the biomimetic visual thought chain cross-modal model, visualization system, and method based on eye tracking proposed in this invention. Those skilled in the art should understand that these embodiments are merely illustrative of the technical principles of this invention and are not intended to limit the scope of protection of this invention.
[0093] A biomimetic visual thought chain cross-modal model based on eye tracking, such as Figure 1 As shown, it includes an input module, a first image encoder, a second image encoder, a first multilayer perceptron, a second multilayer perceptron, a third multilayer perceptron, a fourth multilayer perceptron, a fifth multilayer perceptron, a large language model, an image-text alignment module, and an output module.
[0094] The output of the input module is connected to the inputs of the first image encoder, the second image encoder, and the first input of the large language model, respectively. The input module is used to receive the image to be tested, the eye-tracking fixation region image block of the image to be tested and its fixation sequence and duration, and the text question of the image to be tested, and transmit them to the first image encoder, the second image encoder, and the large language model, respectively.
[0095] The output of the first image encoder is connected to the inputs of the first and second multilayer perceptrons, respectively, for extracting features from the image under test. The output of the first multilayer perceptron is connected to the first input of the image-text alignment module, for mapping the features of the image under test to a feature space compatible with the image-text alignment module. The output of the second multilayer perceptron is connected to the second input of the large language model, for mapping the features of the image under test to a feature space compatible with the large language model.
[0096] The output of the second image encoder is connected to the input of the third multilayer perceptron. It is used to extract features of the eye-tracking fixation region image patch in the image under test, and to fuse these features according to the fixation sequence and fixation duration to obtain the fused features of the eye-tracking fixation region image patch in the image under test. The output of the third multilayer perceptron is connected to the third input of the large language model. It is used to map the fused features of the eye-tracking fixation region image patch in the image under test to a feature space compatible with the large language model.
[0097] The first output of the large language model is connected to the input of the fourth multilayer perceptron. The output of the fourth multilayer perceptron is connected to the second input of the image-text alignment module. The output of the image-text alignment module is connected to the input of the fifth multilayer perceptron. The output of the fifth multilayer perceptron is connected to the fourth input of the large language model. The second and third outputs of the large language model are connected to the first and second inputs of the output module, respectively. The large language model is used to obtain the predicted text response for the image under test, as well as the eye-tracking vector prediction matrix and the eye-tracking predicted text response. The output module is used to output the eye-tracking vector prediction matrix and the eye-tracking predicted text response for the image under test, thereby obtaining the biomimetic visual thought chain and the text thought chain.
[0098] The first and second image encoders are both image feature extraction networks, employing CLIP (Contrastive Language-Image Pre-training, by Alec Radford, Jong Wook Kim, Chris Hallacy, et al. Learning Transferable Visual Models From Natural Language Supervision. OpenAI, arXiv:2103.00020, 2021) or ViT (Vision Transformer, by Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, et al. Vision Transformer (An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale), Google Research, ICLR 2021, arXiv:2010.11929). The first, second, third, fourth, and fifth multilayer perceptrons all utilize residual connections and layer normalization structures. Large language models are GPT (Generative Pre-trained Transformer, by Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever. Improving Language Understanding by Generative Pre-Training. OpenAI, 2018) or LLaMA (Large Language Model Meta AI, by Hugo Touvron, Thibaut Lavril, Gautier Izacard, et al. LLaMA: Open and Efficient Foundation Language Models. Meta AI, arXiv:2302.13971,2023).
[0099] Human cognition is characterized by multi-step visual reasoning. Large language models, when handling visual question-answering tasks involving images and text, lack modeling of the human visual reasoning process. Eye-tracking technology, based on high-speed optical imaging, can accurately capture eye movements and measure the eye's focus point, i.e., the location and object being looked at. Eye-tracking technology can accurately record human visual behavior, providing an important window into understanding human visual cognition. The biomimetic visual thought chain cross-modal model based on eye-tracking provided in this embodiment uses eye-tracking data to construct a biomimetic visual thought chain, which can enhance the understanding and simulation of human visual cognitive processes, enabling large language models to learn realistic human visual thought patterns.
[0100] This embodiment also provides a biomimetic visual thought chain visualization system based on eye tracking, including an eye tracker for acquiring eye tracking data of the image to be tested, a preprocessing module, the aforementioned biomimetic visual thought chain cross-modal model based on eye tracking, and a visual thought chain visualization module.
[0101] The output of the eye tracker is connected to the first input of the preprocessing module. The second input of the preprocessing module receives the image to be tested and the text question associated with the image. Its output is connected to the input of the input module in the aforementioned eye-tracking-based bionic visual thought chain cross-modal model. The preprocessing module extracts the eye-tracking fixation region image patch, its fixation sequence and duration, and the eye-tracking vector matrix from the image to be tested based on the eye-tracking data. It then transmits these, along with the image to be tested and the text question, to the input of the input module in the eye-tracking-based bionic visual thought chain cross-modal model. The output of the output module in the eye-tracking-based bionic visual thought chain cross-modal model is connected to one input of the visual thought chain visualization module. The other input of the visual thought chain visualization module receives the image to be tested, and the visual thought chain visualization module visualizes the bionic visual thought chain on the image to be tested.
[0102] In this embodiment, the sampling frequency of the eye tracker is ≥60Hz, and either Tobii (Tobii AB, Sweden) or EyeLink (EyeLink, SR Research Ltd., Canada) is used.
[0103] This embodiment also provides a biomimetic visual thought chain visualization method based on eye tracking, which uses the above-mentioned biomimetic visual thought chain visualization system based on eye tracking and includes the following steps:
[0104] Step 1: Select a visual reasoning dataset including original images, text questions, and text answers. Then, use an eye tracker to acquire eye-tracking data of the original images in the visual reasoning dataset. Transfer this data, along with the original images, to the preprocessing module for preprocessing. This yields the eye-tracking gaze region image patches corresponding to the original images, their gaze order and duration, and the eye-tracking vector matrix. Annotate the eye-tracking vector matrix and the text answer onto the corresponding original images. Then, use the original images, the corresponding text questions, the eye-tracking gaze region image patches, their gaze order, and gaze duration as the eye-tracking dataset and input it into the biomimetic visual thought chain cross-modal model based on eye-tracking. Specifically:
[0105] Step 1.1: Select a visual reasoning dataset from the public dataset that includes the original images, text questions, and text answers. Then, select multiple original images from the visual reasoning dataset in an equalized manner and randomly and uniformly group them.
[0106] In this embodiment, the visual reasoning dataset selected is the GQA dataset (GQA: A New Dataset for Real-World Visual Reasoning and Compositional Question Answering, from Drew A. Hudson, Christopher D. Manning. 2019 IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), Stanford University, arXiv:1902.09506). The GQA dataset is a large-scale dataset for visual question answering tasks, proposed by Stanford University in 2019, aiming to advance the capabilities of artificial intelligence in visual reasoning and compositional question answering. The GQA dataset contains 113,000 images and 22.66 million questions generated from these images, covering diverse reasoning requirements with varying question lengths and required reasoning steps. This dataset contains a vocabulary of 3097 words and 1878 possible answers. Each image is associated with a scene graph that details the objects, attributes, and relationships. The question types cover complex tasks such as object recognition, spatial reasoning, and logical comparison. Figure 2 The image shown is an original image from the GQA dataset, and the corresponding text question is: How many types of chicks are there in the image?
[0107] Based on the GQA dataset, 10,000 original images were selected evenly from different image categories and randomly divided into 10 groups. Ten testers were invited to view the original images, with each tester viewing one of the groups.
[0108] Step 1.2: Set the eye-tracking calibration accuracy threshold to 1° and the gaze angle speed threshold to 20° / s.
[0109] Step 1.3: Invite testers, such as... Figure 3 As shown, eye-tracking calibration was performed on each image. If the eye-tracking calibration accuracy was ≤1°, the tester was asked to read the text question corresponding to one set of original images. Figure 4 As shown, the original set of images is displayed on the screen for the tester to view, and eye-tracking data is recorded using an eye tracker. The tester then answers a text question to obtain the eye-tracking data and the tester's answer; otherwise, eye-tracking calibration is performed again.
[0110] In step 1.3, text questions must be provided in both Chinese and English; when viewing the original image, the tester must confirm that the viewing has been completed or that the viewing time has reached 30 seconds.
[0111] Step 1.4: Check whether the tester's answer is correct based on the text answer. If it is correct, retain the corresponding eye-tracking data and the corresponding original image, text question and text answer. If not, remove the corresponding eye-tracking data and the corresponding original image, text question and text answer to obtain the initial eye-tracking dataset.
[0112] Step 1.5: Calculate the angular velocity of the gaze at all sampling points in each group of eye-tracking data in the initial eye-tracking dataset. If the angular velocity of a sampling point is less than 20° / s, the sampling point is determined as a fixation point; otherwise, it is determined as a saccade. This process yields the fixation point for each group of eye-tracking data in the initial eye-tracking dataset. The angular velocity of the sampling point is... Calculated using the following formula:
[0113]
[0114] in, The time interval between two adjacent sampling points. The angle between the lines of sight of two adjacent sampling points.
[0115] Step 1.6: Based on the gaze points of each group of eye-tracking data in the initial eye-tracking dataset, obtain the following... Figure 5 The diagram shows six fixation regions, their fixation order, and fixation duration. Based on these six fixation regions and their fixation order, the eye movement vector matrix for each set of eye-tracking data is obtained. Then, the minimum bounding rectangle region for each of the six fixation regions in each set of eye-tracking data is calculated. Specifically:
[0116] Step 1.6.1: Determine whether adjacent fixation points in one set of eye-tracking data in the initial eye-tracking dataset meet the following conditions: time interval < 200ms and gaze angle < 5°. If so, merge them into one fixation point to obtain the merged fixation point; then, based on the merged fixation point, obtain the following... Figure 5 The six fixation areas and their fixation order are shown.
[0117] Step 1.6.2: Count the number of fixation points n in each of the 6 fixation regions, and then calculate the fixation duration t in each of the 6 fixation regions based on the sampling frequency f of the eye tracker, where t = n × f.
[0118] Step 1.6.3: Set all six gaze regions to be ellipses, where the equation of the ellipse is: And satisfy A, B, C, D, E, and F are all parameters of the ellipse equation; then, the ellipse equation parameters of the six fixation regions are solved using the following formulas to obtain the ellipse equations of the six fixation regions:
[0119]
[0120] in, , Let x and y be the x and y coordinates of the i-th fixation point in the k-th fixation region, respectively, where k and i are integers, and 1 ≤ k ≤ 6, 1 ≤ i ≤ N. k , Let A be the number of fixation points in the k-th fixation region. k B k C k D k E k F k All of these are the ellipse equation parameters for the k-th gaze region.
[0121] In step 1.6.3, the parameters of the fixation region ellipse equation can be obtained by using the Lagrange multiplier method or by directly solving the above formula. In other embodiments, the parameters of the fixation region ellipse equation can be obtained by regression, and the prediction accuracy can be optimized by the MSE loss function.
[0122] Step 1.6.4: Calculate the geometric parameters of the center, semi-major axis, semi-minor axis, and rotation angle of the six gaze regions using the following formulas:
[0123]
[0124]
[0125]
[0126]
[0127]
[0128] in, , Let x and y be the geometric center coordinates of the k-th gaze region, respectively. , These are the major and minor semi-axis of the k-th fixation region, respectively. Let be the rotation angle of the k-th gaze region.
[0129] Step 1.6.5: Based on the geometric parameters of the center, semi-major axis, semi-minor axis, and rotation angle of the six fixation regions, obtain the eye movement vectors for each of the six fixation regions. Then, the eye movement vectors are arranged from top to bottom according to the order of the fixation areas, resulting in a 6×5 eye movement vector matrix. The data in the kth row corresponds to the kth gaze region.
[0130] Step 1.6.6: Calculate the minimum bounding rectangle region [x] of each of the six gaze regions using the following formula. mink ,y mink ,w k ,h k ]:
[0131]
[0132]
[0133]
[0134]
[0135]
[0136]
[0137] in, , Let x and y be the x-coordinates of the left and right boundaries of the minimum bounding rectangle of the k-th gaze region, respectively. , These are the ordinates of the lower and upper boundaries of the minimum bounding rectangle of the k-th gaze region, respectively. , These are the width and height of the smallest bounding rectangle of the k-th gaze region, respectively.
[0138] Step 1.6.7: Repeat steps 1.6.1-1.6.6 until the minimum bounding rectangle of the six gaze regions in each group of eye-tracking data in the initial eye-tracking dataset is obtained.
[0139] Step 1.7: Based on the minimum bounding rectangle region of the six fixation areas in each set of eye-tracking data, extract the following from the corresponding original image: Figure 6 The six eye-tracking fixation region image patches shown are used to annotate the eye-tracking vector matrix and the text answer on the corresponding original images. Then, the original images, the corresponding text questions, the eye-tracking fixation region image patches and their fixation order and fixation duration are used as an eye-tracking dataset and input into the biomimetic visual thought chain cross-modal model based on eye-tracking.
[0140] The training process of the biomimetic visual thought chain cross-modal model based on eye tracking, namely steps 2-3, is as follows: Figure 7 As shown.
[0141] Step 2: Based on eye-tracking, the biomimetic visual thought chain cross-modal model learns the features of an original image, its corresponding eye-tracking gaze region image patch, its gaze sequence, and gaze duration from the eye-tracking dataset. Based on these features, it infers the text question, obtaining the eye-tracking predicted text answer and the eye-tracking vector prediction matrix. Specifically:
[0142] Step 2.1: The input module transmits an original image from the eye-tracking dataset to the first image encoder, transmits the eye-tracking gaze region image block corresponding to the original image, along with its gaze order and gaze duration, to the second image encoder, and transmits the text question corresponding to the original image to the large language model.
[0143] Step 2.2: The first image encoder extracts the features of the original image and transmits them to the first multilayer perceptron and the second multilayer perceptron respectively. The first multilayer perceptron maps the features of the original image to a feature space compatible with the image-text alignment module, obtains the original image alignment mapping features, and transmits them to the image-text alignment module. The dimension of the original image alignment mapping features is [1, m], where m is the feature dimension size of the image-text alignment module, which is a hyperparameter set according to the model architecture. The second multilayer perceptron maps the features of the original image to a feature space compatible with the large language model, obtains the original image model mapping features, and transmits them to the large language model.
[0144] The second image encoder extracts features from the eye-tracking fixation region image patch, and then performs feature fusion based on the fixation order and fixation duration of the eye-tracking fixation region image patch to obtain the fused features of the eye-tracking fixation region image patch, which is then transmitted to the third multilayer perceptron. The third multilayer perceptron maps the fused features of the eye-tracking fixation region image patch to a feature space compatible with the large language model to obtain the mapped features of the eye-tracking fixation region image patch, which is then transmitted to the large language model.
[0145] Step 2.3: The large language model infers the text question based on the mapping features of the original image model, obtains the text answer, and transmits it to the fourth multilayer perceptron; the fourth multilayer perceptron maps the text answer to a feature space compatible with the image-text alignment module, obtains the text answer mapping features, and transmits them to the image-text alignment module. The dimension of the text answer mapping features is [1, m].
[0146] Step 2.4: The image-text alignment module aligns and fuses the original image alignment mapping features and text response mapping features to obtain image-text alignment features and transmits them to the fifth multilayer perceptron. The feature dimension of the image-text alignment features is [2, m]. The fifth multilayer perceptron maps the image-text alignment features to a feature space compatible with the large language model to obtain image-text alignment mapping features and transmits them to the large language model.
[0147] Step 2.5: Based on the original image model mapping features and eye-tracking fixation region image patch mapping features obtained in Step 2.2, and the image-text alignment mapping features obtained in Step 2.4, the large language model performs reasoning on the text question, obtaining the eye-tracking predicted text answer and the eye-tracking vector prediction matrix, and transmits them to the output module for output. Among them, the eye-tracking vector prediction matrix is a biomimetic visual thought chain representing the reasoning process of the large language model.
[0148] Step 3: Based on the eye-tracking predicted text answer, the eye-tracking vector prediction matrix and the corresponding text answer, and the eye-tracking vector matrix, calculate the loss function. Then, based on the loss function, adjust the parameters of the first, second, third, fourth, and fifth multilayer perceptrons and the image-text alignment module in the eye-tracking-based bionic visual thought chain cross-modal model. Return to Step 2 and train the eye-tracking-based bionic visual thought chain cross-modal model using the next original image from the eye-tracking dataset, the corresponding text question, the eye-tracking gaze region image patch and its gaze order and gaze duration, until the loss function converges. The trained eye-tracking-based bionic visual thought chain cross-modal model is obtained. The loss function is calculated using the following formula:
[0149]
[0150] in, For loss function, The visual question answering loss is calculated based on the predicted text answer and the text answer. The eye-tracking prediction loss is calculated based on the eye-tracking vector prediction matrix and the eye-tracking vector matrix. The thought chain constraint loss is calculated based on the eye-tracking vector prediction matrix and the eye-tracking vector matrix. α, β, and γ are the weights of the question-answering loss, eye-tracking prediction loss, and thought chain constraint loss, respectively, and are all adjustable hyperparameters that can be optimized according to specific task requirements. In this embodiment, α, β, and γ are initially set to 0.3, 0.3, and 0.4, and are dynamically adjusted during training.
[0151] The loss functions during training include eye-tracking prediction loss, visual question-answering loss, and thought chain constraint loss. Specifically, the eye-tracking prediction loss uses KL divergence to calculate the difference between the predicted and actual eye-tracking distributions, yielding the loss between eye-tracking data and predicted eye-tracking data. The visual question-answering loss uses the text answer as a reference benchmark. The thought chain constraint loss employs a group-relative strategy optimization to dynamically adjust the strategy for multiple steps in the text-based reasoning process, combining policy gradient loss and KL divergence constraints to calculate the loss function for the multi-step inference process.
[0152] In this embodiment, end-to-end training provides a unified optimization objective, adjusts parameters according to the loss function, and backpropagates, which can balance the learning effects of multiple types of losses.
[0153] Step 4: Use an eye tracker to collect eye tracking data of the image to be tested, and then transmit it along with the image to be tested to the preprocessing module. The preprocessing module extracts the eye-tracking region image block, its gaze sequence, and gaze duration from the image to be tested based on the eye tracking data of the image to be tested, and transmits it along with the image to be tested and the text question of the image to be tested to the input module of the trained bionic visual thought chain cross-modal model based on eye tracking.
[0154] In other embodiments, the eye-tracking-based bionic visual thought chain cross-modal model supports online learning and incremental training, and can continuously learn from new eye-tracking data to improve the generation quality of the bionic visual thought chain.
[0155] Step 5: The trained eye-tracking-based bionic visual thinking chain cross-modal model learns the features of the test image, the eye-tracking gaze region image block of the test image, and its gaze sequence and gaze duration. Based on these features, it infers the text question of the test image to obtain the eye-tracking predicted text answer and eye-tracking vector prediction matrix of the test image. Then, the eye-tracking vector prediction matrix of the test image is transmitted to the visual thinking chain visualization module.
[0156] Step 6: Transmit the image to be tested to the visual thinking chain visualization module. The visual thinking chain visualization module draws the gaze region on the image to be tested according to the eye movement vector prediction matrix of the image to be tested, and marks the row number of the gaze region in the eye movement vector prediction matrix at the center of the gaze region, thus obtaining the image to be tested containing the visual thinking chain and completing the visualization of the bionic visual thinking chain.
[0157] In other embodiments, the visualization of the visual thought chain can be achieved by overlaying semi-transparent ellipses and displaying the gaze areas sequentially according to the gaze order, thus intuitively presenting the visual thought process.
[0158] When humans process textual and image information, they engage in reasoning processes. On one hand, they reason about textual information; on the other hand, when viewing images, the brain performs reasoning processes to interpret the images. This reasoning process is more difficult to explain than that of text, but it can be represented by eye-tracking trajectories. In this embodiment, eye-tracking fixation region image patches, their fixation order, and fixation duration are used to enable a large language model to understand the human reasoning process for images. Specifically, the human observation process of an image is represented by an eye-tracking vector matrix, which includes the fixation region and fixation order. During training, eye-tracking fixation region image patches, their fixation order, and fixation duration, obtained from real-world eye-tracking data, are input into a bionic visual thought chain cross-modal model based on eye-tracking. During reasoning, the bionic visual thought chain cross-modal model based on eye-tracking outputs the calculated bionic visual thought chain, and the visual reasoning process can be visualized on the image.
[0159] The data flow of the reasoning process is divided into two stages, such as... Figure 7 As shown, the first stage involves aligning the text and images. Figure 7 The blue line in the middle represents two parts: text problem → large language model → fourth multilayer perceptron → image-text alignment module; and original image → first image encoder → first multilayer perceptron → image-text alignment module. The second stage performs deep inference based on alignment features and eye-tracking information. Figure 7 (The red line in the middle) consists of five parts: First image encoder → Second multilayer perceptron → Large language model; Eye-tracking fixation region image block and its fixation sequence and fixation duration → Second image encoder → Third multilayer perceptron → Large language model; Image-text alignment module → Fifth multilayer perceptron → Large language model; Large language model → Eye-tracking predicted text answer; Large language model → Eye-tracking vector prediction matrix → Bionic visual thought chain.
Claims
1. A biomimetic visual thought chain cross-modal device based on eye tracking, characterized in that: It includes an input module, a first image encoder, a second image encoder, a first multilayer perceptron, a second multilayer perceptron, a third multilayer perceptron, a fourth multilayer perceptron, a fifth multilayer perceptron, a large language model, an image-text alignment module, and an output module; The output of the input module is connected to the input of the first image encoder, the second image encoder, and the first input of the large language model, respectively. The input module is used to receive the image to be tested, the image block of the eye-tracking fixation region of the image to be tested and its fixation sequence and fixation duration, and the text question of the image to be tested, and transmit them to the first image encoder, the second image encoder, and the large language model respectively; The output of the first image encoder is connected to the input of the first multilayer perceptron and the second multilayer perceptron, respectively, for extracting features of the image to be tested; the output of the first multilayer perceptron is connected to the first input of the image-text alignment module, for mapping the features of the image to be tested to a feature space compatible with the image-text alignment module; the output of the second multilayer perceptron is connected to the second input of the large language model, for mapping the features of the image to be tested to a feature space compatible with the large language model. The output of the second image encoder is connected to the input of the third multilayer perceptron, which is used to extract the features of the eye-tracking fixation region image block of the image under test, and fuse the features according to the fixation order and fixation duration to obtain the fused features of the eye-tracking fixation region image block of the image under test. The output of the third multilayer perceptron is connected to the third input of the large language model, and is used to map the fusion features of the eye-tracking fixation region of the image to be tested to a feature space compatible with the large language model. The first output of the large language model is connected to the input of the fourth multilayer perceptron. The output of the fourth multilayer perceptron is connected to the second input of the image-text alignment module. The output of the image-text alignment module is connected to the input of the fifth multilayer perceptron. The output of the fifth multilayer perceptron is connected to the fourth input of the large language model. The second and third outputs of the large language model are connected to the first and second inputs of the output module, respectively. The large language model is used to obtain the predicted text response of the image under test, as well as the eye-tracking vector prediction matrix and the eye-tracking predicted text response. The output module is used to output the eye-tracking vector prediction matrix and the eye-tracking predicted text response of the image under test, thereby obtaining the biomimetic visual thought chain and the text thought chain.
2. The biomimetic visual thought chain cross-modal device based on eye tracking according to claim 1, characterized in that: Both the first image encoder and the second image encoder are image feature extraction networks; The first, second, third, fourth, and fifth multilayer perceptrons all employ residual connections and layer normalization structures.
3. The biomimetic visual thought chain cross-modal device based on eye tracking according to claim 2, characterized in that: The first and second image encoders employ CLIP or ViT. The large language model is GPT or LLaMA.
4. A biomimetic visual thought chain visualization system based on eye tracking, characterized in that: The invention includes an eye tracker for acquiring eye-tracking data of an image to be tested, a preprocessing module, a bionic visual thought chain cross-modal device based on eye tracking as described in any one of claims 1-3, and a visual thought chain visualization module. The output of the eye tracker is connected to the first input of the preprocessing module. The second input of the preprocessing module is used to receive the image to be tested and the text question of the image to be tested. The output is connected to the input of the input module in the bionic visual thinking chain cross-modal device based on eye tracking. The preprocessing module is used to obtain the eye-tracking gaze region image block, its gaze order and gaze duration, and the eye-tracking vector matrix from the image to be tested based on the eye-tracking data of the image to be tested. It then transmits the image to be tested and the text question of the image to be tested to the input of the input module in the bionic visual thinking chain cross-modal device based on eye tracking. In the eye-tracking-based bionic visual thinking chain cross-modal device, the output end of the output module is connected to one input end of the visual thinking chain visualization module. The other input end of the visual thinking chain visualization module is used to receive the image to be tested, and the visual thinking chain visualization module is used to visualize the bionic visual thinking chain on the image to be tested.
5. The biomimetic visual thought chain visualization system based on eye tracking according to claim 2, characterized in that: The sampling frequency of the eye tracker is ≥60Hz.
6. A biomimetic visual thought chain visualization method based on eye tracking, based on the biomimetic visual thought chain visualization system based on eye tracking as described in claim 4 or 5, characterized in that, Includes the following steps: Step 1: Select a visual reasoning dataset including original images, text questions, and text answers. Then, use an eye tracker to acquire eye tracking data of the original images in the visual reasoning dataset. Transmit the data along with the original images to the preprocessing module for preprocessing to obtain the eye-tracking gaze region image blocks, their gaze order, gaze duration, and eye-tracking vector matrix corresponding to the original images. Label the eye-tracking vector matrix and the text answer on the corresponding original images. Then, use the original images, the corresponding text questions, the eye-tracking gaze region image blocks, their gaze order, and gaze duration as the eye-tracking dataset and input it into the bionic visual thinking chain cross-modal device based on eye tracking. Step 2: The bionic visual thinking chain cross-modal device based on eye tracking learns the features of an original image, the corresponding eye-tracking gaze region image patch, its gaze sequence, and gaze duration from the eye-tracking dataset, and uses these features to reason about text questions to obtain eye-tracking predicted text answers and eye-tracking vector prediction matrices. Step 3: Calculate the loss function based on the eye-tracking predicted text answer, the eye-tracking vector prediction matrix and the corresponding text answer, and the eye-tracking vector matrix. Then, adjust the parameters of the first multilayer perceptron, the second multilayer perceptron, the third multilayer perceptron, the fourth multilayer perceptron, the fifth multilayer perceptron and the image-text alignment module in the eye-tracking-based bionic visual thinking chain cross-modal device according to the loss function. Then, return to Step 2 and use the next original image and the corresponding text question, the eye-tracking gaze region image patch and its gaze order and gaze duration in the eye-tracking dataset to train the eye-tracking-based bionic visual thinking chain cross-modal device until the loss function converges, and obtain the trained eye-tracking-based bionic visual thinking chain cross-modal device. Step 4: Use an eye tracker to collect eye-tracking data of the image to be tested, and then transmit it along with the image to be tested to the preprocessing module; The preprocessing module extracts the eye-tracking region image blocks, their gaze order, and gaze duration from the eye-tracking data of the image under test, and transmits them, along with the image under test and the text question in the image under test, to the input module of the trained eye-tracking-based bionic visual thought chain cross-modal device. Step 5: The trained eye-tracking-based bionic visual thinking chain cross-modal device learns the features of the test image, the eye-tracking gaze region image block of the test image, and its gaze sequence and gaze duration. Based on these features, it infers the text question of the test image to obtain the eye-tracking predicted text answer and eye-tracking vector prediction matrix of the test image. Then, it transmits the eye-tracking vector prediction matrix of the test image to the visual thinking chain visualization module. Step 6: Transmit the image to be tested to the visual thinking chain visualization module. The visual thinking chain visualization module draws the gaze region on the image to be tested according to the eye movement vector prediction matrix of the image to be tested, and marks the row number of the gaze region in the eye movement vector prediction matrix at the center of the gaze region, thus obtaining the image to be tested containing the visual thinking chain and completing the visualization of the bionic visual thinking chain.
7. The biomimetic visual thought chain visualization method based on eye tracking according to claim 6, characterized in that, Step 2 is as follows: Step 2.1: The input module transmits an original image from the eye-tracking dataset to the first image encoder, transmits the eye-tracking gaze region image block corresponding to the original image, along with its gaze order and gaze duration, to the second image encoder, and transmits the text question corresponding to the original image to the large language model; Step 2.2: The first image encoder extracts the features of the original image and transmits them to the first multilayer perceptron and the second multilayer perceptron respectively; the first multilayer perceptron maps the features of the original image to a feature space compatible with the image-text alignment module to obtain the original image alignment mapping features and transmits them to the image-text alignment module; the second multilayer perceptron maps the features of the original image to a feature space compatible with the large language model to obtain the original image model mapping features and transmits them to the large language model. The second image encoder extracts features from the eye-tracking fixation region image patch, and then performs feature fusion based on the fixation order and fixation duration of the eye-tracking fixation region image patch to obtain the eye-tracking fixation region image patch fused features, which are then transmitted to the third multilayer perceptron; the third multilayer perceptron maps the eye-tracking fixation region image patch fused features to a feature space compatible with the large language model to obtain the eye-tracking fixation region image patch mapped features, which are then transmitted to the large language model. Step 2.3: The large language model infers the text question based on the mapping features of the original image model, obtains the text answer, and transmits it to the fourth multilayer perceptron; the fourth multilayer perceptron maps the text answer to a feature space compatible with the image-text alignment module, obtains the text answer mapping features, and transmits them to the image-text alignment module. Step 2.4: The image-text alignment module aligns and fuses the original image alignment mapping features and text response mapping features to obtain image-text alignment features, which are then transmitted to the fifth multilayer perceptron. The fifth multilayer perceptron maps the image-text alignment features to a feature space compatible with the large language model, obtains the image-text alignment mapping features, and transmits them to the large language model. Step 2.5: Based on the original image model mapping features and eye-tracking fixation region image patch mapping features obtained in Step 2.2, and the image-text alignment mapping features obtained in Step 2.4, the large language model performs reasoning on the text question to obtain the eye-tracking predicted text answer and eye-tracking vector prediction matrix, and transmits them to the output module for output.
8. The biomimetic visual thought chain visualization method based on eye tracking according to claim 6 or 7, characterized in that, In step 3, the loss function is calculated using the following formula: ; in, For loss function, The visual question-answering loss is calculated based on eye-tracking predicted text responses and text answers. The eye-tracking prediction loss is calculated based on the eye-tracking vector prediction matrix and the eye-tracking vector matrix. The thought chain constraint loss is calculated based on the eye-tracking vector prediction matrix and the eye-tracking vector matrix; α, β, and γ are the weights of the question-answering loss, eye-tracking prediction loss, and thought chain constraint loss, respectively.
9. The biomimetic visual thought chain visualization method based on eye tracking according to claim 8, characterized in that, Step 1 is as follows: Step 1.1: Select a visual reasoning dataset from the public dataset that includes the original images, text questions, and text answers. Then, select multiple original images from the visual reasoning dataset in an equalized manner and randomly and uniformly group them. Step 1.2: Set the eye-tracking calibration accuracy threshold and the gaze angle speed threshold; Step 1.3: Invite testers and perform eye-tracking calibration on them individually. If the eye-tracking calibration accuracy is less than or equal to the eye-tracking calibration accuracy threshold, have the testers read the text questions corresponding to one set of original images and watch the set of original images. At the same time, use an eye tracker to record eye-tracking data. Then have the testers answer the text questions to obtain the eye-tracking data and the testers' answers. Otherwise, recalibrate the eye-tracking. Step 1.4: Check whether the tester's answer is correct based on the text answer. If it is correct, retain the corresponding eye-tracking data and the corresponding original image, text question and text answer. If not, remove the corresponding eye-tracking data and the corresponding original image, text question and text answer to obtain the initial eye-tracking dataset. Step 1.5: Calculate the gaze velocity of all sampling points in each group of eye tracking data in the initial eye tracking dataset. If the gaze velocity of a sampling point is lower than the gaze velocity threshold, the sampling point is determined as a fixation point; otherwise, it is determined as a saccade. This way, the fixation points of each group of eye tracking data in the initial eye tracking dataset are obtained. Step 1.6: Based on the fixation points of each group of eye-tracking data in the initial eye-tracking dataset, obtain M fixation regions, their fixation order, and fixation duration. Based on the M fixation regions and their fixation order, obtain the eye movement vector matrix of each group of eye-tracking data. Then, calculate the minimum bounding rectangle region of the M fixation regions in each group of eye-tracking data, where M is an integer and M≥1. Step 1.7: Based on the minimum bounding rectangle of the M gaze regions in each set of eye-tracking data, extract M eye-tracking gaze region image blocks from the corresponding original images, and label the eye-tracking vector matrix and text answer on the corresponding original images. Then, use the original images, the corresponding text questions, the eye-tracking gaze region image blocks and their gaze order and gaze duration as the eye-tracking dataset, and input it into the bionic visual thinking chain cross-modal device based on eye-tracking.
10. The biomimetic visual thought chain visualization method based on eye tracking according to claim 9, characterized in that, Step 1.6 specifically involves: Step 1.6.1: Determine whether adjacent fixation points in one set of eye-tracking data in the initial eye-tracking dataset meet the following conditions: time interval < 200ms and gaze angle < 5°. If so, merge them into one fixation point to obtain the merged fixation point; then obtain M fixation regions and their fixation order based on the merged fixation point. Step 1.6.2: Count the number of fixation points in each of the M fixation regions, and then calculate the fixation duration of each of the M fixation regions based on the sampling frequency of the eye tracker. Step 1.6.3: Assume that all M fixation regions are ellipses, and then solve for the elliptical equation parameters of the M fixation regions using the following formulas to obtain the elliptical equations of the M fixation regions: ; in, , Let x and y be the x and y coordinates of the i-th fixation point in the k-th fixation region, respectively, where k and i are integers, and 1 ≤ k ≤ M, 1 ≤ i ≤ N. k , Let A be the number of fixation points in the k-th fixation region. k B k C k D k E k F k All of these are the ellipse equation parameters for the k-th gaze region; Step 1.6.4: Calculate the geometric parameters of the center, semi-major axis, semi-minor axis, and rotation angle of the M gaze regions using the following formulas: ; ; ; ; ; in, , Let x and y be the geometric center coordinates of the k-th gaze region, respectively. , These are the major and minor semi-axis of the k-th fixation region, respectively. Let be the rotation angle of the k-th gaze region; Step 1.6.5: Based on the geometric parameters of the M fixation regions—center, semi-major axis, semi-minor axis, and rotation angle—obtain the eye movement vectors for each of the M fixation regions. Then, the eye movement vectors are arranged from top to bottom according to the fixation order of the fixation areas, resulting in an eye movement vector matrix of size M×5. The data in the k-th row corresponds to the k-th gaze region. Step 1.6.6: Calculate the minimum bounding rectangle region [x] for each of the M gaze regions using the following formula. mink ,y mink ,w k ,h k ]: ; ; ; ; ; ; in, , Let x and y be the x-coordinates of the left and right boundaries of the minimum bounding rectangle of the k-th gaze region, respectively. , These are the ordinates of the lower and upper boundaries of the minimum bounding rectangle of the k-th gaze region, respectively. , These are the width and height of the minimum bounding rectangle of the k-th gaze region, respectively; Step 1.6.7: Repeat steps 1.6.1-1.6.6 until the minimum bounding rectangle of the M gaze regions in each group of eye-tracking data in the initial eye-tracking dataset is obtained.
Citation Information
Patent Citations
Non-perception MR glasses man-machine identification method, system and device, and storage medium
CN111966223A
Multi-modal big language model attribute prediction method based on multi-modal thinking chain
CN119693768A