Visual function detection method and system based on 3D visual imaging technology

Through multi-dimensional data fusion and intelligent algorithm analysis, a realistic virtual scene model is built, combined with deep learning and inertial sensor data, the accuracy of visual environment and eyeball dynamic capture in 3D vision imaging technology is solved, and high-precision and low-cost visual function detection is achieved.

CN120259371APending Publication Date: 2025-07-04ZHEJIANG UNIV OF TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510338252.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing 3D vision imaging technology is difficult to take into account the fidelity of the visual environment and the accuracy of eye dynamic capture in visual function detection, and the equipment is complex and cost-effective, and the user experience is poor.

Method used

Through multi-dimensional data fusion and intelligent algorithm analysis, a virtual scene model similar to the real scene is built, combined with deep learning and inertial sensor data, the eye position is corrected in real time, and adaptive filtering and visual attention mechanism are used to identify user visual behavior characteristics.

Benefits of technology

It realizes high-precision, low-cost and convenient eye dynamic capture, provides personalized vision detection, and promptly detect potential visual abnormalities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259371A_ABST
    Figure CN120259371A_ABST
Patent Text Reader

Abstract

The invention discloses a visual function detection method and system based on a 3D visual imaging technology, and the method comprises the steps: obtaining three-dimensional visual imaging data, carrying out the modeling and rendering of the illumination, texture and geometric structure in a visual environment, and constructing a virtual scene model which is highly similar to a real scene; acquiring eyeball movement characteristics and pupil change data of the user, analyzing visual characteristics of the user according to the data, and adaptively designing a personalized visual stimulation mode matched with the visual characteristics; and processing captured eyeball image data by using a deep learning algorithm, extracting features of the eyeball image through a convolutional neural network, and modeling the eyeball motion trajectory by using a long-short-term memory network, thereby realizing recognition and prediction of the eyeball motion mode. According to the method, the visual behaviors of the user are accurately captured and analyzed through multi-dimensional data fusion and intelligent algorithm analysis, and an effective means is provided for timely discovering potential visual function abnormalities.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of visual function detection, and in particular relates to a visual function detection method and system based on 3D visual imaging technology. Background Art

[0002] In the visual function detection system based on 3D visual imaging technology, how to achieve high-precision capture of the user's eye dynamics while ensuring the fidelity of the visual environment is a key technical problem. Due to the complexity of the physiological structure and visual characteristics of the human eye, it is difficult for traditional visual function detection methods to take into account both the realism of the visual environment and the accuracy of eye dynamic capture. Although 3D imaging technology can build a realistic visual environment, it is easily disturbed by factors such as ambient lighting and slight movements of the user's head when capturing eye dynamics, resulting in distorted or inaccurate measurement data. At the same time, high-precision eye dynamic capture usually requires the use of complex hardware equipment and algorithms, which not only increases the cost and complexity of the system, but also brings inconvenience to users. Therefore, how to achieve high-precision, low-cost, and convenient eye dynamic capture in a 3D visual environment is a technical problem that needs to be solved urgently. This requires the development of innovative eye dynamic capture methods based on 3D imaging technology, and comprehensive consideration of multiple factors such as the physiological structure of the human eye, environmental factors, and user experience, in order to achieve a visual function detection system that takes into account both the fidelity of the visual environment and the accuracy of eye dynamic capture, and provide users with high-quality, personalized vision detection services. Summary of the invention

[0003] In order to solve the problems existing in the prior art, the present invention provides a visual function detection method and system based on 3D visual imaging technology, which aims to accurately capture and analyze the user's visual behavior through multi-dimensional data fusion and intelligent algorithm analysis, and provide an effective means for timely detection of potential visual function abnormalities. It can be widely used in fields such as vision health management.

[0004] To achieve the above object, the present invention provides the following solutions:

[0005] A visual function detection method based on 3D visual imaging technology, the method comprising:

[0006] Acquire three-dimensional visual imaging data, and construct a virtual scene model that is highly similar to the real scene based on the three-dimensional visual imaging data;

[0007] Acquire the user's eye movement characteristics and pupil change data, analyze the user's visual characteristics based on the eye movement characteristics and pupil change data, and adaptively design a personalized visual stimulation mode that matches the visual characteristics;

[0008] The captured eye image data is processed using deep learning algorithms. The features of the eye image are extracted through a convolutional neural network, and a long short-term memory network is used to model the eye movement trajectory, thereby realizing the recognition and prediction of eye movement patterns;

[0009] The inertial sensor data and eye image data of the head-mounted device are obtained, and a visual-inertial fusion algorithm is used to fuse and process the inertial sensor data and eye image data, and the eye position offset caused by head movement is estimated in real time to compensate and correct the eye dynamic data;

[0010] According to the eye dynamic measurement data, an adaptive Kalman filtering algorithm is used for filtering and smoothing processing, an eye movement state space model is established, and a Kalman filter is used to recursively estimate the eye movement state;

[0011] The visual environment data is obtained, and a visual attention mechanism is used to extract and analyze the significant regions and edge contours therein, a visual attention map is constructed, the user's fixation points and attention distribution are predicted, and combined with the filtered eye movement trajectory and the attention map, the user's visual behavior is comprehensively analyzed to obtain eye dynamic features for personalized vision detection and identify potential visual function abnormalities.

[0012] Preferably, three-dimensional visual imaging data is obtained, and according to the three-dimensional visual imaging data, constructing a virtual scene model highly similar to the real scene includes:

[0013] A multi-view image sequence and depth map data containing the target scene are obtained, and the visual data of the real scene is collected through a visual sensor;

[0014] According to the obtained image sequence data, feature points in the image are extracted, and through the principles of feature matching and triangulation, the three-dimensional spatial coordinates of the feature points are calculated to obtain the sparse three-dimensional point cloud data of the scene;

[0015] According to the obtained depth map data, the depth information of the scene surface is obtained through a depth camera, and the depth information is registered with the corresponding color image to obtain the dense three-dimensional point cloud data of the scene;

[0016] According to the sparse point cloud data and the dense point cloud data, a point cloud fusion algorithm is used to fuse the two types of point cloud data to obtain a fine three-dimensional point cloud model;

[0017] According to the three-dimensional point cloud model, through a surface reconstruction algorithm, the point cloud data is converted into a triangular mesh surface model to obtain the geometric structure model of the scene;

[0018] According to the obtained image sequence data, the texture information in the image is extracted, and through a texture mapping algorithm, the texture information is mapped onto the geometric structure model to obtain a three-dimensional scene model with real textures;

[0019] According to the lighting conditions of the real scene, an image-based lighting estimation algorithm is used to estimate the lighting distribution of the scene, and the estimated lighting information is applied to the 3D scene model to obtain a virtual scene model with a realistic lighting effect.

[0020] Preferably, the eye movement characteristics and pupil change data of the user are obtained, and the visual characteristics of the user are analyzed based on the eye movement characteristics and pupil change data. The personalized visual stimulation pattern matching the visual characteristics is adaptively designed, including:

[0021] Obtain the user's eye movement and pupil change data, and use an eye tracker to collect the physiological parameters of the user's fixation point coordinates and pupil diameter in real time;

[0022] Preprocess the collected eye movement and pupil data, remove outliers and noise, and extract stable feature vectors;

[0023] According to the extracted eye movement and pupil characteristics, a clustering algorithm is used to classify the visual characteristics of the user to obtain different categories of user visual patterns;

[0024] For each visual pattern, a machine learning algorithm is used to train a matching personalized visual stimulation model to determine the optimal combination of stimulation parameters;

[0025] In practical applications, according to the user's instant eye movement and pupil data, the category of the visual pattern is judged, and the corresponding personalized visual stimulation model is dynamically called;

[0026] The generated personalized visual stimulation pattern includes parameters such as the brightness, contrast, color, and movement speed of the stimulation image, as well as the time series of the stimulation;

[0027] The generated personalized visual stimulation pattern is presented to the user in real time, and the eye movement and pupil feedback data of the user are continuously collected to achieve dynamic matching and optimization adjustment of the stimulation pattern and the user's visual characteristics.

[0028] Preferably, a deep learning algorithm is used to process the captured eye image data, the features of the eye image are extracted through a convolutional neural network, and a long short-term memory network is used to model the eye movement trajectory, so as to realize the recognition and prediction of the eye movement pattern, including:

[0029] Obtain a set of eye image data, preprocess the image data to obtain a standardized eye image data set;

[0030] Use a convolutional neural network to extract features from the standardized eye image data set. Through the combination of convolutional layers and pooling layers, the key features of the eye image are automatically learned and extracted to obtain a compact feature representation vector;

[0031] Input the extracted sequence of eye image feature vectors into a long short-term memory network. Through the gating mechanism and memory unit, perform temporal modeling on the eye movement trajectory to capture the temporal dependence and long-term patterns of eye movement;

[0032] At the output layer of the long short-term memory network, set a classifier or a regressor. According to the learned eye movement pattern features, perform pattern recognition or movement prediction on the new eye movement trajectory.

[0033] Preferably, obtain the inertial sensor data and eye image data of the head-mounted device, and use a visual-inertial fusion algorithm to fuse and process the inertial sensor data and eye image data, and estimate the eye position offset caused by head movement in real time. The compensation and correction of eye dynamic data include:

[0034] Obtain the inertial sensor data and eye image data of the head-mounted device;

[0035] According to the inertial sensor data, estimate the three-dimensional spatial position change caused by head movement;

[0036] According to the eye image data, obtain the two-dimensional position coordinates of the eye in the image;

[0037] Through the visual-inertial fusion algorithm, fuse the three-dimensional spatial position change caused by head movement with the eye two-dimensional coordinates to estimate the three-dimensional offset of the eye position caused by head movement;

[0038] According to the estimated eye position offset, perform compensation and correction on the eye dynamic data to eliminate the eye position offset caused by head movement.

[0039] Preferably, according to the eye dynamic measurement data, use an adaptive Kalman filtering algorithm for filtering and smoothing processing, establish an eye movement state space model, and use a Kalman filter to recursively estimate the eye movement state, including:

[0040] Obtain the eye dynamic measurement data. According to the eye dynamic measurement data, establish an eye movement state space model, which includes eye movement state variables and an observation equation;

[0041] According to the established eye movement state space model, use the Kalman filtering algorithm to recursively estimate the optimal estimated value of the eye movement state variables, and obtain the filtered eye movement state estimated value;

[0042] Obtain the measurement noise statistical characteristics of the eye dynamic measurement data. According to the obtained measurement noise statistical characteristics, adaptively adjust the filtering parameters of the Kalman filter, including the process noise covariance and the measurement noise covariance;

[0043] Substitute the adaptively adjusted filtering parameters into the Kalman filter to filter the dynamic eye measurement data and suppress the high-frequency noise and distortion components in the measurement data;

[0044] Smooth the estimated value of the eye movement state after Kalman filtering;

[0045] Output the smoothed estimated value of the eye movement state as the final dynamic eye capture signal.

[0046] Preferably, obtain visual environment data, use a visual attention mechanism to extract and analyze the significant regions and edge contours therein, construct a visual attention map, predict the user's fixation points and attention distribution, combine the filtered eye movement trajectory and the attention map, comprehensively analyze the user's visual behavior, and obtain dynamic eye characteristics for personalized vision detection to identify potential visual function abnormalities, including:

[0047] Obtain the original visual environment data, use a trained convolutional neural network model to extract image features, construct a feature map, and perform an upsampling operation on the feature map to obtain a saliency map with the same resolution as the original image;

[0048] According to the saliency map, perform edge detection using the Canny operator to obtain an edge feature map and determine the edge contour;

[0049] Adopt a region growing algorithm, combine the saliency map and the edge feature map to segment different significant regions;

[0050] Construct a visual attention map through the significant regions and edge contours, and sort according to the saliency values to generate a saliency list;

[0051] Combine a pre-established deep learning model and the saliency list to predict the probability map of the user's fixation points and attention distribution;

[0052] Obtain the original eye movement data and perform Kalman filtering on the original eye movement data to obtain smooth eye movement trajectory data;

[0053] Calculate eye movement metrics based on the smooth eye movement trajectory data to obtain the duration of fixation points;

[0054] Determine the saccade speed according to the distance and time difference between adjacent fixation points;

[0055] Construct a heat map of fixation point distribution through the coordinates and time information of the fixation points in space;

[0056] Combine the attention map and the heat map of fixation point distribution to judge the user's visual behavior pattern;

[0057] If there are significant differences between the user's visual behavior pattern and the statistical model of the normal population, it is marked as potential visual function abnormality;

[0058] According to the type and degree of the abnormality, a personalized visual acuity detection report is generated.

[0059] The present invention also provides a visual function detection system based on 3D visual imaging technology. The system is used to implement any one of the methods described above. The system includes: a virtual scene construction module, a personalized visual stimulation module, an eye movement pattern recognition module, an eye movement dynamic compensation and correction module, an eye movement dynamic filtering and smoothing module, and a visual behavior analysis module;

[0060] The virtual scene construction module is used to obtain three-dimensional visual imaging data and construct a virtual scene model highly similar to the real scene according to the three-dimensional visual imaging data;

[0061] The personalized visual stimulation module is used to obtain the eye movement characteristics and pupil change data of the user, analyze the visual characteristics of the user according to the eye movement characteristics and pupil change data, and adaptively design a personalized visual stimulation pattern matching the visual characteristics;

[0062] The eye movement pattern recognition module is used to process the captured eye image data by using a deep learning algorithm, extract the features of the eye image through a convolutional neural network, and use a long short-term memory network to model the eye movement trajectory, so as to realize the recognition and prediction of the eye movement pattern;

[0063] The eye movement dynamic compensation and correction module is used to obtain the inertial sensor data and eye image data of the head-mounted device, perform fusion processing on the inertial sensor data and eye image data by using a visual inertial fusion algorithm, estimate the eye position offset caused by head movement in real time, and perform compensation and correction on the eye movement dynamic data;

[0064] The eye movement dynamic filtering and smoothing module is used to perform filtering and smoothing processing on the eye movement dynamic measurement data by using an adaptive Kalman filtering algorithm, establish an eye movement state space model, and use a Kalman filter to recursively estimate the eye movement state;

[0065] The visual behavior analysis module is used to obtain visual environment data, extract and analyze the significant regions and edge contours therein by using a visual attention mechanism, construct a visual attention map, predict the user's fixation points and attention distribution, and comprehensively analyze the user's visual behavior in combination with the filtered eye movement trajectory and the attention map to obtain eye movement dynamic characteristics for personalized visual acuity detection and identify potential visual function abnormalities.

[0066] Compared with the prior art, the beneficial effects of the present invention are:

[0067] The present invention discloses a visual function detection method and system based on 3D vision imaging technology. The method first acquires three-dimensional vision imaging data and constructs a virtual scene model, while capturing the user's eye movement and pupil change data. The eye images are processed by a deep learning algorithm to extract features and model the eye movement trajectory. Combining the inertial sensor data of the head-mounted device, a visual-inertial fusion algorithm is used to correct the eye position in real time. Adaptive Kalman filtering is performed on the eye dynamic measurement data to suppress noise and improve the signal quality. In addition, the present invention also analyzes the significant regions of the visual environment and constructs an attention map to predict the user's fixation point. Finally, by synthesizing the eye movement trajectory and attention distribution, the dynamic eye features are extracted for personalized vision detection. Through multi-dimensional data fusion and intelligent algorithm analysis, the method realizes the accurate capture and analysis of the user's visual behavior, provides an effective means for timely discovery of potential visual function abnormalities, and can be widely applied to fields such as vision health management. BRIEF DESCRIPTION OF THE DRAWINGS

[0068] In order to more clearly illustrate the technical solutions of the present invention, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.

[0069] Figure 1 Schematic flow diagram of a visual function detection method based on 3D vision imaging technology according to an embodiment of the present invention;

[0070] Figure 2 Schematic diagram of the prediction model established according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0071] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some of the embodiments of the present invention, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0072] In order to make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the drawings and specific embodiments.

[0073] Embodiment 1

[0074] As Figure 1 shown, the present invention provides a visual function detection method based on 3D vision imaging technology, and the method includes:

[0075] Obtain three-dimensional visual imaging data, and construct a virtual scene model highly similar to the real scene according to the three-dimensional visual imaging data;

[0076] Obtain the eye movement characteristics and pupil change data of the user, analyze the visual characteristics of the user according to the eye movement characteristics and pupil change data, and adaptively design a personalized visual stimulation pattern matching the visual characteristics;

[0077] Use a deep learning algorithm to process the captured eye image data, extract the features of the eye image through a convolutional neural network, and use a long short-term memory network to model the eye movement trajectory, so as to realize the recognition and prediction of the eye movement pattern;

[0078] Obtain the inertial sensor data and eye image data of the head-mounted device, use a visual-inertial fusion algorithm to fuse and process the inertial sensor data and eye image data, estimate the eye position offset caused by head movement in real time, and compensate and correct the eye dynamic data;

[0079] According to the eye dynamic measurement data, use an adaptive Kalman filter algorithm for filtering and smoothing processing, establish an eye movement state space model, and use a Kalman filter to recursively estimate the eye movement state;

[0080] Obtain visual environment data, use a visual attention mechanism to extract and analyze the significant regions and edge contours therein, construct a visual attention map, predict the user's fixation points and attention distribution, combine the filtered eye movement trajectory and the attention map, comprehensively analyze the user's visual behavior, obtain eye dynamic characteristics, and be used for personalized vision detection and identify potential visual function abnormalities.

[0081] In this embodiment, obtaining three-dimensional visual imaging data, modeling and rendering the light, texture and geometric structure in the visual environment, and constructing a virtual scene model highly similar to the real scene includes:

[0082] Obtain a multi-view image sequence and depth map data containing the target scene, and collect the visual data of the real scene through a visual sensor. According to the obtained image sequence data, extract the feature points in the images, and calculate the three-dimensional spatial coordinates of the feature points through feature matching and triangulation principles to obtain the sparse three-dimensional point cloud data of the scene. According to the obtained depth map data, obtain the depth information of the scene surface through a depth camera, register the depth information with the corresponding color image to obtain the dense three-dimensional point cloud data of the scene. According to the sparse point cloud data and the dense point cloud data, adopt a point cloud fusion algorithm, such as moving least squares method, to fuse the two kinds of point cloud data to obtain a fine three-dimensional point cloud model. According to the three-dimensional point cloud model, through a surface reconstruction algorithm, such as Poisson reconstruction algorithm, convert the point cloud data into a triangular mesh surface model to obtain the geometric structure model of the scene. According to the obtained image sequence data, extract the texture information in the images, and map the texture information to the geometric structure model through a texture mapping algorithm to obtain a three-dimensional scene model with real textures. According to the lighting conditions of the real scene, adopt an image-based lighting estimation algorithm, such as spherical harmonic lighting model, to estimate the lighting distribution of the scene, and apply the estimated lighting information to the three-dimensional scene model to obtain a virtual scene model with realistic lighting effects.

[0083] In this embodiment, obtain the user's eye movement characteristics and pupil change data, and analyze the user's visual characteristics according to the data. The personalized visual stimulation mode adaptively designed to match the visual characteristics includes:

[0084] Obtain the user's eye movement and pupil change data, and use devices such as an eye tracker to collect physiological parameters such as the user's fixation point coordinates and pupil diameter in real time. Preprocess the collected eye movement and pupil data, remove outliers and noise, and extract stable feature vectors. According to the extracted eye movement and pupil characteristics, use a clustering algorithm to classify the user's visual characteristics to obtain different categories of user visual patterns. For each visual pattern, use a machine learning algorithm to train a personalized visual stimulation model that matches it, and determine the best combination of stimulation parameters. In practical applications, according to the user's instant eye movement and pupil data, judge the category of the visual pattern to which it belongs, and dynamically call the corresponding personalized visual stimulation model. The generated personalized visual stimulation mode includes parameters such as the brightness, contrast, color, and movement speed of the stimulation image, as well as the time series of the stimulation. Present the generated personalized visual stimulation mode to the user in real time, and continuously collect the user's eye movement and pupil feedback data to achieve dynamic matching and optimization adjustment of the stimulation mode and the user's visual characteristics.

[0085] Specifically, an eye tracker is a device that can track eye movements and pupil changes, and can collect physiological parameters such as the coordinates of the user's fixation points and pupil diameter in real time. For example, an eye tracker can record the specific position coordinates on the screen where the user's eyes fixate when viewing the screen, as well as the changes in pupil size. These data can reflect information such as the user's visual attention and cognitive load. Through the eye tracker, the fixation point trajectory and pupil diameter sequence of the user within a specific time period can be obtained. For example, when a user watches a one-minute video, the eye tracker can record the fixation point coordinates of the user's eyes on the screen at each moment within this minute, as well as the values of the pupil diameter.

[0086] In this embodiment, a deep learning algorithm is used to process the captured eye image data, extract the features of the eye image through a convolutional neural network, and use a long short-term memory network to model the eye movement trajectory, so as to realize the recognition and prediction of eye movement patterns, including:

[0087] Obtain a set of eye image data, preprocess the image data, including operations such as image denoising and normalization, to obtain a standardized eye image data set. Use a convolutional neural network to extract features from the standardized eye image data set. Through the combination of convolutional layers and pooling layers, automatically learn and extract the key features of the eye image to obtain a compact feature representation vector. Input the extracted sequence of eye image feature vectors into a long short-term memory network. Through the gating mechanism and memory unit, perform temporal modeling on the eye movement trajectory to capture the time-dependent relationship and long-term pattern of eye movement. At the output layer of the long short-term memory network, set a classifier or a regressor. According to the learned eye movement pattern features, perform pattern recognition or motion prediction on the new eye movement trajectory. Construct a training data set, including eye image data and corresponding motion pattern labels, and use the supervised learning method to optimize the parameters of the convolutional neural network and the long short-term memory network to minimize the prediction error. In the test stage, input the newly collected eye image data into the trained deep learning model, extract image features through the convolutional neural network, and use the long short-term memory network for motion pattern recognition and prediction. According to the recognition and prediction results, the user's fixation behavior, interest preferences, etc. can be further analyzed to provide decision-making support for related applications, such as in the fields of gaze tracking and human-computer interaction.

[0088] Specifically, the long short-term memory network LSTM is a double long short-term memory network Double LSTM that introduces an attention mechanism Attention.

[0089] The established model includes: a CNN layer, an LSTM-ATT-LSTM layer, and an output layer; as Figure 2 shown.

[0090] The CNN layer is used to extract features of the training set based on a one-dimensional convolutional layer, and perform downsampling on the extracted features using max pooling to obtain key features; the calculation formula is as follows:

[0091] X conv [i, j, k] = ReLU(X input [i, j:j + k, :] * W conv [k, :, :] + b conv ),

[0092] X pool [i, j] = max(X conv [i, 2j, :], X conv [i, 2j + 1, :]),

[0093] In the formula, X input is the data sequence of the training set input to the model, N is the number of samples, T is the sequence length, and F is the number of features of each sequence; W conv is the convolution kernel, with a shape of k×F×F conv , F conv is the number of convolution kernels, b conv is the bias term; ReLU(z) = max(0, z) is the ReLU activation function, X conv [i, j, k] is the output of the convolutional layer, and X pool [i, j] is the key feature output after pooling;

[0094] The LSTM-ATT-LSTM layer includes a first LSTM layer, an attention mechanism layer, and a second LSTM layer, and is used to process the key features to obtain feature vectors; among them, the flow of information is controlled by the internal gate structure of the double-layer LSTM, and the attention mechanism Attention dynamically assigns weights to the key features; among them, the gate structure includes a forgetting gate, an input gate, and an output gate, and the gate structure controls the flow of information through the sigmoid function and the tanh function;

[0095] The output layer is used to obtain a prediction result based on the feature vector, and use the categorical cross-entropy loss function to obtain the difference between the prediction result and the true value.

[0096] In this embodiment, obtaining inertial sensor data and eye image data of a head-mounted device, and using a visual-inertial fusion algorithm to perform fusion processing on the data, and real-time estimating the offset of the eye position caused by head movement, and compensating and correcting the eye movement data to eliminate the influence of head movement on the accuracy of eye movement capture includes:

[0097] Obtain the inertial sensor data and eye image data of the head-mounted device; estimate the three-dimensional spatial position change caused by head movement according to the inertial sensor data; obtain the two-dimensional position coordinates of the eyes in the image according to the eye image data; through a visual-inertial fusion algorithm, fuse the three-dimensional spatial position change caused by head movement with the two-dimensional eye coordinates to estimate the three-dimensional offset of the eye position caused by head movement; compensate and correct the eye dynamic data according to the estimated eye position offset to eliminate the eye position offset caused by head movement; smooth the compensated and corrected eye dynamic data to reduce data noise; output the processed eye dynamic data to improve the accuracy of eye dynamic capture and eliminate the influence of head movement on the accuracy of eye dynamic capture.

[0098] Specifically, it is necessary to synchronously obtain the inertial sensor data of the head-mounted device and the eye image data captured by the eye-tracking camera. Inertial sensors usually include accelerometers and gyroscopes, which can measure the linear acceleration and angular velocity of the device in three-dimensional space. For example, a user wears a head-mounted device for a virtual reality experience. When the user's head turns to the left, the gyroscope will record the corresponding angular velocity data, and at the same time, the accelerometer will also record the acceleration generated by the head movement. At the same time, the eye-tracking camera will capture a series of eye images. For example, when the user is looking at the center of the screen, the center of the pupil in the eye image is near the center of the image coordinate system.

[0099] Fusing the three-dimensional spatial position change caused by head movement with the two-dimensional eye coordinates through a visual-inertial fusion algorithm to estimate the three-dimensional offset of the eye position caused by head movement includes: obtaining the three-dimensional spatial position change data and the two-dimensional eye coordinate data caused by head movement, and inputting the two types of data into the visual-inertial fusion algorithm for processing. The visual-inertial fusion algorithm fuses the three-dimensional spatial position change data of head movement with the two-dimensional eye coordinate data through a Kalman filter to establish a mapping relationship model between head movement and eye position. According to the established mapping relationship model between head movement and eye position, obtain the three-dimensional spatial position change data of the head in real time, input it into the mapping model, and estimate the three-dimensional spatial position of the eyes under the current head movement. Compare the position of the eyes in three-dimensional space with the position of the eyes in the static state, and calculate the three-dimensional position offset of the eyes caused by head movement. Determine whether the three-dimensional position offset of the eyes exceeds the preset threshold range. If it exceeds, trigger a warning message to prompt the user that the head movement amplitude is too large, which may affect the visual experience. According to the magnitude of the three-dimensional position offset of the eyes, adaptively adjust the content of the virtual reality scene, perform position offset compensation on the screen, and reduce the influence of head movement on the visual experience. Continuously track the change of the three-dimensional position offset of the eyes caused by head movement, and adjust the virtual reality scene screen in real time to ensure that during the head movement, a stable and clear immersive experience is always provided to the user.

[0100] In this embodiment, for the dynamic measurement data of the eyeball, an adaptive Kalman filtering algorithm is used for filtering and smoothing, an eyeball motion state space model is established, the Kalman filter is used to recursively estimate the eyeball motion state, and the filter parameters are adaptively adjusted according to the statistical characteristics of the measurement noise to suppress the noise and distortion in the measurement data and improve the quality of the dynamically captured eyeball signal, including:

[0101] Obtain the dynamic measurement data of the eyeball. For this measurement data, establish an eyeball motion state space model, which includes eyeball motion state variables and an observation equation. According to the established eyeball motion state space model, use the Kalman filtering algorithm to recursively estimate the optimal estimated value of the eyeball motion state variables, and obtain the filtered estimated value of the eyeball motion state. Obtain the statistical characteristics of the measurement noise of the dynamic measurement data of the eyeball. According to the obtained statistical characteristics of the measurement noise, adaptively adjust the filtering parameters of the Kalman filter, including the process noise covariance and the measurement noise covariance. Substitute the adaptively adjusted filtering parameters into the Kalman filter to filter the dynamic measurement data of the eyeball and suppress the high-frequency noise and distortion components in the measurement data. Smooth the estimated value of the eyeball motion state after Kalman filtering, and use smoothing algorithms such as moving average or polynomial fitting to further reduce the fluctuations and jumps of the estimated value. Output the smoothed estimated value of the eyeball motion state as the final dynamically captured eyeball signal. Compared with the original measurement data, this output signal has a higher signal-to-noise ratio and smoothness. Evaluate the quality of the dynamically captured eyeball signal before and after adaptive Kalman filtering, and quantify and evaluate the noise suppression and signal enhancement effects of the filtering algorithm by calculating indicators such as the mean square error and signal-to-noise ratio of the signal.

[0102] In this embodiment, obtain visual environment data, use a visual attention mechanism to extract and analyze the significant regions and edge contours therein, construct a visual attention map, and predict the user's fixation points and attention distribution. Combine the filtered eyeball movement trajectory and the attention map to comprehensively analyze the user's visual behavior, obtain eyeball dynamic features, including fixation duration, saccade speed, and fixation point distribution, for personalized vision detection and identification of potential visual function abnormalities.

[0103] Obtain the original visual environment data, use the trained convolutional neural network model to extract image features, construct a feature map, and perform an upsampling operation on the feature map to obtain a saliency map with the same resolution as the original image. According to the saliency map, use the Canny operator for edge detection to obtain an edge feature map and determine the edge contour. Adopt the region growing algorithm, combine the saliency map and the edge feature map to segment different salient regions. Through the salient regions and the edge contour, construct a visual attention map, and sort according to the saliency values to generate a saliency list. Combine the pre-established deep learning model and the saliency list to predict the user's fixation points and the probability map of attention distribution. Obtain the original eye movement data, and perform Kalman filtering on the original eye movement data to remove noise interference and obtain smooth eye movement trajectory data. According to the smooth eye movement trajectory data, calculate eye movement metrics to obtain the fixation duration. Determine the saccade speed according to the distance and time difference between adjacent fixation points. Construct a fixation point distribution heat map through the coordinates and time information of the fixation points in space. Combine the attention map and the fixation point distribution heat map to judge the user's visual behavior pattern. If the user's visual behavior pattern has a significant difference from the statistical model of the normal population, it is marked as a potential visual function abnormality. Generate a personalized vision detection report according to the type and degree of the abnormality.

[0104] Embodiment 2

[0105] The present invention also provides a visual function detection system based on 3D visual imaging technology. The system is used to implement any one of the above methods. The system includes: a virtual scene construction module, a personalized visual stimulation module, an eye movement pattern recognition module, an eye movement dynamic compensation and correction module, an eye movement dynamic filtering and smoothing module, and a visual behavior analysis module;

[0106] The virtual scene construction module is used to obtain three-dimensional visual imaging data and construct a virtual scene model highly similar to the real scene according to the three-dimensional visual imaging data;

[0107] The personalized visual stimulation module is used to obtain the user's eye movement characteristics and pupil change data, analyze the user's visual characteristics according to the eye movement characteristics and pupil change data, and adaptively design a personalized visual stimulation pattern matching the visual characteristics;

[0108] The eye movement pattern recognition module is used to process the captured eye image data by using a deep learning algorithm, extract the features of the eye image through a convolutional neural network, and use a long short-term memory network to model the eye movement trajectory, so as to realize the recognition and prediction of the eye movement pattern;

[0109] The eye movement dynamic compensation and correction module is used to obtain the inertial sensor data and eye image data of the head-mounted device, and perform fusion processing on the inertial sensor data and eye image data by using a visual-inertial fusion algorithm to estimate in real time the offset of the eye position caused by head movement, and perform compensation and correction on the eye movement dynamic data;

[0110] The eye movement dynamic filtering and smoothing module is used to perform filtering and smoothing processing according to the eye movement dynamic measurement data by using an adaptive Kalman filtering algorithm, establish an eye movement state space model, and recursively estimate the eye movement state by using a Kalman filter;

[0111] The visual behavior analysis module is used to obtain visual environment data, extract and analyze the significant regions and edge contours therein by using a visual attention mechanism, construct a visual attention map, predict the user's fixation points and attention distribution, and comprehensively analyze the user's visual behavior in combination with the filtered eye movement trajectory and the attention map to obtain eye movement dynamic features for personalized vision detection and identify potential visual function abnormalities.

[0112] The above embodiments are only descriptions of the preferred embodiments of the present invention, and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope determined by the claims of the present invention.

Claims

1. A visual function detection method based on 3D vision imaging technology, characterized in that, The method includes: Obtaining three-dimensional visual imaging data, and constructing a virtual scene model highly similar to the real scene according to the three-dimensional visual imaging data; Obtaining the eye movement characteristics and pupil change data of the user, analyzing the visual characteristics of the user according to the eye movement characteristics and pupil change data, and adaptively designing a personalized visual stimulation pattern matching the visual characteristics; Processing the captured eye image data by using a deep learning algorithm, extracting the features of the eye image through a convolutional neural network, and modeling the eye movement trajectory by using a long short-term memory network, so as to realize the recognition and prediction of the eye movement pattern; Obtaining the inertial sensor data and eye image data of the head-mounted device, and performing fusion processing on the inertial sensor data and eye image data by using a visual-inertial fusion algorithm, and estimating the eye position offset caused by head movement in real time to compensate and correct the eye dynamic data; According to the eye dynamic measurement data, performing filtering and smoothing processing by using an adaptive Kalman filtering algorithm, establishing an eye movement state space model, and recursively estimating the eye movement state by using a Kalman filter; Obtaining visual environment data, extracting and analyzing the significant regions and edge contours therein by using a visual attention mechanism, constructing a visual attention map, predicting the user's fixation points and attention distribution, combining the filtered eye movement trajectory and the attention map, comprehensively analyzing the user's visual behavior, obtaining eye dynamic characteristics for personalized vision detection, and identifying potential visual function abnormalities.

2. The method according to claim 1, wherein Obtaining three-dimensional visual imaging data, and constructing a virtual scene model highly similar to the real scene according to the three-dimensional visual imaging data includes: Obtaining a multi-view image sequence and depth map data including the target scene, and collecting the visual data of the real scene through a visual sensor; According to the obtained image sequence data, extracting the feature points in the image, and calculating the three-dimensional spatial coordinates of the feature points through the principles of feature matching and triangulation to obtain the sparse three-dimensional point cloud data of the scene; According to the obtained depth map data, obtaining the depth information of the scene surface through a depth camera, registering the depth information with the corresponding color image to obtain the dense three-dimensional point cloud data of the scene; According to the sparse point cloud data and the dense point cloud data, adopting a point cloud fusion algorithm to fuse the two kinds of point cloud data to obtain a fine three-dimensional point cloud model; According to the three-dimensional point cloud model, converting the point cloud data into a triangular mesh surface model through a surface reconstruction algorithm to obtain the geometric structure model of the scene; According to the obtained image sequence data, extracting the texture information in the image, and mapping the texture information to the geometric structure model through a texture mapping algorithm to obtain a three-dimensional scene model with real texture; According to the illumination condition of the real scene, adopting an image-based illumination estimation algorithm to estimate the illumination distribution of the scene, and applying the estimated illumination information to the three-dimensional scene model to obtain a virtual scene model with a realistic illumination effect.

3. The method according to claim 1, characterized in that Obtaining the eye movement characteristics and pupil change data of the user, analyzing the visual characteristics of the user according to the eye movement characteristics and pupil change data, and adaptively designing a personalized visual stimulation pattern matching the visual characteristics includes: Obtain user eye movement and pupil change data, and use an eye tracker to collect the physiological parameters of the user's fixation point coordinates and pupil diameter in real time; Preprocess the collected eye movement and pupil data, remove outliers and noise, and extract stable feature vectors; According to the extracted eye movement and pupil features, use a clustering algorithm to classify the user's visual characteristics and obtain different categories of user visual patterns; For each visual pattern, use a machine learning algorithm to train a matching personalized visual stimulation model and determine the optimal combination of stimulation parameters; In practical applications, based on the user's instant eye movement and pupil data, judge the category of the visual pattern to which it belongs, and dynamically call the corresponding personalized visual stimulation model; The generated personalized visual stimulation pattern includes parameters such as the brightness, contrast, color, and movement speed of the stimulation image, as well as the time series of the stimulation; Present the generated personalized visual stimulation pattern to the user in real time, and continuously collect the user's eye movement and pupil feedback data to achieve dynamic matching and optimization adjustment of the stimulation pattern and the user's visual characteristics.

4. The method according to claim 1, wherein Use a deep learning algorithm to process the captured eye image data, extract the features of the eye image through a convolutional neural network, and use a long short-term memory network to model the eye movement trajectory, so as to realize the recognition and prediction of the eye movement pattern, including: Obtain a set of eye image data, preprocess the image data, and obtain a standardized eye image data set; Use a convolutional neural network to extract features from the standardized eye image data set. Through the combination of convolutional layers and pooling layers, automatically learn and extract the key features of the eye image, and obtain a compact feature representation vector; Input the extracted sequence of eye image feature vectors into a long short-term memory network. Through the gating mechanism and memory unit, perform temporal modeling on the eye movement trajectory to capture the time-dependent relationship and long-term pattern of eye movement; At the output layer of the long short-term memory network, set a classifier or a regressor. According to the learned eye movement pattern features, perform pattern recognition or motion prediction on the new eye movement trajectory.

5. The method according to claim 1, wherein Obtain the inertial sensor data and eye image data of the head-mounted device, and use a visual-inertial fusion algorithm to fuse the inertial sensor data and eye image data to estimate the eye position offset caused by head movement in real time, and compensate and correct the eye dynamic data, including: Obtain the inertial sensor data and eye image data of the head-mounted device; Estimate the three-dimensional spatial position change caused by head movement according to the inertial sensor data; Obtain the two-dimensional position coordinates of the eye in the image according to the eye image data; Through the visual-inertial fusion algorithm, fuse the three-dimensional spatial position change caused by head movement with the eye two-dimensional coordinates to estimate the three-dimensional offset of the eye position caused by head movement; According to the estimated eye position offset, compensate and correct the eye dynamic data to eliminate the eye position offset caused by head movement.

6. The method according to claim 1, wherein According to the eye dynamic measurement data, use an adaptive Kalman filter algorithm for filtering and smoothing processing, establish an eye movement state space model, and use the Kalman filter to recursively estimate the eye movement state, including: Obtain dynamic eye measurement data, and establish an eye movement state space model based on the dynamic eye measurement data. This model includes eye movement state variables and an observation equation; According to the established eye movement state space model, use the Kalman filter algorithm to recursively estimate the optimal estimated value of the eye movement state variables, and obtain the filtered eye movement state estimated value; Obtain the measurement noise statistical characteristics of the dynamic eye measurement data, and adaptively adjust the filtering parameters of the Kalman filter according to the obtained measurement noise statistical characteristics, including the process noise covariance and the measurement noise covariance; Substitute the adaptively adjusted filtering parameters into the Kalman filter to perform filtering processing on the dynamic eye measurement data, and suppress the high-frequency noise and distortion components in the measurement data; Perform smoothing processing on the eye movement state estimated value after Kalman filtering; Output the smoothed eye movement state estimated value as the final eye movement dynamic capture signal.

7. The method according to claim 1, wherein Obtain visual environment data, use a visual attention mechanism to extract and analyze the significant regions and edge contours therein, construct a visual attention map, predict the user's fixation points and attention distribution, and combine the filtered eye movement trajectory and the attention map to comprehensively analyze the user's visual behavior, obtain eye movement dynamic characteristics for personalized vision detection, and identify potential visual function abnormalities including: Obtain the original visual environment data, use a trained convolutional neural network model to extract image features, construct a feature map, and perform an upsampling operation on the feature map to obtain a saliency map with the same resolution as the original image; According to the saliency map, perform edge detection using the Canny operator to obtain an edge feature map and determine the edge contour; Adopt a region growing algorithm, combine the saliency map and the edge feature map to segment different significant regions; Construct a visual attention map through the significant regions and edge contours, and sort according to the saliency values to generate a saliency list; Combine a pre-established deep learning model and the saliency list to predict the probability map of the user's fixation points and attention distribution; Obtain the original eye movement data, and perform Kalman filtering on the original eye movement data to obtain smooth eye movement trajectory data; Calculate eye movement metrics according to the smooth eye movement trajectory data to obtain the duration of the fixation points; Determine the saccade speed according to the distance and time difference between adjacent fixation points; Construct a fixation point distribution heat map through the coordinates and time information of the fixation points in space; Combine the attention map and the fixation point distribution heat map to judge the user's visual behavior pattern; If the user's visual behavior pattern is significantly different from the statistical model of the normal population, mark it as a potential visual function abnormality; Generate a personalized vision detection report according to the type and degree of the abnormality.

8. A visual function detection system based on 3D vision imaging technology, the system is used to implement the method described in any one of claims 1-7, characterized in that, The system includes: a virtual scene construction module, a personalized visual stimulation module, an eye movement pattern recognition module, an eye movement dynamic compensation and correction module, an eye movement dynamic filtering and smoothing module, and a visual behavior analysis module; The virtual scene construction module is used to obtain three-dimensional visual imaging data and construct a virtual scene model highly similar to the real scene according to the three-dimensional visual imaging data; The personalized visual stimulation module is used to obtain the eye movement characteristics and pupil change data of the user, analyze the visual characteristics of the user according to the eye movement characteristics and pupil change data, and adaptively design a personalized visual stimulation pattern matching the visual characteristics; The eye movement pattern recognition module is used to process the captured eye image data by using a deep learning algorithm, extract the features of the eye image through a convolutional neural network, and use a long short-term memory network to model the eye movement trajectory, so as to realize the recognition and prediction of the eye movement pattern; The eye movement dynamic compensation and correction module is used to obtain the inertial sensor data and eye image data of the head-mounted device, perform fusion processing on the inertial sensor data and eye image data by using a visual inertial fusion algorithm, estimate the eye position offset caused by head movement in real time, and compensate and correct the eye movement dynamic data; The eye movement dynamic filtering and smoothing module is used to perform filtering and smoothing processing according to the eye movement dynamic measurement data by using an adaptive Kalman filtering algorithm, establish an eye movement state space model, and recursively estimate the eye movement state by using a Kalman filter; The visual behavior analysis module is used to obtain visual environment data, extract and analyze the significant regions and edge contours therein by using a visual attention mechanism, construct a visual attention map, predict the user's fixation points and attention distribution, combine the filtered eye movement trajectory and the attention map, comprehensively analyze the user's visual behavior, obtain eye movement dynamic characteristics, and be used for personalized vision detection and identify potential visual function abnormalities.

Citation Information

Cited By

  • Asymmetric scattering myopia control lens based on visual behaviors of human eyes

    CN121325435A