Vestibular function detection system, data processing method, electronic equipment and program product
By combining eye-tracking video and body position data into a neural network model for identification, the problems of accuracy and ease of use in vestibular function detection in existing technologies have been solved, achieving a more accurate and easy-to-understand vestibular function assessment.
Patent Information
- Application Number
- CN202511307418.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-12-19
AI Technical Summary
Existing methods for testing vestibular function rely on eye-tracking videos or trajectory analysis, which cannot accurately reflect the patient's actual vestibular function status. Furthermore, due to the influence of multiple factors such as changes in body position, a single data source cannot provide comprehensive diagnostic evidence.
By combining eye-tracking videos and auxiliary detection data (such as body position data and nystagmus feature parameters), a pre-trained neural network model is used to perform data fusion and recognition, identify nystagmus information under specific body positions, and output vestibular function data.
It improves the accuracy and ease of vestibular function assessment, and the output vestibular function data is more accurate and easier to understand, making it suitable for rapid diagnosis by non-professional medical staff.
Smart Images

Figure CN121154089A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of medical technology, and in particular to a vestibular function testing system and data processing method, electronic equipment and program products. Background Technology
[0002] Vestibular dysfunction is a common manifestation of balance disorder in clinical practice, and its detection and assessment are of great significance for patient diagnosis and treatment. Eye-tracking video recording and analysis of a patient's eye movement trajectory can assess vestibular function. For example, nystagmus imaging is currently used clinically. However, the information provided by eye-tracking videos or trajectories has significant limitations and cannot directly reflect the core conclusions of the examination. Furthermore, because eye-tracking trajectories contain a large amount of complex information, much of which is additional information unrelated to diagnosis, it is difficult to accurately reflect the patient's actual vestibular functional status through manual observation or simple analysis alone. In addition, vestibular function testing is often affected by multiple factors such as changes in body position, and analysis of a single data source may not provide comprehensive diagnostic evidence. Summary of the Invention
[0003] To address the aforementioned technical problems, this invention provides a vestibular function detection system and data processing method, electronic device, and computer program product.
[0004] In a first aspect, this application discloses a method for processing vestibular function test data, comprising: acquiring eye movement videos of a subject during vestibular function testing; extracting the subject's eye movement trajectory from the eye movement videos; the eye movement trajectory including at least one of the following: horizontal eye movement trajectory, vertical eye movement trajectory, and torsional eye movement trajectory; acquiring auxiliary test data of the subject; the auxiliary test data including the subject's body position data during the vestibular function test; inputting the eye movement trajectory and the auxiliary test data into a pre-trained neural network model, wherein the neural network model is configured to: based on the time synchronization association between the eye movement trajectory and the body position data, identify the changes in the subject's eye movement in each specific body position, and output the subject's vestibular function data.
[0005] In some implementations, the changes in eye movement of the subject in each specific body position include: the nystagmus that occurs when the subject completes the change in body position and remains in the specific body position.
[0006] In some embodiments, the positional data is time-series data of the subject's head position, including at least one or both of the pitch and yaw axis trajectory sequences.
[0007] In some embodiments, the vestibular function test data processing method further includes: performing position recognition on the position data to determine the specific positions of the subject in the vestibular function test and the corresponding start and end times.
[0008] In some embodiments, the vestibular function detection data processing method further includes: marking the corresponding time period of the eye movement trajectory according to the identified specific body position and its start and end time, so as to increase the attention weight of the neural network model to the marked time period.
[0009] In some implementations, the neural network model employs a multi-input channel structure, where different input channels receive eye-tracking trajectories in different directions and body position data along different axes.
[0010] In some implementations, the vestibular function data output by the neural network model includes structured results for "specific body position - eye movement changes", which include at least the specific body position, the presence or absence of nystagmus, and the direction of nystagmus.
[0011] In some implementations, the auxiliary detection data input to the neural network model further includes nystagmus feature parameters; the acquisition of the nystagmus feature parameters includes: identifying nystagmus information in eye movement trajectories in each direction and obtaining the corresponding nystagmus feature parameters; the nystagmus feature parameters include any one or more of the following: nystagmus slow phase angular velocity, nystagmus direction, nystagmus time, and nystagmus change trend.
[0012] In some embodiments, the vestibular function data output by the neural network model is the nystagmus information of the subject in various specific body positions during vestibular function testing; the vestibular function testing data processing method further includes: performing data optimization processing on the identified nystagmus information in various specific body positions, and outputting processed structured vestibular function text data.
[0013] In some implementations, the data optimization process includes any one or more of the following: data representation simplification, invalid information filtering, feature visualization, and effective information merging; wherein: the data representation simplification includes: mapping the identified nystagmus latency and duration from numerical ranges to staged text descriptions; the invalid information filtering includes at least one of the following: uniformly classifying nystagmus events with nystagmus intensity below a set threshold as non-significant nystagmus; uniformly describing nystagmus maintained throughout the detection process as spontaneous nystagmus; and retaining only nystagmus events with intensity reaching a set threshold in the nystagmus identification results based on eye movement data in the horizontal, vertical, and torsional directions; the feature visualization process includes: for nystagmus identified from horizontal eye movement trajectories, according to the subject's specific body position and the relationship between the nystagmus direction and the ground... The system generates corresponding descriptions of geotropic or geotropic nystagmus; the effective information merging includes at least one of the following: based on body position information, performing merging processing on nystagmus data of the same diagnostic element detected under different body positions, and outputting merged nystagmus feature data; performing merging processing on nystagmus data of multiple diagnostic elements corresponding to the same diagnosis, and generating merged diagnostic information; based on nystagmus intensity parameters, sorting processing on nystagmus data corresponding to different diagnoses, and outputting diagnostic data sorted by intensity; based on timestamp information, performing alignment and deduplication processing on nystagmus data collected at different time points under the same body position, and generating merged body position-nystagmus data; based on the time series of the diagnosis and treatment process, associating and comparing nystagmus feature data collected before, during and after treatment, and generating a merged result including differences in treatment effects.
[0014] Secondly, this application provides an electronic device, including a screen, a memory, one or more data processors, and one or more programs; wherein the one or more programs are stored in the memory; and when the one or more data processors execute the one or more programs, the electronic device implements the vestibular function detection data processing method described in any of the preceding claims.
[0015] Thirdly, this application provides a computer program product that, when run on a computer, causes the computer to execute the vestibular function detection data processing method described in any of the above claims.
[0016] Finally, this application also provides a vestibular function testing system, comprising: an eye-tracking module equipped with at least one camera for acquiring eye-tracking videos of a subject during a vestibular function test; a body position acquisition submodule for acquiring body position data of the subject during the vestibular function test; and one or more data processors, wherein the body position acquisition submodule and the eye-tracking module are operatively coupled to the one or more data processors, and the one or more data processors are configured to perform the vestibular function testing data processing method described in any of the preceding claims.
[0017] In some embodiments, the one or more data processors are located in the same or different devices; the devices include one or more of eye-tracking acquisition devices, body position acquisition devices, near-field data processing terminals, and servers.
[0018] In some embodiments, the eye-tracking imaging module and the body position acquisition submodule are housed in the same or different devices; wherein: if the eye-tracking imaging module and the body position acquisition submodule are integrated into the same device, then the same device is an eye-tracking acquisition device, which is a head-mounted device or an external camera device fixed near the subject by a bracket. If the eye-tracking imaging module and the body position acquisition submodule are respectively housed in different devices, then the different devices include an eye-tracking acquisition device and a body position acquisition device.
[0019] In some implementations, the eye-tracking module is located in the eye-tracking acquisition device, which can be any of the following product forms: wearable eye-tracking acquisition device, head-mounted virtual reality or augmented reality device, or non-wearable shooting device.
[0020] Compared with the prior art, the present invention has at least one of the following beneficial effects:
[0021] 1. This application identifies nystagmus information in specific body positions based on eye movement data and auxiliary detection data (position data and / or nystagmus feature parameters). Compared with only considering eye movement video or eye movement trajectory data, the addition of position data or nystagmus feature parameters makes the assessed vestibular function data more accurate.
[0022] 2. The vestibular function detection system of this application, combined with artificial intelligence technology, achieves the fusion and recognition of input auxiliary detection data and eye movement trajectory data through a trained neural network model. In particular, the neural network model mainly adopts a model architecture capable of processing time-series data, thereby effectively obtaining the correlation and dependencies between time-series data (such as body position data and eye movement trajectory data), capturing complex patterns in body position data and eye movement trajectory data, and thus making better predictions. More preferably, in addition to body position data and eye movement data in various directions, the model input also includes corresponding nystagmus feature parameters obtained based on eye movement trajectory data in various directions, especially the slow phase angular velocity of nystagmus, thereby significantly improving the accuracy of the output vestibular function data.
[0023] 3. The vestibular function data output by the vestibular function testing system of this application can be text information. Compared with traditional parameter data types, the output vestibular function text information is more easy to understand and removes a lot of complicated information that is not related to diagnosis. It only presents the core diagnostic information of characteristic nystagmus in a specific body position. In particular, it is presented in the form of text, so that medical staff without professional training can understand it directly and quickly obtain key clues for vestibular function diagnosis without having to compare and view complex curves. Attached Figure Description
[0024] The preferred embodiments will now be described in a clear and easy-to-understand manner, in conjunction with the accompanying drawings, to further explain the above-mentioned characteristics, technical features, advantages, and implementation methods of the present invention.
[0025] Figure 1 This is a structural block diagram of one embodiment of the vestibular function testing system of this application;
[0026] Figure 2 This is a flowchart of an embodiment of the method for processing vestibular function test data according to this application;
[0027] Figure 3 This is a schematic diagram illustrating body position data in one embodiment of this application;
[0028] Figure 4 This is a schematic diagram of the architecture of a neural network model in one embodiment of this application;
[0029] Figure 5 This is a schematic diagram showing the eye movement trajectory and head position data of a subject in one embodiment of this application. Detailed Implementation
[0030] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the specific implementation methods of the present invention will be described below with reference to the accompanying drawings. Obviously, the drawings described below are merely some embodiments of the present invention. For those skilled in the art, other drawings and other implementation methods can be obtained based on these drawings without any creative effort.
[0031] To keep the drawings concise, each figure only schematically shows the parts relevant to the invention, and these do not represent the actual structure of the product. Furthermore, to facilitate understanding, in some figures, only one of components with the same structure or function is schematically depicted, or only one is labeled. In this document, "one" not only means "only one," but can also mean "more than one."
[0032] It should also be further understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0033] The invention is described herein by way of example with reference to computer system architecture and exemplary processes performed by a computer system. In one or more embodiments, the functionality described herein can be implemented by computer system instructions. These computer program instructions may be directly loaded onto the internal data storage device of a computing device (e.g., the internal data storage device of a smartphone or laptop computing device). Alternatively, these computer program instructions may be stored on a portable computer-readable medium (e.g., a flash drive, etc.) and then subsequently loaded onto the computing device so that the instructions can be executed thereon. In other embodiments, these computer program instructions may be embodied in the hardware of the computing device rather than in its software. The computer program instructions may also be embodied in a combination of both hardware and software. Furthermore, in this disclosure, when referring to a computing device “configured to,” “arranged to,” and / or “configured and arranged to” perform a particular function (e.g., a data acquisition / data processing device), it should be understood that, in one or more embodiments of the invention, this means that the computing device is specifically programmed to perform a particular function (e.g., a data acquisition / data processing device is specifically programmed to perform a particular function).
[0034] This description provides a general overview of the computer programs required for analyzing vestibular function examination information. Any competent programmer in the field of information technology can use the description presented herein to develop such a system.
[0035] For the sake of brevity, traditional computer system components, traditional data networks, and traditional software coding will not be described in detail here. Furthermore, it should be understood that the connecting lines shown in the block diagrams included herein are intended to represent functional relationships and / or operational couplings between the various components. In addition to what is explicitly described, it should be understood that many alternative or additional functional relationships and / or physical connections may be incorporated into the actual application of the system.
[0036] Vestibular function testing assesses the vestibular system's function, including the semicircular canals, otolith organs, or central vestibular system, by observing and analyzing the system's responses to movement, postural changes, and balance control. There are many methods for vestibular function testing, such as the Dix-Hallpike test, the video head impulse test (vHIT), the Roll test, and the Romberg test, to help diagnose vestibular-related disorders (such as vertigo, balance disorders, and benign paroxysmal positional vertigo) and develop treatment plans. Taking benign paroxysmal positional vertigo (BPPV) detection as an example, clinical practice currently relies heavily on nystagmus imaging for BPPV diagnosis and treatment. This involves recording the nystagmus throughout the entire examination process via video. Manual repositioning and swivel chair repositioning can be performed by having the patient wear an infrared video goggles equipped with a gyroscope, thus enabling nystagmus imaging.
[0037] Currently used clinical nystagmus visualization provides limited information and cannot directly reflect the core conclusions of the examination. Although nystagmus visualization can provide complete video playback and analyze eye movement trajectories obtained from video analysis, this information is extremely complex, containing numerous additional details unrelated to diagnosis. The examination and treatment of a single BPPV patient may involve more than ten examinations and treatments, hundreds of sub-movements, and tens of minutes of data. If the operator needs to review the subject's specific condition, they still need to fully replay the nystagmus information, curves, or videos for each position. Furthermore, the core diagnostic information most important to the doctor cannot be directly obtained by reading the results data. Even doctors with complex professional training need to review the complex curves and videos to arrive at a final diagnostic conclusion.
[0038] This application provides a method and system for processing vestibular function test data, which uses artificial intelligence combined with comprehensive consideration of multi-source data to make vestibular function assessment more intelligent and accurate.
[0039] The vestibular function testing system of this application is as follows: Figure 1 As shown, it includes the following functional modules:
[0040] The eye-tracking module 10 includes at least one camera for capturing eye-tracking videos of the subject during a vestibular function test.
[0041] The eye-tracking trajectory acquisition module 20 is communicatively connected to the eye-tracking capture module 10. It extracts eye-tracking trajectory data based on the eye-tracking video captured by the camera. This eye-tracking trajectory data includes eye-tracking trajectories in at least one of the horizontal, vertical, and torsional directions. Of course, the eye-tracking trajectory acquisition module 20 can be integrated into any device, such as an eye-tracking acquisition device, a data processing device, or a cloud server, and processed by the data processor of that device. For example, after the camera of an eye-tracking device captures eye-tracking video, the data processor of the eye-tracking device extracts eye-tracking trajectory features from the video.
[0042] The auxiliary data acquisition module 30 is used to acquire auxiliary detection data of the subject, including the subject's body position data and / or nystagmus characteristic parameters; this module includes at least one of the following:
[0043] The body position acquisition submodule 31 is used to acquire the body position data of the subject during the vestibular function test; preferably, it includes a body position acquisition unit and a body position processing unit; wherein, the body position acquisition unit is used to acquire the body position data of the subject, such as acquiring the subject's head position data through a gyroscope or inertial measurement unit; while the body position processing unit preprocesses the body position data acquired by the body position acquisition unit, such as denoising, data simplification, coordinate transformation, etc., to obtain body position data that conforms to the input of the subsequent neural network model; if the acquired body position data is video data, image processing is required to extract the subject's body position change trajectory from the body position image video frame, that is, what we finally input into the neural network model can be one-dimensional data or two-dimensional data, but not three-dimensional video data.
[0044] The eye-tracking processing submodule 32 is used to identify nystagmus in the subject's eye-tracking trajectory and obtain the corresponding nystagmus feature parameters.
[0045] The intelligent recognition module 40 has a pre-trained neural network model 41 built in. This module is communicatively connected to the eye movement trajectory acquisition module 20 and the data acquisition module 30, respectively, and is used to input the eye movement trajectory and the auxiliary detection data into the neural network model 41 to obtain the subject's vestibular function data. Preferably, the eye movement trajectory and the auxiliary detection data are temporally correlated. Specifically, the intelligent recognition module 40 identifies the nystagmus under specific body positions through the neural network model 41, thereby obtaining the subject's vestibular function data. This specific body position refers to the designated position in the vestibular function test; different test items will have different specific body positions. Furthermore, even for the same test item, there may be multiple different specific body positions, depending on the vestibular function test item, or they can be directly identified by the neural network model (the specific body position needs to be labeled for learning during model training).
[0046] Among the modules mentioned above, the eye-tracking acquisition module 20, the eye-tracking processing submodule 32, the body position processing unit, and the intelligent recognition module 40 are all data processing function modules, and each can achieve its corresponding function by executing corresponding instructions through one or more data processors.
[0047] In addition to the above, the system may also include the following data processing functional modules:
[0048] The preprocessing submodule is used to perform data cleaning and data transformation (such as formatting) on the collected eye-tracking trajectory data or body position data;
[0049] The data synchronization submodule is used to timestamp and synchronize eye-tracking trajectory and auxiliary detection data.
[0050] The post-processing submodule is used to receive the output of the neural network model and perform post-processing on it, such as converting the output into structured data for later display or storage.
[0051] In this embodiment, each data processing module can be integrated into at least one data processor, meaning that data processing and model inference are completed in the same processor, thereby simplifying the system architecture and reducing the complexity of module interactions. More preferably, the intelligent recognition module with a built-in neural network model and other data processing modules are implemented separately through different processors. One data processor is responsible for data preprocessing (such as synchronization, cleaning, and formatting) and post-processing (such as result display and storage), serving as the logical module responsible for general data operations in the system. The other data processor, which integrates the neural network model, focuses on complex feature extraction, fusion, and prediction tasks, acting as a dedicated analysis and inference module. Together, they process the vestibular function detection data.
[0052] Furthermore, the system for processing vestibular function test data also includes:
[0053] Display devices are used to show data processing information, including eye movement trajectory curves, vestibular function data of subjects (such as nystagmus information in specific body positions, vestibular function assessment results, rehabilitation training recommendations, etc.).
[0054] The storage module is used to save the results of data processing, such as the evaluation report generated from the final analysis of the model.
[0055] In the above system embodiments, each module (including sub-modules) can be set up or integrated in the same or different devices. The functions of each module can be realized by configuring the data processor of the device where each functional module is located. Generally, the modules of the vestibular function detection system can be set up in the acquisition device (such as eye-tracking acquisition device, body position acquisition device), data processing terminal device and / or server; the following are some exemplary descriptions:
[0056] An eye-tracking module can be installed in an eye-tracking acquisition device to capture eye-tracking videos of the subject;
[0057] The eye-tracking acquisition module can be installed in any device among eye-tracking acquisition devices, data processing terminal devices, and servers;
[0058] The body position acquisition submodule includes a body position acquisition unit that can be set in a body position acquisition device or integrated into an eye-tracking acquisition device to acquire the body position data of the subject; and a body position processing unit that can be set in a data processing terminal device or a server.
[0059] The eye-tracking processing submodule can be installed in any device among eye-tracking acquisition devices, data processing terminal devices, and servers;
[0060] The intelligent recognition module can be installed in any device, including eye-tracking acquisition devices, data processing terminal devices, and servers.
[0061] It is worth noting that the aforementioned body position acquisition device can be a device that is independent of the eye movement acquisition device. For example, the body position acquisition device may use a device that moves the subject's body position. Alternatively, the body position acquisition device and the eye movement acquisition device may be combined into one device. For example, a head-mounted device with a built-in gyroscope and camera can be used to collect the subject's head movement information and eye movement video separately. Another example is an external adjustable camera, which is set near the subject by a bracket to capture video of the subject's body position changes and eye movement during the vestibular function test. The video data is then processed to obtain the corresponding body position data (not video data, but body position trajectory data in various directions) and eye movement trajectory data.
[0062] The aforementioned data processing terminal equipment can be computers, tablets, mobile phones, medical terminals, or platform devices, or other terminal devices with data processing capabilities. The server can be a local server or a cloud server, etc.
[0063] Preferably, in one example, the vestibular function testing system comprises an eye tracker, a proximal data processing terminal device, and a cloud server. The eye tracker is equipped with at least one camera and at least one motion sensor (such as a gyroscope) to collect eye-tracking video and head position change data, which is then transmitted to the proximal data processing terminal device. The proximal data processing terminal device processes the received data, such as extracting eye-tracking trajectory data in various directions from the eye-tracking video and preprocessing the head position data to conform to the input format of the neural network model before sending it to the cloud server. The cloud server intelligently identifies nystagmus information in specific body positions (specified positions) using the neural network model, obtains the subject's vestibular function data, and feeds it back to the proximal data processing terminal device.
[0064] Of course, if the model's input source also includes nystagmus feature parameters, then the step of obtaining the nystagmus feature parameters can be set on a near-end data processing terminal device or on a cloud server. This embodiment does not limit this.
[0065] In this system embodiment, the near-end data processing terminal is located close to the data acquisition source, enabling it to process the acquired data, reduce irrelevant information, decrease data volume, focus on information content, save transmission bandwidth, accelerate transmission speed, and improve transmission quality. Simultaneously, the eye-tracking trajectory transmitted to the cloud server does not contain privacy information such as pupil images, effectively preventing privacy leaks and ensuring personal data security.
[0066] The near-field data processing terminal (NFC) sits between the data acquisition unit (eye tracker) and the cloud server, serving as a crucial data processing and relay hub. From a hardware perspective, it typically possesses data storage capacity, a computing unit, and a network communication module. Taking a common medical testing scenario as an example, the NFC can be a moderately configured medical workstation computer. It synchronizes data in real-time with the eye tracker via wired or wireless connection. After the patient completes the test, the workstation computer quickly processes the data, converting the raw eye-tracking video into eye-tracking trajectory data in various directions, preprocessing head position change data, and then sending the processed results to the cloud server for in-depth analysis via a wireless network. This approach satisfies the need for rapid local data processing in primary healthcare units while leveraging the powerful computing capabilities of the cloud for more accurate vestibular function assessment, ensuring efficient and secure data transmission. Of course, large medical institutions can also use local servers, enabling rapid vestibular function assessment without an external network connection.
[0067] Another embodiment of this application provides a vestibular function detection system, including: at least one data processor (e.g., a data processor of a computing device such as a mobile device); at least one posture acquisition submodule and at least one eye-tracking module (including at least one camera); the posture acquisition submodule and the eye-tracking module are operatively coupled to the data processor, which can be configured to perform steps of a vestibular function detection data processing method; specifically including:
[0068] Receives the subject's body position data and eye movement video during vestibular function testing; specifically, it receives eye movement video captured by the eye movement imaging module and body position data captured by the body position acquisition submodule.
[0069] Based on eye-tracking videos, eye-tracking data of the subjects is extracted, which includes eye-tracking trajectories of the pupil in at least one of the three directions: horizontal, vertical, and torsional.
[0070] Eye movement trajectory and body position data are input into a trained neural network model, which then identifies and outputs the subject's vestibular function data. The eye movement trajectory and body position data are synchronized over time.
[0071] In this embodiment, the output vestibular function data is derived from the processing of body position data and eye movement data. Compared to the conventional method of predicting vestibular function data solely based on eye movement videos or eye movement trajectories, this embodiment closely integrates the subject's body position information during vestibular function testing. This fusion of body position information significantly improves the accuracy of vestibular function assessment. Taking benign paroxysmal positional vertigo (BPPV) as an example, during positional testing, the subject's body position needs to be changed based on the test type (e.g., the Dix-Hallpike test). The physician observes for the presence of characteristic nystagmus in specific positions. For instance, in diagnosing canaloid type BPPV in the right posterior semicircular canal, a rightward torsional nystagmus with upward eye movement at the superior pole of the eye needs to be observed in the right posterior Dix-Hallpike position for confirmation.
[0072] In this embodiment, the body position acquisition submodule is mainly used to collect the body position data of the subject; the eye movement imaging module includes at least one camera for collecting eye movement video of the subject; the data processor is configured to receive the eye movement video and extract the eye movement trajectory of the subject's pupil in at least one of the three directions: horizontal, vertical, and torsional; the data processor or another data processor is configured to receive the body position data and input the body position data and the extracted eye movement trajectory into a trained neural network model, and identify the nystagmus information under each specific body position through the neural network model, and output vestibular function examination information.
[0073] In one example of a posture acquisition submodule, the submodule includes at least one of an accelerometer configured to detect linear acceleration and a gyroscope configured to detect angular velocity. For example, in an illustrative embodiment, the posture acquisition submodule may include a 3-axis accelerometer, a 3-axis gyroscope, and a 3-axis magnetometer to enable the use of motion fusion algorithms. Of course, pose sensors such as linear accelerometers or gyroscopes can generally be integrated into a head-mounted device to facilitate the acquisition of changes in the subject's head position.
[0074] In another exemplary embodiment, the body position acquisition submodule is a body position change conversion device, such as a BPPV vertigo treatment rotating chair. During the detection or reset process, the BPPV vertigo treatment rotating chair transmits the subject's head position change data at various time points during this process to a data processor for body position recognition, etc. For example, Figure 3 The figure shows the head position data curves of the subjects sitting in the BPPV swivel chair, including the trajectory data of the subjects' head position on the Pitch axis and the trajectory data on the Yaw axis.
[0075] In another exemplary embodiment, the body position acquisition submodule is a camera. The camera can be set up near the subject through a fixed bracket or other device to capture images of the subject's body position changes during the test. The images are then processed to obtain the subject's body position data corresponding to each video frame (corresponding to different time points).
[0076] In one embodiment, a device integrating an eye-tracking module (with eye-tracking functionality) can take any of the following product forms:
[0077] The wearable eye-tracking acquisition device has a built-in eye-tracking imaging module and is convenient for subjects to wear on their heads so that the eye-tracking imaging module can capture the subject's eyes. It can communicate with external devices through wired or wireless communication protocols (such as Bluetooth or Wi-Fi) so that the acquired eye-tracking video or processed eye-tracking trajectory can be transmitted to the device where the subsequent data processing module is located.
[0078] Head-mounted virtual reality (VR) / augmented reality (AR) devices integrate eye-tracking capabilities, utilizing their built-in sensor array and camera combination to capture the user's eye movements and communicate with external independent data processing devices via the device's built-in high-speed internal bus.
[0079] The eye-tracking acquisition device, fixed to the headrest of the testing chair, uses a multi-angle adjustable bracket to fix the camera to accommodate subjects of different heights and sitting postures. It transmits data directly to a nearby desktop computer as the data processing module via wireless or wired network connection.
[0080] Mobile devices with data processors are selected from the group consisting of: (i) smartphones, (ii) tablet computing devices, (iii) laptop computing devices, (iv) smartwatches, and (v) head-mounted displays. For example, in an illustrative embodiment, the body position acquisition submodule and / or eye-tracking module may be a built-in position sensor and / or camera of a video nystagmography device. A video nystagmography device is generally worn in front of the user's eyes, has a built-in camera to capture video of the user's eye movements, and also integrates a position sensor or motion sensor (such as a gyroscope) to capture the user's head movement position data. In another illustrative embodiment, a different type of computing device is used instead of a mobile computing device. For example, the other type of computing device may be a desktop computing device, a server computing device, or a small personal computer. In yet another illustrative embodiment, the user's body position change data can be obtained by the swivel chair change data output by a control unit that controls the movement of the swivel chair the user is sitting in, and by capturing the user's eye video data through a camera of a video nystagmography device worn on the user's head.
[0081] In one illustrative embodiment, after acquiring eye movement video of the subject during the vestibular function test, an eye movement trajectory extraction is performed on the video by a data processor (eye movement trajectory acquisition module) to obtain eye movement trajectory data of the subject's pupil in the horizontal, vertical and torsional directions. Figure 5 The diagram shows the pupillary eye movement trajectories in the horizontal, vertical, and torsional directions, as well as the head position on the pitch and yaw axes during the left-posterior-right-anterior Dix-Hallpike test. The horizontal axis represents time, the vertical axis of the eye movement trajectory represents the eye movement amplitude, and the vertical axis of the head movement trajectory represents the head movement amplitude.
[0082] Regarding data preprocessing for eye-tracking trajectories and auxiliary detection data, timestamps can be used to synchronize the auxiliary detection data and eye-tracking trajectories before inputting them into the model. Taking body position data as an example, both the acquired body position data and eye-tracking data contain time information. To more accurately assess the subject's vestibular function, body position data and eye-tracking data can be correlated through time synchronization. This allows the neural network model to obtain the correlation features between body position data and eye-tracking data during recognition and processing, such as eye-tracking changes caused by changes in body position. Generally, body position data and eye-tracking data are first correlated through time synchronization before being input into the neural network model for recognition and processing. Of course, this time synchronization step can be implemented within the neural network model. For example, after receiving the input body position data and eye-tracking data, the neural network model can preprocess these two types of data through a data preprocessing layer, mainly including time alignment and normalization operations. Data from each sensor is used as an independent input channel to ensure that the data is synchronized in time and that the data format of different sensors is consistent. This is the foundation for subsequent processing.
[0083] Based on the same technical concept, another embodiment of this application provides a method for processing vestibular function test data, such as... Figure 2 As shown, it includes the following steps:
[0084] S101, Acquire eye-tracking videos of the subject during vestibular function testing;
[0085] S102, extract the subject's eye movement trajectory from the eye-tracking video, the eye movement trajectory including at least one of the following: horizontal eye movement trajectory, vertical eye movement trajectory, and torsional eye movement trajectory;
[0086] S103, Acquire auxiliary detection data of the subject; including the subject's positional data in the vestibular function test and / or nystagmus characteristic parameters obtained from eye movement trajectory;
[0087] S104 inputs eye movement trajectory and auxiliary detection data into a pre-trained neural network model, and the neural network model identifies and outputs the subject's vestibular function data.
[0088] In this embodiment, each trajectory point in the eye movement trajectory data contains eye position information and associated time information (for example, horizontal eye movement trajectory data contains sampling time information and the position information of the pupil in the horizontal direction at that time point); similarly, auxiliary detection data also contains time information, which facilitates subsequent time synchronization or association of different types of data.
[0089] Preferably, the data type of the auxiliary detection data is structured parameter data and / or static image data. This auxiliary detection data does not include video data, that is, it does not include dynamic image sequences represented in the form of continuous image frames. Among them, the structured parameter data includes, but is not limited to, sequence data or parameter data, such as body position sequences, head movement trajectory sequences, nystagmus feature parameters, etc.; while the static image data includes, such as body position trajectory diagrams, head movement trajectory diagrams, nystagmus feature diagrams, etc., but does not include dynamic video data composed of continuous image frames.
[0090] Preferably, before inputting the eye-tracking trajectory and auxiliary detection data into the model, data preprocessing is performed: the eye-tracking trajectory and auxiliary detection data are synchronously correlated; furthermore, both the eye-tracking trajectory and auxiliary detection data are time-series data.
[0091] In the above embodiments, the acquired auxiliary detection data may be body position data and / or nystagmus characteristic parameters; taking the auxiliary detection data including body position data and nystagmus characteristic parameters as an example, another embodiment of this application provides a vestibular function detection data processing method including the following steps:
[0092] 1. Raw data acquisition steps; including:
[0093] S201, Acquire eye-tracking videos of the subjects during vestibular function testing;
[0094] S202, Obtain the subject's postural data during the vestibular function test;
[0095] The two steps mentioned above mainly involve acquiring raw data for vestibular function testing. This acquisition can be achieved by directly receiving eye-tracking videos and body position data collected by external acquisition devices, or by actively controlling external acquisition devices (such as body position acquisition devices and eye-tracking acquisition devices).
[0096] 2. Preliminary data processing steps; specifically including:
[0097] 2.1 Eye-tracking data acquisition:
[0098] S203, extract the subject's eye movement trajectory from the eye-tracking video, including at least one of the following three items: horizontal eye movement trajectory, vertical eye movement trajectory, and torsional eye movement trajectory;
[0099] 2.2 Acquisition of auxiliary detection data:
[0100] S204, Analyze and process the subject's eye movement trajectory to obtain the corresponding nystagmus characteristic parameters; specifically, the nystagmus characteristic parameters include any one or more of the following: slow phase angular velocity of nystagmus, direction of nystagmus, duration of nystagmus, and trend of nystagmus change.
[0101] S205 involves preprocessing the subject's posture data to obtain posture data that matches the input of the neural network model; for example, denoising, data simplification, and coordinate transformation of the acquired posture data. Of course, if the acquired posture data is video data, image processing is also required to extract the subject's posture change trajectory from the posture image video frames. In other words, what we ultimately input into the neural network model can be one-dimensional or two-dimensional data, but not three-dimensional video data.
[0102] 3. Intelligent recognition steps; specifically including:
[0103] S206, the eye movement trajectory and auxiliary detection data (nystagmus characteristic parameters and body position data in this embodiment) are input into the trained neural network model, and the neural network model identifies and outputs the subject's vestibular function data. Preferably, the eye movement trajectory and auxiliary detection data are temporally correlated;
[0104] In the above method embodiments, the auxiliary detection data includes nystagmus feature parameters and body position data, which, together with eye movement trajectory data, serve as input to the neural network model. The model then identifies and outputs vestibular function data. Of course, the auxiliary detection data can also be body position data or nystagmus feature parameters.
[0105] Regarding the combination of input data for neural network models:
[0106] In the above system or method embodiments, the auxiliary detection data includes body position data from vestibular function testing and / or nystagmus features obtained from the eye movement trajectory. Therefore, the data combination input to the model includes at least the following three schemes:
[0107] Option 1: The model's input data includes: eye movement trajectory + body position data; preferably, the body position data is head position data (data on changes in the subject's head position).
[0108] This approach focuses on the correlation between body position and eye movements, specifically addressing situations where changes in body position cause changes in eye movements (such as nystagmus). It is particularly suitable for detecting certain conditions where dizziness is induced by changes in body position. Positional tests such as the Dix-Hallpike test and the Roll test, by having the patient move their head to a specific position (i.e., a specific body position) and observing whether nystagmus occurs, can analyze the correlation between nystagmus and changes in body position, thus providing doctors with crucial diagnostic information for diseases such as benign paroxysmal positional vertigo (BPPV).
[0109] Option 2: The model's input data includes: eye movement trajectory + nystagmus feature parameters;
[0110] This scheme focuses on the details of eye movement features—nystagmus feature parameters. On the one hand, eye movement trajectories provide comprehensive eye movement information, and on the other hand, the nystagmus feature parameters extracted from eye movement trajectories can provide important basis for vestibular function assessment.
[0111] Option 3: The model's input data includes: eye movement trajectory, body position data, and nystagmus characteristic parameters. This option integrates these three data points, thus enabling a more comprehensive and accurate assessment of the subject's vestibular function.
[0112] The aforementioned eye movement trajectories mainly include the subject's horizontal eye movement trajectory data, and / or vertical eye movement trajectory data, and / or torsional eye movement trajectory data. Generally, the eye movement data input into the model includes horizontal and vertical eye movement trajectory data, and preferably, it further includes torsional eye movement trajectory data, thereby improving the comprehensiveness of vestibular function diagnosis and providing more key data for vestibular function diagnosis.
[0113] Both Scheme 2 and Scheme 3 above use nystagmus characteristic parameters as one of the input data for the model. These nystagmus characteristic parameters include: slow-phase angular velocity of nystagmus, nystagmus direction, nystagmus duration (duration or onset time), nystagmus latency, and any one or more of the following: nystagmus slow-phase angular velocity, nystagmus direction, nystagmus duration (duration or onset time), nystagmus latency, and nystagmus trend. These nystagmus characteristic parameters are crucial for the diagnosis or assessment of vestibular function. Typical nystagmus often exhibits variations in intensity and a delayed onset. For example, in the right posterior semicircular canal BPPV canaloid type, nystagmus usually appears within 40 seconds after the head is in position and disappears within 1 minute, meaning the nystagmus intensity changes from weak to strong and then back to weak, with a latency of less than 40 seconds. Therefore, combining these nystagmus characteristic parameters with postural data can significantly improve the accuracy of vestibular function assessments and / or recommendations for vestibular rehabilitation training programs based on vestibular function data. Ideally, it is not better to select more nystagmus feature parameters. It is better to select 1-2 nystagmus feature parameters. For example, the addition of the slow phase angular velocity of nystagmus as an input factor greatly improves the accuracy of vestibular function diagnosis or assessment, and thus also improves the accuracy of vestibular function rehabilitation training program recommendations.
[0114] Compared to conventional methods that rely solely on eye-tracking trajectories or eye-tracking videos to assess vestibular function, while eye-tracking trajectories can reveal patterns in eye movement and provide a preliminary assessment, a single data source cannot offer accurate results. The inclusion of auxiliary detection data in this application effectively addresses this deficiency. The selection of auxiliary detection data is also unique; more data types are not necessarily better. More types mean a more complex model, posing significant challenges to model training and accuracy improvement. In this case, positional data was chosen as one of the auxiliary detection data primarily because it complements eye-tracking trajectories, providing richer vestibular function information. This helps the model more accurately capture the causal relationship between positional changes and nystagmus. Eye-tracking trajectories alone cannot reveal the specific positional changes that induce nystagmus and are difficult to correlate with the etiology of vestibular dysfunction (e.g., eye-tracking trajectories alone cannot distinguish between bilateral semicircular canal abnormalities), thus compensating for the limitations of eye-tracking trajectories.
[0115] The reason for using nystagmus feature parameters as an alternative auxiliary detection data is primarily because these parameters provide a deeper quantification and supplement to the dynamic characteristics of eye movement trajectories, enabling a more accurate reflection of the functional state of the vestibular system. Taking slow-phase angular velocity of nystagmus as an example, this core parameter represents the speed of the eyeball during slow-phase motion. Compared to the path information provided by eye movement trajectories, slow-phase angular velocity quantifies the dynamic intensity of eye movement, serving as a crucial supplement to eye movement trajectory data. Furthermore, since it is core information extracted from eye movement trajectories, it avoids information redundancy and noise interference. However, selecting more nystagmus feature parameters is not necessarily better. Ideally, the slow-phase angular velocity of nystagmus should be combined with the eye movement trajectory as model input.
[0116] In summary, this application uses eye-tracking trajectory, body position data, and / or nystagmus feature parameters as model input. On the one hand, eye-tracking trajectory preserves the detailed information of the original data, allowing the model to access global information. On the other hand, body position data and / or nystagmus feature parameters supplement the information, thereby improving the model's accuracy. Furthermore, by introducing a neural network model, automated feature extraction, fusion, and evaluation of multimodal data are achieved, offering higher accuracy and comprehensiveness compared to traditional rule-based and statistical methods. Its end-to-end processing significantly improves efficiency and robustness, while supporting flexible expansion to meet diverse clinical diagnostic and research needs, providing real-time and intuitive evaluation results.
[0117] Regarding the identification of nystagmus and the acquisition of nystagmus feature parameters:
[0118] The acquisition of nystagmus feature parameters mainly involves identifying nystagmus contained in eye movement trajectories in each direction, thereby obtaining the corresponding nystagmus feature parameters. Nystagmus identification can be achieved through software algorithms or artificial intelligence methods; this application does not limit this approach. Figure 5 As shown, after nystagmus identification is performed based on eye movement trajectories in various directions, the identified nystagmus can be marked on the eye movement trajectory diagram. For example, in the schematic diagram of horizontal eye movement trajectory data, if nystagmus occurs in the eye movement trajectory, it will be marked in the part of the eye movement trajectory where nystagmus occurs (short lines near the trajectory in the figure indicate nystagmus).
[0119] In an illustrative example, nystagmus recognition and the acquisition of nystagmus feature parameters include the following steps:
[0120] Obtain the motion features of each data point in the eye-tracking data for each direction;
[0121] Determine whether periodic eye movements exist in eye movement data based on motion characteristics;
[0122] If periodic eye movements are present in the eye movement data, then the periodic eye movements in the eye movement data will be identified as nystagmus;
[0123] The eye movement trajectory of the data segment identified as nystagmus is analyzed to determine the slow phase and fast phase of nystagmus, thereby obtaining the characteristic parameters of nystagmus.
[0124] Periodic motion generally refers to regular, repetitive movements. A key characteristic of periodic motion is its continuous repetition over a period of time. In nystagmus, the eyeball makes a slow movement in a specific direction (slow phase), followed by a rapid return to its original position (fast phase), and then repeats this pattern again. That is, the fast and slow phases alternate.
[0125] In the above embodiments, nystagmus recognition primarily involves identifying periodic eye movements from the eye movement trajectory. The nystagmus recognition conditions include: periodic eye movements, specifically, alternating fast and slow phases of nystagmus; more preferably, the duration of the periodic eye movements exceeds a first time threshold; and / or the duration of the slow phase of nystagmus conforms to a preset time range; and / or the acceleration of the slow phase of nystagmus conforms to a preset acceleration range. The slow phase has a relatively long duration and a slower movement speed, while the fast phase has a shorter duration and a faster movement speed. Therefore, by combining a time window and a velocity curve, the two can be effectively distinguished.
[0126] In another illustrative example, the data processor is configured to execute the following instructions to obtain nystagmus feature parameters:
[0127] Preprocess the eye movement trajectory data in each direction of the eye movement data, such as noise reduction and / or simplification (extract eye movement trajectory data points at equal intervals to reduce the amount of data processing);
[0128] Iterate through each data point in the preprocessed eye-tracking data, compare the positional change trend of each data point with the previous data point, and find the data points whose change trend changes as key data points.
[0129] Based on all the key data points found, the eye-tracking trajectory data is divided into N data segments; each data segment starts from one key data point and ends at the next key data point; N is a positive integer greater than 1.
[0130] Calculate the slope of each data segment, and based on the calculated slope, identify the first data segment with a slope greater than a first set value and the second data segment with a slope less than a second set value; wherein the first set value is greater than or equal to the second set value.
[0131] Traverse all data segments and filter out the first and second data segments that meet the nystagmus recognition criteria as nystagmus. Among them, the first data segment marked as nystagmus is the fast phase of nystagmus, indicating the direction of nystagmus; the second data segment is the slow phase of nystagmus, indicating the intensity of nystagmus, which is the slow phase angular velocity of nystagmus.
[0132] In the above example, based on the target eye movement trajectory data after preprocessing the eye movement data, the positional change trends of the data points before and after are compared to identify the data points whose positional change trends have changed as key data points. After finding all key data points, adjacent key data points are connected to form line segments. The slope of each line segment is calculated to identify the first and second data segments. Preferably, after calculating the slope of each data segment, the data segments with a slope greater than a set first slope are identified as the first data segment, and the data segments with a slope less than a second slope are identified as the second data segment. Then, the first and second data segments that meet the nystagmus recognition criteria are selected and marked as nystagmus. Regarding the nystagmus recognition criteria, as mentioned before, for example, the first and second data segments appear alternately; and / or the duration of the second data segment conforms to a preset time range.
[0133] In another illustrative embodiment, in addition to identifying nystagmus in the eye movement data in each direction, further, position recognition processing is performed based on the collected time-series position data to obtain specific positions and their start and end time points in the vestibular function test. Then, based on the start and end time points of the specific positions, eye movement trajectory feature data within the corresponding time period is obtained to determine whether there is nystagmus in the eye movement trajectory feature data. If there is, other nystagmus feature parameters under the specific position are further obtained, including nystagmus latency, nystagmus trend change, nystagmus duration, etc.
[0134] Regarding the model architecture of neural network models:
[0135] Depending on the input data, the specific network architecture of a neural network model will vary, but the overall structural framework is similar and may include: an input layer (which receives external data into the network), a feature extraction layer (which extracts features from the input data), a feature fusion layer (which fuses different types of features) / feature interaction layer (which captures the correlation between different types of data / features), and an output layer (which outputs the final classification result or generated text information); among which:
[0136] The neural network model of this application can have a single-input channel in its input layer, but more preferably multiple input channels. If a single-input channel is used, the input data are combined, for example, body position data and eye movement trajectory are combined to form a comprehensive input format containing both body position and eye movement trajectory. This can be achieved by concatenating or creating a structured input vector. If the input layer uses multiple input channels, for example, if body position data in two directions and eye movement trajectory in three directions are used as model input, the input layer can have five input channels, with one input channel receiving body position / eye movement trajectory data in one direction.
[0137] After receiving input data through multiple input channels, a feature extraction layer extracts features from the data input from the multiple input channels, and a feature fusion layer performs feature fusion processing. Taking the auxiliary data processing network and eye-tracking data processing network as examples of the feature extraction layer in this application, the neural network model includes:
[0138] The auxiliary data processing network is used to process the input auxiliary detection data and extract auxiliary features;
[0139] An eye-tracking data processing network is used to process the input eye-tracking trajectories and extract eye-tracking trajectory features;
[0140] The feature fusion layer is used to fuse the auxiliary features output by the auxiliary data processing network and the eye movement trajectory features output by the eye movement data processing network to obtain comprehensive feature data related to vestibular function.
[0141] The output layer is used to generate corresponding vestibular function data based on the output data of the feature fusion layer.
[0142] Since the auxiliary detection data includes body position data and / or nystagmus feature parameters, the corresponding auxiliary data processing network also includes a body position data processing network and / or a nystagmus feature processing network; where:
[0143] The body position data processing network is used to process the input body position data and extract body position features;
[0144] The nystagmus feature processing network is used to receive input nystagmus feature parameters.
[0145] Since the auxiliary detection data contains different data, the input data of the model will also be different. Below, we will describe the neural network model architecture under different input schemes:
[0146] Taking Scheme 1 above (input data being body position data and eye movement trajectory) as an example, the feature extraction layer of the neural network model in this embodiment includes a body position data processing network and an eye movement data processing network; wherein:
[0147] The body position data processing network is used to process the input body position data and extract body position features;
[0148] An eye-tracking data processing network is used to process the input eye-tracking trajectories and extract eye-tracking trajectory features;
[0149] The feature fusion layer is used to fuse and correlate the postural features output by the postural data processing network and the eye movement trajectory features output by the eye movement data processing network to obtain comprehensive feature data related to vestibular function.
[0150] The output layer is used to generate corresponding vestibular function data based on the output data of the feature fusion layer.
[0151] More preferably, in an exemplary embodiment, the eye-tracking data processing network includes at least two sub-networks, each sub-network being used to receive and process eye-tracking trajectory data in one direction; specifically, for example, the eye-tracking data includes the eye-tracking trajectory of the subject's pupil in three directions: horizontal eye-tracking trajectory, vertical eye-tracking trajectory, and torsional eye-tracking trajectory; then the eye-tracking trajectories in these three directions are respectively processed by feature extraction through the three sub-networks of the eye-tracking data processing network.
[0152] In another exemplary embodiment, the body position data processing network includes at least two sub-networks, each sub-network being used to receive and process body position data in one direction; specifically, for example, body position data of the subject's head position on the Pitch axis is processed for feature extraction through one sub-network of the body position data processing network, while body position data of the subject's head position on the Yaw axis is processed for feature extraction through another sub-network of the body position data processing network.
[0153] In another embodiment of this application, as described in Scheme 3, the input to the neural network model includes three data types: eye movement trajectory data in various directions, body position data, and nystagmus feature parameters. The eye movement trajectory data and body position data are obtained through preliminary processing of the collected raw data, while the nystagmus feature parameters are higher-dimensional feature data further extracted from the eye movement trajectory data. Therefore, when using these three types of data as input, the architecture of the neural network model can be further adjusted and optimized. Specifically, the neural network model includes three network structures (these three network structures can be configured into different structures according to the characteristics of the data to be processed). The body position data processing network and the eye movement data processing network can be described with reference to the previous embodiments and will not be repeated here. The nystagmus feature processing network is used to receive the input nystagmus feature parameters. Subsequently, the feature fusion layer fuses the features output by each network, and finally outputs vestibular function data through the output layer. The feature fusion layer in this embodiment can also adopt different implementation forms; the following are some examples:
[0154] (1) Example 1 of Feature Fusion Layer
[0155] like Figure 4 As shown, the body position features in each direction output by the body position data processing network, the eye movement trajectory features in each direction output by the eye movement data processing network, and the nystagmus feature parameters in each direction output by the nystagmus feature processing network are fused together to obtain the fused overall feature vector. Finally, the vestibular function data is obtained based on the fused overall feature vector through the output layer.
[0156] (2) Example 2 of Feature Fusion Layer
[0157] In this example, the feature fusion layer comprises a first fusion layer and a second fusion layer. The first fusion layer fuses the positional features extracted by the positional data processing network and the eye movement trajectory features extracted by the eye movement data processing network to obtain preliminary comprehensive feature data. Then, the second fusion layer fuses this preliminary comprehensive feature data with nystagmus feature parameters. More preferably, a self-attention mechanism is also employed in the feature fusion layer. Specifically, a self-attention mechanism can be set in the first and / or second fusion layers to assign weights to different features, thereby improving the accuracy of the final evaluation data.
[0158] In this example, multi-level feature fusion is used. By gradually fusing low-level features (such as eye movement trajectory and body position data) with high-level features (such as nystagmus feature parameters), a more complete feature representation is achieved to support the preparation, judgment and assessment of vestibular function.
[0159] The goal of multi-level feature fusion is to enable the model to progressively extract more representative high-level features from low-level raw data by integrating features at different levels. Specifically, in this example, the steps to implement multi-level feature fusion include the following:
[0160] S1, Low-level feature extraction:
[0161] Eye movement feature data is extracted from eye movement trajectory data using a body position data processing network (such as a convolutional neural network and / or a recurrent neural network); dynamic features of body position changes are extracted from body position data using an eye movement data processing network (such as an LSTM or RNN) to obtain body position feature data.
[0162] S2, Mid-level Feature Construction
[0163] In the intermediate layer, the low-level features are initially integrated through the first fusion layer; specifically, the eye movement trajectory features in each direction and the body position features of the body position data are integrated to form an overall intermediate feature representation.
[0164] S3, Introduction of High-Level Features
[0165] The nystagmus feature parameters are introduced into the model as high-level features. In the feature fusion layer, the high-level features are further fused with the mid-level features obtained in the previous step through the second fusion layer. Furthermore, a weighted fusion method can be used so that the model can automatically adjust the importance of features according to the specific situation of the input feature data.
[0166] S4, based on the final fused feature data, passes through a fully connected layer to output the vestibular function assessment results.
[0167] (3) Example 3 of Feature Fusion Layer
[0168] Similarly, in this example, the feature fusion layer also contains a first fusion layer and a second fusion layer. The first fusion layer is used to fuse the eye movement trajectory features extracted by the eye movement data processing network and the nystagmus feature parameters output by the nystagmus feature processing network. The second fusion layer further fuses the data output by the first fusion layer with the postural feature data output by the postural data processing network. Finally, the data fused by the second fusion layer is used for evaluation and analysis to obtain the vestibular function data of the subject.
[0169] In this example, eye movement trajectory features are first extracted based on the eye movement trajectory itself, and then fused with nystagmus feature parameters. This approach fully utilizes the spatiotemporal characteristics of the eye movement trajectory data, while combining it with nystagmus feature parameters to obtain a more complete expression of the eye movement trajectory features. Since nystagmus feature parameters are closely related to eye movement trajectories, fusing them first helps to extract eye movement trajectory features with greater diagnostic value. Furthermore, fusing eye movement trajectory features first, and then fusing them with body position features, allows for hierarchical feature processing, reducing interference between different features and making it easier for the model to extract unique features in each direction.
[0170] Preferably, in addition to this, a self-attention mechanism is set in the first and / or second fusion layer. The self-attention mechanism sets different weights for each feature in the fused data, thereby making the final recognition output data more accurate.
[0171] The neural network model architecture in the above embodiments uses a feature fusion layer for feature fusion. In another embodiment of this application, the neural network model uses a feature interaction layer. The specific implementation steps are as follows:
[0172] Multiple input channels: Receives data from multiple input channels, such as eye-tracking trajectory data in three directions (horizontal, vertical, and torsional) and body position data in two directions (Pitch axis and Yaw axis).
[0173] Independent feature extraction: Preliminary feature extraction is performed on the eye movement trajectory data and body position data input from each input channel. For example, LSTM layers are used to extract time series features, or CNN layers are used to extract spatiotemporal features. Feature representations are extracted for eye movement and body position data in each direction.
[0174] Feature Interaction Layer: This layer applies self-attention or bilinear interaction to process the extracted eye-tracking and body position features, generating new interactive features. Ideally, bilinear interaction (Multiply) can be used: the features in each eye-tracking direction are element-wise multiplied with the body position features to generate interactive features. Then, a self-attention mechanism is applied to these interactive features to dynamically adjust the weights of each feature, allowing the model to focus more on diagnostically meaningful feature combinations.
[0175] Feature fusion and output: The interactive features are finally fused through the feature fusion layer, input into the fully connected layer, and output the evaluation results of vestibular function.
[0176] Preferably, if the input model data also includes nystagmus feature parameters, then after the body position data processing network outputs body position feature data, the eye movement data processing network outputs eye movement trajectory feature data, and the nystagmus feature processing network outputs nystagmus feature parameters, the feature interaction layer is used to perform feature interaction on the body position feature data, eye movement trajectory feature data, and nystagmus feature parameters.
[0177] The core function of the feature interaction layer is to directly generate new interactive features between nystagmus and / or eye movement and postural features, capturing their interrelationships and dependencies. Through element-level operations (such as multiplication) or self-attention mechanisms, the interaction layer can dynamically adjust and combine different features to generate more diagnostically meaningful combined features. Because the feature interaction layer can reveal the relationships between different features, it makes the model easier to understand, such as the influence of specific body positions on eye movement trajectories.
[0178] Although the feature interaction layer already provides some feature fusion effect, retaining the feature fusion layer can still further improve the model's performance in some complex vestibular function assessment tasks for the following reasons:
[0179] Further integration of multi-level features: If the relationship between eye-tracking features and body position features is very complex, the interactive features generated by the feature interaction layer may not be sufficient to capture all the information. In this case, an additional feature fusion layer can further integrate the multi-level information of the interactive features.
[0180] Introducing higher-level feature representation: After feature interaction, the feature fusion layer can serve as a higher-level processing layer to integrate the combined results of different interactive features and generate a more representative comprehensive feature.
[0181] Therefore, an architecture of feature interaction layer + feature fusion layer can be adopted. The features generated by the feature interaction layer are further fused with body position and eye movement data, enabling the model to capture more comprehensive contextual information.
[0182] The above mainly describes the main architecture of the neural network model of this application. In actual application, different models can be adopted according to the actual situation. Generally, eye movement trajectory data and auxiliary detection data contain time information, so time correlation or alignment can be performed later. Taking Scheme 1 as an example, the body position data and eye movement trajectory data in the input data are time series data. Therefore, models suitable for processing time series data can be adopted, such as long short-term memory networks, combinations of convolutional neural networks and LSTM, Transformer models, etc.
[0183] Taking Scheme 3 as an example, assuming the input data includes: horizontal eye movement trajectory, vertical eye movement trajectory, pitch axis head position data, yaw axis head position data, horizontal slow phase angular velocity data of eye movement, and vertical slow phase angular velocity data of eye movement; the network architecture example of the neural network model in this embodiment is as follows:
[0184] (1) Input layer, responsible for receiving all types of input data. We can use six input channels to receive the above six types of input data. Preferably, these input data need to be preprocessed, such as standardization or normalization, time window segmentation, etc. Each input can be a sequence of time-series data (for example, eye movement trajectory and nystagmus information are sequences that change over time, while head position data is coordinate information of the time series).
[0185] (2) Feature extraction layer, mainly used to extract spatial and temporal features from input data. Convolutional layers or other network structures can usually be used. Here is an example:
[0186] Eye movement trajectory data processing: For horizontal and vertical eye movement trajectory data, LSTM networks are used to capture changes in eye movement over time and extract eye movement trajectory features; independent LSTM networks can be used for processing in each direction.
[0187] Head position data processing: For pitch axis head position data and yaw axis head position data, LSTM or CNN can also be used to capture the changes in head position data over time and extract the corresponding head position features;
[0188] Nystagmus information processing: Since nystagmus information is actually a higher-level feature information extracted based on eye movement trajectory, the received nystagmus information can be directly entered into the subsequent feature fusion layer without further extraction processing; of course, nystagmus features can also be further extracted through CNN, or lightweight mapping processing can be performed through a layer of MLP, which is not limited in this embodiment.
[0189] Ideally, a self-attention mechanism can also be introduced in the feature extraction layer to help the model focus on key moments, such as determining the head position change or maintenance phase based on needs.
[0190] (3) Feature Fusion Layer: This layer fuses the features extracted by each network in the feature extraction layer, merging different types of data features and capturing the relationships between them to generate comprehensive feature data. Preferably, a self-attention mechanism can also be introduced in the feature fusion layer to dynamically adjust the weights of different features.
[0191] Specifically, we fuse features extracted from different inputs (eye movement trajectory, head position, nystagmus, etc.) and assign a dynamic weight to each feature using an attention mechanism. This attention mechanism ensures the model can flexibly adjust its focus on different inputs based on task requirements by weighting different features. For example, nystagmus features are crucial for BPPV diagnosis; therefore, in the BPPV domain, higher weights can be assigned to nystagmus features.
[0192] (4) Output layer: Based on the comprehensive feature data output by the feature fusion layer, identify and output the vestibular function test results.
[0193] Using the above example models, the output classification data of vestibular function includes, but is not limited to: one or more nystagmus parameters indicating nystagmus status, and / or one or more diagnostic parameters indicating vestibular function, and / or one or more recommended parameters indicating vestibular rehabilitation training. The specific parameters output by the neural network model can be used to indicate or characterize the corresponding vestibular function information.
[0194] Of course, if the desired output is textual information rather than simple classification results, then the neural network model used cannot be the same as the one described above. Instead, a neural network model capable of generating coherent text is required. For example, an encoder-decoder architecture with an attention mechanism can be used. This architecture can transform the model's input data into detailed diagnostic text. Taking a sequence-to-sequence (Seq2Seq) model with an attention mechanism as an example, a Seq2Seq model typically consists of two parts: an encoder and a decoder. The encoder processes the input data, and the decoder generates the output text. Adding an attention mechanism allows the model to focus more on the relevant parts of the input sequence when generating each word, thereby improving the relevance and accuracy of the generated text. A specific architecture example is shown below:
[0195] Sequence-to-sequence (Seq2Seq) models with attention mechanisms include:
[0196] Input layer: Receives preprocessed gyroscope and eye-tracking data, which may be converted into a format more suitable for model processing after feature extraction.
[0197] Encoder: Uses multi-layer LSTM or GRU units to process time series inputs. The encoder learns to understand the context of the input sequence and compresses this understanding into a fixed-length context vector.
[0198] Attention mechanism: When computing the decoder's output at each step, the attention mechanism dynamically selects information from the input that is most relevant to the current output. This helps the model better understand which parts of the input data are most important for generating the current output vocabulary.
[0199] Decoder: Also using LSTM or GRU units, it uses the context vector and attention mechanism of the encoder to weight the input information and gradually generate each word of the diagnostic text.
[0200] Output layer: Usually a fully connected layer, which uses the softmax function to output the probability of words at each step, and selects the word with the highest probability as the output of the current step.
[0201] In addition to the model architectures mentioned above, an enhanced version of the Transformer model can also be used to generate text. Because the Transformer model excels at processing sequential data, especially in natural language processing, it is also a good choice for generating diagnostic text. The Transformer is entirely based on attention mechanisms and has no recurrent layers, making it highly efficient in parallel processing and capturing long-range dependencies.
[0202] Enhanced Transformer model architecture example:
[0203] Input layer: Same as above, with preprocessed gyroscope and eye-tracking data as input.
[0204] Location encoding: Add location information because Transformer does not naturally handle the time-series nature of data.
[0205] Transformer encoder: Composed of multiple self-attention layers and feedforward networks, it processes input data and encodes comprehensive contextual information of the input sequence.
[0206] Transformer decoder: It also consists of multiple layers, but in addition to processing the encoder's output, it also predicts the next word, with each prediction based on the previous output.
[0207] Output layer: Generates the final diagnostic text, selecting the word with the highest probability from the decoder output each time.
[0208] Whether using a Seq2Seq model with attention or a Transformer model, the key lies in effectively converting gyroscope and eye-tracking data into a form the model can understand, and training the model to accurately generate textual information reflecting BPPV diagnosis. This may require a large amount of labeled data to train the model, as well as meticulous parameter tuning and optimization.
[0209] Example of combining two neural network models:
[0210] The first neural network model mainly adopts a network model that can process time series data. For example, it can adopt any model architecture such as a long short-term memory network, a combination of convolutional neural network and LSTM, or a Transformer model. It is used to identify and output nystagmus information under specific body positions based on the input body position data and eye movement data.
[0211] The second neural network model mainly adopts a model capable of text processing, such as an existing language model, to output text information of vestibular function data based on the nystagmus information of each specific body position output by the first neural network model.
[0212] Compared to using only one model architecture to achieve text output, the combination of two neural network models in this example greatly reduces the complexity of the neural network model, the difficulty of data processing, and the difficulty of model training, while improving the accuracy of the model output.
[0213] The vestibular function data output in this embodiment is vestibular function text information, including but not limited to nystagmus text data identified in specific body positions, and / or vestibular function diagnosis or assessment text data, and / or vestibular rehabilitation training program recommendation text data. The neural network model used in this embodiment has the function of generating corresponding vestibular function text information based on input data. For example, the sequence-to-sequence model with attention mechanism mentioned above, the enhanced Transformer model, etc., can also be used to process the data. The latter is specifically used for text processing to simplify the model complexity. Since the model's input data contains multiple types, if multiple input channels are used to receive different types of input data, the feature extraction and subsequent feature fusion of each input data within the model can also refer to the previous embodiments, which will not be repeated here.
[0214] Specifically, if the trained neural network model aims to output key nystagmus text information to provide doctors with a reference for vestibular function diagnosis, then during training, the vestibular function examination information labeled in the medical records used as training samples will be the characteristic nystagmus text information of this vestibular function examination.
[0215] In one illustrative embodiment, during training, the neural network model uses a large number of different medical records to form a training sample dataset. A complete medical record should include auxiliary detection data (such as body position data and / or nystagmus characteristic parameters), eye movement trajectory data, and the results of the subject's vestibular function test given by the doctor during at least one vestibular function test.
[0216] In another embodiment of this application, based on any of the above embodiments, before inputting the body position data into the neural network model, the following is further included:
[0217] The body position data is processed for body position recognition to obtain specific body positions and their start and end time points in the vestibular function test.
[0218] Based on the start and end times of each specific body position, the corresponding time periods of the input data of the neural network model are marked. For example, if the start and end times of a specific body position A are detected to be in the time period a1-a2, then the data in the body position data, and / or eye movement trajectory data, and / or nystagmus feature parameter data with time points in the time period a1-a2 are marked so that the neural network model can focus on the data in this time period and increase the recognition weight of this time period when it is being processed.
[0219] Of course, the above-mentioned identification and processing of body position data can be implemented in various ways, and there is no limitation to this. In one exemplary implementation, a trained body position recognition neural network model is used to identify the body position data and output the specific body position information contained in the body position data. In another exemplary implementation, the specific body position and its start and end times can be obtained through the following steps:
[0220] Extracting the coordinates of each data point from time-series body position data;
[0221] Based on the coordinate data of all data points, consecutive data points with the same coordinates are identified and grouped into a single body position.
[0222] Obtain the duration of the held position and select the held position with a duration greater than a set threshold as the target held position;
[0223] Obtain the start and end times of the target's maintained position, as well as the corresponding coordinate data;
[0224] Based on the coordinate data of the target's maintained position, and combined with the corresponding reference relationship between each specific position type and coordinates, the specific position type corresponding to each target's maintained position is identified.
[0225] In another exemplary embodiment, if head position and posture change data—head position data—is collected via a gyroscope built into the head-mounted device, then the head position data collected by the gyroscope is processed for head position and posture recognition to identify the subject's head position changes throughout the vestibular function test. Taking the right posterior Dix-hallpike test as an example, the subject's body position (head position) change process is as follows:
[0226] Initially, the subject sits on the examination bed (sitting position). The doctor supports the subject's head with both hands and rotates it 45 degrees to the right, maintaining this position (hereinafter referred to as the 45-degree right turn position). Then, the subject is quickly changed to a supine position with the head hanging off the bed at a 30-degree angle to the horizontal plane, maintaining the 45-degree head position. This position (hereinafter referred to as the hanging-supine position) should be maintained for 30 seconds. The subject then slowly returns to the sitting position.
[0227] Throughout the experiment, the subjects wore head-mounted devices with built-in gyroscopes. Starting from an initial pose, the gyroscopes captured changes in head position as the subjects' posture shifted. By processing this series of head position data—many existing techniques rely on gyroscope data for position / pose recognition, which will not be elaborated upon here—the specific changes in the subjects' head positions could be identified.
[0228] Besides identifying specific body positions based on collected posture data, another illustrative embodiment of identifying specific body positions involves filtering data segments where the body position remains unchanged for a set time to determine the specific body position. For example, in the right posterior Dix-Hallpike test, a specific body position mainly refers to a position that remains unchanged for a period of time after a change in posture. In the above test, the patient's head rotated 45 degrees to the right and the body position remained unchanged; therefore, this 45-degree right turn can be considered a specific body position. More importantly, in the suspended-supine position (head hanging backward off the bed at a 30-degree angle to the horizontal plane, while maintaining a 45-degree head position), the doctor observes whether the subject experiences dizziness or nystagmus. The subject needs to maintain this position for a certain period of time; therefore, this suspended-supine position is a specific body position. We can see that the above-mentioned specific body positions must be positions where the body position remains unchanged, while any changing movements will not qualify as specific body positions. Therefore, specific body positions can be identified by filtering body positions that remain unchanged for a set time.
[0229] Regarding the identification and judgment of specific body positions, in another illustrative embodiment, in addition to the above-mentioned screening of specific body positions based on the body position holding time, further screening can be performed by combining the body position identification results. Specifically, the body position data is identified to obtain the body position type corresponding to each data point, and based on the starting time period corresponding to the body position where the body position holding time reaches a set time, the body position type corresponding to that time period is searched. If the body position type also belongs to the specific body position type, then the body position corresponding to that time period is further determined to be the specific body position.
[0230] Preferably, after identifying each body position data, especially each specific body position, the test type of this vestibular function test is identified based on the time sequence of the identified body positions / specific body positions.
[0231] In another example, such as a video nystagmography device, because it has a built-in pose sensor and a camera to capture the user's eyes, it can function as a posture acquisition submodule to collect posture data and also capture the user's eye movement video to obtain eye movement trajectories. By associating and synchronizing the posture information with the eye movement trajectory data, eye movement trajectories for each posture can be obtained. Based on the eye movement trajectory corresponding to the start and end times of that posture, characteristic nystagmus data for that specific posture can be calculated. This characteristic nystagmus data includes nystagmus type, nystagmus intensity, nystagmus latency, nystagmus duration, and nystagmus trend. Compared to the previous embodiment where nystagmus identification was based solely on eye movement trajectories, this example, by using the start and end times of a specific posture, can quickly identify whether nystagmus exists within that time period. If it exists, further determination of nystagmus characteristic parameters is possible, thus significantly reducing the amount of data processing and improving data processing efficiency and accuracy.
[0232] Regarding data optimization and processing:
[0233] In another embodiment of this application, based on any of the above embodiments, the vestibular function data consists of nystagmus information under specific body positions in the current vestibular function detection. After identifying the nystagmus information under specific body positions through the neural network model or other algorithms, it further includes:
[0234] The identified nystagmus information in various specific body positions is subjected to one or more of the following data optimization processes: information simplification, professional expression processing, and structured text processing, to generate vestibular function text data. The data optimization processing steps include, but are not limited to, one or more of the following: information simplification (such as simplifying the data representation, filtering invalid information, merging valid information, etc.), professional expression processing (such as feature visualization processing), and structured text processing.
[0235] In one illustrative embodiment, the data processor is configured to perform one or more of the following steps to simplify the data representation:
[0236] (1) Based on the duration of the nystagmus latency in the identified nystagmus information under each specific body position, and combined with several duration stages of the set nystagmus latency, a classification description is performed. Specifically, the time value is linked to the diagnosis, and is no longer presented directly as a numerical value, but is modified according to the latency duration to no latency, less than 10 seconds, less than 40 seconds, and more than 40 seconds.
[0237] (2) Based on the duration of nystagmus in the nystagmus information identified in each specific body position, and combined with several duration stages of the set nystagmus duration, classify and describe it; for example, divide the duration of nystagmus into: less than half a minute, less than one minute and more than one minute.
[0238] In one illustrative embodiment, at least one data processor is configured to filter invalid information by performing one or more of the following steps:
[0239] (1) Nystagmus with intensity below a set threshold is uniformly described as no obvious nystagmus;
[0240] (2) Those who maintain nystagmus throughout the entire testing process are uniformly described as spontaneous nystagmus;
[0241] (3) Based on the nystagmus recognition results in the horizontal, vertical and torsional directions, only the nystagmus recognition results in the corresponding direction where the nystagmus intensity reaches the set threshold are retained.
[0242] In one illustrative embodiment, at least one data processing is configured to perform feature representation processing by performing the following steps: for nystagmus identified from horizontal eye-tracking data, to provide a specialized description of the nystagmus direction based on the current specific body position and the relationship between the nystagmus direction and the ground.
[0243] Specifically, in the diagnosis of horizontal semicircular canal BPPV, nystagmus is usually described as geotropic nystagmus and geotropic nystagmus based on its direction relative to the ground, rather than horizontal left or right nystagmus, as this makes it easier to associate with the diagnosis. In the right lateral decubitus position, horizontal right nystagmus transforms into geotropic nystagmus, and horizontal left nystagmus is geotropic nystagmus; in the left lateral decubitus position, the opposite is true.
[0244] In the Dix-Hallpike test, if there is typical nystagmus when lying down and typical nystagmus when sitting up, it is described as nystagmus in the opposite direction, rather than being described directly by direction.
[0245] In one illustrative embodiment, at least one data processing module is configured to perform effective information merging by performing any one or more of the following steps:
[0246] (1) Combine the descriptions of nystagmus induced by the same diagnostic element under different body positions. For example, in the supine position and the right posterior Dix-Hallpike position, an upward judder accompanied by torsional nystagmus to the right is observed, with a latency of less than 10 seconds and a duration of less than half a minute.
[0247] (2) All diagnostic elements for the same diagnosis should be combined into a single description. For example, if the Roll test shows left nystagmus when lying on the left side and right nystagmus when lying on the right side, then this should be combined into a single description of bilateral geotropic nystagmus in the Roll test, followed by a description of other features.
[0248] (3) When there are multiple diagnostic elements for different diagnoses, describe them in order of nystagmus intensity. For example, for right posterior to right horizontal, the right posterior Dix-Hallpike test will show an upward jump accompanied by torsional nystagmus to the right, while the right horizontal test will show horizontal nystagmus to the left and to the right.
[0249] (4) Combine and describe the nystagmus information detected at different times in the same body position during the complete vestibular function test. Specifically, establish connections between the complete examination process. For example, if a specific nystagmus is observed during the first roll test, describe it. During the second test, directly describe whether there are characteristic changes in the nystagmus, rather than directly repeating the description of the characteristics. If there are no changes in nystagmus, combine the descriptions from multiple tests.
[0250] (5) Establish a link between diagnosis and treatment, and combine information on nystagmus characteristics before and after diagnosis and treatment. Specifically, if a patient undergoes a complete diagnostic process, then treatment, and finally a follow-up examination, when combining effective information, summarize and describe the information before treatment, summarize and describe the characteristics during treatment, and compare the characteristics after treatment with the characteristics before treatment.
[0251] In one illustrative embodiment, structured text processing includes:
[0252] Based on the set test type and body position sequence, a segmented report on nystagmus for each body position is generated;
[0253] Based on the segmented reports of nystagmus corresponding to each body position, and combined with the vestibular function assessment criteria, a vestibular function assessment text is generated.
[0254] In this embodiment, the key information required for vestibular function diagnosis is output through the aforementioned data optimization processing, while invalid data is filtered out. Highlighting the core information most relevant to the diagnosis and filtering out a large amount of complex and invalid information makes the test results clearer and more conducive to the doctor's judgment; thereby improving the efficiency and accuracy of the doctor's diagnosis and treatment, and effectively reducing the possibility of missed diagnosis and misdiagnosis.
[0255] The embodiments of the vestibular function data processing method and the vestibular function data processing system of this application correspond to each other, and their technical details can be referred to each other. To avoid repetition, they will not be described again.
[0256] Another embodiment of this application provides an electronic device, including a screen, a memory, one or more data processors, and one or more programs; wherein, the one or more programs are stored in the memory; when the one or more data processors execute the one or more programs, the electronic device implements the vestibular function detection data processing method of this application, and the specific steps of the vestibular function detection data processing method of any of the above embodiments can be referred to.
[0257] Another embodiment of this application discloses a storage medium storing computer-executable instructions. When executed by a processor, the computer-executable instructions cause the processor to perform the steps of the vestibular function data processing method in any of the above method embodiments or the steps configured to be executed by the data processor in any of the system embodiments. To avoid repetition, further details are omitted here.
[0258] Another embodiment of this application also discloses a computer program product that, when run on a computer, causes the computer to execute the vestibular function detection data processing method of any embodiment of this application.
[0259] It should be noted that the above embodiments can be freely combined as needed. The above are merely preferred embodiments of the present invention. It should be pointed out that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A method for processing vestibular function test data, characterized in that, include: Acquire eye-tracking videos of subjects during vestibular function testing; The subject's eye movement trajectory is extracted from the eye-tracking video; the eye movement trajectory includes at least one of the following: horizontal eye movement trajectory, vertical eye movement trajectory, and torsional eye movement trajectory; Acquire the subject's positional data during the vestibular function test; The eye movement trajectory and the body position data are input into a pre-trained neural network model, and the eye movement trajectory and the body position data are time-synchronized and correlated before or after being input into the neural network model. The neural network model is configured to: based on the time synchronization of the eye movement trajectory and the body position data, identify the changes in the subject's eye movement in each specific body position, and output the subject's vestibular function data.
2. The vestibular function testing data processing method according to claim 1, characterized in that, The changes in eye movement of the subject in each specific body position include: the nystagmus that occurs when the subject completes the change in body position and maintains the specific body position.
3. The vestibular function testing data processing method according to claim 1, characterized in that, The body position data is time-series data of the subject's head position, including at least one or both of the trajectory sequences of the Pitch axis and the Yaw axis.
4. The vestibular function testing data processing method according to claim 1, characterized in that, Also includes: The body position data is used to identify the specific body positions of the subject in the vestibular function test and the corresponding start and end times.
5. The vestibular function testing data processing method according to claim 4, characterized in that, Also includes: Based on the identified specific body position and its start and end times, the corresponding time period of the eye movement trajectory is marked in order to increase the attention weight of the neural network model to the marked time period.
6. The vestibular function testing data processing method according to claim 1, characterized in that, The neural network model adopts a multi-input channel structure, with different input channels receiving eye movement trajectories in different directions and body position data along different axes.
7. The vestibular function testing data processing method according to claim 1, characterized in that, The vestibular function data output by the neural network model includes structured results for "specific body position - eye movement changes", and the structured results include at least the specific body position, the presence or absence of nystagmus, and the direction of nystagmus.
8. The vestibular function testing data processing method according to claim 1, characterized in that, The data input to the neural network model also includes nystagmus feature parameters; The acquisition of the nystagmus feature parameters includes: Identify nystagmus information in eye movement trajectories in various directions to obtain corresponding nystagmus feature parameters; the nystagmus feature parameters include any one or more of the following: slow phase angular velocity of nystagmus, direction of nystagmus, nystagmus time, and trend of nystagmus change.
9. The vestibular function testing data processing method according to claim 1, characterized in that, The vestibular function data output by the neural network model is the nystagmus information of the subject in various specific body positions during the vestibular function test. The vestibular function detection data processing method further includes: performing data optimization processing on the identified nystagmus information under each specific body position, and outputting processed structured vestibular function text data.
10. The vestibular function testing data processing method according to claim 9, characterized in that, The data optimization process includes any one or more of the following: simplification of data representation, filtering of invalid information, visualization of features, and merging of valid information; wherein: The data representation simplification includes mapping the identified nystagmus latency and / or duration from numerical ranges to staged text descriptions; The invalid information filtering includes at least one of the following: Nystagmus events with intensity below a set threshold are collectively classified as non-significant nystagmus. Nystagmus that is maintained throughout the entire testing process is uniformly described as spontaneous nystagmus; In the nystagmus recognition results based on eye movement data in the horizontal, vertical and torsional directions, only nystagmus events with intensity reaching a set threshold are retained; The feature visualization process includes: For nystagmus identified from horizontal eye movement trajectories, a corresponding description of geotropic or geotropic nystagmus is generated based on the subject's specific body position and the relationship between the nystagmus direction and the ground. The effective information merging includes at least one of the following: Based on body position information, the nystagmus data of the same diagnostic element detected under different body positions are merged and the merged nystagmus feature data is output. Perform merging processing on nystagmus data of multiple diagnostic elements corresponding to the same diagnosis to generate merged diagnostic information; Based on the nystagmus intensity parameter, the nystagmus data corresponding to different diagnoses are sorted and the diagnostic data are output in order of intensity. Based on timestamp information, alignment and deduplication processing are performed on nystagmus data collected at different time points under the same body position to generate merged body position-nystagmus data; Based on the time series of the diagnosis and treatment process, the nystagmus feature data collected before, during and after treatment are correlated and compared to generate a merged result that includes differences in treatment effects.
11. An electronic device comprising a screen, a memory, one or more data processors, and one or more programs; wherein, The one or more programs are stored in the memory; characterized in that, when the one or more data processors execute the one or more programs, the electronic device implements the vestibular function detection data processing method as described in any one of claims 1 to 10.
12. A computer program product, characterized in that, When the computer program product is run on a computer, the computer performs the vestibular function detection data processing method as described in any one of claims 1 to 10.
13. A vestibular function testing system, characterized in that, include: An eye-tracking module, equipped with at least one camera, is used to capture eye-tracking videos of subjects during vestibular function tests; The body position acquisition submodule is used to acquire the body position data of the subject in the vestibular function test; One or more data processors, wherein the body position acquisition submodule and the eye movement imaging module are operatively coupled to the one or more data processors, the one or more data processors being configured to perform the vestibular function detection data processing method according to any one of claims 1-9.
14. The vestibular function detection system according to claim 13, characterized in that, The one or more data processors are located in the same or different devices; The device includes one or more of the following: an eye-tracking acquisition device, a body position acquisition device, a near-end data processing terminal, and a server.
15. The vestibular function detection system according to claim 13, characterized in that, The eye-tracking imaging module and the body position acquisition submodule may be located in the same or different devices; wherein: If the eye-tracking module and the body position acquisition submodule are integrated into the same device, then the same device is an eye-tracking acquisition device, which is a head-mounted device or an external camera device fixed near the subject by a bracket. If the eye-tracking imaging module and the body position acquisition submodule are respectively located in different devices, then the different devices include an eye-tracking acquisition device and a body position acquisition device.