Data processing system, method, device, medium and product

By collecting and fusing eye movement and gesture data through a data processing system and using lightweight neural networks for cognitive assessment, the problem of existing tests requiring professionals is solved, and the accuracy and convenience of the tests are improved.

CN120808440APending Publication Date: 2025-10-17BEIJING RUIKANGFU MEDICAL TECHNOLOGY CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510928942.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-07
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Existing neuropsychological tests such as MMSE and MoCA require well-trained testers to operate, with high barriers to entry and low efficiency. Test subjects may conceal information, resulting in inaccurate evaluation results.

Method used

A data processing system is used to display test tasks through a screen display device. By combining eye-tracking data acquisition and gesture data acquisition, features are generated and weighted and fused, and then input into a lightweight neural network model for cognitive assessment.

Benefits of technology

The test can be completed without professional personnel, which improves the accuracy and operability of the test and the accuracy of the test results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120808440A_ABST
    Figure CN120808440A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a data processing system and method, equipment, a medium and a product, and the system comprises a picture display device which displays a test task to a first object; the eye movement data acquisition device acquires first eye movement data; the gesture data acquisition device acquires gesture action data and a first image selected based on gesture actions; the data processing device generates a first feature according to the first eye movement data, generates a second feature according to the gesture action data and a result of whether the first image is matched with the second image, and determines a third feature according to the first feature and the second feature; and inputting the third feature into a data processing model to obtain a data processing result. According to the technical scheme, the problem of poor test accuracy at present is solved, the test result can be obtained by collecting the eye movement and gesture data in the test process and conducting feature extraction and model processing, the accuracy and operability of eye movement data and gesture data processing in the test process are improved, and the accuracy of the test result is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments of the present application relate to the technical field of computer technology, and particularly relate to a data processing system, method, device, medium and product. BACKGROUND

[0002] Currently, data processing to determine cognitive level, such as determining normal cognition, mild cognitive impairment and dementia, usually uses neuropsychological tests such as Mini-Mental State Examination (MMSE) and Montreal Cognitive Assessment (MoCA). Although these traditional neuropsychological tests are effective and reliable, they need to be implemented by trained testers, have high threshold, are inconvenient and inefficient to operate, are not conducive to large-scale promotion, and the tested person sometimes intentionally conceals, resulting in inaccurate evaluation results. SUMMARY

[0003] Embodiments of the present application provide a data processing system, method, device, medium and product, which can improve the accuracy and operability of eye movement data and gesture data processing in the test process, and improve the accuracy of test results.

[0004] In a first aspect, embodiments of the present application provide a data processing system, which comprises:

[0005] a picture display device, worn on the eye of a first object, configured to display a first picture containing each test task to the first object according to a preset task display order of the test task;

[0006] an eye movement data acquisition device, configured to acquire first eye movement data of the first object in the display process of the first picture;

[0007] a gesture data acquisition device, configured to acquire gesture action data of the first object in the display process of the first picture, and a first image selected based on the gesture action;

[0008] a data processing device, configured to generate a first feature according to the first eye movement data, generate a second feature according to the gesture action data and a result of whether the first image matches a second image, and perform attention weighted fusion on the first feature and the second feature to obtain a third feature, wherein the second image is a reference image of the gesture selection result of the test task;

[0009] The data processing device is further configured to input the third feature into a data processing model to obtain a data processing result of the first object.

[0010] In a second aspect, embodiments of the present application provide a data processing method, which comprises:

[0011] Displaying a first image corresponding to each test task to the first subject by means of an image display device according to a preset task display order of the test tasks, wherein the image display device is worn on the eyes of the first subject;

[0012] collecting, by an eye movement data collection device, first eye movement data of the first subject during the display of the first picture;

[0013] Collecting, by a gesture data collection device, gesture action data of the first object during presentation of the first screen, and a first image selected based on the gesture action;

[0014] A first feature is generated by a data processing device based on the first eye movement data, a second feature is generated based on the gesture action data and whether the first image matches the second image, and the first feature and the second feature are subjected to attention-weighted fusion to obtain a third feature. The third feature is input into a data processing model to obtain a data processing result of the first object, wherein the second image is a reference image of the gesture selection result.

[0015] In a third aspect, an embodiment of the present invention further provides a computer device, comprising:

[0016] one or more processors;

[0017] a memory for storing one or more programs;

[0018] When the one or more programs are executed by one or more processors, the one or more processors implement the data processing method provided by any embodiment of the present invention.

[0019] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the data processing method provided by any embodiment of the present invention.

[0020] In a fifth aspect, an embodiment of the present invention further provides a computer program product, comprising a computer program, which, when executed by a processor, implements the data processing method provided by any embodiment of the present invention.

[0021] The embodiments of the above invention have the following advantages or beneficial effects:

[0022] In an embodiment of the present invention, a screen display device is worn on the eyes of a first subject and is used to display a first screen containing each test task to the first subject in a preset task display order. An eye movement data acquisition device is used to collect first eye movement data of the first subject during the display of the first screen. A gesture data acquisition device is used to collect gesture data of the first subject during the display of the first screen, as well as a first image selected based on the gesture. A data processing device is used to generate a first feature based on the first eye movement data, a second feature based on the gesture data and whether the first image matches a second image, and an attention-weighted fusion of the first and second features to obtain a third feature, wherein the second image is a reference image for the gesture selection result of the test task. The data processing device is further used to input the third feature into a data processing model to obtain a data processing result for the first subject. The technical solution of the embodiment of the present invention solves the problem of poor test accuracy in current testing. Test results can be obtained by collecting eye movement and gesture data during the test process and performing feature extraction and model processing. This improves the accuracy and operability of processing eye movement and gesture data during the test process, thereby improving the accuracy of test results. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 is a structural diagram of a data processing system provided by an embodiment of the present invention;

[0024] Figure 2 is a flow chart of a data processing method provided by an embodiment of the present invention;

[0025] Figure 3 is a flow chart of a data processing method provided by an embodiment of the present invention;

[0026] Figure 4 It is a structural diagram of a computer device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0027] The present invention will be further described in detail below with reference to the accompanying drawings and examples. It will be understood that the specific embodiments described herein are intended only to illustrate the present invention and are not intended to limit the present invention. It should also be noted that, for ease of description, the accompanying drawings only illustrate portions relevant to the present invention, not all structures.

[0028] Figure 1 This is a structural diagram of a data processing system provided by an embodiment of the present invention. This embodiment is applicable to data processing scenarios, especially data processing scenarios for assisting judgment in the screening of cognitive impairment diseases in clinical cognitive assessment.

[0029] like Figure 1As shown, the data processing system comprises a picture display device 110, an eye movement data acquisition device 120, and a data processing device 130.

[0030] The picture display device 110 is worn on the eye of the first object, and is used to display a first picture containing a corresponding test task to the first object in a preset task display order of the test task. The eye movement data acquisition device 120 is used to acquire first eye movement data of the first object in the display process of the first picture. The gesture data acquisition device 130 is used to acquire gesture action data of the first object in the display process of the first picture, and a first image selected based on the gesture action. The data processing device 140 is used to generate a first feature according to the first eye movement data, generate a second feature according to the gesture action data and a result of whether the first image matches a second image, and perform attention weighted fusion on the first feature and the second feature to obtain a third feature, wherein the second image is a reference image of a gesture selection result of the test task. The data processing device 140 is also used to input the third feature into a data processing model to obtain a data processing result of the first object.

[0031] The first object can be a tested object, and specifically can be an object of a cognitive assessment test in cognitive assessment clinical screening of cognitive impairment diseases. The picture display device 110 can be a virtual reality (VR) head-mounted device, such as an external head-mounted device, a mobile terminal head-mounted device, and an integrated head-mounted device, worn on the eye of the first object, and used to display a first picture containing a corresponding test task to the first object in a preset task display order of the test task. The test task of the embodiment is a test task for cognitive assessment. Specifically, each test task is composed of text prompt content and image content containing at least one image. The text prompt content can be prompt content indicating that the user selects a second image in the image content. The second image is a reference image of a gesture selection result of the test task, that is, a correct image corresponding to a test passing state of the displayed test task. After the user selects the image by gesture clicking, the first picture of the next test task is displayed in the task display order until all the test tasks are tested and the test is completed.

[0032] The eye movement data acquisition device 120 can be composed of a near-infrared light source, a high-frame-rate camera, an image processing algorithm, and integrated hardware. It can be combined with the VR head-mounted device in a built-in or external manner. In the display process of the first picture, the near-infrared light is used to irradiate the eyes of the first object, the camera captures the corneal reflection and pupil changes in real time, and the algorithm analyzes the eye movement data such as the line of sight direction, the fixation point, and the blink frequency. The line of sight direction and the fixation point can reflect the attention allocation and information processing, and the blink frequency is related to the cognitive load or fatigue, so it can assist in assessing the cognitive state.

[0033] The gesture data acquisition device 130 can be an infrared camera deployed in the VR headset or environment to capture the gesture contour through binocular vision or structured light technology, and analyze the finger joint movement and gesture posture through computer vision algorithm, for collecting the gesture action data of the first object in the display process of the first picture, and the first image selected based on the gesture action.

[0034] The data processing device 140 is configured to generate first features according to the first eye movement data, and generate second features according to the gesture action data and the result of whether the first image matches the second image, i.e., whether the correct image is selected. The first features can be the gaze point position, gaze duration, saccade amplitude, blink frequency, pupil diameter change, and saccade speed, etc. The second features can include the result of whether the first image matches the second image, click accuracy, reaction time, and trajectory complexity, etc. The second features can include the result of whether the first image matches the second image, which can directly reflect the individual's understanding, judgment, and decision-making ability of information, the click accuracy reflects the fine motor control and attention, the reaction time is related to the information processing speed, and the trajectory complexity embodies the decision-making planning ability, so as to assist cognitive evaluation.

[0035] In the embodiment, the gradient boosting tree (XGBoost) can also be used to screen key features. The XGBoost can be used to process nonlinear relationships, sort the importance of features, and eliminate redundant or irrelevant features in the first features and the second features to reduce noise interference.

[0036] The first features and the second features are fused to obtain third features through attention weighting. Specifically, the first features and the second features are fused through attention weighting by using a preset weight coefficient. The preset weight coefficient can be set according to actual needs.

[0037] The features are integrated through a weighted summation formula to highlight key information, and the third features are obtained. For example, the first features can be a 15-dimensional feature vector including spatial, temporal, and physiological indicators, and the second features can be a 12-dimensional feature vector including kinematic and dynamic indicators. The first features and the second features are fused through attention weighting by using a preset weight coefficient to obtain 20-dimensional third features.

[0038] The third feature is input into a data processing model to obtain a data processing result of the first object. The data processing model can be a neural network model, for example, a lightweight neural network MobileNet. The lightweight neural network learns the mapping relationship between the input feature and the cognitive state during training, including the weight distribution between features, the nonlinear combination mode, and the time sequence dependence relationship. Finally, these mappings are solidified through network parameters such as convolution kernels and fully connected layer weights, so as to output a cognitive evaluation score for the third feature. According to the sum of the cognitive evaluation scores of the third feature of all test tasks, and a preset mapping relationship between scores and cognitive levels, it is determined that the cognitive evaluation result of the first object belongs to normal, mild cognitive impairment, or dementia.

[0039] The embodiment automatically collects eye movement data and gesture data during the test process for analysis to obtain a test result, without the need for professional physicians and professional knowledge. The first object can complete the test process and obtain the test result alone, thereby improving the operability and convenience of the test.

[0040] The technical scheme of the embodiment includes a picture display device worn on the eye of the first object, which is used to display a first picture corresponding to each test task to the first object according to a preset task display order of the test task; an eye movement data collection device used to collect first eye movement data of the first object during the display of the first picture; a gesture data collection device used to collect gesture action data of the first object during the display of the first picture, and a first image selected based on the gesture action; a data processing device used to generate a first feature based on the first eye movement data, generate a second feature based on the gesture action data and a result of whether the first image matches a second image, and perform attention weighted fusion on the first feature and the second feature to obtain a third feature, wherein the second image is a reference image of the gesture selection result of the test task; and the data processing device is further used to input the third feature into a data processing model to obtain a data processing result of the first object. The technical scheme of the embodiment solves the problem of poor test accuracy. The eye movement and gesture data during the test process are collected, and feature extraction and model processing are performed to obtain the test result, thereby improving the accuracy and operability of the eye movement data and gesture data processing during the test process, and improving the accuracy of the test result.

[0041] In an optional implementation, the data processing device 140 is specifically configured to:

[0042] According to the position of the second image of the gesture selection result of the test task, a feature position of the test task is determined; according to the feature position, second eye movement data of the first object gazing at the feature position within a preset time length is determined from the first eye movement data; and feature extraction is performed on the second eye movement data to obtain the first feature.

[0043] The positions of the second images corresponding to different test tasks can be different. According to the position of the second image corresponding to the gesture selection result of the test task, the center point or the preset feature point of the image region of the second image is taken as a feature position, that is, a focus point, second eye movement data in a preset time length of the first object gaze feature position is determined, feature extraction is performed on the second eye movement data, and first features are obtained.

[0044] In an optional implementation, the first features at least include at least one of a gaze time length, a saccade speed and a counter-saccade accuracy. The three indexes respectively provide quantitative basis for cognitive evaluation from three dimensions of attention depth, information processing efficiency and cognitive regulation.

[0045] In an optional implementation, the data processing model includes at least one sub-processing model, and the data processing apparatus 140 is further configured to:

[0046] The third features corresponding to each test task are input into the data processing model, the at least one sub-processing model is used for model processing to obtain a sub-processing result corresponding to the third features, and an average sub-processing result of all sub-processing results is determined. The sum value of the average sub-processing results corresponding to the third features of all test tasks is determined, and the data processing result of the first object is determined according to a mapping relationship between the sum value and a preset sum value and data processing result.

[0047] The sub-processing model can be MobileNet. The third features corresponding to each test task are converted into a tensor format according to the input requirements of MobileNet, for example, the feature sequence is encoded into a two-dimensional image matrix, the convolution calculation of the network is adapted, the input features are extracted and abstracted layer by layer through the deep separable convolution structure of MobileNet, and the parameter quantity and the calculation quantity are reduced. After the dimension is reduced through the global average pooling layer, the fully connected layer is connected to output the sub-processing result. After the processing of all tasks is completed, the average value of the sub-processing results of each task is calculated, and then the sum value of all task average results is obtained. Finally, according to the preset mapping relationship between the sum value and the data processing result, the final data processing result of the first object is determined, and efficient and accurate cognitive evaluation is realized.

[0048] In an optional implementation, the data processing apparatus 140 is further configured to:

[0049] The first sample features of the first sample are determined, the data processing initial model is trained according to the first sample features, the loss function value is calculated according to the first model output result and the first sample label, and the model parameter is adjusted, wherein the first sample includes sample eye movement data, sample gesture action data and third images selected based on gesture actions; when the loss function value is less than a preset convergence threshold, the data processing model is obtained.

[0050] The first sample feature is obtained by attention-weighted fusion of sample eye movement features corresponding to sample eye movement data, sample gesture action data, and sample gesture features corresponding to the third image selected based on the gesture action.

[0051] In an example, in the training process of the data processing model, first, data enhancement is performed on the input sample features. Gaussian noise is added to the eye movement trajectory to simulate errors, and the gesture time axis is randomly stretched and contracted to adapt to different action rhythms, thereby expanding the data set. Then, the processed data is input into the data processing initial model. During training, L1 / L2 mixed regularization constraint parameters are used, and neurons are randomly inactivated through a Dropout rate of 0.3 to prevent overfitting. The optimizer selects Adam, updates the parameters at a learning rate of 0.001, and uses the early stopping strategy to monitor the validation set loss. If there is no improvement for multiple consecutive rounds, the training is terminated to ensure the generalization ability of the model.

[0052] In an alternative embodiment, the data processing apparatus 140 is further configured to:

[0053] determine a second sample feature of a second sample, input the second sample feature into the data processing model to obtain a second model output result and a confidence, determine a third sample feature in the second sample feature, the confidence of which is less than a preset confidence threshold, and train the data processing model through the third sample feature, calculate a loss function value according to a third model output result and a third sample label, and adjust the model parameters, and obtain an optimized data processing model when the loss function value is less than a preset convergence threshold.

[0054] The second sample is an unlabeled sample. After the third sample feature corresponding to the sample with a smaller confidence in the second sample is labeled by a doctor, the data processing model is trained according to the third sample feature and the label information labeled by the doctor, and the model parameters are adjusted until the loss function value is less than the preset convergence threshold, thereby obtaining the optimized data processing model. This embodiment actively learns and efficiently selects key sample labels, reduces the data volume requirement, accelerates the improvement of model performance, and adapts to complex scenarios.

[0055] In an alternative embodiment, different sub-processing models can be used for different test questions. Each sub-screening model is configured to predict a third feature corresponding to a test task to obtain a sub-prediction result. A fusion output layer is configured to determine a prediction category according to multiple sub-prediction results. The prediction category is divided into normal, mild, and dementia.

[0056] For different categories of test tasks, the generated first features and second features can also be different. For a spatial navigation test task corresponding to the spatial cognition domain, the first features can include path efficiency and back view times, and the second features can include turning speed and collision frequency; for a memory test corresponding to the episodic memory cognition domain, the first features can include target fixation preference and scanning mode, and the second features can include dragging accuracy and sequence accuracy; for an attention test task corresponding to the executive function cognition domain, the first features can include interference item fixation suppression rate, and the second features can include click reaction time and error recovery time; for a language understanding test task corresponding to the language ability cognition domain, the first features can include text scanning speed and pupil dilation change, and the second features can include selection reaction time and option dwell time. The embodiment can precisely match the task characteristics, improve the processing efficiency and result accuracy, and enhance the model pertinence and generalization ability by using different models and key features for different tasks.

[0057] Figure 2 A flowchart of a data processing method provided by the embodiment of the application, which can be applied to the scene of data processing.

[0058] As shown in Figure 2 The data processing method of the embodiment includes the following steps:

[0059] S210, a picture display device displays a first picture containing a test task corresponding to each test task to a first object in a preset task display order of the test task, and the picture display device is worn on the eye of the first object.

[0060] S220, an eye movement data acquisition device acquires first eye movement data of the first object in the display process of the first picture.

[0061] S230, a gesture data acquisition device acquires gesture action data of the first object in the display process of the first picture, and a first image selected based on the gesture action.

[0062] S240, a data processing device generates a first feature according to the first eye movement data, generates a second feature according to the gesture action data and whether the first image matches a second image, and performs attention weighted fusion on the first feature and the second feature to obtain a third feature, inputs the third feature into a data processing model, and obtains a data processing result of the first object.

[0063] The second image is a reference image of the gesture selection result.

[0064] The technical scheme of the embodiment, through the picture display device, the first object is displayed in the first picture corresponding to each test task according to the preset test task display order, and the picture display device is worn on the eye of the first object; the first eye movement data of the first object in the display process of the first picture is collected through the eye movement data collection device; the gesture action data of the first object in the display process of the first picture and the first image selected based on the gesture action are collected through the gesture data collection device; the first feature is generated according to the first eye movement data through the data processing device, the second feature is generated according to the gesture action data and the result of whether the first image matches the second image, and the first feature and the second feature are weighted and fused to obtain the third feature, and the third feature is input into the data processing model to obtain the data processing result of the first object; wherein the second image is a reference image of the gesture selection result. The technical scheme of the embodiment solves the problem of poor test accuracy, can obtain the test result by collecting the eye movement and gesture data in the test process and performing feature extraction and model processing, improves the accuracy and operability of the eye movement data and gesture data processing in the test process, and improves the test result accuracy.

[0065] Figure 3 The flowchart of the data processing method provided by the embodiment is the same as the data processing method in the above-mentioned embodiment, and the process of determining the first feature is further described. As shown in Figure 3 The data processing method of the embodiment comprises the following steps:

[0066] S310, the first object is displayed in the first picture corresponding to each test task according to the preset test task display order through the picture display device, and the picture display device is worn on the eye of the first object.

[0067] S320, the first eye movement data of the first object in the display process of the first picture is collected through the eye movement data collection device.

[0068] In an optional embodiment, the first feature at least includes at least one of the fixation duration, the saccade speed and the anti-saccade accuracy.

[0069] S330, the gesture action data of the first object in the display process of the first picture and the first image selected based on the gesture action are collected through the gesture data collection device.

[0070] S340, the feature position of the test task is determined according to the position of the second image of the gesture selection result of the test task through the data processing device.

[0071] S350, determining, by the data processing apparatus, second eye movement data in a preset time length of the first object gazing at the feature position in the first eye movement data according to the feature position.

[0072] S360, performing feature extraction on the second eye movement data by the data processing apparatus to obtain a first feature.

[0073] S370, generating, by the data processing apparatus, a second feature according to the gesture action data and a result of whether the first image matches the second image, and performing attention weighted fusion on the first feature and the second feature to obtain a third feature, and inputting the third feature into the data processing model to obtain a data processing result of the first object.

[0074] The second image is a reference image of the gesture selection result.

[0075] In an optional embodiment, the data processing model includes at least one sub-processing model, and the third feature corresponding to each test task is input into the data processing model, the model processing is performed by the at least one sub-processing model to obtain a sub-processing result corresponding to the third feature, and an average sub-processing result of all sub-processing results is determined; a sum value of the average sub-processing result of the third feature corresponding to all test tasks is determined, and the data processing result of the first object is determined according to a mapping relationship between the sum value and a preset sum value and the data processing result.

[0076] In an optional embodiment, a first sample feature of a first sample is determined by the data processing apparatus, the data processing initial model is trained according to the first sample feature, a loss function value is calculated according to a first model output result and a first sample label, and model parameter adjustment is performed, wherein the first sample contains sample eye movement data, sample gesture action data and a third image selected based on the gesture action.

[0077] When the loss function value is less than a preset convergence threshold value, the data processing model is obtained.

[0078] In an optional embodiment, a second sample feature of a second sample is determined by the data processing apparatus, the second sample feature is input into the data processing model to obtain a second model output result and a confidence, third sample features with a confidence less than a preset confidence threshold value in the second sample feature are determined, the data processing model is trained by the third sample features, a loss function value is calculated according to a third model output result and a third sample label, and model parameter adjustment is performed, and when the loss function value is less than a preset convergence threshold value, an optimized data processing model is obtained.

[0079] The technical scheme of the embodiment, through the picture display device, the first object is displayed in the first picture containing each test task according to the preset task display order of the test task, and the picture display device is worn on the eye of the first object; The first eye movement data of the first object in the display process of the first picture is collected through the eye movement data collection device; The gesture action data of the first object in the display process of the first picture and the first image selected based on the gesture action are collected through the gesture data collection device; The feature position of the test task is determined according to the position of the second image of the gesture selection result of the test task through the data processing device; The second eye movement data in the preset time length of the first object staring at the feature position in the first eye movement data is determined according to the feature position; The first feature is obtained by extracting the second eye movement data; The second feature is generated according to the gesture action data and the result of whether the first image matches the second image, and the first feature and the second feature are weighted and fused to obtain the third feature, and the third feature is input into the data processing model to obtain the data processing result of the first object; Wherein, the second image is the reference image of the gesture selection result. The technical scheme of the embodiment solves the problem of poor data processing effect at present, can determine the driving parameter through the motion data collection, and display the picture of driving the three-dimensional model motion to the user, improve the data processing effect of the user's limbs, and determine the motion plan through the plan determination module, improve the accuracy, individualization degree and efficiency of data processing. The technical scheme of the embodiment solves the problem of poor test accuracy at present, can obtain the test result by collecting the eye movement and gesture data in the test process and performing feature extraction and model processing, improve the accuracy and operability of the eye movement data and gesture data processing in the test process, improve the test result accuracy, and generate features by eye movement data of the staring feature position, which can focus on key features, further improve the test result accuracy.

[0080] Figure 4 A structural schematic diagram of a computer device provided by the embodiment of the present application is provided. Figure 4 A block diagram of an exemplary computer device 12 suitable for implementing embodiments of the present application is shown. Figure 4 The displayed computer device 12 is only an example, and should not bring any limitation to the function and use range of the embodiment of the present application. The computer device 12 can be any terminal device with computing capability, such as a smart controller, a server, a mobile phone and other terminal devices.

[0081] As shown in Figure 4 The computer device 12 is in the form of a general-purpose computing device. The components of the computer device 12 can include but are not limited to one or more processors or processing units 16, system memory 28, and bus 18 connecting different system components, including system memory 28 and processing unit 16.

[0082] Bus 18 represents one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, a graphics acceleration bus, a processor or local bus using any of a variety of bus architectures. By way of example, these architectures include Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MCA) bus, Enhanced ISA (EISA) bus, Video Electronics Standards Association (VESA) local bus, and Peripheral Component Interconnect (PCI) bus.

[0083] Computer device 12 typically includes a variety of computer system readable media. Such media can be any available media that is located either internally or externally to computer device 12, including both volatile and nonvolatile media, removable and non-removable media.

[0084] System memory 28 can include computer system readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache 32. Computer device 12 can further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, storage system 34 can be provided for reading from and writing to a non-removable, non-volatile magnetic media (e.g., a "hard drive"). Figure 4 not shown, a magnetic hard disk drive for reading from and writing to non-removable, non-volatile magnetic media (e.g., a "hard drive"). Although not specifically shown, such computer system can further include other removable / non-removable, volatile / non-volatile computer system storage media including, but not limited to, a magnetic floppy disk drive for reading from and writing to a removable, non-volatile magnetic floppy disk, and an optical disk drive for reading from or writing to a removable, non-volatile optical disk such as a CD ROM or other optical media. Figure 4 In such a scenario, each of the drives can be connected to the bus 18 by one or more data media interfaces. The system memory 28 can include at least one program product having a set (e.g., at least one) of program modules that are configured to carry out the functions of embodiments of the application.

[0085] Program / utility 40, having a set (at least one) of program modules 42, can be stored in system memory 28 by way of example, such as an operating system, one or more application programs, other program modules, and program data, each of which

[0086] Computer device 12 can also communicate with one or more external devices 14 such as a keyboard, a pointing device, a display 24, etc.; one or more devices that enable a user to interact with computer device 12; and / or any devices (e.g., network card, modem, etc.) that enable computer device 12 to communicate with one or more other computing devices. Such communication can occur via Input / Output (I / O) interface(s) 22. Still yet, computer device 12 can communicate with one or more networks, such as a local area network (LAN), a general wide area network (WAN), and / or a public network such as the Internet, via network adapter 20. As depicted, network adapter 20 communicates with the other components of computer device 12 via bus 18. It should be appreciated that although not shown, other hardware and / or software modules could be used in conjunction with computer device 12. Examples, include, but are not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, Artificial Intelligence (AI) systems, tape drives, and data archival storage systems, etc. Figure 4

[0087] Processing unit 16 performs various function applications and data processing by running programs stored in system memory 28, such as implementing the data processing method provided by the embodiments of the present application, which includes:

[0088] The picture display device displays the first picture containing each test task to the first object according to the preset task display order of the test task, and the picture display device is worn on the eye of the first object;

[0089] The eye movement data acquisition device acquires the first eye movement data of the first object in the display process of the first picture;

[0090] The gesture data acquisition device acquires the gesture action data of the first object in the display process of the first picture, and the first image selected based on the gesture action;

[0091] The data processing device generates a first feature according to the first eye movement data, generates a second feature according to the gesture action data and the result of whether the first image matches the second image, and performs attention weighted fusion on the first feature and the second feature to obtain a third feature, and inputs the third feature into the data processing model to obtain the data processing result of the first object, wherein the second image is a reference image of the gesture selection result.

[0092] The embodiments of the present application also provide a computer readable storage medium, which stores a computer program, and the program is executed by a processor to implement the data processing method provided by any of the embodiments of the present application, and the method includes:

[0093] ​The picture display device displays the first picture corresponding to each test task to the first object according to the preset task display sequence of the test task, and the picture display device is worn on the eye part of the first object;

[0094] The eye movement data acquisition device acquires the first eye movement data of the first object in the display process of the first picture;

[0095] The gesture data acquisition device acquires the gesture action data of the first object in the display process of the first picture, and the first image selected based on the gesture action;

[0096] The data processing device generates the first feature according to the first eye movement data, generates the second feature according to the gesture action data and the result of whether the first image matches the second image, and performs attention weighted fusion on the first feature and the second feature to obtain the third feature, and inputs the third feature into the data processing model to obtain the data processing result of the first object, wherein the second image is a reference image of the gesture selection result.

[0097] The computer storage medium of the embodiment of the application can adopt any combination of one or more computer readable media. The computer readable medium can be a computer readable signal medium or a computer readable storage medium. The computer readable storage medium may, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or component, or any combination of the above. More specific examples (non-exhaustive list) of computer readable storage media include: electrical connections with one or more conductors, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this document, the computer readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device or component.

[0098] The computer readable signal medium can include a data signal propagating in a baseband or as part of a carrier wave, carrying computer readable program code. Such a propagating data signal can take various forms, including but not limited to electromagnetic signals, optical signals or any suitable combination of the above. The computer readable signal medium can also be any computer readable medium other than the computer readable storage medium, which can send, propagate or transmit a program for use by or in conjunction with an instruction execution system, device or component.

[0099] The computer readable media on which the program code can be carried can be any media suitable for carrying computer program code, including but not limited to a wireless, a wired, optical, cable, RF, etc. or any suitable combination of the foregoing.

[0100] Computer program code for carrying out operations of the present application can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, Python, C++, or the like, and conventional procedural programming languages, such as the "C" programming language, or the like. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).

[0101] The embodiments of the present application also provide a computer program product, comprising a computer program which, when executed by a processor, implements the data processing method provided by any of the embodiments of the present application.

[0102] The computer program product in the implementation can be written in one or more programming languages or combinations thereof for executing the operations of the present application, including an object oriented programming language such as Java, Smalltalk, Python, C++, or the like, and conventional procedural programming languages, such as the "C" programming language, or the like. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).

[0103] Those skilled in the art should understand that the modules or steps of the present application described above can be realized by general computing devices, which can be centralized on a single computing device or distributed on a network composed of multiple computing devices, and can be realized by computer executable program codes, which can be stored in a storage device and executed by a computing device, or realized by individual integrated circuit modules, or realized by multiple modules or steps as a single integrated circuit module. Thus, the present application is not limited to any specific combination of hardware and software.

[0104] It is noted that the above merely describes the preferred embodiments of the present application and the principles of the applied technology. Those skilled in the art will understand that the present application is not limited to the specific embodiments described herein, and that various obvious changes, re-adjustments and substitutions can be made without departing from the scope of the present application. Therefore, although the present application has been described in detail by the above embodiments, the present application is not limited to the above embodiments, and can include more other equivalent embodiments without departing from the concept of the present application, and the scope of the present application is determined by the appended claims.

Claims

1. A data processing system, characterized in that: include: A picture display device, worn on the eyes of the first subject, for displaying a first picture corresponding to each of the test tasks to the first subject in a preset task display order of the test tasks; an eye movement data collection device, configured to collect first eye movement data of the first subject during the display of the first picture; a gesture data collection device for collecting gesture action data of the first object during the display of the first screen, and a first image selected based on the gesture action; a data processing device, configured to generate a first feature based on the first eye movement data, generate a second feature based on the gesture action data and a result of whether the first image matches a second image, and perform attention-weighted fusion on the first and second features to obtain a third feature, wherein the second image is a reference image for the gesture selection result of the test task; The data processing device is further configured to input the third feature into a data processing model to obtain a data processing result of the first object.

2. The system according to claim 1, wherein: The data processing device is specifically used for: determining a characteristic position of the test task according to the position of the second image of the gesture selection result of the test task; determining, according to the characteristic position, second eye movement data in the first eye movement data, which is a data within a preset time period when the first subject gazes at the characteristic position; Feature extraction is performed on the second eye movement data to obtain a first feature.

3. The system according to claim 2, characterized in that The first feature includes at least one of gaze duration, scanning speed and anti-saccade accuracy.

4. The system according to claim 1, wherein: The data processing model includes at least one sub-processing model, and the data processing device is further configured to: Inputting the third feature corresponding to each of the test tasks into a data processing model, performing model processing through at least one sub-processing model to obtain a sub-processing result corresponding to the third feature, and determining an average sub-processing result of all the sub-processing results; Determine the sum of the average sub-processing results of the third feature corresponding to all the test tasks, and determine the data processing result of the first object based on the sum and a mapping relationship between the sum and a preset sum and data processing result.

5. The system according to claim 1, wherein: The data processing device is further configured to: Determining a first sample feature of a first sample, training an initial data processing model based on the first sample feature, calculating a loss function value based on an output result of the first model and a first sample label, and adjusting model parameters, wherein the first sample includes sample eye movement data, sample gesture action data, and a third image selected based on the gesture action; When the loss function value is less than a preset convergence threshold, a data processing model is obtained.

6. The system according to claim 1, wherein: The data processing device is further configured to: Determining a second sample feature of the second sample, inputting the second sample feature into the data processing model, and obtaining a second model output result and a confidence level; Determining, among the second sample features, a third sample feature whose confidence level is less than a preset confidence threshold; The data processing model is trained using the third sample feature, the loss function value is calculated based on the third model output result and the third sample label, and the model parameters are adjusted. When the loss function value is less than a preset convergence threshold, an optimized data processing model is obtained.

7. A data processing method, characterized in that: include: Displaying a first screen corresponding to each test task to the first subject by means of a screen display device according to a preset task display order of the test tasks, wherein the screen display device is worn on the eyes of the first subject; collecting, by an eye movement data collection device, first eye movement data of the first subject during the display of the first picture; Collecting, by a gesture data collection device, gesture action data of the first object during the display of the first screen, and a first image selected based on the gesture action; A first feature is generated according to the first eye movement data by a data processing device, a second feature is generated according to the gesture action data and whether the first image matches the second image, and the first feature and the second feature are subjected to attention-weighted fusion to obtain a third feature. The third feature is input into a data processing model to obtain a data processing result of the first object, wherein the second image is a reference image of the gesture selection result.

8. A computer device, characterized in that: The computer device comprises: one or more processors; a memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the data processing method according to claim 7.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the data processing method according to claim 7 is implemented.

10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the computer program implements the data processing method according to claim 7.