Gaze point guided child visual cognition classification method, system and equipment
By using gaze point data and prompt learning methods in deep learning models, the shortcomings of traditional machine learning models in capturing nonlinear relationships and long-term dependencies of eye movement data are solved, and higher classification accuracy and generalization capabilities are achieved.
Patent Information
- Application Number
- CN202510645199.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-06-17
AI Technical Summary
Traditional machine learning models are difficult to fully capture the nonlinear relationships and long-term dependencies in eye movement data, resulting in insufficient generalization capabilities of the model and it is difficult to maintain stable performance in different populations or scenarios.
Through the gaze point data of autistic children, a prompt learning method is used to guide the deep learning model to focus on distinguishing key image areas of the autistic population, thereby achieving accurate classification.
The model's perception of core features is strengthened, classification accuracy is improved, interference with irrelevant information is effectively suppressed, nonlinear relationships and long-term dependencies in eye movement data are fully captured, and the model's interpretability and generalization ability is enhanced.
Smart Images

Figure CN120164049A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of eye movement data processing, and particularly relates to a method, system and device for classifying children's visual cognition guided by fixation points. Background Art
[0002] In recent years, with the development of artificial intelligence, methods based on eye tracking and machine learning have gradually become an emerging research direction for the early identification and evaluation of autism. Eye tracking technology reveals abnormalities in attention allocation and social information processing by recording the eye movement patterns of children when watching social scenes or specific visual stimuli; machine learning constructs prediction models by analyzing a large amount of behavioral data to help identify early characteristics of autism. In addition, researchers have found that autistic children show obvious differences in eye movement data, mainly reflected in aspects such as fixation time, fixation trajectory, and selection of regions of interest. Compared with typically developing children, autistic children tend to have shorter fixation times on social stimuli (such as human faces and eye regions), while paying more attention to non-social objects (such as backgrounds or inanimate objects). In addition, their eye movement trajectories are more scattered, lacking focused attention on key information regions, and it is difficult to follow others' gazes or notice social interactions when watching dynamic social videos. In addition, autistic children also have weaker fixation responses in situations such as name calling and social competition. These characteristics can provide objective biomarkers for the early screening of autism, and can also combine the attention patterns and neural activity characteristics of children under different experimental paradigms to provide a more accurate scientific basis for the early identification and intervention of autism.
[0003] Although machine learning shows great potential in extracting eye movement features and identifying autism, it still faces some significant challenges and deficiencies in practical applications. First, the performance of the model highly depends on high-quality large-scale datasets. However, the acquisition of autism-related data is often restricted by various factors, such as small sample sizes, large individual differences, and subjectivity in data annotation. These restrictions lead to deficiencies in the generalization ability of the trained model, making it difficult to maintain stable performance in different populations or scenarios. In addition, as a highly heterogeneous neurodevelopmental disorder, the diverse manifestations of autism further increase the complexity of data collection and model training. Second, eye movement data has strong time series characteristics and individual differences. The eye movement trajectory not only reflects the attention allocation pattern of the subject in a specific task, but also contains rich dynamic information, such as the transfer speed of fixation points, dwell time, and saccade behavior. The complex temporal patterns pose high requirements on traditional machine learning models, and many traditional machine learning methods are difficult to fully capture the non-linear relationships and long-term dependencies in eye movement data. Summary of the Invention
[0004] The embodiments of this application provide a method, system, and device for classifying children's visual cognition guided by fixation points, which can solve the problem that traditional machine learning models are difficult to fully capture the non-linear relationships and long-term dependencies in eye movement data, resulting in insufficient generalization ability of the models.
[0005] In a first aspect, the embodiments of this application provide a method for classifying children's visual cognition guided by fixation points, including: Obtain the fixation point map data of the target child; Use a children's visual cognition classification model to process the fixation point map data and the preset visual stimulus image data to obtain the visual cognition classification result of the target child; Wherein, the children's visual cognition classification model is a machine learning model pre-trained by prompt learning using the sample fixation point map data and sample visual stimulus image data of sample children.
[0006] In the above technical solutions of the embodiments of this application, there are at least the following technical effects: The method for classifying children's visual cognition guided by fixation points provided by the embodiments of this application uses the fixation point data of children with autism and adopts the prompt learning method to guide the deep learning model to focus on the key image regions for distinguishing the autism group, thereby achieving accurate classification, strengthening the model's perception ability of core features, improving the classification accuracy, effectively suppressing the interference of irrelevant information, fully capturing the non-linear relationships and long-term dependencies in eye movement data, and enhancing the interpretability and generalization ability of the model.
[0007] In a second aspect, the embodiments of this application provide a system for classifying children's visual cognition guided by fixation points, including: An acquisition unit for obtaining the fixation point map data of the target child; A classification unit for using a children's visual cognition classification model to process the fixation point map data and the preset visual stimulus image data to obtain the visual cognition classification result of the target child; Wherein, the children's visual cognition classification model is a machine learning model pre-trained by prompt learning using the sample fixation point map data and sample visual stimulus image data of sample children.
[0008] In a third aspect, the embodiments of this application provide an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the method described in any one of the above aspects is implemented.
[0009] In a fourth aspect, the embodiments of this application provide a computer program product, which, when running on an electronic device, causes the electronic device to execute the method described in any one of the above aspects.
[0010] It is understandable that for the beneficial effects of the second to fourth aspects above, reference may be made to the relevant descriptions in the above aspects, which will not be elaborated herein. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0012] Figure 1 is a schematic flowchart of a gaze point-guided visual cognitive classification method for children provided by an embodiment of the present application; Figure 2 is a schematic operation diagram of a gaze point-guided visual cognitive classification method for children provided by an embodiment of the present application; Figure 3 is a schematic structural diagram of a gaze point-guided visual cognitive classification system for children provided by an embodiment of the present application; Figure 4 is a schematic structural diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0013] In the following description, specific details such as specific system structures and technologies are proposed for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.
[0014] It should be understood that when used in the specification and the appended claims of the present application, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0015] It should also be understood that the term " / and" as used in the specification and the appended claims of the present application refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0016] As used in the specification and claims of this application, the term "if" may be construed contextually as "when" or "once" or "in response to determining" or "in response to detecting". Similarly, the phrases "if determined" or "if the described condition or event is detected" may be construed contextually to mean "once determined" or "in response to determining" or "once the described condition or event is detected" or "in response to detecting the described condition or event".
[0017] In addition, in the description of the specification and claims of this application, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be construed as indicating or implying relative importance.
[0018] The reference to "one embodiment" or "some embodiments" or the like described in the specification of this application means that a specific feature, structure, or characteristic described in connection with the embodiment is included in one or more embodiments of this application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized. The terms "comprising", "including", "having", and their variants all mean "including but not limited to", unless otherwise specifically emphasized.
[0019] Although machine learning has shown great potential in extracting eye movement features and identifying autism, it still faces some significant challenges and deficiencies in practical applications. First, the performance of the model highly depends on high-quality large-scale datasets. However, the acquisition of autism-related data is often restricted by various factors, such as small sample sizes, large individual differences, and subjectivity in data annotation. These restrictions lead to deficiencies in the generalization ability of the trained model and make it difficult to maintain stable performance in different populations or scenarios. In addition, as a highly heterogeneous neurodevelopmental disorder, the diverse manifestations of autism further increase the complexity of data collection and model training. Second, eye movement data has strong time series characteristics and individual differences. The eye movement trajectory not only reflects the attention allocation pattern of the subject in a specific task but also contains rich dynamic information, such as the transfer speed of fixation points, dwell time, and saccade behavior. The complex temporal patterns pose high requirements on traditional machine learning models, and many traditional machine learning methods are difficult to fully capture the non-linear relationships and long-term dependencies in eye movement data.
[0020] To solve the above problems, an embodiment of the present application provides a gaze-point-guided visual cognitive classification method for children. In this method, based on the gaze-point data of children with autism, a prompting learning method is adopted to guide the deep learning model to focus on the key image regions for distinguishing the autism group, so as to achieve accurate classification, strengthen the model's perception ability of core features, improve the classification accuracy, effectively suppress the interference of irrelevant information, fully capture the non-linear relationship and long-term dependence in the eye movement data, and enhance the interpretability and generalization ability of the model.
[0021] The gaze-point-guided visual cognitive classification method provided by the embodiment of the present application can be applied to an electronic device. At this time, the electronic device is the execution subject of the gaze-point-guided visual cognitive classification method provided by the embodiment of the present application. The embodiment of the present application does not impose any restrictions on the specific type of the electronic device.
[0022] For example, the electronic device can be a mobile phone, a tablet computer, a notebook computer, an ultra-mobile personal computer (UMPC), a netbook, a desktop computer, a smart large screen, a smart TV, a handheld device with wireless communication function, a computing device or other processing devices connected to a wireless modem, a vehicle-mounted device, a vehicle networking terminal, a computer, a laptop computer, a communication device, a computing device, etc.
[0023] To better understand the gaze-point-guided visual cognitive classification method provided by the embodiment of the present application, the following provides an exemplary introduction to the specific implementation process of the gaze-point-guided visual cognitive classification method provided by the embodiment of the present application.
[0024] Figure 1 The schematic flowchart of the gaze-point-guided visual cognitive classification method provided by the embodiment of the present application is shown. Figure 2 The operation flowchart of the gaze-point-guided visual cognitive classification method provided by the embodiment of the present application is shown. The gaze-point-guided visual cognitive classification method includes: S100, obtaining the gaze-point map data of the target child.
[0025] It can be understood that the fixation point map is a graph that visually presents the fixation point-related data in the eye movement data of children when they are viewing visual stimulus images. The eye movement trajectory of the target child when viewing the preset visual stimulus image data can be recorded in real time by using a Tobii eye tracker, an iMotions eye tracker, an SRResearch eye tracker, etc., and the eye movement trajectory of the child can be monitored in real time. The collected eye movement data is divided according to the visual stimulus image data to ensure that the fixation data corresponding to each visual stimulus image data is stored and processed independently. The effective fixation point coordinates and the corresponding fixation duration can be extracted from the original eye movement data, and the abnormal data caused by instantaneous saccades or interference factors can be removed to improve the data quality. The fixation points can be smoothed using a large kernel Gaussian filter, and the discrete fixation points can be converted into a continuous fixation point map to visually present the attention distribution pattern of the child on different images.
[0026] S200. Use the child visual cognitive classification model to process the fixation point map data and the preset visual stimulus image data to obtain the visual cognitive classification result of the target child; among them, the child visual cognitive classification model is a machine learning model pre-trained by prompt learning using the sample fixation point map data and sample visual stimulus image data of sample children.
[0027] It can be understood that the obtained fixation map data of the target child can be associated one by one with the preset visual stimulus image data according to the image number. Each set of corresponding fixation map data and visual stimulus image data is used as input features and input into the child visual cognitive classification model. The fixation map, as a direct reflection of the subject's eye movement characteristics, can provide additional visual cues for the model, helping the model make more accurate judgments during the recognition process. Especially when distinguishing different groups (such as the autism group and the healthy control group), it can provide more effective support. The visual stimulus image data can be fused with the corresponding fixation map of the subject to obtain a model input image with local discriminability. By combining the fixation map with the original visual stimulus image and using the child's fixation data to guide the deep learning model to focus on the discriminative regions in the image, it can highlight the regions that the subject pays attention to during viewing, not only retaining the overall semantic information of the image but also effectively capturing the local features related to the fixation points, thereby improving the attention of the classification model to the key regions. The child visual cognitive classification model is a machine learning model pre-trained using the sample fixation map data and sample visual stimulus image data of sample children through prompt learning, such as convolutional neural networks, recurrent neural networks (RNN) and their variants (such as LSTM, GRU). The multi-layer neural network structure inside the child visual cognitive classification model processes and extracts features from the input data layer by layer. In the convolutional layer, the model extracts the local features of the fixation map and the visual stimulus image, capturing feature information at different scales in the image through convolutional kernels of different sizes. For example, smaller convolutional kernels can capture the detailed features in the image, while larger convolutional kernels can obtain more macroscopic structural features. In the pooling layer, the model performs downsampling operations on the feature maps output by the convolutional layer, reducing the data dimension while retaining important feature information to improve the computational efficiency and generalization ability of the model. After multiple convolutional and pooling operations, the model sends the extracted features to the fully connected layer for integration and classification. The neurons in the fully connected layer perform weighted summation on the input features and perform non-linear transformation through the activation function, outputting the visual cognitive category scores of the target child under the current visual stimulus image. According to the output result of the model, the scores are converted into probability values by setting a classification threshold or using the Softmax function to determine the visual cognitive classification result of the target child for each visual stimulus image. For example, if among the probability values output by the Softmax function, the probability corresponding to a certain category is the largest and exceeds the preset threshold (such as 0.5), it is determined that the visual cognition of the target child for this image belongs to this category. Finally, the visual cognitive classification results corresponding to all visual stimulus images are summarized to obtain the complete visual cognitive classification result of the target child.
[0028] In a possible implementation, before using the child visual cognitive classification model to process the fixation point map data and visual stimulus image data to obtain the visual cognitive classification result of the target child, the method further includes: S300, obtaining the sample fixation point map data of the sample child.
[0029] It can be understood that by using a Tobii eye tracking device, an iMotions eye tracking device, an SRResearch eye tracker, etc., the eye movement trajectory of the sample child when viewing the preset visual stimulus image data can be recorded in real time, and key data such as the fixation point and fixation duration can be accurately captured, so as to ensure the integrity and accuracy of the data and provide a reliable basis for subsequent analysis. During the acquisition process, the sample child views multiple pieces of visual stimulus image data preset by the active-guided free-view paradigm in sequence, each picture is shown for 3 seconds, and then the line of sight is calibrated through a central cross for 1.2 seconds to reduce eye movement data drift and ensure experimental consistency, systematically analyzing the attention preferences of children with autism when facing social-related or competitive visual information, and further revealing their abnormal performance in eye movement characteristics, providing a basis for subsequent automated screening research. The eye movement trajectory of the child can be monitored in real time, and the collected eye movement data can be divided according to the visual stimulus image data to ensure that the fixation data corresponding to each piece of visual stimulus image data is stored and processed independently. Effective fixation point coordinates and corresponding fixation durations can be extracted from the original eye movement data, and abnormal data caused by instantaneous saccades or interference factors can be removed to improve the data quality. The fixation points can be smoothed using a Gaussian filter with a kernel size of 245×245 and a standard deviation of 35, converting the discrete fixation points into continuous sample fixation point map data to visually present the attention distribution pattern of the sample child on different images. The formula is as follows: where represents the fixation point data on the j-th image of the i-th subject, represents a Gaussian smoothing with a kernel size of 245*245 and a standard deviation of 35, represents the smoothed fixation point data, and the generated fixation point map will be stored to provide high-quality visual input for subsequent attention feature analysis and deep learning modeling.
[0030] S400, determining the visual cognitive classification label of the sample child.
[0031] It can be understood that professional psychological and medical assessment tools, such as the Autism Diagnostic Observation Schedule (ADOS), the Children's Developmental Assessment Scale, etc., can be used in advance to comprehensively and systematically evaluate the behavior, language expression, social interaction ability, etc. of the sample child, understand the development level of the child as a whole, divide the child sample into social attention disorder and social attention healthy control groups, and provide accurate and appropriate visual cognitive classification labels based on the macro basis.
[0032] The S500 uses the sample fixation point map data of the sample children and the preset sample visual stimulus image data as inputs, and the corresponding visual cognitive classification labels as the expected outputs to train the initial machine learning model, and obtains the trained children's visual cognitive classification model.
[0033] It can be understood that the sample fixation point map data and the preset sample visual stimulus image data can be combined, and after appropriate preprocessing such as normalization and resizing, they are input into the model; the visual cognitive classification labels reflecting the visual cognitive state of the sample children are used as the expected outputs to provide goals for the model to learn. Taking a convolutional neural network as an example, after receiving the input data, the model will extract features through the convolutional layer, reduce the dimension through the pooling layer, integrate and classify through the fully connected layer, and then calculate the difference between the predicted output and the expected output using a measure such as the cross-entropy loss function to measure the gap. The parameters of each layer are updated through the backpropagation algorithm to make the model prediction closer to the expectation. After repeated iterative training with a large amount of sample data until the relevant indicators reach the preset threshold on the validation set, the initial model learns the relationship between the data and the labels, and finally obtains the trained children's visual cognitive classification model. This model can process the data of new target children to predict their visual cognitive classification results, providing a powerful tool for children's visual cognitive assessment and related research.
[0034] Exemplarily, the sample fixation point map data of children and the sample visual stimulus image data can be fused as the input image of the model to construct the local difference features of different groups on the sample visual stimulus image data, so as to highlight the unique pattern of autistic children in visual attention. Then, the ResNet34 deep learning model is used to extract the high-level semantic information for classification from the input image, and the fixation point map is introduced at different stages of the model. Through the fusion with the stage features, the expression ability of the local discriminative features is further enhanced. Finally, feature mapping is performed through the classification head to achieve the accurate identification of the autistic group. Compared with the traditional machine learning method, the fusion of the sample fixation point map data of children and the sample visual stimulus image data can fully exploit the deep information of the fixation point data, capture the non-linear relationship and long-term dependence in the eye movement data, avoid the complexity of manual feature construction, improve the adaptive ability and recognition accuracy of the model, and provide an efficient and intelligent technical means for the early screening of autistic children.
[0035] In a possible implementation manner, for S300, obtaining the sample fixation point map data of the sample children includes: S310. Present the preset active-guided free-view paradigm to the sample children according to the preset dynamic stimulus timing sequence to obtain the eye movement data of the sample children. The active-guided free-view paradigm includes different sample visual stimulus image data containing socially competitive semantic elements. In the sample visual stimuli images, social cues and non-social elements are symmetrically distributed, which is used to actively amplify the gaze avoidance behavior of children with autism.
[0036] It can be understood that the preset dynamic stimulus timing sequence refers to the presentation order and time rhythm of visual stimuli planned before the experiment. For example, in the order from easy to difficult and from simple scenes to complex scenes, control the display duration of each image and the interval between image switches, etc., so that the conditions for each sample child to receive stimuli are consistent, which is convenient for horizontal comparison of experimental results and data analysis. The active-guided free-view paradigm is an experimental mode of preset sample visual stimulus image data, which not only gives children the space to freely view images, but also actively guides children's attention through the sample visual stimulus image data containing socially competitive semantic elements. The active-guided free-view paradigm can include elements such as human interaction, competitive scenes, and cooperative tasks, aiming to simulate real social situations and stimulate different visual cognitive responses of children. At the same time, social cues (such as facial expressions, eye contact, body movements) and non-social elements (such as background objects, decorative patterns) in the images are symmetrically distributed. This design is to balance the attractiveness of various elements in the images and avoid interfering with the experimental results due to layout differences. For children with autism, there is generally a tendency to avoid gazing at social cues. This symmetric distribution design can more significantly amplify this behavioral characteristic, making it easier for researchers to observe and capture the differences in gaze behavior between children with autism and normal children, thus providing key data support for the early screening of autism and visual cognitive research.
[0037] S320. Divide the eye movement data based on the sample visual stimulus image data to obtain the independent eye movement data corresponding to each sample visual stimulus image data.
[0038] It can be understood that when the sample children are watching multiple groups of visual stimulus images, the eye movement tracking device continuously collects their eye movement data, and the eye movement data is continuously recorded. The eye movement data can be divided according to the sample visual stimulus image data. According to the display order and time nodes of the images, the continuous eye movement data is split into independent segments corresponding to each sample visual stimulus image data. For example, if in the experiment, sample visual stimulus image data A is shown first, and then sample visual stimulus image data B, and the eye movement tracking device records a complete eye movement trajectory data from start to end, through the start and end times of the display of sample visual stimulus image data A and sample visual stimulus image data B, this data can be divided into the eye movement data corresponding to image A and the eye movement data corresponding to image B, which helps to clearly analyze the eye movement characteristics of children when watching each specific sample visual stimulus image, avoid the confusion of eye movement data of different images, ensure that the subsequent eye movement analysis results of a single image are more accurate and targeted, and provide a basis for in-depth research on the response patterns of children to different visual stimuli.
[0039] S330. Based on the independent eye movement data corresponding to each sample visual stimulus image data, obtain the independent sample fixation point map data corresponding to each sample visual stimulus image data, and determine each sample fixation point map data as the sample fixation point map data of the sample children.
[0040] It can be understood that the independent eye movement data contains information such as the fixation points, fixation durations, and saccade paths of children when watching a single image, but the independent eye movement data is relatively scattered and abstract. To more intuitively present the attention distribution of children, independent sample fixation point map data can be generated based on the independent eye movement data. Effective fixation point coordinates and corresponding fixation durations can be extracted from the independent eye movement data, and abnormal data caused by interference factors such as instantaneous saccades, blinks, and head movements can be removed to ensure data quality. Then, algorithms such as Gaussian filtering are used to smooth the discrete fixation points and convert them into continuous visual graphs.
[0041] Optionally, in step S330, obtaining the independent sample fixation point map data corresponding to each sample visual stimulus image data based on the independent eye movement data corresponding to each sample visual stimulus image data includes: S331. Extract the fixation point coordinates and the fixation duration corresponding to each fixation point coordinate based on the independent eye movement data corresponding to each sample visual stimulus image data.
[0042] It can be understood that the independent eye movement data records the eye movement of the sample children when viewing a single sample visual stimulus image, but this independent eye movement data is original and messy. The gaze point coordinates can clearly indicate the specific position where the children's eyes fixate on the image, and the fixation duration corresponding to each gaze point coordinate reflects the degree of attention concentration of the children at that position. By extracting these two key pieces of information from the independent eye movement data, it is possible to provide basic data for subsequent analysis of the children's visual attention points and attention distribution patterns, which helps to deeply understand the children's interests and cognitive characteristics towards different image contents.
[0043] Exemplarily, S331, extracting the gaze point coordinates and the fixation duration corresponding to each gaze point coordinate based on the independent eye movement data corresponding to each sample visual stimulus image data, includes: S3311, based on the independent eye movement data corresponding to each sample visual stimulus image data, determining the eye movement trajectory corresponding to each sample visual stimulus image data.
[0044] It can be understood that the independent eye movement data is composed of a series of discrete eye movement sampling points, and the eye movement sampling points contain the position information of the eyes at different times. The discrete sampling points can be connected in chronological order to form a continuous path, that is, the eye movement trajectory. The eye movement trajectory intuitively shows the movement process of the children's eyes during the image viewing process, and it can reflect the children's visual exploration mode, such as starting to browse from the upper left corner of the image or first focusing on the central area of the image, etc., providing an overall framework for subsequent further analysis of eye movement behavior.
[0045] S3312, identifying the direction mutation points of the eye movement trajectory, and using the direction mutation points of the eye movement trajectory as segmentation nodes to divide the independent eye movement data into multiple gaze data segments.
[0046] It can be understood that during the eye movement process, when the movement direction of the eyes suddenly changes, it often means that the eyes have moved from one fixation point to another. By identifying the direction mutation points in the eye movement trajectory, it is possible to determine the conversion moment of the eye fixation state. Using the direction mutation points as segmentation nodes, the continuous independent eye movement data can be divided into multiple relatively independent gaze data segments. Each gaze data segment represents the fixation process of the children at a relatively stable position, and such a division helps to more carefully analyze the eye movement characteristics of the children at different fixation stages.
[0047] S3313, sorting the gaze data segments in chronological order, and calculating the stability index of the eye movement trajectory within each gaze data segment; where the stability index is used to reflect the eye movement change situation of the gaze data segment.
[0048] It can be understood that the gaze data segments are sorted in chronological order to ensure the logic and coherence of subsequent analysis, which is consistent with the actual eye movement process of children. The stability index is calculated because even within a gaze data segment, the eyes are not completely still and may move slightly. The stability index can be used to quantify the degree of change in this eye movement and determine whether the gaze data segment truly represents the child's effective gaze. If the stability index is low, it means that the eye movement changes little, and the child may have a relatively stable gaze at this position; conversely, if the stability index is high, it may mean that there is more interference in the gaze data segment or it is not a real gaze.
[0049] Exemplarily, S3313, sorting the gaze data segments in chronological order, and calculating the stability index of the eye movement trajectory in each gaze data segment, includes: S33131, sorting the gaze data segments in chronological order, respectively calculating the offsets of all adjacent sampling points in each gaze data segment on the horizontal and vertical coordinates, and obtaining the moving distances of all adjacent sampling points according to the offsets.
[0050] It can be understood that within a gaze data segment, the position change of adjacent sampling points reflects the movement of the eyes in a short period of time. The horizontal and vertical movement amplitudes of the eyes can be obtained by calculating the offsets of adjacent sampling points on the horizontal and vertical coordinates. According to the Pythagorean theorem, the offsets of the horizontal and vertical coordinates are comprehensively calculated to obtain the movement distances of adjacent sampling points. The calculation of these movement distances is the basis for the subsequent determination of stability indicators, and can more accurately describe the dynamic changes of eye movements within the gaze data segment.
[0051] S33132, obtaining a stability index of the eye movement trajectory in each gaze data segment according to the moving distances of all adjacent sampling points.
[0052] It can be understood that the moving distances of all adjacent sampling points reflect the fluctuation of eye movement in the entire gaze data segment. By performing statistical analysis on these moving distances, such as calculating the mean value, standard deviation and other statistics, an index that can represent the stability of eye movement in the gaze data segment can be obtained. The stability index can filter out truly stable gaze data segments, exclude unstable eye movement data that may be caused by accidental factors or interference, and improve the accuracy of subsequent extraction of gaze point coordinates and fixation duration.
[0053] S3314, screening the gaze data segments based on the stability index to obtain stable gaze data segments and extracting gaze point coordinates and gaze duration corresponding to each gaze point coordinate according to the stable gaze data segments.
[0054] It can be understood that since the original gaze data segments may contain some unstable eye movement conditions, such as rapid saccades of the eyes or brief jitters when affected by external interference, the unstable data will affect the accurate extraction of gaze point coordinates and fixation durations. Screening can be performed through stability indicators to remove those unstable gaze data segments and only retain the parts that truly represent the stable fixation of children, obtaining stable gaze data segments. In the stable gaze data segments, the average value of the sampling point coordinates within the data segment is taken as the gaze point coordinate, and the duration of this data segment is the corresponding fixation duration. The information extracted in this way can more accurately reflect the true fixation behavior of children.
[0055] S332. Perform Gaussian smoothing on the gaze point coordinates and the fixation duration corresponding to each gaze point coordinate to obtain independent sample gaze point map data corresponding to each sample visual stimulus image data.
[0056] It can be understood that the extracted gaze point coordinates and fixation durations are discrete data points, and directly presenting them may be rather messy and not conducive to intuitively observing the attention distribution of children. Gaussian smoothing is a commonly used signal processing method. It makes the data smoother and more continuous by performing weighted averaging on each gaze point and the points around it. After Gaussian smoothing processing, the noise and mutations in the data can be eliminated, and the discrete gaze points are converted into a continuous graph with a certain distribution pattern, that is, independent sample gaze point map data.
[0057] In a possible implementation manner, S500. Using the sample gaze point map data of the sample children and the preset sample visual stimulus image data as inputs and the corresponding visual cognitive classification labels as the expected outputs, train the initial machine learning model to obtain a trained children's visual cognitive classification model, including: S510. Fuse the sample gaze point map data with the preset sample visual stimulus image data to obtain a discriminative input image corresponding to the sample gaze point map data; among them, the specific fusion formula for obtaining the discriminative input image is: , where Img represents the j-th sample visual stimulus image data, represents the sample gaze point map data of sample child i under the stimulation of the j-th sample visual stimulus image data, is the discriminative input image.
[0058] It can be understood that by combining the fixation point map with the original visual stimulus image, the areas that the subject focuses on during viewing can be highlighted. This not only preserves the overall semantic information of the image but also effectively captures the local features related to the fixation points, thereby increasing the attention of the classification model to key regions. As a direct reflection of the subject's eye movement characteristics, the fixation point map can provide additional visual cues for the model, helping the model make more accurate judgments during recognition. Especially when distinguishing different groups (such as the autism group and the healthy control group), it can provide more effective support. The formula for the discriminative input image is as follows, and the specific fusion formula for the discriminative input image is: , where Img represents the visual stimulus image data of the j-th sample, represents the sample fixation point map data of sample child i under the stimulation of the visual stimulus image data of sample j, is the discriminative input image.
[0059] S520. Perform stage semantic extraction on the discriminative input image through the ResNet34 backbone network to obtain the discriminative semantic information corresponding to different network stages.
[0060] It can be understood that to effectively extract the semantic information for classification from the image, the ResNet34 model can be selected as the backbone network. ResNet34 performs excellently in image classification tasks due to its deep network structure and residual connections, which can effectively avoid the problem of gradient disappearance and deepen the network depth to extract more abundant feature information. The deep structure of ResNet34 enables it to capture the complex semantic features in the image, providing strong representation ability, and is particularly suitable for feature extraction in image classification and visual tasks. Therefore, this method uses ResNet34 as the backbone network to extract the semantic information of each stage , and this network can extract the low-level and high-level semantic features of the image layer by layer, ensuring that the discriminative features required for the classification task are fully retained, and making appropriate adjustments and optimizations on this basis to improve the classification accuracy, this method.
[0061] Optionally, S520. Perform stage semantic extraction on the discriminative input image through the ResNet34 backbone network to obtain the discriminative semantic information corresponding to different network stages, including: S521. Perform stage semantic extraction on the discriminative input image through the ResNet34 backbone network to obtain the semantic information corresponding to each network stage.
[0062] It can be understood that the semantic information corresponding to each network stage can be extracted using the ResNet34 backbone network , the ResNet34 backbone network can extract low-level and high-level semantic features of images layer by layer, ensuring that the discriminative features required for the classification task are fully retained. Among them, the low-level features mainly contain basic information such as edges and textures, while the high-level features capture more complex semantic patterns to support the final classification.
[0063] S522, fuse the semantic information corresponding to each network stage with the sample fixation map data respectively to obtain the discriminative semantic information corresponding to different network stages; among them, the specific fusion formula for obtaining the discriminative semantic information is: , represents the processing operations of each stage of ResNet34, represents the sample fixation map data of sample child i under the stimulation of the j sample visual stimulus image data, is the discriminative input image, represents the semantic information obtained at the network stage, represents the discriminative semantic information, is the pooling operation.
[0064] It can be understood that in order to enhance the effectiveness of the fixation map in the feature processing of each stage, this method fuses the sample fixation map data with the semantic information corresponding to each extracted network stage at different network stages, realizes the progressive attention modulation of the fixation map on the features of each layer of the neural network, and enables the entire feature extraction process of the model from the bottom layer to the top layer to closely fit the real fixation behavior of children (especially children with autism), rather than the general image semantics. Through the hierarchical fixation logic embedding, the accurate modeling of specific visual behaviors such as gaze avoidance is realized. Compared with the traditional model that only processes in a single stage or ignores the fixation behavior, it can capture the visual cognitive features of the target group more efficiently and accurately. Obtain discriminative semantic information; , the specific fusion formula is as follows: , , where represents the processing operations of each stage of ResNet34, represents the extracted semantic information, represents the discriminative semantic information, is the pooling operation, which not only improves the sensitivity of the model to local discriminative regions, but also enhances the interpretability of the classification task. By combining multi-level semantic information with fixation data, classification is performed more accurately, improving the overall performance of the model.
[0065] S530, based on the discriminative semantic information corresponding to different network stages, map and classify the visual cognitive classification labels of sample children to obtain the trained visual cognitive classification model of children.
[0066] It can be understood that based on the discriminative semantic information corresponding to different network stages, the discriminative features of the last layer can be aggregated through the global average pooling layer to reduce the feature dimension and retain key information. The fully connected layer is used to further map the pooled features to classification categories, and the class probabilities are calculated through the Softmax activation function. Finally, the classification result, that is, the visual cognitive classification label, is obtained, and the trained child visual cognitive classification model is obtained, making full use of the semantic information of different levels, and at the same time combining the gaze point features of the sample children and the active stimuli of the visual stimulus images themselves to achieve accurate classification.
[0067] Optionally, in S530, based on the discriminative semantic information corresponding to different network stages, the visual cognitive classification labels of the sample children are mapped and classified to obtain the trained child visual cognitive classification model, including: In S531, based on the discriminative semantic information corresponding to different network stages, the discriminative semantic information of the last network stage is aggregated in the global average pooling layer to obtain the reduced-dimensional discriminative features corresponding to the discriminative semantic information.
[0068] It can be understood that after obtaining the discriminative semantic information by fusing the gaze point map and semantic information through each network stage, the global average pooling layer can be used to process the discriminative semantic information of the last network stage. By calculating the global average value of each channel of the feature map, the spatial dimension is compressed, redundant local spatial details are removed, and the core features that can reflect the overall fixation pattern of children and the semantic association of images are retained, so as to obtain compact reduced-dimensional discriminative features, reducing the computational burden for subsequent classification and strengthening the feature robustness.
[0069] In S532, based on the reduced-dimensional discriminative features corresponding to the discriminative semantic information, the visual cognitive classification labels of the sample children are mapped and classified in the fully connected layer, and the probability value of each visual cognitive classification label is calculated through the Softmax activation function to determine the visual cognitive classification label of the sample children, and the trained child visual cognitive classification model is obtained; among them, the specific formula for determining the visual cognitive classification label is: where represents the discriminative semantic information of the last network stage, represents the global average pooling layer, Linear is the fully connected layer, Soft represents the softmax function, is a vector containing the probability values of each visual cognitive classification label.
[0070] It can be understood that the fully connected layer can be used to further map the pooled features to classification categories, and the class probabilities are calculated through the Softmax activation function. Finally, the classification result is obtained, and the formula is as follows: where Represents the discriminative semantic information of the last layer, Represents the global average pooling layer, Is a fully connected layer, Represents the softmax function, Is the final classification result, making full use of the semantic information of different levels and combining the fixation point features of the subjects to achieve accurate classification.
[0071] Finally, in terms of model training, the input images can be preprocessed by adjusting the original resolution of 1920×1080 to 480×270 to reduce the computational cost and adapt to the model input size. In terms of the optimization strategy, AdamW can be selected as the optimizer, and the weight decay can be set to 0.05 to suppress overfitting and improve the generalization ability of the model. In addition, during the training process, the learning rate can be set to , the batch size is 8 to ensure stable training. After 50 epochs, the iteration is completed, and the best model parameters are saved to obtain the trained child visual cognitive classification model, which realizes guiding the deep learning model to focus on the key image regions for distinguishing the autism group, thereby achieving accurate classification, strengthening the model's perception ability of the core features, improving the classification accuracy, effectively suppressing the interference of irrelevant information, fully capturing the non-linear relationships and long-term dependencies in the eye movement data, and enhancing the interpretability and generalization ability of the model.
[0072] Corresponding to the fixation point-guided child visual cognitive classification method in the above embodiments, the present application embodiment also provides a fixation point-guided child visual cognitive classification system, and each unit of the system can implement each step of the fixation point-guided child visual cognitive classification method. Figure 3 FIG. shows the structural block diagram of the fixation point-guided child visual cognitive classification system provided by the embodiment of the present application. For the sake of convenience of description, only the parts related to the embodiment of the present application are shown.
[0073] Refer to Figure 3 , the fixation point-guided child visual cognitive classification system includes: An acquisition unit for acquiring the fixation point map data of the target child; A classification unit for processing the fixation point map data and the preset visual stimulus image data using the child visual cognitive classification model to obtain the visual cognitive classification result of the target child; Wherein, the child visual cognitive classification model is a machine learning model pre-trained by prompt learning using the sample fixation point map data and sample visual stimulus image data of the sample child.
[0074] It should be noted that, regarding the information interaction, execution process, etc. between the above-mentioned systems / units, since they are based on the same concept as the method embodiments of the present application, for their specific functions and the technical effects brought about, reference can be specifically made to the method embodiment section, and details will not be elaborated here.
[0075] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the system is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated in a processing unit, or each unit module exists physically alone, or two or more unit modules are integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present application. The specific working process of the units and modules in the above system can refer to the corresponding process in the foregoing method embodiment, and details will not be elaborated here.
[0076] The embodiment of the present application also provides an electronic device, Figure 4 which is a schematic structural diagram of the electronic device provided by an embodiment of the present application. As Figure 4 shown, the electronic device 6 in this embodiment includes: at least one processor 60 ( Figure 4 only one is shown here), at least one memory 61 ( Figure 4 only one is shown here), and a computer program 62 stored in the at least one memory 61 and executable on the at least one processor 60. When the processor 60 executes the computer program 62, the electronic device 6 implements the steps in any of the above-mentioned embodiments of the child visual cognitive classification method guided by gaze points, or the functions of each unit in the above-mentioned system embodiments.
[0077] Exemplarily, the computer program 62 can be divided into one or more units. The one or more units are stored in the memory 61 and executed by the processor 60 to complete the present application. The one or more units can be a series of computer program instruction segments capable of completing specific functions, and these instruction segments are used to describe the execution process of the computer program 62 in the electronic device 6.
[0078] The electronic device 6 can be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server, etc. The electronic device may include, but is not limited to, a processor 60 and a memory 61. Those skilled in the art can understand, Figure 4This is only an example of the electronic device 6, and does not constitute a limitation on the electronic device 6. It may include more or fewer components than those shown in the figure, or combine certain components, or different components. For example, it may also include input / output devices, network access devices, buses, etc.
[0079] The processor 60 may be a central processing unit (CPU), and the processor 60 may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0080] In some embodiments, the memory 61 may be an internal storage unit of the electronic device 6, such as the hard disk or memory of the electronic device 6. In other embodiments, the memory 61 may also be an external storage device of the electronic device 6, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc., equipped on the electronic device 6. Further, the memory 61 may also include both the internal storage unit and the external storage device of the electronic device 6. The memory 61 is used to store an operating system, application programs, a boot loader, data, and other programs, such as the program code of the computer program, etc. The memory 61 may also be used to temporarily store data that has been output or is to be output.
[0081] The embodiments of the present application also provide a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps in any of the above method embodiments are implemented.
[0082] The embodiments of the present application provide a computer program product, and when the computer program product runs on an electronic device, the electronic device implements the steps in any of the above method embodiments.
[0083] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-mentioned embodiment methods of this application, a computer program can be used to instruct relevant hardware to complete. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can at least include: any entity or device that can carry the computer program code to an electronic device, a recording medium, a computer memory, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), an electrical carrier signal, a telecommunication signal, and a software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk, or an optical disc, etc. In some jurisdictions, according to legislation and patent practice, the computer-readable medium cannot be an electrical carrier signal and a telecommunication signal.
[0084] In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0085] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. A professional technician can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0086] In the embodiments provided in this application, it should be understood that the disclosed fixation-point-guided children's visual cognitive classification system / electronic device and method can be implemented in other ways. For example, the fixation-point-guided children's visual cognitive classification system / electronic device embodiments described above are only illustrative. For example, the division of the units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of devices or units can be in electrical, mechanical or other forms.
[0087] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or distributed across multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0088] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. A gaze point guided children's visual cognition classification method, characterized in that: include: Obtain the gaze map data of the target child; Using a children's visual cognition classification model to process the gaze point map data and preset visual stimulation image data to obtain a visual cognition classification result of the target child; The children's visual cognition classification model is a machine learning model obtained by using sample gaze point map data and sample visual stimulation image data of sample children through prompt learning training.
2. The gaze point guided children's visual cognition classification method according to claim 1, characterized in that: Before using the children's visual cognition classification model to process the gaze point map data and the visual stimulation image data to obtain the visual cognition classification result of the target child, the method further includes: Obtain sample gaze map data of sample children; Determining visual cognitive classification labels of the sample children; The sample gaze point map data of the sample children and the preset sample visual stimulation image data are used as input, and the corresponding visual cognition classification label is used as the expected output, and the initial machine learning model is trained to obtain the trained children's visual cognition classification model.
3. The gaze point guided children's visual cognition classification method according to claim 2, characterized in that: The step of obtaining sample gaze point graph data of the sample child includes: A preset active guided free view paradigm is displayed to the sample child according to a preset dynamic stimulation timing sequence to obtain eye movement data of the sample child; the active guided free view paradigm includes different sample visual stimulation image data containing social competitive semantic elements, and the social cues and non-social elements in the sample visual stimulation images are symmetrically distributed, so as to actively amplify the gaze avoidance behavior of the autistic child; Dividing the eye movement data according to the sample visual stimulation image data to obtain independent eye movement data corresponding to each piece of the sample visual stimulation image data; Based on the independent eye movement data corresponding to each of the sample visual stimulation image data, independent sample gaze point map data corresponding to each of the sample visual stimulation image data is obtained, and each of the sample gaze point map data is determined as the sample gaze point map data of the sample child.
4. The gaze point guided children's visual cognition classification method according to claim 3, characterized in that: Based on the independent eye movement data corresponding to each piece of the sample visual stimulation image data, independent sample gaze point map data corresponding to each piece of the sample visual stimulation image data is obtained, including: Extracting gaze point coordinates and gaze duration corresponding to each gaze point coordinate based on the independent eye movement data corresponding to each sample visual stimulation image data; Gaussian smoothing is performed on the gaze point coordinates and the gaze duration corresponding to each of the gaze point coordinates to obtain independent sample gaze point map data corresponding to each of the sample visual stimulation image data.
5. The gaze point guided children's visual cognition classification method according to claim 4, characterized in that: The extracting of gaze point coordinates and the gaze duration corresponding to each gaze point coordinate based on the independent eye movement data corresponding to each sample visual stimulation image data comprises: Determine the eye movement trajectory corresponding to each piece of the sample visual stimulation image data based on the independent eye movement data corresponding to each piece of the sample visual stimulation image data; Identifying a direction mutation point of the eye movement trajectory, and using the direction mutation point of the eye movement trajectory as a segmentation node to divide the independent eye movement data into a plurality of gaze data segments; Sorting the gaze data segments in chronological order, and calculating a stability index of the eye movement trajectory in each of the gaze data segments; wherein the stability index is used to reflect the eye movement change of the gaze data segment; The gaze data segments are screened based on the stability index to obtain stable gaze data segments, and gaze point coordinates and gaze durations corresponding to each gaze point coordinate are extracted according to the stable gaze data segments.
6. The gaze point guided children's visual cognition classification method according to claim 5, characterized in that: The step of sorting the gaze data segments in chronological order and calculating the stability index of the eye movement trajectory in each of the gaze data segments includes: Sorting the gaze data segments in chronological order, respectively calculating the offsets of all adjacent sampling points in each gaze data segment on the horizontal and vertical coordinates, and obtaining the moving distances of all adjacent sampling points according to the offsets; A stability index of the eye movement trajectory in each of the gaze data segments is obtained according to the moving distances of all the adjacent sampling points.
7. The gaze point guided children's visual cognition classification method according to claim 2, characterized in that: The method takes the sample gaze point map data of the sample child and the preset sample visual stimulation image data as input, takes the corresponding visual cognition classification label as expected output, trains the initial machine learning model, and obtains the trained child visual cognition classification model, including: The sample gaze point map data is fused with the preset sample visual stimulation image data to obtain a discriminative input image corresponding to the sample gaze point map data; wherein the specific fusion formula for obtaining the discriminative input image is: , where Img represents the jth sample visual stimulus image data, represents the sample gaze point map data of sample child i under the stimulation of sample visual stimulus image data j, is the discriminative input image; Performing stage semantic extraction on the discriminative input image through the ResNet34 backbone network to obtain discriminative semantic information corresponding to different network stages; Based on the discriminative semantic information corresponding to different network stages, the visual cognition classification labels of the sample children are mapped and classified to obtain the trained children's visual cognition classification model.
8. The gaze point guided children's visual cognition classification method according to claim 7, characterized in that: The step of performing stage semantic extraction on the discriminative input image through the ResNet34 backbone network to obtain discriminative semantic information corresponding to different network stages includes: Performing stage semantic extraction on the discriminative input image through the ResNet34 backbone network to obtain semantic information corresponding to each network stage; The semantic information corresponding to each network stage is fused with the sample gaze point map data to obtain the discriminative semantic information corresponding to different network stages; wherein the specific fusion formula for obtaining the discriminative semantic information is: , Represents the processing operations of each stage of ResNet34, represents the sample gaze point map data of sample child i under the stimulation of sample visual stimulus image data j, is the discriminative input image, represents the semantic information obtained in the network stage, Represents discriminative semantic information, For pooling operation.
9. The gaze point guided children's visual cognition classification method according to claim 7, characterized in that: The mapping and classification of the visual cognition classification labels of the sample children based on the discriminative semantic information corresponding to the different network stages to obtain the trained children's visual cognition classification model includes: Based on the discriminative semantic information corresponding to different network stages, the discriminative semantic information of the last layer of the network stage is aggregated in a global average pooling layer to obtain a reduced dimension discriminative feature corresponding to the discriminative semantic information; The visual cognition classification labels of the sample children are mapped and classified in the fully connected layer according to the dimensionality reduction discriminant features corresponding to the discriminative semantic information, and the probability value of each visual cognition classification label is calculated through the Softmax activation function to determine the visual cognition classification label of the sample child, and obtain the trained visual cognition classification model of the child; wherein the specific formula for determining the visual cognition classification label is: in Represents the discriminative semantic information of the network stage described in the last layer, represents the global average pooling layer, Linear represents the fully connected layer, Soft represents the softmax function, is a vector containing the probability values of each of the visual cognition classification labels.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method according to any one of claims 1 to 9 is implemented.
Citation Information
Patent Citations
Eye movement track identification method and system
CN113469053A
Autism early screening eye tracker detection system and method, medium, equipment and terminal
CN115530830A
Diagnostic processing method and system for autism spectrum disorder
CN119418908A