Eye movement tracking system for aphasia rehabilitation training, vocabulary difficulty adaptive adjustment method, equipment and medium
Through the eye tracking system, the patient's gaze behavior is identified and analyzed, and combined with semantic characteristics and model evaluation, adaptively regulates the voice content of aphasia rehabilitation training, solving the problem of lack of objective feedback and poor adaptability in the existing system, and improving the accuracy and effectiveness of rehabilitation training.
Patent Information
- Application Number
- CN202511093470.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-06
- Publication Date
- 2025-09-05
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing aphasia rehabilitation training system lacks objective data feedback methods, making it difficult to accurately grasp the rehabilitation progress, and the training adaptability is poor, so it is impossible to deeply correlate the patient's eye movement and judging characteristics and the speech sequence structure of the training speech content.
The eye tracking system recognizes the patient's gaze focus on the task display interface, records the gaze time data and response sequences, combines semantic diffusion characteristics and semantic recognition gradients, and uses the interactive evaluation model to adaptively adjust the vocabulary difficulty of training speech, and adjusts the next stage of rehabilitation training based on the cognitive efficacy index.
It achieves a deep correlation between the speech recognition ability and gaze behavior of aphasia patients, adaptively adjusts the training difficulty, improves the accuracy and adaptability of rehabilitation training, and provides objective training evaluation and feedback.
Smart Images

Figure CN120600244A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of eye tracking technology, and in particular to an eye tracking system, a vocabulary difficulty adaptive adjustment method, a device, and a medium for aphasia rehabilitation training. Background Art
[0002] Aphasia is a language disorder caused by damage to the brain's language center, which manifests as varying degrees of impairment in language expression, comprehension, naming, reading, and writing, seriously affecting patients' communication ability and quality of life. Traditional aphasia rehabilitation treatment is mainly based on systematic language training under the guidance of speech therapists, covering naming training, vocabulary association, sentence imitation, and contextual dialogue, which has obvious limitations in practical application.
[0003] The language comprehension ability of aphasic patients during rehabilitation training is usually subjectively assessed by therapists based on the content of their answers or their performance. There is a lack of objective data feedback, making it difficult to accurately grasp the progress of rehabilitation. In addition, existing auxiliary training systems based on eye tracking mostly focus on the collection of behavioral data and basic statistical analysis. Although objective physiological signals are introduced to evaluate the language processing ability of aphasic patients, in rehabilitation training practice, training systems that only rely on basic statistical analysis are too simple in function and have poor training adaptability. Therefore, how to deeply associate the patient's eye gaze characteristics with the word order structure of the training speech content, and then adaptively adjust the training speech during rehabilitation training has become a difficult problem facing the industry. Summary of the Invention
[0004] Based on this, the present application provides an eye tracking system, vocabulary difficulty adaptive adjustment method, equipment and medium for aphasia rehabilitation training, which is used to deeply associate the patient's eye gaze characteristics with the word order structure of the training speech content, and then adaptively adjust the training speech during the rehabilitation training process.
[0005] In a first aspect, the present application provides a method for adaptively adjusting vocabulary difficulty during aphasia rehabilitation training, which is applied to an eye tracking system for aphasia rehabilitation training. When the eye tracking system plays rehabilitation training speech to a target user, the corresponding vocabulary in the task display interface is synchronously highlighted. The method comprises the following steps: Play the initial rehabilitation training speech of the current stage to the target user and identify the target user's gaze focus in the task display interface; When the gaze focus remains on the same word, candidate words matching the current word are arranged in a circular pattern in the task display interface, and the target user's gaze time data on the candidate words is recorded. The semantic diffusion feature representing the target user's attention to the current word is extracted from the gaze time data; When the gaze focus is in a scanning mode in the task display interface, the target user's time response sequence to all scanned words is recorded, and then the target user's semantic recognition gradient for the initial rehabilitation training speech is determined based on the response characteristics of the words corresponding to the initial rehabilitation training speech in the time response sequence; performing response feedback analysis on the speech recognition of the target user during the initial rehabilitation training based on the semantic diffusion feature and the semantic recognition gradient to obtain the semantic response deviation of the target user at the current stage; The pre-built interactive evaluation model combines the gaze characteristics of each word in the target user's corresponding historical eye movement trajectory data with the semantic structure characteristics of the corresponding rehabilitation training speech to map the target user's speech recognition ability and gaze behavior, and obtains the target user's gaze association relationship under different semantic structure characteristics; The cognitive efficacy index of the target user at the current stage is determined according to the semantic response deviation and all gaze association relationships, and the vocabulary difficulty of the rehabilitation training speech at the next stage is adaptively adjusted based on the cognitive efficacy index.
[0006] In some embodiments, identifying the gaze focus of the target user in the task display interface specifically includes: Acquiring eye movement signals of the target user when performing the initial rehabilitation training speech; performing noise removal processing on the eye movement signal to obtain an eye movement signal after noise removal; The gaze focus of the target user in the task display interface is identified based on a preset gaze determination algorithm combined with the eye movement signal after noise removal.
[0007] In some embodiments, extracting the semantic diffusion feature representing the target user's attitude toward the current word from the gaze time data specifically includes: Pre-build semantic diffusion model; extracting the target user's gaze time distribution on the candidate vocabulary from the gaze time data; The gaze time distribution is input into the semantic diffusion model, and then the semantic diffusion characteristics of the target user for the current word are output.
[0008] In some embodiments, determining the target user's semantic recognition gradient for the initial rehabilitation training speech based on the response characteristics of the vocabulary corresponding to the initial rehabilitation training speech in the time response sequence specifically includes: Extracting the time response unit of each word in the time response sequence in the initial rehabilitation training speech; The corresponding lexical response feature vector is constructed based on the fixation start time, end time, duration, time delay of the first fixation and the number of repeated fixations during the scanning process in each time response unit; All vocabulary response feature vectors are input into a preset recognition gradient evaluation model to output the semantic recognition gradient of the target user for the initial rehabilitation training speech.
[0009] In some embodiments, based on the semantic diffusion feature and the semantic recognition gradient, a response feedback analysis is performed on the speech recognition of the target user during the initial rehabilitation training process to obtain the semantic response deviation of the target user at the current stage, specifically including: The word corresponding to the time when the focus of attention continuously stays on the same word is regarded as the stay word; Obtain the semantic diffusion features of the target user for the current word and other words that remain; Matching all the retained words with the initial rehabilitation training speech of the current stage to obtain multiple matching words and multiple non-matching words; Based on the weight distribution mechanism, semantic response weights are distributed to all matching words and corresponding semantic diffusion features, all non-matching words and the target user's semantic recognition gradient for the initial rehabilitation training speech, and then the semantic response deviation of the target user at the current stage is calculated based on the weight factors corresponding to all matching words, non-matching words and semantic recognition gradients.
[0010] In some embodiments, the speech recognition ability and gaze behavior of the target user are correlated and mapped by combining the gaze features of each word in the target user's historical eye movement trajectory data and the semantic structure features of the corresponding rehabilitation training speech through a pre-built interaction evaluation model. The gaze correlation relationship of the target user under different semantic structure features is obtained, specifically including: Obtain the target user's eye movement trajectory data corresponding to multiple historical rehabilitation training tasks, and extract the gaze features corresponding to each word from it; Acquire rehabilitation training speech corresponding to each historical rehabilitation training task, and then determine the semantic structure characteristics of each rehabilitation training speech; The gaze features corresponding to each extracted word and the semantic structure features of each rehabilitation training speech are aligned on a word-by-word basis to construct data pairs containing the relationship between the word playback order and gaze intensity in multiple training tasks; The above data pairs are input into the pre-built interaction evaluation model to obtain the gaze association relationship of the target user under different semantic structure features.
[0011] In some embodiments, determining the cognitive efficacy index of the target user at the current stage based on the semantic response deviation and all gaze association relationships specifically includes: Based on the gaze association relationship of the target user under different semantic structure features, the training evaluation coefficients corresponding to different semantic structure features are set; Determine the semantic structure characteristics of the initial rehabilitation training speech at the current stage; Matching the semantic structure features of the initial rehabilitation training speech with the semantic structure features of the target user in the historical rehabilitation training task, thereby obtaining a training evaluation coefficient corresponding to the semantic response deviation; The cognitive efficacy index of the target user at the current stage is calculated based on the semantic response deviation and the corresponding training evaluation coefficient.
[0012] In a second aspect, the present application provides an eye tracking system for aphasia rehabilitation training, which includes a vocabulary difficulty adaptive adjustment unit, and the vocabulary difficulty adaptive adjustment unit includes: A recognition module is used to play the initial rehabilitation training speech of the current stage to the target user and identify the target user's gaze focus in the task display interface; a processing module configured to, when the gaze focus remains on the same word, arrange candidate words matching the current word in a circular manner in the task display interface, record the target user's gaze time data on the candidate words, and extract semantic diffusion features representing the target user's attention to the current word from the gaze time data; The processing module is further configured to record, when the gaze focus is in a scanning mode in the task display interface, a time response sequence of the target user to all scanned words, and then determine the target user's semantic recognition gradient for the initial rehabilitation training speech based on the response characteristics of the words corresponding to the initial rehabilitation training speech in the time response sequence; The processing module is further configured to perform response feedback analysis on the speech recognition of the target user during the initial rehabilitation training based on the semantic diffusion feature and the semantic recognition gradient to obtain the semantic response deviation of the target user at the current stage; The processing module is further configured to perform correlation mapping between the target user's speech recognition ability and gaze behavior by combining the gaze features of each word in the target user's corresponding historical eye movement trajectory data and the semantic structure features of the corresponding rehabilitation training speech through a pre-built interactive evaluation model, thereby obtaining the target user's gaze correlation relationship under different semantic structure features; The execution module is used to determine the cognitive efficacy index of the target user in the current stage according to the semantic response deviation and all gaze association relationships, and then adaptively adjust the vocabulary difficulty of the rehabilitation training speech in the next stage based on the cognitive efficacy index.
[0013] In a third aspect, the present application provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, it implements the steps of the above-mentioned method for adaptively adjusting vocabulary difficulty in the aphasia rehabilitation training process.
[0014] In a fourth aspect, the present application provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the above-mentioned method for adaptively adjusting vocabulary difficulty in the aphasia rehabilitation training process.
[0015] The technical solutions provided by the embodiments disclosed in this application have the following beneficial effects: In the eye tracking system, vocabulary difficulty adaptive adjustment method, device and medium for aphasia rehabilitation training provided by the present application, the initial rehabilitation training voice of the current stage is first played to the target user, and the target user's gaze focus in the task display interface is identified; when the gaze focus continues to stay on the same vocabulary, candidate vocabulary matching the current vocabulary is arranged in a circular manner in the task display interface, and the target user's gaze time data on the candidate vocabulary is recorded, and the semantic diffusion characteristics representing the target user's attitude towards the current vocabulary are extracted from the gaze time data; when the gaze focus is in a scanning mode in the task display interface, the target user's time response sequence to all scanned vocabulary is recorded, and then the target user's attention is determined based on the response characteristics of the vocabulary corresponding to the initial rehabilitation training voice in the time response sequence. The method comprises the following steps: first, analyzing the target user's semantic recognition gradient of the initial rehabilitation training speech; performing response feedback analysis on the target user's speech recognition during the initial rehabilitation training based on the semantic diffusion feature and the semantic recognition gradient, and obtaining the semantic response deviation of the target user at the current stage; performing correlation mapping on the target user's speech recognition ability and gaze behavior through a pre-constructed interactive evaluation model combined with the gaze features of each vocabulary in the target user's corresponding historical eye movement trajectory data and the semantic structure features of the corresponding rehabilitation training speech, and obtaining the gaze association relationship of the target user under different semantic structure features; determining the cognitive efficacy index of the target user at the current stage based on the semantic response deviation and all the gaze association relationships, and then adaptively adjusting the vocabulary difficulty of the rehabilitation training speech in the next stage based on the cognitive efficacy index.
[0016] It can be seen that the present application determines the cognitive efficiency index of the target user in the current stage based on the semantic response deviation and all gaze association relationships, and then adaptively adjusts the vocabulary difficulty of the rehabilitation training speech in the next stage based on the cognitive efficiency index; first, the semantic diffusion feature representing the target user's current vocabulary is extracted from the gaze time data. The semantic diffusion feature is a numerical value reflecting the target user's semantic association range and semantic transfer ability, which can help characterize the user's association breadth and semantic transfer ability of the core vocabulary, reflect the user's word meaning understanding depth, and provide a basis for identifying semantic processing deviations; secondly, the semantic recognition gradient of the target user for the initial rehabilitation training speech is determined according to the response characteristics of the corresponding vocabulary of the initial rehabilitation training speech in the time response sequence. The semantic recognition gradient is used to quantify the cognitive load level of the target user at different semantic depth levels, and can judge the degree of the user's obstacle in speech recognition; then, based on the semantic diffusion feature and the semantic recognition gradient, the speech recognition of the target user in the initial rehabilitation training process is analyzed for response feedback, and the semantic response deviation of the target user in the current stage is obtained. The semantic response deviation is used to quantify the user's level of cognitive load at different semantic depth levels, and can judge the degree of obstacle of the user in speech recognition. The semantic attention shift or understanding error generated in the process, the comprehensive matching and non-matching word gaze behavior and recognition ability, the degree of deviation between the user's actual recognition performance and the expected semantics is measured, and used to feedback the training effect; then, the speech recognition ability and gaze behavior of the target user are correlated and mapped through the pre-constructed interactive evaluation model combined with the gaze characteristics of each word in the target user's corresponding historical eye movement trajectory data and the semantic structure characteristics of the corresponding rehabilitation training speech, and the gaze association relationship of the target user under different semantic structure characteristics is obtained. The gaze association relationship represents the influence relationship between the playback order of each word in the rehabilitation training speech and the target user's gaze degree on each word in the rehabilitation training semantics, reflecting the user's adaptability to the semantic structure and providing an important reference dimension for the training evaluation index; finally, the cognitive efficacy index of the target user in the current stage is determined based on the semantic response deviation and all the gaze association relationships, and the vocabulary difficulty of the rehabilitation training speech in the next stage is adaptively adjusted based on the cognitive efficacy index; in summary, the solution of the present application deeply associates the patient's eye movement gaze characteristics with the word order structure of the training speech content, and then adaptively adjusts the training speech during the rehabilitation training process. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 is an exemplary flow chart of a method for adaptively adjusting vocabulary difficulty during aphasia rehabilitation training according to some embodiments of the present application; Figure 2 is a schematic diagram of an application scenario of an eye tracking system for aphasia rehabilitation training according to some embodiments of the present application; Figure 3is a schematic diagram of a process for determining semantic response deviation according to some embodiments of the present application; Figure 4 is a schematic structural diagram of a vocabulary difficulty adaptive adjustment unit according to some embodiments of the present application; Figure 5 It is a structural diagram of a computer device for implementing a method for adaptively adjusting vocabulary difficulty during aphasia rehabilitation training according to some embodiments of the present application. DETAILED DESCRIPTION
[0018] In order to better understand the above technical solution, the above technical solution will be described in detail below with reference to the accompanying drawings and specific implementation methods.
[0019] refer to Figure 1 , which is an exemplary flow chart of a method for adaptively adjusting vocabulary difficulty during aphasia rehabilitation training according to some embodiments of the present application. The method mainly includes the following steps: In step 101, the initial rehabilitation training speech of the current stage is played to the target user, and the gaze focus of the target user in the task display interface is identified.
[0020] In specific implementation, playing the initial rehabilitation training voice of the current stage to the target user can be achieved in the following way: first, according to the training plan corresponding to the current stage, a voice clip that matches the training plan is retrieved from a preset voice material library, and the voice material library contains standard training voice content classified by vocabulary difficulty, semantic type and syntactic structure.
[0021] It should be noted that this application uses an integrated audio playback module to synchronously play the voice clip on the task display interface, and ensures that the audio playback is consistent with the vocabulary content displayed on the interface to maintain the alignment of visual and auditory information.
[0022] In some embodiments, identifying the target user's gaze focus in the task display interface may be achieved by using the following steps, namely: Acquiring eye movement signals of the target user when performing the initial rehabilitation training speech; performing noise removal processing on the eye movement signal to obtain an eye movement signal after noise removal; The gaze focus of the target user in the task display interface is identified based on a preset gaze determination algorithm combined with the eye movement signal after noise removal.
[0023] In specific implementation, obtaining the eye movement signal of the target user when executing the initial rehabilitation training voice can be achieved in the following manner, for example: the eye movement data of the target user when executing the initial rehabilitation training voice can be captured in real time through an eye tracking device, and the eye movement data includes: horizontal and vertical displacement of the eye, position of the gaze point, gaze time and blinking frequency, etc., and the eye movement data is used as the eye movement signal of the target user when executing the initial rehabilitation training voice; wherein, the eye tracking device accurately records the eye movement through a high-frequency sampling rate (such as: 200 times per second or higher).
[0024] In a specific implementation, the eye movement signal is subjected to noise removal processing to obtain the eye movement signal after noise removal, which can be achieved in the following manner, namely: for the collected eye movement signal, the data is first cleaned using a preset noise removal algorithm. Commonly used noise removal techniques include low-pass filtering and high-pass filtering methods, among which low-pass filtering can be used to remove high-frequency noise in the eye movement signal, while high-pass filtering is used to remove low-frequency drift; further, a smoothing algorithm (such as the weighted moving average method) can also be applied to smooth the eye movement signal to reduce interference caused by equipment errors or environmental changes; in addition, the characteristics of the eye movement signal (such as the degree of concentration of the gaze point) can be combined, and eye movement signals with excessive fluctuations can be eliminated by setting a threshold to ensure that the remaining eye movement signal is stable and consistent with the user's actual eye movement behavior; other methods can also be used in other embodiments, which will not be repeated here.
[0025] In specific implementation, the following method can be used to identify the target user's gaze focus in the task display interface based on a preset gaze determination algorithm combined with the eye movement signal after noise removal, namely: based on the eye movement signal after noise removal, the gaze determination algorithm is called to process the eye movement signal to obtain the target user's gaze focus in the task display interface; wherein the processing includes extracting the target user's eye movement trajectory in each time segment, including the eye movement position coordinates and timestamp of each frame, and then calculating the eye movement speed according to the Euclidean distance of the eye movement position between consecutive frames divided by the time interval. When the speed is less than the set threshold (such as 50 pixels / second), it is regarded as stable gaze; otherwise, it is judged as a saccade or jump; at the same time, in order to determine whether the gaze point forms an effective focus, the consecutive frame gazes can be counted within a radius of 30 pixels. Whether the viewpoints are concentrated in the same area, and the degree of focus is calculated (the degree of focus can be quantified by the frequency of the gaze point falling into the area per unit time). When the degree of focus is higher than the set threshold (e.g., 80% of the frames are concentrated in the same visual area) and the gaze duration exceeds the set time threshold (e.g., 150 milliseconds), the position is determined to be a valid gaze focus. If the target user's eye movement signal shows multiple short-term gazes at a certain area (e.g., multiple revisits within 500 milliseconds), the area can also be judged as a potential gaze focus based on the scan density and the frequency of re-gaze in a short period of time. In addition, during the judgment process, the current gaze behavior can be dynamically adapted and adjusted based on historical eye movement behavior patterns. For example, based on whether the user has the habit of repeatedly scanning the same type of words, the gaze judgment threshold can be automatically corrected to avoid misjudgment or missed judgment.
[0026] In some embodiments, reference Figure 2 As shown, this figure is a schematic diagram of the application scenario of the eye tracking system for aphasia rehabilitation training shown in some embodiments of the present application. The figure includes three main components: an acquisition device, a server and a data storage device. The acquisition device is responsible for collecting the gaze time data and time response sequence of the target user, and sending the collected gaze time data and time response sequence to the server through the communication network. The eye tracking system for aphasia rehabilitation training is running in the server, and the server stores the processing results in the data storage device and visualizes them.
[0027] In step 102, when the gaze focus continues to stay on the same word, candidate words matching the current word are arranged in a circular manner in the task display interface, and the target user's gaze time data on the candidate words are recorded, and the semantic diffusion characteristics representing the target user's gaze time data for the current word are extracted from the gaze time data.
[0028] In a specific implementation, when the gaze focus continuously remains on the same word, candidate words matching the current word are arranged in a circular manner in the task display interface. This can be achieved by automatically triggering a candidate word arrangement mechanism upon recognizing that the target user's gaze focus continuously remains on an area corresponding to a word in the task display interface, and the duration of the gaze stay exceeds a preset recognition threshold (e.g., 300 milliseconds). The candidate word arrangement mechanism includes: first, performing an extended match based on synonyms or related words in the semantic space of the word corresponding to the gaze focus. A preset word vector embedding model (e.g., Word2Vec, GloVe) can be used to retrieve multiple words whose semantic similarity with the corresponding word exceeds a preset threshold (e.g., a semantic similarity greater than 0.75); then, selecting an appropriate number (e.g., 5 to 8) of representative words from all the words as candidate words; then, dynamically drawing a circular layout area on the interface with the central word as the center of the circle, and arranging the candidate words on the circular trajectory in a clockwise direction at equal intervals to ensure that the display position of each candidate word is evenly distributed within the user's field of view to avoid visual clustering interference.
[0029] It should be noted that in order to facilitate subsequent gaze time statistics, the area corresponding to each candidate word in this application has a clear gaze judgment boundary, and ensures that the circular word layout is aligned with the eye movement recording timing, thereby providing accurate spatial positioning and time basis for subsequent semantic diffusion feature extraction.
[0030] In specific implementation, recording the target user's gaze time data on the candidate words can be achieved in the following manner, namely: after completing the circular arrangement of the candidate words matching the current gaze word in the task display interface, the target user's gaze time data on the candidate words can be recorded through the integrated eye tracking device; more specifically, the display area corresponding to the candidate words in the interface can be bound to the coordinate system, and the gaze determination boundary area of each candidate word can be delineated, so as to monitor in real time whether the gaze focus falls into any candidate word area; when it is identified that the user's gaze focus enters the boundary area of a candidate word, the gaze start time of the candidate word is immediately recorded, and the end time is recorded when the gaze point leaves, so as to calculate the effective duration of the gaze event through the time difference; for multiple gaze behaviors, the gaze durations of all the candidate words are superimposed to form the cumulative gaze time of the candidate word.
[0031] It should be noted that in order to ensure the accuracy and reliability of the data, a minimum gaze time threshold (for example, 80 milliseconds) can be set to exclude unintentional rapid glances; in addition, the timestamp and sequence number of each gaze event are recorded synchronously, and the above information is bound to the semantic identification and arrangement position of the candidate vocabulary and stored in a local or cloud database; the eye tracking device can be a commercial near-infrared optical eye tracker, a front-facing camera facial tracking device or an integrated wearable device, and the data sampling rate is not less than 60 Hz.
[0032] In some embodiments, extracting the semantic diffusion feature representing the target user's attitude toward the current word from the gaze time data may be achieved by using the following steps, namely: Pre-build semantic diffusion model; extracting the target user's gaze time distribution on the candidate vocabulary from the gaze time data; The gaze time distribution is input into the semantic diffusion model, and then the semantic diffusion characteristics of the target user for the current word are output.
[0033] In specific implementation, the pre-construction of the semantic diffusion model can be achieved in the following ways. For example, a semantic diffusion model can be pre-constructed based on historical user data or experimental results, wherein the model can infer the user's semantic understanding level based on the gaze time distribution; more specifically, the model can be implemented by jointly training based on the lexical semantic space and the user's historical gaze time data, namely: first, a pre-trained word vector model is constructed in a large-scale text corpus, common words are embedded in a high-dimensional semantic space, and the lexical semantic distance matrix is obtained by calculating the cosine similarity between word vectors; then, a training sample set is constructed based on the gaze time data of a large number of users collected in historical aphasia rehabilitation training tasks, each sample consists of "central word + a group of candidate words + corresponding gaze time distribution", and the semantic diffusion level label presented by each group of samples is annotated (for example, the semantic diffusion level labels are divided into "strong diffusion", "medium diffusion", and "weak diffusion" according to expert annotation, Gaussian distribution fitting or clustering); In terms of feature design, each sample not only contains the gaze time vector of the candidate word, but also integrates the semantic distance between words, the spatial variance of the user's gaze trajectory, the gaze sequence characteristics, and the user's basic cognitive parameters such as age, training stage, etc.; among them, the training stage adopts a supervised learning method, for example: a model structure that supports semantic regression capabilities (such as: multi-layer perceptron and attention mechanism network, etc.) can be selected, with the gaze feature vector and the semantic distance as input, and the semantic diffusion level or semantic diffusion value as output. The optimization process adopts a labeled classification loss or a regression loss function based on time distribution deviation; it should be noted that in order to improve the generalization ability of the model, this application introduces cross-validation, feature perturbation, and normalization and enhancement of user features during the training process. After the model training is completed, the semantic diffusion features can be generalized and predicted in the user's task execution, which is used to judge the user's semantic association breadth and word meaning understanding transfer ability. In other embodiments, other methods can also be used to construct a semantic diffusion model, which will not be repeated here.
[0034] In specific implementation, the following method can be used to extract the target user's gaze time distribution on the candidate words from the gaze time data, namely: classify and organize the gaze time data corresponding to each candidate word, count the total gaze time of each candidate word, and normalize the total gaze time of all candidate words, and express the gaze time of each candidate word as a relative proportion value or a normalized distribution weight, so as to construct a set of gaze time distribution vectors (the gaze time distribution vector is: gaze time distribution) that reflect the target user's attention level to each candidate word in the current task. In other embodiments, other methods can also be used for implementation, which are not limited here.
[0035] It should be noted that the semantic diffusion feature described in this application is a numerical value reflecting the semantic association range and semantic transfer ability of the target user. As a preferred embodiment, the gaze time distribution is input into the semantic diffusion model, and then the semantic diffusion feature of the target user for the current word is output. It can be achieved in the following way, namely: the gaze time distribution is input into the semantic diffusion model, and the model automatically identifies the diffusion law between the time concentration degree and the semantic distance. For example: when the user has a long and balanced gaze behavior on multiple semantically adjacent words, the model outputs "strong diffusion"; on the contrary, if the user mainly focuses on a small number of words with a close semantic distance, the model outputs "weak diffusion"; finally, the semantic diffusion feature is encoded as a numerical value reflecting the user's semantic association range and semantic transfer ability for output. In other embodiments, other methods can also be used to achieve it, which will not be repeated here.
[0036] In step 103, when the gaze focus is in a scanning mode in the task display interface, the time response sequence of the target user to all the scanned words is recorded, and then the semantic recognition gradient of the target user to the initial rehabilitation training speech is determined based on the response characteristics of the words corresponding to the initial rehabilitation training speech in the time response sequence.
[0037] In a specific implementation, when the gaze focus is in a scanning mode in the task display interface, recording the target user's time response sequence to all scanned words can be achieved in the following manner, namely: when the target user's gaze focus is in a scanning mode in the task display interface, recording the timestamp corresponding to each scanned word, wherein the timestamp includes: the start time of the gaze, the end time, the duration, the time delay of the first gaze, and the number of repeated gazes during the scanning process, and then combining to construct a time response unit corresponding to each word; wherein the time delay of the first gaze is used to measure the response lag of the first recognition of the word, and the number of repeated gazes is used to reflect the user's possible hesitation or recognition difficulty for the word, and all time response units are sorted according to the order of the gaze start time to construct a complete time response sequence.
[0038] In some embodiments, determining the target user's semantic recognition gradient for the initial rehabilitation training speech based on the response characteristics of the vocabulary corresponding to the initial rehabilitation training speech in the time response sequence can be achieved by using the following steps, namely: Extracting the time response unit of each word in the time response sequence in the initial rehabilitation training speech; The corresponding lexical response feature vector is constructed based on the fixation start time, end time, duration, time delay of the first fixation and the number of repeated fixations during the scanning process in each time response unit; All vocabulary response feature vectors are input into a preset recognition gradient evaluation model to output the semantic recognition gradient of the target user for the initial rehabilitation training speech.
[0039] In specific implementation, the following method can be used to construct the corresponding word response feature vector based on the gaze start time, end time, duration, time delay of the first gaze, and the number of repeated gazes during the scanning process in each time response unit, namely: the five feature data of the gaze start time, end time, duration, time delay of the first gaze, and the number of repeated gazes during the scanning process in each time response unit are standardized, including: normalization of the time field, range compression processing of the number of repeated gazes, etc., so as to enhance the adaptability of the subsequent model to different feature dimensions. Subsequently, the five feature data after standardization are combined into a high-dimensional response feature vector in a fixed order, wherein each dimension in the vector corresponds to a certain feature data. The vector is used to express the user's eye movement response pattern to a specific word in the speech recognition task, and can more comprehensively reflect the user's response intensity and recognition stability in the semantic understanding process.
[0040] In specific implementation, all vocabulary response feature vectors are input into a preset recognition gradient evaluation model, and the output of the target user's semantic recognition gradient for the initial rehabilitation training speech can be achieved by the following steps, namely: each constructed vocabulary response feature vector is input into the recognition gradient evaluation model in sequence according to the sequence of the corresponding time response units in the time response sequence. The model can use a pre-trained multidimensional behavior scoring network or a lightweight regression neural structure, combined with the semantic understanding performance samples that have been annotated by the recognition gradient evaluation model during the training phase, to perform response scoring on the input vocabulary response feature vector. Finally, based on the recognition gradient scoring results of all vocabulary (i.e., vocabulary recognition gradient values), weighted fusion is performed to generate the semantic recognition gradient score of the target user for the entire initial rehabilitation training speech.
[0041] It should be noted that the recognition gradient evaluation model in this application can be preset in a supervised learning-based manner, and a large amount of collected rehabilitation training sample data is usually used as the training basis, where each set of training sample data includes: the target user's response characteristics to multiple words during the eye movement scanning process when performing rehabilitation training speech (such as: gaze duration, first gaze delay, number of repeated scans, etc.) and the actual semantic recognition ability label shown in the semantic understanding test; during the training process, a lightweight model structure with good feature expression ability (such as: multi-layer perceptron, recurrent neural network or support vector machine, etc.) can be selected, and the user vocabulary response feature vector is used as input, and the semantic recognition score result is used as output to construct a mapping relationship of the model; preferably, the training process can introduce a sample balancing mechanism, feature normalization processing and model regularization strategy to enhance the generalization ability of the model under small samples or unbalanced data; after the model training is completed, it can be deployed as a preset evaluation model in the rehabilitation training system for online real-time evaluation of the semantic recognition gradient of the target user.
[0042] It should be noted that the semantic recognition gradient in this application is used to quantify the cognitive load level of the target user at different semantic depth levels.
[0043] In step 104 , a response feedback analysis is performed on the speech recognition of the target user during the initial rehabilitation training based on the semantic diffusion feature and the semantic recognition gradient to obtain the semantic response deviation of the target user at the current stage.
[0044] In some embodiments, reference Figure 3 As shown in FIG, this figure is a flow chart of determining the semantic response deviation in some embodiments of the present application. In this embodiment, based on the semantic diffusion feature and the semantic recognition gradient, the speech recognition response feedback analysis of the target user during the initial rehabilitation training is performed to obtain the semantic response deviation of the target user at the current stage. The following steps can be used to achieve this, namely: In step 1031, the corresponding word when the focus of attention continuously stays on the same word is regarded as the stay word; In step 1032, semantic diffusion features of the target user for the current word and other words that remain are obtained respectively; In step 1033, all the remaining words are matched with the initial rehabilitation training speech of the current stage to obtain a plurality of matching words and a plurality of non-matching words; In step 1034, semantic response weights are allocated to all matching words and corresponding semantic diffusion features, all non-matching words and the target user's semantic recognition gradient for the initial rehabilitation training speech based on a weight allocation mechanism, and then the semantic response deviation of the target user at the current stage is calculated based on the weight factors corresponding to all matching words, non-matching words and semantic recognition gradients.
[0045] It should be noted that in this application, if it is identified that the target user's gaze focus continues to stay on other words in the current stage, the other words are also treated in step 102 when the gaze focus continues to stay on the same word, and the candidate words matching the current word are arranged in a circular manner in the task display interface, and the target user's gaze time data on the candidate words are recorded. A specific implementation method is extracted from the gaze time data to characterize the semantic diffusion characteristics of the target user for the current word, thereby obtaining the semantic diffusion characteristics of the target user for other remaining words in the current stage.
[0046] In specific implementation, all the staying words are matched with the initial rehabilitation training voice of the current stage, and multiple matching words and multiple non-matching words can be obtained by the following method, namely: each word in the initial rehabilitation training voice of the current stage is obtained, and all the staying words are matched with each word in the initial rehabilitation training voice of the current stage. If the staying word is the same as a word in the initial rehabilitation training voice of the current stage, then the word is used as a matching word. If the staying word is different from all the words in the initial rehabilitation training voice of the current stage, then the word is used as a non-matching word, thereby obtaining multiple matching words and multiple non-matching words.
[0047] In specific implementation, semantic response weights are allocated to all matching words and corresponding semantic diffusion features, all non-matching words and the semantic recognition gradient of the target user to the initial rehabilitation training speech based on the weight allocation mechanism. This can be achieved in the following way, namely: for matching words, positive response weights are assigned according to the concentration in their semantic diffusion features (such as: most gazes are concentrated on candidate words with high similarity) to reflect their guiding role as the focus of attention in the current semantic recognition process; for non-matching words, relatively low or negative weights are set based on the degree of dispersion and interference of their semantic diffusion paths to reflect the possible deviation tendency of users in semantic recognition; for semantic recognition gradients, the semantic recognition gradient is used as an indicator of the user's overall semantic perception ability at the current stage, and a weight setting mechanism is independently introduced. This mechanism can be based on historical performance stability, current language The confidence interval of the sound task complexity and the recognition gradient is used to assign a high or low weighted value to the gradient; wherein, the semantic recognition gradient can also be introduced as a weight factor of the overall semantic perception ability, and the weight distribution results formed by the above-mentioned matching and non-matching words are dynamically adjusted to enhance the connectivity of the importance weights of matching words and non-matching words, that is: when the recognition gradient is high, the weight contribution of matching words is enhanced and the influence of non-matching words is suppressed; when the recognition gradient is low, the participation of non-matching words in the response deviation is increased, reflecting the risk characteristics of unstable semantic recognition ability of the user; finally, through the fusion distribution of the above-mentioned multi-dimensional weight factors, the weights corresponding to the semantic recognition gradients of all matching words, all non-matching words and the target user for the initial rehabilitation training speech are constructed. In other embodiments, other methods can also be used to achieve semantic response weight distribution, which will not be repeated here.
[0048] In specific implementation, the semantic response deviation of the target user at the current stage is calculated based on the weight factors corresponding to all matching words, non-matching words and semantic recognition gradients, which can be achieved in the following way, namely: weighted fusion is performed based on all matching words, all non-matching words and the semantic recognition gradient in combination with the corresponding weights, and the output value is used as the value of the semantic response deviation of the target user at the current stage. In other embodiments, other methods can also be used for implementation, which will not be repeated here.
[0049] It should be noted that the semantic response deviation in this application is used to quantify the semantic attention shift or understanding error generated by the user during the speech comprehension process.
[0050] In step 105, the speech recognition ability and gaze behavior of the target user are correlated and mapped by combining the gaze features of each word in the target user's corresponding historical eye movement trajectory data and the semantic structure features of the corresponding rehabilitation training speech through a pre-built interaction evaluation model to obtain the gaze correlation relationship of the target user under different semantic structure features.
[0051] It should be noted that the gaze association relationship in this application represents the influence relationship between the playback sequence of each word in the rehabilitation training speech and the degree of the target user's gaze on each word in the rehabilitation training semantics.
[0052] In some embodiments, the speech recognition ability and gaze behavior of the target user are correlated and mapped by combining the gaze features of each word in the target user's corresponding historical eye movement trajectory data and the semantic structure features of the corresponding rehabilitation training speech through a pre-built interactive evaluation model. Obtaining the gaze correlation relationship of the target user under different semantic structure features can be achieved by the following steps, namely: Obtain the target user's eye movement trajectory data corresponding to multiple historical rehabilitation training tasks, and extract the gaze features corresponding to each word from it; Acquire rehabilitation training speech corresponding to each historical rehabilitation training task, and then determine the semantic structure characteristics of each rehabilitation training speech; The gaze features corresponding to each extracted word and the semantic structure features of each rehabilitation training speech are aligned on a word-by-word basis to construct data pairs containing the relationship between the word playback order and gaze intensity in multiple training tasks; The above data pairs are input into the pre-built interaction evaluation model to obtain the gaze association relationship of the target user under different semantic structure features.
[0053] In a specific implementation, the eye movement trajectory data corresponding to the target user in multiple historical rehabilitation training tasks is obtained, and the gaze features corresponding to each word are extracted therefrom. This can be achieved in the following manner, namely: the eye movement trajectory data of the target user is recorded by a preset eye tracking device during the process of the target user performing the historical rehabilitation training tasks. The eye movement trajectory data includes the target user's gaze time and gaze response time on each word, wherein the gaze time is the difference between the target user's gaze start time and gaze end time on the word, and the gaze response time is the difference between the playback time and the target user's gaze start time on the word when playing the word in the rehabilitation training voice; the gaze time and gaze response time corresponding to each word are combined into the gaze feature of each word.
[0054] It should be noted that the gaze feature in this application represents the user's gaze intensity on each word.
[0055] In specific implementation, obtaining the rehabilitation training speech corresponding to each historical rehabilitation training task can be achieved in the following manner, namely: the text of the rehabilitation training speech corresponding to each historical rehabilitation training task can be retrieved from the text library of the rehabilitation training speech, or the text of the rehabilitation training speech corresponding to each historical rehabilitation training task can be obtained through a speech recognition device; it should be noted that the semantic structure features described in this application reflect the different semantic logical relationships caused by the different order of appearance of each word in the rehabilitation training speech. As a preferred embodiment, the semantic structure features of each rehabilitation training speech are determined, for example: the existing dependency syntactic analysis tools (such as: Stanford Parser, spaCy) analyzes the speech text, extracts syntactic dependency relationships such as subject-predicate, verb-object, and modifiers between words, and forms a semantic dependency graph with words as nodes and grammatical relationships as edges, thereby constructing a semantic structure path within the speech, and using the constructed semantic structure path as the semantic structure feature corresponding to the rehabilitation training speech. For example, a semantic graph can be constructed based on the semantic similarity, co-occurrence frequency or contextual co-reference relationship between each word in the rehabilitation training speech, and then a graph algorithm (such as the shortest path, centrality analysis, etc.) is used to identify the semantic structure features of the rehabilitation training speech. In other embodiments, other methods can also be used for implementation, which is not limited here.
[0056] In specific implementation, the gaze features corresponding to each extracted word and the semantic structure features of each rehabilitation training speech are aligned in units of words, and a data pair containing the relationship between the playback order of words and the gaze intensity in multiple training tasks is constructed. This can be achieved in the following way, namely: first, based on the word order in the speech text, the temporal position index of each word appearing in the speech is extracted in turn, and used as the word playback order identifier. At the same time, combined with the eye movement trajectory data recorded by the target user in the corresponding rehabilitation training task, a match is made to determine whether each word in the speech text is being gazed at during playback. If the user's gaze point is on the corresponding word, If a word stays in the display area for more than a set threshold (such as 100 milliseconds), the gaze feature of the word is recorded. Subsequently, each word in the voice playback order is mapped one by one with its corresponding gaze feature to construct an initial correspondence between the word playback order and the gaze intensity. In other embodiments, the word gaze feature can also be structurally modified or normalized according to the semantic dependency, role label or context aggregation degree in the semantic structure feature, so that the gaze behavior can accurately reflect the user's visual attention path when understanding the semantic structure. Finally, data pairs corresponding to "word playback order" and "gaze intensity" are formed in multiple historical training tasks.
[0057] In a specific implementation, the aforementioned data pairs are input into a pre-built interaction evaluation model to obtain the target user's gaze association relationships under different semantic structural features. This can be achieved in the following manner: first, based on the previously constructed data pairs, the input form of the interaction evaluation model is set to be a serialized vocabulary playback order index and its corresponding gaze intensity feature vector. Subsequently, a structural model capable of learning temporal and sequential dependencies, such as a multi-layer perceptron network with a temporal attention mechanism or a lightweight recurrent neural network with sequence-to-value mapping, is used. The sequence data constructed in each training task is used as training input. The model learns the correspondence between the position of vocabulary during speech playback and the user's gaze response, thereby extracting the weight of the influence of playback order on the degree of gaze. During the training process, the model aims to minimize the error between the predicted gaze intensity and the actual gaze intensity. Temporal precedence weights, contextual semantic dependencies, etc. can be introduced as additional features to enhance the sequential modeling capability. Finally, the trained model can be used as an evaluation engine to achieve quantitative evaluation of the target user's gaze association relationships, thereby obtaining the target user's gaze association relationships under different semantic structural features.
[0058] It should be noted that during the training process of the interactive evaluation model, semantic structure features can also be used as auxiliary input to model the regulatory effect of semantic logic on gaze behavior.
[0059] In step 106, the cognitive efficacy index of the target user at the current stage is determined based on the semantic response deviation and all gaze association relationships, and the vocabulary difficulty of the rehabilitation training speech at the next stage is adaptively adjusted based on the cognitive efficacy index.
[0060] In some embodiments, determining the cognitive efficacy index of the target user at the current stage based on the semantic response deviation and all gaze association relationships can be achieved by using the following steps, namely: Based on the gaze association relationship of the target user under different semantic structure features, the training evaluation coefficients corresponding to different semantic structure features are set; Determine the semantic structure characteristics of the initial rehabilitation training speech at the current stage; Matching the semantic structure features of the initial rehabilitation training speech with the semantic structure features of the target user in the historical rehabilitation training task, thereby obtaining a training evaluation coefficient corresponding to the semantic response deviation; The cognitive efficacy index of the target user at the current stage is calculated based on the semantic response deviation and the corresponding training evaluation coefficient.
[0061] It should be noted that the training evaluation coefficient in this application reflects the historical performance ability of the target user in the gaze dimension under different semantic structure conditions, and can reflect the target user's ability to understand and adapt to a specific semantic structure. It is a continuous value and is set between 0.5 and 1.5. As a preferred embodiment, the training evaluation coefficient corresponding to different semantic structure features is set based on the gaze association relationship of the target user under different semantic structure features. This can be achieved in the following way, namely: when the gaze association relationship is manifested as a high degree of consistency between the gaze behavior and the vocabulary playback, a small time delay, and a low sequence deviation, it indicates that the user is under this type of semantic structure. The semantic understanding performance is good, and the corresponding training evaluation coefficient can be set between 1.2 and 1.5; when there is a certain degree of sequence deviation or delay in the gaze behavior but the overall trend is still regular, the corresponding evaluation coefficient can be set between 0.9 and 1.2; when the gaze behavior is weakly related to the playback order, and there are features such as frequent jumps, delays or omissions, the corresponding training evaluation coefficient is set between 0.5 and 0.9; finally, the training evaluation coefficient is used as a weight factor for calculating the user cognitive efficacy index in combination with the semantic response deviation in the subsequent stage, so as to reflect the user's historical performance ability in the gaze dimension under different semantic structure conditions.
[0062] In specific implementation, the semantic structure features of the initial rehabilitation training speech at the current stage can be determined by the following methods: Parser, spaCy) analyzes the speech text of the initial rehabilitation training speech, extracts syntactic dependency relationships such as subject-predicate, verb-object, and modifier between words, and forms a semantic dependency graph with words as nodes and grammatical relationships as edges, thereby constructing a semantic structure path within the speech. The constructed semantic structure path is used as the semantic structure feature corresponding to the rehabilitation training speech. For example, a semantic graph can be constructed based on the semantic similarity, co-occurrence frequency, or contextual coreference relationship between each word in the initial rehabilitation training speech, and then a graph algorithm (such as shortest path, centrality analysis, etc.) is used to identify the semantic structure features of the initial rehabilitation training speech. The semantic structure features of the initial rehabilitation training speech are matched with the semantic structure features of the target user in historical rehabilitation training tasks to obtain a training evaluation coefficient corresponding to the semantic response deviation. This can be achieved by matching the semantic structure features of the initial rehabilitation training speech with the semantic structure features of the target user in historical rehabilitation training tasks based on the semantic structure path, and using the training evaluation coefficient corresponding to the semantic structure feature with the highest similarity to the semantic structure path of the initial rehabilitation training speech as the training evaluation coefficient of the initial rehabilitation training speech.
[0063] In specific implementation, the cognitive efficacy index of the target user at the current stage is calculated based on the semantic response deviation and the corresponding training evaluation coefficient, that is: the numerical value of the semantic response deviation is multiplied by the corresponding training evaluation coefficient, and the value obtained by the product is used as the cognitive efficacy index of the target user at the current stage.
[0064] It should be noted that the cognitive efficacy index in this application represents the comprehensive performance score of the target user's semantic recognition ability in the current stage of rehabilitation training speech tasks, reflecting the stability and adaptability of their understanding of the training content under the current semantic structure.
[0065] In specific implementation, adaptive adjustment of the vocabulary difficulty of the next stage of rehabilitation training speech based on the cognitive efficacy index can be achieved in the following manner, namely: judging the level range of the semantic understanding ability of the target user according to the numerical value of the cognitive efficacy index of the current stage, and the level range can be divided into multiple levels, such as: "low recognition level", "medium recognition level" and "high recognition level"; then, according to the recognition level of the target user, the rehabilitation training tasks matching the level are screened out from the speech material library, wherein the low recognition level corresponds to the basic semantic structure and high-frequency vocabulary, the medium recognition level corresponds to common conjunctions, synonymous word groups and sentence structures of medium complexity, and the high recognition level corresponds to abstract concept vocabulary, long sentence nested structure and corpus content with higher requirements for semantic generalization ability; finally, the next stage of rehabilitation training speech content matching the user's current semantic recognition ability is generated and output to the speech playback module for use in the training task, thereby realizing adaptive vocabulary difficulty adjustment based on the cognitive efficacy index.
[0066] It should be noted that, preferably, in some embodiments, in order to improve vocabulary adaptability, in the project of screening out rehabilitation training tasks that match the level from the voice material library, it is also possible to further combine the target user's recognition performance of specific semantic category vocabulary in historical training, and dynamically adjust the core vocabulary composition in the next stage of rehabilitation training speech. For example: if the user's recognition of "spatial position related" vocabulary is weak in previous training, then the training of this type of vocabulary will be given priority under the condition of high evaluation index.
[0067] In addition, in another aspect of the present application, in some embodiments, the present application provides an eye tracking system for aphasia rehabilitation training, the system comprising a vocabulary difficulty adaptive adjustment unit, reference Figure 4 , which is a schematic diagram of exemplary hardware and / or software of a vocabulary difficulty adaptive adjustment unit according to some embodiments of the present application. The vocabulary difficulty adaptive adjustment unit 400 includes: a recognition module 401, a processing module 402, and an execution module 403, which are described as follows: Identification module 401, in this application, is mainly used to play the initial rehabilitation training voice of the current stage to the target user and identify the target user's gaze focus in the task display interface; Processing module 402, in this application, is mainly used to arrange candidate words matching the current word in a circular manner in the task display interface when the gaze focus continues to remain on the same word, record the target user's gaze time data on the candidate words, and extract semantic diffusion features representing the target user's attention to the current word from the gaze time data; The processing module 402 of the present application is further configured to record the target user's time response sequence to all scanned words when the gaze focus is in a scanning mode in the task display interface, and then determine the target user's semantic recognition gradient for the initial rehabilitation training speech based on the response characteristics of the words corresponding to the initial rehabilitation training speech in the time response sequence; The processing module 402 in the present application is further configured to perform response feedback analysis on the speech recognition of the target user during the initial rehabilitation training based on the semantic diffusion feature and the semantic recognition gradient to obtain the semantic response deviation of the target user at the current stage; The processing module 402 in the present application is further configured to perform correlation mapping between the speech recognition ability and gaze behavior of the target user by combining the gaze features of each word in the target user's corresponding historical eye movement trajectory data and the semantic structure features of the corresponding rehabilitation training speech through a pre-built interactive evaluation model, thereby obtaining the gaze correlation relationship of the target user under different semantic structure features; Execution module 403, in this application, execution module 403 is mainly used to determine the cognitive efficiency index of the target user in the current stage based on the semantic response deviation and all gaze association relationships, and then adaptively adjust the vocabulary difficulty of the rehabilitation training speech in the next stage based on the cognitive efficiency index.
[0068] Each module in the aforementioned eye tracking system for aphasia rehabilitation training can be implemented in whole or in part through software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a computer device's memory in the form of software, so that the processor can call and execute the corresponding operations of each module.
[0069] In addition, in one embodiment, the present application provides a computer device, which may be a server, and its internal structure diagram may be as follows: Figure 5As shown. The computer device includes a processor, a memory and a network interface connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store vocabulary difficulty adaptive adjustment data. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, the vocabulary difficulty adaptive adjustment method in the above-mentioned aphasia rehabilitation training process is implemented.
[0070] Those skilled in the art will understand that Figure 5 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0071] In one embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the embodiment of the method for adaptively adjusting vocabulary difficulty in the aphasia rehabilitation training process are implemented.
[0072] In one embodiment, a computer-readable storage medium is provided, storing a computer program. When the computer program is executed by a processor, the steps of the embodiment of the method for adaptively adjusting vocabulary difficulty in the aphasia rehabilitation training process are implemented.
[0073] In one embodiment, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the steps of the embodiment of the method for adaptively adjusting vocabulary difficulty during aphasia rehabilitation training.
[0074] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory or optical memory, etc. Volatile memory may include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).
[0075] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0076] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art could make various modifications and improvements without departing from the spirit of the present application, all of which fall within the scope of protection of the present application. Therefore, the scope of protection of the present patent application shall be determined by the appended claims.
Claims
1. A method for adaptively adjusting vocabulary difficulty during aphasia rehabilitation training, applied to an eye tracking system for aphasia rehabilitation training, wherein: When the eye tracking system plays rehabilitation training speech to the target user, the corresponding words in the task display interface are highlighted synchronously, and the method is characterized in that it includes the following steps: Play the initial rehabilitation training speech of the current stage to the target user and identify the target user's gaze focus in the task display interface; When the gaze focus remains on the same word, candidate words matching the current word are arranged in a circular pattern in the task display interface, and the target user's gaze time data on the candidate words is recorded. The semantic diffusion feature representing the target user's attention to the current word is extracted from the gaze time data; When the gaze focus is in a scanning mode in the task display interface, the target user's time response sequence to all scanned words is recorded, and then the target user's semantic recognition gradient for the initial rehabilitation training speech is determined based on the response characteristics of the words corresponding to the initial rehabilitation training speech in the time response sequence; performing response feedback analysis on the speech recognition of the target user during the initial rehabilitation training based on the semantic diffusion feature and the semantic recognition gradient to obtain the semantic response deviation of the target user at the current stage; The pre-built interactive evaluation model combines the gaze characteristics of each word in the target user's corresponding historical eye movement trajectory data with the semantic structure characteristics of the corresponding rehabilitation training speech to map the target user's speech recognition ability and gaze behavior, and obtains the target user's gaze association relationship under different semantic structure characteristics; The cognitive efficacy index of the target user at the current stage is determined according to the semantic response deviation and all gaze association relationships, and the vocabulary difficulty of the rehabilitation training speech at the next stage is adaptively adjusted based on the cognitive efficacy index.
2. The method according to claim 1, wherein Identifying the target user's gaze focus in the task display interface specifically includes: Acquiring eye movement signals of the target user when performing the initial rehabilitation training speech; performing noise removal processing on the eye movement signal to obtain an eye movement signal after noise removal; The gaze focus of the target user in the task display interface is identified based on a preset gaze determination algorithm combined with the eye movement signal after noise removal.
3. The method according to claim 1, wherein Extracting semantic diffusion features representing the target user's attitude towards the current word from the gaze time data specifically includes: Pre-build semantic diffusion model; extracting the target user's gaze time distribution on the candidate vocabulary from the gaze time data; The gaze time distribution is input into the semantic diffusion model, and then the semantic diffusion characteristics of the target user for the current word are output.
4. The method according to claim 1, wherein Determining the semantic recognition gradient of the target user to the initial rehabilitation training speech according to the response characteristics of the vocabulary corresponding to the initial rehabilitation training speech in the time response sequence specifically includes: Extracting the time response unit of each word in the time response sequence in the initial rehabilitation training speech; The corresponding lexical response feature vector is constructed based on the fixation start time, end time, duration, time delay of the first fixation and the number of repeated fixations during the scanning process in each time response unit; All vocabulary response feature vectors are input into a preset recognition gradient evaluation model to output the semantic recognition gradient of the target user for the initial rehabilitation training speech.
5. The method according to claim 1, wherein Based on the semantic diffusion feature and the semantic recognition gradient, the speech recognition response feedback analysis of the target user during the initial rehabilitation training is performed to obtain the semantic response deviation of the target user at the current stage, specifically including: The word corresponding to the time when the focus of attention continuously stays on the same word is regarded as the stay word; Obtain the semantic diffusion features of the target user for the current word and other words that remain; Matching all the retained words with the initial rehabilitation training speech of the current stage to obtain multiple matching words and multiple non-matching words; Based on the weight distribution mechanism, semantic response weights are distributed to all matching words and corresponding semantic diffusion features, all non-matching words and the target user's semantic recognition gradient for the initial rehabilitation training speech, and then the semantic response deviation of the target user at the current stage is calculated based on the weight factors corresponding to all matching words, non-matching words and semantic recognition gradients.
6. The method according to claim 1, wherein The pre-built interactive evaluation model combines the gaze features of each word in the target user's historical eye movement data with the semantic structure features of the corresponding rehabilitation speech to map the target user's speech recognition ability and gaze behavior. The gaze association relationships of the target user under different semantic structure features are as follows: Obtain the target user's eye movement trajectory data corresponding to multiple historical rehabilitation training tasks, and extract the gaze features corresponding to each word from it; Acquire rehabilitation training speech corresponding to each historical rehabilitation training task, and then determine the semantic structure characteristics of each rehabilitation training speech; The gaze features corresponding to each extracted word and the semantic structure features of each rehabilitation training speech are aligned on a word-by-word basis to construct data pairs containing the relationship between the word playback order and gaze intensity in multiple training tasks; The above data pairs are input into the pre-built interaction evaluation model to obtain the gaze association relationship of the target user under different semantic structure features.
7. The method according to claim 1, wherein Determining the cognitive efficacy index of the target user at the current stage based on the semantic response deviation and all gaze association relationships specifically includes: Based on the gaze association relationship of the target user under different semantic structure features, the training evaluation coefficients corresponding to different semantic structure features are set; Determine the semantic structure characteristics of the initial rehabilitation training speech at the current stage; Matching the semantic structure features of the initial rehabilitation training speech with the semantic structure features of the target user in the historical rehabilitation training task, thereby obtaining a training evaluation coefficient corresponding to the semantic response deviation; The cognitive efficacy index of the target user at the current stage is calculated based on the semantic response deviation and the corresponding training evaluation coefficient.
8. An eye tracking system for aphasia rehabilitation training, the system comprising a vocabulary difficulty adaptive adjustment unit, characterized in that: The vocabulary difficulty adaptive adjustment unit includes: A recognition module is used to play the initial rehabilitation training speech of the current stage to the target user and identify the target user's gaze focus in the task display interface; a processing module configured to, when the gaze focus remains on the same word, arrange candidate words matching the current word in a circular manner in the task display interface, record the target user's gaze time data on the candidate words, and extract semantic diffusion features representing the target user's attention to the current word from the gaze time data; The processing module is further configured to record, when the gaze focus is in a scanning mode in the task display interface, a time response sequence of the target user to all scanned words, and then determine the target user's semantic recognition gradient for the initial rehabilitation training speech based on the response characteristics of the words corresponding to the initial rehabilitation training speech in the time response sequence; The processing module is further configured to perform response feedback analysis on the speech recognition of the target user during the initial rehabilitation training based on the semantic diffusion feature and the semantic recognition gradient to obtain the semantic response deviation of the target user at the current stage; The processing module is further configured to perform correlation mapping between the target user's speech recognition ability and gaze behavior by combining the gaze features of each word in the target user's corresponding historical eye movement trajectory data and the semantic structure features of the corresponding rehabilitation training speech through a pre-built interactive evaluation model, thereby obtaining the target user's gaze correlation relationship under different semantic structure features; The execution module is used to determine the cognitive efficacy index of the target user in the current stage according to the semantic response deviation and all gaze association relationships, and then adaptively adjust the vocabulary difficulty of the rehabilitation training speech in the next stage based on the cognitive efficacy index.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method for adaptively adjusting vocabulary difficulty in the aphasia rehabilitation training process according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method for adaptively adjusting vocabulary difficulty in an aphasia rehabilitation training process according to any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Early dementia screening system and method based on eye movement tracking technology
CN122030886A