Intelligent reading assistance method, apparatus, device, storage medium, and program product
By acquiring users' eye-tracking behavior data, and using eye-tracking models and large language models to identify areas of difficulty in reading and provide supplementary knowledge, the problem of AI interactive interfaces being unable to judge reading difficulties in real time has been solved, achieving more accurate and timely knowledge assistance and improving reading effectiveness.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA MOBILE (SUZHOU) SOFTWARE TECH CO LTD
- Filing Date
- 2026-02-12
- Publication Date
- 2026-06-09
Smart Images

Figure CN122172961A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to an intelligent reading assistance method, device, equipment, storage medium, and program product. Background Technology
[0002] Users without a professional background often encounter comprehension barriers when reading product documentation in specialized fields, requiring them to acquire relevant background knowledge for successful understanding. When faced with these knowledge gaps, users typically seek assistance from search engines. However, with the development of AI (Artificial Intelligence) technology, knowledge assistance through AI-powered interactive interfaces is gradually replacing traditional search engines. Nevertheless, the knowledge provided by AI still contains concepts that are difficult for users to understand; therefore, users must actively seek further information and ask follow-up questions to ultimately gain a clear understanding of the subject matter.
[0003] While existing solutions can quickly help users overcome difficulties reading professional texts by using AI interactive interfaces, they cannot accurately and timely provide knowledge assistance because users need to encounter problems, continuously input questions, and obtain relevant knowledge to resolve their confusion through in-depth questioning. Summary of the Invention
[0004] To address the problems existing in the prior art, embodiments of the present invention provide an intelligent reading assistance method, device, equipment, storage medium, and program product, which can improve the accuracy and timeliness of knowledge assistance and enhance the reading assistance effect.
[0005] In a first aspect, embodiments of the present invention provide an intelligent reading assistance method, comprising: Acquire eye-tracking data of users while they are reading text; Based on the eye-tracking data, a pre-trained eye-tracking model is used to predict reading difficulties and identify the difficult reading areas of the text. The first pre-trained model supplements the knowledge of reading difficulties in the aforementioned reading difficulty areas.
[0006] As an improvement to the above solution, the acquisition of eye-tracking behavior data of the user while reading text includes: The system detects the temporal and spatial metrics of the user's eyes on the text while the user is reading it. The temporal metrics include one or more of the following: fixation time, first fixation time, number of fixations, number of regressions, total fixation time, and pupil diameter. The spatial metrics include one or more of the following: fixation position, saccade amplitude, saccade angle, fixation density, fixation distribution, and saccade target selection. The eye movement behavior data is generated based on the time and spatial metrics.
[0007] As an improvement to the above solution, based on the eye-tracking behavior data, a pre-trained eye-tracking model is used to predict reading difficulties and determine the difficult reading areas of the text, including: The eye-tracking behavior data is input into the eye-tracking model to predict reading difficulties and obtain reading difficulty information of the text; wherein, the reading difficulty information includes the real-time difficulty of the text and its first probability value, the user difficulty and its second probability value, and the text difficulty and its third probability value; Based on the first probability value of the real-time difficulty corresponding to the text, the second probability value of the user difficulty, and the third probability value of the text difficulty, the words in the text are identified as having difficulty, and the difficult reading areas of the text are determined.
[0008] As an improvement to the above solution, a pre-trained first large model is used to supplement the knowledge of reading difficulties in the aforementioned reading difficulty areas, including: Determine whether the reading difficulties in the reading difficulty area are auxiliary reading content; wherein, the auxiliary reading content is generated based on the first large model; If not, the first major model is used to simplify the sentence containing the reading difficulty and generate supplementary reading content for the reading difficulty. If so, supplementary knowledge content on the reading difficulties is generated through the first major model.
[0009] As an improvement to the above scheme, the training process of the eye-tracking model includes: Several eye-tracking data samples were collected from test users as they read pre-marked regions of interest in professional texts; wherein, the regions of interest were obtained by dividing the professional vocabulary in the professional texts into regions using a second major model; Based on the eye movement behavior data samples, the pre-built eye movement tracking model is trained to obtain a trained eye movement tracking model; The eye-tracking model is constructed based on a random forest model that integrates recursive feature elimination algorithm and a bidirectional long short-term memory network model that integrates conditional random field.
[0010] As an improvement to the above scheme, based on the eye-tracking behavior data samples, a pre-built eye-tracking model is trained to obtain a trained eye-tracking model, including: By using a random forest model that incorporates a recursive feature elimination algorithm, feature filtering is performed on each of the eye-tracking behavior data samples to obtain the real-time difficulties and their first probability values, and the user difficulties and their second probability values in each of the eye-tracking behavior data samples. Based on the real-time difficulties and their first probability values, user difficulties and their second probability values in the eye-tracking behavior data samples, the text difficulties of each eye-tracking behavior data sample are predicted by the bidirectional long short-term memory network model to obtain the third probability value of the text difficulties in each eye-tracking behavior data sample. Based on the first probability value of the real-time difficulty, the second probability value of the user difficulty, and the third probability value of the text difficulty corresponding to the eye-tracking behavior data sample, the reading difficulty area of the professional text under each eye-tracking behavior data sample is determined. Based on the reading difficulties of the professional text and the pre-labeled interest areas, the eye-tracking model is iteratively optimized until the eye-tracking model converges, thus obtaining a trained eye-tracking model. During the training process of the eye-tracking model, the model performance of the eye-tracking model under different threshold combinations is verified based on the first probability value of the real-time difficulty, the second probability value of the user difficulty, and the third probability value of the text difficulty corresponding to each eye-tracking behavior data sample. The threshold combination conditions corresponding to the best model performance are then selected. The threshold combination conditions include a combination of two or three of the first probability value of the real-time difficulty, the second probability value of the user difficulty, and the third probability value of the text difficulty. The threshold combination conditions are used to help determine the reading difficulty area.
[0011] Secondly, embodiments of the present invention provide an intelligent reading assistance device, comprising: The eye-tracking data acquisition module is used to acquire eye-tracking behavior data of users while reading text; The reading difficulty region determination module is used to predict reading difficulties based on the eye movement behavior data using a pre-trained eye tracking model, and determine the reading difficulty region of the text. The knowledge supplementation module is used to supplement the reading difficulties in the reading difficulty area using the pre-trained first model.
[0012] Thirdly, embodiments of the present invention provide an intelligent reading assistance device, comprising: a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor executes the computer program to implement the intelligent reading assistance method as described in any one of the first aspects.
[0013] Fourthly, embodiments of the present invention provide a computer-readable storage medium storing a computer program, wherein the computer program, when running, controls the device where the computer-readable storage medium is located to execute the intelligent reading assistance method as described in any one of the first aspects.
[0014] Fifthly, embodiments of the present invention provide a computer program product, including a computer program or instructions, which, when executed by a processor, implement the intelligent reading assistance method as described in any one of the first aspects.
[0015] Compared to existing technologies, the intelligent reading assistance method, device, equipment, storage medium, and program product provided in this invention first acquires eye-tracking behavior data of a user while reading text; then, based on the eye-tracking behavior data, a pre-trained eye-tracking model is used to predict reading difficulties and determine the difficult reading areas of the text; subsequently, a pre-trained first-level model is used to supplement the reading difficulties in the difficult reading areas with knowledge. This invention combines eye-tracking behavior data of the user while reading text to locate the difficult reading areas of the currently read text, and uses a first-level model to supplement the difficult reading areas with professional knowledge, which can effectively improve the accuracy and timeliness of knowledge assistance and enhance the reading assistance effect. Attached Figure Description
[0016] To more clearly illustrate the technical solution of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 This is a flowchart of an intelligent reading assistance method provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of the overall process of the intelligent reading assistance method provided in the embodiments of the present invention; Figure 3 This is a functional diagram of the eye-tracking model and the first major model in the intelligent reading assistance provided in this embodiment of the invention; Figure 4 This is a structural block diagram of an intelligent reading assistance device provided in an embodiment of the present invention; Figure 5 This is a structural block diagram of an intelligent reading assistance device provided in an embodiment of the present invention. Detailed Implementation
[0018] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0019] It is understood that the various numerical designations used in the embodiments of this invention are merely for descriptive convenience and are not intended to limit the scope of this application. The order of the process numbers does not imply the order of execution; the execution order of each process should be determined by its function and internal logic.
[0020] In embodiments of the invention, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, without necessarily requiring or implying any such actual relationship or order between these entities or operations. The terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising" does not exclude the presence of additional identical elements in the process, method, article, or apparatus that includes said element. The term "a plurality or several" refers to two or more.
[0021] See Figure 1 , Figure 1 This is a flowchart of an intelligent reading assistance method provided in an embodiment of the present invention. The intelligent reading assistance method specifically includes: S11: Obtain eye movement data of the user while reading text; This includes acquiring eye-tracking data of users while reading text, including: The system detects the temporal and spatial metrics of the user's eyes on the text while the user is reading it. The temporal metrics include one or more of the following: fixation time, first fixation time, number of fixations, number of regressions, total fixation time, and pupil diameter. The spatial metrics include one or more of the following: fixation position, saccade amplitude, saccade angle, fixation density, fixation distribution, and saccade target selection. The eye movement behavior data is generated based on the time and spatial metrics.
[0022] S12: Based on the eye movement data, a pre-trained eye tracking model is used to predict reading difficulties and determine the difficult reading areas of the text; S13: Supplement the reading difficulties in the reading difficulty area by using the first pre-trained model.
[0023] It should be noted that the embodiments of the present invention are executed by a reading device integrating an eye-tracking device, such as a computer, mobile phone, large screen, or tablet. An eye-tracking device is a device that tracks eye movements, collects physiological signals of the eyes, and converts them into quantitative data. When a user reads text using the reading device, the eye-tracking device can detect whether the user's eyes are lingering on the text. If no lingering occurs, it indicates that the user is not experiencing reading difficulties and no action is needed; if lingering occurs, it indicates that the user may be experiencing reading difficulties at the point where the text is lingered. The device acquires the user's temporal and spatial indicators of the eyes and generates corresponding eye-tracking behavior data. In the embodiments of the present invention, the method for detecting whether the user's eyes are lingering on the text is not specifically limited. For example, a lingering time threshold can be set. If the user's eyes linger on the text for a time exceeding the threshold, it is considered that a lingering has occurred; otherwise, no lingering has occurred.
[0024] The acquired eye-tracking data is then input into the eye-tracking model for difficulty prediction, resulting in regions of reading difficulty encountered when reading text. The difficulty identification and prediction uses a lexical-level granularity, meaning one difficulty corresponds to one word. The word containing the difficulty, the sentence containing the difficulty, or the paragraph containing the difficulty is considered the region of reading difficulty, but this is not specifically limited in this embodiment of the invention.
[0025] The reading difficulties in the affected areas are then input into the first major model for knowledge supplementation, expanding the professional knowledge information related to these difficulties. This first major model is built on a Large Language Model (LLM) and uses a pre-built professional knowledge base as the data source for knowledge supplementation. This professional knowledge base includes professional knowledge information related to vocabulary in specific professional fields; corresponding professional knowledge information can be matched based on vocabulary from each professional field through this base.
[0026] This invention combines eye-tracking data from when a user is reading text to locate areas of difficulty in the text. For these areas, a first-class model is used to supplement the knowledge with relevant information, which can effectively improve the accuracy and timeliness of knowledge assistance and enhance the reading assistance effect.
[0027] In one optional embodiment, knowledge supplementation is performed on the reading difficulties in the reading hardship region using a pre-trained first large model, including: Determine whether the reading difficulties in the reading difficulty area are auxiliary reading content; wherein, the auxiliary reading content is generated based on the first large model; If not, the first major model is used to simplify the sentence containing the reading difficulty and generate supplementary reading content for the reading difficulty. If so, supplementary knowledge content on the reading difficulties is generated through the first major model.
[0028] In this embodiment of the invention, for the identified reading difficulty Yn, it is first determined whether it is the auxiliary reading content generated by the first model, that is, whether the reading difficulty Yn is the original vocabulary in the text or the new vocabulary output by the model to assist in text understanding (i.e., the new vocabulary generated by AI, referred to as AI-generated content). When the reading difficulty Yn is not part of the auxiliary reading content generated by the first model, i.e., when the reading difficulty Yn is an existing word in the text, the difficulty count n is incremented by 1. Here, the difficulty count n represents the number of difficulties in the text that belong to the existing words, and its initial value is 0. At the same time, the first model is used to simplify the sentence containing the reading difficulty Yn and interpret the context of the reading difficulty Yn to obtain the professional knowledge information corresponding to the reading difficulty Yn, which is output as auxiliary reading content.
[0029] When the reading difficulty Yn is the auxiliary reading content generated by the first model, that is, when the reading difficulty Yn is a word generated by AI, the difficulty count value n remains unchanged. The first model is used to interpret the context of the reading difficulty Yn, obtain the professional knowledge information corresponding to the reading difficulty Yn as supplementary knowledge content k, and add it to the reading difficulty Yn as the new auxiliary reading content output.
[0030] Understandably, if the first model fails to find professional knowledge information corresponding to the reading difficulty Yn, meaning the output auxiliary reading content is empty, then the process continues to the next reading difficulty. Otherwise, the output auxiliary reading content is overlaid and displayed in the reading difficulty area where the corresponding reading difficulty is located to help the user understand the reading difficulty. Figure 2 As shown.
[0031] It should be noted that the embodiments of the present invention do not specifically limit the method by which the first major model obtains the professional knowledge information corresponding to the reading difficulty. For example, the first major model performs similarity matching between the input reading difficulty and professional domain words in the professional knowledge base, and selects the professional domain word with the highest similarity from the similarity scores greater than a set similarity threshold based on the similarity between the reading difficulty and the professional domain words in the professional knowledge base. This professional domain word is then used as the professional domain word that matches the reading difficulty, and the professional knowledge information corresponding to this professional domain word is then obtained from the knowledge base. The sentence simplification processing of the first major model can be found in the text summarization extraction, and will not be elaborated here.
[0032] This invention takes into account the randomness and incompleteness of the content generated by the large model. To ensure that the output auxiliary reading content can be fully understood by the user, a real-time supplementation method for the content generated by the large model is introduced. When the user views the auxiliary reading content, the eye-tracking difficulty assessment of the content generated by the large model is performed again, and the content generated by the large model is supplemented in real time. This makes the content generated by the large model more complete, provides more comprehensive auxiliary reading content for the text, ensures that the user fully understands the difficulties in reading, and improves the reading assistance effect.
[0033] In one optional embodiment, based on the eye-tracking behavior data, a pre-trained eye-tracking model is used to predict reading difficulties and determine the difficult reading regions of the text, including: The eye-tracking behavior data is input into the eye-tracking model to predict reading difficulties and obtain reading difficulty information of the text; wherein, the reading difficulty information includes the real-time difficulty of the text and its first probability value, the user difficulty and its second probability value, and the text difficulty and its third probability value; Based on the first probability value of the real-time difficulty corresponding to the text, the second probability value of the user difficulty, and the third probability value of the text difficulty, the words in the text are identified as having difficulty, and the difficult reading areas of the text are determined.
[0034] Among them, real-time difficulties are used to indicate difficult words that users may encounter that pose comprehension obstacles during the reading of text; user difficulties are used to indicate difficult words that users encounter that pose comprehension obstacles during the reading of text, selected from real-time difficulties; and text difficulties are used to indicate difficult words that may pose comprehension obstacles during subsequent reading of text, predicted by the combined effect of user difficulties, real-time difficulties, the sentences containing difficulties, and their contextual understanding.
[0035] In this embodiment of the invention, the eye-tracking model predicts the real-time difficulty and its first probability value, the user difficulty and its second probability value, and the text difficulty and its third probability value of the words the user is currently looking at while reading text, based on the currently input eye-tracking behavior data. This is combined with threshold judgment conditions, such as multiple threshold combinations. These threshold combinations include two or three of the following: the first probability value of the real-time difficulty, the second probability value of the user difficulty, and the third probability value of the text difficulty. These threshold judgment conditions indicate the probability value combination of the difficulty corresponding to the best model performance selected by the eye-tracking model during training. For example, assuming the third probability value of the text difficulty is T, the first probability value of the real-time difficulty is d, and the second probability value of the user difficulty is D, then the threshold judgment conditions include at least one of the following combinations: T > the first probability threshold (such as 0.8), the second probability threshold (such as 0.5) < d < the third probability threshold (such as 1): used to determine as a reading difficulty point; T > the first probability threshold (such as 0.8), d <= the second probability threshold (such as 0.5): used to determine as not a reading difficulty point; T <= the first probability threshold (such as 0.8), d >= D: used to determine as a reading difficulty point; T <= the first probability threshold (such as 0.8), d < D: used to determine as not a reading difficulty point.
[0036] Through the above judgment conditions, it can be determined whether the word that the user is fixating on when reading the current text is a reading difficulty point. If it is a reading difficulty point, it can further determine the reading difficulty area where the reading difficulty point is located.
[0037] In the embodiment of the present invention, the user's difficulty points are initially identified through eye movement behavior data, text difficulty discrimination is also introduced, and then combined with the threshold judgment conditions to further determine the reading difficulty points, which can supplement the fuzzy data that appears when discriminating the user's difficulty points and improve the accuracy of reading difficulty point discrimination.
[0038] Further, the training process of the eye movement tracking model includes: Collect eye movement behavior data samples of several test users when reading the pre - marked interest areas in the professional text; wherein, the interest areas are obtained by dividing the professional vocabulary in the professional text through a second large model; According to the eye movement behavior data samples, train the pre - constructed eye movement tracking model to obtain a trained eye movement tracking model; Among them, the eye movement tracking model is constructed based on a random forest model integrating a recursive feature elimination algorithm and a bidirectional long - short - term memory network model integrating a conditional random field.
[0039] Among them, according to the eye movement behavior data samples, training the pre - constructed eye movement tracking model to obtain a trained eye movement tracking model includes: Perform feature screening on each of the eye movement behavior data samples through a random forest model integrating a recursive feature elimination algorithm to obtain real - time difficulty points and their first probability values, and user difficulty points and their second probability values in each of the eye movement behavior data samples; According to the real - time difficulty points and their first probability values, and user difficulty points and their second probability values in the eye movement behavior data samples, perform text difficulty prediction on each of the eye movement behavior data samples through the bidirectional long - short - term memory network model to obtain the third probability values of the text difficulty points in each of the eye movement behavior data samples; Based on the first probability value of the real-time difficulty, the second probability value of the user difficulty, and the third probability value of the text difficulty corresponding to the eye-tracking behavior data sample, the reading difficulty area of the professional text under each eye-tracking behavior data sample is determined. Based on the reading difficulties of the professional text and the pre-labeled interest areas, the eye-tracking model is iteratively optimized until the eye-tracking model converges, thus obtaining a trained eye-tracking model. During the training process of the eye-tracking model, the model performance of the eye-tracking model under different threshold combinations is verified based on the first probability value of the real-time difficulty, the second probability value of the user difficulty, and the third probability value of the text difficulty corresponding to each eye-tracking behavior data sample. The threshold combination conditions corresponding to the best model performance are then selected. The threshold combination conditions include a combination of two or three of the first probability value of the real-time difficulty, the second probability value of the user difficulty, and the third probability value of the text difficulty. The threshold combination conditions are used to help determine the reading difficulty area.
[0040] In this embodiment of the invention, the training of the eye-tracking model mainly includes: constructing a feature set of test user eye-tracking behavior data samples when reading difficulties occur; based on the feature set, using a Random Forest (RF) model that integrates Recursive Feature Elimination (RFE) and a Bidirectional Long Short-Term Memory-Conditional Random Field (Bi-LSTM-CRF) model to train and optimize the model, enabling the system to have reading difficulty assistance functions. The trained eye-tracking model includes functions such as eye-tracking data input (e.g., inputting data into the model), eye-tracking data processing (e.g., data cleaning), and user difficulty discrimination (e.g., model difficulty prediction), etc. Figure 3 As shown. The training process of the eye-tracking model specifically includes: Sample data collection. Different roles within the same company who need to read professional texts, or different roles in other groups who frequently come into contact with professional texts, such as decision-makers, product managers, operations managers, marketing managers, customer service managers, technical experts, financial analysts, supply chain managers, R&D managers, project managers, compliance managers, and business analysts in B2B work scenarios, can be selected as test users. One representative from each role can be selected to participate in the test. The professional texts used in the test cover all professional areas that these roles may encounter in their work.
[0041] During the testing process, eye-tracking data samples were collected from all test users when they encountered difficulties reading professional texts, including temporal and spatial metrics. For temporal metrics, since reading assistance needs to provide timely prompts after difficulties occur, earlier metrics from the reading time series were selected. Because the reading assistance in this example targets vocabulary difficulties, word-based metrics were used for spatial metrics. The temporal and spatial metrics in the eye-tracking data samples are detailed above and will not be repeated here.
[0042] The aforementioned eye-tracking behavior data samples were used as training data and underwent data cleaning; for example, standard deviation was used to identify outliers in the eye-tracking behavior data samples, and outliers were removed. Eye-tracking region of interest (ROI) segmentation was performed on the selected professional text. A second-largest model (such as LLM) was used to segment ROIs on the selected professional text, dividing all content of the professional text into units based on professional domain vocabulary. Each region containing a professional domain vocabulary was designated as an ROI. The goal was to track the specific eye-tracking behavior of the user's gaze on professional domain vocabulary (i.e., eye-tracking behavior data samples). Simultaneously, the user was required to select whether knowledge prompts were needed for the selected professional text; prompts were labeled 1, and no prompts were labeled 0.
[0043] Construct a feature set. When constructing the feature set, the RFE algorithm is combined with a random forest model. Initially, all feature values are selected. RFE recursively constructs multiple random forest models and trains the currently constructed random forest model. In each iteration, the feature value that contributes the least to the random forest model is removed, and the random forest model is retrained until the predetermined number of features is reached or the performance of the random forest model no longer improves significantly.
[0044] Each feature value corresponds to one of the following indicators in the eye movement behavior data sample: fixation time, first fixation time, number of fixations, number of regressions, total fixation time, pupil diameter, fixation position, saccade amplitude, saccade angle, fixation density, fixation distribution, and saccade target selection.
[0045] Understandably, during the RFE process, after each iteration, the remaining feature values are used to train the random forest model, and the contribution of each feature value to the prediction results of the random forest model is calculated, such as the Gini coefficient of the random forest model.
[0046] During the training of the random forest model, the input is the feature value extracted in each round, the output is the prediction of reading difficulty, and the label is the professional text of the pre-labeled region of interest; the difficulty prediction is a binary classification problem: difficult (requires hints) or not difficult (does not require hints).
[0047] Specifically, the random forest model outputs a first probability value d (ranging from 0 to 1) of a real-time difficulty point for each eye movement behavior data sample, indicating the likelihood that the word corresponding to the sample is a difficult point. By adjusting the threshold judgment condition corresponding to this probability value d, it is determined when to classify the words gazed at by the user in the professional text as difficult points.
[0048] Use different threshold judgment conditions to optimize the model performance, so as to balance indicators such as the recall rate, F1 score, and AUG (Area Under the Gain Curve) value of the model; for example, when the AUG value is greater than 0.8, the model performance is considered good.
[0049] Use the cross-validation algorithm to evaluate the model performance under different thresholds (such as the above-mentioned first probability threshold, second probability threshold, and probability threshold boundary for distinguishing difficult points) corresponding to the first probability value d. According to the results of cross-validation, select a threshold under the optimal model performance as the threshold corresponding to the second probability value D of the user's difficult points. This threshold maximizes the recall rate while maintaining high accuracy to ensure that difficult points are not missed.
[0050] For example, based on the above model training, the probability threshold boundary of the first probability value can be determined. For example, when 0.5 < d < 1, it is classified as a difficult point. Then, an optimal model performance corresponding threshold can be selected from the threshold interval (0.5, 1) as the threshold of the second probability value D, so that when d >= D, it is classified as a difficult point; Furthermore, a certain proportion of samples can be extracted from the eye movement behavior data samples as a test set to evaluate the model performance of the trained random forest model under the selected threshold, ensuring the effectiveness of the model in practical applications.
[0051] Connect the trained random forest model to the bidirectional long short-term memory network model Bi-LSTM-CRF that integrates conditional random fields to learn the text difficulties, the sentences where the text difficulties are located, and the context understanding of the individual in the current professional text, so as to predict the text difficulties that may be encountered in the subsequent text content and the threshold (such as the first probability threshold) corresponding to its third probability value T, and obtain a trained eye movement tracking model.
[0052] With the trained eye movement tracking model and the thresholds corresponding to the probability values of each determined difficult point, the threshold judgment condition can be obtained and applied to subsequent reading texts to automatically identify the user's difficult points. For specific threshold judgment conditions, refer to the above description and will not be repeated here.
[0053] Furthermore, a certain proportion of samples can be extracted from the eye-tracking behavior data samples as a test set or validation set to evaluate the model performance of the eye-tracking model, for example, based on metrics such as accuracy, recall, and F1 score. Based on the evaluation results, the threshold corresponding to the probability values output by the model can be adjusted to balance the sensitivity and specificity of difficulty determination. The specific model performance evaluation process is existing technology and will not be described in detail here.
[0054] The trained eye-tracking model interacts with the primary model via an API interface. This integrates the eye-tracking model's API with the primary model's API, enabling the primary model to access eye-tracking functionality.
[0055] In this embodiment of the invention, the eye-tracking model establishes a mechanism to identify reading difficulties for users. The eye-tracking model uses a combination of recursive feature elimination and a random forest model to obtain real-time probability thresholds for difficulties and probability thresholds for user difficulties in professional texts labeled with difficulty tags. These are then input into a Bi-LSTM-CRF model to obtain text difficulty probability thresholds, thus deriving a reading difficulty identification mechanism, i.e., threshold judgment conditions. This reading difficulty identification mechanism can be used to determine difficulties in words, sentences, and paragraphs. Besides identifying user difficulties through eye-tracking behavior data, this mechanism also introduces text difficulty identification, addressing the problem of fuzzy data when identifying user difficulties. This improves the accuracy of reading difficulty identification in both deep and shallow reading scenarios and avoids the problem of chaotic auxiliary content output.
[0056] Furthermore, the first major model, once fully trained, includes functions such as question answering, content output, and knowledge supplementation. The training process for this first major model includes: The above eye-tracking data samples are input into the trained eye-tracking model to obtain the reading difficulties identified by the model, which are then used as training samples for the LLM model. Automatic question-and-answer standards, namely prompt templates, are developed based on users' questioning habits and methods for professional content.
[0057] Use prompt templates to formulate standard answers that can be executed by the LLM model for questions. For example, obtain professional knowledge information corresponding to patent domain terms from the professional knowledge base and feed it into the prompt template to generate standard answers with corresponding professional domain terms, and construct question-answer pairs, including professional domain terms and their standard answers.
[0058] The input of the constructed answers is used to train the LLM model. After each round of training iterations, the prompt word template and the parameters of the LLM model are adjusted according to the model's output answers so that the LLM model's output answers meet the requirements. For example, if the information accuracy (relative to patent knowledge information) of the LLM model's output answers exceeds a set threshold, the first trained model is obtained, which enables the first trained model to have the function of answering questions with professional knowledge.
[0059] The trained primary model also possesses content output and knowledge supplementation functions. The content output function is integrated with the primary model and is used to define the content output area within the reading text. This area is filled with supplementary reading content output by the primary model after supplementing the reading difficulties encountered in the current text. The subsequent eye-tracking model further determines whether the user's gaze falls within this area, i.e., whether the user's gaze remains within that area. Once it is determined whether the gaze point falls within the designated area, the knowledge supplementation function is triggered. The eye-tracking model is invoked to predict the difficulty level of the content in that area. The process for difficulty prediction is described above and will not be repeated here. If a reading difficulty is identified, the first model S1~S4 is invoked again to supplement the knowledge about the reading difficulty in that area. This process is repeated until the user's gaze no longer lingers in that area. During subsequent text reading, this process is repeated to refine the explanation of the reading difficulties. The knowledge supplementation function stores question-and-answer pairs, expanding the LLM model's proprietary knowledge base. Furthermore, it can also store the user's eye-tracking behavior data and question-and-answer data (i.e., reading difficulties and their corresponding supplementary reading content), forming a user knowledge base.
[0060] Based on the powerful training and learning capabilities of the two models, it can long-term memory and deep learning of users' eye-tracking behavior data, reading habits, and auxiliary content needs. This approach forms a personal knowledge base and personalized reading assistance for users during use, adapting to changes in users' reading habits, knowledge reserves, and reading comprehension.
[0061] Compared to existing technologies, this invention uses eye-tracking data to identify difficulties in text reading, the sentences containing the difficulties, and the contextual information of the difficulties. It also uses an LLM model to simplify the sentences containing the difficulties based on the identified difficulties and provide explanations of the difficulties, i.e., supplementary reading content. Compared to simply outputting explanations of the difficulties, this invention can output simplified sentences and explanations of the difficulties simultaneously, which can solve the problem of reading difficult professional texts, effectively improve the accuracy and timeliness of knowledge assistance, and enhance the reading assistance effect.
[0062] This invention utilizes a combination of eye-tracking and LLM models to generate supplementary reading content. Compared to existing technologies based on pre-set corpora and word indexes, this approach is more efficient, comprehensive, and accurate. It addresses the issues of incomplete semantic coverage, inadequate understanding, and inefficient searching for complex professional vocabulary, achieving more accurate and efficient explanations and responses. Furthermore, the LLM model provides more personalized responses and can perform full-network searches and long-term learning, resulting in more comprehensive and reliable content responses.
[0063] See Figure 4 , Figure 4 This is a structural block diagram of an intelligent reading assistance device provided in an embodiment of the present invention. The intelligent reading assistance device includes: The eye-tracking data acquisition module 11 is used to acquire eye-tracking behavior data of the user while reading text; The reading difficulty region determination module 12 is used to predict reading difficulties in the text based on the eye movement behavior data and a pre-trained eye tracking model, thereby determining the reading difficulty region of the text. The knowledge supplementation module 13 is used to supplement the reading difficulties in the reading difficulty area through the pre-trained first large model.
[0064] In one optional embodiment, the eye-tracking data acquisition module 11 includes: An eye-tracking detection unit is used to detect the temporal and spatial metrics of the user's eyes on the text while the user is reading the text; wherein, the temporal metrics include one or more of the following: fixation time, first fixation time, number of fixations, number of regressions, total fixation time, and pupil diameter; the spatial metrics include one or more of the following: fixation position, saccade amplitude, saccade angle, fixation density, fixation distribution, and saccade target selection. An eye-tracking data generation unit is used to generate the eye-tracking behavior data based on the time indicators and the spatial indicators.
[0065] In one optional embodiment, the reading difficulty region determination module 12 includes: The difficulty prediction unit is used to input the eye-tracking behavior data into the eye-tracking model to predict reading difficulties and obtain reading difficulty information of the text; wherein, the reading difficulty information includes the real-time difficulty of the text and its first probability value, the user difficulty and its second probability value, and the text difficulty and its third probability value. The first difficult region determination unit is used to identify the difficult regions of the text by recognizing the words in the text based on the first probability value of the real-time difficulty corresponding to the text, the second probability value of the user difficulty, and the third probability value of the text difficulty.
[0066] In one optional embodiment, the knowledge supplementation module 13 includes: The difficulty identification unit is used to determine whether the reading difficulty in the reading difficulty area is auxiliary reading content; wherein, the auxiliary reading content is generated based on the first large model; The first auxiliary content generation unit is used to, if not, simplify the sentence containing the reading difficulty through the first large model and generate auxiliary reading content for the reading difficulty; The second auxiliary content generation unit is used, if so, to generate supplementary knowledge content on the reading difficulties through the first large model.
[0067] In an optional embodiment, the device further includes: The data acquisition module is used to collect eye-tracking behavior data samples from several test users when reading pre-marked regions of interest in professional texts; wherein, the regions of interest are obtained by dividing the professional vocabulary in the professional texts into regions using a second major model; The training process module is used to train the pre-built eye-tracking model based on the eye-tracking behavior data samples to obtain a trained eye-tracking model. The eye-tracking model is constructed based on a random forest model that integrates recursive feature elimination algorithm and a bidirectional long short-term memory network model that integrates conditional random field.
[0068] In one optional embodiment, the training process module includes: The feature filtering unit is used to perform feature filtering on each eye-tracking behavior data sample by using a random forest model that integrates a recursive feature elimination algorithm, and to obtain the real-time difficulties and their first probability values, and the user difficulties and their second probability values in each eye-tracking behavior data sample. The text difficulty prediction unit is used to predict the text difficulty of each eye movement data sample based on the real-time difficulty and its first probability value, the user difficulty and its second probability value in the eye movement behavior data sample, and through the bidirectional long short-term memory network model, to obtain the third probability value of the text difficulty in each eye movement behavior data sample. The second difficulty region determination unit is used to determine the reading difficulty region of the professional text under each eye movement behavior data sample based on the first probability value of the real-time difficulty corresponding to the eye movement behavior data sample, the second probability value of the user difficulty, and the third probability value of the text difficulty. The model iterative optimization unit is used to iteratively optimize the eye-tracking model based on the difficult reading areas of the professional text and the pre-labeled interest areas until the eye-tracking model converges and a trained eye-tracking model is obtained. During the training process of the eye-tracking model, the model performance of the eye-tracking model under different threshold combinations is verified based on the first probability value of the real-time difficulty, the second probability value of the user difficulty, and the third probability value of the text difficulty corresponding to each eye-tracking behavior data sample. The threshold combination conditions corresponding to the best model performance are then selected. The threshold combination conditions include a combination of two or three of the first probability value of the real-time difficulty, the second probability value of the user difficulty, and the third probability value of the text difficulty. The threshold combination conditions are used to help determine the reading difficulty area.
[0069] It should be noted that the working process of each module in the intelligent reading assistance device described in the embodiments of the present invention can refer to the working process of the intelligent reading assistance method described in the above embodiments, and the technical effect achieved is the same as that of the intelligent reading assistance method described in the above embodiments, so it will not be repeated here.
[0070] See Figure 5 , Figure 5 This is a structural block diagram of an intelligent reading assistance device provided in an embodiment of the present invention. The intelligent reading assistance device includes a processor 21, a memory 22, and a computer program stored in the memory 22 and executable on the processor 21. When the processor 21 executes the computer program, it implements the steps in the various intelligent reading assistance method embodiments described above, such as steps S11 to S13.
[0071] For example, the computer program may be divided into one or more modules or units, which are stored in the memory 22 and executed by the processor 21 to complete the present invention. The one or more modules or units may be a series of computer program instruction segments capable of performing specific functions, which describe the execution process of the computer program in the intelligent reading assistance device.
[0072] The intelligent reading aid device may include, but is not limited to, a processor 21 and a memory 22. Those skilled in the art will understand that the schematic diagram is merely an example of an intelligent reading aid device and does not constitute a limitation on the intelligent reading aid device. It may include more or fewer components than illustrated, or combine certain components, or different components. For example, the intelligent reading aid device may also include input / output devices, network access devices, buses, etc.
[0073] The processor 21 can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor. The processor 21 is the control center of the intelligent reading aid device, connecting all parts of the intelligent reading aid device via various interfaces and lines.
[0074] The memory 22 can be used to store the computer programs and / or modules. The processor 21 implements various functions of the intelligent reading aid device by running or executing the computer programs and / or modules stored in the memory 22 and calling the data stored in the memory 22. The memory 22 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the mobile phone (such as audio data, phonebook, etc.). In addition, the memory 22 may include high-speed random access memory, and may also include non-volatile memory, such as hard disk, memory, plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, at least one disk storage device, flash memory device, or other volatile solid-state storage device.
[0075] If the modules or units integrated into the intelligent reading assistance device are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by the processor 21, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc.
[0076] It should be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Furthermore, in the accompanying drawings of the device embodiments provided by this invention, the connection relationships between modules indicate that they have communication connections, which can be specifically implemented as one or more communication buses or signal lines. Those skilled in the art can understand and implement this without any creative effort.
[0077] The above description represents the preferred embodiments of the present invention. It should be noted that, for those skilled in the art, various improvements and modifications can be made without departing from the principles of the present invention, and these improvements and modifications are also considered to be within the scope of protection of the present invention.
Claims
1. An intelligent reading assistance method, characterized in that, include: Acquire eye-tracking data of users while they are reading text; Based on the eye-tracking data, a pre-trained eye-tracking model is used to predict reading difficulties and identify the difficult reading areas of the text. The first pre-trained model supplements the knowledge of reading difficulties in the aforementioned reading difficulty areas.
2. The intelligent reading assistance method as described in claim 1, characterized in that, The acquisition of eye-tracking data of users while reading text includes: The system detects the temporal and spatial metrics of the user's eyes on the text while the user is reading it. The temporal metrics include one or more of the following: fixation time, first fixation time, number of fixations, number of regressions, total fixation time, and pupil diameter. The spatial metrics include one or more of the following: fixation position, saccade amplitude, saccade angle, fixation density, fixation distribution, and saccade target selection. The eye movement behavior data is generated based on the time and spatial metrics.
3. The intelligent reading assistance method as described in claim 1, characterized in that, Based on the eye-tracking data, a pre-trained eye-tracking model is used to predict reading difficulties and identify the difficult reading areas of the text, including: The eye-tracking behavior data is input into the eye-tracking model to predict reading difficulties and obtain reading difficulty information of the text; wherein, the reading difficulty information includes the real-time difficulty of the text and its first probability value, the user difficulty and its second probability value, and the text difficulty and its third probability value; Based on the first probability value of the real-time difficulty corresponding to the text, the second probability value of the user difficulty, and the third probability value of the text difficulty, the words in the text are identified as having difficulty, and the difficult reading areas of the text are determined.
4. The intelligent reading assistance method as described in claim 1, characterized in that, The first pre-trained model supplements the reading difficulties in the aforementioned reading hardship areas with knowledge, including: Determine whether the reading difficulties in the reading difficulty area are auxiliary reading content; wherein, the auxiliary reading content is generated based on the first large model; If not, the first major model is used to simplify the sentence containing the reading difficulty and generate supplementary reading content for the reading difficulty. If so, supplementary knowledge content on the reading difficulties is generated through the first major model.
5. The intelligent reading assistance method as described in claim 1, characterized in that, The training process of the eye-tracking model includes: Several eye-tracking data samples were collected from test users as they read pre-marked regions of interest in professional texts; wherein, the regions of interest were obtained by dividing the professional vocabulary in the professional texts into regions using a second major model; Based on the eye movement behavior data samples, the pre-built eye movement tracking model is trained to obtain a trained eye movement tracking model; The eye-tracking model is constructed based on a random forest model that integrates recursive feature elimination algorithm and a bidirectional long short-term memory network model that integrates conditional random field.
6. The intelligent reading assistance method as described in claim 5, characterized in that, Based on the eye-tracking behavior data samples, a pre-built eye-tracking model is trained to obtain a trained eye-tracking model, including: By using a random forest model that incorporates a recursive feature elimination algorithm, feature filtering is performed on each of the eye-tracking behavior data samples to obtain the real-time difficulties and their first probability values, and the user difficulties and their second probability values in each of the eye-tracking behavior data samples. Based on the real-time difficulties and their first probability values, user difficulties and their second probability values in the eye-tracking behavior data samples, the text difficulties of each eye-tracking behavior data sample are predicted by the bidirectional long short-term memory network model to obtain the third probability value of the text difficulties in each eye-tracking behavior data sample. Based on the first probability value of the real-time difficulty, the second probability value of the user difficulty, and the third probability value of the text difficulty corresponding to the eye-tracking behavior data sample, the reading difficulty area of the professional text under each eye-tracking behavior data sample is determined. Based on the reading difficulties of the professional text and the pre-labeled interest areas, the eye-tracking model is iteratively optimized until the eye-tracking model converges, thus obtaining a trained eye-tracking model. During the training process of the eye-tracking model, the model performance of the eye-tracking model under different threshold combinations is verified based on the first probability value of the real-time difficulty, the second probability value of the user difficulty, and the third probability value of the text difficulty corresponding to each eye-tracking behavior data sample. The threshold combination conditions corresponding to the best model performance are then selected. The threshold combination conditions include a combination of two or three of the first probability value of the real-time difficulty, the second probability value of the user difficulty, and the third probability value of the text difficulty. The threshold combination conditions are used to help determine the reading difficulty area.
7. An intelligent reading aid device, characterized in that, include: The eye-tracking data acquisition module is used to acquire eye-tracking behavior data of users while reading text; The reading difficulty region determination module is used to predict reading difficulties based on the eye movement behavior data using a pre-trained eye tracking model, and determine the reading difficulty region of the text. The knowledge supplementation module is used to supplement the reading difficulties in the reading difficulty area using the pre-trained first model.
8. An intelligent reading assistance device, characterized in that, include: A processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor, when executing the computer program, implements the intelligent reading assistance method as described in any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, wherein, when the computer program is executed, it controls the device where the computer-readable storage medium is located to perform the intelligent reading assistance method as described in any one of claims 1 to 6.
10. A computer program product, comprising a computer program or instructions, characterized in that, When the computer program or instructions are executed by the processor, they implement the intelligent reading assistance method according to any one of claims 1 to 6.