Image magnification display system
The image enlargement display system uses reinforcement learning to determine and manage image enlargement based on user gaze, addressing the inaccuracies and delays of prior systems by promptly and accurately enlarging difficult characters or images.
Patent Information
- Application Number
- JP2024014289
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-02-01
- Publication Date
- 2025-08-14
AI Technical Summary
Existing image enlargement systems fail to accurately and quickly respond to user gaze fixation to enlarge difficult-to-read characters or images, often requiring user installation of new software and lacking the ability to learn user gaze fixation states for timely character enlargement.
An image enlargement display system using a reinforcement learning unit to determine image enlargement based on gaze fixation position and time, incorporating an image enlargement control unit to manage character and image enlargement on an image display device, with components like a data input unit, reward determination unit, function update unit, and enlargement condition creation unit to learn and apply appropriate enlargement conditions.
The system accurately and promptly enlarges difficult-to-identify characters or images based on user gaze, ensuring timely and user-satisfactory enlargement without unnecessary delays, overcoming the limitations of prior systems.
Smart Images

Figure 2025119409000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an image magnification display system that magnifies an image based on a user's line of sight. [Background technology]
[0002] Development is underway on an image enlargement system that automatically enlarges and displays difficult-to-read text or images when the user fixates their gaze on the screen of a website, tablet, smartphone, or other device. In this system, the user communicates to the machine (agent) a request to enlarge text or images by fixating their gaze, and it is essential for the agent to quickly grasp the user's request during the time that the gaze is fixed. For example, Non-Patent Document 1 discloses a web browser that allows users with motor disabilities to control mouse operations with their gaze. Non-Patent Document 2 discloses a system that uses a Windows magnifying glass to automatically enlarge the area around a character when the gaze stays there for a long period of time. Non-Patent Document 3 discloses a system that predicts the occurrence of character discrimination difficulties from gaze fixation time based on reinforcement learning. [Prior art documents] [Non-patent literature]
[0003] [Non-Patent Document 1] Menges et al., 2019. Improving User Experience of Eye Tracking-Based Interaction: Introspecting and Adapting Interfaces. ACM Transactions on Computer-Human Interaction, 26, 6 (October 2019), Article 37:1-46 [Non-patent document 2] Ishida, 2020, Development of an automatic enlargement and presentation system for difficult-to-identify characters on the web based on gaze, Kanazawa Institute of Technology PDIII Project Report, 2020. [Non-patent document 3] Saito and Matsumori, 2020, Prediction system for character identification difficulties when browsing the web, Kanazawa Institute of Technology PDIII Project Report, 2020. Summary of the Invention [Problem to be solved by the invention]
[0004] However, Non-Patent Document 1 has the problem that it requires the user to install a new browser, and it also does not include the concept of enlarging web page elements such as difficult-to-identify characters or images. Non-Patent Document 2 does not determine whether the user is experiencing difficulty in character identification, which may result in unnecessary enlargement of characters. Non-Patent Document 3 merely devises a method for predicting the occurrence of difficulty in character identification because it sets the gaze point as a state. In other words, it has the problem of not being a system in which the agent learns the user's gaze fixation state and performs quick and accurate character enlargement.
[0005] SUMMARY OF THE INVENTION In consideration of the above problems, an object of the present invention is to provide an image magnification display system that magnifies an image based on the user's line of sight. [Means for solving the problem]
[0006] The image enlargement display system of the present invention comprises an image display device that displays documents and images, and a control device that controls the enlargement of characters and images on the image display device. The control device comprises a reinforcement learning unit that determines whether or not to enlarge an image based on the gaze fixation position and gaze fixation time calculated from the user's gaze point coordinates collected by an eye gaze sensor installed on the image display device, and an image enlargement control unit. The reinforcement learning unit is composed of a data input unit that reads time series data consisting of an eye movement characteristic vector and an indicator of the appropriateness of enlargement stored in the image enlargement control unit, a reward determination unit that determines a reward based on the eye movement characteristic vector of the data input unit, a function update unit that updates an action value function while gradually developing the state of the gaze fixation time based on the reward determined by the reward determination unit and the time series data of the data input unit, and an enlargement condition creation unit that creates image enlargement conditions from the learning results obtained by repeatedly updating the action value function. The control device controls the image display device so as to satisfy the image enlargement conditions learned by the reinforcement learning unit. The early reinforcement learning unit is characterized in that it learns by successively evolving the state of the gaze fixation time and evaluating the appropriateness of image enlargement as needed. Furthermore, the eye movement characteristic vector x other than the gaze fixation time at each gaze point i on the image display screen of the image display device is i is used as external information for the state. Further, when the gaze fixation time is equal to or longer than the time predicted to make character identification difficult, the character is always enlarged, and when it is shorter than that time, the character is always not enlarged. [Effects of the Invention]
[0007] The present invention is a system that guides the user's gaze fixation to quickly and accurately enlarge difficult-to-identify characters or images. Therefore, gaze fixation time cannot be used as an explanatory variable. This is because if gaze fixation time were used as an explanatory variable, the difficulty of identification would be predicted based on the results of a certain period of gaze fixation, making it impossible to predict the difficulty of identification and enlarge the characters at each time step from the start to the end of the gaze fixation. In other words, the present invention uses the user's gaze fixation time as a control variable for enlarging the characters, and waiting for the results until the fixation ends would prevent timely enlargement. Therefore, unlike general reinforcement learning, which seeks the correct action plan at each gaze point, this invention is characterized by deriving a uniform criterion for determining character enlargement across all gaze points in a way that satisfies the user, even at the expense of the results of some gaze points. [Brief explanation of the drawings]
[0008] [Figure 1] Block diagram showing the configuration of an image enlargement display system [Figure 2] Figures (a) and (b) show examples of image enlargement. [Figure 3] Photograph showing an example of an image display device and an eye-gaze sensor [Figure 4] Flowchart of image enlargement control unit [Figure 5] Reinforcement learning flow chart [Figure 6] Table showing the increase / decrease coefficients [Figure 7] Action-value function table [Figure 8] Reinforcement learning convergence results DETAILED DESCRIPTION OF THE INVENTION
[0009] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS An embodiment of an image enlargement display system according to the present invention will be described. Note that the following embodiment is an easy-to-understand example and is not intended to limit the present invention, its application method, and application range. As shown in FIG. 1, the image enlargement display system 1 includes an image display device 10 and a control device 20. The image display device 10 displays characters and images on an image display surface 11. Examples of the image display device 10 include, but are not limited to, a personal computer, a tablet terminal, a smartphone, etc. When characters or images that are difficult for a user to distinguish appear on the image display surface 11 and the user fixates their gaze on the characters, etc., the image display device 10 automatically enlarges and displays the characters, etc., based on instructions from the control device 20, as shown by the arrows in Figures 2(a) and (b).
[0010] (Image display device 10) As shown in Fig. 3, the image display device 10 is equipped with an eye-gaze sensor 12. The eye-gaze sensor 12 detects the coordinates of the user's gaze point on the image display surface 11 at regular time intervals (e.g., 1 / 60 seconds). This allows calculation of eye movement characteristics such as gaze fixation time, gaze movement speed, and movement distance. The eye-gaze sensor 12 may be incorporated into the image display device 10.
[0011] (Control device 20) The control device 20 will be described with reference to FIG.
[0012] The control device 20 includes an image enlargement control unit 30 and a reinforcement learning unit 40. The control device 20 is configured to be able to communicate with the image display device 10 via the internal circuitry of a computer or wirelessly. The control device 20 receives as input gaze point coordinates collected by the gaze sensor 12, and causes the image display device 10 to display or remove enlarged windows of characters or images as shown in Figures 2(a) and (b).
[0013] (Image enlargement control unit 30) Based on the gaze point coordinates collected by the gaze sensor 12, the image enlargement control unit 30 calculates a vector x consisting of gaze fixation time and eye movement characteristic values other than the gaze fixation time for each gaze point i. i Calculate (eye movement speed, movement distance, pupil diameter, etc.).
[0014] FIG. 4 shows a flowchart in which the image enlargement control unit 30 operates the image display device 10 to enlarge an image that is difficult to distinguish. Data exchange between the image display device 10 and the image enlargement control unit 30 is performed online. The image enlargement control unit 30 uses the eye movement characteristic vector x derived by the reinforcement learning unit 40. i When the image enlargement condition labeled with is satisfied, the image enlargement window appears on the image display device 10, and then when the enlargement window deletion condition is satisfied, the enlargement window is deleted. i where a i When enlarging the text is the correct answer (necessary for the user), i = 1, and if the answer is incorrect, a i = 0. If the character enlargement is incorrect, a button indicating this is provided on the image display surface 11, and the user touches or clicks the button with their eyes or a mouse. In order to reflect this in learning, the eye movement characteristic vector x at each fixation point i until the end of browsing is calculated. i and the suitability index for expansion a i The time series data (described below) is saved and transmitted to the reinforcement learning unit 40. This transmission may be performed sequentially for each gaze point, or several data may be batch processed together.
number
[0015] (Reinforcement Learning Section 40) The learning update algorithm in the reinforcement learning unit 40 may be the Monte Carlo method, SARSA method, Q-learning method, function approximation method, DQN (Deep Q-Network) method, or the like. In this embodiment, reinforcement learning using the SARSA method will be described on the premise that when character identification becomes difficult during web browsing, the relevant character portion is enlarged and displayed. The reinforcement learning unit 40 corresponds to an agent in reinforcement learning. The reinforcement learning unit 40 learns the appropriate conditions for image enlargement by evaluating which action the agent should take in a certain state to maximize the action value function. The "state" of the reinforcement learning unit 40 is the eye movement characteristic vector x iThe "actions" are enlarging or not enlarging the letters, and if the letters are enlarged when they are difficult to distinguish, a reward is given, while if they are not enlarged when they are difficult to distinguish (no enlargement) or if they are enlarged when they are not difficult to distinguish, a negative reward (loss) is given. To appropriately enlarge characters, the occurrence of character discrimination difficulties is predicted from the user's eye movements. To achieve this, the gaze sensor 12 collects the user's gaze data while viewing, and the image enlargement control unit 30 calculates an index for the appropriateness of enlargement. However, this cannot be used as complete training data. This is because character discrimination difficulties may or may not occur even if all eye movement characteristic values are the same. Therefore, it is appropriate to use reinforcement learning, which learns the correct answer through trial and error. 5 shows a flowchart of the reinforcement learning unit 40. The reinforcement learning unit 40 is independent of the image enlargement control unit 30, and can be either a batch type or an online type.
[0016] The reinforcement learning unit 40 includes, as functional components, a data input unit 41, a reward determination unit 42, a function update unit 43, and an expansion condition creation unit 44.
[0017] The data input unit 41 receives time-series data that reflects the user's most recent browsing behavior. This time-series data is used by the agent to determine whether to grant a reward or a loss, and as a basis for updating the action-value function.
[0018] The reward determination unit 42 determines the reward as the eye movement characteristic vector x i It is determined depending on the category number k (see the action value function below) to which it belongs and the time step t of the gaze fixation. Therefore, this reward is k,t This definition is intended to minimize the contradiction that one screen viewing contains many fixation points, and even with the same fixation time, one fixation point may be enlarged and another may not be enlarged. Specifically, for each fixation point i, the increase / decrease coefficient c k,t The reward and loss are calculated by multiplying this coefficient by the following equation (1).
number
[0019] The function update unit 43 gradually develops the gaze fixation time for each fixation point. The state change and the result of the action selection during one fixation point correspond to the move and the outcome of a game (Go or Shogi). The user conveys a request to enlarge the text or image to the agent by fixating on it, and it is important for the agent to quickly grasp the user's request during the time that the fixation time is passing. Therefore, the function update unit 43 defines the "state" as the time development of the fixation: if the time increment is Δd, then the state (gaze fixation time) S at the tth step is t is expressed by the following equation (2).
number
[0020] (Action Value Function) Next, the function update unit 43 calculates the next action value function for each dwell time that evolves according to the formula (2).
number
[0021] The enlargement condition creation unit 44 creates conditions for image enlargement based on the learning results. However, significantly different from general reinforcement learning, a constraint is added that "when the gaze fixation time is longer than the predicted character discrimination difficulty, always enlarge the characters, and when it is shorter, always select not to enlarge the characters." This is an essential constraint for guiding the character enlargement based on the user's gaze fixation. Therefore, unlike general reinforcement learning, it is not enough to simply find the correct action plan at each gaze point; rather, a uniform criterion for determining character enlargement across all gaze points must be derived, even if it means sacrificing the results of some gaze points. The eye movement characteristic vector x at gaze point i i Let k be the category number to which the eye gaze belongs. k is determined and the criterion of the following equation (3) is derived as the learning result.
number
[0022] As shown in the flowchart in Figure 5, learning is repeated in a triple loop. As shown in Figure 8, learning is repeated until the action value function converges. Note that Figure 8 shows the difference in the action value function for each iteration. Focusing on the state of the i-th fixation point and the t-th step, the minimum configuration of sample data (empirical data) is described as in the following equation (4).
number
[0023] (Update Algorithm) As mentioned above, the learning update algorithms used are the Monte Carlo method, the SARSA method, the Q-learning method, the function approximation method, and the DQN (Deep Q-Network) method. The SARSA method and the Q-learning method require reference to a table of action value functions (Q-tables), but as mentioned above, the eye movement characteristic values x other than fixation time are also used. i into appropriate categories and calculate Q for each combination. k It is sufficient to create a table of the above. As the number of eye movement characteristics and categories increases, the number of Q tables increases explosively, so function approximation methods and DQN methods based on function approximation become effective. In fact, eye movement characteristic values other than fixation time can be treated as continuous values, and the DQN method can also create a large amount of sample data for learning, so it is possible to derive highly accurate criteria for character enlargement. Furthermore, the increase / decrease coefficient c k,tcategory number x instead of k i In reality, when the time series data is relatively small, it is appropriate to use the SARSA method or Q-learning method that uses a Q-table, and when the data becomes large, it is appropriate to switch to the function approximation method or DQN method. [Industrial Applicability]
[0024] The present invention is an image magnification display system that magnifies an image based on the user's line of sight, and has industrial applicability. [Explanation of symbols]
[0025]
number
Claims
1. The device includes an image display device that displays documents and images, and a control device that controls the enlargement of characters and images on the image display device, The control device includes a reinforcement learning unit and an image enlargement control unit that determine whether or not an image needs to be enlarged based on a gaze fixation position and a gaze fixation time calculated from coordinates of a user's gaze point collected by a gaze sensor installed on the image display device, and the reinforcement learning unit is comprised of a data input unit that reads time-series data consisting of an eye movement characteristic vector and an index of whether enlargement is appropriate, which are stored in the image enlargement control unit; a reward determination unit that determines a reward according to the eye movement characteristic vector of the data input unit; a function update unit that updates an action value function while gradually evolving the state of the gaze fixation time based on the reward determined by the reward determination unit and the time-series data of the data input unit; and an enlargement condition creation unit that creates image enlargement conditions from learning results obtained by repeatedly updating the action value function, The image enlargement display system is characterized in that the control device controls the image display device so as to satisfy the image enlargement condition learned by the reinforcement learning unit.
2. The image enlargement display system according to claim 1, characterized in that the early reinforcement learning unit sequentially develops the state of the gaze fixation time, and evaluates the appropriateness of image enlargement at any time to learn.
3. The eye movement characteristic vector x other than the gaze fixation time at each gaze point i on the image display screen of the image display device i 2. The image enlargement display system according to claim 1, wherein the image enlargement display system uses the above as external information for the state.
4. The image enlargement display system according to claim 1, characterized in that when the gaze fixation time is equal to or longer than the time predicted to make character identification difficult, the system always selects to enlarge the characters, and when it is shorter than that time, the system always selects not to enlarge the characters.
Citation Information
Cited By
Method and system for adjusting vibration optical fiber alarm threshold
CN121482981A