Image display processing method and terminal device

By identifying regions of interest to the user and then magnifying the image, the problem of poor user experience in existing technologies is solved, achieving accurate image magnification and reducing computational load.

WO2026045387A1PCT designated stage Publication Date: 2026-03-05HONOR DEVICE CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/095118
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-08-30
Filing Date
2025-05-15
Publication Date
2026-03-05

AI Technical Summary

Technical Problem

When existing terminal devices control image zooming via eye movements, they typically zoom mechanically from the image center, failing to meet users' specific zooming needs and resulting in a poor user experience.

Method used

By acquiring user gaze information, identifying the user's region of interest, and using this as a reference area for image magnification, combined with semantic segmentation technology to identify multiple semantic regions, the region of interest that reflects the user's true intent is accurately magnified.

Benefits of technology

It achieves more precise image magnification, meets users' needs for magnifying specific areas, improves user experience, and reduces computational load and power consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025095118_05032026_PF_FP_ABST
    Figure CN2025095118_05032026_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed in the embodiments of the present application are an image display processing method and a terminal device. The method comprises: displaying a target image; acquiring a gaze region and a gaze duration of a user gazing at the target image; when the gaze duration exceeds a threshold, determining a target region of interest in the gaze region; and performing magnification processing on the image on the basis of the target region of interest, and displaying the magnified image on a display screen. In the embodiments of the present application, gaze-controlled picture magnification can be realized on the basis of a region of interest of users, thereby improving user experience and reducing power consumption.
Need to check novelty before this filing date? Find Prior Art

Description

An image display processing method and terminal device

[0001] This application claims priority to Chinese Patent Application No. 202411221310.6, filed on August 30, 2024, entitled "An Image Display Processing Method and Terminal Device", the entire contents of which are incorporated herein by reference. Technical Field

[0002] This invention relates to the field of image processing, and more particularly to an image display processing method and terminal device. Background Technology

[0003] With the development of terminal technology, terminal devices can now perform many image editing operations. Enlarging images on a terminal device is a relatively important and fundamental operation.

[0004] Existing terminal devices can enable users to control image zooming with their eyes without touching the device, freeing up their hands and meeting the needs of specific scenarios. However, current image zooming control with eyes usually involves mechanical zooming based on the image center, which cannot meet the user's specific zooming needs and results in a poor user experience. Summary of the Invention

[0005] This invention provides an image display processing method and a terminal device that uses the user's region of interest as a reference area and controls image magnification through eye tracking.

[0006] In a first aspect, embodiments of the present invention provide an image display processing method, the method being applied to a terminal device, the terminal device including a display screen, comprising: displaying a target image on the display screen; acquiring first gaze information of a user gazing at the target image, the first gaze information including a first gaze region of the user gazing at the target image and a first gaze time of the user gazing at the first gaze region; determining whether the first gaze time exceeds a first threshold; if it exceeds the first threshold, determining a first target region of interest within the first gaze region based on the first gaze region and the content of the first gaze region; performing a first magnification processing on the target image based on the first target region of interest, and displaying the target image after the first magnification processing on the display screen.

[0007] In this embodiment, the terminal device can acquire the user's gaze time and determine whether the user is consciously gazing at the image using a threshold, thereby more accurately judging whether the user may have the intention and need to zoom in on the photo. Furthermore, during image zooming, this embodiment first determines a larger area of ​​user gaze, and then, based on this gaze area and its content (e.g., a person, scenery, or local features of a person), determines a smaller area of ​​user interest that more closely reflects the user's true intention. Then, the image is zoomed in on this area of ​​interest, thus solving the problems of high computational load and inaccurate zooming caused by zooming in on the entire image or the entire area. This embodiment can more accurately identify the user's target zoom area while reducing computational load and improving zooming efficiency. For example, when a user views an image using a terminal device, precise zooming based on eye gaze can be achieved, meeting the user's need to zoom in on a specific area. The operation is simple and convenient, improving the user experience.

[0008] In one possible implementation of the first aspect, determining the first target region of interest based on the first gaze region includes: performing semantic segmentation on the target image to identify M first-level semantic regions in the target image, where M is an integer greater than 1; determining N first-level semantic regions that intersect with the first gaze region from the M first-level semantic regions; and determining the first target region of interest based on a preset magnification priority of the N first-level semantic regions and the area occupied by the N first-level semantic regions in the first gaze region. Implementing the embodiments of this application allows for the identification of regions of interest that intersect with the user's gaze region and that the user may have, improving the accuracy of gaze detection. Furthermore, by comprehensively determining the target region that the user most likely wants to magnify during gaze by using a preset magnification priority and area, it avoids erroneously magnifying interference regions in the image that have high magnification priority but small area.

[0009] In one possible implementation of the first aspect, semantic segmentation of the target image to identify M first-level semantic regions in the target image includes: performing semantic segmentation on the target image according to a first semantic category to identify Q second-level semantic regions in the target image, where Q is a positive integer less than M; performing semantic segmentation on each of the Q second-level semantic regions according to a second semantic category to identify one or more first-level semantic regions in each second-level semantic region; and determining the first-level semantic regions corresponding to the Q second-level semantic regions as the M first-level semantic regions. In this embodiment, when a user views an image displaying multiple objects or human bodies and details of those objects or human bodies, multiple semantic segmentations can identify and segment different objects or human bodies. Simultaneously, more subdivided regions can be identified and segmented from the segmented objects or human bodies, improving the accuracy of semantic region recognition, uncovering more information in the image, and more accurately determining the user's likely desired area of ​​interest for zooming in.

[0010] In one possible implementation of the first aspect, the first semantic category includes one or more of people, animals, plants, and scenery; the second semantic category includes one or more of human facial features, hair, arms, legs, hands, feet, clothing, tree crowns, tree trunks, flower crowns, plant stems, animal facial features, and limbs. In the embodiments of this application, when the image viewed by the user includes multiple different categories of objects and human bodies, semantic segmentation of the target image according to the first semantic category can first identify different objects or human bodies, more accurately distinguishing the subject from the background. Furthermore, when different categories of objects or human bodies contain a lot of details, semantic segmentation of the target image segmented according to the first semantic category according to the second semantic category can identify semantic regions such as eyes, noses, arms, and palms of the identified objects or human bodies in the target image, improving the accuracy of semantic segmentation.

[0011] In one possible implementation of the first aspect, the first target region of interest is a portion or the entirety of a first-level semantic region that intersects with the first gaze region. In this embodiment, when the first target region of interest is a portion of a first-level semantic region that intersects with the first gaze region, the first-level semantic region is partially covered by the first gaze region, and the covered portion is the first target region of interest. The electronic device uses this first target region of interest as the area the user might want to zoom in on. This avoids calculating the area of ​​a first-level semantic region that has high priority or a large area but only a small area intersecting with the first gaze region, as the target area the user might want to zoom in on. This allows the device to focus only on content contained within the user's gaze region, reducing computational load.

[0012] In one possible implementation of the first aspect, performing a first magnification process on the target image based on the first target region of interest includes: if the first target region of interest is semantically complete, then using the first target region of interest as a first reference region for the first magnification operation, and performing the first magnification process on the target image based on the first reference region; if the first target region of interest is semantically incomplete, then determining a second reference region based on the first target region of interest, the first-level semantic region, and the second-level semantic region, and performing the first magnification process on the target image based on the second reference region. In this embodiment, when the semantics of the target region that the user may wish to magnify is complete, the target region can be directly used as the reference region for magnification, and the target region will become the center region of the magnified image. When the semantics of the target region that the user may wish to magnify is incomplete, the reference region for magnification can be adjusted, and a larger region containing the target region that the user may wish to magnify can be used as the reference region for magnification, and the larger region containing the target region that the user may wish to magnify will become the center region of the magnified image. This ensures that each magnification process is based on the user's region of interest, and the image displayed on the screen will present the complete semantics of the user's region of interest.

[0013] In one possible implementation of the first aspect, after displaying the target image after the first magnification process on the display screen, the method further includes: when the target image after the first magnification process meets a first preset condition, acquiring second gaze information of the user gazing at the target image after the first magnification process, the second gaze information including a second gaze region gazed at by the user in the target image after the first magnification process and a second gaze time of the user gazing at the second gaze region; determining whether the second gaze time exceeds a second threshold; if it exceeds the second threshold, determining a second target interest region within the second gaze region based on the second gaze region and the content of the second gaze region; performing a second magnification process on the target image based on the second target interest region, and displaying the target image after the second magnification process on the display screen. Therefore, in this embodiment of the application, when the image clarity after magnification is still high and contains a lot of information, the image can be magnified again through gazing, helping the user to further explore image information or process the image.

[0014] In one possible implementation of the first aspect, the method further includes: obtaining the current resolution of the display screen and the resolution of the target image after the first magnification process; if the resolution of the target image after the first magnification process is greater than or equal to the current resolution of the display screen, then the target image after the first magnification process satisfies the first preset condition; if the resolution of the target image after the first magnification process is lower than the current resolution of the display screen, then the target image after the first magnification process does not satisfy the first preset condition. Implementing this embodiment, after the target image undergoes the first magnification process and is displayed on the display screen, the resolution of the target image after the first magnification process is compared with the current resolution of the terminal device's display screen. When the resolution of the target image after the first magnification process is greater than or equal to the current resolution of the terminal device's display screen, further magnification of the target image can still obtain a relatively clear image and more detailed information, therefore it is determined that the image can be magnified again, i.e., it meets the first preset condition; when the resolution of the target image after the first magnification process is lower than the current resolution of the terminal device's display screen, further magnification of the target image may result in blurriness, making it impossible to accurately obtain the information needed by the user, thereby weakening the actual utility and value of the image magnification function and reducing the overall user experience, therefore it is determined that the image is not suitable for further magnification, i.e., it does not meet the first preset condition.

[0015] In one possible implementation of the first aspect, after displaying the target image after the second magnification process on the display screen, the method further includes: stopping image magnification when the target image after the second magnification process meets a second preset condition. The second preset condition includes: obtaining the current resolution of the display screen and the resolution of the target image after the second magnification process; if the resolution of the target image after the second magnification process is less than the current resolution of the display screen, then the target image after the second magnification process meets the second preset condition; if the resolution of the target image after the second magnification process is greater than or equal to the current resolution of the display screen, then the target image after the second magnification process does not meet the second preset condition. Implementing the embodiments of this application, after the target image has undergone multiple magnification processes and is displayed on the display screen, the resolution of the target image after multiple magnification processes is compared with the current resolution of the terminal device's display screen. When the resolution of the target image after multiple magnification processes is already lower than the current resolution of the terminal device's display screen, further magnification may significantly reduce the image's clarity, failing to provide more useful information and practical value. Therefore, it is determined that the image is not suitable for further magnification, and based on this determination, the eye-controlled image magnification is stopped. This avoids blurring of the target image displayed on the display screen after further magnification, reducing user experience, and also avoids unlimited magnification, reducing computational load.

[0016] In one possible implementation of the first aspect, the terminal device further includes a camera; acquiring first gaze information of a user gazing at the target image includes: capturing one or more frames of eye images of the user gazing at the target image using the camera, and determining the first gaze information based on the one or more frames of user eye images. Implementing the embodiments of this application can achieve continuous monitoring of user eye images and accurately acquire the user's gaze area and gaze time.

[0017] In one possible implementation of the first aspect, the method further includes: using the one or more frames of eye images as input to an artificial intelligence (AI) model, analyzing the one or more frames of eye images through the AI ​​model, and outputting the first gaze information. In this embodiment, the AI ​​model can flexibly and accurately obtain the gaze region and gaze duration from user eye images.

[0018] In one possible implementation of the first aspect, the method further includes: detecting whether a user is gazing at the display screen using always-on AO technology, and invoking the camera to acquire an image of the user's eyes. By implementing embodiments of this application, the electronic device can automatically and continuously monitor the user's eye activity using a low-power camera. Upon detecting a user's gaze, the system can respond quickly, initiating eye tracking to accurately acquire the user's gaze information, thus eliminating the need for manual operation by the user. Simultaneously, it effectively reduces energy consumption during the monitoring process, ensuring a balance between device battery life and user experience.

[0019] Secondly, embodiments of the present invention provide a smart terminal device, including: a touch screen, a camera, one or more processors, and one or more memories. The one or more processors are coupled to the touch screen, the camera, and the one or more memories. The one or more memories are used to store computer program code, which includes computer instructions. When the one or more processors execute the computer instructions, the electronic device performs the method in any possible implementation of any of the above aspects.

[0020] Thirdly, this application provides a computer storage medium storing a computer program that, when executed by a processor, implements the image display processing method flow described in any one of the first aspects above.

[0021] Fourthly, embodiments of the present invention provide a computer program product including instructions that, when executed by a computer, enable the computer to perform the image display processing method flow described in any of the first aspects above.

[0022] Fifthly, this application provides an image display processing apparatus that has the function of implementing any of the above-described image display processing methods. This function can be implemented in hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above-described functions.

[0023] Sixthly, this application provides a chip system including a processor for implementing the functions involved in the image display processing method flow described in any one of the first aspects above. In one possible design, the chip system further includes a memory for storing program instructions and data necessary for the image display processing method. This chip system may be composed of chips or may include chips and other discrete devices. Attached Figure Description

[0024] To more clearly illustrate the technical solutions in the embodiments of the present invention or the background art, the accompanying drawings used in the embodiments of the present invention or the background art will be described below.

[0025] Figure 1 is a schematic diagram of the structure of an electronic device provided in an embodiment of this application;

[0026] Figure 2 is a schematic diagram of the system architecture of an electronic device provided in an embodiment of this application;

[0027] Figure 3 is a schematic diagram of an eye-tracking principle provided in an embodiment of this application;

[0028] Figures 4A-4F are schematic diagrams illustrating application scenarios of a set of image display processing methods provided in the embodiments of this application;

[0029] Figure 5 is a desktop schematic diagram of an electronic device provided in an embodiment of this application;

[0030] Figures 6A-6D are schematic diagrams of a set of image gallery viewing interfaces provided in the embodiments of this application;

[0031] Figure 7 is a schematic diagram of an electronic device acquiring a user's eye image according to an embodiment of this application;

[0032] Figures 8A-8G are schematic diagrams illustrating a possible first gaze region division of a set of target images provided in an embodiment of this application;

[0033] Figures 9A-9D are a set of semantic segmentation diagrams provided in the embodiments of this application;

[0034] Figures 10A-10F are another set of semantic segmentation diagrams provided in the embodiments of this application;

[0035] Figure 11 is a schematic diagram of an interface for displaying prompt information on an electronic device according to an embodiment of this application;

[0036] Figures 12A-12D are a set of schematic diagrams of the second enlarged processing procedure provided in the embodiments of this application;

[0037] Figure 13 is a flowchart of an image display processing method provided in an embodiment of this application;

[0038] Figure 14 is a flowchart of another image display processing method provided in an embodiment of this application. Detailed Implementation

[0039] This application provides a method and terminal device for magnifying images by eye gaze.

[0040] The method for magnifying images by eye gaze provided in this application allows users to magnify images by relying on eye gaze. Specifically, after the image is displayed on the display interface, the smart terminal device detects the time the user gazes at a certain area of ​​the image. When the user gazes at a certain area for a longer period than a preset threshold, the image is magnified and displayed on the display interface using the Region of Interest (ROI) within the user's gaze area as the reference area for magnification. This achieves image magnification by eye gaze.

[0041] This method can be applied to electronic devices. The electronic device is a smart terminal device, and this application does not limit the specific type of smart terminal device. For example, the electronic device can be a mobile phone, but it can also include tablet computers, desktop computers, desktop computers with cameras, laptop computers, handheld computers, wearable devices (such as smartwatches, smart bracelets, etc.), augmented reality (AR) devices, virtual reality (VR) devices, artificial intelligence (AI) devices, in-vehicle systems, game consoles, etc.

[0042] The terminology used in the following embodiments of this application is for the purpose of describing particular embodiments only and is not intended to limit the scope of this application. As used in the specification and appended claims of this application, the singular expressions “a,” “an,” “some,” “the,” “the,” “the,” and “this” are intended to include the plural expressions as well, unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in this application refers to any or all possible combinations that include one or more of the listed items.

[0043] Figure 1 shows a schematic diagram of the structure of the electronic device 100 provided in an embodiment of this application.

[0044] Electronic device 100 may include a processor 101, a memory 102, a wireless communication module 103, a mobile communication module 104, an antenna 103A, an antenna 104A, a power switch 105, a sensor module 106, a focusing motor 107, a camera 108, a display screen 109, etc. The sensor module 106 may include a gyroscope sensor 106A, an accelerometer sensor 106B, an ambient light sensor 106C, an image sensor 106D, a proximity sensor 106E, etc. The wireless communication module 103 may include a WLAN communication module, a Bluetooth communication module, etc. All of the above components can transmit data via a bus.

[0045] Processor 101 may include one or more processing units, such as: application processor (AP), modem processor, graphics processing unit (GPU), image signal processor (ISP), controller, video codec, digital signal processor (DSP), baseband processor, and / or neural network processing unit (NPU), etc. Different processing units may be independent devices or integrated into one or more processors.

[0046] Memory 102 can be used to store computer executable program code, which may include instructions. Processor 101 executes various functional applications and data processing of electronic device 100 by running the instructions stored in memory 102. Memory 102 may include a program storage area and a data storage area. In specific implementations, memory 102 may include high-speed random access memory, and may also include non-volatile memory, such as one or more disk storage devices, flash memory devices, or other non-volatile solid-state storage devices.

[0047] The wireless communication function of the electronic device 100 can be implemented through antenna 103A, antenna 104A, mobile communication module 104, wireless communication module 103, modem processor, and baseband processor.

[0048] Antennas 103A and 104A can be used to transmit and receive electromagnetic wave signals. Each antenna in electronic device 100 can be used to cover one or more communication frequency bands. Different antennas can also be reused to improve antenna utilization.

[0049] The mobile communication module 104 can provide wireless communication solutions, including 2G / 3G / 4G / 5G, for use on electronic devices 100.

[0050] The wireless communication module 103 can provide solutions for wireless communication applications on electronic devices 100, including wireless local area networks (WLAN), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), and infrared (IR).

[0051] The gyroscope sensor 106A can be used to determine the motion attitude of the electronic device 100.

[0052] Accelerometer 106B can detect the magnitude of acceleration of electronic device 100 in various directions (generally three axes).

[0053] The 106F infrared optical sensor can capture infrared light signals to enable eye-tracking functionality, analyzing the user's gaze direction and fixation point.

[0054] Electronic device 100 can perform shooting functions through ISP, camera 108, video codec, GPU, display screen 109 and application processor.

[0055] Electronic device 100 can implement display functions through a GPU, display screen 109, and application processor. The GPU is a microprocessor for image processing, connected to the display screen 109 and the application processor. The GPU is used to perform mathematical and geometric calculations and for graphics rendering. Processor 101 may include one or more GPUs, which execute program instructions to generate or modify display information.

[0056] The display screen 109 is used to display images, videos, etc. The display screen 109 includes a display panel. In some embodiments, the electronic device 100 may include one or N display screens 109, where N is a positive integer greater than 1.

[0057] The structures illustrated in the embodiments of this application do not constitute a specific limitation on the electronic device 100. In other embodiments of this application, the electronic device 100 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0058] Figure 2 is a schematic diagram of the system architecture of the electronic device 100 provided in an embodiment of the present invention.

[0059] A layered architecture divides the system into several layers, each with a clear role and function. Layers communicate with each other through software interfaces. In some embodiments, the system is divided into five layers, from top to bottom: application layer, application framework layer, hardware abstraction layer, driver layer, and hardware layer.

[0060] The application layer may include a series of application packages. In this embodiment, the application package may include a camera, a gallery, etc.

[0061] The application framework layer provides application programming interfaces (APIs) and programming frameworks for applications in the application layer. The application framework layer includes some predefined functions.

[0062] As shown in Figure 2, the application framework layer may include a window manager, content provider, view system, resource manager, notification manager, etc.

[0063] The window manager is used to manage windowed applications. It can retrieve screen size, determine the presence of a status bar, lock the screen, and capture screenshots, among other things.

[0064] Content providers store and retrieve data, making that data accessible to applications. This data may include videos, images, audio, made and received phone calls, browsing history and bookmarks, phone books, etc.

[0065] A view system includes visual controls, such as controls for displaying text and controls for displaying images. View systems can be used to build applications. A display interface can consist of one or more views. For example, a display interface including a text notification icon could include views for displaying text and views for displaying images.

[0066] The file explorer provides applications with various resources, such as localized strings, icons, images, layout files, video files, and more.

[0067] The notification manager allows applications to display notifications in the status bar. These notifications can be used to deliver informational messages and can disappear automatically after a short pause, requiring no user interaction. For example, the notification manager can be used to notify users of completed downloads or message alerts. The notification manager can also display notifications as icons or scrolling text in the top status bar, such as notifications from background applications, or as dialog-style notifications on the screen. Examples include displaying text messages in the status bar, emitting sounds, vibrating electronic devices, and flashing indicator lights.

[0068] In this embodiment of the application, the gallery application can utilize the interfaces and services (such as content providers) provided by the application framework layer to effectively access and manage media files (including pictures, videos, etc.) stored on the device, and at the same time use the view system to build a user interface to display these media contents in an intuitive way.

[0069] The hardware abstraction layer is an interface layer located between the application framework layer and the driver layer, providing a virtual hardware platform for the operating system. In this embodiment, the hardware abstraction layer may include a camera hardware abstraction layer, a camera algorithm library, and a sensor control center.

[0070] The camera hardware abstraction layer can provide virtual hardware for camera device 1, camera device 2, or more camera devices. The sensor control center may include an always-on sensor module that continuously monitors environmental conditions and user eye images, triggering image magnification. For example, when the always-on sensor module detects that the user's eye activity meets preset trigger conditions, it activates the gaze algorithm module. The gaze algorithm module receives eye image data or signals from the always-on sensor module and analyzes them to identify the user's gaze direction, gaze duration, and other eye expression information. It then matches the identified eye expression information with preset eye expression control commands; if a match is successful, it triggers the image magnification operation.

[0071] The driver layer is the layer between hardware and software. It includes drivers for various hardware components, such as camera drivers, image processor drivers, display drivers, sensor drivers, and digital signal processor drivers.

[0072] The camera device driver is used to drive the camera sensor to acquire images and to drive the image signal processor to preprocess the images. The digital signal processor driver is used to drive the digital signal processor to process images. The image processor driver is used to drive the graphics processor to process images.

[0073] The hardware layer is the most fundamental layer in a computer system or embedded system, directly involving the existence and operation of physical hardware devices. The hardware layer is the foundation upon which software can run, including all physically tangible and visible computer components, as well as those invisible but equally crucial components such as integrated circuits and circuit boards. The hardware layer can include sensors, image signal processors, digital signal processors, and image processors. Sensors can include sensor 1, time-of-flight (TOF) sensors, infrared optical sensors, and multispectral sensors.

[0074] The terms used in the embodiments of this application are explained below.

[0075] (1) Eye tracking: Eye tracking is a technique that tracks eye movements by measuring the position of the eye's gaze point or the movement of the eyeball relative to the head. This technique aims to monitor the user's eye movements and gaze direction when looking at a specific target, providing important data for understanding human visual behavior and attention allocation. Commonly used techniques include pupil-center corneal reflection (PCCR) and video oculography (VOG), which will be described below.

[0076] ① The Pupil-Center Corneal Reflection (PCCR), also known as the Pukenye image tracking method, is one of the most commonly used methods in eye-tracking technology. The following explanation of the Pupil-Center Corneal Reflection method, with reference to Figure 3 (a schematic diagram of eye-tracking principles), will illustrate this method in detail:

[0077] As shown in Figure 3, the eyeball includes the vitreous body 1, lens 2, cornea 3, pupil 7, and iris 8, among which:

[0078] Vitreous body 1 is a colorless, transparent, gel-like vitreous body that fills the space between the lens and the retina.

[0079] Lens 2 is a biconvex transparent tissue that is fixed and suspended behind the iris and in front of the vitreous body by the suspensory ligaments. It is the only refractive media with accommodation capabilities.

[0080] The cornea 3 is the transparent part at the front of the eyeball and is the first barrier for light to enter the eyeball. In the eyeball structure model provided in the embodiments of this application, the cornea 3 is assumed to be a spherical arc surface.

[0081] The pupil is a small round opening in the center of the iris of an animal or human eye, serving as a passage for light to enter the eye.

[0082] The iris (8) is a disc-shaped membrane with a central opening called the pupil (7).

[0083] In addition, here is a brief introduction to some important concepts or components in eye tracking:

[0084] The pupil center 4 is the geometric center of the pupil 7, which is the center of the ring structure of the pupil 7. In eye-tracking technology, by capturing the movement trajectory of the pupil center, the user's gaze direction and fixation point can be analyzed.

[0085] The Purchin spot is the reflected image formed on the outer surface of the cornea when light (usually infrared or other visible light) shines on the eye. In eye-tracking technology, the Purchin spot is often used as a stable reference point to track eye movements and gaze direction.

[0086] The corneal reflection center 6 refers to the corresponding position on the visual or image sensor of the reflection point or reflection area formed after light (such as infrared light) is reflected on the corneal surface when it shines on the eye.

[0087] Infrared rays, also known as infrared radiation, are electromagnetic waves in the infrared band with wavelengths ranging from 0.76 to 1000 micrometers, which are between visible light and microwaves. They are invisible light with frequencies lower than red light.

[0088] The pupil center line (line of sight) 10 is an imaginary straight line passing through the center of the pupil, which roughly represents the direction in which the eye is looking.

[0089] Image sensor 11, for example, can be image sensor 106D in the electronic device 100 of Figure 1 above. It is a device that converts optical images into electronic signals and is widely used in electronic products such as digital cameras and mobile phones. In mobile phone terminals, the image sensor is mainly responsible for receiving light entering through the lens and converting it into digital signals for subsequent processing and storage.

[0090] Camera 12, for example, may be camera 1 to N108 in the electronic device 100 of Figure 1, is used to acquire eye images.

[0091] Infrared light source 13, electronic device 100, is a non-lighting electric light source whose main purpose is to generate infrared radiation.

[0092] Implementation steps:

[0093] Image acquisition: Using a terminal device equipped with an infrared light source 13 and a camera 12, images of the pupil center 4 and the corneal reflection center 6 are clearly captured;

[0094] Image processing: The acquired images are preprocessed to improve image quality, and then image processing algorithms are used to identify the positions of pupil center 4 and corneal reflection center 6;

[0095] Calculate the relative position: Based on the coordinate information of the pupil center 4 and the corneal reflex center 6, calculate the relative position vector between them. This vector can represent the direction and angle of eyeball rotation;

[0096] Gaze mapping: A calibration process is used to establish a mapping relationship between the pupil-corneal reflectance vector and the gaze point on the computer screen. During calibration, the user needs to observe points appearing at specific locations on the screen (calibration points). The eye-tracking device records the pupil-corneal reflectance vector information of these points and constructs a mapping model.

[0097] Real-time tracking: During the real-time tracking phase, the eye-tracking device continuously acquires images of the pupil and cornea, calculates the pupil-cornea reflection vector, and calculates the user's gaze point position based on the mapping model.

[0098] The pupil-corneal reflex method can accurately track eye movements and calculate the user's gaze point. This method does not require direct contact with the eye and is harmless to the user. Because it uses infrared light sources and image processing technology, the external environment has little impact on the tracking effect.

[0099] It should be noted that the schematic diagram of the eyeball model shown in Figure 3 is an exemplary illustration of the embodiments of this application, and may also be other different forms of models, which are not limited in the embodiments of this application.

[0100] ② Retinal Image Localization Method: The retinal image localization method for eye tracking utilizes the unique physiological structures on the retina, such as the patterns formed by irregular capillaries and the fovea, to track eye movements by calculating changes in the retinal image. This method relies on high-precision image capture and processing technology to achieve accurate tracking of eye movements.

[0101] Basic Principles: Retinal Structure: The retina is a thin membrane specifically responsible for photoreception and image formation. Light entering through the pupil is refracted by the lens and converges onto the retina. The resolving power on the retina is uneven; the macula is the most sensitive area for light perception, and the small fovea in its center is called the fovea centralis, which contains a large number of photoreceptor cells.

[0102] Image capture: Eye-tracking devices capture images on the retina using high-precision cameras. These images contain information about the unique physiological structures of the retina.

[0103] Image processing: Using image processing techniques, the captured retinal images are analyzed and processed to calculate the changes in feature points in the image, thereby inferring the movement trajectory of the eyeball and the position of the fixation point.

[0104] Retinal image localization utilizes the unique physiological structure of the retina for eye tracking, offering high precision and accuracy. It also does not require contact with the eyeball, causing no discomfort or harm to the user. It is suitable for eye tracking in various scenarios and conditions, including eye movement monitoring during tasks such as natural viewing, reading, and driving.

[0105] (2) Semantic Segmentation: Semantic segmentation is an important task in the field of computer vision, aiming to assign each pixel in an image to a specific category label. Unlike image classification and object detection, semantic segmentation requires fine-grained classification of each pixel to achieve a deeper understanding of the image content. This technique can identify different objects in an image and segment them with pixel-level precision.

[0106] (3) Region of Interest (ROI): In image processing, computer vision, and user interaction, ROI specifically refers to the region of interest or focus that a user pays particular attention to when browsing an image. It is a region of the image being processed, delineated with a specific shape (such as a rectangle, circle, ellipse, or irregular polygon) that requires special attention or processing. This region is the focus of image analysis, processing, or user interaction. ROI allows users or systems to select specific areas of interest from the entire image, rather than processing the entire image indiscriminately. This approach can significantly reduce data processing volume, improve processing efficiency, and potentially increase the accuracy of the processing results. Furthermore, the shape, size, and position of the ROI can be flexibly adjusted according to actual needs to adapt to different application scenarios.

[0107] (4) User Interface: The user interface is the medium through which applications or operating systems interact and exchange information with users. It converts the internal form of information into a form that users can understand. The user interface is source code written in specific computer languages ​​such as Java or Extensible Markup Language (XML). This source code is parsed and rendered on the electronic device, ultimately presenting content that the user can recognize. A common form of user interface is the graphical user interface (GUI), which refers to a user interface related to computer operation displayed graphically. It can be visible interface elements such as text, icons, buttons, menus, tabs, text boxes, dialog boxes, status bars, navigation bars, and widgets displayed on the screen of an electronic device.

[0108] First, to facilitate understanding of the embodiments of this application, the specific technical problem to be solved by this application is further analyzed and proposed. In some implementations of the prior art, users can control the zoom in or out of an image by looking at it and blinking. For example, as shown in Figures 4A and 4B, when a user clicks to view image 402 using a gallery application, the electronic device displays the gallery viewing interface 1 shown in Figure 4B. When the user looks at image 402, the electronic device uses the pupil-corneal reflection method, utilizing a front-facing camera 407 and an infrared light source 406, to acquire the user's eye image and gaze position, determine whether the user is looking at the image, and capture the user's conscious blinking behavior. In response to this behavior, the electronic device zooms in on the image with the center of the image as the reference area for zooming in, and displays it on the display screen 411 shown in the gallery viewing interface 2.

[0109] In the aforementioned existing technologies, when a user controls image zooming by eye contact, the electronic device uses the image center as the reference area for zooming, which is likely to be mismatched with the area the user needs to zoom in on, thus failing to meet the user's need to zoom in on a specific area and resulting in a poor user experience.

[0110] In some existing solutions, as shown in Figure 4B, after image 412 is displayed on screen 411, the electronic device can continue to monitor the user's eyes through front-facing camera 407 and infrared light source 406. If the device detects the user blinking consciously again, it triggers image scaling down, restoring the image to its state before magnification, i.e., image 412 is restored to image 402. However, the image resolution may be one or more times higher than the resolution of the electronic device's screen, and the image contains rich details. Magnifying it only once cannot fully extract the information contained in the image, resulting in a waste of image resolution.

[0111] In other solutions of the prior art, electronic devices can also use retinal image localization to obtain information about the user's eye gaze.

[0112] Therefore, based on the problems existing in the prior art, the technical problem to be solved by this application may include the following aspects: providing an image processing and display method that, while enabling eye-controlled image magnification, magnifies the image based on the area of ​​interest most likely to be of interest to the user, and enables multiple magnifications of the image by eye-controlled eye, thereby improving the user experience and reducing power consumption.

[0113] The reference area for magnification refers to the region selected by the user as the starting point for magnification or the center of the magnified image during the image magnification process. The reference area can be selected in various ways, depending on the user's specific actions, such as two-finger swipe magnification, double-tap magnification, or magnification based on eye gaze. For example, when using two fingers to slide on the touchscreen (usually the reverse of a pinch gesture, i.e., fingers spread apart) to adjust the image size, the center point between the two fingers is considered the reference area. As the fingers spread apart, the image is magnified with this center point as the reference, keeping the center point in the center of the magnified image. When the user taps a point on the screen twice consecutively to trigger the partial magnification function, the second tap becomes the reference area. Larger operations will magnify the surrounding image area with this point as the center, making that point the center of the magnified image.

[0114] The following describes the application scenarios involved in the embodiments of this application.

[0115] Scenario 1: Using the image viewing method in this application, images can be zoomed in on via eye gaze in a gallery application.

[0116] Referring to Figures 4A and 4B, Figure 4A illustrates an application scenario of magnifying images through eye gaze in a gallery application, as provided in this embodiment of the application. In this scenario, the user opens the gallery application on the electronic device, selects and views an image, and the electronic device displays a display interface 401 as shown in the gallery viewing interface 1 in Figure 4B. As shown in the gallery viewing interface 1, the display interface 401 may include an image 402, a more information viewing control 403, basic information 404, a return control 405, and an operation menu bar 408. The operation menu bar 408 may include operation options such as share, favorite, edit, delete, and more. When the user consciously gazes at a certain area or object in the image, the electronic device can use pupil-corneal reflex or retinal image localization to track the user's eye movements, thus magnifying the image through the user's eye gaze.

[0117] Scenario 2: Using the image viewing method described in this application, we can zoom in on images in the Weibo application based on eye gaze.

[0118] Referring to Figure 4C, Figure 4C is a schematic diagram of an application scenario in the "Weibo" application provided by an embodiment of the present invention, illustrating the magnification of images through eye gaze. Figure 4C exemplarily shows the browsing interface 421 of the "Weibo" application on an electronic device. The browsing interface 421 may include posts 422 published by other users or the user themselves and an interface menu bar 425. Posts 422 may include images 423 and comment viewing controls 424. Users can click to view images 423 on the electronic device's display screen. Users can also click the comment viewing control 424 to enter comment details, as shown in Figure 4D. The comment interface 431 may include comments 432 and an operation menu bar 434. Comments 432 may include comment images 433, which users can click to view. When a user clicks to view an image in the Weibo application, the electronic device displays the image on the display screen, as shown in Figure 4E. The image viewing interface 441 may include images 442, a collapse control 443, and other operation controls 444. When a user consciously gazes at a certain area or object in the image, the electronic device can use pupil-corneal reflex or retinal image localization methods to track the user's eye movements and achieve image magnification through eye gaze. Optionally, the electronic device may include an infrared light source 445.

[0119] Scenario 3: Using the image viewing method in this application, the paused frame can be enlarged by eye gaze when the video is paused.

[0120] Referring to Figure 4F, which is a schematic diagram of an application scenario for magnifying a paused frame by eye gaze during video pause according to an embodiment of the present invention, the figure includes a video pause interface 1 and a video pause interface 2. In this scenario, video playback can be implemented in other applications capable of playing videos on a display screen, such as a gallery application. When the video is paused, the display screen does not show any content other than the paused frame (image). The video can be played in portrait or landscape mode. After the video pauses, the user can consciously gaze at a certain area or object of the paused frame 451 or 452. The electronic device can use pupil-corneal reflex or retinal image localization to track the user's eye movements and achieve image magnification through eye gaze. Optionally, the electronic device may include an infrared light source 453.

[0121] It is understood that the application scenarios shown in Figures 4A-4F are merely several exemplary implementations in the embodiments of the present invention, and the application scenarios in the embodiments of the present invention include, but are not limited to, the above application scenarios. Optionally, the method in this application can also be applied to, for example, scenarios of image processing and retouching. When a user needs to process a local area of ​​an image, it is usually necessary to first magnify the local area before processing. Applying this method allows the user to zoom in on the image with the local area to be processed as the center by focusing their gaze, and then further edit the local area to be processed. Other scenarios and examples will not be listed or described in detail.

[0122] The user interface involved in the image processing and display method provided in the embodiments of this application is described below with reference to the accompanying drawings.

[0123] The image processing and display method provided in this application embodiment can be applied to the image viewing scenarios shown in Figures 5, 6A, and 6B. This scenario can involve using an electronic device to view a single image through a gallery application or viewing a frame of an image on the screen when a video is paused. Figure 5 is a desktop schematic diagram of an electronic device provided in this application embodiment, wherein the electronic device includes a front-facing camera 504 and a display screen 505 as shown in Figure 5. The display screen 505 of the electronic device displays the desktop shown in Figure 5, which includes a status bar 501 and an application menu bar 502. The status bar 501 includes the operator, current time, network status, signal status, and battery level. As shown in Figure 5, the operator is China Mobile; the current time is 08:08; the network status is Wi-Fi; the signal status is full signal, indicating a strong signal; the black portion of the battery level indicates the remaining battery power of the electronic device. The application menu bar 502 includes icons for at least one application, with the corresponding application name below each icon, such as: Camera, Gallery, Contacts, Dialer, Messages, Weather, etc. The positions of the application icons and their corresponding names can be adjusted according to user preferences, and this application embodiment does not limit this.

[0124] It should be noted that the desktop schematic diagram of the electronic device shown in Figure 5 is an exemplary illustration of the embodiments of this application. The desktop schematic diagram of the electronic device may also be of other styles, and the embodiments of this application do not limit it.

[0125] In some embodiments, a user can click the gallery application icon 503 in the application menu bar 502 shown in Figure 5. The electronic device responds to this operation and displays the interface shown in Figure 6A, which is a schematic diagram of the gallery application interface in an electronic device provided in this application embodiment. As shown in Figure 6A, the interface includes an image search bar 601 and an album menu bar 602. The image search bar 601 can display prompts such as "People, Locations...", allowing users to search for desired images. The album menu bar 602 can include multiple album options such as camera, screenshot, and video. Each album option control uses a thumbnail of the first image in the corresponding album as its cover. By clicking different album option covers, users can view all the corresponding images or videos under that album option on the display screen. Different album options also display the number of images that can be viewed under that album option.

[0126] In some embodiments, as shown in FIG6A, a user can click the camera album option. In response to this operation, the electronic device displays the interface shown in FIG6B, which is a schematic diagram of the interface corresponding to the camera album option provided in this embodiment. As shown in FIG6B, the interface includes a return control 611 and an image option 612. The image option 612 may include one or more image thumbnails corresponding to the original images. The user can swipe down or up on the display screen to browse the image thumbnails and find the image they want to view. Once the user finds the image they want to view, they can click on the thumbnail corresponding to the image they want to view. The electronic device receives the user's click input and further responds to the input by displaying the interface shown in FIG6C, which is a schematic diagram of the interface for viewing images through the gallery application provided in this embodiment.

[0127] As shown in Figure 6C, the interface includes a view image 621, a basic information bar 622, an image operation menu bar 623, a return control 624, and an information viewing control 625. The basic information bar 622 can include time and address information, such as "August 5, 2024, Nanshan District, Shenzhen". The image operation menu bar 623 can include operations such as "share", "favorite", "edit", "delete", and "more". Clicking the information viewing control 625 allows users to view more information about the image, including but not limited to storage path, image size, and image resolution.

[0128] It should be noted that the user interface diagrams of the electronic devices shown in Figures 6A-6C are exemplary illustrations of the embodiments of this application. The user interface diagrams of the electronic devices may also be of other styles, and the embodiments of this application do not limit them.

[0129] The following section, with reference to Figures 5-12D, 13, and 14, explains how electronic devices can control image magnification by eye movements.

[0130] The electronic device can be the electronic device 100 shown in Figure 1, which includes a front-facing camera and a display screen, and may also include an infrared light source, and the electronic device can support the system architecture shown in Figure 2.

[0131] First, the method for implementing eye-controlled image magnification using an electronic device will be explained with reference to Figures 5-10F and Figure 13. Figure 13 is a flowchart of an image display processing method provided by an embodiment of this application.

[0132] S101: The electronic device displays the target image on the screen.

[0133] Specifically, the process of viewing and displaying an image on the screen can be seen in Figures 5, 6A, and 6B and their corresponding text descriptions, which will not be repeated here. It is understood that the interface illustrated in the embodiments of the present invention does not constitute a specific limitation on the electronic device, and the process of viewing and displaying an image on the screen is not limited to the process shown in Figures 5, 6A, and 6B.

[0134] S102: Obtain the first gaze region and the first gaze time of the target image being gazed at by the user.

[0135] The first fixation time can be the time during which a user continuously fixates on a specific area (by analyzing multiple consecutive frames of user eye images acquired by a camera to determine that the user's eye fixation position has not changed). Furthermore, the first fixation time can also include the time during which the user primarily fixates on a fixed area over a longer period, even if the area of ​​fixation changes briefly during this time, the user's primary area of ​​fixation remains relatively fixed within this period.

[0136] Electronic devices can first determine the first fixation area by capturing the first few frames of user eye images, and then determine the first fixation time, or they can first determine the first fixation time, and then determine the first fixation area based on the acquired user eye images.

[0137] Specifically, the electronic device displays a target image on the screen. When the user has a need to zoom in, they consciously look at a certain area of ​​the target image. The electronic device continuously captures multiple frames of the user's eye images at relatively long time intervals using a front-facing camera. The electronic device analyzes the acquired multi-frame eye images with long time intervals to determine whether the user is looking at the image. If the user's gaze area does not shift significantly, it is determined that the user is looking at the image. The electronic device then captures multiple high-resolution images of the user's eye at shorter time intervals using a front-facing camera. The electronic device analyzes the acquired multi-frame eye images with shorter time intervals to determine first gaze information, which includes a first gaze area and a first gaze time.

[0138] In one possible implementation, when the user completes the click input of the image thumbnail as shown in FIG6B, and the electronic device displays the interface shown in FIG6C, the electronic device begins to acquire first gaze information. Exemplarily, the electronic device first determines the first gaze time, and then determines the first gaze region based on the acquired user eye images. As shown in FIG7, the electronic device displays the target image 702 on the display screen, and continuously captures multiple frames of the user's eye images per second using the front-facing camera 701 at a low sampling rate of 50 frames per second, acquiring 50 frames of user eye images. The electronic device analyzes these 50 frames of user eye images. If the distance between the center points of the gaze region obtained after analyzing more than 30 consecutive frames of user eye images does not exceed the maximum pixel distance, then it is determined that the user is gazing at the image. Here, the maximum pixel distance refers to the maximum permissible pixel offset used to determine whether the user's gaze region has undergone a significant change. If no significant displacement of the gaze area is obtained after analyzing more than 30 consecutive frames of the user's eye images, the front-facing camera 701 will continue to capture multiple frames of the user's eye images at a low sampling rate of 50 frames per second, acquiring the next set of 50 frames of the user's eye images every second, until it is detected that no significant displacement of the gaze area is obtained after analyzing more than 30 consecutive frames of the user's eye images.

[0139] For example, after the electronic device determines that the user is gazing at an image, as shown in Figure 7, it continuously captures multiple frames of the user's eye images for 1.2 seconds using the front-facing camera 701 at a high sampling rate of 100 frames per second, acquiring 120 frames of user eye images. The 120 frames are analyzed to obtain the area the user is gazing at in each frame. Images with 50 or more consecutive frames corresponding to which the user's gaze area has not undergone significant displacement are detected, and the number of consecutive image frames is counted. The first gaze time is calculated using the largest consecutive image frame count. Specifically, the calculation method is: T1=(N1-1)·0.01

[0140] Where T1 is the first gaze time, and N1 is the first gaze region of the target image obtained by executing step S202 and the maximum number of consecutive image frames detected during the first gaze time.

[0141] For example, the electronic device can determine the first gaze region after determining the first gaze time. Specifically, the electronic device can determine the first gaze region of the target image on the display screen based on the first gaze region of the target image obtained in step S202 and the continuous image corresponding to the largest number of consecutive image frames detected in the first gaze time.

[0142] The first gaze region can be defined by coarse-precise coordinates. Before determining the first gaze region, the electronic device divides the image into possible first gaze regions based on the aspect ratio, and then analyzes the captured image of the user's eyes to determine the first gaze region.

[0143] For example, as shown in Figures 8A-8G, for images with common aspect ratios in electronic devices, the electronic device can divide the possible first gaze region as shown in Figures 8A-8G:

[0144] 1) 1:1

[0145] As shown in Figure 8A, an image with an aspect ratio of 1:1 is divided into four possible first viewing regions 801, 802, 803, and 804 by a straight line connecting the centers of two parallel sides of the image.

[0146] 2) 16:9

[0147] As shown in Figure 8B, for an image with an aspect ratio of 16:9, the center of the long side (the side with a length ratio of 16:9 to the adjacent side) is connected to the center of the other long side. Similarly, the center of the short side (the side with a length ratio of 9:16 to the adjacent side) is connected to the center of the other short side, dividing the image into four possible first viewing regions 811, 812, 813, and 814.

[0148] 3) 9:16

[0149] As shown in Figure 8C, for an image with an aspect ratio of 9:16, the center of the short side (the side with a length ratio of 9:16 to the adjacent side) is connected to the center of the other short side. Similarly, the center of the long side (the side with a length ratio of 16:9 to the adjacent side) is connected to the center of the other long side, dividing the image into four possible first viewing regions 821, 822, 823, and 824.

[0150] 4) 4:3

[0151] As shown in Figure 8D, for an image with an aspect ratio of 4:3, the center of the long side (the side with a length ratio of 4:3 to the adjacent side) is connected to the center of the other long side. Similarly, the center of the short side (the side with a length ratio of 3:4 to the adjacent side) is connected to the center of the other short side, dividing the image into four possible first viewing regions 831, 832, 833, and 834.

[0152] 5) 3:4

[0153] As shown in Figure 8E, for an image with an aspect ratio of 3:4, the center of the short side (the side with a length ratio of 3:4 to the adjacent side) is connected to the center of the other short side. Similarly, the center of the long side (the side with a length ratio of 4:3 to the adjacent side) is connected to the center of the other long side, dividing the image into four possible first viewing regions 841, 842, 843, and 844.

[0154] 6) 3:2

[0155] As shown in Figure 8F, for an image with an aspect ratio of 3:2, the center of the long side (the side with a length ratio of 3:2 to the adjacent side) is connected to the center of the other long side. Similarly, the center of the short side (the side with a length ratio of 2:3 to the adjacent side) is connected to the center of the other short side, dividing the image into four possible first viewing regions 851, 852, 853, and 854.

[0156] 7)2:3

[0157] As shown in Figure 8G, for an image with an aspect ratio of 2:3, the center of the shorter side (the side with a length ratio of 2:3 to the adjacent side) is connected to the center of the other shorter side. Similarly, the center of the longer side (the side with a length ratio of 3:2 to the adjacent side) is connected to the center of the other longer side, dividing the image into four possible first viewing regions 861, 862, 863, and 864.

[0158] The images viewed by users can also be images with any aspect ratio, and this embodiment of the invention does not impose any limitations.

[0159] In one possible implementation, the electronic device can calculate and determine the first gaze region of the target image on the user's viewing display screen based on all continuous images corresponding to the maximum number of continuous image frames in the above implementation, or it can calculate and determine the first gaze region of the target image on the user's viewing display screen based on the first frame, one or two middle frames and the last frame of the continuous images, or it can calculate and determine the first gaze region of the target image on the user's viewing display screen based on the last frame of the continuous images corresponding to the maximum number of continuous image frames.

[0160] In one possible implementation, the electronic device can use an artificial intelligence (AI) model to analyze the user's eye images. The electronic device takes multiple frames of the user's eye images captured by the front-facing camera at a high or low sampling rate as input to the AI ​​model, and outputs the first gaze region and the first gaze time after calculation.

[0161] It should be noted that the schematic diagrams of possible first gaze regions of the target image shown in Figures 8A-8G are exemplary illustrations of embodiments of this application. Other division methods may be used when dividing possible first gaze regions, and this application does not limit them.

[0162] In the above implementation, after the user clicks on the image thumbnail as shown in Figure 6B and displays the interface shown in Figure 6C, the user can again click on the image 621 as shown in Figure 6C. In response to this input, the electronic device displays a full-screen image viewing interface as shown in Figure 6D. The display screen does not show any content other than the image 631. This input will interrupt the process of detecting whether the user is looking at the image triggered by the click on the image thumbnail as shown in Figure 6B, and will trigger a new process of detecting whether the user is looking at the image. That is, the electronic device re-executes step S202 to obtain the first gaze area and the first gaze time of the target image being gazed at by the user.

[0163] S103: Determine whether the first fixation time is greater than or equal to the first threshold.

[0164] Specifically, after acquiring the first gaze time, the electronic device compares the first gaze time with a preset first threshold. If the first gaze time is less than the first threshold, the electronic device repeats step S202 to acquire the first gaze region and the first gaze time of the user's gaze on the target image. If the first gaze time is greater than or equal to the first threshold, step S204 is executed to determine the first target region of interest within the first gaze region. The first threshold is used to determine whether the user is consciously gazing at a certain area of ​​the image displayed on the screen in order to achieve image magnification.

[0165] In one possible implementation, the electronic device can use an AI model to determine whether the user's first gaze time is greater than or equal to a first threshold. The electronic device uses the first gaze time as input to the AI ​​model, analyzes it, and outputs the relationship between the first gaze time and the first threshold. The relationship includes two possibilities: greater than or equal to, and less than.

[0166] For example, the first threshold can be set to 0.9 seconds. When the first gaze time is less than 0.9 seconds, the electronic device repeatedly executes step S202 to obtain the first gaze region and the first gaze time of the user's gaze on the target image; when the first gaze time is greater than or equal to 0.9 seconds, step S204 is executed. If the first threshold is exceeded, the first target interest region within the first gaze region is determined based on the first gaze region and the content of the first gaze region.

[0167] S104: If the first threshold is exceeded, then based on the first gaze region and the content of the first gaze region, determine the first target region of interest within the first gaze region.

[0168] First, the electronic device performs semantic segmentation on the target image displayed on the screen. Specifically, the electronic device performs semantic segmentation on the target image according to a first semantic category, identifying Q second-level semantic regions in the target image, where Q is a positive integer less than M; it then performs semantic segmentation on each of the Q second-level semantic regions according to a second semantic category, identifying one or more first-level semantic regions within each second-level semantic region; finally, the first-level semantic regions corresponding to the Q second-level semantic regions are determined as the M first-level semantic regions.

[0169] In one possible implementation, the first semantic category includes one or more of people, animals, plants, and scenery; the second semantic category includes one or more of human facial features, hair, arms, legs, hands, feet, clothing, tree crowns, tree trunks, flower crowns, plant stems, animal facial features, and limbs.

[0170] In one possible implementation, when determining Q second-level semantic regions based on a first semantic category, a single target image may include multiple objects corresponding to the first semantic category. When objects corresponding to the first semantic category overlap, among every two overlapping objects, the object with relatively complete semantics and positioned higher in the image corresponding to the first semantic category is called the occluder, and the object whose portion is obscured by the occluder and positioned lower in the image is called the occluded object. Specifically, when the occluder only partially covers the occluded object, causing part of the occluded object to be invisible, but not segmented into two or more independent regions that are not directly connected in space, simple image analysis techniques, such as color thresholding and edge detection, are first used to preliminarily identify the occluder in the image. Based on the detected occluder, an edge detection algorithm is further used to extract the precise edge of the occluder. The edge of the occluded object is determined based on the precise edge of the occluder and the unoccluded portion of the occluded object. The closed regions enclosed by the edges of the occluder and the occluded object are respectively determined as Q second-level semantic regions.

[0171] For example, the semantic segmentation diagram 1 shown in Figure 9A is a schematic diagram of the result obtained by semantically segmenting the target image 901 according to the first semantic category. As shown in Figure 9A, the target image 901 contains two objects corresponding to the first semantic category, including a person and a plant. The person and the plant corresponding to the first semantic category have overlapping parts. The person is identified as an occluder, and the plant is identified as an occluded object. The electronic device performs semantic segmentation on the target image 901 displayed on the screen according to the first semantic category, and obtains two second-level semantic regions, namely the closed regions enclosed by the dashed boxes in Figure 9A, namely the second-level semantic region 902 and the second-level semantic region 903, where Q equals 2. The semantic segmentation of the electronic device first ensures the semantic integrity of the second-level semantic region 902, and determines the edge of the second-level semantic region 903 based on the segmentation edge of the second-level semantic region 902 and the unoccluded part of the second-level semantic region 903, and tries to ensure the integrity of the second-level semantic region 903 without destroying the edge of the second-level semantic region 902.

[0172] In one possible implementation, when the target image contains multiple objects corresponding to a first semantic category, and the objects corresponding to the first semantic category overlap with two or more other objects, but the overlapping portion still satisfies that the occluding object only partially covers the occluded object in every two objects corresponding to the first semantic category, causing part of the occluded object to be invisible, but it is not segmented into two or more independent regions that are not directly connected in space, and is still segmented according to the semantic segmentation method described above.

[0173] For example, the semantic segmentation diagram 2 shown in Figure 9B is a schematic diagram of the result obtained by semantically segmenting the target image 911 according to the first semantic category. As shown in Figure 9B, the target image 911 contains three objects corresponding to the first semantic category, including a male figure, a female figure, and a tree. The male figure and the female figure have overlapping parts, with the male figure acting as an occluder and the female figure as the occluded object; the female figure and the tree also have overlapping parts, with the tree acting as an occluder and the female figure as the occluded object. The female figure has overlapping parts with the two objects corresponding to the first semantic category, and both of these overlapping parts satisfy the condition that the occluder only partially covers the occluded object, making part of the occluded object invisible, but not segmenting it into two or more independent regions that are not directly connected in space. The second-level semantic region corresponding to the male character is determined, and semantic segmentation is performed according to the first semantic category to obtain the second-level semantic region 912. Simultaneously, the edges of the second-level semantic region 912 in the overlapping area between the male and female characters are obtained. The second-level semantic region corresponding to the tree is determined, and semantic segmentation is performed according to the first semantic category to obtain the second-level semantic region 913. Simultaneously, the edges of the second-level semantic region 913 in the overlapping area between the tree and the female character are obtained. Based on the edges of the second-level semantic region 912 in the overlapping area between the male and female characters, the edges of the second-level semantic region 913 in the overlapping area between the tree and the female character, and the unoccluded portion of the female character (i.e., the portion not overlapping with other objects corresponding to the first semantic category), the second-level semantic region 914 corresponding to the female character is determined. At this point, Q equals 3.

[0174] In one possible implementation, if the occlusion not only partially covers the occluded object, but also covers it to a degree sufficient to segment the originally continuous occluded object into two or more spatially disjoint regions, a deep learning model is used to perform semantic segmentation on the target image according to a first semantic category. Specifically, the image to be processed is input into a trained model, which performs a series of forward propagation calculations, including feature extraction, feature fusion, and contextual understanding, to comprehensively analyze the occlusion situation in the image. Based on its understanding of the image content, the model infers the edges of the occluded object using prior knowledge. The model outputs Q determined second semantic regions.

[0175] In one possible implementation, each of the Q second-level semantic regions is semantically segmented according to the second semantic category, and M first-level semantic regions are identified in each second-level semantic region, where M is a positive integer greater than Q.

[0176] For example, the semantic segmentation diagram 3 shown in Figure 9C is a schematic diagram of the result obtained by semantically segmenting the target image 921 according to the second semantic category. The target image 921 is the same image as the target image 901 in Figure 9A. Furthermore, the semantic segmentation of the target image 921 according to the second semantic category is performed based on the semantic segmentation diagram 1 shown in Figure 9A; that is, semantic segmentation is performed on the result image obtained by segmentation according to the first semantic category according to the second semantic category. In Figure 9A, Q is 2, meaning there are two second-level semantic regions. In the second-level semantic region 902 of Figure 9A, according to the second semantic category, the first-level semantic regions 922, 923, 924, 925, 926, 927, 928, 929, and 9210 shown in Figure 9C were identified, totaling nine first-level semantic regions. Among them, first-level semantic region 922 is the left eye, first-level semantic region 923 is the right eye (here, left and right are determined by the relative positions of the two eyes in the image), first-level semantic region 924 is the nose, first-level semantic region 925 is the mouth, and first-level semantic region 926 is the nose. The image shows the right ear, the left hand in first-level semantic region 927, the right hand in first-level semantic region 928 (the left and right hands are determined by their relative positions in the image), the jacket pocket in first-level semantic region 929, and the bow tie in first-level semantic region 9210. In second-level semantic region 903 of Figure 9A, first-level semantic regions 9211 and 9212 (as shown in Figure 9C) are identified according to the second semantic category, totaling two first-level semantic regions. First-level semantic region 9211 represents the tree trunk, and first-level semantic region 9212 represents the tree crown. At this point, M is 11. Therefore, the two second-level semantic regions in Figure 9A can be further divided into 11 first-level semantic regions.

[0177] It should be noted that the semantic segmentation diagrams of the target images shown in Figures 9A-9C are exemplary illustrations of embodiments of this application. Other semantic segmentation methods can be used when performing semantic segmentation on the target images, and this application does not limit them.

[0178] Next, from the M first-level semantic regions, N first-level semantic regions that intersect with the first gaze region determined in step S202 are identified. These N first-level semantic regions that intersect with the first gaze region may include first-level semantic regions segmented according to the second semantic category and completely covered by the first gaze region, and first-level semantic regions segmented according to the second semantic category and partially covered by the gaze region.

[0179] Based on the preset magnification priority of the N first-level semantic regions and the area occupied by the N first-level semantic regions in the first gaze region, the first target region of interest is determined.

[0180] The preset magnification priority is a set of priority rules for object magnification display pre-set in electronic devices based on factors such as product design concepts and user experience. These rules can be based on various factors such as object type (e.g., facial features, limbs, clothing, plants, animals), size, position, and color. The magnification priority rules are encoded and stored in the firmware of the electronic device or an updatable database so that the electronic device can call them at any time.

[0181] In one possible implementation, after determining the N first-level semantic regions covered within the user's first gaze region, the electronic device traverses the N first-level semantic regions covered within the first gaze region, which are semantically segmented according to a second semantic category. Based on a preset magnification priority, it determines the magnification priority of the first-level semantic regions covered within each first gaze region and counts the pixels of each first-level semantic region to obtain the pixel count as the area corresponding to each first-level semantic region. When the first-level semantic region covered by the first gaze region is a first-level semantic region segmented according to the second semantic category and completely covered by the first gaze region, the total number of pixels in this first-level semantic region is counted as the area corresponding to this first-level semantic region. When the first-level semantic region covered by the first gaze region is a first-level semantic region segmented according to the second semantic category and partially covered by the gaze region, only the pixel count of the portion of the first-level semantic region covered by the first gaze region is counted as the area of ​​this first-level semantic region. The total number of pixels in the first gaze region is then counted as the area of ​​the first gaze region.

[0182] In one possible implementation, the electronic device calculates the magnification weight of each of the N first-level semantic regions based on the magnification priority and area of ​​the N first-level semantic regions covered by the first gaze region, and selects the first-level semantic region with the largest magnification weight as the first target region of interest. If the first-level semantic region with the largest magnification weight is a first-level semantic region segmented according to the second semantic category and completely covered by the first gaze region, then this first-level semantic region is selected as the first target region of interest. If the first-level semantic region with the largest magnification weight is a first-level semantic region segmented according to the second semantic category and partially intersects with the first gaze region, then the intersection of this first-level semantic region and the first gaze region is selected as the first target region of interest.

[0183] For example, Figure 9D is a semantic segmentation diagram 4. Figure 9D shows the position of the first gaze region in the target image and the first-level semantic regions that intersect with the first gaze region. As shown in Figure 9D, the target image 931 has a resolution of 2340x1080 pixels and can be divided into four possible first gaze regions. The first gaze region 932 that the user gazes at is the upper left rectangular area of ​​the target image 931. The first-level semantic regions covered by the first gaze region include first-level semantic regions 933, 934, 935, 936, and 937. Among them, the first-level semantic region 934 is the right eye (here, the left and right in the left and right eyes are determined by the relative positions of the two eyes in the image), the first-level semantic region 935 is the nose, the first-level semantic region 936 is the mouth, and the first-level semantic region 937 is the bow tie. The first-level semantic regions (right eye) 934 and (mouth) 936 are both first-level semantic regions that intersect with the first gaze region. When counting pixels, only the intersection with the first gaze region 932 is counted, in which case N is 5. The electronic device presets the magnification priority of the first-level semantic regions as an integer from 0 to 10, and each first-level semantic region corresponds to an integer representing the magnification priority. The first-level semantic region (left eye) 933 has a magnification priority of 7 and an area of ​​60,000 pixels. The first-level semantic region (right eye) 934 has a magnification priority of 7, and the area of ​​its intersection with the first gaze region is 40,000 pixels. The first-level semantic region (nose) 935 has a magnification priority of 6 and an area of ​​90,000 pixels. The first-level semantic region (mouth) 936 has a magnification priority of 6, and the area of ​​its intersection with the first gaze region is 100,000 pixels. The first-level semantic region (bow tie) 937 has a magnification priority of 3 and an area of ​​170,000 pixels; the area of ​​the first gaze region is 631,800 pixels.

[0184] Optionally, the amplification weight W1 of the first-level semantic region is calculated as follows: W1 = L n1 ·V l1 +A n1 ·V a1

[0185] Where L n1 V represents the amplification priority of the first-level semantic region after normalization. l1 The weight of the priority amplification for the first-level semantic region is a pre-set value for the electronic device, A n1 V represents the area of ​​the first-level semantic region after normalization. a1 The area weight is a preset value for electronic devices.

[0186] Optionally, L n1 The calculation formula is:

[0187] Where L r1 L is the magnification priority determined after the electronic device traverses the first-level semantic region covered by the first gaze region. min1 The maximum value of the amplification priority of the first-level semantic region preset for electronic devices, L max1 The maximum value of the amplification priority of the first-level semantic region preset for electronic devices.

[0188] Optionally, A n1 The calculation formula is:

[0189] Where A o1 A is the area of ​​each first-level semantic region covered by the first gaze region traversed by the electronic device. r1 This represents the area of ​​the first gaze region.

[0190] For example, the weight V of the amplification priority l1 =0.6, area weight V a1 =0.4, L corresponding to the first-level semantic region (left eye) 933 shown in Figure 9D r1 It is 7, L max1 For 10, L min1 If it is 1, then Corresponding A o1 For 60000, A r1 If it is 631800, then The corresponding amplification weight W1 = L n1 ·V l1 +A n1 ·V a1 =0.666666·0.6+0.094967·0.4=0.437986. Similarly, the magnification weights of other first-level semantic regions covered within the first gaze region 932 shown in Figure 9D are calculated in the same way. The magnification weights are: first-level semantic region (right eye) 934 W1=0.425324, first-level semantic region (nose) 935 W1=0.391984, first-level semantic region (mouth) 936 W1=0.398501, and first-level semantic region (bow tie) 937 W1=0.244118. The first-level semantic region (left eye) 933 with the largest magnification weight is selected as the first target region of interest.

[0191] In one possible implementation, the electronic device can provide a user feedback mechanism that allows users to adjust the magnification priority rules or select specific objects to be magnified according to their preferences.

[0192] S105: Perform a first magnification process on the target image based on the first target region of interest, and display it on the display screen.

[0193] Specifically, the first magnification process can be the first magnification process of the target image, or it can be any magnification process of the target image other than the first magnification process.

[0194] Specifically, when performing a first magnification process on the target image based on the first target region of interest, the first target region of interest can be directly used as the reference region for magnification, meaning the magnified target image is centered on the first target region of interest. Alternatively, the reference region for magnification can be determined based on the first target region of interest and a larger region containing the first target region of interest, and the magnified target image can be centered on the larger region containing the first target region of interest.

[0195] In one possible implementation, if the semantics of the first target region of interest are complete, then the first target region of interest is used as the first reference region for the first magnification process, and the target image is magnified at a fixed magnification based on the first reference region.

[0196] The definition of the reference area can be found in the description of the reference area in the application scenarios involved in the embodiments of this application, and will not be repeated here.

[0197] For example, as shown in FIG10A, the electronic device displays a target image 1001 that has not undergone the first magnification process on the display screen. The electronic device determines the first gaze region 1002 that the user is looking at and determines the first target region of interest 1003. The first target region of interest 1003 is semantically complete and is directly used as the reference region for the first magnification process. The target image is magnified by a fixed magnification of 1.5 times and displayed on the display screen, as shown in FIG10B. At this time, the electronic device displays a target image 1011 that has undergone the first magnification process on the display screen. The target image 1011 takes the first target region of interest 1003 as the center of the image.

[0198] In one possible implementation, if the semantics of the first target region of interest are incomplete, it is determined whether the first target region of interest belongs to a part of a first-level semantic region or the whole region. If the first target region of interest belongs to a part of a first-level semantic region, then the first-level semantic region to which the first target region of interest belongs is used as the second reference region for the first amplification process.

[0199] For example, as shown in FIG10C, the electronic device displays a target image 1021 without the first magnification processing on the display screen. The electronic device determines a first gaze region 1022 that the user is looking at, and determines a first target region of interest 1023 in the first gaze region 1022. The first target region of interest is an arc-shaped area that is cut out by the edge of the first-level semantic region 1024 and covered by the first gaze region. The first target region of interest 1023 is semantically incomplete, and the first target region of interest 1023 belongs to a part of the first-level semantic region 1024. Then, the electronic device determines the first-level semantic region 1024 as the reference area for the first magnification processing, and performs the first magnification processing on the target image 1021 at a fixed magnification of 1.5 times, and displays it on the display screen, as shown in FIG10D. At this time, the electronic device displays a target image 1031 after the first magnification processing on the display screen, with the first-level semantic region 1024 as the center of the image.

[0200] In one optional implementation, if the semantics of the first target region of interest are incomplete and the first target region of interest belongs to the entire region of a certain first-level semantic region, the second-level semantic region to which this first-level semantic region belongs is used as the second reference region for the first magnification process; the target image is magnified based on the second reference region.

[0201] For example, as shown in FIG10E, the electronic device displays a target image 1041 without the first magnification processing on the display screen. The electronic device determines the first gaze region 1042 that the user is looking at, and determines the first target interest region 1043 in the first gaze region 1042. The first target interest region 1043 is semantically incomplete, and the first target interest region 1043 belongs to the entire region of the first-level semantic region, that is, the first target interest region 1043 is also the first-level semantic region 1043. Then the electronic device determines the second-level semantic region 1044 to which the first target interest region 1043 belongs as the reference region for the first magnification processing. The second-level semantic region 1044 belongs to the occluded object, and is partially occluded and does not become multiple unconnected regions. The target image 1041 is magnified by a fixed magnification of 1.5 times and displayed on the display screen, as shown in FIG10F. At this time, the electronic device displays a target image 1051 after the first magnification processing on the display screen. The target image 1051 takes the second-level semantic region 1044 as the center of the image.

[0202] In one optional implementation, if the semantics of the first target region of interest are incomplete and the first target region of interest belongs to the entire region of a certain first-level semantic region, the detection step S204 determines whether the first-level semantic region with the largest magnification weight calculated in the first target region of interest within the first gaze region is semantically complete. If this first-level semantic region is semantically complete, it is used as the second reference region for the first magnification process. If this first-level semantic region is semantically incomplete, the detection step S204 determines whether the first-level semantic region with the second largest magnification weight calculated in the first target region of interest within the first gaze region is semantically complete. If this first-level semantic region is semantically complete, it is used as the second reference region for the first magnification process. If it is incomplete, the same method is used to traverse the N first-level semantic regions covered in the first gaze region from high to low magnification weight. If all N first-level semantic regions covered in the first gaze region are semantically incomplete, the second-level semantic region to which the first target region of interest belongs is used as the second reference region for the first magnification process. Then, based on the determined second reference region, the target image is magnified at a fixed magnification.

[0203] The following describes a method for implementing eye-controlled image magnification using an electronic device, with reference to Figures 11-12D and 14. Figure 14 is a flowchart of another image display processing method provided in an embodiment of this application.

[0204] S201: The electronic device displays the target image on the display screen.

[0205] S202: Obtain the first gaze region and the first gaze time of the target image being gazed at by the user.

[0206] S203: Determine whether the first fixation time is greater than or equal to the first threshold.

[0207] S204: If the first threshold is exceeded, then based on the first gaze region and the content of the first gaze region, determine the first target region of interest within the first gaze region.

[0208] S205: Perform a first magnification process on the target image based on the first target region of interest, and display it on the display screen.

[0209] The specific process of steps S201-S205 above can be found in the relevant description of steps S101-S105 in Figure 13, and will not be repeated here.

[0210] S206: Determine whether the resolution of the target image after the first magnification process is greater than or equal to the resolution of the display screen.

[0211] The electronic device acquires the current resolution of the display screen, which is determined by multiplying the width and height of the display screen in pixels, i.e., by multiplying the number of pixels in the horizontal direction by the number of pixels in the vertical direction. The electronic device also acquires the resolution of the first magnified target image, similarly determined by multiplying the width and height of the target image in pixels, i.e., by multiplying the number of pixels in the horizontal direction by the number of pixels in the vertical direction. The resolution of the first magnified target image is compared to whether it is greater than or equal to the current resolution of the display screen. If the resolution of the first magnified target image is greater than or equal to the current resolution of the display screen, step S208 is triggered to acquire the second gaze time and the second gaze area of ​​the user gazing at the first magnified target image; if the resolution of the first magnified target image is lower than the current resolution of the display screen, step S207 is triggered to stop the eye-controlled image magnification process.

[0212] In one possible implementation, when determining whether the resolution of the target image after the first magnification process is greater than or equal to the current resolution of the electronic device display screen, the width pixel count of the target image resolution after the first magnification process is compared with the width pixel count of the current resolution of the electronic device display screen, and the height pixel count of the target image resolution after the first magnification process is compared with the height pixel count of the current resolution of the electronic device display screen. Only when the width pixel count of the target image resolution after the first magnification process is greater than or equal to the width pixel count of the current resolution of the electronic device display screen and the height pixel count of the target image resolution after the first magnification process is also greater than or equal to the height pixel count of the current resolution of the electronic device display screen is the resolution of the target image after the first magnification process considered to be greater than or equal to the current resolution of the electronic device display screen.

[0213] For example, if the resolution of the electronic device display screen is 2400×1600 pixels, and the resolution of the target image after the first magnification process is 3264×2448 pixels, the width and height pixel counts of the two are compared respectively. The width pixel count of the target image after the first magnification process is 3264, which is greater than the width pixel count of the electronic device (2400). At the same time, the height pixel count of the target image after the first magnification process is 2448, which is greater than the width pixel count of the electronic device (1600). This satisfies the condition that the width pixel count of the target image after the first magnification process is greater than or equal to the width pixel count of the current resolution of the electronic device display screen, and the height pixel count of the target image after the first magnification process is also greater than or equal to the height pixel count of the current resolution of the electronic device display screen. Then, step S208 is executed to obtain the second gaze time and the second gaze area of ​​the user gazing at the target image after the first magnification process.

[0214] If the resolution of the target image after the first magnification process is 3200×1200 pixels, the width of the target image after the first magnification process is 3200 pixels, which is greater than the width of the electronic device (2400 pixels), and the height of the target image after the first magnification process is 1200 pixels, which is less than the width of the electronic device (1600 pixels), then the conditions are not met: the width of the target image after the first magnification process is greater than or equal to the width of the electronic device's current display screen resolution, and the height of the target image after the first magnification process is also greater than or equal to the height of the electronic device's current display screen resolution. Then, step S207 is executed to stop the eye-controlled image magnification process.

[0215] If the resolution of the target image after the first magnification process is 1200×1600 pixels, the width of the target image after the first magnification process is 1200 pixels, which is less than the width of the electronic device (2400 pixels), and the height of the target image after the first magnification process is 1600 pixels, which is equal to the width of the electronic device (1600 pixels), then the conditions are not met: the width of the target image after the first magnification process is greater than or equal to the width of the electronic device's display screen, and the height of the target image after the first magnification process is also greater than or equal to the height of the electronic device's display screen. Then, step S207 is executed to stop the eye-controlled image magnification process.

[0216] If the resolution of the target image after the first magnification process is 1600×1200 pixels, the width of the target image after the first magnification process is 1600 pixels, which is less than the width of the electronic device (2400 pixels), and the height of the target image after the first magnification process is 1200 pixels, which is less than the width of the electronic device (1600 pixels), then the conditions that the width of the target image after the first magnification process is greater than or equal to the width of the current resolution of the electronic device's display screen and the height of the target image after the first magnification process is greater than or equal to the height of the current resolution of the electronic device's display screen are not met, then step S207 is executed to stop the eye-controlled image magnification process.

[0217] S207: Stop eye-tracking image zoom-in process.

[0218] Specifically, when the target image meets preset conditions, the electronic device stops the eye-controlled image magnification process. The target image can be an image displayed on the screen after a first magnification process, or it can be a target image displayed on the screen after a second magnification process. The preset conditions can be a first preset condition, i.e., the resolution of the target image is greater than or equal to the current resolution of the display screen, or a second preset condition, i.e., the resolution of the target image is less than the current resolution of the display screen.

[0219] In one possible implementation, while the electronic device stops the currently executing eye-controlled image zooming process, it can display a prompt message on the electronic device's screen, informing the user that zooming cannot continue.

[0220] For example, as shown in Figure 11, the electronic device stops the eye-controlled image zooming process and simultaneously pops up a prompt window on top of the target image, displaying the text message "Cannot continue zooming". The prompt window is displayed on the screen of the electronic device for 1.5 seconds and then cancels the display.

[0221] In one possible implementation, the electronic device stops the currently executing eye-controlled image magnification process, but it can still continuously monitor the user's eyes using AO (Aggregate Eye Movement) technology via the front-facing camera. If the user consciously blinks, the electronic device detects this blink, exits the image magnification process, and restores the display interface of the target image before magnification. In some possible implementations, the user can also manually click on any location on the magnified target image on the display screen. The electronic device receives the user's click input, exits the image magnification process, and restores the display interface of the target image before magnification.

[0222] S208: Obtain the second gaze time and the second gaze region of the target image after the user gazes at it in the first magnified processing.

[0223] Specifically, the second fixation time can be the time during which a user continuously fixates on a specific area (by analyzing multiple consecutive frames of user eye images acquired by a camera to determine that the user's eye fixation position has not changed). Furthermore, the second fixation time can also include the time during which the user primarily fixates on a fixed area over a longer period, even if the area of ​​fixation changes briefly during this time, the user's primary area of ​​fixation remains relatively fixed within this period.

[0224] Electronic devices can first determine the second fixation area by capturing the first few frames of user eye images, and then determine the second fixation time, or they can first determine the second fixation time, and then determine the second fixation area based on the acquired user eye images.

[0225] In one possible implementation, the electronic device displays a target image that has undergone a first magnification process on the screen. The user experiences a need for magnification and consciously focuses on a specific area of ​​the magnified target image. The electronic device first determines a second fixation time and then determines a second fixation area. The electronic device captures multiple high-resolution images of the user's eyes at short time intervals using a front-facing camera. The electronic device analyzes the acquired, short-interval-range eye images to determine second fixation information, which includes a second fixation time and a second fixation area.

[0226] For example, Figure 12A is a schematic diagram of the second magnification process 1. As shown in Figure 12A, the electronic device displays the target image 1201 without the first magnification process on the display screen. The electronic device has determined the first target region of interest 1202 and performed the first magnification process on the first target region of interest 1202 as the reference region, and displays it on the display screen of the electronic device shown in Figure 12B. Figure 12B is a schematic diagram of the second magnification process 2. As shown in Figure 12B, the electronic device displays the target image 1211 after the first magnification process on the display screen. The electronic device continuously captures multiple frames of user eye images for 0.8 seconds at a high sampling rate of 100 frames per second through the front-facing camera 1212, acquiring 80 frames of user eye images. The 80 frames of eye images are analyzed to obtain the area of ​​user gaze in each frame. Images with more than or equal to 50 consecutive frames corresponding to user gaze areas without significant displacement are detected, and the number of consecutive image frames is counted. The second gaze time is calculated by the largest number of consecutive image frames. Specifically, the calculation method is: T2=(N2-1)·0.01

[0227] Where T2 is the second gaze time, N2 is the second gaze time of obtaining the target image after the user gazes at the first magnified image in step S208, and the maximum number of consecutive image frames detected in the second gaze region.

[0228] In one possible implementation, the electronic device determines the second gaze region of the target image on the display screen based on the second gaze time and the consecutive images corresponding to the largest number of consecutive image frames detected in the second gaze region obtained in step S208 after the user gazes at the first magnified target image. The second gaze region can be coarse-precision coordinates. Before determining the second gaze region, the electronic device divides the possible second gaze regions according to the aspect ratio of the image, and then determines the second gaze region based on the captured image of the user's eyes. The electronic device's division and determination of the possible second gaze regions of the target image after the first magnification process can be found in the description of the electronic device's division of the possible first gaze regions in step S202, which describes the acquisition of the first gaze region and first gaze time of the user gazing at the target image.

[0229] S209: Determine whether the second fixation time is greater than or equal to the second threshold.

[0230] Specifically, after acquiring the second gaze time, the electronic device compares the second gaze time with a preset second threshold. If the second gaze time is less than the second threshold, the electronic device continues to execute step S208 to acquire the second gaze time and the second gaze region of the user's gaze at the first magnified target image. If the second gaze time is greater than or equal to the second threshold, step S210 is triggered to determine the second target region of interest within the second gaze region. The second threshold is used to determine whether the user is consciously gazing at a certain area of ​​the image displayed on the screen in order to achieve image magnification, and the second threshold can be less than the first threshold.

[0231] In one possible implementation, the electronic device can use an AI model to determine whether the user's second gaze time is greater than or equal to a second threshold. The electronic device uses the second gaze time as input to the AI ​​model, analyzes it, and outputs the relationship between the second gaze time and the second threshold. The relationship includes two possibilities: greater than or equal to, and less than.

[0232] In one possible implementation, the second threshold can be set to 0.5 seconds. When the second gaze time is less than 0.5 seconds, the electronic device repeatedly executes step S208 to acquire the second gaze time and the second gaze region of the target image after the user gazes at the first magnified image; when the first gaze time is greater than or equal to 0.5 seconds, step S210 is triggered to determine the second target region of interest within the second gaze region.

[0233] S210: Determine the region of interest of the second target within the second gaze area.

[0234] First, the electronic device performs semantic segmentation on the target image displayed on the screen after the first magnification process. Specifically, the electronic device performs semantic segmentation on the target image displayed on the screen after the first magnification process according to a third semantic category, identifying P third-level semantic regions in the target image, where P is a positive integer greater than or equal to 1.

[0235] In one possible implementation, the third semantic category includes one or more of the following: eye, eyelash, eyelid, pupil, nose, nostril, lips, upper lip, lower lip, teeth, palm, thumb, index finger, middle finger, ring finger, little finger, arm, elbow, forearm, upper arm, blood vessel, leg, toe, clothing, clothing folds, button, plant, petal, stamen, calyx, flower stem, bark texture, and stone.

[0236] In one possible implementation, when determining P third-level semantic regions based on semantic segmentation according to a third semantic category, a single target image may include multiple objects corresponding to the third semantic category. When objects corresponding to the third semantic category overlap, among every two overlapping objects, the object with relatively complete semantics and located at a higher level in the image corresponding to the third semantic category is called the occluder, and the object whose portion is obscured by the occluder and located at a lower level in the image is called the occluded object. Specifically, when the occluder only partially covers the occluded object, causing part of the occluded object to be invisible, but not segmented into two or more independent regions that are not directly connected in space, simple image analysis techniques, such as color thresholding and edge detection, are first used to preliminarily identify the occluder in the image. Based on the detected occluder, an edge detection algorithm is further used to extract the precise edge of the occluder. The edge of the occluded object is determined based on the precise edge of the occluder and the unoccluded portion of the occluded object. The closed regions enclosed by the edges of the occluder and the occluded object are respectively determined as P third-level semantic regions.

[0237] For example, as shown in FIG12B, the electronic device displays the target image 1211 after the first magnification process on the display screen. The electronic device performs semantic segmentation on the target image 1211 after the first magnification process according to the third semantic category, and obtains 10 third-level semantic regions: 1213, 1214, 1215, 1216, 1217, 1218, 1219, 12110, 12111, and 12112. Among them, the third-level semantic region 1 213 is the left eyebrow, 1214 is the left eye, 1215 is the right eyebrow, 1216 is the right eye, 1217 is the right ear, 1218 is the upper lip, 1219 is part 1 of the bow tie, 12110 is part 2 of the bow tie, 12111 is the nose, and 12112 is the lower lip. At this point, P is 10.

[0238] In one possible implementation, when the target image after the first magnification process has an overlap between the object corresponding to the third semantic category and other objects, the segmentation processing method of the third-level semantic region refers to step S204, which determines the method of semantic segmentation of the target image without the first magnification process by electronic device in the first target interest region within the first gaze region according to the first category and the second category.

[0239] It should be noted that the semantic segmentation diagrams shown in Figures 12A and 12B are exemplary illustrations of embodiments of this application. Other semantic segmentation methods may be used when performing semantic segmentation on the target image after the first magnification process, and this application embodiment does not limit this.

[0240] Next, from the P third-level semantic regions, determine the second gaze time of the target image after the user gazes at it in step S208, and the T third-level semantic regions that intersect with the second gaze region determined in the second gaze region. The T third-level semantic regions that intersect with the second gaze region may include third-level semantic regions segmented according to a third semantic category that are completely covered by the second gaze region, and third-level semantic regions segmented according to a third semantic category that are partially covered by the gaze region.

[0241] Based on the preset magnification priority of the T third-level semantic regions and the area occupied by the T third-level semantic regions in the second gaze region, the second target region of interest is determined.

[0242] The preset magnification priority is a set of priority rules for object magnification display that are pre-set in electronic devices based on factors such as product design concept and user experience. Specifically, it is involved in step S204, which describes the preset magnification priority in the first target region of interest within the first gaze area. This will not be elaborated here.

[0243] In one possible implementation, after determining T third-level semantic regions that intersect with the second gaze region, the electronic device traverses the T third-level semantic regions covered by the second gaze region, which are semantically segmented according to a third semantic category. Based on a preset magnification priority rule, the magnification priority of each third-level semantic region intersecting with the second gaze region is determined, and the pixels of each third-level semantic region are counted to obtain the pixel count of each third-level semantic region as the area corresponding to each third-level semantic region. When a third-level semantic region intersecting with the second gaze region is a third-level semantic region segmented according to a third semantic category and completely covered by the second gaze region, the total number of pixels in this third-level semantic region is counted as the area corresponding to this third-level semantic region. When a third-level semantic region intersecting with the second gaze region is a third-level semantic region segmented according to a third semantic category and partially covered by the gaze region, only the pixel count of the portion of the third-level semantic region covered by the second gaze region is counted as the area of ​​this third-level semantic region; the pixels of the second gaze region are counted as the area of ​​the third gaze region.

[0244] In one possible implementation, the electronic device calculates the magnification weight of each of the T third-level semantic regions based on the magnification priority and area of ​​the T third-level semantic regions covered by the second gaze region, and selects the third-level semantic region with the largest magnification weight as the second target region of interest. If the third-level semantic region with the largest magnification weight is a third-level semantic region segmented according to the third semantic category and completely covered by the second gaze region, then this third-level semantic region is selected as the second target region of interest. If the third-level semantic region with the largest magnification weight is a third-level semantic region segmented according to the third semantic category and partially covered by the second gaze region, then the portion of this third-level semantic region covered by the second gaze region is selected as the second target region of interest.

[0245] For example, as shown in Figure 12B, the target image after the first magnification process has a resolution of 1560x720 pixels and can be divided into four possible second gaze regions. The second gaze region 12113 gazed upon by the user is the upper right rectangular area of ​​the target image 1211 after the first magnification process. The third-level semantic regions covered by the second gaze region include third-level semantic regions 12114, 12115, 12116, 12117, and 12118. Among them, third-level semantic regions 12117 (part of the left eyebrow) and 12118 (part of the left eye) are partially covered by the second gaze region. When counting pixels, only the part covered by the second gaze region 12113 is counted. The electronic device presets the magnification priority of the third-level semantic regions to an integer from 0 to 10, and each third-level semantic region corresponds to an integer representing the magnification priority. The magnification priority of the third-level semantic region (right eyebrow) 12114 is 6, with an area of ​​8000 pixels; the magnification priority of the third-level semantic region (right eye) 12115 is 7, with an area of ​​20000 pixels; the magnification priority of the third-level semantic region (right ear) 12116 is 6, with an area of ​​28000 pixels; the magnification priority of the third-level semantic region (part of the left eyebrow) 12117 is 6, with an area of ​​3000 pixels; the magnification priority of the third-level semantic region (part of the left eye) 12118 is 7, with an area of ​​9000 pixels; the area of ​​the second gaze region is 280800 pixels.

[0246] Optionally, the amplification weight W2 of the third-level semantic region is calculated as follows: W2 = L n2 ·V l2 +A n2 ·V a2

[0247] Where L n2 V represents the amplification priority of the third-level semantic region after normalization. l2 The weight of the priority amplification for the third-level semantic region is a pre-set value for the electronic device, A.n2 V represents the area of ​​the third-level semantic region after normalization. a2 The area weight is a preset value for electronic devices.

[0248] Optionally, L n2 The calculation formula is:

[0249] Where L r2 L is the magnification priority determined after the electronic device traverses the third-level semantic region covered by the second gaze region. min2 The maximum value of the amplification priority of the third-level semantic region preset for electronic devices, L max2 The maximum value of the amplification priority of the third-level semantic region preset for electronic devices.

[0250] Optionally, A n2 The calculation formula is:

[0251] Where A o2 A is the area of ​​each third-level semantic region covered by the second gaze region traversed by the electronic device. r2 This represents the area of ​​the second gaze region.

[0252] For example, the weight V of the amplification priority l2 =0.6, area weight V a2 =0.4, A r2 The L value is 280,800 pixels, corresponding to the third-level semantic region (right eyebrow) 12114 shown in Figure 12B. r2 It is 6, L max2 For 10, L min2 If it is 1, then Corresponding A o2 If it is 8000 pixels, then The corresponding amplification weights W2 = W2 = L n2 ·V l2 +A n2 ·V a2=0.555555·0.6+0.009971·0.4=0.373219. Similarly, the magnification weights of other third-level semantic regions covered within the second fixation region 12113 shown in Figure 12B are calculated using the same method. The magnification weights are: W2 = 0.428490 for the third-level semantic region (right eye) 12115, W2 = 0.373219 for the third-level semantic region (right ear) 12116, W2 = 0.337607 for the third-level semantic region (partial left eyebrow) 12117, and W2 = 0.412821 for the third-level semantic region (partial left eye) 12118. The third-level semantic region (right eye) 12115 with the largest magnification weight is selected as the second target region of interest.

[0253] S211: Perform a second magnification process on the target image based on the second target region of interest, and display it on the display screen.

[0254] In one possible implementation, if the semantics of the second target region of interest are complete, then the second target region of interest is used as the second reference region for the second magnification process, and the target image after the first magnification process is magnified at a fixed magnification based on the second reference region.

[0255] For example, as shown in FIG12B, the electronic device determines the second gaze region 12113 of the user's gaze. The electronic device has determined the third-level semantic region (right eye) 12115 as the second target region of interest, and the third-level semantic region (right eye) 12115 is semantically complete within the second gaze region 12113. The electronic device uses the third-level semantic region (right eye) 12115 as the reference region for the second magnification process, and performs a second magnification process on the target image displayed on the display screen after the first magnification process at a fixed magnification of 1.5 times, and displays it on the display screen, as shown in FIG12C. FIG12C is a schematic diagram of the second magnification process. At this time, the target image 1221 after the second magnification process is displayed on the display screen of the electronic device. The target image 1221 takes the second target region of interest, i.e., the third-level semantic region (right eye) 12115 in FIG12B, as the center of the image.

[0256] The definition of the reference area can be found in the description of the reference area in the application scenarios involved in the embodiments of this application in the specification, and will not be repeated here.

[0257] In one possible implementation, if the semantics of the second target region of interest are incomplete and the second target region of interest belongs to a part of a third-level semantic region, then the third-level semantic region to which the second target region of interest belongs is used as the second reference region for the second amplification process.

[0258] For example, as shown in Figure 12B, if the electronic device calculates and determines that the third-level semantic region (part of the left eyebrow) 12117 is the second target region of interest, where the second target region of interest is an irregular region that is cut out by the edge of the third-level semantic region (left eyebrow) 1213 and covered by the second gaze region 12113, and the third-level semantic region (part of the left eyebrow) 12117 is semantically incomplete, and the second target region of interest belongs to a part of the third-level semantic region (left eyebrow) 1213, then the electronic device determines that the third-level semantic region (left eyebrow) 1213 is the reference region for the second magnification processing, and performs a second magnification processing on the target image 1211 after the first magnification processing at a fixed magnification of 1.5 times, and displays it on the display screen, as shown in Figure 12D, which is a schematic diagram of the second magnification processing process. At this time, the electronic device displays the target image 1231 after the second magnification processing, with the third-level semantic region (left eye) 1214 shown in Figure 12B as the center of the image.

[0259] In one possible implementation, if the second target region of interest is semantically incomplete and the second target region of interest belongs to the entire region of a certain third-level semantic region, the electronic device can detect whether the third-level semantic region with the largest magnification weight calculated in the second target region of interest within the second gaze region, excluding the second target region of interest, is semantically complete, as determined in step S210. If this third-level semantic region is semantically complete, it is used as the second reference region for the second magnification process. If this third-level semantic region is semantically incomplete, the electronic device can detect whether the third-level semantic region with the second largest magnification weight calculated in the second target region of interest within the second gaze region, excluding the second target region of interest, is semantically complete, as determined in step S210. If this third-level semantic region is semantically complete, it is used as the second reference region for the second magnification process. If it is incomplete, the same method is used to traverse the T third-level semantic regions covered within the second gaze region according to the magnification weight from high to low. If all T third-level semantic regions covered within the second gaze region are semantically incomplete, the second magnification process is performed with the center of the target image after the first magnification process as the reference region.

[0260] S212: Determine whether the resolution of the target image after the second magnification process is lower than the resolution of the display screen.

[0261] In step S206, the electronic device determines whether the resolution of the first magnified target image is greater than or equal to the resolution of the display screen. The electronic device has already obtained the current resolution of the electronic device's display screen. The electronic device then obtains the resolution of the second magnified target image. The resolution of the second magnified target image is determined by multiplying the width pixel count and the height pixel count, just like the current resolution of the electronic device's display screen. That is, it is represented by multiplying the number of pixels in the horizontal direction of this target image by the number of pixels in the vertical direction. The electronic device then compares whether the resolution of the second magnified target image is lower than the current resolution of the display screen. If the resolution of the second magnified target image is lower than the current resolution of the display screen, then step S207 is executed to stop the eye-controlled image magnification process.

[0262] In one possible implementation, when determining whether the resolution of the target image after the second magnification process is lower than the current resolution of the electronic device display screen, the number of pixels in the width of the target image after the second magnification process is compared with the number of pixels in the width of the current resolution of the electronic device display screen, and the number of pixels in the height of the target image after the second magnification process is compared with the number of pixels in the height of the current resolution of the electronic device display screen. When the number of pixels in the width of the target image after the second magnification process is lower than the number of pixels in the width of the current resolution of the electronic device display screen, or when the number of pixels in the height of the target image after the second magnification process is lower than the number of pixels in the height of the current resolution of the electronic device display screen, it is considered that the resolution of the target image after the second magnification process is lower than the current resolution of the electronic device display screen.

[0263] For example, if the resolution of the electronic device display screen is 2400×1600 pixels, and the resolution of the target image after the second magnification process is 1600×1200 pixels, the width and height pixel counts of the two are compared respectively. The width pixel count of the target image after the second magnification process is 1600, which is less than the width pixel count of the electronic device (2400). At the same time, the height pixel count of the target image after the second magnification process is 1200, which is less than the width pixel count of the electronic device (1200). This satisfies the condition that the width pixel count of the target image resolution after the second magnification process is less than the width pixel count of the current resolution of the electronic device display screen, or the height pixel count of the target image resolution after the second magnification process is less than the height pixel count of the current resolution of the electronic device display screen. Then, step S207 is executed to stop the eye-controlled image magnification process.

[0264] If the resolution of the target image after the second magnification process is 3200×1200 pixels, the width of the target image after the second magnification process is 3200 pixels, which is greater than the width of the electronic device (2400 pixels), and the height of the target image after the second magnification process is 1200 pixels, which is less than the width of the electronic device (1600 pixels), then the conditions are met: the width of the target image after the second magnification process is less than the width of the electronic device's current display screen resolution, or the height of the target image after the second magnification process is less than the height of the electronic device's current display screen resolution. Then, step S207 is executed to stop the eye-controlled image magnification process.

[0265] If the resolution of the target image after the second magnification process is 1200×1600 pixels, the width of the target image after the second magnification process is 1200 pixels, which is less than the width of the electronic device (2400 pixels), and the height of the target image after the second magnification process is 1600 pixels, which is equal to the width of the electronic device (1600 pixels), then the condition that the width of the target image after the second magnification process is less than the width of the electronic device's display screen or the height of the target image after the second magnification process is less than the height of the electronic device's display screen is met, then step S207 is executed to stop the eye-controlled image magnification process.

[0266] It should be understood that each step in the above method embodiments can be completed by integrated logic circuits in the processor hardware or by instructions in software form. The method steps disclosed in the embodiments of this application can be directly manifested as being executed by a hardware processor, or being executed by a combination of hardware and software modules in the processor.

[0267] This application also provides an electronic device, which may include a memory and a processor. The memory may be used to store a computer program; the processor may be used to invoke the computer program in the memory, causing the electronic device to execute the method executed by the electronic device in any of the above embodiments.

[0268] This application also provides a chip system including at least one processor for implementing the functions involved in the electronic device in any of the above embodiments.

[0269] In one possible design, the chip system also includes a memory for storing program instructions and data, which may be located within or outside the processor.

[0270] The chip system can consist of chips or include chips and other discrete components.

[0271] Optionally, the chip system may contain one or more processors. These processors can be implemented in hardware or software. When implemented in hardware, the processor can be a logic circuit, an integrated circuit, etc. When implemented in software, the processor can be a general-purpose processor, implemented by reading software code stored in memory.

[0272] Optionally, the chip system may contain one or more memories. The memory may be integrated with the processor or disposed separately from it; this application embodiment does not limit this. For example, the memory may be a non-transient processor, such as a read-only memory (ROM), which may be integrated with the processor on the same chip or disposed separately on different chips. This application embodiment does not specifically limit the type of memory or the arrangement of the memory and processor.

[0273] For example, the chip system may be a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on chip (SoC), a central processor unit (CPU), a network processor (NP), a digital signal processor (DSP), a micro controller unit (MCU), a programmable logic device (PLD), or other integrated chips.

[0274] This application also provides a computer program product comprising: a computer program (also referred to as code or instructions) that, when run, causes a computer to perform the method executed by the electronic device in any of the above embodiments.

[0275] This application also provides a computer-readable storage medium storing a computer program (also referred to as code or instructions). When the computer program is run, it causes the computer to perform the method executed by the electronic device in any of the above embodiments.

[0276] The various embodiments of this application can be combined arbitrarily to achieve different technical effects.

[0277] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state disk (SSD)).

[0278] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This program can be stored in a computer-readable storage medium, and when executed, it can include the processes described in the above method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM or random access memory (RAM), magnetic disks, or optical disks.

[0279] In summary, the above description is merely an embodiment of the technical solution of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, improvements, etc., made based on the disclosure of this application should be included within the scope of protection of this application.

Claims

1. A method for image display processing, characterized in that, Applied to a terminal device, the terminal device including a display screen, the method includes: Display a target image on the display screen and obtain first gaze information of the user looking at the target image. The first gaze information includes a first gaze area of ​​the user looking at the target image and a first gaze time of the user looking at the first gaze area. Determine whether the first gaze duration exceeds the first threshold; If the first threshold is exceeded, then based on the first gaze region and the content of the first gaze region, a first target region of interest within the first gaze region is determined; The target image is magnified based on the first target region of interest, and the magnified target image is displayed on the display screen.

2. The method as described in claim 1, characterized in that, Determining the first target region of interest within the first gaze region includes: The target image is semantically segmented to identify M first-level semantic regions in the target image, where M is an integer greater than 1; From the M first-level semantic regions, determine N first-level semantic regions that intersect with the first gaze region; Based on the preset magnification priority of the N first-level semantic regions and the area occupied by the N first-level semantic regions in the first gaze region, the first target region of interest is determined.

3. The method as described in claim 2, characterized in that, The step of semantic segmentation of the target image, identifying M first-level semantic regions in the target image, includes: The target image is semantically segmented according to the first semantic category to identify Q second-level semantic regions in the target image, where Q is a positive integer less than M; Semantically segment each of the Q second-level semantic regions according to the second semantic category, and identify one or more first-level semantic regions in each second-level semantic region; The first-level semantic regions corresponding to the Q second-level semantic regions are determined as the M first-level semantic regions.

4. The method as described in claim 3, characterized in that, The first semantic category includes one or more of people, animals, plants, and scenery; the second semantic category includes one or more of human facial features, hair, arms, legs, hands, feet, clothing, tree crowns, tree trunks, flower crowns, plant stems, animal facial features, and limbs.

5. The method according to any one of claims 1-4, characterized in that, The first target region of interest is part or all of a first-level semantic region that intersects with the first gaze region.

6. The method according to any one of claims 1-5, characterized in that, The first magnification process of the target image based on the first target region of interest includes: If the semantics of the first target region of interest are complete, then the first target region of interest is used as the first reference region for the first magnification operation, and the target image is magnified based on the first reference region. If the semantics of the first target region of interest are incomplete, a second reference region is determined based on the first target region of interest, the first-level semantic region, and the second-level semantic region, and the target image is magnified based on the second reference region.

7. The method according to any one of claims 1-6, characterized in that, After displaying the target image after the first magnification process on the display screen, the method further includes: When the target image after the first magnification process meets the first preset condition, the second gaze information of the user looking at the target image after the first magnification process is obtained. The second gaze information includes the second gaze area of ​​the user looking at the target image after the first magnification process and the second gaze time of the user looking at the second gaze area. Determine whether the second fixation time exceeds the second threshold; If the second threshold is exceeded, then based on the second gaze region and the content of the second gaze region, a second target region of interest within the second gaze region is determined; The target image is magnified a second time based on the second target region of interest, and the magnified target image is displayed on the display screen.

8. The method as described in claim 7, characterized in that, Also includes: Obtain the current resolution of the display screen and the resolution of the target image after the first magnification process; If the resolution of the target image after the first magnification process is greater than or equal to the current resolution of the display screen, then the target image after the first magnification process satisfies the first preset condition. If the resolution of the target image after the first magnification process is lower than the current resolution of the display screen, then the target image after the first magnification process does not meet the first preset condition.

9. The method as described in claim 7, characterized in that, After displaying the target image after the second magnification process on the display screen, the method further includes: When the target image after the second magnification process meets the second preset condition, the image magnification stops. The second preset condition includes: Obtain the current resolution of the display screen and the resolution of the target image after the second magnification process; If the resolution of the target image after the second magnification process is less than the current resolution of the display screen, then the target image after the second magnification process satisfies the second preset condition. If the resolution of the target image after the second magnification process is greater than or equal to the current resolution of the display screen, then the target image after the second magnification process does not meet the second preset condition.

10. The method according to any one of claims 1-9, characterized in that, The terminal device further includes a camera; the first gaze information obtained by acquiring the user's gaze at the target image includes: The camera captures one or more frames of eye images of the user gazing at the target image, and the first gaze information is determined based on the one or more frames of user eye images.

11. The method as described in claim 10, characterized in that, Also includes: The one or more frames of eye images are used as input to an artificial intelligence (AI) model. The AI ​​model analyzes the one or more frames of eye images and outputs the first gaze information.

12. The method according to any one of claims 1-11, characterized in that, The method further includes: detecting whether the user is looking at the display screen using always-on AO technology, and using the camera to acquire the user's eye image.

13. A smart terminal device, characterized in that, include: A memory, and one or more processors; the memory is coupled to the one or more processors, the memory being used to store computer program code, the computer program code including computer instructions, the one or more processors invoking the computer instructions to cause the electronic device to perform the method as described in any one of claims 1-12.

14. A computer storage medium, characterized in that, The computer storage medium stores a computer program that, when executed by a processor, implements the method described in any one of claims 1-12.

15. A computer program product, characterized in that, The computer program product includes instructions that, when executed by a computer, cause the computer to perform the method as described in any one of claims 1-12.

Citation Information

Patent Citations

  • Image processing method and device

    CN108255299A

  • Eye movement tracking method and device, electronic equipment and computer readable storage medium

    CN116339510A

  • A composition for emitting glucose comprising hydrogel and epidermal growth factor receptor ligand as an active ingredient

    KR1020210003052A