Image display processing method and terminal equipment
By acquiring user gaze information and using semantic segmentation technology, terminal devices can accurately identify and zoom in on areas of interest for users, solving the problem of poor user experience in existing technologies and achieving more efficient image zooming and reduced power consumption.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-30
- Publication Date
- 2026-03-17
AI Technical Summary
When existing terminal devices control image zooming via eye movements, they typically zoom mechanically from the image center, failing to meet users' specific zooming needs and resulting in a poor user experience.
By acquiring user gaze information, determining whether the gaze time exceeds a threshold, identifying the user's region of interest, and performing image magnification based on this region, combined with semantic segmentation technology to identify semantic regions in the image, the system can accurately determine the region that the user wants to magnify.
It achieves more precise image magnification, meets users' needs for magnifying specific areas, improves user experience, and reduces computational load and power consumption.
Smart Images

Figure CN121680697A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of image processing, and in particular to an image display processing method and a terminal device. BACKGROUND
[0002] With the development of terminal technology, terminal devices can implement many editing operations on pictures. Enlarging a picture is a relatively important and basic operation implemented by a terminal device.
[0003] Existing terminal devices can support a user to control picture enlargement through eye contact without contacting the terminal device, which can free the user's hands and meet the needs in specific scenarios. However, the picture enlargement through eye contact usually mechanically enlarges the center of an image, which cannot meet the specific enlargement needs of the user and the user experience is poor. SUMMARY
[0004] Embodiments of the present application provide an image display processing method and a terminal device, which implement a user interested area as a reference area and control enlarged pictures through eye tracking.
[0005] In a first aspect, embodiments of the present application provide an image display processing method, which is applied to a terminal device including a display screen, and includes: displaying a target image on the display screen, obtaining first gaze information of a user gazing at the target image, the first gaze information including a first gaze area gazed at by the user in the target image and a first gaze time of the user gazing at the first gaze area; determining whether the first gaze time exceeds a first threshold; if the first gaze time exceeds the first threshold, determining a first target interested area in the first gaze area based on the first gaze area and content of the first gaze area; and performing first enlargement processing on the target image based on the first target interested area, and displaying the target image after the first enlargement processing on the display screen.
[0006] In the embodiments of the present application, the terminal device can acquire the gaze time of the user and determine whether the user consciously gazes at the picture through a threshold, so as to more accurately determine whether the user has the intention and demand to zoom in the picture. Further, in the process of zooming in the picture, the embodiments of the present application first determine a larger user gaze area, and then determine a smaller user interested area close to the real intention of the user based on the gaze area and the content (for example, whether it is a person, a scene, or a local feature in a person) in the gaze area, and then zoom in the interested area, so as to solve the problems of large calculation amount and inaccurate zoomed-in area caused by zooming in the whole picture or the whole area. The embodiments of the present application can more accurately identify the target zoom-in area of the user, reduce the calculation amount, and improve the zoom-in efficiency. For example, when the user uses the terminal device to view the picture, the user can accurately zoom in the picture by eye gaze, the demand of the user for zooming in a specific area can be met, the operation is simple and convenient, and the user experience can be improved.
[0007] In a possible implementation form of the first aspect, the first target interested area is determined based on the first gaze area, including: performing semantic segmentation on the target picture to identify M first-level semantic areas in the target picture, M being an integer greater than 1; determining N first-level semantic areas having an intersection with the first gaze area from the M first-level semantic areas; and determining the first target interested area based on a preset zoom-in priority of the N first-level semantic areas and an area occupied by the N first-level semantic areas in the first gaze area. By performing semantic segmentation, the embodiments of the present application can identify the user's possible interested area having an intersection with the user's gaze area, improve the accuracy of gaze detection, and determine the target area that the user hopes to zoom in most likely by comprehensively considering the preset zoom-in priority and the area, so as to avoid zooming in the interference area having high zoom-in priority but small area in the picture.
[0008] In a possible implementation manner of the first aspect, the semantic segmentation on the target image identifies M first-level semantic regions in the target image, comprising: performing semantic segmentation on the target image according to a first semantic category, to identify Q second-level semantic regions in the target image, Q being a positive integer smaller than M; performing semantic segmentation on each of the Q second-level semantic regions according to a second semantic category, to identify one or more first-level semantic regions in each of the second-level semantic regions; and determining the first-level semantic regions corresponding to the Q second-level semantic regions as the M first-level semantic regions. In the embodiment of the present application, when the picture viewed by the user shows multiple objects or human bodies and details of the multiple objects or human bodies, different objects or human bodies can be identified and segmented out through multiple semantic segmentations, and more subdivided regions can be identified and segmented out from the different objects or human bodies, so as to improve the accuracy of semantic region identification, mine more information in the picture, and more accurately determine the possible zoom-in region of interest of the user.
[0009] In a possible implementation manner of the first aspect, the first semantic category comprises one or more of a person, an animal, a plant, and a scene, and the second semantic category comprises one or more of human facial features, hair, arms, legs, hands, feet, clothes, tree crowns, tree trunks, flower crowns, stems of plants, animal facial features, and limbs. In the embodiment of the present application, when the picture viewed by the user includes objects of different categories and human bodies, performing semantic segmentation on the target image according to the first semantic category can first identify different objects or human bodies, and more accurately distinguish the subject from the background, and when the objects or human bodies of different categories contain more details, performing semantic segmentation on the target image segmented according to the first semantic category according to the second semantic category can identify semantic regions such as eyes, noses, arms, and palms from the identified objects or human bodies in the target image, to improve the accuracy of semantic segmentation.
[0010] In a possible implementation manner of the first aspect, the first target region of interest is part or the whole of the first-level semantic region intersecting with the first gaze region. In the embodiment of the present application, when the first target region of interest is part of the first-level semantic region intersecting with the first gaze region, the first-level semantic region is partially covered by the first gaze region, and the covered part is the first target region of interest. The electronic device takes the first target region of interest as the region possibly zoomed in by the user, avoids taking the whole of the first-level semantic region with high priority or a large area but only a small area intersecting with the first gaze region as the target region possibly zoomed in by the user to calculate the area, and can realize focusing on only the content contained in the gaze region of the user, to reduce the amount of calculation.
[0011] In one possible implementation of the first aspect, performing a first magnification process on the target image based on the first target region of interest includes: if the first target region of interest is semantically complete, then using the first target region of interest as a first reference region for the first magnification operation, and performing the first magnification process on the target image based on the first reference region; if the first target region of interest is semantically incomplete, then determining a second reference region based on the first target region of interest, the first-level semantic region, and the second-level semantic region, and performing the first magnification process on the target image based on the second reference region. In this embodiment, when the semantics of the target region that the user may wish to magnify is complete, the target region can be directly used as the reference region for magnification, and the target region will become the center region of the magnified image. When the semantics of the target region that the user may wish to magnify is incomplete, the reference region for magnification can be adjusted, and a larger region containing the target region that the user may wish to magnify can be used as the reference region for magnification, and the larger region containing the target region that the user may wish to magnify will become the center region of the magnified image. This ensures that each magnification process is based on the user's region of interest, and the image displayed on the screen will present the complete semantics of the user's region of interest.
[0012] In one possible implementation of the first aspect, after displaying the target image after the first magnification process on the display screen, the method further includes: when the target image after the first magnification process meets a first preset condition, acquiring second gaze information of the user gazing at the target image after the first magnification process, the second gaze information including a second gaze region gazed at by the user in the target image after the first magnification process and a second gaze time of the user gazing at the second gaze region; determining whether the second gaze time exceeds a second threshold; if it exceeds the second threshold, determining a second target interest region within the second gaze region based on the second gaze region and the content of the second gaze region; performing a second magnification process on the target image based on the second target interest region, and displaying the target image after the second magnification process on the display screen. Therefore, in this embodiment of the application, when the image clarity after magnification is still high and contains a lot of information, the image can be magnified again through gazing, helping the user to further explore image information or process the image.
[0013] In one possible implementation of the first aspect, the method further includes: obtaining the current resolution of the display screen and the resolution of the target image after the first magnification process; if the resolution of the target image after the first magnification process is greater than or equal to the current resolution of the display screen, then the target image after the first magnification process satisfies the first preset condition; if the resolution of the target image after the first magnification process is lower than the current resolution of the display screen, then the target image after the first magnification process does not satisfy the first preset condition. Implementing this embodiment, after the target image undergoes the first magnification process and is displayed on the display screen, the resolution of the target image after the first magnification process is compared with the current resolution of the terminal device's display screen. When the resolution of the target image after the first magnification process is greater than or equal to the current resolution of the terminal device's display screen, further magnification of the target image can still obtain a relatively clear image and more detailed information, therefore it is determined that the image can be magnified again, i.e., it meets the first preset condition; when the resolution of the target image after the first magnification process is lower than the current resolution of the terminal device's display screen, further magnification of the target image may result in blurriness, making it impossible to accurately obtain the information needed by the user, thereby weakening the actual utility and value of the image magnification function and reducing the overall user experience, therefore it is determined that the image is not suitable for further magnification, i.e., it does not meet the first preset condition.
[0014] In one possible implementation of the first aspect, after displaying the target image after the second magnification process on the display screen, the method further includes: stopping image magnification when the target image after the second magnification process meets a second preset condition. The second preset condition includes: obtaining the current resolution of the display screen and the resolution of the target image after the second magnification process; if the resolution of the target image after the second magnification process is less than the current resolution of the display screen, then the target image after the second magnification process meets the second preset condition; if the resolution of the target image after the second magnification process is greater than or equal to the current resolution of the display screen, then the target image after the second magnification process does not meet the second preset condition. Implementing the embodiments of this application, after the target image has undergone multiple magnification processes and is displayed on the display screen, the resolution of the target image after multiple magnification processes is compared with the current resolution of the terminal device's display screen. When the resolution of the target image after multiple magnification processes is already lower than the current resolution of the terminal device's display screen, further magnification may significantly reduce the image's clarity, failing to provide more useful information and practical value. Therefore, it is determined that the image is not suitable for further magnification, and based on this determination, the eye-controlled image magnification is stopped. This avoids blurring of the target image displayed on the display screen after further magnification, reducing user experience, and also avoids unlimited magnification, reducing computational load.
[0015] In one possible implementation of the first aspect, the terminal device further includes a camera; acquiring first gaze information of a user gazing at the target image includes: capturing one or more frames of eye images of the user gazing at the target image using the camera, and determining the first gaze information based on the one or more frames of user eye images. Implementing the embodiments of this application can achieve continuous monitoring of user eye images and accurately acquire the user's gaze area and gaze time.
[0016] In one possible implementation of the first aspect, the method further includes: using the one or more frames of eye images as input to an artificial intelligence (AI) model, analyzing the one or more frames of eye images through the AI model, and outputting the first gaze information. In this embodiment, the AI model can flexibly and accurately obtain the gaze region and gaze duration from user eye images.
[0017] In one possible implementation of the first aspect, the method further includes: detecting whether a user is gazing at the display screen using always-on AO technology, and invoking the camera to acquire an image of the user's eyes. By implementing embodiments of this application, the electronic device can automatically and continuously monitor the user's eye activity using a low-power camera. Upon detecting a user's gaze, the system can respond quickly, initiating eye tracking to accurately acquire the user's gaze information, thus eliminating the need for manual operation by the user. Simultaneously, it effectively reduces energy consumption during the monitoring process, ensuring a balance between device battery life and user experience.
[0018] Secondly, embodiments of the present invention provide a smart terminal device, including: a touch screen, a camera, one or more processors, and one or more memories. The one or more processors are coupled to the touch screen, the camera, and the one or more memories. The one or more memories are used to store computer program code, which includes computer instructions. When the one or more processors execute the computer instructions, the electronic device performs the method in any possible implementation of any of the above aspects.
[0019] Thirdly, this application provides a computer storage medium storing a computer program that, when executed by a processor, implements the image display processing method flow described in any one of the first aspects above.
[0020] Fourthly, embodiments of the present invention provide a computer program product including instructions that, when executed by a computer, enable the computer to perform the image display processing method flow described in any of the first aspects above.
[0021] Fifthly, this application provides an image display processing apparatus that has the function of implementing any of the above-described image display processing methods. This function can be implemented in hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above-described functions.
[0022] Sixthly, this application provides a chip system including a processor for implementing the functions involved in the image display processing method flow described in any one of the first aspects above. In one possible design, the chip system further includes a memory for storing program instructions and data necessary for the image display processing method. This chip system may be composed of chips or may include chips and other discrete devices. Attached Figure Description
[0023] To more clearly illustrate the technical solutions in the embodiments of the present invention or the background art, the accompanying drawings used in the embodiments of the present invention or the background art will be described below.
[0024] Figure 1 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application;
[0025] Figure 2 This is a schematic diagram of the system architecture of an electronic device provided in an embodiment of this application;
[0026] Figure 3 This is a schematic diagram of an eye-tracking principle provided in an embodiment of this application;
[0027] Figures 4A-4F This is a schematic diagram illustrating an application scenario of a set of image display processing methods provided in the embodiments of this application;
[0028] Figure 5 This is a desktop illustration of an electronic device provided in an embodiment of this application;
[0029] Figures 6A-6D This is a schematic diagram of a set of image gallery viewing interfaces provided in an embodiment of this application;
[0030] Figure 7 This is a schematic diagram illustrating an electronic device acquiring a user's eye image, provided in an embodiment of this application.
[0031] Figures 8A-8G A schematic diagram illustrating the possible first gaze region division of a set of target images provided in this application embodiment;
[0032] Figures 9A-9D A set of semantic segmentation diagrams provided for embodiments of this application;
[0033] Figures 10A-10FAnother set of semantic segmentation diagrams provided for embodiments of this application;
[0034] Figure 11 This is a schematic diagram of an interface for displaying prompt information on an electronic device, provided in an embodiment of this application.
[0035] Figures 12A-12D A set of schematic diagrams illustrating a second enlarged processing procedure provided for embodiments of this application;
[0036] Figure 13 A flowchart of an image display processing method provided in an embodiment of this application;
[0037] Figure 14 A flowchart illustrating another image display processing method provided in this application embodiment. Detailed Implementation
[0038] This application provides a method and terminal device for magnifying images by eye gaze.
[0039] The method for magnifying images by eye gaze provided in this application allows users to magnify images by relying on eye gaze. Specifically, after the image is displayed on the display interface, the smart terminal device detects the time the user gazes at a certain area of the image. When the user gazes at a certain area for a longer period than a preset threshold, the image is magnified and displayed on the display interface using the Region of Interest (ROI) within the user's gaze area as the reference area for magnification. This achieves image magnification by eye gaze.
[0040] This method can be applied to electronic devices. The electronic device is a smart terminal device, and this application does not limit the specific type of smart terminal device. For example, the electronic device can be a mobile phone, and may also include tablet computers, desktop computers, desktop computers with cameras, laptop computers, handheld computers, wearable devices (such as smartwatches, smart bracelets, etc.), augmented reality (AR) devices, virtual reality (VR) devices, artificial intelligence (AI) devices, in-vehicle systems, game consoles, etc.
[0041] The terminology used in the following embodiments of this application is for the purpose of describing particular embodiments only and is not intended to limit the scope of this application. As used in the specification and appended claims of this application, the singular expressions “a,” “an,” “some,” “the,” “the,” “the,” and “this” are intended to include the plural expressions as well, unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in this application refers to any or all possible combinations that include one or more of the listed items.
[0042] Figure 1 A schematic diagram of the structure of the electronic device 100 provided in an embodiment of this application is shown.
[0043] Electronic device 100 may include a processor 101, a memory 102, a wireless communication module 103, a mobile communication module 104, an antenna 103A, an antenna 104A, a power switch 105, a sensor module 106, a focusing motor 107, a camera 108, a display screen 109, etc. The sensor module 106 may include a gyroscope sensor 106A, an accelerometer sensor 106B, an ambient light sensor 106C, an image sensor 106D, a proximity sensor 106E, etc. The wireless communication module 103 may include a WLAN communication module, a Bluetooth communication module, etc. All of the above components can transmit data via a bus.
[0044] Processor 101 may include one or more processing units, such as an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural network processing unit (NPU). These different processing units may be independent devices or integrated into one or more processors.
[0045] Memory 102 can be used to store computer executable program code, which may include instructions. Processor 101 executes various functional applications and data processing of electronic device 100 by running the instructions stored in memory 102. Memory 102 may include a program storage area and a data storage area. In specific implementations, memory 102 may include high-speed random access memory, and may also include non-volatile memory, such as one or more disk storage devices, flash memory devices, or other non-volatile solid-state storage devices.
[0046] The wireless communication function of the electronic device 100 can be implemented through antenna 103A, antenna 104A, mobile communication module 104, wireless communication module 103, modem processor, and baseband processor.
[0047] Antennas 103A and 104A can be used to transmit and receive electromagnetic wave signals. Each antenna in electronic device 100 can be used to cover one or more communication frequency bands. Different antennas can also be reused to improve antenna utilization.
[0048] The mobile communication module 104 can provide wireless communication solutions, including 2G / 3G / 4G / 5G, for use on electronic devices 100.
[0049] The wireless communication module 103 can provide solutions for wireless communication applications on electronic devices 100, including wireless local area networks (WLAN), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), and infrared (IR).
[0050] The gyroscope sensor 106A can be used to determine the motion attitude of the electronic device 100.
[0051] Accelerometer 106B can detect the magnitude of acceleration of electronic device 100 in various directions (generally three axes).
[0052] The 106F infrared optical sensor can capture infrared light signals to enable eye-tracking functionality, analyzing the user's gaze direction and fixation point.
[0053] Electronic device 100 can perform shooting functions through ISP, camera 108, video codec, GPU, display screen 109 and application processor.
[0054] Electronic device 100 can implement display functions through a GPU, display screen 109, and application processor. The GPU is a microprocessor for image processing, connected to the display screen 109 and the application processor. The GPU is used to perform mathematical and geometric calculations and for graphics rendering. Processor 101 may include one or more GPUs, which execute program instructions to generate or modify display information.
[0055] The display screen 109 is used to display images, videos, etc. The display screen 109 includes a display panel. In some embodiments, the electronic device 100 may include one or N display screens 109, where N is a positive integer greater than 1.
[0056] The structures illustrated in the embodiments of this application do not constitute a specific limitation on the electronic device 100. In other embodiments of this application, the electronic device 100 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0057] Figure 2 This is a schematic diagram of the system architecture of the electronic device 100 provided in an embodiment of the present invention.
[0058] A layered architecture divides the system into several layers, each with a clear role and function. Layers communicate with each other through software interfaces. In some embodiments, the system is divided into five layers, from top to bottom: application layer, application framework layer, hardware abstraction layer, driver layer, and hardware layer.
[0059] The application layer may include a series of application packages. In this embodiment, the application package may include a camera, a gallery, etc.
[0060] The application framework layer provides application programming interfaces (APIs) and programming frameworks for applications in the application layer. The application framework layer includes some predefined functions.
[0061] like Figure 2 As shown, the application framework layer may include a window manager, content provider, view system, resource manager, notification manager, etc.
[0062] The window manager is used to manage windowed applications. It can retrieve screen size, determine the presence of a status bar, lock the screen, and capture screenshots, among other things.
[0063] Content providers store and retrieve data, making that data accessible to applications. This data may include videos, images, audio, made and received phone calls, browsing history and bookmarks, phone books, etc.
[0064] A view system includes visual controls, such as controls for displaying text and controls for displaying images. View systems can be used to build applications. A display interface can consist of one or more views. For example, a display interface including a text notification icon could include views for displaying text and views for displaying images.
[0065] The file explorer provides applications with various resources, such as localized strings, icons, images, layout files, video files, and more.
[0066] The notification manager allows applications to display notifications in the status bar. These notifications can be used to deliver informational messages and can disappear automatically after a short pause, requiring no user interaction. For example, the notification manager can be used to notify users of completed downloads or message alerts. The notification manager can also display notifications as icons or scrolling text in the top status bar, such as notifications from background applications, or as dialog-style notifications on the screen. Examples include displaying text messages in the status bar, emitting sounds, vibrating electronic devices, and flashing indicator lights.
[0067] In this embodiment of the application, the gallery application can utilize the interfaces and services (such as content providers) provided by the application framework layer to effectively access and manage media files (including pictures, videos, etc.) stored on the device, and at the same time use the view system to build a user interface to display these media contents in an intuitive way.
[0068] The hardware abstraction layer is an interface layer located between the application framework layer and the driver layer, providing a virtual hardware platform for the operating system. In this embodiment, the hardware abstraction layer may include a camera hardware abstraction layer, a camera algorithm library, and a sensor control center.
[0069] The camera hardware abstraction layer can provide virtual hardware for camera device 1, camera device 2, or more camera devices. The sensor control center may include an always-on sensor module that continuously monitors environmental conditions and user eye images, triggering image magnification. For example, when the AlwaysOn sensor module detects that the user's eye activity meets preset trigger conditions, it activates the gaze algorithm module. The gaze algorithm module receives eye image data or signals from the AlwaysOn sensor module and analyzes them to identify the user's gaze direction, gaze duration, and other eye expression information. It then matches the identified eye expression information with preset eye expression control commands; if a match is successful, it triggers the image magnification operation.
[0070] The driver layer is the layer between hardware and software. It includes drivers for various hardware components, such as camera drivers, image processor drivers, display drivers, sensor drivers, and digital signal processor drivers.
[0071] The camera device driver is used to drive the camera sensor to acquire images and to drive the image signal processor to preprocess the images. The digital signal processor driver is used to drive the digital signal processor to process images. The image processor driver is used to drive the graphics processor to process images.
[0072] The hardware layer is the most fundamental layer in a computer system or embedded system, directly involving the existence and operation of physical hardware devices. The hardware layer is the foundation upon which software can run, including all physically tangible and visible computer components, as well as those invisible but equally crucial components such as integrated circuits and circuit boards. The hardware layer can include sensors, image signal processors, digital signal processors, and image processors. Sensors can include sensor 1, time-of-flight (TOF) sensors, infrared optical sensors, and multispectral sensors.
[0073] The terms used in the embodiments of this application are explained below.
[0074] (1) Eye tracking: Eye tracking is a technique that tracks eye movements by measuring the position of the eye's gaze point or the movement of the eyeball relative to the head. This technique aims to monitor the user's eye movements and gaze direction when looking at a specific target, providing important data for understanding human visual behavior and attention allocation. Commonly used techniques include pupil-center corneal reflection (PCCR) and video occultation (VOG), which will be described below.
[0075] ① Pupil-Center Corneal Reflection (PCCR), also known as Pukenye image tracking, is the most commonly used method in eye-tracking technology. The following section combines... Figure 3 A schematic diagram illustrating the principle of eye tracking provides a detailed explanation of the pupil-corneal reflex method:
[0076] like Figure 3 As shown, the eyeball includes the vitreous body 1, lens 2, cornea 3, pupil 7, and iris 8, wherein:
[0077] Vitreous body 1 is a colorless, transparent, gel-like vitreous body that fills the space between the lens and the retina.
[0078] Lens 2 is a biconvex transparent tissue that is fixed and suspended behind the iris and in front of the vitreous body by the suspensory ligaments. It is the only refractive media with accommodation capabilities.
[0079] The cornea 3 is the transparent part at the front of the eyeball and is the first barrier for light to enter the eyeball. In the eyeball structure model provided in the embodiments of this application, the cornea 3 is assumed to be a spherical arc surface.
[0080] The pupil is a small round opening in the center of the iris of an animal or human eye, serving as a passage for light to enter the eye.
[0081] The iris (8) is a disc-shaped membrane with a central opening called the pupil (7).
[0082] In addition, here is a brief introduction to some important concepts or components in eye tracking:
[0083] The pupil center 4 is the geometric center of the pupil 7, which is the center of the ring structure of the pupil 7. In eye-tracking technology, by capturing the movement trajectory of the pupil center, the user's gaze direction and fixation point can be analyzed.
[0084] The Purchin spot is the reflected image formed on the outer surface of the cornea when light (usually infrared or other visible light) shines on the eye. In eye-tracking technology, the Purchin spot is often used as a stable reference point to track eye movements and gaze direction.
[0085] The corneal reflection center 6 refers to the corresponding position on the visual or image sensor of the reflection point or reflection area formed after light (such as infrared light) is reflected on the corneal surface when it shines on the eye.
[0086] Infrared rays, also known as infrared radiation, are electromagnetic waves in the infrared band with wavelengths ranging from 0.76 to 1000 micrometers, which are between visible light and microwaves. They are invisible light with frequencies lower than red light.
[0087] The pupil center line (line of sight) 10 is an imaginary straight line passing through the center of the pupil, which roughly represents the direction in which the eye is looking.
[0088] Image sensor 11, for example, can be the one described above. Figure 1 The image sensor 106D in electronic device 100 is a device that converts optical images into electronic signals and is widely used in electronic products such as digital cameras and mobile phones. In mobile phone terminals, the image sensor is mainly responsible for receiving light entering through the lens and converting it into digital signals for subsequent processing and storage.
[0089] Camera 12, for example, could be Figure 1 Cameras 1 to N108 in electronic device 100 are used to acquire images of eyes.
[0090] Infrared light source 13, electronic device 100, is a non-lighting electric light source whose main purpose is to generate infrared radiation.
[0091] Implementation steps:
[0092] Image acquisition: Using a terminal device equipped with an infrared light source 13 and a camera 12, images of the pupil center 4 and the corneal reflection center 6 are clearly captured;
[0093] Image processing: The acquired images are preprocessed to improve image quality, and then image processing algorithms are used to identify the positions of the pupil center 4 and the corneal reflection center 6;
[0094] Calculate the relative position: Based on the coordinate information of the pupil center 4 and the corneal reflex center 6, calculate the relative position vector between them. This vector can represent the direction and angle of eyeball rotation;
[0095] Gaze mapping: A calibration process is used to establish a mapping relationship between the pupil-corneal reflectance vector and the gaze point on the computer screen. During calibration, the user needs to observe points appearing at specific locations on the screen (calibration points). The eye-tracking device records the pupil-corneal reflectance vector information of these points and constructs a mapping model.
[0096] Real-time tracking: During the real-time tracking phase, the eye-tracking device continuously acquires images of the pupil and cornea, calculates the pupil-cornea reflection vector, and calculates the user's gaze point position based on the mapping model.
[0097] The pupil-corneal reflex method can accurately track eye movements and calculate the user's gaze point. This method does not require direct contact with the eye and is harmless to the user. Because it uses infrared light sources and image processing technology, the external environment has little impact on the tracking effect.
[0098] It should be noted that, Figure 3 The schematic diagram of the eyeball model shown is an exemplary illustration of an embodiment of this application. Other different forms of models may also be used, and this application does not limit them.
[0099] ② Retinal Image Localization Method: The retinal image localization method for eye tracking utilizes the unique physiological structures on the retina, such as the patterns formed by irregular capillaries and the fovea, to track eye movements by calculating changes in the retinal image. This method relies on high-precision image capture and processing technology to achieve accurate tracking of eye movements.
[0100] Basic Principles: Retinal Structure: The retina is a thin membrane specifically responsible for photoreception and image formation. Light entering through the pupil is refracted by the lens and converges onto the retina. The resolving power on the retina is uneven; the macula is the most sensitive area for light perception, and the small fovea in its center is called the fovea centralis, which contains a large number of photoreceptor cells.
[0101] Image capture: Eye-tracking devices capture images on the retina using high-precision cameras. These images contain information about the unique physiological structures of the retina.
[0102] Image processing: Using image processing techniques, the captured retinal images are analyzed and processed to calculate the changes in feature points in the image, thereby inferring the movement trajectory of the eyeball and the position of the fixation point.
[0103] Retinal image localization utilizes the unique physiological structure of the retina for eye tracking, offering high precision and accuracy. It also does not require contact with the eyeball, causing no discomfort or harm to the user. It is suitable for eye tracking in various scenarios and conditions, including eye movement monitoring during tasks such as natural viewing, reading, and driving.
[0104] (2) Semantic Segmentation: Semantic segmentation is an important task in the field of computer vision, aiming to assign each pixel in an image to a specific category label. Unlike image classification and object detection, semantic segmentation requires fine-grained classification of each pixel to achieve a deeper understanding of the image content. This technique can identify different objects in an image and segment them with pixel-level precision.
[0105] (3) Region of Interest (ROI): In image processing, computer vision, and user interaction, the region of interest specifically refers to the area that a user is particularly interested in or pays special attention to when browsing an image. It is a region in the image being processed that is delineated with a specific shape (such as a rectangle, circle, ellipse, or irregular polygon) and requires special attention or processing. This region is the focus of image analysis, processing, or user interaction. The ROI allows users or systems to select a specific area from the entire image for focus, rather than processing the entire image indiscriminately. This approach can significantly reduce the amount of data processed, improve processing efficiency, and potentially increase the accuracy of the processing results. Furthermore, the shape, size, and position of the ROI can be flexibly adjusted according to actual needs to adapt to different application scenarios.
[0106] (4) User Interface: The user interface is the medium through which applications or operating systems interact and exchange information with users. It converts the internal form of information into a form that users can understand. The user interface is source code written in specific computer languages such as Java or Extensible Markup Language (XML). This source code is parsed and rendered on the electronic device, ultimately presenting content that the user can recognize. A common form of user interface is the graphical user interface (GUI), which refers to a user interface related to computer operation displayed graphically. It can be visible interface elements such as text, icons, buttons, menus, tabs, text boxes, dialog boxes, status bars, navigation bars, and widgets displayed on the screen of an electronic device.
[0107] First, to facilitate understanding of the embodiments of this application, the specific technical problem to be solved by this application is further analyzed and proposed. In some implementations of the prior art, users can control the zoom in or out of an image by looking at it and blinking. For example, as... Figure 4A and Figure 4B As shown, when a user clicks to view an image using the gallery app, a 402 error occurs; the electronic device displays... Figure 4B In the image gallery viewing interface 1 shown, when the user looks at image 402, the electronic device uses the pupil-corneal reflection method, with the front-facing camera 407 and infrared light source 406, to acquire the user's eye image and gaze position, determine whether the user is looking at the image, and capture the user's conscious blinking behavior. In response to this behavior, the electronic device magnifies the image with the center of the image as the reference area for magnification and displays it on the display screen 411 shown in the image gallery viewing interface 2.
[0108] In the aforementioned existing technologies, when a user controls image zooming by eye contact, the electronic device uses the image center as the reference area for zooming, which is likely to be mismatched with the area the user needs to zoom in on, thus failing to meet the user's need to zoom in on a specific area and resulting in a poor user experience.
[0109] In some existing technological solutions, such as Figure 4B As shown, after image 412 is displayed on screen 411, the electronic device can continue to monitor the user's eyes through front-facing camera 407 and infrared light source 406. If it detects the user blinking consciously again, it triggers image scaling down, restoring the image to its state before magnification, i.e., image 412 reverts to image 402. However, the image resolution may be one or more times higher than the resolution of the electronic device's screen, and the image contains rich details. Zooming in only once cannot fully extract the information contained in the image, resulting in a waste of image resolution.
[0110] In other solutions of the prior art, electronic devices can also use retinal image localization to obtain information about the user's eye gaze.
[0111] Therefore, based on the problems existing in the prior art, the technical problem to be solved by this application may include the following aspects: providing an image processing and display method that, while enabling eye-controlled image magnification, magnifies the image based on the area of interest most likely to be of interest to the user, and enables multiple magnifications of the image by eye-controlled eye, thereby improving the user experience and reducing power consumption.
[0112] The reference area for magnification refers to the region selected by the user as the starting point for magnification or the center of the magnified image during the image magnification process. The reference area can be selected in various ways, depending on the user's specific actions, such as two-finger swipe magnification, double-tap magnification, or magnification based on eye gaze. For example, when using two fingers to slide on the touchscreen (usually the reverse of a pinch gesture, i.e., fingers spread apart) to adjust the image size, the center point between the two fingers is considered the reference area. As the fingers spread apart, the image is magnified with this center point as the reference, keeping the center point in the center of the magnified image. When the user taps a point on the screen twice consecutively to trigger the partial magnification function, the second tap becomes the reference area. Larger operations will magnify the surrounding image area with this point as the center, making that point the center of the magnified image.
[0113] The following describes the application scenarios involved in the embodiments of this application.
[0114] Scenario 1: Using the image viewing method in this application, we can enable image zooming in on a gallery application based on eye gaze.
[0115] See Figure 4A and Figure 4B , Figure 4A This illustration shows an application scenario of zooming in on an image through eye gaze in a gallery application, as provided in an embodiment of this application. In this scenario, the user opens the gallery application on their electronic device, selects and views an image, and the electronic device displays as follows: Figure 4B The image viewing interface 1 in the image library shows a display interface 401. As shown in the image viewing interface 1, the display interface 401 may include an image 402, an information viewing control 403, basic information 404, a return control 405, and an operation menu bar 408. The operation menu bar 408 may include operation options such as share, favorite, edit, delete, and more. When a user consciously gazes at a certain area or object in an image, the electronic device can use pupil-corneal reflex or retinal image localization to track the user's eye movements and magnify the image based on the user's gaze.
[0116] Scenario 2: Using the image viewing method described in this application, we can zoom in on images in the Weibo application based on eye gaze.
[0117] See Figure 4C , Figure 4C This is a schematic diagram illustrating an application scenario in the "Weibo" application where images are magnified by eye contact, as provided in an embodiment of the present invention. Figure 4CAn exemplary illustration shows a browsing interface 421 of a "microblogging" application on an electronic device. The browsing interface 421 may include posts 422 posted by other users or the user themselves, and an interface menu bar 425. Posts 422 may include images 423 and comment viewing controls 424. Users can click to view images 423 on the electronic device's display screen, and can also click the comment viewing control 424 to access comment details, such as... Figure 4D As shown, the comment interface 431 may include comments 432 and an operation menu bar 434. Comments 432 may include comment images 433, which users can click to view. When a user clicks to view an image in the Weibo application, the electronic device displays the image on its screen. Figure 4E As shown, the image viewing interface 441 may include an image 442, a collapse control 443, and other operation controls 444. When a user consciously gazes at a certain area or object in the image, the electronic device can use pupil-corneal reflex or retinal image localization to track the user's eye movements and magnify the image by focusing the gaze. Optionally, the electronic device may include an infrared light source 445.
[0118] Scenario 3: Using the image viewing method in this application, the paused frame can be enlarged by eye gaze when the video is paused.
[0119] See Figure 4F , Figure 4F This invention provides a schematic diagram of an application scenario where eye-tracking magnifies a paused frame during video pause, including a video pause interface 1 and a video pause interface 2. In this scenario, video playback can be implemented in applications capable of playing videos on a display screen, such as a gallery application. When the video is paused, the display screen does not show any content other than the paused frame (image). The video can be played in portrait or landscape mode. After the video pauses, the user can consciously gaze at a certain area or object of the paused frame 451 or 452. The electronic device can use pupil-corneal reflex or retinal image localization to track the user's eye movements and achieve eye-tracking magnification of the image. Optionally, the electronic device may include an infrared light source 453.
[0120] Understandable, Figures 4A-4F The application scenarios described above are merely a few exemplary implementations in the embodiments of the present invention. The application scenarios in the embodiments of the present invention include, but are not limited to, the above-described application scenarios. Optionally, the method in this application can also be applied to scenarios such as image processing and retouching. When a user needs to process a local area of an image, it is usually necessary to first magnify the local area before processing. Applying this method allows the user to zoom in on the image centered on the local area to be processed through gaze, and then further edit the local area. Other scenarios and examples will not be listed or elaborated upon.
[0121] The user interface involved in the image processing and display method provided in the embodiments of this application is described below with reference to the accompanying drawings.
[0122] The image processing and display method provided in this application embodiment can be applied to... Figure 5 , Figure 6A , Figure 6B The image viewing scenario shown could be using an electronic device to view a single image through a gallery app or viewing a frame of an image on the screen when a video is paused. Figure 5 A desktop illustration of an electronic device provided in this application embodiment, wherein the electronic device includes... Figure 5 The front-facing camera 504 and display screen 505 are shown in the diagram. The display screen 505 of the electronic device displays as shown... Figure 5 The desktop shown includes a status bar 501 and an application menu bar 502. The status bar 501 displays the carrier, current time, network status, signal strength, and battery level. Figure 5 As shown, the operator is China Mobile; the current time is 08:08; the network status is Wi-Fi; the signal strength is full, indicating a strong signal; the black portion of the power indicator represents the remaining battery power of the electronic device. The application menu bar 502 includes icons for at least one application, each with its corresponding application name below it, such as: Camera, Gallery, Contacts, Dialer, Messages, Weather, etc. The positions of the application icons and their names can be adjusted according to user preferences, and this embodiment does not limit this.
[0123] It should be noted that, Figure 5 The desktop illustration of the electronic device shown is an exemplary representation of an embodiment of this application. The desktop illustration of the electronic device may also be of other styles, and this embodiment of the application does not limit it.
[0124] In some embodiments, the user can click Figure 5 The gallery application icon 503 in the application menu bar 502, as shown, is displayed by the electronic device in response to this operation. Figure 6A The interface shown is a schematic diagram of a gallery application interface in an electronic device provided in an embodiment of this application. Figure 6AAs shown, the interface includes an image search bar 601 and an album menu bar 602. The image search bar 601 can display prompts such as "people, places, etc.", allowing users to search for desired images. The album menu bar 602 can include multiple album options such as camera, screenshot, and video. Each album option control uses a thumbnail of the first image in the corresponding album as its cover. By clicking on different album option covers, users can view all the corresponding images or videos under that album option on the display screen. Different album options also display the number of images that can be viewed under that album option.
[0125] In some embodiments, such as Figure 6A As shown, users can click the camera album option, and the electronic device will respond to this action by displaying the following: Figure 6B The interface shown is a schematic diagram of the camera album options provided in this embodiment of the application. Figure 6B As shown, the interface includes a return control 611 and image options 612. Image options 612 can include one or more image thumbnails corresponding to the original image. Users can swipe down or up on the screen to browse the image thumbnails and find the image they want to view. Once the user finds the desired image, they can click on the thumbnail corresponding to that image to view it. The electronic device receives the user's click input and further responds to the input by displaying... Figure 6C The interface shown is a schematic diagram of the interface for viewing images through a gallery application provided in an embodiment of this application.
[0126] like Figure 6C As shown, the interface includes a view image (621), a basic information bar (622), an image operation menu bar (623), a return control (624), and an information viewing control (625). The basic information bar (622) can include time and address information, such as "August 5, 2024, Nanshan District, Shenzhen". The image operation menu bar (623) can include operations such as "share", "favorite", "edit", "delete", and "more". Clicking the information viewing control (625) allows users to view more information about the image, including but not limited to storage path, image size, and image resolution.
[0127] It should be noted that, Figures 6A-6C The illustrated user interface diagram of the electronic device is an exemplary representation of an embodiment of this application. The user interface diagram of the electronic device may also be of other styles, and this application does not limit them.
[0128] The following is combined Figures 5-12D , Figure 13 and Figure 14 Explain how electronic devices can control image magnification by eye movements.
[0129] Among them, electronic devices can be Figure 1 The illustrated electronic device 100 includes a front-facing camera and a display screen, and may also include an infrared light source; furthermore, the electronic device can support… Figure 2 The system architecture is shown.
[0130] First, combine Figures 5-10F and Figure 13 Explain how electronic devices can control image magnification by eye movements. Figure 13 This is a flowchart of an image display processing method provided in an embodiment of this application.
[0131] S101: The electronic device displays the target image on the screen.
[0132] Specifically, the process of viewing and displaying an image on a screen can be found in [link to relevant documentation]. Figure 5 , Figure 6A , Figure 6B The corresponding textual descriptions are not elaborated here. It is understood that the interface illustrated in the embodiments of the present invention does not constitute a specific limitation on the electronic device, and the process of viewing and displaying images on the screen is not limited to... Figure 5 , Figure 6A , Figure 6B The process is shown.
[0133] S102: Obtain the first gaze region and the first gaze time of the target image being gazed at by the user.
[0134] The first fixation time can be the time during which a user continuously fixates on a specific area (by analyzing multiple consecutive frames of user eye images acquired by a camera to determine that the user's eye fixation position has not changed). Alternatively, the first fixation time can also include the time during which the user primarily fixates on a fixed area over a longer period, even if the area of fixation changes briefly during this time; the user's primary area of fixation remains relatively fixed during this period.
[0135] Electronic devices can first determine the first fixation area by capturing the first few frames of user eye images, and then determine the first fixation time, or they can first determine the first fixation time, and then determine the first fixation area based on the acquired user eye images.
[0136] Specifically, the electronic device displays a target image on the screen. When the user has a need to zoom in, they consciously look at a certain area of the target image. The electronic device continuously captures multiple frames of the user's eye images at relatively long time intervals using a front-facing camera. The electronic device analyzes the acquired multi-frame eye images with long time intervals to determine whether the user is looking at the image. If the user's gaze area does not shift significantly, it is determined that the user is looking at the image. The electronic device then captures multiple high-resolution images of the user's eye at shorter time intervals using a front-facing camera. The electronic device analyzes the acquired multi-frame eye images with shorter time intervals to determine first gaze information, which includes a first gaze area and a first gaze time.
[0137] In one possible implementation, when the user completes... Figure 6B The image thumbnail is shown as an input method for clicking, and the electronic device displays... Figure 6C As shown in the interface, the electronic device begins acquiring first gaze information. For example, the electronic device first determines the first gaze duration, and then determines the first gaze region based on the acquired image of the user's eyes. Figure 7 As shown, the electronic device displays the target image 702 on the screen and continuously captures 50 frames of user eye images per second for 1 second using the front-facing camera 701 at a low sampling rate of 50 frames per second. The electronic device analyzes these 50 frames of user eye images. If the distance between the center points of the gaze area obtained after analyzing more than 30 consecutive frames of user eye images does not exceed the maximum pixel distance, it is determined that the user is gazing at the image. The maximum pixel distance refers to the maximum allowable pixel offset used to determine whether the user's gaze area has changed significantly. If no significant displacement is found in the gaze area after analyzing more than 30 consecutive frames of user eye images, the front-facing camera 701 continues to capture 50 frames of user eye images per second for 1 second until no significant displacement is found in the gaze area after analyzing more than 30 consecutive frames of user eye images.
[0138] For example, after the electronic device determines that the user is looking at the image, such as Figure 7 As shown, the front-facing camera 701 continuously captures multiple frames of the user's eye images at a high sampling rate of 100 frames per second for 1.2 seconds, acquiring 120 frames of user eye images. The 120 frames are analyzed to obtain the area the user is looking at in each frame. Images with 50 or more consecutive frames where the user's gaze area has not shifted significantly are detected, and the number of consecutive frames is counted. The first gaze time is calculated using the largest consecutive frame count. Specifically, the calculation method is as follows:
[0139] T1=(N1-1)·0.01
[0140] Where T1 is the first gaze time, and N1 is the first gaze region of the target image obtained by executing step S202 and the maximum number of consecutive image frames detected during the first gaze time.
[0141] For example, the electronic device can determine the first gaze region after determining the first gaze time. Specifically, the electronic device can determine the first gaze region of the target image on the display screen based on the first gaze region of the target image obtained in step S202 and the continuous image corresponding to the largest number of consecutive image frames detected in the first gaze time.
[0142] The first gaze region can be defined by coarse-precise coordinates. Before determining the first gaze region, the electronic device divides the image into possible first gaze regions based on the aspect ratio, and then analyzes the captured image of the user's eyes to determine the first gaze region.
[0143] For example, such as Figures 8A-8G As shown, for images with common aspect ratios in electronic devices, the electronic device can follow... Figures 8A-8G The diagram shows the possible first fixation region: 1) 1:1
[0145] like Figure 8A As shown, for an image with an aspect ratio of 1:1, a straight line connecting the centers of two parallel sides of the image divides the image into four possible first gaze regions 801, 802, 803, and 804, which are squares. 2) 16:9
[0147] like Figure 8B As shown, for an image with an aspect ratio of 16:9, the center of the long side (the side with a length ratio of 16:9 to the adjacent side) is connected to the center of the other long side. Similarly, the center of the short side (the side with a length ratio of 9:16 to the adjacent side) is connected to the center of the other short side, dividing the image into four possible first viewing regions 811, 812, 813, and 814. 3) 9:16
[0149] like Figure 8C As shown, for an image with an aspect ratio of 9:16, the center of the shorter side (the side with a length ratio of 9:16 to the adjacent side) is connected to the center of the other shorter side. Similarly, the center of the longer side (the side with a length ratio of 16:9 to the adjacent side) is connected to the center of the other longer side, dividing the image into four possible first gaze regions 821, 822, 823, and 824. 4) 4:3
[0151] like Figure 8DAs shown, for an image with an aspect ratio of 4:3, the center of the longer side (the side with a length ratio of 4:3 to the adjacent side) is connected to the center of the other longer side. Similarly, the center of the shorter side (the side with a length ratio of 3:4 to the adjacent side) is connected to the center of the other shorter side, dividing the image into four possible first gaze regions 831, 832, 833, and 834. 5) 3:4
[0153] like Figure 8E As shown, for an image with an aspect ratio of 3:4, the center of the shorter side (the side with a length ratio of 3:4 to the adjacent side) is connected to the center of the other shorter side. Similarly, the center of the longer side (the side with a length ratio of 4:3 to the adjacent side) is connected to the center of the other longer side, dividing the image into four possible first viewing regions 841, 842, 843, and 844. 6) 3:2
[0155] like Figure 8F As shown, for an image with an aspect ratio of 3:2, the center of the longer side (the side with a length ratio of 3:2 to the adjacent side) is connected to the center of the other longer side. Similarly, the center of the shorter side (the side with a length ratio of 2:3 to the adjacent side) is connected to the center of the other shorter side, dividing the image into four possible first gaze regions 851, 852, 853, and 854. 7)2:3
[0157] like Figure 8G As shown, for an image with an aspect ratio of 2:3, the center of the shorter side (the side with a length ratio of 2:3 to the adjacent side) is connected to the center of the other shorter side. Similarly, the center of the longer side (the side with a length ratio of 3:2 to the adjacent side) is connected to the center of the other longer side, dividing the image into four possible first gaze regions 861, 862, 863, and 864.
[0158] The images viewed by users can also be images with any aspect ratio, and this embodiment of the invention does not impose any limitations.
[0159] In one possible implementation, the electronic device can calculate and determine the first gaze region of the target image on the user's viewing display screen based on all continuous images corresponding to the maximum number of continuous image frames in the above implementation, or it can calculate and determine the first gaze region of the target image on the user's viewing display screen based on the first frame, one or two middle frames and the last frame of the continuous images, or it can calculate and determine the first gaze region of the target image on the user's viewing display screen based on the last frame of the continuous images corresponding to the maximum number of continuous image frames.
[0160] In one possible implementation, the electronic device can use an artificial intelligence (AI) model to analyze the user's eye images. The electronic device takes multiple frames of the user's eye images captured by the front-facing camera at a high or low sampling rate as input to the AI model, and outputs the first gaze region and the first gaze time after calculation.
[0161] It should be noted that, Figures 8A-8G The schematic diagram of the possible first gaze region division of the target image shown is an exemplary illustration of the embodiments of this application. Other division methods may be used when dividing the possible first gaze region, and the embodiments of this application do not limit this.
[0162] In the above embodiments, the user performs actions on the electronic device. Figure 6B The input shown is for clicking on an image thumbnail, and it is displayed. Figure 6C After the interface shown, you can input the following into the electronic device again: Figure 6C As shown, when a click is made to view image 621, the electronic device responds to this input and displays the following: Figure 6D The image is displayed in full-screen mode. The screen does not show anything other than image 631. This input will interrupt the view. Figure 6B The process of detecting whether the user is looking at the image is triggered by clicking on the image thumbnail, and will trigger a new process of detecting whether the user is looking at the image, that is, the electronic device re-executes step S202 to obtain the first gaze area and the first gaze time of the user looking at the target image.
[0163] S103: Determine whether the first fixation time is greater than or equal to the first threshold.
[0164] Specifically, after acquiring the first gaze time, the electronic device compares the first gaze time with a preset first threshold. If the first gaze time is less than the first threshold, the electronic device repeats step S202 to acquire the first gaze region and the first gaze time of the user's gaze on the target image. If the first gaze time is greater than or equal to the first threshold, step S204 is executed to determine the first target region of interest within the first gaze region. The first threshold is used to determine whether the user is consciously gazing at a certain area of the image displayed on the screen in order to achieve image magnification.
[0165] In one possible implementation, the electronic device can use an AI model to determine whether the user's first gaze time is greater than or equal to a first threshold. The electronic device uses the first gaze time as input to the AI model, analyzes it, and outputs the relationship between the first gaze time and the first threshold. The relationship includes two possibilities: greater than or equal to, and less than.
[0166] For example, the first threshold can be set to 0.9 seconds. When the first gaze time is less than 0.9 seconds, the electronic device repeatedly executes step S202 to obtain the first gaze region and the first gaze time of the user's gaze on the target image; when the first gaze time is greater than or equal to 0.9 seconds, step S204 is executed. If the first threshold is exceeded, the first target interest region within the first gaze region is determined based on the first gaze region and the content of the first gaze region.
[0167] S104: If the first threshold is exceeded, then based on the first gaze region and the content of the first gaze region, determine the first target region of interest within the first gaze region.
[0168] First, the electronic device performs semantic segmentation on the target image displayed on the screen. Specifically, the electronic device performs semantic segmentation on the target image according to a first semantic category, identifying Q second-level semantic regions in the target image, where Q is a positive integer less than M; it then performs semantic segmentation on each of the Q second-level semantic regions according to a second semantic category, identifying one or more first-level semantic regions within each second-level semantic region; finally, the first-level semantic regions corresponding to the Q second-level semantic regions are determined as the M first-level semantic regions.
[0169] In one possible implementation, the first semantic category includes one or more of people, animals, plants, and scenery; the second semantic category includes one or more of human facial features, hair, arms, legs, hands, feet, clothing, tree crowns, tree trunks, flower crowns, plant stems, animal facial features, and limbs.
[0170] In one possible implementation, when determining Q second-level semantic regions based on a first semantic category, a single target image may include multiple objects corresponding to the first semantic category. When objects corresponding to the first semantic category overlap, among every two overlapping objects, the object with relatively complete semantics and positioned higher in the image corresponding to the first semantic category is called the occluder, and the object whose portion is obscured by the occluder and positioned lower in the image is called the occluded object. Specifically, when the occluder only partially covers the occluded object, causing part of the occluded object to be invisible, but not segmented into two or more independent regions that are not directly connected in space, simple image analysis techniques, such as color thresholding and edge detection, are first used to preliminarily identify the occluder in the image. Based on the detected occluder, an edge detection algorithm is further used to extract the precise edge of the occluder. The edge of the occluded object is determined based on the precise edge of the occluder and the unoccluded portion of the occluded object. The closed regions enclosed by the edges of the occluder and the occluded object are respectively determined as Q second-level semantic regions.
[0171] For example, Figure 9AThe semantic segmentation diagram shown Figure 1 This is a schematic diagram illustrating the result of semantic segmentation of the target image 901 according to the first semantic category. For example... Figure 9A As shown, target image 901 contains two objects corresponding to the first semantic category, including a person and a plant. The person and plant, corresponding to the first semantic category, overlap; the person is identified as an occluder, and the plant as the occluded object. The electronic device performs semantic segmentation on target image 901 displayed on the screen according to the first semantic category, obtaining two second-level semantic regions, namely... Figure 9A The closed regions enclosed by the dashed boxes are the second-level semantic region 902 and the second-level semantic region 903, respectively, where Q equals 2. The semantic segmentation of the electronic device first ensures the semantic integrity of the second-level semantic region 902. Furthermore, it determines the edges of the second-level semantic region 903 based on the segmentation edges of 902 and the unoccluded portions of 903, while striving to maintain the integrity of the second-level semantic region 903 without compromising its edges.
[0172] In one possible implementation, when the target image contains multiple objects corresponding to a first semantic category, and the objects corresponding to the first semantic category overlap with two or more objects, but the overlapping parts still satisfy that the occluder in each pair of objects corresponding to the first semantic category only partially covers the occluded object, causing part of the occluded object to be invisible, but it is not divided into two or more independent regions that are not directly connected in space, and is still segmented according to the semantic segmentation method described above.
[0173] For example, Figure 9B The semantic segmentation diagram shown Figure 2 This is a schematic diagram illustrating the result of semantic segmentation of the target image 911 according to the first semantic category. For example... Figure 9BAs shown, target image 911 contains three objects corresponding to the first semantic category, including a male figure, a female figure, and a tree. The male figure and female figure overlap, with the male figure acting as an occluder and the female figure as the occluded object. The female figure and the tree also overlap, with the tree acting as an occluder and the female figure as the occluded object. The female figure overlaps with two objects corresponding to the first semantic category, and in both instances, the occluder only partially covers the occluded object, making a portion of the occluded object invisible, but without dividing it into two or more spatially unconnected independent regions. The second-level semantic region corresponding to the male character is determined, and semantic segmentation is performed according to the first semantic category to obtain the second-level semantic region 912. Simultaneously, the edges of the second-level semantic region 912 in the overlapping area between the male and female characters are obtained. The second-level semantic region corresponding to the tree is determined, and semantic segmentation is performed according to the first semantic category to obtain the second-level semantic region 913. Simultaneously, the edges of the second-level semantic region 913 in the overlapping area between the tree and the female character are obtained. Based on the edges of the second-level semantic region 912 in the overlapping area between the male and female characters, the edges of the second-level semantic region 913 in the overlapping area between the tree and the female character, and the unoccluded portion of the female character (i.e., the portion not overlapping with other objects corresponding to the first semantic category), the second-level semantic region 914 corresponding to the female character is determined. At this point, Q equals 3.
[0174] In one possible implementation, if the occlusion not only partially covers the occluded object, but also covers it to a degree sufficient to segment the originally continuous occluded object into two or more spatially disjoint regions, a deep learning model is used to perform semantic segmentation on the target image according to a first semantic category. Specifically, the image to be processed is input into a trained model, which performs a series of forward propagation calculations, including feature extraction, feature fusion, and contextual understanding, to comprehensively analyze the occlusion situation in the image. Based on its understanding of the image content, the model infers the edges of the occluded object using prior knowledge. The model outputs Q determined second semantic regions.
[0175] In one possible implementation, each of the Q second-level semantic regions is semantically segmented according to the second semantic category, and M first-level semantic regions are identified in each second-level semantic region, where M is a positive integer greater than Q.
[0176] For example, Figure 9C The semantic segmentation diagram shown Figure 3 This is a schematic diagram showing the result of semantic segmentation of target image 921 according to the second semantic category, where target image 921 and... Figure 9A The target image 901 is the same image, and semantic segmentation of target image 921 according to the second semantic category is performed inFigure 9A The semantic segmentation diagram shown Figure 1 This is based on the first semantic category segmentation, specifically, semantic segmentation of the resulting image obtained according to the first semantic category according to the second semantic category. Figure 9A In this case, Q is 2, meaning there are two second-level semantic regions. Figure 9A In the second-level semantic region 902, it was identified according to the second semantic category. Figure 9C The first-level semantic regions shown are 922, 923, 924, 925, 926, 927, 928, 929, and 9210, a total of nine first-level semantic regions. Among them, first-level semantic region 922 is the left eye, first-level semantic region 923 is the right eye (here, left and right are determined by the relative positions of the two eyes in the image), first-level semantic region 924 is the nose, first-level semantic region 925 is the mouth, first-level semantic region 926 is the right ear, first-level semantic region 927 is the left hand, first-level semantic region 928 is the right hand (here, left and right are determined by the relative positions of the two hands in the image), first-level semantic region 929 is the jacket pocket, and first-level semantic region 9210 is the bow tie; Figure 9A In the second-level semantic region 903, it was identified according to the second semantic category. Figure 9C The first-level semantic regions shown are 9211 and 9212, a total of two first-level semantic regions. First-level semantic region 9211 represents the tree trunk, and first-level semantic region 9212 represents the tree crown. In this case, M is 11. Then... Figure 9A The consensus of two second-level semantic regions is derived from 11 first-level semantic regions.
[0177] It should be noted that, Figures 9A-9C The semantic segmentation diagram of the target image shown is an exemplary illustration of an embodiment of this application. Other semantic segmentation methods can be used when performing semantic segmentation on the target image, and this application embodiment does not limit them.
[0178] Next, from the M first-level semantic regions, N first-level semantic regions that intersect with the first gaze region determined in step S202 are identified. These N first-level semantic regions that intersect with the first gaze region may include first-level semantic regions segmented according to the second semantic category and completely covered by the first gaze region, and first-level semantic regions segmented according to the second semantic category and partially covered by the gaze region.
[0179] Based on the preset magnification priority of the N first-level semantic regions and the area occupied by the N first-level semantic regions in the first gaze region, the first target region of interest is determined.
[0180] The preset magnification priority is a set of priority rules for object magnification display pre-set in electronic devices based on factors such as product design concepts and user experience. These rules can be based on various factors such as object type (e.g., facial features, limbs, clothing, plants, animals), size, position, and color. The magnification priority rules are encoded and stored in the firmware of the electronic device or an updatable database so that the electronic device can call them at any time.
[0181] In one possible implementation, after determining the N first-level semantic regions covered within the user's first gaze region, the electronic device traverses the N first-level semantic regions covered within the first gaze region, which are semantically segmented according to a second semantic category. Based on a preset magnification priority, it determines the magnification priority of the first-level semantic regions covered within each first gaze region and counts the pixels of each first-level semantic region to obtain the pixel count as the area corresponding to each first-level semantic region. When the first-level semantic region covered by the first gaze region is a first-level semantic region segmented according to the second semantic category and completely covered by the first gaze region, the total number of pixels in this first-level semantic region is counted as the area corresponding to this first-level semantic region. When the first-level semantic region covered by the first gaze region is a first-level semantic region segmented according to the second semantic category and partially covered by the gaze region, only the pixel count of the portion of the first-level semantic region covered by the first gaze region is counted as the area of this first-level semantic region. The total number of pixels in the first gaze region is then counted as the area of the first gaze region.
[0182] In one possible implementation, the electronic device calculates the magnification weight of each of the N first-level semantic regions based on the magnification priority and area of the N first-level semantic regions covered by the first gaze region, and selects the first-level semantic region with the largest magnification weight as the first target region of interest. If the first-level semantic region with the largest magnification weight is a first-level semantic region segmented according to the second semantic category and completely covered by the first gaze region, then this first-level semantic region is selected as the first target region of interest. If the first-level semantic region with the largest magnification weight is a first-level semantic region segmented according to the second semantic category and partially intersects with the first gaze region, then the intersection of this first-level semantic region and the first gaze region is selected as the first target region of interest.
[0183] For example, Figure 9D This is a semantic segmentation diagram (4). Figure 9D This shows the location of the first gaze region in the target image and the first-level semantic regions that intersect with it. For example... Figure 9DAs shown, the target image 931 has a resolution of 2340x1080 pixels and can be divided into four possible first gaze regions. The first gaze region 932, which is the upper left rectangular area of the target image 931, is the first gaze region. The first-level semantic regions covered by the first gaze region include first-level semantic regions 933, 934, 935, 936, and 937. Among them, first-level semantic region 934 is the right eye (here, left and right in the left and right eyes are determined by the relative positions of the two eyes in the image), first-level semantic region 935 is the nose, first-level semantic region 936 is the mouth, and first-level semantic region 937 is the bow tie. First-level semantic regions (right eye) 934 and first-level semantic regions (mouth) 936 are both first-level semantic regions that intersect with the first gaze region. When counting pixels, only the intersection with the first gaze region 932 is counted, and N is 5 in this case. The electronic device's preset magnification priority for the first-level semantic region is an integer from 0 to 10, with each first-level semantic region corresponding to an integer representing the magnification priority. The first-level semantic region (left eye) 933 has a magnification priority of 7 and an area of 60,000 pixels; the first-level semantic region (right eye) 934 has a magnification priority of 7, with an intersection area of 40,000 pixels with the first gaze region; the first-level semantic region (nose) 935 has a magnification priority of 6 and an area of 90,000 pixels; the first-level semantic region (mouth) 936 has a magnification priority of 6, with an intersection area of 100,000 pixels with the first gaze region; and the first-level semantic region (bow tie) 937 has a magnification priority of 3 and an area of 170,000 pixels; the area of the first gaze region is 631,800 pixels.
[0184] Optionally, the formula for calculating the magnification weight W1 of the first-level semantic region is:
[0185] W1 = L n1 ·V l1 +A n1 ·V a1
[0186] Where L n1 V represents the amplification priority of the first-level semantic region after normalization. l1 The weight of the priority amplification for the first-level semantic region is a pre-set value for the electronic device, A n1 V represents the area of the first-level semantic region after normalization. a1 The area weight is a preset value for electronic devices.
[0187] Optionally, L n1 The calculation formula is:
[0188]
[0189] Where L r1 L is the magnification priority determined after the electronic device traverses the first-level semantic region covered by the first gaze region. min1 The maximum value of the amplification priority of the first-level semantic region preset for electronic devices, L max1 The maximum value of the amplification priority of the first-level semantic region preset for electronic devices.
[0190] Optionally, A n1 The calculation formula is:
[0191]
[0192] Where A o1 A is the area of each first-level semantic region covered by the first gaze region traversed by the electronic device. r1 The area of the first gaze region.
[0193] For example, the weight V of the amplification priority l1 =0.6, area weight V a1 =0.4, Figure 9D The first-level semantic region (left eye) 933 shown corresponds to L r1 It is 7, L max1 For 10, L min1 If it is 1, then Corresponding A o1 For 60000, A r1 If it is 631800, then The corresponding amplification weight W1 = l n1 ·V l1 +A n1 ·V a1 = 0.666666 * 0.6 + 0.094967 * 0.4 = 0.437986. Similarly, Figure 9D The magnification weights of other first-level semantic regions covered within the first gaze region 932 shown are calculated in the same way. The magnification weights are: first-level semantic region (right eye) 934 W1 = 0.425324, first-level semantic region (nose) 935 W1 = 0.391984, first-level semantic region (mouth) 936 W1 = 0.398501, and first-level semantic region (bow tie) 937 W1 = 0.244118. The first-level semantic region (left eye) 933 with the largest magnification weight is selected as the first target region of interest.
[0194] In one possible implementation, the electronic device can provide a user feedback mechanism that allows users to adjust the magnification priority rules or select specific objects to be magnified according to their preferences.
[0195] S105: Perform a first magnification process on the target image based on the first target region of interest, and display it on the display screen.
[0196] Specifically, the first magnification process can be the first magnification process of the target image, or it can be any magnification process of the target image other than the first magnification process.
[0197] Specifically, when performing a first magnification process on the target image based on the first target region of interest, the first target region of interest can be directly used as the reference region for magnification, meaning the magnified target image is centered on the first target region of interest. Alternatively, the reference region for magnification can be determined based on the first target region of interest and a larger region containing the first target region of interest, and the magnified target image can be centered on the larger region containing the first target region of interest.
[0198] In one possible implementation, if the semantics of the first target region of interest are complete, then the first target region of interest is used as the first reference region for the first magnification process, and the target image is magnified at a fixed magnification based on the first reference region.
[0199] The definition of the reference area can be found in the description of the reference area in the application scenarios involved in the embodiments of this application, and will not be repeated here.
[0200] For example, such as Figure 10A As shown, the electronic device displays a target image 1001 on the screen without the first magnification process. The electronic device determines the first gaze region 1002 of the user's gaze and determines the first target region of interest 1003. The first target region of interest 1003 is semantically complete and is directly used as the reference region for the first magnification process. The target image is magnified at a fixed magnification of 1.5 times and displayed on the screen. Figure 10B As shown, at this time, the electronic device displays a target image 1011 that has undergone the first magnification process, with the first target region of interest 1003 as the center of the image.
[0201] In one possible implementation, if the semantics of the first target region of interest are incomplete, it is determined whether the first target region of interest belongs to a part of a first-level semantic region or the whole region. If the first target region of interest belongs to a part of a first-level semantic region, then the first-level semantic region to which the first target region of interest belongs is used as the second reference region for the first amplification process.
[0202] For example, such as Figure 10C As shown, the electronic device displays the target image 1021 without the first magnification processing on the display screen. The electronic device determines the first gaze region 1022 that the user is looking at, and determines the first target interest region 1023 within the first gaze region 1022. The first target interest region is an arc-shaped region of the first-level semantic region 1024 that is cut out by the edge of the first gaze region and covered by the first gaze region. The first target interest region 1023 is semantically incomplete, and the first target interest region 1023 belongs to a part of the first-level semantic region 1024. Therefore, the electronic device determines the first-level semantic region 1024 as the reference region for the first magnification processing, and performs the first magnification processing on the target image 1021 at a fixed magnification of 1.5 times, and displays it on the display screen. Figure 10D As shown, at this time, the electronic device display screen shows the target image 1031 after the first magnification process, with the first-level semantic region 1024 as the center of the image.
[0203] In one optional implementation, if the semantics of the first target region of interest are incomplete and the first target region of interest belongs to the entire region of a certain first-level semantic region, the second-level semantic region to which this first-level semantic region belongs is used as the second reference region for the first magnification process; the target image is magnified based on the second reference region.
[0204] For example, such as Figure 10E As shown, the electronic device displays the target image 1041 without the first magnification processing on the display screen. The electronic device determines the first gaze region 1042 that the user is looking at, and determines the first target region of interest 1043 in the first gaze region 1042. The first target region of interest 1043 is semantically incomplete, and the first target region of interest 1043 belongs to the entire area of the first-level semantic region, that is, the first target region of interest 1043 is also the first-level semantic region 1043. Then the electronic device determines the second-level semantic region 1044 to which the first target region of interest 1043 belongs as the reference area for the first magnification processing. The second-level semantic region 1044 belongs to the occluded object, and is partially occluded and does not become multiple disconnected regions. The target image 1041 is magnified at a fixed magnification of 1.5 times and displayed on the display screen. Figure 10F As shown, at this time, the electronic device display screen shows the target image 1051 after the first magnification process, with the second-level semantic region 1044 as the center of the image.
[0205] In one optional implementation, if the semantics of the first target region of interest are incomplete and the first target region of interest belongs to the entire region of a certain first-level semantic region, the detection step S204 determines whether the first-level semantic region with the largest magnification weight calculated in the first target region of interest within the first gaze region is semantically complete. If this first-level semantic region is semantically complete, it is used as the second reference region for the first magnification process. If this first-level semantic region is semantically incomplete, the detection step S204 determines whether the first-level semantic region with the second largest magnification weight calculated in the first target region of interest within the first gaze region is semantically complete. If this first-level semantic region is semantically complete, it is used as the second reference region for the first magnification process. If it is incomplete, the same method is used to traverse the N first-level semantic regions covered in the first gaze region from high to low magnification weight. If all N first-level semantic regions covered in the first gaze region are semantically incomplete, the second-level semantic region to which the first target region of interest belongs is used as the second reference region for the first magnification process. Then, based on the determined second reference region, the target image is magnified at a fixed magnification.
[0206] The following is combined Figures 11-12D and Figure 14 Explain how electronic devices can control image magnification by eye movements. Figure 14 A flowchart illustrating another image display processing method provided in this application embodiment.
[0207] S201: The electronic device displays the target image on the display screen.
[0208] S202: Obtain the first gaze region and the first gaze time of the target image being gazed at by the user.
[0209] S203: Determine whether the first fixation time is greater than or equal to the first threshold.
[0210] S204: If the first threshold is exceeded, then based on the first gaze region and the content of the first gaze region, determine the first target region of interest within the first gaze region.
[0211] S205: Perform a first magnification process on the target image based on the first target region of interest, and display it on the display screen.
[0212] For details of steps S201-S205 above, please refer to [link / reference]. Figure 13 The relevant descriptions of steps S101-S105 are not repeated here.
[0213] S206: Determine whether the resolution of the target image after the first magnification process is greater than or equal to the resolution of the display screen.
[0214] The electronic device acquires the current resolution of the display screen, which is determined by multiplying the width and height of the display screen in pixels, i.e., by multiplying the number of pixels in the horizontal direction by the number of pixels in the vertical direction. The electronic device also acquires the resolution of the first magnified target image, similarly determined by multiplying the width and height of the target image in pixels, i.e., by multiplying the number of pixels in the horizontal direction by the number of pixels in the vertical direction. The resolution of the first magnified target image is compared to whether it is greater than or equal to the current resolution of the display screen. If the resolution of the first magnified target image is greater than or equal to the current resolution of the display screen, step S208 is triggered to acquire the second gaze time and the second gaze area of the user gazing at the first magnified target image; if the resolution of the first magnified target image is lower than the current resolution of the display screen, step S207 is triggered to stop the eye-controlled image magnification process.
[0215] In one possible implementation, when determining whether the resolution of the target image after the first magnification process is greater than or equal to the current resolution of the electronic device display screen, the width pixel count of the target image resolution after the first magnification process is compared with the width pixel count of the current resolution of the electronic device display screen, and the height pixel count of the target image resolution after the first magnification process is compared with the height pixel count of the current resolution of the electronic device display screen. Only when the width pixel count of the target image resolution after the first magnification process is greater than or equal to the width pixel count of the current resolution of the electronic device display screen and the height pixel count of the target image resolution after the first magnification process is also greater than or equal to the height pixel count of the current resolution of the electronic device display screen is the resolution of the target image after the first magnification process considered to be greater than or equal to the current resolution of the electronic device display screen.
[0216] For example, if the resolution of the electronic device display screen is 2400×1600 pixels, and the resolution of the target image after the first magnification process is 3264×2448 pixels, the width and height pixel counts of the two are compared respectively. The width pixel count of the target image after the first magnification process is 3264, which is greater than the width pixel count of the electronic device (2400). At the same time, the height pixel count of the target image after the first magnification process is 2448, which is greater than the width pixel count of the electronic device (1600). This satisfies the condition that the width pixel count of the target image after the first magnification process is greater than or equal to the width pixel count of the current resolution of the electronic device display screen, and the height pixel count of the target image after the first magnification process is also greater than or equal to the height pixel count of the current resolution of the electronic device display screen. Then, step S208 is executed to obtain the second gaze time and the second gaze area of the user gazing at the target image after the first magnification process.
[0217] If the resolution of the target image after the first magnification process is 3200×1200 pixels, the width of the target image after the first magnification process is 3200 pixels, which is greater than the width of the electronic device (2400 pixels), and the height of the target image after the first magnification process is 1200 pixels, which is less than the width of the electronic device (1600 pixels), then the conditions are not met: the width of the target image after the first magnification process is greater than or equal to the width of the electronic device's current display screen resolution, and the height of the target image after the first magnification process is also greater than or equal to the height of the electronic device's current display screen resolution. Then, step S207 is executed to stop the eye-controlled image magnification process.
[0218] If the resolution of the target image after the first magnification process is 1200×1600 pixels, the width of the target image after the first magnification process is 1200 pixels, which is less than the width of the electronic device (2400 pixels), and the height of the target image after the first magnification process is 1600 pixels, which is equal to the width of the electronic device (1600 pixels), then the conditions are not met: the width of the target image after the first magnification process is greater than or equal to the width of the electronic device's display screen, and the height of the target image after the first magnification process is also greater than or equal to the height of the electronic device's display screen. Then, step S207 is executed to stop the eye-controlled image magnification process.
[0219] If the resolution of the target image after the first magnification process is 1600×1200 pixels, the width of the target image after the first magnification process is 1600 pixels, which is less than the width of the electronic device (2400 pixels), and the height of the target image after the first magnification process is 1200 pixels, which is less than the width of the electronic device (1600 pixels), then the conditions that the width of the target image after the first magnification process is greater than or equal to the width of the current resolution of the electronic device's display screen and the height of the target image after the first magnification process is greater than or equal to the height of the current resolution of the electronic device's display screen are not met, then step S207 is executed to stop the eye-controlled image magnification process.
[0220] S207: Stop eye-tracking image zoom-in process.
[0221] Specifically, when the target image meets preset conditions, the electronic device stops the eye-controlled image magnification process. The target image can be an image displayed on the screen after a first magnification process, or it can be a target image displayed on the screen after a second magnification process. The preset conditions can be a first preset condition, i.e., the resolution of the target image is greater than or equal to the current resolution of the display screen, or a second preset condition, i.e., the resolution of the target image is less than the current resolution of the display screen.
[0222] In one possible implementation, while the electronic device stops the currently executing eye-controlled image zooming process, it can display a prompt message on the electronic device's screen, informing the user that zooming cannot continue.
[0223] For example, such as Figure 11 As shown, the electronic device stops the eye-controlled image zooming process and simultaneously pops up a prompt window on top of the target image, displaying the text message "Cannot continue zooming". The prompt window remains displayed on the electronic device's screen for 1.5 seconds before being canceled.
[0224] In one possible implementation, the electronic device stops the currently executing eye-controlled image magnification process, but it can still continuously monitor the user's eyes using AO (Aggregate Eye Movement) technology via the front-facing camera. If the user consciously blinks, the electronic device detects this blink, exits the image magnification process, and restores the display interface of the target image before magnification. In some possible implementations, the user can also manually click on any location on the magnified target image on the display screen. The electronic device receives the user's click input, exits the image magnification process, and restores the display interface of the target image before magnification.
[0225] S208: Obtain the second gaze time and the second gaze region of the target image after the user gazes at it in the first magnified processing.
[0226] Specifically, the second fixation time can be the time during which a user continuously fixates on a specific area (by analyzing multiple consecutive frames of user eye images acquired by a camera to determine that the user's eye fixation position has not changed). Furthermore, the second fixation time can also include the time during which the user primarily fixates on a fixed area over a longer period, even if the area of fixation changes briefly during this time, the user's primary area of fixation remains relatively fixed within this period.
[0227] Electronic devices can first determine the second fixation area by capturing the first few frames of user eye images, and then determine the second fixation time, or they can first determine the second fixation time, and then determine the second fixation area based on the acquired user eye images.
[0228] In one possible implementation, the electronic device displays a target image that has undergone a first magnification process on the screen. The user experiences a need for magnification and consciously focuses on a specific area of the magnified target image. The electronic device first determines a second fixation time and then determines a second fixation area. The electronic device captures multiple high-resolution images of the user's eyes at short time intervals using a front-facing camera. The electronic device analyzes the acquired, short-interval-range eye images to determine second fixation information, which includes a second fixation time and a second fixation area.
[0229] For example, Figure 12A This is a schematic diagram of the second amplification process. Figure 1 ,like Figure 12A As shown, the electronic device displays a target image 1201 on the screen without the first magnification process. The electronic device has determined the first target region of interest 1202 and performed the first magnification process on the first target region of interest 1202 as a reference region, and then displayed it. Figure 12B On the display screen of the electronic device shown. Figure 12B This is a schematic diagram of the second amplification process. Figure 2 ,like Figure 12B As shown, the electronic device displays the target image 1211 after the first magnification process on the display screen. The electronic device continuously captures multiple frames of the user's eye images for 0.8 seconds using the front-facing camera 1212 at a high sampling rate of 100 frames per second, acquiring 80 frames of user eye images. The 80 frames of eye images are analyzed to obtain the area the user is looking at in each frame. Images with 50 or more consecutive frames corresponding to which the user's gaze area has not undergone significant displacement are detected, and the number of consecutive image frames is counted. The second gaze time is calculated using the largest consecutive image frame count. Specifically, the calculation method is as follows:
[0230] T2=(N2-1)·0.01
[0231] Where T2 is the second gaze time, N2 is the second gaze time of obtaining the target image after the user gazes at the first magnified image in step S208, and the maximum number of consecutive image frames detected in the second gaze region.
[0232] In one possible implementation, the electronic device determines the second gaze region of the target image on the display screen based on the second gaze time and the consecutive images corresponding to the largest number of consecutive image frames detected in the second gaze region obtained in step S208 after the user gazes at the first magnified target image. The second gaze region can be coarse-precision coordinates. Before determining the second gaze region, the electronic device divides the possible second gaze regions according to the aspect ratio of the image, and then determines the second gaze region based on the captured image of the user's eyes. The electronic device's division and determination of the possible second gaze regions of the target image after the first magnification process can be found in the description of the electronic device's division of the possible first gaze regions in step S202, which describes the acquisition of the first gaze region and first gaze time of the user gazing at the target image.
[0233] S209: Determine whether the second fixation time is greater than or equal to the second threshold.
[0234] Specifically, after acquiring the second gaze time, the electronic device compares the second gaze time with a preset second threshold. If the second gaze time is less than the second threshold, the electronic device continues to execute step S208 to acquire the second gaze time and the second gaze region of the user's gaze at the first magnified target image. If the second gaze time is greater than or equal to the second threshold, step S210 is triggered to determine the second target region of interest within the second gaze region. The second threshold is used to determine whether the user is consciously gazing at a certain area of the image displayed on the screen in order to achieve image magnification, and the second threshold can be less than the first threshold.
[0235] In one possible implementation, the electronic device can use an AI model to determine whether the user's second gaze time is greater than or equal to a second threshold. The electronic device uses the second gaze time as input to the AI model, analyzes it, and outputs the relationship between the second gaze time and the second threshold. The relationship includes two possibilities: greater than or equal to, and less than.
[0236] In one possible implementation, the second threshold can be set to 0.5 seconds. When the second gaze time is less than 0.5 seconds, the electronic device repeatedly executes step S208 to acquire the second gaze time and the second gaze region of the target image after the user gazes at the first magnified image; when the first gaze time is greater than or equal to 0.5 seconds, step S210 is triggered to determine the second target region of interest within the second gaze region.
[0237] S210: Determine the region of interest of the second target within the second gaze area.
[0238] First, the electronic device performs semantic segmentation on the target image displayed on the screen after the first magnification process. Specifically, the electronic device performs semantic segmentation on the target image displayed on the screen after the first magnification process according to a third semantic category, identifying P third-level semantic regions in the target image, where P is a positive integer greater than or equal to 1.
[0239] In one possible implementation, the third semantic category includes one or more of the following: eye, eyelash, eyelid, pupil, nose, nostril, lips, upper lip, lower lip, teeth, palm, thumb, index finger, middle finger, ring finger, little finger, arm, elbow, forearm, upper arm, blood vessel, leg, toe, clothing, clothing folds, button, plant, petal, stamen, calyx, flower stem, bark texture, and stone.
[0240] In one possible implementation, when determining P third-level semantic regions based on semantic segmentation according to a third semantic category, a single target image may include multiple objects corresponding to the third semantic category. When objects corresponding to the third semantic category overlap, among every two overlapping objects, the object with relatively complete semantics and located at a higher level in the image corresponding to the third semantic category is called the occluder, and the object whose portion is obscured by the occluder and located at a lower level in the image is called the occluded object. Specifically, when the occluder only partially covers the occluded object, causing part of the occluded object to be invisible, but not segmented into two or more independent regions that are not directly connected in space, simple image analysis techniques, such as color thresholding and edge detection, are first used to preliminarily identify the occluder in the image. Based on the detected occluder, an edge detection algorithm is further used to extract the precise edge of the occluder. The edge of the occluded object is determined based on the precise edge of the occluder and the unoccluded portion of the occluded object. The closed regions enclosed by the edges of the occluder and the occluded object are respectively determined as P third-level semantic regions.
[0241] For example, such as Figure 12BAs shown, the electronic device displays the target image 1211 after the first magnification process on the screen. The electronic device performs semantic segmentation on the target image 1211 after the first magnification process according to the third semantic category, and obtains 10 third-level semantic regions: 1213, 1214, 1215, 1216, 1217, 1218, 1219, 12110, 12111, and 12112. Among them, the third-level semantic region 1213 is the left eyebrow, the third-level semantic region 1214 is the left eye, the third-level semantic region 1215 is the right eyebrow, the third-level semantic region 1216 is the right eye, the third-level semantic region 1217 is the right ear, the third-level semantic region 1218 is the upper lip, the third-level semantic region 1219 is part 1 of the bow tie, the third-level semantic region 12110 is part 2 of the bow tie, the third-level semantic region 12111 is the nose, and the third-level semantic region 12112 is the lower lip. At this time, P is 10.
[0242] In one possible implementation, when the target image after the first magnification process has an overlap between the object corresponding to the third semantic category and other objects, the segmentation processing method of the third-level semantic region refers to step S204, which determines the method of semantic segmentation of the target image without the first magnification process by electronic device in the first target interest region within the first gaze region according to the first category and the second category.
[0243] It should be noted that, Figures 12A-12B The semantic segmentation diagram shown is an exemplary illustration of an embodiment of this application. Other semantic segmentation methods may be used when performing semantic segmentation on the target image after the first magnification process, and this embodiment of the application does not limit this method.
[0244] Next, from the P third-level semantic regions, determine the second gaze time of the target image after the user gazes at it in step S208, and the T third-level semantic regions that intersect with the second gaze region determined in the second gaze region. The T third-level semantic regions that intersect with the second gaze region may include third-level semantic regions segmented according to a third semantic category that are completely covered by the second gaze region, and third-level semantic regions segmented according to a third semantic category that are partially covered by the gaze region.
[0245] Based on the preset magnification priority of the T third-level semantic regions and the area occupied by the T third-level semantic regions in the second gaze region, the second target region of interest is determined.
[0246] The preset magnification priority is a set of priority rules for object magnification display that are pre-set in electronic devices based on factors such as product design concept and user experience. Specifically, it is involved in step S204, which describes the preset magnification priority in the first target region of interest within the first gaze area. This will not be elaborated here.
[0247] In one possible implementation, after determining T third-level semantic regions that intersect with the second gaze region, the electronic device traverses the T third-level semantic regions covered by the second gaze region, which are semantically segmented according to a third semantic category. Based on a preset magnification priority rule, the magnification priority of each third-level semantic region intersecting with the second gaze region is determined, and the pixels of each third-level semantic region are counted to obtain the pixel count of each third-level semantic region as the area corresponding to each third-level semantic region. When a third-level semantic region intersecting with the second gaze region is a third-level semantic region segmented according to a third semantic category and completely covered by the second gaze region, the total number of pixels in this third-level semantic region is counted as the area corresponding to this third-level semantic region. When a third-level semantic region intersecting with the second gaze region is a third-level semantic region segmented according to a third semantic category and partially covered by the gaze region, only the pixel count of the portion of the third-level semantic region covered by the second gaze region is counted as the area of this third-level semantic region; the pixels of the second gaze region are counted as the area of the third gaze region.
[0248] In one possible implementation, the electronic device calculates the magnification weight of each of the T third-level semantic regions based on the magnification priority and area of the T third-level semantic regions covered by the second gaze region, and selects the third-level semantic region with the largest magnification weight as the second target region of interest. If the third-level semantic region with the largest magnification weight is a third-level semantic region segmented according to the third semantic category and completely covered by the second gaze region, then this third-level semantic region is selected as the second target region of interest. If the third-level semantic region with the largest magnification weight is a third-level semantic region segmented according to the third semantic category and partially covered by the second gaze region, then the portion of this third-level semantic region covered by the second gaze region is selected as the second target region of interest.
[0249] For example, such as Figure 12BAs shown, the target image after the first magnification process has a resolution of 1560x720 pixels and can be divided into four possible second gaze regions. The second gaze region 12113, which is the upper right rectangular area of the target image 1211 after the first magnification process, is the second gaze region. The third-level semantic regions covered by the second gaze region include third-level semantic regions 12114, 12115, 12116, 12117, and 12118. Among them, third-level semantic regions 12117 (part of the left eyebrow) and 12118 (part of the left eye) are partially covered by the second gaze region. When counting pixels, only the part covered by the second gaze region 12113 is counted. The electronic device presets the magnification priority of the third-level semantic regions to an integer from 0 to 10, and each third-level semantic region corresponds to an integer representing the magnification priority. The magnification priority of the third-level semantic region (right eyebrow) 12114 is 6, with an area of 8000 pixels; the magnification priority of the third-level semantic region (right eye) 12115 is 7, with an area of 20000 pixels; the magnification priority of the third-level semantic region (right ear) 12116 is 6, with an area of 28000 pixels; the magnification priority of the third-level semantic region (part of the left eyebrow) 12117 is 6, with an area of 3000 pixels; the magnification priority of the third-level semantic region (part of the left eye) 12118 is 7, with an area of 9000 pixels; the area of the second gaze region is 280800 pixels.
[0250] Optionally, the formula for calculating the magnification weight W2 of the third-level semantic region is:
[0251] W2 = L n2 ·V l2 +A n2 ·V a2
[0252] Where L n2 V represents the amplification priority of the third-level semantic region after normalization. l2 The weight of the priority amplification for the third-level semantic region is a pre-set value for the electronic device, A. n2 V represents the area of the third-level semantic region after normalization. a2 The area weight is a preset value for electronic devices.
[0253] Optionally, L n2 The calculation formula is:
[0254]
[0255] Where L r2 L is the magnification priority determined after the electronic device traverses the third-level semantic region covered by the second gaze region. min2The maximum value of the amplification priority of the third-level semantic region preset for electronic devices, L max2 The maximum value of the amplification priority of the third-level semantic region preset for electronic devices.
[0256] Optionally, A n2 The calculation formula is:
[0257]
[0258] Where A o2 A is the area of each third-level semantic region covered by the second gaze region traversed by the electronic device. r2 This represents the area of the second gaze region.
[0259] For example, the weight V of the amplification priority l2 =0.6, area weight V a2 =0.4, A r2 280,800 pixels Figure 12B The L corresponding to the third-level semantic region (right eyebrow) 12114 shown is... r2 It is 6, L max2 For 10, L min2 If it is 1, then Corresponding A o2 If it is 8000 pixels, then The corresponding amplification weights W2 = W2 = L n2 ·V l2 +A n2 ·V a2 =0.555555·0.6 + 0.009971·0.4 = 0.373219. Similarly, Figure 12B The magnification weights of the other third-level semantic regions covered within the second fixation region 12113 shown are calculated using the same method. The magnification weights are: W2 = 0.428490 for the third-level semantic region (right eye) 12115, W2 = 0.373219 for the third-level semantic region (right ear) 12116, W2 = 0.337607 for the third-level semantic region (partial left eyebrow) 12117, and W2 = 0.412821 for the third-level semantic region (partial left eye) 12118. The third-level semantic region (right eye) 12115 with the largest magnification weight is selected as the second target region of interest.
[0260] S211: Perform a second magnification process on the target image based on the second target region of interest, and display it on the display screen.
[0261] In one possible implementation, if the semantics of the second target region of interest are complete, then the second target region of interest is used as the second reference region for the second magnification process, and the target image after the first magnification process is magnified at a fixed magnification based on the second reference region.
[0262] For example, such as Figure 12B As shown, the electronic device determines the second gaze region 12113 of the user's gaze. The electronic device has determined the third-level semantic region (right eye) 12115 as the second target region of interest, and the third-level semantic region (right eye) 12115 is semantically complete within the second gaze region 12113. The electronic device uses the third-level semantic region (right eye) 12115 as the reference region for the second magnification process, and performs a second magnification process on the target image displayed on the display screen after the first magnification process at a fixed magnification of 1.5 times, and displays it on the display screen, as shown. Figure 12C As shown, Figure 12C This is a schematic diagram of the second amplification process. Figure 3 At this time, the electronic device display shows the target image 1221 after the second magnification process. The target image 1221 is... Figure 12B The second target region of interest, namely the third-level semantic region (right eye), is taken as the center of the image.
[0263] The definition of the reference area can be found in the description of the reference area in the application scenarios involved in the embodiments of this application in the specification, and will not be repeated here.
[0264] In one possible implementation, if the semantics of the second target region of interest are incomplete and the second target region of interest belongs to a part of a third-level semantic region, then the third-level semantic region to which the second target region of interest belongs is used as the second reference region for the second amplification process.
[0265] For example, such as Figure 12B As shown, if the electronic device calculates and determines that the third-level semantic region (part of the left eyebrow) 12117 is the second target region of interest, where the second target region of interest is an irregular region that is cut out by the edge of the third-level semantic region (left eyebrow) 1213 and covered by the second gaze region 12113, and the third-level semantic region (part of the left eyebrow) 12117 is semantically incomplete, and the second target region of interest belongs to a part of the third-level semantic region (left eyebrow) 1213, then the electronic device determines that the third-level semantic region (left eyebrow) 1213 is the reference region for the second magnification processing, and performs a second magnification processing on the target image 1211 after the first magnification processing at a fixed magnification of 1.5 times, and displays it on the display screen, as shown. Figure 12D As shown, Figure 12DThis is a schematic diagram of the second magnification process, 4. At this time, the electronic device's display screen shows the target image 1231 after the second magnification process. The target image 1231 after the second magnification process... Figure 12B The third-level semantic region (left eye) 1214 shown is used as the center of the image.
[0266] In one possible implementation, if the second target region of interest is semantically incomplete and the second target region of interest belongs to the entire region of a certain third-level semantic region, the electronic device can detect whether the third-level semantic region with the largest magnification weight calculated in the second target region of interest within the second gaze region, excluding the second target region of interest, is semantically complete, as determined in step S210. If this third-level semantic region is semantically complete, it is used as the second reference region for the second magnification process. If this third-level semantic region is semantically incomplete, the electronic device can detect whether the third-level semantic region with the second largest magnification weight calculated in the second target region of interest within the second gaze region, excluding the second target region of interest, is semantically complete, as determined in step S210. If this third-level semantic region is semantically complete, it is used as the second reference region for the second magnification process. If it is incomplete, the same method is used to traverse the T third-level semantic regions covered within the second gaze region according to the magnification weight from high to low. If all T third-level semantic regions covered within the second gaze region are semantically incomplete, the second magnification process is performed with the center of the target image after the first magnification process as the reference region.
[0267] S212: Determine whether the resolution of the target image after the second magnification process is lower than the resolution of the display screen.
[0268] In step S206, the electronic device determines whether the resolution of the first magnified target image is greater than or equal to the resolution of the display screen. The electronic device has already obtained the current resolution of the electronic device's display screen. The electronic device then obtains the resolution of the second magnified target image. The resolution of the second magnified target image is determined by multiplying the width pixel count and the height pixel count, just like the current resolution of the electronic device's display screen. That is, it is represented by multiplying the number of pixels in the horizontal direction of this target image by the number of pixels in the vertical direction. The electronic device then compares whether the resolution of the second magnified target image is lower than the current resolution of the display screen. If the resolution of the second magnified target image is lower than the current resolution of the display screen, then step S207 is executed to stop the eye-controlled image magnification process.
[0269] In one possible implementation, when determining whether the resolution of the target image after the second magnification process is lower than the current resolution of the electronic device display screen, the number of pixels in the width of the target image after the second magnification process is compared with the number of pixels in the width of the current resolution of the electronic device display screen, and the number of pixels in the height of the target image after the second magnification process is compared with the number of pixels in the height of the current resolution of the electronic device display screen. When the number of pixels in the width of the target image after the second magnification process is lower than the number of pixels in the width of the current resolution of the electronic device display screen, or when the number of pixels in the height of the target image after the second magnification process is lower than the number of pixels in the height of the current resolution of the electronic device display screen, it is considered that the resolution of the target image after the second magnification process is lower than the current resolution of the electronic device display screen.
[0270] For example, if the resolution of the electronic device display screen is 2400×1600 pixels, and the resolution of the target image after the second magnification process is 1600×1200 pixels, the width and height pixel counts of the two are compared respectively. The width pixel count of the target image after the second magnification process is 1600, which is less than the width pixel count of the electronic device (2400). At the same time, the height pixel count of the target image after the second magnification process is 1200, which is less than the width pixel count of the electronic device (1200). This satisfies the condition that the width pixel count of the target image resolution after the second magnification process is less than the width pixel count of the current resolution of the electronic device display screen, or the height pixel count of the target image resolution after the second magnification process is less than the height pixel count of the current resolution of the electronic device display screen. Then, step S207 is executed to stop the eye-controlled image magnification process.
[0271] If the resolution of the target image after the second magnification process is 3200×1200 pixels, the width of the target image after the second magnification process is 3200 pixels, which is greater than the width of the electronic device (2400 pixels), and the height of the target image after the second magnification process is 1200 pixels, which is less than the width of the electronic device (1600 pixels), then the conditions are met: the width of the target image after the second magnification process is less than the width of the electronic device's current display screen resolution, or the height of the target image after the second magnification process is less than the height of the electronic device's current display screen resolution. Then, step S207 is executed to stop the eye-controlled image magnification process.
[0272] If the resolution of the target image after the second magnification process is 1200×1600 pixels, the width of the target image after the second magnification process is 1200 pixels, which is less than the width of the electronic device (2400 pixels), and the height of the target image after the second magnification process is 1600 pixels, which is equal to the width of the electronic device (1600 pixels), then the condition that the width of the target image after the second magnification process is less than the width of the electronic device's display screen or the height of the target image after the second magnification process is less than the height of the electronic device's display screen is met, then step S207 is executed to stop the eye-controlled image magnification process.
[0273] It should be understood that each step in the above method embodiments can be completed by integrated logic circuits in the processor hardware or by instructions in software form. The method steps disclosed in the embodiments of this application can be directly manifested as being executed by a hardware processor, or being executed by a combination of hardware and software modules in the processor.
[0274] This application also provides an electronic device, which may include a memory and a processor. The memory may be used to store a computer program; the processor may be used to invoke the computer program in the memory, causing the electronic device to execute the method executed by the electronic device in any of the above embodiments.
[0275] This application also provides a chip system including at least one processor for implementing the functions involved in the electronic device in any of the above embodiments.
[0276] In one possible design, the chip system also includes a memory for storing program instructions and data, which may be located within or outside the processor.
[0277] The chip system can consist of chips or include chips and other discrete components.
[0278] Optionally, the chip system may contain one or more processors. These processors can be implemented in hardware or software. When implemented in hardware, the processor can be a logic circuit, an integrated circuit, etc. When implemented in software, the processor can be a general-purpose processor, implemented by reading software code stored in memory.
[0279] Optionally, the chip system may contain one or more memories. The memory may be integrated with the processor or disposed separately from it; this application embodiment does not limit this. For example, the memory may be a non-transient processor, such as a read-only memory (ROM), which may be integrated with the processor on the same chip or disposed separately on different chips. This application embodiment does not specifically limit the type of memory or the arrangement of the memory and processor.
[0280] For example, the chip system may be a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), a system-on-chip (SoC), a central processing unit (CPU), a network processor (NP), a digital signal processing circuit (DSP), a microcontroller unit (MCU), a programmable logic device (PLD), or other integrated chips.
[0281] This application also provides a computer program product comprising: a computer program (also referred to as code or instructions) that, when run, causes a computer to perform the method executed by the electronic device in any of the above embodiments.
[0282] This application also provides a computer-readable storage medium storing a computer program (also referred to as code or instructions). When the computer program is run, it causes the computer to perform the method executed by the electronic device in any of the above embodiments.
[0283] The various embodiments of this application can be combined arbitrarily to achieve different technical effects.
[0284] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state disk (SSD)).
[0285] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This program can be stored in a computer-readable storage medium, and when executed, it can include the processes described in the above method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM or random access memory (RAM), magnetic disks, or optical disks.
[0286] In summary, the above description is merely an embodiment of the technical solution of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, improvements, etc., made based on the disclosure of this application should be included within the scope of protection of this application.
Claims
1. A method of image display processing, characterized by, The method is applied to a terminal device, and the terminal device comprises a display screen, and the method comprises the following steps: displaying a target image on the display screen, obtaining first gaze information of a user gazing at the target image, the first gaze information comprising a first gaze area in which the user gazes at the target image and a first gaze time at which the user gazes at the first gaze area; determining whether the first gaze time exceeds a first threshold value; if the first gaze time exceeds the first threshold value, determining a first target region of interest in the first gaze area based on the first gaze area and content of the first gaze area; performing first zoom-in processing on the target image based on the first target region of interest, and displaying the target image after the first zoom-in processing on the display screen.
2. The method of claim 1, wherein, The determination of the first target region of interest in the first gaze area comprises the following steps: performing semantic segmentation on the target image to identify M first-level semantic regions in the target image, M being an integer greater than 1; determining N first-level semantic regions having an intersection with the first gaze area from the M first-level semantic regions; determining a first target region of interest based on a preset zoom-in priority of the N first-level semantic regions and an area occupied by the N first-level semantic regions in the first gaze area.
3. The method of claim 2, wherein, The semantic segmentation of the target image to identify M first-level semantic regions in the target image comprises the following steps: performing semantic segmentation on the target image according to a first semantic category to identify Q second-level semantic regions in the target image, Q being a positive integer less than M; performing semantic segmentation on each of the Q second-level semantic regions according to a second semantic category to identify one or more first-level semantic regions in each of the second-level semantic regions; determining the first-level semantic regions corresponding to the Q second-level semantic regions as the M first-level semantic regions.
4. The method of claim 3, wherein, The first semantic category comprises one or more of a person, an animal, a plant and a scene, and the second semantic category comprises one or more of human facial features, hair, arms, legs, hands, feet, clothes, tree crowns, tree trunks, flower crowns, stems of plants, animal facial features and limbs.
5. The method according to any one of claims 1 to 4, wherein The first target region of interest is part or the entirety of a first-level semantic region having an intersection with the first gaze area.
6. The method according to any one of claims 1 to 5, wherein, The first zoom-in processing on the target image based on the first target region of interest comprises the following steps: if the first target region of interest is semantically complete, taking the first target region of interest as a first reference area of a first zoom-in operation, and performing first zoom-in processing on the target image based on the first reference area; if the first target region of interest is not semantically complete, determining a second reference area based on the first target region of interest, the first-level semantic regions and the second-level semantic regions, and performing first zoom-in processing on the target image based on the second reference area.
7. The method according to any one of claims 1 to 6, wherein After displaying the target image after the first zoom-in processing on the display screen, the method further comprises the following steps: When the target image after the first zoom-in processing meets a first preset condition, second gaze information of a user gazing at the target image after the first zoom-in processing is acquired, the second gaze information including a second gaze area in which the user gazes at the target image after the first zoom-in processing and a second gaze time at which the user gazes at the second gaze area; It is judged whether the second gaze time exceeds a second threshold value; If the second threshold value is exceeded, a second target region of interest in the second gaze area is determined based on the second gaze area and content of the second gaze area; The target image is subjected to second zoom-in processing based on the second target region of interest, and the target image after the second zoom-in processing is displayed on the display screen.
8. The method of claim 7, wherein, Further comprising: The current resolution of the display screen and the resolution of the target image after the first zoom-in processing are acquired; If the resolution of the target image after the first zoom-in processing is greater than or equal to the current resolution of the display screen, the target image after the first zoom-in processing meets the first preset condition; If the resolution of the target image after the first zoom-in processing is lower than the current resolution of the display screen, the target image after the first zoom-in processing does not meet the first preset condition.
9. The method of claim 7, wherein, After the target image after the second zoom-in processing is displayed on the display screen, further comprising: When the target image after the second zoom-in processing meets a second preset condition, picture zoom-in is stopped, the second preset condition including: The current resolution of the display screen and the resolution of the target image after the second zoom-in processing are acquired; If the resolution of the target image after the second zoom-in processing is less than the current resolution of the display screen, the target image after the second zoom-in processing meets the second preset condition; If the resolution of the target image after the second zoom-in processing is greater than or equal to the current resolution of the display screen, the target image after the second zoom-in processing does not meet the second preset condition.
10. The method of any one of claims 1-9, wherein, The terminal device further comprises a camera; the acquisition of the first gaze information of the user gazing at the target image comprises: The camera is used to capture one or more eye images of the user gazing at the target image, and the first gaze information is determined based on the one or more eye images of the user.
11. The method of claim 10, wherein, Further comprising: The one or more eye images are used as input of an artificial intelligence (AI) model, the AI model is used to analyze the one or more eye images, and the first gaze information is output.
12. The method of any one of claims 1-11, wherein, The method further comprises detecting whether the user gazes at the display screen through always online (AO) technology, and acquiring the eye image of the user through the camera.
13. An intelligent terminal device, characterized by Comprise: A memory and one or more processors; the memory is coupled with the one or more processors, the memory is used to store computer program code, the computer program code includes computer instructions, and the one or more processors invoke the computer instructions to enable the electronic device to perform the method in any one of claims 1-12.
14. A computer storage medium, characterized in that The computer storage medium stores a computer program, which, when executed by a processor, implements the method of any one of claims 1-12.
15. A computer program product, characterised in that, The computer program product comprises instructions, which, when the computer program product is executed by a computer, cause the computer to perform the method of any one of claims 1-12.