Image processing method and device, electronic equipment, storage medium and program product

By combining multimodal information from eye and touch data, the image resolution is dynamically adjusted for rendering, solving the problems of high load and uneven image quality in electronic device image display, and achieving efficient visual experience and low-load image display.

CN121120891AInactive Publication Date: 2025-12-12GUANG DONG MING CHUANG SOFTWARE TECH CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511138063.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-13
Publication Date
2025-12-12
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing technologies for improving the image display effect of electronic devices suffer from high load and uneven image quality, especially when the user's visual focus changes, resulting in wasted hardware load and a decline in visual experience.

Method used

By fusing multimodal information from eye and touch data, the system dynamically partitions images in real time for rendering. Different resolutions are used to render images in both attentive and non-attentive areas, and then the images are stitched together for display, ensuring a good visual experience while reducing device load.

Benefits of technology

This approach achieves the goal of reducing the load on electronic devices while ensuring a good visual experience, avoiding image quality fragmentation and hardware resource waste, and improving the smoothness of the user experience and device efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121120891A_ABST
    Figure CN121120891A_ABST
Patent Text Reader

Abstract

The invention discloses an image processing method and device, electronic equipment, a storage medium and a program product. Obtaining multi-modal data including eyeball data of the user and touch data of the user on the touch screen, determining an attention area of the user on the touch screen according to the eyeball data and the touch data, rendering an image corresponding to the attention area according to a first resolution, and obtaining a rendered image; the first resolution ratio is larger than the first resolution ratio, the image corresponding to a non-attention area except the attention area on the touch screen is rendered according to the second resolution ratio, the first resolution ratio is larger than the second resolution ratio, the rendered image corresponding to the attention area and the rendered image corresponding to the non-attention area are spliced, and the image corresponding to the non-attention area is obtained. And the spliced image is displayed on the touch screen. According to the method and the device, the attention area of the user is determined by fusing the multi-modal data including the eyeball data and the touch data, and image rendering is performed in real-time dynamic partitioning with different resolutions, so that the load of the electronic equipment can be reduced while the visual experience can be ensured.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of image processing, and more particularly, to an image processing method and device, an electronic device, a storage medium, and a program product. BACKGROUND

[0002] With the development of science and technology, electronic devices are increasingly widely used and have more and more functions, and have become one of the necessities in people's daily life. Moreover, people have increasingly high requirements for display effects when using electronic devices to display images. In order to meet the needs of users, manufacturers begin to pursue higher display effects, for example, using high resolution to display images. However, this way will cause a high load of the electronic device. SUMMARY

[0003] In view of the above problems, the present application provides an image processing method and device, an electronic device, a storage medium, and a program product to solve the above problems.

[0004] In a first aspect, an image processing method is provided, applied to an electronic device, the electronic device comprising a touch screen, the method comprising: acquiring multi-modal data, wherein the multi-modal data comprises eye data of a user and touch data of the user on the touch screen; determining an attention area of the user on the touch screen according to the eye data and the touch data; rendering an image corresponding to the attention area according to a first resolution, and rendering an image corresponding to a non-attention area on the touch screen except the attention area according to a second resolution, wherein the first resolution is greater than the second resolution; splicing the rendered image corresponding to the attention area and the rendered image corresponding to the non-attention area, and displaying the spliced image on the touch screen.

[0005] In a second aspect, an image processing device is provided, applied to an electronic device, the electronic device comprising a touch screen, the device comprising: a multi-modal data acquisition module, configured to acquire multi-modal data, wherein the multi-modal data comprises eye data of a user and touch data of the user on the touch screen; an attention area determination module, configured to determine an attention area of the user on the touch screen according to the eye data and the touch data; an image rendering module, configured to render an image corresponding to the attention area according to a first resolution, and render an image corresponding to a non-attention area on the touch screen except the attention area according to a second resolution, wherein the first resolution is greater than the second resolution; and an image display module, configured to splice the rendered image corresponding to the attention area and the rendered image corresponding to the non-attention area, and display the spliced image on the touch screen.

[0006] Thirdly, embodiments of this application provide an electronic device, including a memory and a processor, wherein the memory is coupled to the processor, the memory stores instructions, and when the instructions are executed by the processor, the processor performs the above-described method.

[0007] Fourthly, embodiments of this application provide a computer-readable storage medium storing program code, which can be invoked by a processor to execute the above-described method.

[0008] Fifthly, embodiments of this application provide a computer program product, which includes a computer program that, when executed by a processor, implements the above-described method.

[0009] The image processing method, apparatus, electronic device, storage medium, and program product provided in this application embodiment acquire multimodal data, including user eye data and user touch data on a touchscreen. Based on the eye data and touch data, the user's attention area on the touchscreen is determined. The image corresponding to the attention area is rendered according to a first resolution, and the image corresponding to the non-attention area on the touchscreen (excluding the attention area) is rendered according to a second resolution, wherein the first resolution is greater than the second resolution. The rendered image corresponding to the attention area and the rendered image corresponding to the non-attention area are stitched together, and the stitched image is displayed on the touchscreen. Thus, by fusing multimodal data including eye data and touch data to determine the user's attention area and dynamically partitioning and rendering images at different resolutions in real time, the visual experience can be guaranteed while reducing the load on the electronic device. Attached Figure Description

[0010] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0011] Figure 1 A schematic flowchart of an image processing method provided in an embodiment of this application is shown; Figure 2 A schematic flowchart of an image processing method provided in an embodiment of this application is shown; Figure 3 This application shows Figure 2 The flowchart of step S240 of the image processing method shown is illustrated. Figure 4A schematic flowchart of an image processing method provided in an embodiment of this application is shown; Figure 5 A schematic flowchart of an image processing method provided in an embodiment of this application is shown; Figure 6 A schematic flowchart of an image processing method provided in an embodiment of this application is shown; Figure 7 A block diagram of an image processing apparatus provided in one embodiment of this application is shown; Figure 8 A block diagram of an electronic device for performing an image processing method according to an embodiment of this application is shown; Figure 9 An embodiment of the present application shows a storage unit for storing or carrying program code that implements the image processing method according to the embodiment of the present application. Detailed Implementation

[0012] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.

[0013] To meet user needs, manufacturers typically employ the following methods to improve display quality: First, a uniform, fixed resolution is used for full-screen rendering, such as locking the rendering resolution to 720P in a game scene. This approach reduces the load on the graphics processing unit (GPU) through downsampling. Specifically, during rendering, the image is first generated at a resolution lower than the screen's native resolution (e.g., 720P), and then downsampling is used to adapt it for screen display. This reduces the computational load on the GPU for pixel processing, thus lowering the hardware load. The advantage of this mode is its simplicity, requiring no complex partitioning or dynamic adjustment logic. However, it also has significant limitations: because the full-screen resolution is fixed, it cannot optimize image quality allocation based on the user's visual focus, potentially leading to performance waste in unnecessary areas or a degraded experience in areas requiring high image quality due to the fixed low resolution.

[0014] Secondly, static tile rendering divides the screen into fixed grid partitions (e.g., a 4x4 grid) and applies differentiated resolution strategies to different partitions. Typically, the central area of ​​the screen is set to high resolution for rendering, while the edges or other areas use lower resolution. This approach is based on the assumption that the user's visual focus is more likely to be concentrated in the center of the screen, balancing visual experience and hardware load by prioritizing the image quality of the central area. However, its static nature (fixed tile division and no dynamic adjustment of the partitioning strategy based on user behavior) leads to significant drawbacks: when the user's visual focus moves out of the preset high-resolution central area, the image quality drops sharply due to entering a low-resolution partition, causing a fragmented experience (user surveys show increased discomfort); simultaneously, the static high-resolution area continuously renders at full power, resulting in wasted power; and the tile strategy updates infrequently, failing to track rapid eye movements, with a focus loss rate of 32% in test videos.

[0015] To address the aforementioned problems, the inventors, through long-term research, discovered and proposed the image processing method, apparatus, electronic device, storage medium, and program product provided in the embodiments of this application. By fusing multimodal data, including eye data and touch data, to determine the user's attention region, and dynamically partitioning the image in real time for rendering at different resolutions, the visual experience can be guaranteed while reducing the load on the electronic device. The specific image processing method will be described in detail in the subsequent embodiments.

[0016] Please see Figure 1 , Figure 1 A schematic flowchart of an image processing method according to an embodiment of this application is shown. This method is used to determine the user's attention region by fusing multimodal data including eye data and touch data, and to dynamically partition the image in real time for rendering at different resolutions, thereby ensuring a good visual experience while reducing the load on the electronic device. In a specific embodiment, this image processing method is applied to, for example... Figure 7 The image processing device 200 and the electronic device 100 equipped with the image processing device 200 are shown. Figure 8 The following will use an electronic device as an example to illustrate the specific process of this embodiment. Of course, it is understood that the electronic device may include smartphones, tablets, wearable devices, in-vehicle systems, etc., and is not limited thereto. In this embodiment, the electronic device includes a touchscreen. The following will focus on... Figure 1 The process shown will be described in detail. The image processing method may specifically include the following steps: Step S110: Acquire multimodal data, wherein the multimodal data includes the user's eye data and the user's touch data on the touchscreen.

[0017] In this embodiment, multimodal data can be acquired.

[0018] In some implementations, electronic devices may be equipped with multiple sensors, each of which is used to collect one or more modal data. Accordingly, electronic devices can obtain the modal data collected by the multiple sensors, thereby obtaining multimodal data.

[0019] Optionally, the multimodal data may include the user's eye data. This eye data primarily includes eye coordinates and eye movement signals, directly recording information such as the user's eye movement trajectory, fixation point position, and movement speed, providing the most direct evidence of the user's current visual focus. As one approach, electronic devices can have built-in eye-tracking sensors, which can then collect the user's eye data.

[0020] Optionally, the multimodal data may include user touch data on the touchscreen. This touch data primarily includes touch trajectory sequences, touch pressure, and other information, recording changes in the user's touch position and pressure on the screen. These behaviors are often related to the user's visual attention distribution (for example, the area touched by the user is usually their focus). As one approach, the electronic device may have a built-in pressure sensor, which can then collect the user's touch data on the touchscreen.

[0021] Optionally, the multimodal data may include image / depth map data. Image data refers to two-dimensional visual images, containing information such as the shape, color, and positional distribution of objects in the scene (e.g., characters and props in a game scene, text areas in a document). Depth map data reflects the three-dimensional spatial relationships of objects in the scene, showing the distance and hierarchy of different objects (e.g., determining which object is in front of the user's line of sight and is more likely to become the focus of attention). As one approach, electronic devices may have built-in image acquisition devices (such as cameras), which can then acquire image / depth map data.

[0022] Optionally, the multimodal data may include sound data. Sound data primarily includes various sound signals from the environment, such as voice commands from user operations, sound effects in game scenes (e.g., character movement sounds, skill release sounds), and ambient sounds. This sound information is often related to the scene elements the user is currently focusing on (e.g., the direction of the sound source may point to a potential area of ​​the user's visual focus). As one approach, electronic devices can have built-in sound acquisition devices (such as microphones), and correspondingly, sound data can be acquired through these devices.

[0023] Optionally, the multimodal data may include motion data. Motion data primarily includes information such as the device's trajectory, tilt angle, and rotation speed, reflecting changes in the device's position and attitude adjustments in space. These motion behaviors are often related to shifts in the user's visual attention (e.g., when rotating the device, the user's gaze may focus on a new area as the device's attitude changes). As one approach, electronic devices can have built-in motion sensors (such as accelerometers and gyroscopes), and motion data can be collected through these sensors.

[0024] Step S120: Determine the user's attention area on the touchscreen based on the eye data and the touch data.

[0025] In this embodiment, when multimodal data is acquired, the user's attention area on the touchscreen can be determined based at least on the eye data and touch data in the multimodal data.

[0026] In some implementations, where the multimodal data also includes image data, depth map data, sound data, and motion data, one or more of these data can be incorporated into the determination of the attention region. For example, sound data can be included in the determination of the attention region; in this case, the user's attention region on the touchscreen is determined based on eye data, touch data, and sound data. Alternatively, sound data and motion data can be included in the determination of the attention region; in this case, the user's attention region on the touchscreen is determined based on eye data, touch data, sound data, and motion data.

[0027] In some implementations, electronic devices may have a pre-set method for determining attention regions. Accordingly, when multimodal data is acquired, eye data and touch data can be processed using this method to determine the user's attention region on the touchscreen.

[0028] In some implementations, electronic devices can be pre-trained with models of attention regions. Accordingly, when multimodal data is acquired, eye data and touch data can be processed by the model to determine the user's attention regions on the touchscreen.

[0029] Step S130: Render the image corresponding to the attention region according to the first resolution, and render the image corresponding to the non-attention region on the touch screen other than the attention region according to the second resolution, wherein the first resolution is greater than the second resolution.

[0030] In some implementations, the electronic device may pre-set and store multiple different resolutions, which may include a first resolution and a second resolution, and may also include a third resolution, a fourth resolution, etc., without limitation. Optionally, the first resolution is greater than the second resolution; for example, the first resolution is 1080P and the second resolution is 720P.

[0031] In some implementations, when an electronic device determines the area of ​​user attention on the touchscreen, it can define other display areas on the touchscreen outside the area of ​​user attention as non-attention areas.

[0032] In this embodiment, when the touchscreen is divided into attention areas and non-attention areas, the image corresponding to the attention area can be rendered according to a preset first resolution (higher resolution), and the image corresponding to the non-attention area can be rendered according to a preset second resolution (lower resolution). This dynamically displays a high-resolution image of the user's attention area and a low-resolution image of the non-attention area, effectively balancing the user's visual experience and the device's power consumption.

[0033] As one feasible approach, the rendering execution layer integrates the resolution parameters of the aforementioned partitions (attention areas and first resolution, non-attention areas and second resolution) and submits them to the GPU's Command Buffer. The GPU renders in parallel according to the instructions, dividing the regions: high-resolution areas are processed according to the first resolution, and low-resolution areas are processed according to the second resolution. The final rendering result is then composited and output to the screen, achieving the dynamic effect of focusing attention areas with high resolution and optimizing non-focus areas with low resolution.

[0034] In some implementations, rendering the image corresponding to the attention region according to a first resolution and rendering the image corresponding to the non-attention region on the touchscreen according to a second resolution may include: using graph multisampling anti-aliasing technology (such as MSAA 4x) to render the image corresponding to the attention region according to the first resolution and using bilinear sampling technology to render the image corresponding to the non-attention region according to the second resolution.

[0035] Among them, multi-sampling anti-aliasing technology can effectively eliminate the jagged edges of objects in high-resolution areas (attention areas) (such as game character outlines and text edges) by performing multiple sampling calculations on the sub-sampling points around each pixel, thereby improving the image detail and ensuring that the visual experience at the user's visual focus is not compromised.

[0036] Dual-thread sampling technology generates target pixels by linearly interpolating the color values ​​of adjacent pixels. While the calculation process is simple and efficient, its image quality accuracy is lower than MSAA. Applying it to low-resolution areas allows for rapid rendering while maintaining basic display quality, further reducing GPU resource consumption.

[0037] In some implementations, rendering the image corresponding to the attention region according to a first resolution and rendering the image corresponding to the non-attention region according to a second resolution may include: determining a non-attention region adjacent to the attention region as a transition region from the non-attention region; rendering the image corresponding to the attention region according to the first resolution; rendering the image corresponding to the non-attention region other than the transition region according to the second resolution; and rendering the image corresponding to the transition region according to the first resolution and the second resolution.

[0038] As one feasible approach, when identifying attention-grabbing and non-attention-grabbing areas from the display screen of an electronic device, a transition area can be defined as an adjacent non-attention-grabbing area within the non-attention-grabbing area. Alternatively, an 8-pixel-wide strip region adjacent to the attention-grabbing area can be selected from the non-attention-grabbing area, and this selected strip region can be defined as the transition area.

[0039] As a feasible approach, for the attention region, a first resolution (e.g., 1080P) can be assigned, and MSAA 4x anti-aliasing technology can be enabled to ensure that the image quality in this region is clear and detailed, matching the high experience requirements of the user's visual focus. Then, the image corresponding to the attention region is rendered according to the first resolution and anti-aliasing technology. For the non-attention region, excluding the transition region, a second resolution (e.g., 720P) can be assigned, and Bilinear sampling technology can be used to reduce GPU load and optimize power consumption by reducing resolution and simplifying calculations. Then, the image corresponding to the non-attention region, excluding the transition region, is rendered according to the second resolution and Bilinear sampling technology. For the transition region, a dynamic resolution gradient strategy can be adopted, which integrates the parameters of the first and second resolutions. From the side closer to the attention region to the side farther away, the resolution smoothly transitions from the first resolution (1080P) to the second resolution (720P). At the same time, the processing effects of MSAA anti-aliasing and Bilinear sampling are mixed to avoid abrupt changes in image quality at the boundary. Then, the resolution gradient fusion is completed through the transition band blending module to output a smoothly transitioned image.

[0040] In other implementations, given the defined attention and non-attention areas, the entire screen can be rendered at a low resolution (e.g., second resolution rendering) first, and then super-resolved to the target high resolution (e.g., first resolution) in real time by an NPU (Neural Processing Unit). At the same time, the original high-resolution texture is retained only in the user's attention area, so as to further optimize the GPU load while balancing image quality and power consumption.

[0041] Step S140: The image corresponding to the rendered attention region is stitched together with the image corresponding to the rendered non-attention region, and the stitched image is displayed on the touch screen.

[0042] In this embodiment, after rendering the image corresponding to the attention region, a rendered image corresponding to the attention region can be obtained, and after rendering the image corresponding to the non-attention region, a rendered image corresponding to the non-attention region can be obtained. Then, the rendered image corresponding to the attention region and the rendered image corresponding to the non-attention region can be stitched together to obtain a stitched image, and the stitched image can be displayed on the touch screen.

[0043] This allows for dynamic high-resolution display of areas of interest to the user, while dynamically displaying low-resolution areas of inattention, thus balancing visual appeal and device power consumption. In essence, the stitched image is submitted to the screen display buffer and ultimately presented as a unified display on the touchscreen. At this point, the user sees high-quality images in the areas of interest, while non-interesting areas are rendered in a low-load mode, with smooth transitions between areas and no visual disjointedness due to resolution differences. This retains the performance and power efficiency advantages of differentiated rendering while ensuring a seamless overall visual experience for the user, resolving the image fragmentation problem caused by fixed-block rendering in existing technologies.

[0044] In some implementations, the rendering results of the attention and non-attention areas can be precisely aligned according to their respective spatial locations (such as coordinate ranges) based on the pixel coordinate system of the touchscreen, ensuring that the images of each area are physically aligned (for example, the edge pixels of the attention area are continuous with the adjacent pixels of the non-attention area). Furthermore, during the stitching process, the color parameters (such as brightness and contrast) of each area image can be automatically calibrated to ensure consistent color representation across areas of different resolutions, reducing visual disjointedness.

[0045] In some implementations, for the boundary between the attention region and the non-attention region (i.e. the transition region), a pixel-level fusion algorithm can be executed by the transition band blending module. For each pixel in the transition region, the weight of the high / low resolution image is dynamically adjusted according to its distance from the attention region (pixels closer to the attention region are dominated by the first resolution image, and those farther away are dominated by the second resolution image), so as to achieve a smooth transition of image quality and avoid obvious boundary lines.

[0046] In some implementations, after the stitched image is displayed on the touchscreen, the device status of the electronic device can be obtained, and the first resolution and the second resolution can be adjusted based on the device status of the electronic device to adjust the display of the images corresponding to the attention area and the non-attention area to adapt to different device statuses.

[0047] Optionally, the device status of the electronic device may include its battery status. This battery status can be obtained by a battery status monitoring module, including the remaining battery percentage and charging status (e.g., charging / not charging). For example, when the battery is low, the second resolution of non-attention areas may be further reduced to decrease power consumption; when the battery is sufficient, the first resolution of attention areas may be appropriately increased to optimize image quality.

[0048] Optionally, the device status of an electronic device may include GPU load and temperature. This includes data such as real-time GPU load rate and peak temperature obtained through system monitoring. If the GPU load or temperature is too high, the primary resolution may be reduced or the low-resolution area expanded to reduce load and temperature; if the load is low, the resolution may be maintained or increased to ensure a good user experience.

[0049] Optionally, the device status of an electronic device can include the device's operating scenario. The operating scenario can include the type of application currently running (games, reading, videos, etc.). For example, a game scenario requires high smoothness and may more strictly control the resolution of non-focus areas while ensuring image quality in the focus area; a reading scenario may appropriately expand the high-resolution area to improve text clarity.

[0050] An embodiment of this application provides an image processing method that acquires multimodal data, including user eye data and user touch data on a touchscreen. Based on the eye data and touch data, the user's attention area on the touchscreen is determined. The image corresponding to the attention area is rendered according to a first resolution, and the image corresponding to the non-attention area on the touchscreen (excluding the attention area) is rendered according to a second resolution, wherein the first resolution is greater than the second resolution. The rendered image corresponding to the attention area is stitched together with the rendered image corresponding to the non-attention area, and the stitched image is displayed on the touchscreen. By fusing multimodal data including eye data and touch data to determine the user's attention area and dynamically partitioning the image in real time for rendering at different resolutions, the visual experience can be guaranteed while reducing the load on the electronic device.

[0051] Please see Figure 2 , Figure 2 A schematic flowchart of an image processing method according to an embodiment of this application is shown. The following will focus on... Figure 2 The process shown will be described in detail. The image processing method may specifically include the following steps: Step S210: Acquire multimodal data, wherein the multimodal data includes the user's eye data and the user's touch data on the touchscreen.

[0052] For a detailed description of step S210, please refer to step S110, which will not be repeated here.

[0053] Step S220: Perform signal quality analysis on the eyeball data to obtain a first analysis result, and perform signal quality analysis on the touch data to obtain a second analysis result.

[0054] Optionally, in this embodiment, when multimodal data is acquired, it can be processed by a "dynamic perception fusion device" built into the electronic device. This dynamic perception fusion device acts like a referee, analyzing the signal quality of each modality of data in real time (such as whether the image is blurry or the sound is noisy). Based on the analysis results, it can automatically adjust the "voting power" (weight) of the final decision corresponding to each modality of data, and thus determine the attention area. That is, this embodiment adjusts the signal quality in real time, rather than using fixed rules or simple learning weights, thus solving the core problem of signal quality fluctuations (such as sudden blurring or noise) in real environment.

[0055] In this embodiment, when multimodal data is acquired, signal quality analysis can be performed on the eye data in the multimodal data to obtain a first analysis result, and signal quality analysis can be performed on the touch data in the multimodal data to obtain a second analysis result.

[0056] In some implementations, signal quality analysis of eye data may include analyzing the signal clarity of the eye data. One approach is to analyze the signal clarity of the eye data by determining whether the eye movement coordinates are stable, such as determining whether blinking or occlusion causes a jump in the eye movement coordinates.

[0057] In some implementations, signal quality analysis of eye data may include analyzing the data integrity of the eye data. One approach is to analyze the data integrity of the eye data by examining the loss rate of eye movement signals in consecutive frames, such as whether data interruptions occur during rapid eye movements.

[0058] In some implementations, signal quality analysis of eye data may include analyzing environmental interference related to the eye data. One approach is to analyze environmental interference related to eye data by assessing the impact of changes in lighting and facial posture on eye-tracking accuracy, such as whether sensor recognition errors increase in low light.

[0059] In some implementations, signal quality analysis of touch data may include analyzing the trajectory continuity of the touch data. One approach is to analyze the trajectory continuity of the touch data by determining whether the swipe trajectory is broken due to accidental touch or finger removal from the touchscreen.

[0060] In some implementations, signal quality analysis of touch data may include analyzing the stability of the pressure signal corresponding to the touch data. One approach is to analyze the stability of the pressure signal by checking for abnormal fluctuations in pressure values, such as false pressure alarms caused by screen smudges.

[0061] In some implementations, signal quality analysis of touch data may include analyzing the timeliness of the response to the touch data. One approach is to analyze the timeliness of the response to the touch data by evaluating the delay between the touch signal and the touchscreen feedback, such as whether trajectory lag occurs under high load.

[0062] Step S230: Based on the first analysis result and the second analysis result, determine the first decision weight corresponding to the eye data and the second decision weight corresponding to the touch data.

[0063] In this embodiment, after obtaining the first analysis result and the second analysis result, the first decision weight corresponding to the eye data and the second decision weight corresponding to the touch data can be determined based on the first analysis result and the second analysis result, wherein the sum of the first decision weight and the second decision weight is equal to 1.

[0064] In some implementations, after obtaining a first analysis result and a second analysis result, the first analysis result and the second analysis result can be quantitatively scored to obtain a first score corresponding to the first analysis result and a second score corresponding to the second analysis result. Based on the first score and the second score, a first decision weight corresponding to the eye data and a second decision weight corresponding to the touch data are determined.

[0065] Optionally, the criteria for quantifying and scoring eye data may include, but are not limited to, signal clarity (such as coordinate stability), data integrity (such as data loss rate), and the degree of environmental interference (such as the influence of light).

[0066] Optionally, the criteria for quantifying and scoring touch data may include, but are not limited to, the continuity of the trajectory (e.g., whether it is broken), the stability of the pressure signal (e.g., the fluctuation amplitude), and the timeliness of response (e.g., the delay time).

[0067] In some implementations, based on a first score corresponding to a first analysis result and a second score corresponding to a second analysis result, the electronic device can calculate the weight ratio of eye data and touch data using a built-in algorithm, thereby obtaining a first decision weight and a second decision weight. The rules may include: If the eye data quality score (A) is higher than the touch data quality score (B), then the first decision weight (eye) > the second decision weight (touch). For example, if A = 80 points and B = 60 points, the possible decision weights are 60% (eye): 40% (touch).

[0068] If the touch data quality score (B) is higher than the eye data quality score (A), then the second decision weight (touch) > the first decision weight (eye). For example, if A = 50 points (signal disconnection due to blinking) and B = 90 points (continuous trajectory and low latency), the decision weights could be assigned as 30% (eye): 70% (touch).

[0069] If the two scores are similar (e.g., A=75 points, B=72 points), then the decision weights tend to be balanced (e.g., 55%: 45%).

[0070] It is understood that in this embodiment, the decision weights are not fixed values, but are updated in real time according to changes in signal quality. For example, if a user suddenly turns their head, causing blurred eye data, the first decision weight can be immediately reduced, while the second decision weight of touch data can be increased; if the touch signal is interrupted due to the finger leaving the screen, the weight of eye data can be quickly increased. By responding to signal quality fluctuations in real time, it ensures that in complex scenarios (such as sudden interference), the most reliable signal always dominates the decision-making process, providing a more accurate input basis for subsequent attention area prediction.

[0071] Step S240: Based on the eye data, the first decision weight, the touch data, and the second decision weight, determine the user's attention area on the touchscreen.

[0072] In this embodiment, after obtaining eye data, the first decision weight corresponding to the eye data, touch data, and the second decision weight corresponding to the touch data, the user's attention area on the touch screen can be determined based on the eye data, the first decision weight, the touch data, and the second decision weight.

[0073] In some implementations, the electronic device may pre-set and store a method for determining the attention area. When eye data, a first decision weight corresponding to the eye data, touch data, and a second decision weight corresponding to the touch data are acquired, the eye data, the first decision weight, the touch data, and the second decision weight can be processed by the determination method to determine the user's attention area on the touch screen.

[0074] Please see Figure 3 , Figure 3 This application shows Figure 2 The flowchart shown illustrates step S240 of the image processing method. The following will focus on... Figure 3 The process shown will be described in detail, and the method may specifically include the following steps: Step S241: Determine a first key local area in the touch screen based on the eye data, and determine a second key local area in the touch screen based on the touch data.

[0075] Optionally, in some implementations, after obtaining eye data, the first decision weight corresponding to the eye data, touch data, and the second decision weight corresponding to the touch data, important local areas (such as bright objects in an image or the direction from which sound comes) can be identified within the data of each modality. Then, the determined decision weights are used to guide the alignment and fusion of key areas of data from different modalities (e.g., the depth map indicates that the table is on the left, and the sound indicates that there is a noise on the left, so more attention is paid to the table area on the left side of the image). The fused data is then processed further to generate a preliminary heatmap of "where is most likely to be noticed", thereby determining the attention area.

[0076] In this embodiment, when multimodal data is acquired, a first key local region can be determined on the touchscreen based on eye data in the multimodal data, and a second key local region can be determined on the touchscreen based on touch data in the multimodal data.

[0077] In some implementations, when multimodal data is acquired, electronic devices can identify key local regions from eye data within the multimodal data to determine a first key local region. One approach is to extract the coordinates of the user's real-time focus point (such as the screen pixel position of the gaze point) using an eye-tracking sensor, and then identify the local region of eye focus as the first key local region (such as a button, the main subject of the screen, etc.).

[0078] In some implementations, when multimodal data is acquired, the electronic device can identify key local regions from the touch data within the multimodal data to determine a second key local region. One approach is to analyze touch trajectory sequences (such as swipe paths, click positions, pressure changes, etc.) and extract the local region indicated by the touch behavior as the key local region (such as the endpoint of a continuous swipe, a high-frequency click area, etc.).

[0079] Step S242: Using the first decision weight and the second decision weight, align and fuse the first key local region and the second key local region to obtain the fused information.

[0080] In this embodiment, after determining the first key local region and the second key local region, the first decision weight and the second decision weight can be used to align and fuse the first key local region and the second key local region to obtain the fused information.

[0081] In some implementations, when the first decision weight is greater than the second decision weight (because eye data is more reliable), the spatial correlation (such as distance and overlap) between the second key local region and the first key local region is calculated based on the first key local region. If the touch area is adjacent to the gaze area (e.g., distance < 50 pixels), the touch area is fused with the gaze area as an auxiliary attention area; if the distance is too far (e.g., > 200 pixels), the fusion weight of the touch area is reduced (only some edge features are retained) to avoid irrelevant interference.

[0082] In some implementations, when the second decision weight is greater than the first decision weight (touch data is more reliable), the second key local area is used as the core to verify whether the first key local area has a temporal correlation with the touch trajectory (e.g., whether the gaze follows the finger swipe). If the gaze area moves with the touch trajectory (e.g., the gaze tracks synchronously when the finger swipes), the gaze area is merged with the touch area as an enhanced verification area; if the gaze and touch are not related, only the core features of the touch area are retained, weakening the influence of the gaze area.

[0083] In some implementations, when the difference between the first decision weight and the second decision weight is less than the difference threshold (weight balance), the optimal overlapping or complementary region between the two regions (such as the intersection of the gaze region and the touch region, or the middle region of the line connecting the two) can be calculated by a spatial mapping algorithm. The overlapping region is used as the fusion core, and the non-overlapping part retains features according to the weight ratio.

[0084] In some implementations, after fusing the first and second key local regions, the rationality of the region can be verified (e.g., in a game scene, whether the fused region is a semantically important element such as a character or UI button). If the verification passes, the fused information (e.g., the joint attention region) is output; if it fails (e.g., the fused region is a meaningless screen edge), the alignment strategy is readjusted based on the weights (e.g., expanding the search range or strengthening the features of high-weight regions).

[0085] Understandably, the fused region, as a cross-modal alignment feature, enters the spatial and temporal branches of the model for in-depth processing, providing core input for subsequent heatmap generation (attention region prediction). This process, through dynamic weight adjustment, ensures that the fusion of the two key local regions conforms to both real-time signal quality and scene semantic logic, solving the problem of rigid region alignment in traditional static segmentation.

[0086] Step S243: Based on the fused information, determine the user's attention area on the touchscreen.

[0087] In some implementations, given the fused information, the user's attention area on the touchscreen can be determined based on the fused information.

[0088] As an feasible approach, given the fused information, the user's attention region in different modalities can be obtained based on the fused information, and the attention region can be determined as the user's attention region on the touchscreen.

[0089] Optionally, in some implementations, the electronic device can pre-configure and store an attention state memory bank, which acts like a drawer, storing attention patterns from important past moments (such as reading a document or searching for a key). It can analyze several consecutive frames of data to capture rapid changes in attention (such as the gaze following a moving ball), then search for similar historical attention patterns in the attention state memory bank and incorporate them to aid prediction (e.g., if the current scene resembles searching for a key, it strengthens the search prediction in areas where keys are commonly placed). Understandably, the attention state memory bank stores representative past attention states (not just raw data), and actively retrieves and applies similar historical states during prediction, thereby enabling the model to have long-term memory and the ability to associate specific attention patterns (such as "reading state" or "searching state"), thus avoiding the limitation of only processing recent sequences and significantly improving the prediction accuracy for repetitive or patterned behaviors.

[0090] In some implementations, determining the user's attention area on the touchscreen based on the fused information may include: determining the user's current attention behavior based on eye data and touch data; searching for historical attention patterns in the attention state memory database where the similarity between the corresponding historical attention behavior and the current attention behavior reaches a similarity threshold, wherein the attention state memory database stores multiple attention patterns of the user within a historical time period; and determining the user's attention area on the touchscreen based on the fused information and the historical attention patterns.

[0091] As an feasible approach, eye data (such as real-time gaze coordinates, gaze movement trajectory, and dwell time) and touch data (such as click location, swipe path, pressure change, and operation frequency) are first extracted. The dynamic characteristics of the two types of data are analyzed through models, and the key attributes of the current attention behavior are extracted based on the dynamic characteristics, such as the screen area where attention is focused and the time pattern of the behavior.

[0092] The attention state memory stores the user's historical attention patterns. Each pattern contains features of historical attention behaviors (such as gazing at a text area while reading a document and slightly swiping to turn pages, or gazing at a character in a game and frequently clicking skill buttons). Based on this, the features of the current attention behavior (such as navigation bar operation and clicking immediately after gazing) are compared with the historical patterns in the memory. The matching degree is calculated using algorithms such as feature vector cosine similarity. If the similarity of a certain historical pattern (such as clicking the navigation bar to switch pages in a browser) with the current behavior in terms of spatial attributes (same area), temporal attributes (clicking after gazing), and semantic attributes (navigation interaction) is greater than or equal to a threshold (such as 80%), it is determined to be a similar pattern.

[0093] Next, the spatial region of the current attention behavior (e.g., the detected navigation bar area) is aligned with the corresponding typical region in historical patterns, and the current region boundary is corrected (e.g., extended by 50 pixels to cover historically high-frequency operation points). Combining the temporal patterns of historical patterns (e.g., attention briefly spreads to adjacent buttons during navigation bar operations), the region where current attention may expand is predicted (e.g., the adjacent favorite button may also be noticed in addition to the home button). If the current data quality is high (e.g., clear eye data, no accidental touches), the current information is prioritized; if the current signal contains noise (e.g., occasional gaze shifts), the reference weight of similar historical patterns is increased. After fusion, an attention heatmap is generated, and regions with a probability > 0.7 in the heatmap are considered the final attention region.

[0094] Step S250: Render the image corresponding to the attention region according to the first resolution, and render the image corresponding to the non-attention region on the touch screen other than the attention region according to the second resolution, wherein the first resolution is greater than the second resolution.

[0095] Step S260: The image corresponding to the rendered attention region is stitched together with the image corresponding to the rendered non-attention region, and the stitched image is displayed on the touch screen.

[0096] For a detailed description of steps S250-S260, please refer to steps S130-S140, which will not be repeated here.

[0097] An embodiment of this application provides an image processing method that acquires multimodal data, including user eye data and user touch data on a touchscreen. The method performs signal quality analysis on the eye data to obtain a first analysis result, and performs a second quality analysis on the touch data to obtain a second analysis result. Based on the first and second analysis results, it determines a first decision weight corresponding to the eye data and a second decision weight corresponding to the touch data. Based on the eye data, the first decision weight, the touch data, and the second decision weight, it determines the user's attention region on the touchscreen. It renders the image corresponding to the attention region according to a first resolution, and renders the image corresponding to the non-attention region on the touchscreen according to a second resolution, wherein the first resolution is greater than the second resolution. The rendered image corresponding to the attention region and the rendered image corresponding to the non-attention region are then stitched together, and the stitched image is displayed on the touchscreen. Compared to... Figure 1 The image processing method shown in this embodiment also dynamically senses the signal quality to intelligently fuse data from different modalities, which can improve the accuracy of the determined attention area and the accuracy of image rendering.

[0098] Please see Figure 4 ,Figure 4 A schematic flowchart of an image processing method according to an embodiment of this application is shown. The following will focus on... Figure 4 The process shown will be described in detail. The image processing method may specifically include the following steps: Step S310: Acquire multimodal data, wherein the multimodal data includes the user's eye data and the user's touch data on the touch screen, the eye data includes eye coordinates, and the touch data includes a touch trajectory sequence.

[0099] For a detailed description of step S310, please refer to step S110, which will not be repeated here.

[0100] Step S320: Extract the spatial features of the eye coordinates through a deep residual network, and generate a first feature map based on the spatial features.

[0101] Optionally, eye data may include eye coordinates, which serve as the core input to the spatial flow branch (ResNet-18) and directly reflect the spatial location of the user's gaze (e.g., (x, y) pixel coordinates on the screen). This data can be acquired in real time using an eye-tracking sensor to locate the spatial position of the user's current visual focus, providing the fundamental spatial basis for dividing high-resolution regions.

[0102] Touch data can include a sequence of touch trajectories. This sequence, serving as a key input to the temporal flow branch (LSTM), contains continuous touch positions, swipe directions, and timestamp information. This sequence is used to capture the temporal characteristics of user interactions, helping to determine the dynamic shift of attention (e.g., finger swipe trajectories predicting upcoming eye movements).

[0103] In this embodiment, given the eye coordinates, spatial features of the eye coordinates can be extracted using a deep residual network (ResNet-18), and a first feature map can be generated based on these spatial features. As one approach, given the eye coordinates, spatial location information centered on the eye coordinates can be determined. ResNet-18 (a deep convolutional neural network) can then extract features from the spatial distribution of the eye coordinates through multiple convolutional operations. Specifically, local features of the eye's focus area (such as densely clustered regions of coordinates, i.e., high-frequency gaze points) can be identified, spatial correlations (such as the distance and distribution shape between different gaze points, determining whether attention is focused or scattered) can be captured, and finally, a first feature map is generated. Optionally, the first feature map has a size of 224×224×512, where 224×224 corresponds to the pixel grid of the screen space, and 512 represents the 512 extracted spatial features (such as region edges, density distribution, etc.).

[0104] Step S330: Extract the dynamic features of the touch trajectory sequence over time using a long short-term memory network, and generate a second feature map based on the dynamic features.

[0105] In this embodiment, when a touch trajectory sequence is obtained, dynamic features of the touch trajectory sequence changing over time can be extracted using a Long Short-Term Memory (LSTM) network, and a second feature map can be generated based on these dynamic features. As one approach, when a touch trajectory sequence is obtained, the temporal dynamic information of the touch trajectory sequence can be determined. An LSTM network, through a recurrent neural network structure, analyzes the temporal patterns of touch behavior, capturing the changing trends of the trajectory (such as swiping direction, speed increase / decrease, and determining whether attention follows finger movement), extracting temporal correlations (such as the interval between consecutive clicks, the time difference between touch and gaze, and determining the persistence of the interaction intent), and finally generating a second feature map. Optionally, the size of the second feature map is 224×224×512, where 224×224 also corresponds to the screen space grid, and 512 represents the 512 extracted temporal features (such as trajectory acceleration, dwell time, etc.).

[0106] Step S340: Concatenate the first feature map and the second feature map along the channel dimension to obtain an attention heatmap.

[0107] In this embodiment, given the first feature map and the second feature map, the first feature map and the second feature map can be stitched together along the channel dimension to obtain an attention heatmap.

[0108] In some implementations, given the first feature map and the second feature map, the first feature map of the spatial flow and the second feature map of the temporal flow can be concatenated along the channel dimension using the Concatenate operation (i.e., merging 512+512=1024 features) to achieve complementary fusion of spatial static features and temporal dynamic features.

[0109] In this model, spatial features (such as whether the pixel is being gazed upon) and temporal features (such as whether it is being touched) at the same screen location (a pixel within a 224×224 grid) are correlated and integrated to form a comprehensive attention feature for that location. The fused features are then processed by a heatmap generation module, ultimately outputting an attention heatmap with a size of 224×224×1. The value of each pixel in the attention heatmap represents the probability of the user's attention at that location (the higher the value, the more focused the attention).

[0110] Understandably, this embodiment uses a dual-branch structure to capture spatial information about where the user is looking using ResNet-18 and temporal information about what the user is doing using LSTM. Finally, it fuses these two types of features to generate an attention heatmap, achieving accurate prediction of the user's visual focus. This overcomes the limitations of single-modal data (eye movement only or touch only) and improves the accuracy of attention region prediction in complex scenes.

[0111] Step S350: Based on the attention heatmap, determine the user's attention area on the touchscreen.

[0112] In this embodiment, with an attention heatmap obtained, the user's attention area on the touchscreen can be determined based on the attention heatmap.

[0113] In some implementations, when an attention heatmap is obtained, a threshold can be used to segment the attention heatmap into regions, and continuous regions with attention probabilities greater than the threshold can be selected from the attention heatmap and identified as attention regions.

[0114] Step S360: Render the image corresponding to the attention region according to the first resolution, and render the image corresponding to the non-attention region on the touch screen other than the attention region according to the second resolution, wherein the first resolution is greater than the second resolution.

[0115] Step S370: The image corresponding to the rendered attention region is stitched together with the image corresponding to the rendered non-attention region, and the stitched image is displayed on the touch screen.

[0116] For a detailed description of steps S360-S370, please refer to steps S130-S140, which will not be repeated here.

[0117] An embodiment of this application provides an image processing method that acquires multimodal data, including user eye data and user touch data on a touchscreen. The eye data includes eye coordinates, and the touch data includes touch trajectory sequences. A deep residual network is used to extract spatial features of the eye coordinates, and a first feature map is generated based on these spatial features. A long short-term memory network is used to extract dynamic features of the touch trajectory sequences over time, and a second feature map is generated based on these dynamic features. The first and second feature maps are then concatenated along the channel dimension to obtain an attention heatmap. Based on the attention heatmap, the user's attention region on the touchscreen is determined. The image corresponding to the attention region is rendered according to a first resolution, and the images corresponding to non-attention regions on the touchscreen (excluding the attention region) are rendered according to a second resolution, where the first resolution is greater than the second resolution. The rendered images corresponding to the attention region and the rendered images corresponding to the non-attention region are then concatenated, and the concatenated image is displayed on the touchscreen. Compared to... Figure 1 The image processing method shown in this embodiment also performs spatial and temporal dimension processing on data of different modalities to improve the rationality of data processing and thus improve the accuracy of image processing.

[0118] Please see Figure 5 , Figure 5 A schematic flowchart of an image processing method according to an embodiment of this application is shown. The following will focus on... Figure 5 The process shown will be described in detail. The image processing method may specifically include the following steps: Step S410: Acquire multimodal data, wherein the multimodal data includes the user's eye data, the user's touch data on the touchscreen, and the touchscreen's display data.

[0119] For a detailed description of step S410, please refer to step S110, which will not be repeated here.

[0120] Step S420: Perform semantic analysis on the display data to obtain the display scene.

[0121] Optionally, multimodal data may also include display data.

[0122] The displayed data may include content information rendered by the touchscreen, including but not limited to: application type (such as whether the foreground is a game application, browser or video player); interface elements (such as the layout and attributes of buttons, text areas, images, dynamic video frames, etc.); scene features (such as battle scenes / menu interfaces in games, text paragraphs / chart areas in documents).

[0123] In this embodiment, when multimodal data is obtained, semantic analysis can be performed on the display data in the multimodal data to obtain the display scene.

[0124] In some implementations, when multimodal data is acquired, hierarchical semantic parsing can be performed on the display data in the multimodal dataset to obtain the display scene. As one approach, the type of interface elements (such as text, images, interactive buttons, and dynamic characters) can be identified, and their spatial location and functional attributes (such as button clickability and text readability) can be determined. Based on element combinations and application types, the overall display scene can be judged (e.g., a game battle scene includes elements such as characters, skill effects, and health bars; a reading scene is centered around continuous text blocks). Combining scene characteristics, semantic features of high-priority areas can be marked (e.g., character positions and skill buttons are high priority in a game scene; the current reading behavior is high priority in a reading scene).

[0125] Among them, obtaining the display scene through semantic analysis of the display data can provide scene prior knowledge for multimodal fusion, reduce the dependence on sensor data (such as inferring the area that the user may pay attention to based on scene features when eye movement signals are briefly lost), thereby improving the robustness of attention area prediction, especially in complex dynamic scenes (such as rapidly switching game interfaces).

[0126] Step S430: Determine the user's attention area on the touchscreen based on the eye data, the touch data, and the display scene.

[0127] In this embodiment, given eye data, touch data, and the display scene, the user's attention area on the touchscreen can be determined based on the eye data, touch data, and the display scene.

[0128] In some implementations, eye data can be input into the spatial flow branch to generate a first feature map reflecting spatial focus characteristics, while touch data can be input into the temporal flow branch to generate a second feature map reflecting dynamic interaction characteristics. Semantic features of the display scene (such as scene type and high-priority element coordinates) are used as additional constraints to adjust the weights of eye and touch data. Then, the eye gaze area and touch operation area are spatially matched with high-priority elements in the display scene (such as the position of a game character), filtering for overlapping or adjacent areas (such as the gaze coinciding with the character's position or the touch point being close to a button). Combined with dynamic changes in the display scene (such as character movement in a game), the temporal patterns of eye / touch data are analyzed (such as the gaze following character movement or the timing of touch operations matching skill cooldowns), strengthening the prediction of dynamic attention areas. Finally, the first feature map, second feature map, and scene semantic features are merged through a concatenation operation to generate a comprehensive feature, and based on this comprehensive feature, the user's attention area on the touchscreen is determined.

[0129] Step S440: Render the image corresponding to the attention region according to the first resolution, and render the image corresponding to the non-attention region on the touch screen other than the attention region according to the second resolution, wherein the first resolution is greater than the second resolution.

[0130] Step S450: The image corresponding to the rendered attention region is stitched together with the image corresponding to the rendered non-attention region, and the stitched image is displayed on the touch screen.

[0131] For a detailed description of steps S440-S450, please refer to steps S130-S140, which will not be repeated here.

[0132] An embodiment of this application provides an image processing method that acquires multimodal data, including user eye data, user touch data on a touchscreen, and touchscreen display data. Semantic analysis is performed on the display data to obtain a display scene. Based on the eye data, touch data, and display scene, the user's attention area on the touchscreen is determined. The image corresponding to the attention area is rendered according to a first resolution, and the images corresponding to non-attention areas on the touchscreen (excluding the attention area) are rendered according to a second resolution, wherein the first resolution is greater than the second resolution. The rendered images corresponding to the attention area and the rendered images corresponding to the non-attention area are then stitched together, and the stitched image is displayed on the touchscreen. Compared to... Figure 1 The image processing method shown in this embodiment also integrates eye tracking, touch hotspots, and scene semantic analysis to predict the visual focus area, which can improve the accuracy of the determined attention area and the accuracy of image processing.

[0133] Please see Figure 6 , Figure 6 A schematic flowchart of an image processing method according to an embodiment of this application is shown. The following will focus on... Figure 6 The process shown will be described in detail. The image processing method may specifically include the following steps: Step S510: Acquire multimodal data, wherein the multimodal data includes the user's eye data and the user's touch data on the touchscreen.

[0134] Step S520: Determine the user's attention area on the touchscreen based on the eye data and the touch data.

[0135] For a detailed description of steps S510-S520, please refer to steps S110-S120, which will not be repeated here.

[0136] Step S530: Obtain the pose data corresponding to the head-mounted device.

[0137] Optionally, the electronic device may include a head-mounted device, which a user may wear.

[0138] In this embodiment, the pose data corresponding to the head-mounted device can be obtained.

[0139] As a feasible approach, a head-mounted device can incorporate an inertial measurement unit (IMU) to acquire raw head motion data in real time. This raw data includes acceleration and angular velocity data. Acceleration data is obtained by capturing the linear acceleration changes of the head in three-dimensional space (X, Y, Z axes) using an accelerometer; angular velocity data is obtained by capturing the rotational angular velocities of the head around three-dimensional axes (such as the speed of head turning, tilting, and tilting) using a gyroscope. The raw data acquired by the IMU is then filtered to eliminate noise interference (such as errors caused by device vibration and sensor drift), ensuring data stability. Based on the pre-processed IMU data, kinematic algorithms (such as Kalman filtering and extended Kalman filtering) are used in real time to convert the acceleration and angular velocity data into spatial pose parameters of the head. These pose parameters include attitude information and position information. Attitude information includes the three-dimensional rotation angles of the head (such as pitch, yaw, and roll angles), reflecting the head's orientation; position information includes the relative displacement of the head in three-dimensional space.

[0140] Step S540: Determine the visual center region from the attention region based on the pose data.

[0141] In this embodiment, when the pose data corresponding to the head-mounted device (user's head) is obtained, the visual center region can be determined from the attention region based on the pose data.

[0142] In some implementations, given the pose data of the head-mounted device, a spatial coordinate transformation algorithm can be used to construct a user's current visual field cone model. This model uses the head position as the vertex and the gaze direction corresponding to the pose data as the central axis, simulating the user's actual perceptible visual field boundary. The attention region is projected into a virtual 3D space, and spatial overlap calculations are performed with the visual field cone. Attention sub-regions that are completely or primarily located within the visual field cone are selected, and the spatial distance between each sub-region and the central axis of the visual field cone (the theoretical center of gaze) is calculated; the closer the distance, the higher the priority. Based on this, the final visual field center region can be determined using the central axis of the visual field cone as the core, combined with the heatmap probability of the attention region (a higher probability indicates a greater likelihood of user attention): if an attention sub-region simultaneously satisfies the conditions of being located within the visual field cone and having a heatmap probability > 0.7, it is designated as the visual field center region.

[0143] Step S550: Render the image corresponding to the visual field center region according to the first resolution, and render the images corresponding to the non-attention region and the attention region other than the visual field center region according to the second resolution, wherein the first resolution is greater than the second resolution.

[0144] In this embodiment, when the visual field center region is determined from the attention region, the display area of ​​the touchscreen of the electronic device is equivalent to being divided into: the visual field center region, the attention region excluding the visual field center region, and the non-attention region. In this case, the user's attention region can be considered the visual field center region.

[0145] Based on this, the image corresponding to the center region of the field of view can be rendered at a higher resolution (first resolution), and the images corresponding to the non-attention region and the attention region outside the center region of the field of view can be rendered at a lower resolution (second resolution). This reduces the area rendered at high resolution, thereby reducing device power consumption while ensuring the user's visual experience.

[0146] Step S560: The image corresponding to the rendered attention region is stitched together with the image corresponding to the rendered non-attention region, and the stitched image is displayed on the touch screen.

[0147] For a detailed description of step S560, please refer to step S140, which will not be repeated here.

[0148] An embodiment of this application provides an image processing method that acquires multimodal data, including user eye data and user touch data on a touchscreen. Based on the eye data and touch data, the method determines the user's attention region on the touchscreen, acquires pose data corresponding to a head-mounted device, determines a visual field center region from the attention region based on the pose data, renders the image corresponding to the visual field center region according to a first resolution, and renders images corresponding to non-attention regions and attention regions other than the visual field center region according to a second resolution, wherein the first resolution is greater than the second resolution. The rendered image corresponding to the attention region is then stitched together with the rendered image corresponding to the non-attention region, and the stitched image is displayed on the touchscreen. Compared to... Figure 1 The image processing method shown in this embodiment also predicts the center of the user's field of vision for head-mounted devices and renders only the image within the center of the field of vision at high resolution, which can improve the rendering rationality and reduce rendering power consumption.

[0149] Please see Figure 7 , Figure 7 A block diagram of an image processing apparatus according to an embodiment of this application is shown. The following will focus on... Figure 7 The image processing device 200, illustrated in the block diagram, includes: a multimodal data acquisition module 210, an attention region determination module 220, an image rendering module 230, and an image display module 240, wherein: The multimodal data acquisition module 210 is used to acquire multimodal data, wherein the multimodal data includes the user's eye data and the user's touch data on the touch screen.

[0150] Attention area determination module 220 is used to determine the user's attention area on the touch screen based on the eye data and the touch data.

[0151] Further, the attention region determination module 220 includes: a signal quality analysis submodule, a decision weight determination submodule, and a first attention region determination submodule, wherein: The signal quality analysis submodule is used to perform signal quality analysis on the eye data to obtain a first analysis result, and to perform signal quality analysis on the touch data to obtain a second analysis result.

[0152] The decision weight determination submodule is used to determine the first decision weight corresponding to the eye data and the second decision weight corresponding to the touch data based on the first analysis result and the second analysis result.

[0153] The first attention region determination submodule is used to determine the user's attention region on the touch screen based on the eye data, the first decision weight, the touch data, and the second decision weight.

[0154] Further, the first attention region determination submodule includes: a local region determination unit, a local region fusion unit, and an attention region determination unit, wherein: A local area determination unit is used to determine a first key local area in the touch screen based on the eye data, and to determine a second key local area in the touch screen based on the touch data.

[0155] The local region fusion unit is used to align and fuse the first key local region and the second key local region using the first decision weight and the second decision weight to obtain fused information.

[0156] An attention region determination unit is used to determine the user's attention region on the touchscreen based on the fused information.

[0157] Furthermore, the attention region determination unit includes: an attention behavior determination subunit, an attention pattern search subunit, and an attention region determination subunit, wherein: The attention behavior determination subunit is used to determine the user's current attention behavior based on the eye data and the touch data.

[0158] The attention pattern lookup subunit is used to search for historical attention patterns in the attention state memory that have a similarity threshold with the current attention behavior. The attention state memory stores multiple attention patterns of the user within a historical time period.

[0159] An attention region determination subunit is used to determine the user's attention region on the touchscreen based on the fused information and the historical attention pattern.

[0160] Further, the eye data includes eye coordinates, the touch data includes a touch trajectory sequence, and the attention region determination module 220 includes: a first feature map generation submodule, a second feature map generation submodule, an attention heatmap acquisition submodule, and a second attention region determination submodule, wherein: The first feature map generation submodule is used to extract the spatial features of the eye coordinates through a deep residual network and generate a first feature map based on the spatial features.

[0161] The second feature map generation submodule is used to extract the dynamic features of the touch trajectory sequence over time through a long short-term memory network, and generate a second feature map based on the dynamic features.

[0162] The attention heatmap acquisition submodule is used to concatenate the first feature map and the second feature map along the channel dimension to obtain an attention heatmap.

[0163] The second attention region determination submodule is used to determine the user's attention region on the touch screen based on the attention heatmap.

[0164] Furthermore, the multimodal data also includes the display data of the touchscreen, and the attention region determination module 220 includes: a display scene acquisition submodule and a third attention region determination submodule, wherein: The display scene acquisition submodule is used to perform semantic analysis on the display data to obtain the display scene.

[0165] The third attention area determination submodule is used to determine the user's attention area on the touch screen based on the eye data, the touch data, and the display scene.

[0166] The image rendering module 230 is used to render the image corresponding to the attention area according to a first resolution, and to render the image corresponding to the non-attention area on the touch screen other than the attention area according to a second resolution, wherein the first resolution is greater than the second resolution.

[0167] Furthermore, the image rendering module 230 includes: a transition region determination submodule and a first image rendering submodule, wherein: The transition region determination submodule is used to determine, from the non-attention region, a non-attention region adjacent to the attention region as a transition region.

[0168] The first image rendering submodule is used to render the image corresponding to the attention region according to a first resolution, render the image corresponding to the non-attention region (excluding the transition region) according to a second resolution, and render the image corresponding to the transition region according to the first resolution and the second resolution.

[0169] Furthermore, the image rendering module 230 includes: a second image rendering submodule, wherein: The second image rendering submodule is used to render the image corresponding to the attention region according to the first resolution using multisampling anti-aliasing technology, and to render the image corresponding to the non-attention region according to the second resolution using bilinear sampling technology.

[0170] Furthermore, the electronic device includes a head-mounted device, and the image rendering module 230 includes: a pose data acquisition submodule, a field-of-view center region determination submodule, and a third image rendering submodule, wherein: The pose data acquisition submodule is used to acquire pose data corresponding to the head-mounted device.

[0171] The visual field center region determination submodule is used to determine the visual field center region from the attention region based on the pose data.

[0172] The third image rendering submodule is used to render the image corresponding to the center region of the field of view according to the first resolution, and to render the images corresponding to the non-attention region and the attention region other than the center region of the field of view according to the second resolution.

[0173] The image display module 240 is used to stitch together the image corresponding to the rendered attention area and the image corresponding to the rendered non-attention area, and display the stitched image on the touch screen.

[0174] Furthermore, the image processing device 200 further includes: a device status acquisition module and a resolution adjustment module, wherein: The device status acquisition module is used to acquire the device status of the electronic device.

[0175] A resolution adjustment module is used to adjust the first resolution and the second resolution based on the device status.

[0176] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the above-described device and module can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0177] In the several embodiments provided in this application, the coupling between modules can be electrical, mechanical, or other forms of coupling.

[0178] Furthermore, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module. The integrated modules described above can be implemented in hardware or as software functional modules.

[0179] Please see Figure 8This document illustrates a structural block diagram of an electronic device 100 provided in an embodiment of this application. The electronic device 100 can be a smartphone, tablet computer, e-reader, or other electronic device capable of running applications. The electronic device 100 in this application may include one or more of the following components: a processor 110, a memory 120, and one or more applications, wherein the one or more applications can be stored in the memory 120 and configured to be executed by one or more processors 110, and the one or more applications are configured to perform the methods described in the foregoing method embodiments.

[0180] The processor 110 may include one or more processing cores. The processor 110 connects to various parts of the electronic device 100 using various interfaces and lines, and performs various functions and processes data of the electronic device 100 by running or executing instructions, programs, code sets, or instruction sets stored in the memory 120, and by calling data stored in the memory 120. Optionally, the processor 110 may be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). The processor 110 may integrate one or a combination of several of the following: Central Processing Unit (CPU), Graphics Processing Unit (GPU), and modem. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the content to be displayed; and the modem handles wireless communication. It is understood that the modem may also not be integrated into the processor 110 and may be implemented separately using a communication chip.

[0181] The memory 120 may include random access memory (RAM) or read-only memory (ROM). The memory 120 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 120 may include a program storage area and a data storage area. The program storage area may store instructions for implementing an operating system, instructions for implementing functions (such as touch functionality, sound playback functionality, image playback functionality, etc.), and instructions for implementing the various method embodiments described below. The data storage area may also store data created by the electronic device 100 during use (such as phonebook data, audio and video data, chat log data, etc.).

[0182] In some embodiments, the electronic device 100 may further include a touch screen, which is used to display information input by the user, information provided to the user, and various graphical user interfaces of the electronic device 100. These graphical user interfaces may be composed of graphics, text, icons, numbers, video, and any combination thereof. In one example, the touch screen may be a liquid crystal display (LCD) or an organic light-emitting diode (OLED), without limitation.

[0183] In some embodiments, the electronic device 100 may further include a camera for collecting user motion data. Optionally, the camera may include a front-facing camera, a rear-facing camera, a telescopic camera, a rotating camera, etc., and is not limited thereto.

[0184] In some embodiments, the electronic device 100 may further include sensors, including a light sensor that can be used to turn off the touchscreen display when an object approaches the touchscreen, such as when the body of the electronic device is moved to the ear. The sensor may also include a pressure sensor that can detect pressure generated by pressing on the electronic device; that is, the pressure sensor can detect pressure generated by contact or pressing between the user and the electronic device, such as pressure generated by contact or pressing between the user's ear or finger and the electronic device. The sensor may also include a gravity sensor that can detect the magnitude of acceleration in various directions (generally three axes), and can detect the magnitude and direction of gravity when stationary. This can be used for applications that identify the posture of the electronic device (such as landscape / portrait switching, related games, magnetometer posture calibration), vibration recognition-related functions (such as pedometers, taps), etc. Additionally, the electronic device may also be equipped with other sensors such as gyroscopes, barometers, and hygrometers, which are not limited here.

[0185] In some embodiments, the electronic device 100 may also include an artificial intelligence module, which may be integrated into the processor 110 of the electronic device 100 to improve the intelligence level and performance of the electronic device.

[0186] Of course, electronic devices may include many other components, which will not be elaborated here.

[0187] Please see Figure 9 This diagram illustrates a structural block diagram of a computer-readable storage medium provided in an embodiment of this application. The computer-readable medium 300 stores program code that can be called by a processor to execute the methods described in the above method embodiments.

[0188] The computer-readable storage medium 300 may be an electronic memory such as flash memory, EEPROM (Electrically Erasable Programmable Read-Only Memory), EPROM, hard disk, or ROM. Optionally, the computer-readable storage medium 300 includes a non-transitory computer-readable storage medium. The computer-readable storage medium 300 has storage space for program code 310 that performs any of the method steps described above. This program code can be read from or written to one or more computer program products. The program code 310 may be compressed, for example, in a suitable form.

[0189] In some embodiments, this application provides a computer program product, which includes a computer program that, when executed by a processor, implements the image processing method of this application.

[0190] In summary, the image processing method, apparatus, electronic device, storage medium, and program product provided in this application acquire multimodal data, including user eye data and user touch data on a touchscreen. Based on the eye data and touch data, the user's attention area on the touchscreen is determined. The image corresponding to the attention area is rendered according to a first resolution, and the image corresponding to the non-attention area on the touchscreen (excluding the attention area) is rendered according to a second resolution, wherein the first resolution is greater than the second resolution. The rendered image corresponding to the attention area and the rendered image corresponding to the non-attention area are stitched together, and the stitched image is displayed on the touchscreen. Thus, by fusing multimodal data including eye data and touch data to determine the user's attention area and dynamically partitioning and rendering images at different resolutions in real time, the visual experience can be guaranteed while reducing the load on the electronic device.

[0191] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. An image processing method, characterized in that, Applied to an electronic device, the electronic device including a touchscreen, the method includes: Acquire multimodal data, wherein the multimodal data includes the user's eye data and the user's touch data on the touchscreen; Based on the eye data and the touch data, the user's attention area on the touchscreen is determined; The image corresponding to the attention region is rendered according to a first resolution, and the image corresponding to the non-attention region on the touch screen (excluding the attention region) is rendered according to a second resolution, wherein the first resolution is greater than the second resolution; The image corresponding to the rendered attention region is stitched together with the image corresponding to the rendered non-attention region, and the stitched image is displayed on the touch screen.

2. The method according to claim 1, characterized in that, Determining the user's attention area on the touchscreen based on the eye data and the touch data includes: A first analysis result is obtained by performing signal quality analysis on the eyeball data, and a second analysis result is obtained by performing signal quality analysis on the touch data. Based on the first analysis result and the second analysis result, a first decision weight corresponding to the eyeball data and a second decision weight corresponding to the touch data are determined; Based on the eye data, the first decision weight, the touch data, and the second decision weight, the user's attention area on the touchscreen is determined.

3. The method according to claim 2, characterized in that, Determining the user's attention area on the touchscreen based on the eye data, the first decision weight, the touch data, and the second decision weight includes: A first key local area is determined on the touchscreen based on the eye data, and a second key local area is determined on the touchscreen based on the touch data; Using the first decision weight and the second decision weight, the first key local region and the second key local region are aligned and fused to obtain fused information; Based on the fused information, the user's attention area on the touchscreen is determined.

4. The method according to claim 3, characterized in that, Determining the user's attention area on the touchscreen based on the fused information includes: Based on the eye data and the touch data, the user's current attention behavior is determined; Search the attention state memory for historical attention patterns that have a similarity threshold with the current attention behavior. The attention state memory stores multiple attention patterns of the user within a historical time period. Based on the fused information and the historical attention patterns, the user's attention area on the touchscreen is determined.

5. The method according to claim 1, characterized in that, The eye data includes eye coordinates, and the touch data includes a touch trajectory sequence. Determining the user's attention area on the touchscreen based on the eye data and the touch data includes: The spatial features of the eye coordinates are extracted using a deep residual network, and a first feature map is generated based on the spatial features. The dynamic features of the touch trajectory sequence over time are extracted using a long short-term memory network, and a second feature map is generated based on the dynamic features. The first feature map and the second feature map are concatenated along the channel dimension to obtain an attention heatmap; Based on the attention heatmap, the user's attention area on the touchscreen is determined.

6. The method according to claim 1, characterized in that, The step of rendering the image corresponding to the attention region according to a first resolution, and rendering the image corresponding to the non-attention region on the touchscreen other than the attention region according to a second resolution, includes: From the non-attention regions, determine the non-attention regions adjacent to the attention regions as transition regions; The image corresponding to the attention region is rendered according to a first resolution, the image corresponding to the non-attention region (excluding the transition region) in the non-attention region is rendered according to a second resolution, and the image corresponding to the transition region is rendered according to the first resolution and the second resolution.

7. The method according to claim 1, characterized in that, The step of rendering the image corresponding to the attention region according to a first resolution, and rendering the image corresponding to the non-attention region on the touchscreen other than the attention region according to a second resolution, includes: Multisampling anti-aliasing technology is used to render the image corresponding to the attention region according to the first resolution, and bilinear sampling technology is used to render the image corresponding to the non-attention region according to the second resolution.

8. The method according to any one of claims 1-7, characterized in that, After stitching the image corresponding to the rendered attention region with the image corresponding to the rendered non-attention region and displaying the stitched image on the touchscreen, the method further includes: Obtain the device status of the electronic device; Based on the device status, the first resolution and the second resolution are adjusted.

9. The method according to any one of claims 1-7, characterized in that, The multimodal data also includes the display data of the touchscreen. Determining the user's attention area on the touchscreen based on the eye data and the touch data includes: Perform semantic analysis on the displayed data to obtain the display scene; Based on the eye data, the touch data, and the display scene, the user's attention area on the touchscreen is determined.

10. The method according to any one of claims 1-7, characterized in that, The electronic device includes a head-mounted device. The step of rendering the image corresponding to the attention area according to a first resolution, and rendering the image corresponding to the non-attention area on the touchscreen (excluding the attention area) according to a second resolution, includes: Obtain the pose data corresponding to the head-mounted device; The visual center region is determined from the attention region based on the pose data; The image corresponding to the visual field center region is rendered according to the first resolution, and the images corresponding to the non-attention region and the attention region other than the visual field center region are rendered according to the second resolution.

11. An image processing apparatus, characterized in that, Applied to an electronic device, the electronic device including a touch screen, the device includes: A multimodal data acquisition module is used to acquire multimodal data, wherein the multimodal data includes the user's eye data and the user's touch data on the touch screen; An attention area determination module is used to determine the user's attention area on the touchscreen based on the eye data and the touch data. An image rendering module is used to render the image corresponding to the attention region according to a first resolution, and to render the image corresponding to the non-attention region on the touch screen other than the attention region according to a second resolution, wherein the first resolution is greater than the second resolution; The image display module is used to stitch together the image corresponding to the rendered attention area and the image corresponding to the rendered non-attention area, and then display the stitched image on the touch screen.

12. An electronic device, characterized in that, The method includes a memory and a processor, the memory being coupled to the processor, the memory storing instructions, and when the instructions are executed by the processor, the processor performing the method as described in any one of claims 1-10.

13. A computer-readable storage medium, characterized in that, The computer-readable storage medium contains program code that can be invoked by a processor to execute the method as described in any one of claims 1-10.

14. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the method as described in any one of claims 1-10.