AI-BASED VIDEO TRANSFER SYSTEM

TR202614362A2Pending Publication Date: 2026-09-21TURKCELL TEKNOLOJI ARASTIRMA & GELISTIRME AS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
TR202614362
Authority / Receiving Office
TR · TR
Patent Type
Applications
Current Assignee / Owner
Filing Date
2026-08-25
Publication Date
2026-09-21
Patent Text Reader

Abstract

This invention relates to a system (1) that analyzes the user's eye movements and the data transmission conditions of the mobile network to determine the regions on the image that the user is looking at and is expected to look at shortly, transmits these regions in high quality while transmitting the image regions that the user is not looking at in lower quality, thereby reducing network usage.
Need to check novelty before this filing date? Find Prior Art

Description

1 TARIFF AI-BASED VIDEO TRANSFER SYSTEM Technical Area This invention analyzes the data transmission conditions of a mobile network based on the user's eye movements. by pointing out the image the user is looking at and is expected to look at shortly. Identifying regions, delivering high-quality information about these regions while avoiding what the user is looking for. By transmitting image regions at a lower quality, it reduces network usage. It is related to a system. 10 Previous Technique In current focal region-priority imaging methods, the user's gaze By determining its location, the viewed area of ​​the image is high, while the surrounding areas are 15. is transmitted with lower image quality. In these systems, the region of interest (Region) Image data is generated using bits of Interest (ROI) based on viewing position. They are transmitted with their speeds and transmission priorities. However, the perspective is... sudden changes in location, network latency, decreased data transfer speed, or packet Due to the loss, the high-quality image was transferred to the new focus area in a timely manner. It cannot be delivered. Therefore, depending on the change in ROI perspective, it was previously... inability to update and current network image encoding and transmission parameters The shortcomings arise in that they cannot be organized together according to the conditions. Therefore, considering the studies and shortcomings in the current technology, 25 when present, the user's eye movements as they look at the image and Identifying the areas it is expected to check shortly, mobile network latency, packet By evaluating data loss and available data transfer capacity, these regions are given high priority. transmitting in high quality and lowering the quality level of image regions that the user is not looking at. by reducing bandwidth usage, the image in the focused area is 30. It is clear that a system is needed to ensure the preservation of quality. 2 United Kingdom patent number GB2640153A, which is included in the prior art. in the document, the processes of creating and encoding video images We are talking about a system that adapts to the user's perspective. The invention is a split-view client, a split-view server, display device, virtual camera and communication that enables data transmission between them 5 It includes the network. The client receives the location, orientation, and information of the virtual camera associated with it. Metadata such as the display area is split into a display server. The server sends this metadata and the perspective-based optimization profile. using it to create the video frame and transmit the created frame to the client. It codes accordingly. The gaze-based optimization profile, image generation 10 the first spatial quality map used in the process and in the video encoding process It includes the second spatial quality map used. The first spatial quality the map creates different images based on the areas the user is looking at. It allows different image qualities to be assigned to its sections. Second spatial The quality map shows that the parts of the video frame the user is looking at have higher quality, 15 This results in lower quality coding of environmental components. Optimization profile; client device operating conditions, network conditions, estimated experience quality, service quality, device capabilities, applications run, and It can be modified according to user input. Server and client, display can discuss the optimization profile to be used for the session and profile 20 The parameters can be adapted based on metadata collected during the session. Brief Description of the Invention The aim of this invention is to analyze the user's eye movements and make 25 changes on the image. Identifying the regions it has looked at and is expected to look at soon, these regions The reliability of determining this depends on latency, packet loss, and availability on the mobile network. By analyzing data transfer capacity, it identifies the image regions the user is viewing. the image it doesn't look at while transmitting in high quality and with higher transmission priority the reduction in the quality level of the regions and the delay that occurs during image transmission, 30 3 image based on re-buffering and quality loss in the focus region results. The goal is to create a system that readjusts transmission settings in terms of quality. Detailed Description of the Invention The purpose of this invention is to create "Artificial Intelligence-Based Video". The "Transmission System" is shown in the attached figure; Figure 1. Schematic view of the system described in the invention. The parts shown in the figure are individually numbered, and the corresponding numbers correspond to these numbers. given below: 1. System 2. Electronic device 15 3. Eye tracking module 4. Server By analyzing the user's eye movements and the mobile network's data transmission conditions The areas on the image that the user is looking at and is expected to look at shortly are 20. Identifying and transmitting these regions in high quality while displaying images the user isn't looking at. reducing network usage by transmitting signals to those regions at a lower quality The system in question, developed for the purpose of invention (1); - the user's augmented reality, virtual reality, mixed reality, cloud gaming or Enables viewing of high-resolution interactive video content, eye 25 at least one electronic device capable of tracking (2), - the gaze vector, representing the user's gaze direction, on the image. The fixation point refers to the eye's ability to focus on a single point; a point where the eye concentrates on an image. jerky eye movement, which refers to the rapid movement of the eye from one region to another, movement speed, separate focal points of the two eyes, and error in eye tracking measurement and 30 4 the most structured to determine stability information at specific time intervals a small eye tracking module (3) and - gaze vector, fixation point, jump eye taken from eye tracking module (3) by analyzing movement, eye movement speed and eye tracking stability data Determining the user's current and expected viewing area in the near future, prediction 5 reliability is assessed by mobile network latency, packet loss, effective data transfer speed, and service. By analyzing quality data, we can see which images the user is looking at and which they are not looking at. by assigning different quality and bit rate levels to its regions, the resulting image configured to transmit sections for display on an electronic device (2) It contains at least one server (4). 10 The electronic device (2) included in the system (1) which is the subject of the invention, augmented reality, virtual reality, mixed reality, cloud gaming, or high-resolution interactive video to display the content to the user, determined by the eye tracking module (3) gaze vector, fixation point, jerky eye movement, eye movement velocity, and eye 15 To ensure that the monitoring stability data is transmitted to the server (4) and from the server (4) The image fragments obtained at different quality levels are ROI (Region of Interest – (Region of Interest) based on mask metadata and image frame timestamps. It is structured to present to the user by combining them. Electronic device (2), Decoding delay during image transmission, image creation 20 delay, buffer occupancy, rebuffering, and foveal escape rate By identifying client telemetry and feedback data in this form, we can analyze this data. updating subsequent image transmission decisions by the server (4) It is configured to be transmitted for use. The eye tracking module (3) in the system (1) which is the subject of the invention, tracks the user’s eye representing the direction of view by tracking its movements at specific time intervals The gaze vector is the fixation vector that represents the focal point on an image. jumps refer to the rapid transition of the eye from one visual region to another. eye movement, eye movement speed, separate focal points for both eyes, and eye tracking 30 to determine the error and stability information of the measurement and the measurements with time data to generate timestamped eye movement data by correlating them is configured. The eye tracking module (3) processes the eye movement data in question. Determining the user's current gaze region, stable eye movement. fixation, jerky transition, unstable measurement, or diopter mismatch Assessment of their status, 5 that the user is expected to view shortly Predicting the region using artificial intelligence and using this data with network telemetry by evaluating the areas of the image that the user is looking at and not looking at. to transmit to the server (4) for the creation of subsequent image transmission decisions It is being structured. The server (4) in the system (1) which is the subject of the invention, any remote communication using the protocol to communicate with the electronic device (2) and the eye tracking module (3). to establish, and to exchange data through this established communication. It is configured. The server (4) receives the eye tracking module (3) at certain times. gaze vector, fixation point, jerky eye movement created at intervals, 15 eye movement speed, focal points of both eyes, and eye tracking error and stability data. to receive, decoding delay, image generation from electronic device (2) delay, buffer occupancy, rebuffering, and foveal escape rate to receive client feedback data in this form and to convert that data into an image. 20 in determining the quality and transmission conditions under which the data will be transmitted to the user It is configured to use. The server (4) receives data from the eye tracking module (3). by processing timestamped eye movement data within short time windows The user's eye movement can be stable fixation, jerky eye movement transition, or unstable. categorizing situations such as measurement or eye misalignment, past view artificial intelligence using sequence, fixation information, jerky eye movement and eye movement speed 25 Through an intelligence-based gaze prediction model, the user's brief overview of the image... to create a probability for the location that is expected to be looked at after some time and the said Numerical confidence or uncertainty indicates how reliable the prediction is. It is configured to determine the value. The server (4) determines the user's current value. The next viewing position predicted by AI based on the viewing position is 30. Converting the image to a coordinate system to create an ROI mask; this mask 6 through the fovea region, the area where the user perceives the image with high detail. Identifying peripheral regions that it is not looking at, eye tracking, and ensuring a stable prediction result. If this is the case, the foveal region should be kept narrower; if the uncertainty increases, then... to prevent the actual viewing position from remaining outside the high-quality area Expanding the ROI area to create a safety margin around the fovea or 5 to prepare multiple possible focus areas in a short period of time with high quality is structured. Server (4), confidence or uncertainty of the view estimate. its value is 5G (Fifth Generation Mobile Communications) or 6G (Sixth Generation Mobile Communications) RTT provided from the Generation – Sixth Generation Mobile Communications) network (Round-Trip Time), jitter (delay variability), packet 10 data loss, throughput (effective data transfer rate), QoS (Quality of Service). Quality) class, indirect indications of radio conditions and, if applicable, MEC (Multi- Access Edge Computing (Access Edge Computing) - network latency. to evaluate telemetry data on a common time axis and an image Network transmission 15 valid in the same time interval as the viewing status of the frame It is structured to enable the analysis of the conditions. Server (4), The video footage is displayed independently according to the ROI mask created. dividing into manageable tiles (image segments) or quality layers, fovea image segments corresponding to the region at higher resolution and lower quality. With compression, the image segments corresponding to the peripheral regions are reduced to a lower 20. encoding with resolution or higher compression and image processing in these operations QP with square timestamps, tile or layer indexes, and resolution levels. (Quantization Parameter), GOP (Group of Pictures – Video encoding parameters such as Image Group) and in-frame refresh using 25 different quality levels between the fovea and peripheral image regions It is configured to allow for independent control. Server (4), estimated view position, estimate uncertainty, current effective data transfer rate, packet loss, delay variability, QoS class and received from electronic device (2) Quality of experience by evaluating feedback data on the fovea and peripheral regions. Specifying the fovea protection level with separate bit rate budgets, 30 of the network capacity instead of reducing the overall image quality proportionally if the image quality decreases 7 By applying asymmetric resource reduction, the bit rate allocated to peripheral regions is increased faster. to reduce and allocate the bit rate of the image to the foveal region the user is looking at It is structured to maintain its quality. The server (4) corresponds to the fovea region. incoming tile or stream packets are less affected by delay and packet loss To provide this, FEC (Forward Error Correction) is applied to the packets in question. Correction), ARQ (Automatic Repeat Request), more frequent In-frame refresh, smaller image group, or higher transport priority. Determining fault tolerance and protection levels based on network conditions, The importance levels of fovea and peripheral packets in 5G / 6G network QoS transport The encoded image segments are matched with profiles and the ROI mask metadata is segment 10. by associating their identities with image frame timestamps on the electronic device (2) is configured to transmit for merging. Server (4), electronic Decoding delay, image generation delay, transmitted back from the device (2), re-buffering, bit rate saving, and the user's actual viewing position. 15 refers to the area being left outside of the region which was prepared to a high standard in a timely manner. Performing closed-loop control by evaluating the foveal leakage rate, If foveal leakage or rebuffering values ​​increase, then the following ROI width, view estimation, fovea-peripheral bit rate distribution or in cycles readjusting the protection level and artificially adjusting eye movement or network conditions This refers to the way an intelligence model differentiates over time from the conditions it has learned from. 20 When drift (data / model distribution shift) is detected, the model is recalibrated. by enabling the changing use of image quality and network usage. It is configured to be continuously updated according to the terms and conditions. Industrial Application of the Invention 25 Thanks to the system (1) which is the subject of the invention, telecommunications, augmented reality, virtual reality, mixed reality, cloud gaming, and high-resolution video streaming. in these areas, more transmission sources to the image region the user is looking at. By separating and reducing the amount of data in unmonitored environmental areas, the network 30 8 efficient use of capacity and providing the user with high quality at low latency. The presentation of the image is ensured. Based on these fundamental concepts, the invention is an "AI-Based Video Transmission System". It is possible to develop a wide variety of applications related to (1)”, and the invention here is 5 It cannot be limited to the examples given; it is essentially as stated in the claims.

Claims

9 REQUESTS 1. Analyzing the user's eye movements and the mobile network's data transmission conditions. by looking at the image the user is looking at and shortly after 5 Identifying the expected regions, delivering high-quality information to these regions. by transmitting image regions that the user is not looking at at a lower quality, the network that helps reduce its use; - augmented reality, virtual reality, mixed reality, cloud for the user games or high-resolution interactive video content at least one electronic device that enables viewing and eye tracking. device (2), - the gaze vector, representing the user's gaze direction, on the image The fixation point refers to the eye's ability to focus on a specific point. Jump eye refers to rapid movement of the image from one region to another. movement, eye movement speed, separate focal points of the two eyes, and eye tracking 15 measurement error and stability information at specific time intervals containing at least one eye-tracking module (3) configured to determine and - gaze vector, fixation point, received from the eye tracking module (3), jerky eye movement, eye movement speed, and eye tracking stability data. by analyzing the user's current and expected outlook in the near future 20 Determining the region, predicting reliability, mobile network latency, packet by analyzing data on data loss, effective data transfer speed, and quality of service. Different qualities and bit rates are applied to the image areas that the user is looking at and those that they are not looking at. electronically allocates speed levels and creates image segments. at least one 25 configured to transmit for display on the device (2). a payment system characterized by server (4) (1).

2. Augmented reality, virtual reality, mixed reality, cloud gaming, or high-resolution interactive video content to the user to display, the gaze vector determined by the eye tracking module (3), 30 fixation point, jerky eye movement, eye movement speed, and eye tracking. to ensure that the stability data is transmitted to the server (4) and from the server (4) ROI mask the received image fragments at different quality levels. By combining image frame timestamps with metadata. characterized by an electronic device (2) configured to present to the user A system like the one in Request 1 (1). 5 3. Decoding delay that occurs during image transmission, image creation. delay, buffer occupancy, rebuffering, and foveal escape rate By identifying client telemetry and feedback data in this form, data server (4) subsequent image transmission decisions 10 electronically configured to be transmitted for use in updating any of the above claims characterized by the device (2) such a system (1).

4. By tracking the user's eye movements at specific time intervals, gaze detection is performed. the gaze vector representing the direction, the position of focus on the image The point of fixation indicates the rapid movement of the eye from one visual region to another. jerky eye movement, which describes the transition, eye movement speed, bilateral separate focal points and error and stability information of eye-tracking measurement to determine and correlate measurements with time data, timestamped 20 Eye tracking module configured to generate eye movement data (3) as in any of the above claims characterized by system (1).

5. The eye movement data in question is calculated based on the user's current gaze area (25°). determination of eye movement: stable fixation, jerky transition, unstable conditions such as measurement or two-eye discrepancy the assessment of the area the user is expected to view shortly prediction using artificial intelligence and network telemetry of this data By evaluating the image, 30 areas are shown both where the user is looking and where they are not looking. (4) to the server in order to make subsequent image transmission decisions. 11 characterized by the eye-tracking module (3) configured to transmit a system like any of the above requests (1).

6. Using any remote communication protocol, electronic device (2) and communicate with the eye tracking module (3), and through this communication 5 with the server (4) configured to perform data exchange a system like any of the above characterized claims (1).

7. The gaze created at specific time intervals from the eye tracking module (3) 10 vector, fixation point, jerky eye movement, eye movement velocity, bipolar electronically obtaining eye-tracking error and stability data with respect to focal points. from the device (2) decoding delay, image generation delay, buffer occupancy, rebuffering, and foveal leakage rate to receive client feedback data and to use that data in image 15 the quality and transmission conditions under which it will be transmitted to the user characterized by the server configured for use in determining (4) a system like any of the above-mentioned requests (1).

8. Time-stamped eye movement data received from the eye tracking module (3) in short 20 By processing within time windows, it accurately tracks the user's eye movement. fixation, jerky eye movement transition, unstable measurement, or binoculars to categorize situations as incompatibility, past perspective sequence, fixation artificial intelligence that uses information, jerky eye movements, and eye movement speed through a perspective-based prediction model, the user can briefly observe the image for 25 minutes. to create a probability for the location that is expected to be looked at after some time and promise numerical confidence, which shows how reliable the prediction is. with the server (4) configured to determine the uncertainty value a system like any of the above characterized claims (1). 30 12 9. Predicted by artificial intelligence based on the user's current viewing position. Converting the next viewpoint to the image coordinate system ROI creating a mask, and through this mask, displaying the image to the user at a high level the fovea region, which it perceives in detail, and the peripheral regions it does not look at. to determine, if eye tracking and prediction result is stable, fovea 5 keeping the scope narrower, and a more realistic view if uncertainty increases. to prevent its location from remaining outside the high-quality area of ​​the fovea To expand the ROI area in a way that creates a margin of confidence around it, or High-quality rendering of multiple possible focal regions in a short period of time. The above 10 is characterized by the server (4) configured to prepare for it. a system like any of the requests (1).

10. Confidence or uncertainty value for the outlook forecast (5G or 6G) RTT, jitter, packet loss, throughput, QoS class provided by the network, 15 indirect indications of radio conditions and, if applicable, the MEC delay. to evaluate network telemetry data on a common timeline and a the network that is valid in the same time interval as the viewing position of the image frame Server configured to enable analysis of transfer conditions (4) like any of the above-mentioned claims characterized by system (1). 20 11. Video images are separated according to the generated ROI mask. dividing into manageable tile or quality layers, into the fovea region the corresponding image segments at higher resolution and lower With compression, the image segments corresponding to the peripheral regions are further reduced to 25. encoding with low resolution or higher compression and this In operations, image frame timestamps, tile or layer indexes, Video resolution levels with QP, GOP, and in-frame refresh rates. Using encoding parameters, the fovea and peripheral image regions 30 13 The above is characterized by the server (4) configured to provide a system like any of the requests (1).

12. Estimated view position, forecast uncertainty, current effective data transfer. speed, packet loss, latency variability, QoS class, and 5 from the electronic device (2) by evaluating the quality of experience feedback data received, fovea and Foveal protection level with separate bit rate budgets for peripheral regions. to determine if the entire image will be displayed in case of reduced network capacity. Asymmetric resource reduction instead of reducing quality proportionally. by implementing it, the bit rate allocated to peripheral regions can be reduced more quickly and 10 image quality is determined by the bit rate allocated to the fovea region that the user is looking at. the above characterized by the server (4) configured to protect a system like any of the requests (1).

13. Delay and packet 15 of tile or stream packets corresponding to the fovea region. To ensure that it is less affected by the loss, the FEC included these packages in its plans. ARQ requires more frequent intra-frame refreshes, smaller image groups, or more high priority transmission fault tolerance and protection levels in the network to determine the importance levels of foveal and peripheral packets according to the conditions Image 20 mapped and encoded with the QoS transport profiles of the 5G / 6G network. The components are represented by ROI mask metadata, component IDs, and image frame time. to be combined in the electronic device (2) by associating with their stamps The above is characterized by the server (4) configured to transmit the above. a system like any of the requests (1).

14. Decoding delay transmitted back from the electronic device (2), image render delay, re-buffering, bitrate saving, and user the actual point of view prepared in a timely and high-quality manner foveal abruption rate, which refers to remaining outside the region Performing closed-loop control by evaluating, foveal evacuation 30 or in subsequent cycles if buffering values ​​are increased again 14 ROI width, view estimation, fovea-peripheral bit rate distribution, or readjusting the protection level and eye movement or mesh conditions over time, from the conditions that the artificial intelligence model has learned When drift, which indicates a difference in performance, is detected, the model is recalibrated. by enabling image quality and network usage to change. 5 Server configured to be continuously updated according to its conditions (4) like any of the above-mentioned claims characterized by system (1). 15 25