Browsing area determination method and device based on eye movement tracking, equipment and medium
By processing infrared images and visual offset models of target users and using eye-tracking technology, the problem of low accuracy of browsing tracking points in large-screen mobile devices has been solved, enabling accurate acquisition of the user's current focus area and improving the accuracy of browsing tracking points.
Patent Information
- Application Number
- CN202511017027.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-23
- Publication Date
- 2025-10-31
AI Technical Summary
In existing technologies, the accuracy of browsing tracking data on large-screen mobile devices is relatively low, as it cannot accurately obtain the area currently focused by the user, resulting in reduced accuracy of browsing tracking data.
By acquiring infrared images of the target user and the device's visual offset model, data processing is performed using preset image processing rules and visual coordinate axis establishment rules. Combined with eye-tracking technology, visual focus is calculated, thereby determining the user's browsing area.
It enables accurate acquisition of the user's current focus area on large-screen mobile devices, improving the accuracy of browsing tracking.
Smart Images

Figure CN120872155A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method, apparatus, device, and medium for determining browsing areas based on eye tracking. Background Technology
[0002] With the increasing intelligence of mobile devices, screen sizes are getting larger and larger, and more and more content is being displayed on these large-screen devices. Therefore, identifying the user's browsing area on the screen to improve product recommendation conversion rates has become crucial.
[0003] In existing technologies, the browsing area of a user on a small screen is typically determined by pre-setting browsing events. Specifically, each browsing area of a small-screen mobile device is assigned a number. Then, since all areas on the screen are within the user's field of vision, the numbered area where information is displayed can be considered the browsing area.
[0004] However, because the area displayed on a large-screen mobile device may not be the area the user is actually focusing on, the user's gaze may not be fixed on that area. If the user's browsing area is determined by pre-setting browsing events as in existing technologies, the accuracy of browsing tracking will be reduced. Therefore, how to accurately obtain the area the user is currently focusing on and improve the accuracy of browsing tracking is a problem that urgently needs to be solved. Summary of the Invention
[0005] This invention provides a method, apparatus, device, and medium for determining browsing areas based on eye tracking, which can solve the problem of low accuracy of browsing tracking points in large-screen mobile devices.
[0006] According to one aspect of the present invention, a method for determining a browsing area based on eye tracking is provided, comprising:
[0007] Obtain the target infrared image and the target device visual offset model corresponding to the target user; wherein, the target device visual offset model includes the visual focus offset between the target user and each target point in the target terminal device;
[0008] The target infrared image is processed based on preset image processing rules and preset visual coordinate axis establishment rules to determine the visual coordinate axis model set corresponding to the target infrared image; wherein, the visual coordinate axis model set includes a first visual coordinate axis model and a second visual coordinate axis model;
[0009] Based on preset eye-tracking technology, focus calculation is performed on the first visual coordinate axis model and the second visual coordinate axis model to determine the target visual focus corresponding to the target infrared image;
[0010] Based on the target device visual offset model, the target visual focus is located to determine the target browsing area corresponding to the target infrared image.
[0011] According to another aspect of the present invention, a browsing area determination device based on eye tracking is provided, comprising:
[0012] The data acquisition module is used to acquire the target infrared image and the target device visual offset model corresponding to the target user; wherein, the target device visual offset model includes the visual focus offset between the target user and each target point in the target terminal device;
[0013] The data processing module is used to process the target infrared image based on preset image processing rules and preset visual coordinate axis establishment rules, and determine the visual coordinate axis model set corresponding to the target infrared image; wherein, the visual coordinate axis model set includes a first visual coordinate axis model and a second visual coordinate axis model;
[0014] The focus calculation module is used to perform focus calculation on the first visual coordinate axis model and the second visual coordinate axis model based on preset eye-tracking technology to determine the target visual focus corresponding to the target infrared image;
[0015] The region determination module is used to locate the visual focus of the target based on the visual offset model of the target device, and determine the target browsing area corresponding to the infrared image of the target.
[0016] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:
[0017] At least one processor; and
[0018] A memory communicatively connected to the at least one processor; wherein,
[0019] The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the eye-tracking-based browsing area determination method according to any embodiment of the present invention.
[0020] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the eye-tracking-based browsing area determination method according to any embodiment of the present invention.
[0021] According to another aspect of the present invention, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the eye-tracking-based browsing area determination method described in any embodiment of the present invention.
[0022] The technical solution of this invention processes the target infrared image corresponding to the target user using preset image processing rules and preset visual coordinate axis establishment rules to determine a set of visual coordinate axis models corresponding to the target infrared image, including a first visual coordinate axis model and a second visual coordinate axis model. Then, based on preset eye-tracking technology, focus calculation is performed on the first and second visual coordinate axis models to determine the target visual focus corresponding to the target infrared image. Finally, based on the target device visual offset model, the target visual focus is located to determine the target browsing area corresponding to the target infrared image. Because the dynamic calculation of the user's gaze focus is achieved through relevant algorithms, the problem of low accuracy in browsing tracking on large-screen mobile devices is solved, accurately obtaining the area currently focused by the user and improving the accuracy of browsing tracking.
[0023] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0024] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0025] Figure 1 This is a flowchart of a browsing area determination method based on eye tracking according to Embodiment 1 of the present invention;
[0026] Figure 2 This is a flowchart of a browsing area determination method based on eye tracking according to Embodiment 2 of the present invention;
[0027] Figure 3 This is a flowchart of an optional eye-tracking-based browsing area determination method according to Embodiment 2 of the present invention;
[0028] Figure 4 This is a schematic diagram of a target device visual offset model construction process according to Embodiment 2 of the present invention;
[0029] Figure 5This is a schematic diagram of a visual coordinate axis model construction process provided in Embodiment 2 of the present invention;
[0030] Figure 6 This is a schematic diagram of a browsing area determination device based on eye tracking according to Embodiment 3 of the present invention;
[0031] Figure 7 This is a schematic diagram of the structure of an electronic device that implements the eye-tracking-based browsing area determination method according to an embodiment of the present invention. Detailed Implementation
[0032] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0033] It should be noted that the terms "first," "second," "target," "initial," etc., used in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0034] In the technical solution of this application, the information collected is information and data authorized by the user or fully authorized by all parties. The collection, storage, use, processing, transmission, provision, disclosure and application of the relevant data all comply with the relevant laws, regulations and standards of the relevant countries and regions, and necessary confidentiality measures have been taken. The data does not violate public order and good morals, and corresponding operation entry points are provided for users to choose to authorize or refuse. If the user chooses to refuse, the process will proceed to the expert decision-making process.
[0035] Example 1
[0036] Figure 1This is a flowchart of a browsing area determination method based on eye tracking, provided in Embodiment 1 of the present invention. This embodiment is applicable to situations where browsing data is embedded in large-screen mobile devices. The method can be executed by an eye-tracking-based browsing area determination device, which can be implemented in hardware and / or software. This eye-tracking-based browsing area determination device can be configured in an electronic device, for example, in a mobile terminal. Figure 1 As shown, the method includes:
[0037] S110. Obtain the target infrared image and the target device visual offset model corresponding to the target user; wherein, the target device visual offset model includes the visual focus offset between the target user and each target point in the target terminal device.
[0038] In this context, "target user" refers to a user who needs to determine the visual browsing area on a mobile device. Typically, user identification can be determined based on their needs. For example, if a user clicks a button or icon related to browsing area determination on their mobile device, that user is identified as a target user. "Infrared image" refers to a facial image of the user captured by an infrared camera. Typically, an infrared image is a frame image. It is worth noting that in this embodiment of the invention, the mobile device used by the target user needs to be equipped with an infrared camera. "Target infrared image" refers to the infrared image of the currently selected browsing area for determination corresponding to the target user. "Terminal device" refers to a large-screen mobile device currently being used by the user. "Target terminal device" refers to the terminal device corresponding to the target user.
[0039] The target point can refer to any pre-defined adjustment position on the screen of the target terminal device. For example, the target point can be a glowing icon or a brightly colored icon. The visual focus offset refers to the numerical offset between the user's actual visual focus and the target point's coordinates when looking at it. For example, the visual focus offset can be the difference between the actual visual focus and the target point's coordinates. The device visual offset model can refer to a focus model that includes all visual focus offsets and their corresponding target point coordinates. Typically, one user corresponds to one device visual offset model on one terminal device. The target device visual offset model can refer to the device visual offset model corresponding to the target user on the target terminal device.
[0040] In an optional implementation, acquiring the target infrared image corresponding to the target user may include:
[0041] Step a1: Obtain the basic infrared image sequence corresponding to the target user; wherein, the basic infrared image sequence includes candidate infrared images.
[0042] Here, the basic infrared image refers to the infrared image initially captured at the current moment corresponding to the target user. For example, after the target user logs in to the target terminal device, the infrared camera in the target terminal device can capture images of the target user at a pre-set shooting frequency to obtain the user's facial image. The basic infrared image sequence refers to a frame sequence composed of various basic infrared images captured at the same time. The candidate infrared image refers to the basic infrared image initially selected from the basic infrared image sequence according to set filtering conditions. For example, the filtering condition could be to use the basic infrared image at the 100th millisecond of the basic infrared image sequence as the candidate infrared image.
[0043] Step a2: Perform pupil recognition on the basic infrared image sequence based on preset focus recognition rules to determine the pupil focus change information corresponding to the basic infrared image sequence.
[0044] The preset focus recognition rule can refer to a pre-defined algorithm used for pupil recognition in each basic infrared image in a basic infrared image sequence. For example, the preset focus recognition rule could be a pupil detection method based on thresholds and contours. Pupil recognition can refer to the operation of determining the characteristics of the human pupil through detection and analysis. For example, pupil recognition can determine the coordinates of the user's pupil center in an infrared image. Pupil focus change information can refer to the changes in the pupil center coordinates in a basic infrared image sequence. For example, pupil focus change information can be generated by comparing the pupil recognition results of each basic infrared image in the basic infrared image sequence.
[0045] Step a3: If the pupil focus change information satisfies the preset focus change rule, then the candidate infrared image is taken as the target infrared image corresponding to the target user.
[0046] The preset focus change rule can refer to a pre-defined rule for evaluating pupil focus change information. For example, the preset focus change rule could be: if the pupil focus change information is stable, then the candidate infrared image is selected for subsequent operations.
[0047] Specifically, after the target user logs in and verifies their identity on the target terminal device, the infrared camera on the target terminal device first captures images of the target user at a pre-set shooting frequency, obtaining a basic infrared image sequence corresponding to the target user. Then, using a preset focus recognition rule, pupil recognition is performed on the basic infrared image sequence to obtain the pupil center coordinates corresponding to all basic infrared images in the sequence. The pupil center coordinates of all basic infrared images are compared according to the shooting order to determine the pupil focus change information corresponding to the basic infrared image sequence. Further, the preset focus change rule is used to rule-basedly judge the pupil focus change information. If the pupil focus change information tends to stabilize, the candidate infrared image is selected as the target infrared image corresponding to the target user. Therefore, by selecting the final image data for data processing based on the pupil focus change in a continuous frame sequence, the accuracy of subsequent data processing can be guaranteed.
[0048] S120. Based on preset image processing rules and preset visual coordinate axis establishment rules, perform data processing on the target infrared image to determine the visual coordinate axis model set corresponding to the target infrared image; wherein, the visual coordinate axis model set includes a first visual coordinate axis model and a second visual coordinate axis model.
[0049] The preset image processing rules refer to pre-defined rules that define the feature point recognition process in a target infrared image. For example, the preset image processing rules could involve feature point recognition using an open-source vision library. Typically, the preset image processing rules can determine the feature points corresponding to the face, eyes, and pupils in the target infrared image. The visual coordinate axis model refers to a mathematical framework used to digitally represent the eye structure and gaze direction. Typically, the visual coordinate axis model can be a local coordinate axis model between the eye and the pupil.
[0050] The preset visual coordinate axis establishment rules refer to pre-defined rules that define the construction process of the visual coordinate axis model. For example, the preset visual coordinate axis establishment rules may include the construction object and the specific construction process. The first visual coordinate axis model may refer to the visual coordinate axis model corresponding to the left eye region in the target infrared image. The second visual coordinate axis model may refer to the visual coordinate axis model corresponding to the right eye region in the target infrared image. The visual coordinate axis model set may refer to the set of various visual coordinate axis models corresponding to the same target infrared image.
[0051] S130. Based on preset eye-tracking technology, focus calculation is performed on the first visual coordinate axis model and the second visual coordinate axis model to determine the target visual focus corresponding to the target infrared image.
[0052] Among these, preset eye-tracking technology refers to pre-defined algorithms used to dynamically calculate the user's gaze focus. Visual focus refers to the spatial location where the gaze is actually focused. Typically, visual focus can be the spatial location where the left and right eyes of the same user are actually focused. Target visual focus refers to the visual focus corresponding to the target's infrared image.
[0053] S140. Based on the target device visual offset model, the target visual focus is located to determine the target browsing area corresponding to the target infrared image.
[0054] In this context, "regional positioning" refers to the operation of determining the browsing area on a target terminal device based on the target's visual focus and the target device's visual offset model. The browsing area can refer to the region where the user's gaze is currently focused. For example, the browsing area could be a piece of text or a button on the target terminal device's screen. Alternatively, the target browsing area can refer to the browsing area determined based on the target's infrared image.
[0055] In an optional implementation, after determining the target browsing area corresponding to the target infrared image by locating the target visual focus based on the target device's visual offset model, the method may further include: acquiring basic attribute parameters corresponding to the target browsing area; generating browsing behavior data corresponding to the target user based on the target browsing area and the basic attribute parameters; summarizing the browsing behavior data corresponding to the target user based on a preset time threshold to generate a browsing behavior data set corresponding to the target user; and sending the browsing behavior data set to a server so that the server can perform data analysis on the browsing behavior data set, generate a browsing hotspot corresponding to the target user, and perform corresponding push operations based on the browsing hotspot. The basic attribute parameters may refer to attribute information matched to the target browsing area. For example, the basic attribute parameters may include the target user's username and the acquisition time corresponding to the target infrared image. The browsing behavior data may refer to all recordable visual attention information generated during the user's interaction with the digital interface. For example, the browsing behavior data may include basic attribute parameters and the area coordinates of the target browsing area. The preset time threshold may refer to a pre-set time value used to limit the data summarization time. For example, the preset time threshold can be 24 hours. The browsing behavior data set can refer to the collection of browsing behavior data corresponding to each target user within the preset time threshold. Typically, one target user corresponds to one browsing behavior data set within one preset time threshold. The server can refer to the selected computer terminal for user data analysis. Browsing hotspots can refer to the browsing areas that appear most frequently in the browsing behavior data set.
[0056] Specifically, after generating the target browsing area corresponding to the target infrared image, the basic attribute parameters of the target browsing area can be integrated with the target browsing area information through the target terminal device, forming a browsing behavior data point. Then, all browsing behavior data corresponding to the target user within the same preset time threshold are aggregated to generate a browsing behavior data set for the target user, which is then sent to the server. Finally, the server analyzes the browsing behavior data set to determine the browsing area with the highest browsing frequency as the target user's browsing hotspot, and configures key recommended content within the browsing hotspot. Thus, by aggregating the target user's browsing behavior data within a preset time threshold to generate a browsing behavior data set for the target user, and enabling the server to analyze this data set and generate browsing hotspots for the target user, precise operational targets can be provided to each user based on their browsing habits, achieving personalized page operations.
[0057] The technical solution of this invention processes the target infrared image corresponding to the target user using preset image processing rules and preset visual coordinate axis establishment rules to determine a set of visual coordinate axis models corresponding to the target infrared image, including a first visual coordinate axis model and a second visual coordinate axis model. Then, based on preset eye-tracking technology, focus calculation is performed on the first and second visual coordinate axis models to determine the target visual focus corresponding to the target infrared image. Finally, based on the target device visual offset model, the target visual focus is located to determine the target browsing area corresponding to the target infrared image. Because the dynamic calculation of the user's gaze focus is achieved through relevant algorithms, the problem of low accuracy in browsing tracking on large-screen mobile devices is solved, accurately obtaining the area currently focused by the user and improving the accuracy of browsing tracking.
[0058] Example 2
[0059] Figure 2 This is a flowchart of a browsing area determination method based on eye tracking provided in Embodiment 2 of the present invention. This embodiment is a refinement based on the above embodiment. Specifically, this embodiment refines the operation of processing the target infrared image based on preset image processing rules and preset visual coordinate axis establishment rules to determine the visual coordinate axis model set corresponding to the target infrared image. Figure 2 As shown, the method includes:
[0060] S210. Obtain the target infrared image and the target device visual offset model corresponding to the target user; wherein, the target device visual offset model includes the visual focus offset between the target user and each target point in the target terminal device.
[0061] In an optional implementation, before acquiring the target infrared image and the visual offset model of the target device corresponding to the target user, the method may further include: acquiring a function trigger command corresponding to the target user, and determining an initial infrared image of the target guidance point corresponding to the target user based on the function trigger command and a preset set of guidance points; wherein the preset set of guidance points includes each target guidance point. Data processing is performed on the initial infrared image based on preset image processing rules and preset visual coordinate axis establishment rules to determine a set of visual coordinate axis models corresponding to the initial infrared image. Based on the target guidance points and the set of visual coordinate axis models, the focus deviation value of the target guidance point corresponding to the target user is determined, and the focus deviation values of each target guidance point in the preset set of guidance points corresponding to the target user are summarized and processed as the visual offset model of the target device corresponding to the target user.
[0062] The function trigger instruction can refer to the instruction that the target user triggers the function of determining the browsing area. For example, the function trigger instruction can be the instruction that the target user clicks the icon or button corresponding to the function of determining the browsing area on the target terminal device. Typically, the function trigger instruction only needs to appear once. The guide point can refer to the location used to guide the user's gaze. For example, the guide point can be a glowing icon or a brightly colored icon, etc. Typically, only one guide point appears at a time, and the guide point can guide the user's gaze to that guide point, achieving gaze shift. The preset guide point can refer to the guide point preset in the target terminal device. For example, the preset guide point must at least include the four corners and the center of the screen. The preset guide point set can refer to the set of preset guide points corresponding to the same target terminal device. Generally, the more preset guide points included in the preset guide point set, the more accurate the final browsing area position; this embodiment of the invention does not specifically limit this. The target guide point can refer to the preset guide point selected from the preset guide point set for gaze guidance. For example, the target guidance point can be any one of the preset guidance points in the preset guidance point set. The initial infrared image can refer to the infrared image of the target user based on the target guidance point. The focus deviation value can refer to the deviation value between the actual visual focus corresponding to the visual coordinate axis model set and the coordinates of the target guidance point. The actual visual focus can refer to the user's visual focus on the screen based on the target guidance point.
[0063] Specifically, before acquiring the target infrared image and visual offset model of the target device corresponding to the target user, the function trigger command corresponding to the target user can be obtained first. Then, the browsing area determination function corresponding to this function trigger command accesses the operating system of the target terminal device through a hardware interface, prompting the user through the operating system to "authorize the camera to take pictures." Further, after the user authorizes the camera's access control permissions, the infrared camera takes pictures of the target user based on a pre-set shooting frequency, obtaining an initial infrared image at the target guidance point. Then, the initial infrared image is processed using pre-set image processing rules and pre-set visual coordinate axis establishment rules to determine the visual coordinate axis model set corresponding to the initial infrared image. Finally, using the visual coordinate axis model set corresponding to the target guidance point, the focus deviation value at that target guidance point is determined, and the coordinates of each focus deviation value and the corresponding target guidance point are summarized to form the target device visual offset model corresponding to the target user. Thus, by determining the target user's gaze offset at the pre-set guidance position and constructing a user gaze focus model, a valid basis can be provided for subsequently determining the target user's browsing area.
[0064] S220. Based on preset image processing rules, feature point recognition is performed on the target infrared image to determine the set of feature regions corresponding to the target infrared image.
[0065] The feature region can refer to the feature point recognition area corresponding to the infrared image of the target. For example, the feature region can include a facial rectangular region, a left-eye rectangular region, a right-eye rectangular region, and a pupil center region. The facial rectangular region can include the coordinates of the top-left corner of the facial rectangle, the width and height of the rectangle, etc. The left-eye rectangular region can include the coordinates of the top-left corner of the left-eye rectangle, the width and height of the rectangle, etc. The right-eye rectangular region can include the coordinates of the top-left corner of the right-eye rectangle, the width and height of the rectangle, etc. The pupil center region can include the coordinates of the pupil center and its radius, etc. The feature region set can refer to the collection of all feature regions in the infrared image of the same target.
[0066] S230. Determine a set of local feature images corresponding to the target infrared image based on the target feature region in the set of feature regions; wherein, the set of local feature images includes a first local feature image and a second local feature image corresponding to the same target infrared image.
[0067] The target feature region can refer to the selected area within the feature region set for establishing visual coordinate axes. For example, the target feature region can be a rectangular region for the left eye or a rectangular region for the right eye. The local feature image can refer to a local visual image cropped from the target feature region. For example, if the target feature region is a rectangular region for the left eye, then the local feature image can be the left eye feature image. The first local feature image can refer to the local feature image corresponding to the left eye. The second local feature image can refer to the local feature image corresponding to the right eye. The set of local feature images can refer to the collection of various local feature images within the same target infrared image.
[0068] S240. Based on the preset visual coordinate axis establishment rules, visual coordinate axes are established for the target local feature images in the local feature image set, generating the target visual coordinate axis model corresponding to the target local feature image, and summarizing and processing the target visual coordinate axis models corresponding to the target local feature images in the local feature image set as the visual coordinate axis model set corresponding to the target infrared image.
[0069] The target local feature image can refer to the selected local feature image from the set of local feature images used to establish the visual coordinate axis. For example, the target local feature image can be either a first local feature image or a second local feature image. The target visual coordinate axis model can refer to the visual coordinate axis model constructed based on the target local feature image.
[0070] Specifically, after obtaining the target infrared image corresponding to the target user, feature point recognition can be performed on the target infrared image using preset image processing rules to determine the set of feature regions corresponding to the target infrared image. Then, the target feature regions in the feature region set are cropped to obtain a set of local feature images corresponding to the target infrared image. Further, using preset visual coordinate axis establishment rules, visual coordinate axes are established for each target local feature image in the local feature image set, generating a target visual coordinate axis model corresponding to the target local feature image. Thus, by extracting feature points from the infrared image of the target user and using these feature points to construct a visual model, the global feature image can be converted into a local coordinate axis model, providing a valid foundation for subsequent operations.
[0071] In an optional implementation, visual coordinate axes are established for the target local feature images in the local feature image set based on preset visual coordinate axis establishment rules, generating a target visual coordinate axis model corresponding to the target local feature images, including:
[0072] Step b1: Calculate the contour points of the target local feature image based on the preset center point calculation rules to determine the geometric center point corresponding to the target local feature image.
[0073] Here, contour points can refer to a set of discrete points used to describe the shape boundaries in a local feature image of a target. Preset center point calculation rules can refer to pre-defined rules that define the center point calculation process for the local feature image of the target. For example, the preset center point calculation rule could be to calculate the average of all contour points in the local feature image of the target. Geometric center point can refer to the geometric equilibrium point of the local feature image of the target in space. Typically, geometric center points can be represented using coordinates.
[0074] Step b2: Based on the horizontal span between target contour points in the target local feature image, determine the horizontal axis direction corresponding to the target local feature image.
[0075] Here, the target contour point can refer to the leftmost or rightmost point in the local feature image of the target. The horizontal span can refer to the connection direction between two target contour points in the same local feature image of the target. The horizontal axis direction can refer to the direction represented by the horizontal axis in the visual coordinate axis model.
[0076] Step b3: Based on the angle of the line connecting the target contour points in the target local feature image, determine the tilt angle corresponding to the target local feature image.
[0077] The connecting angle can refer to the angle between the lines connecting two target contour points in the local feature image of the same target. The tilt angle can refer to the rotation angle in the visual coordinate axis model.
[0078] Step b4: Based on the preset visual coordinate axis template, summarize the parameters of the geometric center point, horizontal axis direction and tilt angle to generate the basic visual coordinate axis model corresponding to the target local feature image.
[0079] The preset visual coordinate axis template refers to a pre-defined template for constructing a visual coordinate axis model. For example, the preset visual coordinate axis template can be a model constructed with the geometric center point as the origin, the vertical direction of the horizontal axis as the vertical axis, and the tilt angle as the tilt angle. The basic visual coordinate axis model refers to the visual coordinate axis model obtained by summarizing the parameters of the geometric center point, horizontal axis direction, and tilt angle according to the preset visual coordinate axis template. Typically, one local feature image corresponds to one basic visual coordinate axis model; that is, the left eye and the right eye each correspond to one basic visual coordinate axis model.
[0080] Step b5: Based on the basic visual coordinate axis model, perform coordinate transformation on the target pupil center corresponding to the target local feature image, generate coordinate transformation result, and merge the coordinate transformation result into the basic visual coordinate axis model to generate the target visual coordinate axis model corresponding to the target local feature image.
[0081] Here, the target pupil center can refer to the pupil center coordinates corresponding to the local feature image of the target. Typically, the pupil center region can be determined by the set of feature regions corresponding to the target infrared image, and the pupil center coordinates within this region can be extracted as the target pupil center. The coordinate transformation result can refer to the transformation result obtained after performing a coordinate transformation on the target pupil center using a basic visual coordinate axis model. For example, a rotation matrix can be used to transform the target pupil center from global coordinates to local coordinates.
[0082] Specifically, after obtaining the set of local feature images corresponding to the target infrared image, the target local feature image in the set can be selected first. Contour points are calculated for the target local feature image using a preset center point calculation rule, serving as the geometric center point corresponding to the target local feature image. Then, the horizontal span between the target contour points in the target local feature image is determined as the horizontal axis direction corresponding to the target local feature image, and the angle between the lines connecting the target contour points in the target local feature image is determined as the tilt angle corresponding to the target local feature image. Further, a preset visual coordinate axis template is used to summarize the parameters of the geometric center point, horizontal axis direction, and tilt angle, generating a basic visual coordinate axis model corresponding to the target local feature image. Finally, the basic visual coordinate axis model is used to perform coordinate transformation on the target pupil center to generate a coordinate transformation result, which is then merged into the basic visual coordinate axis model to generate the target visual coordinate axis model. Thus, by constructing visual coordinate axis models for the left and right eyes respectively using local feature images, a valid foundation can be provided for subsequent focus calculation.
[0083] S250. Determine the first pupil offset corresponding to the first visual region based on the first visual coordinate axis model, and determine the second pupil offset corresponding to the second visual region based on the second visual coordinate axis model.
[0084] The visual region can refer to the local feature region corresponding to the infrared image of the target. For example, the visual region can be the left eye region or the right eye region. The first visual region can refer to the visual region corresponding to the left eye. The second visual region can refer to the visual region corresponding to the right eye. Pupil offset can refer to the offset coordinates of the pupil center relative to its centered state. Typically, the horizontal axis represents horizontal offset, affecting left-right vision; the vertical axis represents vertical offset, affecting up-down vision. For example, if the pupil offset is (0,0), it indicates that the pupil is centered. The first pupil offset can refer to the left eye pupil offset. Typically, the coordinate transformation result in the first visual coordinate axis model can be used as the first pupil offset. The second pupil offset can refer to the right eye pupil offset. Typically, the coordinate transformation result in the second visual coordinate axis model can be used as the second pupil offset.
[0085] S260. Based on a preset first gaze direction vector, the coordinate transformation of the first pupil offset is performed to generate a first global gaze corresponding to the first visual region, and based on a preset second gaze direction vector, the coordinate transformation of the second pupil offset is performed to generate a second global gaze corresponding to the second visual region.
[0086] Here, the gaze direction vector refers to the vector used to represent the visual direction of the eye. Typically, this can be represented using the horizontal deflection angle θ and the vertical deflection angle. Construct a gaze direction vector V. The preset first gaze direction vector can refer to the first visual region, i.e., the gaze direction vector corresponding to the left eye region. For example, taking the first pupil offset as (pupil_left_x, pupil_left_y), the preset first gaze direction vector can be expressed as: The horizontal deflection angle of the left eye can be expressed as: θ left =k x The left eye vertical deflection angle, `pupil_left_x`, can be represented as: The preset second gaze direction vector can refer to the second visual region, that is, the gaze direction vector corresponding to the right eye region. For example, taking the second pupil offset as (pupil_right_x, pupil_right_y), the preset second gaze direction vector can be expressed as: The horizontal deflection angle of the right eye can be expressed as: θ right =k x ·pupil_right_x, the vertical deflection angle of the right eye is: It is worth noting that k x It can represent the horizontal scaling factor, k y It can represent the scaling factor in the vertical direction. Typically, k... x and k y Pre-calibration is required, k x and k y The specific differences may exist, but this embodiment of the invention does not impose any specific limitations on this.
[0087] The global line of sight can refer to the line of sight at the global level, which comprehensively considers the head pose in the infrared image. For example, the global line of sight can be constructed using the head rotation matrix R corresponding to the target infrared image. Typically, the head rotation matrix R can be calculated by estimating facial key points in the facial rectangular region of the feature region set corresponding to the target infrared image. The first global line of sight can refer to the first visual region, i.e., the global line of sight corresponding to the left eye region. For example, the first global line of sight can be represented as: The second global gaze can refer to the second visual region, that is, the global gaze corresponding to the right eye region. For example, the second global gaze can be represented as:
[0088] S270. Based on a preset parameterized linear equation, focus calculation is performed on the first global line of sight and the second global line of sight to generate the target visual focus corresponding to the target infrared image.
[0089] The preset parametric line equation can refer to a pre-defined equation used to parametrically represent the first or second global line of sight. Typically, the preset parametric line equation can include a left-eye parametric line equation and a right-eye parametric line equation. For example, with the left-eye position as P... left The right eye position is P. right For example, where the positions of the left and right eyes are both in coordinate form, the parametric line equation for the left eye can be expressed as: The parametric equation of the right eye line can be expressed as: Typically, in an infrared image of the same target, the left and right eye lines of a user looking at the same target should be consistent, i.e., r left (t)=r right (s). Therefore, through the... Linear equations, by performing scalar equation decomposition and mixed product calculation, can yield the form: The target visual focus. Here, t and s are solved from a system of scalar equations.
[0090] Specifically, after obtaining the visual coordinate axis model set corresponding to the target infrared image, the first pupil offset corresponding to the first visual region and the second pupil offset corresponding to the second visual region can be determined using the visual coordinate axis model set. Then, using a pre-defined gaze direction vector, coordinate transformations are performed on the first and second pupil offsets to generate the first and second global gaze directions. Finally, focus calculation is performed on the first and second global gaze directions to obtain the target visual focus. Thus, by using relevant algorithms to dynamically calculate the user's gaze focus, the user's gaze focus can be obtained. This accurately identifies the area currently focused by the user, improving the accuracy of browsing tracking. Simultaneously, it avoids the high hardware cost of eye-tracking technology and the inability to integrate eye-tracking technology into mobile devices such as smartphones.
[0091] S280. Based on the target device visual offset model, the target visual focus is located to determine the target browsing area corresponding to the target infrared image.
[0092] Specifically, after obtaining the target user's visual offset model and target visual focus for the target terminal device, the target visual focus can be projected onto the target terminal device to obtain the screen visual focus. Then, a set of relationships containing the target point coordinates and corresponding visual focus offsets is selected in the target device visual offset model. This visual focus offset is added to the screen visual focus to obtain the mapped visual focus. It is then determined whether the mapped visual focus and the target point coordinates meet a preset coordinate threshold. If so, the area corresponding to the target point can be used as the target browsing area; otherwise, another set of relationships in the target device visual offset model is obtained, and the above operation is repeated until the target browsing area is obtained or all relationships in the target device visual offset model have been matched.
[0093] The technical solution of this invention involves identifying feature points in a target infrared image corresponding to a target user using preset image processing rules to determine a set of feature regions corresponding to the target infrared image. Based on the target feature regions in the feature region set, a set of local feature images corresponding to the target infrared image is then determined. Next, visual coordinate axes are established for the target local feature images in the local feature image set based on preset visual coordinate axis establishment rules, generating a target visual coordinate axis model corresponding to the target local feature images. These target visual coordinate axis models are then aggregated and processed to form a set of visual coordinate axis models corresponding to the target infrared image. Further, a first pupil offset corresponding to a first visual region is determined based on a first visual coordinate axis model in the set of visual coordinate axis models, and a second pupil offset corresponding to a second visual region is determined based on a second visual coordinate axis model in the set of visual coordinate axis models. Further, the first pupil offset is coordinate-transformed based on a preset first gaze direction vector to generate a first global gaze corresponding to the first visual region, and the second pupil offset is coordinate-transformed based on a preset second gaze direction vector to generate a second global gaze corresponding to the second visual region. Finally, based on a preset parametric linear equation, the focus of the first and second global lines of sight is calculated to generate the target visual focus corresponding to the target infrared image. Then, based on the target device visual offset model, the target visual focus is located to determine the target browsing area corresponding to the target infrared image. By using relevant algorithms to dynamically calculate the user's gaze focus, the accuracy of browsing tracking points in large-screen mobile devices is solved, accurately capturing the user's current focused area and improving the accuracy of browsing tracking points.
[0094] Figure 3The diagram shows a flowchart of an optional eye-tracking-based browsing area determination method provided by an embodiment of the present invention. Specifically, firstly, after a target user logs in to a target terminal device, the infrared camera in the target terminal device captures images of the target user at a pre-set shooting frequency, obtaining a target infrared image corresponding to the target user. Then, the target infrared image is processed using preset image processing rules and preset visual coordinate axis establishment rules to determine the visual coordinate axis model set corresponding to the target infrared image. Furthermore, preset eye-tracking technology is used to calculate the focus of the first and second visual coordinate axis models in the visual coordinate axis model set, determining the target visual focus corresponding to the target infrared image. Further, a pre-constructed target device visual offset model is used to locate the target visual focus, determining the target browsing area corresponding to the target infrared image. Finally, browsing behavior data corresponding to the target user is generated based on the target browsing area and its corresponding basic attribute parameters. This browsing behavior data is then aggregated according to a preset time threshold to generate a browsing behavior data set corresponding to the target user, which is then sent to the server. This completes the entire process of the eye-tracking-based browsing area determination method.
[0095] Figure 4 The diagram illustrates a process for constructing a visual offset model for a target device according to an embodiment of the present invention. Specifically, firstly, the target user triggers a browsing area determination function on the target terminal device, generating a function trigger command. Then, the operating system of the target terminal device is accessed via a hardware interface. After the user authorizes access control permissions for the camera, the infrared camera captures images of the target user based on a pre-set shooting frequency. Simultaneously, the user's line of sight is guided to change sequentially according to a pre-set set of guidance points, resulting in an initial infrared image of the target user at the target guidance point. Further, the initial infrared image is processed using pre-set image processing rules and pre-set visual coordinate axis establishment rules to determine the visual coordinate axis model set corresponding to the initial infrared image. Finally, using the visual coordinate axis model set corresponding to the target guidance point, the focus deviation value at that target guidance point is determined, and the coordinates of each focus deviation value and the corresponding target guidance point are summarized to form a visual offset model for the target device corresponding to the target user.
[0096] Figure 5The diagram illustrates a visual coordinate axis model construction process according to an embodiment of the present invention. Specifically, firstly, feature point recognition is performed on the target infrared image using preset image processing rules to determine the set of feature regions corresponding to the target infrared image. Based on the target feature regions in the feature region set, a set of local feature images, including a first local feature image and a second local feature image, is determined corresponding to the target infrared image. Next, a target local feature image is selected from the set of local feature images. Contour points are calculated on the target local feature image using preset center point calculation rules to determine the geometric center point corresponding to the target local feature image. The horizontal span between the target contour points in the target local feature image is used to determine the horizontal axis direction corresponding to the target local feature image. The tilt angle corresponding to the target local feature image is determined using the angle of the line connecting the target contour points in the target local feature image. Finally, a preset visual coordinate axis template is used to summarize the parameters of the geometric center point, horizontal axis direction, and tilt angle to generate a basic visual coordinate axis model corresponding to the target local feature image.
[0097] Example 3
[0098] Figure 6 This is a schematic diagram of a browsing area determination device based on eye tracking, provided in Embodiment 3 of the present invention. Figure 6 As shown, the device includes: a data acquisition module 310, a data processing module 320, a focus calculation module 330, and a region determination module 340;
[0099] The data acquisition module 310 is used to acquire the target infrared image and the target device visual offset model corresponding to the target user; wherein the target device visual offset model includes the visual focus offset between the target user and each target point in the target terminal device.
[0100] The data processing module 320 is used to process the target infrared image based on preset image processing rules and preset visual coordinate axis establishment rules, and determine the visual coordinate axis model set corresponding to the target infrared image; wherein, the visual coordinate axis model set includes a first visual coordinate axis model and a second visual coordinate axis model;
[0101] The focus calculation module 330 is used to perform focus calculation on the first visual coordinate axis model and the second visual coordinate axis model based on preset eye-tracking technology, and determine the target visual focus corresponding to the target infrared image;
[0102] The region determination module 340 is used to locate the visual focus of the target based on the visual offset model of the target device, and determine the target browsing area corresponding to the infrared image of the target.
[0103] The technical solution of this invention processes the target infrared image corresponding to the target user using preset image processing rules and preset visual coordinate axis establishment rules to determine a set of visual coordinate axis models corresponding to the target infrared image, including a first visual coordinate axis model and a second visual coordinate axis model. Then, based on preset eye-tracking technology, focus calculation is performed on the first and second visual coordinate axis models to determine the target visual focus corresponding to the target infrared image. Finally, based on the target device visual offset model, the target visual focus is located to determine the target browsing area corresponding to the target infrared image. Because the dynamic calculation of the user's gaze focus is achieved through relevant algorithms, the problem of low accuracy in browsing tracking on large-screen mobile devices is solved, accurately obtaining the area currently focused by the user and improving the accuracy of browsing tracking.
[0104] Optionally, the data acquisition module 310 can be specifically used to: acquire a basic infrared image sequence corresponding to the target user; wherein the basic infrared image sequence includes candidate infrared images; perform pupil recognition on the basic infrared image sequence based on a preset focus recognition rule to determine the pupil focus change information corresponding to the basic infrared image sequence; if the pupil focus change information satisfies the preset focus change rule, then the candidate infrared image is used as the target infrared image corresponding to the target user.
[0105] Optionally, the eye-tracking-based browsing area determination device may further include: an offset model construction module, configured to: acquire a function trigger command corresponding to the target user before acquiring the target infrared image and the target device visual offset model corresponding to the target user; and determine an initial infrared image of the target guidance point corresponding to the target user based on the function trigger command and a preset set of guidance points; wherein the preset set of guidance points includes each target guidance point; perform data processing on the initial infrared image based on preset image processing rules and preset visual coordinate axis establishment rules to determine a set of visual coordinate axis models corresponding to the initial infrared image; and determine the focus deviation value of the target guidance point corresponding to the target user based on the target guidance point and the set of visual coordinate axis models, and summarize and process the focus deviation values of each target guidance point in the preset set of guidance points corresponding to the target user as the target device visual offset model corresponding to the target user.
[0106] Optionally, the data processing module 320 can be specifically used for: performing feature point recognition on the target infrared image based on preset image processing rules to determine the feature region set corresponding to the target infrared image; determining the local feature image set corresponding to the target infrared image based on the target feature regions in the feature region set; wherein, the local feature image set includes a first local feature image and a second local feature image corresponding to the same target infrared image; establishing visual coordinate axes for the target local feature images in the local feature image set based on preset visual coordinate axis establishment rules, generating a target visual coordinate axis model corresponding to the target local feature image, and summarizing and processing the target visual coordinate axis models corresponding to the target local feature images in the local feature image set as the visual coordinate axis model set corresponding to the target infrared image.
[0107] Optionally, the data processing module 320 can be specifically used for: calculating contour points on the target local feature image based on a preset center point calculation rule to determine the geometric center point corresponding to the target local feature image; determining the horizontal axis direction corresponding to the target local feature image based on the horizontal span between target contour points in the target local feature image; determining the tilt angle corresponding to the target local feature image based on the angle of the line connecting the target contour points in the target local feature image; summarizing the parameters of the geometric center point, horizontal axis direction, and tilt angle based on a preset visual coordinate axis template to generate a basic visual coordinate axis model corresponding to the target local feature image; performing coordinate transformation on the target pupil center corresponding to the target local feature image based on the basic visual coordinate axis model to generate a coordinate transformation result, and merging the coordinate transformation result into the basic visual coordinate axis model to generate a target visual coordinate axis model corresponding to the target local feature image.
[0108] Optionally, the focus calculation module 330 can be specifically used to: determine the first pupil offset corresponding to the first visual region based on the first visual coordinate axis model, and determine the second pupil offset corresponding to the second visual region based on the second visual coordinate axis model; perform coordinate transformation on the first pupil offset based on a preset first gaze direction vector to generate a first global gaze corresponding to the first visual region, and perform coordinate transformation on the second pupil offset based on a preset second gaze direction vector to generate a second global gaze corresponding to the second visual region; and perform focus calculation on the first global gaze and the second global gaze based on a preset parameterized straight line equation to generate the target visual focus corresponding to the target infrared image.
[0109] Optionally, the eye-tracking-based browsing area determination device may further include: a post-processing module, configured to: after determining the target browsing area corresponding to the target infrared image by locating the target visual focus based on the target device visual offset model; obtain basic attribute parameters corresponding to the target browsing area; generate browsing behavior data corresponding to the target user based on the target browsing area and the basic attribute parameters; summarize the browsing behavior data corresponding to the target user based on a preset time threshold; generate a browsing behavior data set corresponding to the target user; and send the browsing behavior data set to the server so that the server can perform data analysis on the browsing behavior data set, generate browsing hotspots corresponding to the target user, and perform corresponding push operations based on the browsing hotspots.
[0110] The eye-tracking-based browsing area determination device provided in this embodiment of the invention can execute the eye-tracking-based browsing area determination method provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the method.
[0111] Example 4
[0112] Figure 7 A schematic diagram of an electronic device 410 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of mobile devices, such as smartphones and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0113] like Figure 7As shown, the electronic device 410 includes at least one processor 420 and a memory, such as a read-only memory (ROM) 430 or a random access memory (RAM) 440, communicatively connected to the at least one processor 420. The memory stores computer programs executable by the at least one processor. The processor 420 can perform various appropriate actions and processes based on the computer program stored in the ROM 430 or loaded into the RAM 440 from storage unit 490. The RAM 440 may also store various programs and data required for the operation of the electronic device 410. The processor 420, ROM 430, and RAM 440 are interconnected via a bus 450. An input / output (I / O) interface 460 is also connected to the bus 450. Multiple components in electronic device 410 are connected to I / O interface 460, including: input unit 470, such as a keyboard, mouse, etc.; output unit 480, such as various types of displays, speakers, etc.; storage unit 490, such as a disk, optical disk, etc.; and communication unit 4100, such as a network card, modem, wireless transceiver, etc. Communication unit 4100 allows electronic device 410 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks. Processor 420 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 420 include, but are not limited to, central processing unit (CPU), graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, digital signal processors (DSPs), and any suitable processor, controller, microcontroller, etc. Processor 420 performs the various methods and processes described above, such as eye-tracking-based browsing area determination methods.
[0114] The method includes: acquiring a target infrared image and a target device visual offset model corresponding to a target user; wherein the target device visual offset model includes the visual focus offset between the target user and each target point in the target terminal device; processing the target infrared image based on preset image processing rules and preset visual coordinate axis establishment rules to determine a set of visual coordinate axis models corresponding to the target infrared image; wherein the set of visual coordinate axis models includes a first visual coordinate axis model and a second visual coordinate axis model; calculating the focus of the first visual coordinate axis model and the second visual coordinate axis model based on preset eye-tracking technology to determine the target visual focus corresponding to the target infrared image; and locating the target visual focus based on the target device visual offset model to determine the target browsing area corresponding to the target infrared image.
[0115] In some embodiments, the eye-tracking-based browsing region determination method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 490. In some embodiments, part or all of the computer program may be loaded and / or mounted on electronic device 410 via ROM 430 and / or communication unit 4100. When the computer program is loaded into RAM 440 and executed by processor 420, one or more steps of the eye-tracking-based browsing region determination method described above may be performed. Alternatively, in other embodiments, processor 420 may be configured to perform the eye-tracking-based browsing region determination method by any other suitable means (e.g., by means of firmware).
[0116] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transferring data and instructions to the storage system, the at least one input device, and the at least one output device.
[0117] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, partially on a machine and partially on a remote machine as a standalone software package, or entirely on a remote machine or server. In the context of the present invention, a computer-readable storage medium may be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. Computer-readable storage media may include, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0118] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input). The systems and techniques described herein can be implemented in computing systems that include back-end components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include front-end components (e.g., a user computer with a graphical user interface or a web browser through which the user interacts with embodiments of the systems and techniques described herein), or computing systems that include any combination of such back-end components, middleware components, or front-end components. System components can be interconnected via digital data communication (e.g., communication networks) in any form or medium. Examples of communication networks include Local Area Networks (LANs), Wide Area Networks (WANs), blockchain networks, and the Internet. Computing systems can include clients and servers. Clients and servers are generally geographically separated and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, also known as cloud computing servers or cloud hosts, which are hosting products within the cloud computing service ecosystem to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0119] This application also discloses a computer program product, which includes a computer program that, when executed by a processor, implements the eye-tracking-based browsing area determination method provided in any embodiment of this application. This program product shares the same inventive concept as the eye-tracking-based browsing area determination method disclosed in the embodiments of this application, and therefore will not be described in detail here.
[0120] It should be understood that the various processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein. The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A method for determining browsing areas based on eye tracking, characterized in that, include: Obtain the target infrared image and the target device visual offset model corresponding to the target user; wherein, the target device visual offset model includes the visual focus offset between the target user and each target point in the target terminal device; The target infrared image is processed based on preset image processing rules and preset visual coordinate axis establishment rules to determine the visual coordinate axis model set corresponding to the target infrared image; wherein, the visual coordinate axis model set includes a first visual coordinate axis model and a second visual coordinate axis model; Based on preset eye-tracking technology, focus calculation is performed on the first visual coordinate axis model and the second visual coordinate axis model to determine the target visual focus corresponding to the target infrared image; Based on the target device visual offset model, the target visual focus is located to determine the target browsing area corresponding to the target infrared image.
2. The method according to claim 1, characterized in that, The step of acquiring the target infrared image corresponding to the target user includes: Obtain the basic infrared image sequence corresponding to the target user; wherein, the basic infrared image sequence includes candidate infrared images; Based on a preset focus recognition rule, pupil recognition is performed on the basic infrared image sequence to determine the pupil focus change information corresponding to the basic infrared image sequence; If the pupil focus change information meets the preset focus change rules, then the candidate infrared image is used as the target infrared image corresponding to the target user.
3. The method according to claim 1, characterized in that, Before acquiring the target infrared image and the target device visual offset model corresponding to the target user, the method further includes: Obtain the function trigger command corresponding to the target user, and determine the initial infrared image of the target guidance point corresponding to the target user based on the function trigger command and the preset guidance point set; wherein, the preset guidance point set includes each target guidance point; Based on preset image processing rules and preset visual coordinate axis establishment rules, the initial infrared image is processed to determine the visual coordinate axis model set corresponding to the initial infrared image. Based on the target guidance points and the visual coordinate axis model set, the focus deviation value of the target guidance point corresponding to the target user is determined, and the focus deviation values of each target guidance point in the preset guidance point set corresponding to the target user are summarized and processed as the target device visual offset model corresponding to the target user.
4. The method according to claim 1, characterized in that, The process of processing the target infrared image based on preset image processing rules and preset visual coordinate axis establishment rules to determine the visual coordinate axis model set corresponding to the target infrared image includes: Based on preset image processing rules, feature point recognition is performed on the target infrared image to determine the set of feature regions corresponding to the target infrared image; Based on the target feature regions in the feature region set, a set of local feature images corresponding to the target infrared image is determined; wherein, the set of local feature images includes a first local feature image and a second local feature image corresponding to the same target infrared image; Based on the preset visual coordinate axis establishment rules, visual coordinate axes are established for the target local feature images in the local feature image set, generating the target visual coordinate axis model corresponding to the target local feature image, and summarizing and processing the target visual coordinate axis models corresponding to the target local feature images in the local feature image set as the visual coordinate axis model set corresponding to the target infrared image.
5. The method according to claim 4, characterized in that, The step of establishing visual coordinate axes for target local feature images in the local feature image set based on preset visual coordinate axis establishment rules, and generating a target visual coordinate axis model corresponding to the target local feature image, includes: Based on a preset center point calculation rule, contour points are calculated for the target local feature image to determine the geometric center point corresponding to the target local feature image; Based on the horizontal span between target contour points in the target local feature image, the horizontal axis direction corresponding to the target local feature image is determined; Based on the angle of the line connecting the target contour points in the target local feature image, determine the tilt angle corresponding to the target local feature image; Based on a preset visual coordinate axis template, the parameters of the geometric center point, horizontal axis direction and tilt angle are summarized to generate a basic visual coordinate axis model corresponding to the target local feature image; Based on the basic visual coordinate axis model, the coordinate transformation of the target pupil center corresponding to the target local feature image is performed to generate a coordinate transformation result, and the coordinate transformation result is merged into the basic visual coordinate axis model to generate the target visual coordinate axis model corresponding to the target local feature image.
6. The method according to claim 1, characterized in that, The step of calculating the focus of the first visual coordinate axis model and the second visual coordinate axis model based on preset eye-tracking technology to determine the target visual focus corresponding to the target infrared image includes: The first pupil offset corresponding to the first visual region is determined based on the first visual coordinate axis model, and the second pupil offset corresponding to the second visual region is determined based on the second visual coordinate axis model. Based on a preset first gaze direction vector, the coordinate transformation of the first pupil offset is performed to generate a first global gaze corresponding to the first visual region; and based on a preset second gaze direction vector, the coordinate transformation of the second pupil offset is performed to generate a second global gaze corresponding to the second visual region. Based on a preset parameterized linear equation, focus calculation is performed on the first global line of sight and the second global line of sight to generate the target visual focus corresponding to the target infrared image.
7. The method according to claim 1, characterized in that, After determining the target browsing area corresponding to the target infrared image by locating the target visual focus based on the target device visual offset model, the method further includes: Obtain the basic attribute parameters corresponding to the target browsing area, and generate browsing behavior data corresponding to the target user based on the target browsing area and the basic attribute parameters; Based on a preset time threshold, the browsing behavior data corresponding to the target user is aggregated to generate a browsing behavior data set corresponding to the target user. The browsing behavior data set is then sent to the server so that the server can perform data analysis on the browsing behavior data set, generate browsing hotspots corresponding to the target user, and perform corresponding push operations based on the browsing hotspots.
8. A browsing area determination device based on eye tracking, characterized in that, include: The data acquisition module is used to acquire the target infrared image and the target device visual offset model corresponding to the target user; wherein, the target device visual offset model includes the visual focus offset between the target user and each target point in the target terminal device; The data processing module is used to process the target infrared image based on preset image processing rules and preset visual coordinate axis establishment rules, and determine the visual coordinate axis model set corresponding to the target infrared image; wherein, the visual coordinate axis model set includes a first visual coordinate axis model and a second visual coordinate axis model; The focus calculation module is used to perform focus calculation on the first visual coordinate axis model and the second visual coordinate axis model based on preset eye-tracking technology to determine the target visual focus corresponding to the target infrared image; The region determination module is used to locate the visual focus of the target based on the visual offset model of the target device, and determine the target browsing area corresponding to the infrared image of the target.
9. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, which enables the at least one processor to perform the eye-tracking-based browsing area determination method according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the eye-tracking-based browsing area determination method according to any one of claims 1-7.
11. A computer program product comprising a computer program that, when executed by a processor, implements the eye-tracking-based browsing area determination method according to any one of claims 1-7.