Gaze point interaction method and electronic equipment
By determining that the gaze point is located in the foreground region and smoothing it in the gaze-guided visual question-answering system, a more stable gaze point is generated, which solves the problem of inaccurate positioning caused by gaze point jitter and improves the accuracy and interaction efficiency of the question-answering system.
Patent Information
- Application Number
- CN202511568751.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-30
- Publication Date
- 2025-12-26
AI Technical Summary
Existing gaze-guided visual question-answering systems suffer from jitter in gaze point data due to physiological eye tremors and device noise, making it difficult to stably focus on small-scale objects and affecting the interactive experience.
By acquiring the gaze point and determining that it is located in the foreground region of the image, smoothing is performed to generate a more stable second gaze point, and the question answering results of the visual question answering model are generated based on this gaze point and the local image.
It improves the accuracy and interaction efficiency of visual question answering systems, reduces the amount of data to be analyzed by the model, and enhances the user interaction experience.
Smart Images

Figure CN121209702A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of electronics, and particularly relates to a gaze point interaction method and an electronic device. BACKGROUND
[0002] A visual question answering (VQA) system introduces user gaze point information to improve the accuracy of question answering, especially in scenarios requiring accurate positioning of image details.
[0003] However, the performance of the system is highly dependent on the quality of the gaze data. Due to the inherent physiological tremor of the human eye and the inherent noise of the acquisition device, the obtained original gaze point data of the user has jitter. The jitter makes it difficult for the system to stably focus on small-scale objects in the image, thereby causing the final output result of the system to deviate, affecting the user interaction experience. SUMMARY
[0004] Therefore, the embodiments of the present application at least provide a gaze point interaction method.
[0005] The technical scheme of the embodiments of the present application is implemented as follows: In a first aspect, the embodiments of the present application provide a gaze point interaction method, which comprises the following steps: In response to a first question answering event, a first image corresponding to the first question answering event and a first gaze point are obtained, the first gaze point being used to indicate a position at which a user gazes in the first image; It is determined that the first gaze point is located in a foreground region of the first image; The first gaze point is smoothed to obtain a second gaze point; Based on a second image corresponding to the second gaze point and a visual question answering model, a question answering result corresponding to the first question answering event is obtained; the second image is used to represent a region in which the second gaze point is located in the first image.
[0006] In a second aspect, the embodiments of the present application provide an electronic device, which comprises a display unit, a gaze point tracking unit and a processor: The display unit is configured to display a first image; The gaze point tracking unit is configured to generate a first gaze point; The processor is configured to: in response to a first question and answer event, acquire a first image corresponding to the first question and answer event and a first gaze point, the first gaze point being used to indicate a position at which a user gazes in the first image; determine that the first gaze point is located in a foreground region of the first image; perform smoothing processing on the first gaze point to obtain a second gaze point; and based on a second image corresponding to the second gaze point and a visual question and answer model, obtain a question and answer result corresponding to the first question and answer event; the second image is used to represent a region in which the second gaze point is located in the first image.
[0007] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, but not limiting the technical solutions of the present application. BRIEF DESCRIPTION OF DRAWINGS
[0008] The accompanying drawings, which are incorporated into and form part of the specification, illustrate embodiments consistent with the present application and, together with the specification, serve to explain the technical solutions of the present application.
[0009] Figure 1 An implementation flowchart of a gaze point interaction method provided for an embodiment of the present application Figure One ; Figure 2 A schematic diagram of a segmented first image provided for an embodiment of the present application; Figure 3 A schematic diagram of a first image and a second image provided for an embodiment of the present application; Figure 4 A processing flowchart of a line-of-sight tracking algorithm provided for an embodiment of the present application; Figure 5 An implementation flowchart of a gaze point interaction method provided for an embodiment of the present application Figure Two ; Figure 6 A schematic diagram of a first gaze point, a center point and an edge point provided for an embodiment of the present application; Figure 7 An implementation flowchart of a gaze point interaction method provided for an embodiment of the present application Figure Three ; Figure 8 An implementation flowchart of a gaze point interaction method provided for an embodiment of the present application Figure Four ; Figure 9 An implementation flowchart of a gaze point interaction method provided for an embodiment of the present application Figure Five ; Figure 10 An implementation flowchart of a gaze point interaction method provided for an embodiment of the present application Figure Six ; Figure 11A schematic structural diagram of an electronic device provided in an embodiment of the present application Figure One ; Figure 12 A schematic structural diagram of an electronic device provided in an embodiment of the present application Figure Two ; Figure 13 A schematic diagram of a hardware entity of an electronic device in an embodiment of the present application. DETAILED DESCRIPTION
[0010] In order to make the purposes, technical solutions and advantages of the present application clearer, the technical solutions of the present application are further described in detail below in combination with the drawings and embodiments, and the described embodiments should not be regarded as limiting the present application, and all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.
[0011] In the following description, “some embodiments” are described, which describe a subset of all possible embodiments, but it can be understood that “some embodiments” can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.
[0012] The terms “first / second / third” involved are only to distinguish similar objects, and do not represent a specific order of the objects, and it can be understood that “first / second / third” can interchange specific order or sequence as allowed, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0013] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present application belongs. The terms used herein are only for the purpose of describing the present application and are not intended to limit the present application.
[0014] In order to solve the technical problems in the related art, the gaze point interaction method provided in the embodiments of the present application can be applied to an electronic device, exemplarily, the electronic device includes but is not limited to a smart phone, a tablet computer, a wearable device, a personal computer (PC), a netbook, etc., and the implementation form of the embodiments of the present application is not fixedly limited.
[0015] Figure 1 An implementation flowchart of the gaze point interaction method provided in an embodiment of the present application Figure One As shown in Figure 1 , the gaze point interaction method can be implemented through steps S101-S104: Step S101. In response to the first question-answering event, a first image corresponding to the first question-answering event and a first gaze point are obtained.
[0016] Here, the first gaze point is used to indicate a position at which the user gazes in the first image. The first image refers to an image viewed by the user.
[0017] In some examples, the first question-answering event is used to represent that the user initiates a question through voice or gesture, etc. The first question-answering event at least includes a question corresponding to the first question-answering event.
[0018] For example, when the user views a first image including at least one object, the user can ask a question about the first image, for example, the question is “what is this?”, and based on the question, a first question-answering event can be constituted.
[0019] After the electronic device receives the first question-answering event, the electronic device can obtain a first image corresponding to the first question-answering event and a first gaze point in response to the first question-answering event. For example, the first image can be a scene image captured by a front camera, or a scene image displayed by the electronic device.
[0020] In some examples, the first gaze point can be represented by a two-dimensional coordinate point.
[0021] Specifically, the first gaze point can be determined by a gaze tracking algorithm. The process of determining the first gaze point by the gaze tracking algorithm can include: obtaining a user image including a head pose and an eye region of the user, inputting the user image into the gaze tracking algorithm, and outputting the first gaze point by the gaze tracking algorithm.
[0022] Step S102. It is determined that the first gaze point is located in a foreground region of the first image.
[0023] In some examples, before determining that the first gaze point is located in the foreground region of the first image, the method further includes: performing segmentation processing on the first image by using an instance segmentation algorithm to obtain a foreground region and a background region corresponding to the first image.
[0024] Here, the instance segmentation algorithm can identify each object in the first image, and assign different mask ranges to each object in the first image, and divide the foreground region and the background region based on each mask range.
[0025] The foreground region includes at least one detection box and detection box information, and one detection box corresponds to one object. The detection box information includes a class name and a confidence score of the object. The confidence score is used to represent the confidence degree of the identification result of the instance segmentation algorithm. The confidence score has a value range of [0, 1], and the higher the confidence score, the more reliable the identification result.
[0026] For example, Figure 2 An example of a segmented first image is shown. For example... Figure 2 As shown, in a context containing multiple objects (such as...) Figure 2 In the first image (containing people, dogs, mountains, and roads), the people and dogs are considered the foreground area, while the mountains and roads are classified as the background area.
[0027] Within this foreground region, two people and one dog were identified. The confidence scores for the two people were 0.90 and 0.94, respectively. Confidence scores of 0.90 and 0.94 indicate a high probability that the foreground region was identified as a person. The confidence score for the dog was 0.88, indicating a high probability that the foreground region was identified as a dog.
[0028] The instance segmentation algorithm described above can effectively classify the foreground and background regions of the first image.
[0029] In some examples, after segmenting the first image using an instance segmentation algorithm to obtain the foreground and background regions corresponding to the first image, the method further includes: determining the coordinate set corresponding to the foreground region based on at least one detection box corresponding to the foreground region; and determining the position of the first gaze point based on the coordinate set corresponding to the first gaze point and the foreground region.
[0030] The coordinate set includes the vertex coordinates of at least one detection box. A detection box can correspond to 4 vertex coordinates. Of course, if the detection box is of other shapes, the vertex coordinates corresponding to a detection box can also be 3, 5, etc. This application does not limit this.
[0031] For example, based on the coordinate set corresponding to the first gaze point and the foreground region, when the coordinates of the first gaze point fall within the coordinate set corresponding to the foreground region, the position of the first gaze point is determined as follows: the first gaze point is located in the foreground region of the first image.
[0032] Based on the coordinate set corresponding to the first gaze point and the foreground region, when the coordinates of the first gaze point do not fall within the coordinate set corresponding to the foreground region, the position of the first gaze point is determined as follows: the first gaze point is located in the background region of the first image.
[0033] Understandably, the detection box here can be equivalent to the aforementioned mask range. When the coordinates of the first gaze point are within at least one mask range (i.e., the set of coordinates corresponding to the foreground region), the first gaze point is considered to be located in the foreground region of the first image.
[0034] Step S103. Smooth the first fixation point to obtain the second fixation point.
[0035] After determining that the first gaze point is located in the foreground region, it can be smoothed to obtain the second gaze point. Smoothing can eliminate noise caused by eye movement, making the final gaze point trajectory more stable and smooth.
[0036] Step S104. Based on the second image corresponding to the second gaze point and the visual question-answering model, obtain the question-answering result corresponding to the first question-answering event.
[0037] The second image is used to characterize the region where the second gaze point is located in the first image.
[0038] Understandably, the second image is a local region image extracted from the first image, containing the detection box of the second gaze point. The second image is used to more accurately describe what the user is currently focusing on. For example, if the user is looking at a small object in the first image, then the second image will only contain that small object from the first image.
[0039] The visual question answering model here is a deep learning-based model that can generate the question answering result corresponding to the first question answering event based on the second image and the question corresponding to the first question answering event.
[0040] For example, the second image and the question corresponding to the first question-and-answer event (e.g., what is this?) are input into the visual question-and-answer model. The visual question-and-answer model analyzes the second image and, in conjunction with the question corresponding to the first question-and-answer event, generates the question-and-answer result corresponding to the first question-and-answer event.
[0041] like Figure 3 As shown, in the absence of gaze-based interaction, the input to the visual question-answering model is the first image. For the question corresponding to the first question-answering event (e.g., "What is this?"), the output of the visual question-answering model is: "This is a picture. The picture shows two tables and a girl. One table has a bench, and the other does not. On the table without a bench, there is a cup and a cup holder. The girl is standing near the table with the cup."
[0042] In the case of gaze-based interaction, the second image can be extracted from the first image based on the above method. The input of the visual question answering model is the second image. For the question corresponding to the first question answering event (e.g., What is this?), the output of the visual question answering model for the first question answering event is: This is a cup and a saucer. The cup is placed on the saucer. The cup has a handle.
[0043] Based on the results output by the visual question answering model, it can be seen that by focusing the visual question answering model on the second image according to the user's first gaze point, the output of the visual question answering model is more accurate and the interaction efficiency is higher.
[0044] In the embodiments of the present application, in response to the first question and answer event, the first image and the first gaze point are obtained, and it is determined whether to perform smoothing processing on the first gaze point according to whether the first gaze point is located in the foreground region. Finally, based on the smoothed second gaze point and the second image, the visual question and answer model is used to output the question and answer result. By performing smoothing processing when the user's gaze point enters the foreground region, the problem of inaccurate positioning caused by gaze tracking jitter can be solved. The first gaze point is converted into a second gaze point that is more stable and more in line with the user's visual habits, thereby improving the input quality of subsequent visual question and answer. At the same time, the local region image (i.e., the second image) cropped by the second gaze point is used as the input of the visual question and answer model, which not only makes the model more focused, but also reduces the amount of data analyzed by the visual question and answer model, thereby improving the accuracy and interaction efficiency of the question and answer result.
[0045] In some examples, obtaining the first gaze point in step S101 can include: Step S201. Obtain a fourth image.
[0046] Here, the fourth image is used to indicate a user face image collected by a camera, and the fourth image includes a head posture and an eye region of the user; the fourth image can be used as input data of a gaze tracking algorithm.
[0047] In some examples, the fourth image can be collected in real time by a rear camera of an electronic device (such as smart glasses, a notebook computer, etc.). The head posture of the user can include a head angle and a position. The eye region can include a pupil position and an eye movement direction.
[0048] After obtaining the fourth image, the head posture of the user in the fourth image can be used to determine the approximate gaze direction of the user, and the eye region information of the user in the fourth image can be used to accurately locate the first gaze point.
[0049] In order to improve the accuracy of the first gaze point, a high-resolution camera can be used, and an infrared light source can be combined to enhance the capture ability of eye reflections.
[0050] Step S202. Use a gaze tracking algorithm to identify the gaze point in the head posture and the eye region in the fourth image, and determine the first gaze point.
[0051] It should be noted that the gaze tracking algorithm is a technology based on computer vision and machine learning, which can derive the gaze point coordinates of the user in the scene image from the head posture and the eye region of the user.
[0052] Figure 4 An exemplary processing flow diagram of a gaze tracking algorithm is shown as follows: Figure 4As shown, the gaze tracking algorithm receives the input data as the fourth image. First, the fourth image is subjected to face and key point detection to locate the face and feature points, and target tracking is maintained in consecutive frames. Then, the face region and eye region are extracted, the analysis region is focused, and the eye key point is optimized. After obtaining the eye features, the eye and ear interaction information and spatial position calculation are processed in parallel, and the head pose estimation is performed. Next, the various data are subjected to data normalization processing to provide input for gaze estimation. The preliminary gaze estimation result is corrected, and the correction process integrates the pupil position, gaze direction, and screen pose. The corrected data is mapped from the eye coordinate system to the screen coordinate system through coordinate system conversion, and the process combines the user-specific parameters obtained through individual calibration in advance to finally output the first gaze point.
[0053] To further improve the accuracy, the above process can also be optimized. The optimization module starts with the eye key points, and performs data verification by analyzing the iris diameter difference and pupil position difference. At the same time, it optimizes the head pose in depth, specifically: using the yaw angle data, performing optimized head pose calculation, and further solving the roll angle. In addition, the optimization module also calculates the projected pupil distance to correct the perspective error, and performs optimized eye position to adjust the core coordinates. The system robustness is also enhanced by analyzing the pupil position difference, and the relevant data is subjected to data normalization processing again before calculation to ensure the stability of the gaze tracking algorithm and the accuracy of the output result.
[0054] After identifying the first gaze point, it can be further determined whether the first gaze point belongs to the foreground region, so as to decide whether to perform subsequent smoothing processing. By combining the gaze tracking algorithm with the user's head pose and eye region, the accuracy of gaze point recognition can be effectively improved, and the tracking stability can be maintained at a high level when the user's head frequently turns.
[0055] It can be understood that, in order to improve the accuracy of the first gaze point recognition, the fourth image can also be subjected to pre-processing operation to remove background interference and enhance contrast. Then the pre-processed fourth image can be input into the gaze tracking algorithm to improve the accuracy of the final output first gaze point.
[0056] In the embodiments of the present application, by introducing the gaze tracking algorithm, the gaze point can be recognized by combining the head pose and eye region information, which can improve the accuracy and robustness of gaze point recognition, so as to stably obtain the user's visual focus, and further support more accurate visual question and answer interaction experience.
[0057] Through the foregoing, first, the fourth image is collected to ensure that the fourth image quality meets the requirements of the gaze tracking algorithm; second, based on the gaze tracking algorithm, gaze point recognition is performed on the head posture and eye region of the user to calculate the first gaze point, and the first gaze point is taken as the basis for subsequent gaze point interaction control.
[0058] As shown in FIG. 10, step S103 performs smoothing processing on the first gaze point to obtain the second gaze point, including: Figure 5 Step S501. In a case where the first gaze point is in a first preset range, performing first smoothing processing on the first gaze point to obtain the second gaze point.
[0059] The deviation between the position corresponding to the second gaze point and the position corresponding to the first gaze point is less than or equal to a first numerical value.
[0060] In some examples, the first preset range is a range determined according to the center point of the foreground region. The first numerical value can be data determined based on experience. The deviation between the position corresponding to the second gaze point and the position corresponding to the first gaze point being less than or equal to the first numerical value means that the position corresponding to the second gaze point is not much different from the position corresponding to the first gaze point.
[0061] When the first gaze point is in the first preset range, it means that the first gaze point is near the center point of the foreground region. When the first gaze point is near the center point of the foreground region, it is considered that even if there is jitter, the first gaze point is still likely to fall within the detection frame of the foreground region. Therefore, when the first gaze point is in the first preset range, the first smoothing processing on the first gaze point is weak in processing strength, and the second gaze point tends to be output with the recognized real gaze point (i.e., the first gaze point).
[0062] Step S402. In a case where the first gaze point is in a second preset range, performing second smoothing processing on the first gaze point to obtain the second gaze point.
[0063] The running track of the second gaze point conforms to the visual inertia of the user's gaze, and the smoothing degree of the second smoothing processing is greater than the smoothing degree of the first smoothing processing.
[0064] In some examples, the second preset range is a range determined according to the edge point of the foreground region.
[0065] When the first gaze point is in the second preset range, it is indicated that the first gaze point is near the edge point. When the first gaze point is near the edge point, if there is jitter, the first gaze point is likely to fall outside the detection box. After the first gaze point falls outside the detection box, the final output of the first question and answer event corresponding to the question and answer result is likely to be incorrect. Therefore, when the first gaze point is in the second preset range, the second smoothing processing of the first gaze point has a stronger processing strength, thereby avoiding the case that the gaze point is incorrect due to jitter, and the second gaze point tends to be output with the predicted gaze point.
[0066] In the embodiments of the application, different smoothing processing with different strengths is used to correct the first gaze point according to different preset ranges in which the first gaze point is located. When the first gaze point is near the center point, the correction strength is small to enable the model to respond quickly; and when the first gaze point is near the edge point, the correction strength is large to enhance the stability and smoothness of the gaze point. In this way, the accuracy and reliability of eye movement interaction can be realized, and user experience can be further improved.
[0067] In some examples, before the first gaze point is smoothed in step S103 to obtain the second gaze point, the method further includes: Step S501. Obtain a first preset range according to a distance between the first gaze point and a center point of the foreground region.
[0068] As described above in S102, the foreground region can include at least one detection box.
[0069] In some examples, before the first preset range is obtained according to the distance between the first gaze point and the center point of the foreground region, the method further includes: determining the center point of the foreground region based on the detection box coordinates corresponding to the foreground region; and determining the edge point of the foreground region based on the first gaze point and the center point of the foreground region.
[0070] After the at least one detection box corresponding to the foreground region is identified, the detection box to which the first gaze point belongs can be determined according to the coordinates of the first gaze point. The center point of the detection box to which the first gaze point belongs, i.e., the center point of the foreground region, can be obtained by performing geometric mean calculation on the plurality of vertex coordinates of the detection box to which the first gaze point belongs.
[0071] After the center point of the foreground region is obtained, the edge point of the foreground region can be determined based on the first gaze point and the center point of the foreground region.
[0072] Specifically, the process of determining the edge point of the foreground region based on the first gaze point and the center point of the foreground region can include: taking the center point coordinate of the foreground region as the starting point, and a straight line passing through the first gaze point coordinate as the first ray; and determining the intersection point of the first ray and the edge frame of the foreground region as the first edge point coordinate.
[0073] As shown in Figure 6 , the first gaze point is ; the center point of the foreground region is ; and the edge point of the foreground region is . The distance between the first gaze point and the center point of the foreground region is d. The distance between the center point of the foreground region and the edge point of the foreground region is L.
[0074] Here, the first preset range is a dynamic threshold range defined based on the distance between the first gaze point and the center point of the foreground region. When the first gaze point falls within the first preset range, it indicates that the first gaze point is located near the center point of the foreground region. In this case, even if there is jitter, the first gaze point is likely to fall within the detection frame corresponding to the foreground region, so it is not necessary to exert a large smoothing force on the first gaze point. At this time, the smoothing force on the first gaze point is small.
[0075] Step S502. Obtain a second preset range according to the distance between the first gaze point and the edge point of the foreground region.
[0076] Here, the second preset range is a dynamic threshold range defined based on the distance between the first gaze point and the edge point of the foreground region. When the first gaze point falls within the second preset range, it indicates that the first gaze point is near the edge point of the foreground region. If there is jitter, the first gaze point is likely to fall outside the detection frame, so the output result will be biased, and a large smoothing force needs to be exerted on the first gaze point to avoid sudden changes in the first gaze point, thereby affecting the final visual question and answer result.
[0077] Based on the above technical features, by calculating the first preset range and the second preset range according to the distances between the first gaze point and the center point and the edge point of the foreground region respectively, the response behavior of the electronic device can be more finely controlled, thereby realizing accurate analysis and processing of the first gaze point and improving the accuracy of gaze point interaction.
[0078] In some examples, as shown in Figure 7 , when the first gaze point is within the first preset range, the first gaze point is subjected to first smoothing processing to obtain a second gaze point, including: Step 701. Obtain a first damping coefficient according to the first gaze point and the center point.
[0079] In some examples, the process of obtaining the first damping coefficient according to the first gaze point and the center point can include: determining first length data based on the coordinates of the first gaze point and the coordinates of the center point; determining second length data based on the coordinates of the first edge point and the coordinates of the center point; determining a damping coefficient of the first edge point based on the pixel area of the foreground region; determining the first damping coefficient based on the first length data, the second length data, and the damping coefficient of the first edge point.
[0080] As shown in Figure 6 , the first gaze point is ; the center point of the foreground region is ; and the edge point of the foreground region is . According to the coordinates of the first gaze point and the coordinates of the center point , the first length data can be determined, which can be marked as d. According to the coordinates of the center point and the coordinates of the edge point , the second length data can be determined, which can be marked as L. Based on the pixel area of the foreground region, the damping coefficient of the first edge point is determined. Based on the first length data d, the second length data L, and the damping coefficient of the first edge point , the first damping coefficient is determined.
[0081] Exemplarily, the first length data d, the second length data L, and the first damping coefficient satisfy the following expressions: (1) (2) (3) wherein the value of the damping coefficient of the first edge point is related to the pixel area of the foreground region, the larger the pixel area of the foreground region is, the smaller the value of the damping coefficient of the first edge point is; and the smaller the pixel area of the foreground region is, the larger the value of the damping coefficient of the first edge point is.
[0082] It can be known from the above formula that the value of the first damping coefficient depends on the distance between the first gaze point and the center point and the area of the foreground region. Since the first gaze point is in the first preset range at this time, that is, the position of the first gaze point is close to the position of the center point, the value of the first damping coefficient is small. When the first damping coefficient is small, the smoothing ability of the first damping coefficient to the first gaze point is weak, and the first gaze point tends to be directly output.
[0083] That is, when the user gazes around the center point, it is considered that the error caused by the slight jitter does not affect the final result. Therefore, the second gaze point tends to directly refer to the first gaze point output.
[0084] Step S702. According to the first damping coefficient, the first gaze point is first smoothed to obtain the second gaze point.
[0085] After obtaining the first damping coefficient, the first gaze point can be first smoothed using the first damping coefficient to obtain the second gaze point.
[0086] In some examples, the process of first smoothing the first gaze point according to the first damping coefficient to obtain the second gaze point can include: taking the first damping coefficient as the measurement noise covariance of a Kalman filter, filtering the first gaze point using the Kalman filter, and the Kalman filter outputting the second gaze point. Wherein, the second gaze point can be marked as .
[0087] Here, the first smoothing is a filtering operation based on the Kalman filter. More specifically, the first smoothing is to smooth the first gaze point coordinates by taking the first damping coefficient as the measurement noise covariance parameter, so that the obtained second gaze point is more stable. Exemplarily, the larger the first damping coefficient, the more the output result of the Kalman filter tends to use the predicted value, thereby reducing the influence of actual value fluctuation; on the contrary, the smaller the first damping coefficient, the more the Kalman filter relies on the actual value, thereby maintaining the sensitivity of the model response.
[0088] For example, the second gaze point satisfies the following expression: (4) (5) Wherein, K represents the Kalman gain, represents the gaze point prediction value at the last time, represents a Gaussian noise with a mean of 0 and a covariance of . represents a calibration coefficient, and I represents a unit matrix. represents the first damping coefficient.
[0089] In the embodiments of the present application, the first damping coefficient is calculated through the first gaze point and the center point, and the first gaze point is first smoothed according to the calculated first damping coefficient, so as to effectively control the line of sight trajectory and improve the stability of the line of sight positioning.
[0090] In some examples, as Figure 8As shown, in a case where the first gaze point is in the second preset range, the first gaze point is subjected to second smoothing processing to obtain a second gaze point, including: Step S801. Obtain a second damping coefficient according to the first gaze point and the edge point.
[0091] In some examples, the process of obtaining the second damping coefficient according to the first gaze point and the center point is similar to the foregoing manner of obtaining the first damping coefficient, which will not be described here. The difference between the determination manners of the second damping coefficient and the first damping coefficient is that the first damping coefficient is obtained in a case where the first gaze point is in the first preset range, and the second damping coefficient is obtained in a case where the first gaze point is in the second preset range.
[0092] When the first gaze point is in the second preset range, it indicates that the position of the first gaze point is close to the position of the edge point, and therefore the value of the second damping coefficient is larger. When the second damping coefficient is larger, the smoothing ability of the second damping coefficient on the first gaze point is stronger, and the first gaze point tends to be output according to the predicted value.
[0093] Step S602. Subject the first gaze point to second smoothing processing according to the second damping coefficient to obtain a second gaze point.
[0094] Here, the second damping coefficient is greater than the first damping coefficient.
[0095] After obtaining the second damping coefficient, the first gaze point can be subjected to second smoothing processing by using the second damping coefficient to obtain a second gaze point.
[0096] In some examples, the process of subjecting the first gaze point to second smoothing processing according to the second damping coefficient to obtain a second gaze point can include: taking the second damping coefficient as a measurement noise covariance of a Kalman filter, using the Kalman filter to filter the first gaze point, and outputting the second gaze point by the Kalman filter.
[0097] Here, the second smoothing processing is also a filtering operation of the Kalman filter. More specifically, the second smoothing processing is to take the second damping coefficient as an input of the measurement noise covariance to optimize the first gaze point coordinates, so that the obtained second gaze point is more stable and smoother. Since the second damping coefficient is greater than the first damping coefficient, the Kalman filter will be more inclined to rely on the predicted value, thereby achieving a stronger smoothing effect. A larger second damping coefficient is suitable for a case where the user's line of sight is located near the edge of a small object, and can effectively reduce the instability phenomenon caused by slight movement.
[0098] By employing a larger second damping coefficient within a second preset range for filtering, the stability of the gaze point coordinates can be significantly improved. This reduces misjudgments caused by natural eye tremors, thereby increasing the accuracy of user intent assessment and ultimately enhancing interaction efficiency and user experience.
[0099] In some examples, such as Figure 9 As shown, step S104, based on the second image corresponding to the second gaze point and the visual question-answering model, obtains the question-answering result corresponding to the first question-answering event, including: Step S901. Based on the second gaze point, determine the detection box corresponding to the second gaze point in the first image.
[0100] Here, the detection box corresponding to the second gaze point represents a region bounded by the target object in the first image; the detection box corresponding to the second gaze point is used to identify the target object or region that the user is interested in.
[0101] In some examples, the detection box corresponding to the second gaze point is a region bounding box determined on the first image based on the coordinates of the second gaze point. As can be seen from the aforementioned step S102, the detection box corresponding to the second gaze point is generated by an instance segmentation algorithm to ensure that the second gaze point falls on the object of interest to the user.
[0102] By generating a detection box corresponding to the second gaze point, the range corresponding to the second gaze point can be accurately defined in the first image, thereby improving the recognition efficiency and accuracy of the visual question answering model. The effect is even more obvious in small object recognition scenarios.
[0103] Step S902. Based on the detection box corresponding to the second gaze point, crop the first image to obtain the second image corresponding to the second gaze point.
[0104] In some examples, the process of cropping the first image based on the detection box corresponding to the second gaze point to obtain the second image corresponding to the second gaze point may include: cutting the first image according to the position coordinates of the detection box corresponding to the second gaze point to generate the second image.
[0105] By removing irrelevant background information from the first image and retaining only the part relevant to the user's second gaze (i.e., the second image), the complexity of image processing can be significantly reduced, the amount of computation can be decreased, and the visual question answering model can generate answers faster and more accurately.
[0106] Step S903. Input the second image corresponding to the second gaze point into the visual question answering model to obtain the question answering result corresponding to the first question answering event.
[0107] Exemplarily, the visual question answering model can employ a convolutional neural network (CNN) to extract image features, and use a recurrent neural network (RNN) or a transformer structure to process natural language questions, and finally output a question and answer result corresponding to the first question and answer event related to the content in the second image.
[0108] Here, the input of the visual question answering model is the second image corresponding to the second gaze point after cropping, rather than the complete first image, which can improve the focus of the visual question answering model on key information.
[0109] It can be understood that the detection frame corresponding to the second gaze point generated in step S901 provides the basis for cropping in step S902, and the second image corresponding to the second gaze point cropped in step S902 is used as the input data of the visual question answering model in step S903. Based on the above steps, a closed-loop visual question answering process is formed, thereby efficiently and accurately responding to questions and answers.
[0110] In the embodiments of the present application, the detection frame corresponding to the second gaze point is generated through the second gaze point and is cropped, and then the cropped second image corresponding to the second gaze point is input into the visual question answering model, which can effectively focus on the user's attention area, thereby improving the inference efficiency and accuracy of the visual question answering model, and further improving the human-computer interaction experience.
[0111] In some examples, as shown in Figure 10 The method further includes: Step S1001. Determine that the first gaze point is located in the background region of the first image.
[0112] In some examples, before determining that the first gaze point is located in the background region of the first image, the method further includes: performing segmentation processing on the first image using an instance segmentation algorithm to obtain a foreground region and a background region corresponding to the first image. For details, reference can be made to the foregoing step S102, which will not be repeated here.
[0113] After performing segmentation processing on the first image using the instance segmentation algorithm, the foreground region corresponding to the first image includes at least one detection frame, and when the first gaze point is outside the coordinate set corresponding to the at least one detection frame, it is determined that the first gaze point is located in the background region of the first image. For example, in an image containing a person and a background scenery, the region where the person is located is identified as the foreground, and the scenery part is the background region.
[0114] By extending the gaze point judgment range from the foreground region to the background region, the effective coverage range of the line-of-sight interaction can be expanded, which is especially suitable for scenes in which the user's line of sight is temporarily deviated or there is no obvious foreground object in the image, and the accuracy of identifying the user's intention is improved.
[0115] Step S1002. Obtain a question and answer result corresponding to the first question and answer event based on the first gaze point, a third image corresponding to the first gaze point, and a visual question and answer model.
[0116] The third image is used to represent a region in the first image where the first gaze point is located.
[0117] Here, the third image is a local region image cropped from the first image with the first gaze point as the center. By using the third image, the content of the user's current gaze can be more specific.
[0118] In the visual question and answer scenario, the user mostly asks questions only for the foreground region in the first image. When it is determined that the first gaze point is located in the background region of the first image, it indicates that the user's focus is in the background region. When the user asks questions for the background region, since the background region is usually large, the influence of gaze point jitter on user experience can be ignored, so the first gaze point is directly used to obtain the question and answer result corresponding to the first question and answer event. There is no need to smooth the first gaze point.
[0119] In some examples, the process of obtaining the question and answer result corresponding to the first question and answer event by using the first gaze point can include: cropping a third image corresponding to the first gaze point from the first image by using the coordinates of the first gaze point; inputting the third image and a question corresponding to the first question and answer event into the visual question and answer model, and the visual question and answer model outputs the question and answer result corresponding to the first question and answer event.
[0120] In the embodiments of the present application, when it is identified that the first gaze point of the user falls in the background region, the third image corresponding to the first gaze point is combined with the visual question and answer model for question and answer processing. In this way, the gaze point of the user is determined in the background region, and the question and answer processing is performed in combination with the local region image and the visual question and answer model, which can expand the application range of the line-of-sight interaction, support more types of interactive scenarios, and thus significantly improve the intelligence of the visual question and answer model and the user satisfaction.
[0121] In some examples, after step S1001 determines that the first gaze point is located in the background region of the first image, the method further includes: Step S1003. Determine the running trajectory of the first gaze point.
[0122] Here, the running trajectory of the first gaze point can be the change path of multiple gaze points of the user in a period of time.
[0123] By recording the coordinates of multiple gaze points within a period of time, a continuous motion trajectory can be formed. The motion trajectory of the first gaze point can be used to describe the user's visual line movement behavior. For example, in a visual question and answer scenario, if the user continuously places the gaze point on a small object in the first image, it can be considered that the small object is the focus of the user's attention.
[0124] By identifying the motion trajectory of the first gaze point, not only can the user's focus be identified, but also it can serve as an important basis for subsequent correction of the gaze point. In particular, in the case of jitter or transient errors in eye movement interaction, the user's true intention can be more accurately captured by combining the motion trajectory information of the first gaze point.
[0125] Step S1004. Obtain a second gaze point based on the first gaze point and the motion trajectory of the first gaze point.
[0126] After obtaining the first gaze point and the motion trajectory of the first gaze point, the first gaze point and the motion trajectory of the first gaze point can be processed to generate a corrected gaze position, i.e., a second gaze point.
[0127] In some scenarios, for example, the first gaze point is always in the foreground region within a predetermined time period. According to the visual inertia of the user's gaze point, the next first gaze point should fall in the foreground region, while the detected first gaze point falls in the background region. For such a case, a corrected second gaze point can be generated based on the first gaze point and the motion trajectory of the first gaze point, and the corrected second gaze point falls in the foreground region.
[0128] There is a dependency relationship between the second gaze point and the first gaze point. By analyzing the motion trajectory of the first gaze point, abnormal data points caused by physiological reasons (such as blinking and slight head shaking) can be effectively filtered out while preserving the user's true intention, and the second gaze point can be more accurately determined. At the same time, the interaction experience and recognition accuracy in the visual question and answer task are further improved.
[0129] After obtaining the second gaze point, a second image corresponding to the second gaze point can be cropped from the first image according to the second gaze point, and finally the second image can be input into the visual question and answer model to obtain the question and answer result corresponding to the first question and answer event.
[0130] In the embodiments of the present application, the user's visual line movement behavior is analyzed through the motion trajectory of the first gaze point, and the first gaze point is further optimized to generate a more stable second gaze point. In this way, accurate identification of the user's true intention can be achieved, and the adaptability to complex scenarios is also improved, enhancing the naturalness and efficiency of human-computer interaction.
[0131] In some examples, Figure 11 An example structure of an electronic device is shown schematicallyFigure One As shown in Figure 11 The electronic device includes a display unit 1101, a gaze point tracking unit 1102, and a processor 1103.
[0132] The display unit 1101 is configured to display a first image.
[0133] The gaze point tracking unit 1102 is configured to generate a first gaze point.
[0134] The processor 1103 is configured to, in response to a first question and answer event, acquire a first image corresponding to the first question and answer event and a first gaze point, the first gaze point being used to indicate a position at which a user gazes in the first image; determine that the first gaze point is located in a foreground region of the first image; perform smoothing processing on the first gaze point to obtain a second gaze point; and based on a second image corresponding to the second gaze point and a visual question and answer model, obtain a question and answer result corresponding to the first question and answer event, the second image being used to represent a region in which the second gaze point is located in the first image.
[0135] Of course, the display unit can also be used to display other images or scenes observed by the user. It can be understood that the display unit can be integrated in an electronic device (such as smart glasses, a notebook computer, or other devices with display functions), and can also exist independently of the electronic device.
[0136] When the user gazes at the first image displayed by the electronic device, the gaze point tracking unit 1102 is further configured to acquire a fourth image and generate the first gaze point based on the fourth image.
[0137] In some examples, the gaze point tracking unit 1102 can further include a camera, which can be located at a front position or a rear position of the gaze point tracking unit. The camera is configured to acquire the fourth image, the fourth image including a head posture and an eye region of the user.
[0138] The gaze point tracking unit 1102 is further configured to determine a two-dimensional coordinate of the first gaze point in the first image based on the head posture and the eye region of the user in the fourth image acquired by the camera. In this way, the electronic device can track the content region to which the user pays attention in real time, and provide input data for subsequent gaze point interaction processing.
[0139] In some examples, the processor 1103 is further configured to, in a case where the first gaze point is in a first preset range, perform first smoothing processing on the first gaze point to obtain a second gaze point, a deviation between a position corresponding to the second gaze point and a position corresponding to the first gaze point being less than or equal to a first value; and in a case where the first gaze point is in a second preset range, perform second smoothing processing on the first gaze point to obtain the second gaze point, a running track of the second gaze point conforming to visual inertia of the user's gaze; a smoothing degree of the second smoothing processing being greater than a smoothing degree of the first smoothing processing.
[0140] In some examples, the processor 1103 is further configured to obtain the first preset range according to a distance between the first gaze point and a center point of the foreground region; and obtain the second preset range according to a distance between the first gaze point and an edge point of the foreground region.
[0141] In some examples, the processor 1103 is further configured to, in a case where the first gaze point is in the first preset range, obtain a first damping coefficient according to the first gaze point and the center point; and perform the first smoothing processing on the first gaze point according to the first damping coefficient to obtain the second gaze point.
[0142] In some examples, the processor 1103 is further configured to, in a case where the first gaze point is in the second preset range, obtain a second damping coefficient according to the first gaze point and the edge point; perform the second smoothing processing on the first gaze point according to the second damping coefficient to obtain the second gaze point; and the second damping coefficient is greater than the first damping coefficient.
[0143] In some examples, the processor 1103 is further configured to determine that the first gaze point is located in a background region of the first image; obtain a question and answer result corresponding to the first question and answer event based on the first gaze point, a third image corresponding to the first gaze point, and a visual question and answer model; and the third image is used to represent a region in which the first gaze point is located in the first image.
[0144] In some examples, the processor 1103 is further configured to determine a running track of the first gaze point; and obtain the second gaze point based on the first gaze point and the running track of the first gaze point.
[0145] In some examples, the processor 1103 is further configured to determine a detection frame corresponding to the second gaze point in the first image based on the second gaze point; perform cropping processing on the first image based on the detection frame corresponding to the second gaze point to obtain a second image corresponding to the second gaze point; and input the second image corresponding to the second gaze point into the visual question and answer model to obtain the question and answer result corresponding to the first question and answer event.
[0146] In the embodiments of the present application, the user's gaze point can be accurately captured based on the above-mentioned electronic device, and the accurately captured user's gaze point information can be applied to visual question and answer and other interactive scenarios, thereby further improving the user's interactive experience and the accuracy of the visual question and answer task.
[0147] Figure 12 An example of the composition structure of an electronic device provided in the present application Figure Two As shown in Figure 12 The electronic device 1200 includes an acquisition module 1201, a determination module 1202, and a processing module 1203; wherein, The acquisition module 1201 is configured to acquire a first image and a first gaze point corresponding to a first question and answer event in response to the first question and answer event, the first gaze point being used to indicate a position at which a user gazes in the first image.
[0148] The determination module 1202 is configured to determine that the first gaze point is located in a foreground region of the first image.
[0149] The processing module 1203 is configured to perform smoothing processing on the first gaze point to obtain a second gaze point, and obtain a question and answer result corresponding to the first question and answer event based on a second image corresponding to the second gaze point and a visual question and answer model, the second image being used to represent a region in which the second gaze point is located in the first image.
[0150] In some embodiments, the processing module 1203 is further configured to, in a case where the first gaze point is in a first preset range, perform first smoothing processing on the first gaze point to obtain the second gaze point, a deviation between a position corresponding to the second gaze point and a position corresponding to the first gaze point being less than or equal to a first value; and in a case where the first gaze point is in a second preset range, perform second smoothing processing on the first gaze point to obtain the second gaze point, a running track of the second gaze point conforming to visual inertia of the user's gaze; a smoothing degree of the second smoothing processing being greater than a smoothing degree of the first smoothing processing.
[0151] In some embodiments, the processing module 1203 is further configured to obtain the first preset range according to a distance between the first gaze point and a center point of the foreground region; and obtain the second preset range according to a distance between the first gaze point and an edge point of the foreground region.
[0152] In some embodiments, the processing module 1203 is further configured to obtain a first damping coefficient according to the first gaze point and the center point; and perform the first smoothing processing on the first gaze point according to the first damping coefficient to obtain the second gaze point.
[0153] In some embodiments, the processing module 1203 is further configured to obtain a second damping coefficient according to the first gaze point and the edge point; and perform second smoothing processing on the first gaze point according to the second damping coefficient to obtain a second gaze point; the second damping coefficient is greater than the first damping coefficient.
[0154] In some embodiments, the determining module 1202 is further configured to determine that the first gaze point is located in a background region of the first image.
[0155] The processing module 1203 is further configured to obtain a question and answer result corresponding to the first question and answer event based on the first gaze point, a third image corresponding to the first gaze point, and the visual question and answer model; the third image is used to represent a region in which the first gaze point is located in the first image.
[0156] In some embodiments, the determining module 1202 is further configured to determine a running track of the first gaze point.
[0157] The processing module 1203 is further configured to obtain a second gaze point based on the first gaze point and the running track of the first gaze point.
[0158] In some embodiments, the processing module 1203 is further configured to determine a detection frame corresponding to the second gaze point in the first image based on the second gaze point; perform cropping processing on the first image based on the detection frame corresponding to the second gaze point to obtain a second image corresponding to the second gaze point; and input the second image corresponding to the second gaze point into the visual question and answer model to obtain the question and answer result corresponding to the first question and answer event.
[0159] In some embodiments, the obtaining module 1201 is further configured to obtain a fourth image, the fourth image including a head posture and an eye region of a user; and perform gaze point identification on the head posture and the eye region in the fourth image by using a gaze tracking algorithm to determine the first gaze point.
[0160] The descriptions of the above device embodiments are similar to those of the above method embodiments, and have similar beneficial effects to the method embodiments. In some embodiments, the device provided by the embodiments of the present application has the functions or includes the modules for performing the methods described in the above method embodiments. For technical details not disclosed in the device embodiments of the present application, please refer to the descriptions of the method embodiments of the present application.
[0161] It should be noted that, in the embodiments of the present application, if the above-mentioned gaze point interaction method is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of the present application can be embodied in the form of a software product, which is stored in a storage medium, includes several instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute all or part of the embodiments of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a magnetic disk or an optical disk, and various storage media that can store program codes. Thus, the embodiments of the present application are not limited to any specific hardware, software or firmware, or any combination of hardware, software and firmware.
[0162] The embodiments of the present application provide a computer device, including a memory and a processor, the memory stores a computer program capable of running on the processor, and the processor implements part or all of the steps of the above method when executing the program.
[0163] The embodiments of the present application provide a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement part or all of the steps of the above method. The computer readable storage medium can be transitory or non-transitory.
[0164] The embodiments of the present application provide a computer program, which includes computer readable code, and when the computer readable code runs in a computer device, a processor in the computer device executes part or all of the steps of the above method.
[0165] The embodiments of the present application provide a computer program product, which includes a non-transitory computer readable storage medium storing a computer program, and when the computer program is read and executed by a computer, part or all of the steps of the above method are implemented. The computer program product can be specifically implemented by hardware, software or a combination thereof. In some embodiments, the computer program product is specifically embodied as a computer storage medium, and in other embodiments, the computer program product is specifically embodied as a software product, such as a software development kit (Software Development Kit, SDK) and the like.
[0166] It should be noted that the above description of the various embodiments tends to emphasize differences between the various embodiments, and the same or similar elements can be referred to in each other. The above description of the device, storage medium, computer program and computer program product embodiments is similar to the description of the method embodiments, and has similar beneficial effects as the method embodiments. For technical details not disclosed in the device, storage medium, computer program and computer program product embodiments of the present application, please refer to the description of the method embodiments.
[0167] Figure 13 A hardware entity diagram of an electronic device in the embodiments of the present application is shown in FIG. 13, which includes a processor 1301, a communication interface 1302 and a memory 1303. Among them: Figure 13 The processor 1301 generally controls the overall operation of the electronic device 1300, which can be the implementation of the gaze point interaction method provided by the embodiments of the present application.
[0168] The communication interface 1302 can enable the electronic device 1300 to communicate with other terminals or servers through a network.
[0169] The memory 1303 is configured to store instructions and applications executable by the processor 1301, and can also cache data to be processed by the processor 1301 and each module in the electronic device 1300 (for example, image data, audio data, voice communication data and video communication data) that has been processed or has been processed. It can be realized by FLASH or Random Access Memory (RAM). The processor 1301, the communication interface 1302 and the memory 1303 can transmit data through the bus 1304.
[0170] The embodiments of the present application provide a computer storage medium, and the computer storage medium stores one or more programs, which can be executed by one or more processors to implement the steps of the gaze point interaction method according to any one of the above embodiments.
[0171] It should be noted that the above description of the storage medium and device embodiments is similar to the description of the method embodiments, and has similar beneficial effects as the method embodiments. For technical details not disclosed in the storage medium and device embodiments of the present application, please refer to the description of the method embodiments.
[0172] The processor can be at least one of an Application Specific Integrated Circuit (ASIC), a Digital Signal Processor (DSP), a Digital Signal Processing Device (DSPD), a Programmable Logic Device (PLD), a Field Programmable Gate Array (FPGA), a Central Processing Unit (CPU), a controller, a microcontroller, or a microprocessor. It can be understood that the electronic device for implementing the functions of the processor can also be other devices, and the embodiments of the present application are not limited.
[0173] The computer storage medium / memory can be a Read Only Memory (ROM), a Programmable Read-Only Memory (PROM), an Erasable Programmable Read-Only Memory (EPROM), an Electrically Erasable Programmable Read-Only Memory (EEPROM), a Ferromagnetic Random Access Memory (FRAM), a Flash Memory, a magnetic surface memory, an optical disc, a Compact Disc Read-Only Memory (CD-ROM), or the like. It can also be various terminals including one or any combination of the above memories, such as a mobile phone, a computer, a tablet device, a personal digital assistant, and the like.
[0174] It should be understood that the term "one embodiment" or "an embodiment" as used herein means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the application. Thus, the appearances of the phrase "in one embodiment" or "in an embodiment" in various places throughout the specification are not necessarily referring to the same embodiment. Furthermore, the particular features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. It should be understood that, in various embodiments of the application, the order of the steps / processes can be changed without altering the essence of the application described. The sequence of the steps / processes should be determined by their functions and the inherent logic, and should not constitute any limitation on the implementation of the embodiments of the application. The sequence of the embodiments of the application is only for the purpose of description, and does not represent the advantages or disadvantages of the embodiments.
[0175] It should be noted that, as used in this document, the terms "includes," "including," "has," "having," "contains," "containing," or any other variation thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements is not necessarily limited to only those elements, but can include other elements not expressly listed or inherent to such process, method, article, or apparatus.
[0176] The above merely describes the embodiments of the application, but the protection scope of the application is not limited thereto, and any person skilled in the art can easily think of changes or replacements within the technical scope disclosed by the application, which should be covered within the protection scope of the application.
Claims
1. A gaze point interaction method, comprising: in response to a first question and answer event, obtaining a first image corresponding to the first question and answer event and a first gaze point, the first gaze point being used to indicate a position at which a user gazes in the first image; determining that the first gaze point is located in a foreground region of the first image; performing smoothing processing on the first gaze point to obtain a second gaze point; obtaining a question and answer result corresponding to the first question and answer event based on a second image corresponding to the second gaze point and a visual question and answer model, the second image being used to represent a region in which the second gaze point is located in the first image.
2. The method of claim 1, wherein the performing smoothing processing on the first gaze point to obtain a second gaze point comprises: in a case where the first gaze point is in a first preset range, performing first smoothing processing on the first gaze point to obtain the second gaze point, a deviation between a position corresponding to the second gaze point and a position corresponding to the first gaze point being less than or equal to a first value; in a case where the first gaze point is in a second preset range, performing second smoothing processing on the first gaze point to obtain the second gaze point, a running track of the second gaze point conforming to visual inertia of user gaze, and a smoothing degree of the second smoothing processing being greater than a smoothing degree of the first smoothing processing.
3. The method of claim 2, further comprising: obtaining the first preset range according to a distance between the first gaze point and a center point of the foreground region; and obtaining the second preset range according to a distance between the first gaze point and an edge point of the foreground region.
4. The method of claim 2 or 3, wherein the performing first smoothing processing on the first gaze point to obtain the second gaze point in a case where the first gaze point is in the first preset range comprises: obtaining a first damping coefficient according to the first gaze point and the center point; and performing first smoothing processing on the first gaze point according to the first damping coefficient to obtain the second gaze point.
5. The method of claim 2 or 3, wherein the performing second smoothing processing on the first gaze point to obtain the second gaze point in a case where the first gaze point is in the second preset range comprises: obtaining a second damping coefficient according to the first gaze point and the edge point; and performing second smoothing processing on the first gaze point according to the second damping coefficient to obtain the second gaze point; and the second damping coefficient is greater than the first damping coefficient.
6. The method of claim 1, further comprising: determining that the first gaze point is located in a background region of the first image; and obtaining a question and answer result corresponding to the first question and answer event based on the first gaze point, a third image corresponding to the first gaze point, and the visual question and answer model, the third image being used to represent a region in which the first gaze point is located in the first image.
7. The method of claim 6, wherein the method further comprises, after the determining that the first gaze point is located in the background region of the first image: determine a running track of the first gaze point; obtain the second gaze point based on the first gaze point and the running track of the first gaze point.
8. The method of any one of claims 1-7, wherein the obtaining the answer result corresponding to the first question and answer event based on the second image corresponding to the second gaze point and the visual question and answer model comprises: determining a detection box corresponding to the second gaze point in the first image based on the second gaze point; performing cropping processing on the first image based on the detection box corresponding to the second gaze point to obtain a second image corresponding to the second gaze point; and inputting the second image corresponding to the second gaze point into the visual question and answer model to obtain the answer result corresponding to the first question and answer event.
9. The method of any one of claims 1-8, wherein the obtaining the first gaze point comprises: obtaining a fourth image, the fourth image comprising a head pose and an eye region of a user; and determining the first gaze point by performing gaze point identification on the head pose and the eye region in the fourth image using a line-of-sight tracking algorithm.
10. An electronic device comprising a display unit, a gaze point tracking unit, and a processor, wherein: the display unit is configured to display a first image; the gaze point tracking unit is configured to generate a first gaze point; and the processor is configured to, in response to a first question and answer event, obtain a first image corresponding to the first question and answer event and a first gaze point, the first gaze point indicating a position at which a user gazes in the first image; determine that the first gaze point is located in a foreground region of the first image; perform smoothing processing on the first gaze point to obtain a second gaze point; and obtain an answer result corresponding to the first question and answer event based on a second image corresponding to the second gaze point and a visual question and answer model, the second image representing a region in which the second gaze point is located in the first image.