Automatic task control method and device based on eye movement tracking and vehicle
The eye tracking technology is used to process user eye images, combine the multi-level space-time graph behavior representation module to learn user intentions and generate task rules, which solves the problems of insufficient perception and low intelligence of the on-board automated task control system, and realizes high-precision user intention understanding and personalized services.
Patent Information
- Application Number
- CN202510455514.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-07-11
AI Technical Summary
In the prior art, the on-board automation task control system has insufficient perception capabilities, low intelligence level, weak coordination ability, lack of security, and cannot provide precise and personalized services.
By obtaining the user's eye image, using infrared and visible camera fusion processing, combining three-axis acceleration and angular velocity data, eliminating jitter, calculating the gaze point area, and using a lightweight vision-language understanding model and multi-level spatio-temporal graph behavior representation module to learn user behavior patterns, generate task rules, match based on user intentions and issue control instructions.
It realizes high-precision user intention understanding, improves the intelligence level of on-board automation task control, enhances collaboration capabilities, and improves user experience and security.
Smart Images

Figure CN120295478A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent vehicles, and in particular, to an automated task control method, device, and vehicle based on eye movement tracking. Background Art
[0002] With the continuous improvement of the intelligent level of automobiles, in-vehicle systems have evolved from the initial simple information entertainment systems to comprehensive systems integrating multi-functional features such as personalized settings and automated task management. The demand of vehicle owners for personalized driving experiences is increasing day by day, especially the intelligent requirements in aspects such as automated task management, environmental adjustment, and device coordination are prominent. Providing intelligent services by analyzing the daily behavior habits of vehicle owners has become a key direction for enhancing user experience.
[0003] Currently, there are mainly three types of implementation solutions on the market: The first type is an automated system based on preset rules, which can adjust the seat position, air conditioning temperature, etc. according to fixed rules, but lacks learning ability; the second type is a simple habit memory system, which can record the user's commonly used settings and support one-key calling, but cannot understand and predict the user's needs; the third type is a primary machine learning system, although basic learning algorithms are introduced, its understanding of complex scenarios and multi-device coordination capabilities are limited.
[0004] The existing technical solutions have the following common problems: First, the perception ability is insufficient, the sensors are used singly, and it is difficult to comprehensively collect and analyze user behavior data; second, the intelligent level is low, the application of deep learning and reinforcement learning technologies is lacking, and accurate personalized services cannot be provided; third, the coordination ability is weak, the cross-device synchronization and multi-user sharing mechanisms are imperfect, affecting the use convenience; finally, the security is lacking, and the measures in aspects such as data encryption storage, access control, and privacy protection are not perfect enough to meet the user's requirements for data security. Summary of the Invention
[0005] In view of the above defects of the prior art, the present invention provides an automated task control method, device, and vehicle based on eye movement tracking to solve many technical problems such as low intelligent level, weak coordination ability, and inconvenience in automated task control in the prior art.
[0006] To achieve the above object and other related objects, the present invention provides an automated task control method based on eye movement tracking, including: obtaining an eye image of a user, and processing the eye image to obtain gaze point area data of the user; using a trained lightweight vision-language understanding model to understand the gaze point area data to obtain a user intention; based on the user intention, matching with the learned task rules to obtain a matching result; creating a task according to the matching result and issuing a control instruction.
[0007] In an embodiment of the present invention, obtaining an eye image of a user and processing the eye image to obtain gaze point region data of the user includes: obtaining an infrared image of the user's eyes using an infrared camera and obtaining a visible light image of the user's eyes using a visible light camera; obtaining a fused eye image according to the infrared image and the visible light image; obtaining triaxial acceleration and angular velocity data of a vehicle; processing the fused eye image according to the triaxial acceleration and angular velocity data and according to a preset mapping relationship between vibration and image movement to obtain a jitter-eliminated eye image; and obtaining the gaze point region data of the user according to the jitter-eliminated eye image.
[0008] In an embodiment of the present invention, obtaining the gaze point region data of the user according to the jitter-eliminated eye image includes: calculating the position of the user's eyeball and the pupil direction vector according to the jitter-eliminated eye image; projecting the pupil direction vector into a three-dimensional space model of a key interaction area in the vehicle constructed in advance according to the position of the user's eyeball to obtain the gaze point of the user; and obtaining the gaze point region data according to the gaze point of the user and the three-dimensional space model.
[0009] In an embodiment of the present invention, a Kalman filter is used to smooth the gaze point of the user, and a final gaze point of the user is obtained based on a focus confirmation mechanism of gaze duration.
[0010] In an embodiment of the present invention, the lightweight vision-language understanding model is obtained by using a complete BLIP model as a teacher model to train a lightweight student model, and is aligned through an attention mechanism to ensure that the lightweight student model focuses on the same key areas as the teacher model.
[0011] In an embodiment of the present invention, the task rules are learned according to the following steps: learning the behavior pattern of the user by using a multi-level spatio-temporal graph behavior representation module; and generating the task rules by using a ternary graph causal reasoning and task generation module according to the behavior pattern of the user.
[0012] In an embodiment of the present invention, the multi-level spatio-temporal graph behavior representation module includes a perception layer and a behavior layer; learning the behavior pattern of the user by using the multi-level spatio-temporal graph behavior representation module includes: obtaining the operation behavior of the user, the interaction behavior of the user within a preset time before the operation behavior, and the current state data of the vehicle by using the perception layer, where the interaction behavior at least includes the user's needs understood based on the eye image of the user; and aggregating the data obtained by the perception layer by using the behavior layer to obtain the behavior pattern of the user.
[0013] In an embodiment of the present invention, the current state data of the vehicle includes time information, position information, and environment information.
[0014] In an embodiment of the present invention, according to the user's behavior pattern, the task rules are generated by using the ternary graph causal reasoning and task generation module, including: constructing triples by using a three-layer spatio-temporal graph structure, where the triples include an execution subject, an automated task, and a task execution condition; configuring a confidence level and a frequency for each of the triples according to the user's behavior pattern; and obtaining the task rules according to each of the triples and its corresponding confidence level and frequency.
[0015] In an embodiment of the present invention, based on the user intention, it is matched with the learned task rules to obtain a matching result, including: obtaining the current state data of the vehicle; based on the user intention, traversing the task execution conditions of each of the triples to obtain a number of triples including the user intention; based on the current state data of the vehicle, traversing the task execution conditions of the number of triples including the user intention to obtain a number of triples that match the current state of the vehicle; and obtaining a number of matching triples according to the confidence levels and frequencies of the number of triples that match the current state of the vehicle, as well as a preset confidence level threshold and frequency threshold.
[0016] In an embodiment of the present invention, a task is created and a control instruction is issued according to the matching result, including: creating an automatic task and an inquiry task according to the automated tasks of the number of matching triples and the confidence levels and frequencies of these triples; issuing the control instruction according to the automatic task; presenting the inquiry task to the user, and issuing the control instruction according to the response result of the user to the inquiry task.
[0017] To achieve the above object and other related objects, the present invention also provides an automated task control device based on eye movement tracking, including: an image processing unit, configured to obtain an eye image of a user and process the eye image to obtain the fixation point area data of the user; an intention processing unit, configured to use a trained lightweight vision-language understanding model to understand the fixation point area data to obtain a user intention; a matching unit, configured to match the user intention with the learned task rules to obtain a matching result; and an instruction generation unit, configured to create a task and issue a control instruction according to the matching result.
[0018] To achieve the above object and other related objects, the present invention also provides a vehicle, including the automated task control device based on eye movement tracking as described above.
[0019] Advantages of the present invention: An automated task control method, device, and vehicle based on eye movement tracking proposed by the present invention start from the user's eye image, obtain the user's fixation point area through processing the eye image, and obtain the final matching result through intelligent analysis of the fixation point area, thereby realizing the automatic creation and execution of tasks. Since the task is triggered by the user's fixation area, a task trigger mechanism is introduced, avoiding the inaccurate intention judgment caused by fully automated tasks and further improving the user experience. Description of the Drawings
[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0021] Figure 1 The first flowchart of the task control method provided by an embodiment of the present invention;
[0022] Figure 2 The detailed flowchart of step S100 provided by an embodiment of the present invention;
[0023] Figure 3 The detailed flowchart of step S150 provided by an embodiment of the present invention;
[0024] Figure 4 The structural diagram of the multi-band adaptive eye movement tracking module provided by an embodiment of the present invention;
[0025] Figure 5 The architecture diagram of the lightweight vision-language understanding model provided by an embodiment of the present invention;
[0026] Figure 6 The learning flowchart of the task rules provided by an embodiment of the present invention;
[0027] Figure 7 The detailed flowchart of step S310 provided by an embodiment of the present invention;
[0028] Figure 8 The detailed flowchart of step S320 provided by an embodiment of the present invention;
[0029] Figure 9 The structural diagram of the ternary graph causal reasoning and task generation module provided by an embodiment of the present invention;
[0030] Figure 10Detailed flowchart of step S300 provided by an embodiment of the present invention;
[0031] Figure 11 Detailed flowchart of step S400 provided by an embodiment of the present invention;
[0032] Figure 12 Second flowchart of the task control method provided by an embodiment of the present invention;
[0033] Figure 13 Schematic diagram of the task control device provided by an embodiment of the present invention.
[0034] Explanation of reference numerals: 101, image processing unit; 102, intention processing unit; 103, matching unit; 104, instruction generation unit. Detailed implementation manners
[0035] The following describes the embodiments of the present invention through specific specific examples. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other. Except for the specific methods, devices, and materials used in the embodiments, according to the knowledge of those skilled in the art in the prior art and the description of the present invention, any methods, devices, and materials similar or equivalent to those described in the embodiments of the present invention can also be used to implement the present invention.
[0036] It should be understood that the terms used in the embodiments of the present invention are for the purpose of describing specific specific implementation manners, rather than for limiting the protection scope of the present invention. Unless otherwise defined, all technical and scientific terms used in the present invention have the same meaning as commonly understood by those skilled in the art of this technical field.
[0037] In the following description, a large number of details are explored to provide a more thorough explanation of the embodiments of the present invention. However, it is obvious to those skilled in the art that the embodiments of the present invention can be implemented without these specific details. In some of these embodiments, well-known structures and devices are shown in the form of block diagrams rather than in detail to avoid making the embodiments of the present invention difficult to understand.
[0038] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of methods and computer program products that may be implemented according to various embodiments disclosed in the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system that performs the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0039] Please refer to Figure 1 , Figure 1 An automated task control method based on eye tracking provided for an embodiment of the present invention includes steps S100 to S400.
[0040] Step S100: Obtain an eye image of the user and process the eye image to obtain data on the user's fixation point area. In the present invention, the user's fixation point is obtained based on the user's eye image, and subsequent tasks are automatically executed based on this. Therefore, first, it is necessary to obtain the user's eye image and process it to obtain data on the user's fixation point area.
[0041] The data on the user's fixation point area is used to characterize the specific information of the user's fixation point. For example, it may be an image of the fixation point area, or the specific position information of the fixation point, or the position code of the fixation point, etc. In a specific embodiment of the present invention, the data on the user's fixation point area is an image of the fixation point area. The advantage of selecting an image is that on the one hand, the position of the user's fixation point can be known, and on the other hand, the state of the interactive component in the user's fixation point area can also be obtained, providing a more accurate basis for understanding the user's intention in the subsequent process.
[0042] Please refer to Figure 2 , in a specific embodiment of the present invention, step S100 includes steps S110 to S150.
[0043] Step S110: Use an infrared camera to obtain an infrared image of the user's eyes, and use a visible light camera to obtain a visible light image of the user's eyes. The reason for choosing two cameras is to obtain the user's eye images more accurately. The infrared camera is mainly used for eye movement tracking in low-light environments, and its output data can be, for example, an infrared grayscale image sequence, or a high-contrast image of the pupil and iris in low-light environments. The visible light camera is mainly used for scene understanding under normal lighting conditions, and its output data can be, for example, a high-definition color image sequence.
[0044] Step S120: Obtain a fused eye image based on the infrared image and the visible light image. In this step, the two images can be fused together through a preset weight ratio, or the weights of the two images can be dynamically adjusted according to the environmental lighting conditions through an image fusion algorithm to generate an optimal eye image. This can obtain a relatively clear eye image under different environmental lighting conditions, providing a good basis for subsequent eye movement data analysis.
[0045] Step S130: Obtain the three-axis acceleration and angular velocity data of the vehicle. This data can be obtained, for example, through a three-axis accelerometer and a three-axis gyroscope. The reason for obtaining the vehicle's motion data is to compensate for the vehicle's vibration to eliminate the interference of vehicle motion on eye movement data.
[0046] Step S140: Process the fused eye image according to the three-axis acceleration and angular velocity data and according to the preset mapping relationship between vibration and image movement to obtain a jitter-eliminated eye image. In this step, the mapping relationship between vibration and image movement can be obtained through experiments and saved in the system as a known quantity. Subsequently, the fused eye image can be directly processed according to the three-axis acceleration and angular velocity data of the vehicle obtained in Step S130 in combination with this mapping relationship to eliminate jitter.
[0047] Step S150: Obtain the user's fixation point area data based on the jitter-eliminated eye image. Theoretically, the user's fixation point area data can also be obtained based on the fused eye image, but it may be inaccurate due to vehicle jitter. To further improve the accuracy of subsequent processing, it is best to eliminate jitter first and then perform image processing to obtain more accurate user fixation point area data.
[0048] Please refer to Figure 3 , in a specific embodiment of the present invention, Step S150 includes Steps S151 to S153.
[0049] Step S151: Calculate the user's eyeball position and pupil direction vector based on the eye image after jitter elimination. In this step, for example, the following processing can be performed: (1) First, identify the pupil boundary through the ellipse fitting algorithm; (2) Perform corneal reflection detection by identifying the reflection points of the light source on the eyeball surface, and obtain stable reference points through extraction of the eye corner feature points; (3) Estimate the eyeball center based on the anatomical model and the geometric relationship of the feature points to obtain the eyeball position information; (4) Then, obtain the line-of-sight vector, i.e., the pupil direction vector, according to the vector from the eyeball center to the pupil center.
[0050] Step S152: Project the pupil direction vector into the three-dimensional space model of the key in-vehicle interaction area pre-constructed according to the user's eyeball position to obtain the user's fixation point. The three-dimensional space model of the key in-vehicle interaction area can be pre-constructed. When specifically constructing it, for example, a depth sensor (based on structured light or ToF technology) can be used to construct the in-vehicle three-dimensional space model. It should be noted that visible light sensors and infrared sensors are generally set in front of the user to capture the user's facial image; the depth sensor is different, and it is generally set obliquely behind the user to capture the key in-vehicle interaction area, such as areas like the central control screen, instrument panel, control buttons, steering wheel, driver and passenger doors, etc.
[0051] It can be understood that the construction of the above three-dimensional space model can be carried out by using the depth sensor to comprehensively scan and model the entire cockpit when the vehicle is first started, or the three-dimensional space model of the entire cockpit constructed during vehicle design can be imported. If scanning and modeling are performed through the depth sensor, for example, the following steps can be used to achieve it: First, downsample the point cloud data collected by the depth sensor, which can reduce the data volume and improve the processing efficiency; second, perform noise filtering to remove outliers and measurement errors; finally, generate a continuous surface through an algorithm (such as Poisson surface reconstruction) to achieve surface reconstruction, thereby obtaining the three-dimensional space model of the key in-vehicle interaction area.
[0052] When constructing the three-dimensional space model of the key in-vehicle interaction area, a unique ID and function label can be assigned to each key in-vehicle interaction area, and a new space coordinate system can be established with the driver's seat as the center, and then the bounding box of the interaction area can be constructed to obtain its coordinate information. After such processing, the key in-vehicle interaction area can include multiple pieces of information: ID, function label, position information in the new space coordinate system, and so on.
[0053] Step S153: Obtain the fixation point area data according to the user's fixation point and the three-dimensional space model.
[0054] In a specific embodiment of the present invention, a Kalman filter is used to smooth the user's fixation point, and based on a focus confirmation mechanism of the fixation duration, the user's final fixation point is obtained. According to each captured user face image, the fixation point corresponding to the image can be obtained; by processing a series of consecutive images, a series of user fixation points can be obtained. Generally speaking, the changes in the user's fixation points have a certain continuity. If there are some sudden change behaviors, it is very likely to be noise or error. Therefore, a Kalman filter can be used to smooth the user's fixation points. In addition, a parameter such as the fixation duration can be introduced for focus confirmation, which can distinguish the user's unconscious saccade behavior and focus on the user's conscious fixation behavior, so as to obtain a more accurate user fixation point.
[0055] Please refer to Figure 4 , the above-mentioned various sensors and the subsequent data processing process can be used as a whole to form a multi-band adaptive eye movement tracking module, which includes a data acquisition layer, a processing layer, and an analysis layer. Each layer also includes several units to achieve the function of eye movement tracking. Although the expressions are different, the overall meaning is still certain.
[0056] Step S200: Use the trained lightweight vision-language understanding model to understand the data in the fixation point area and obtain the user's intention. In order to fully understand the data in the user's fixation point area, in the present invention, the data in the fixation point area is processed by the trained lightweight vision-language understanding model.
[0057] The lightweight vision-language understanding model can be obtained, for example, by using the complete BLIP model (Bootstrapping Language-Image Pre-training) as the teacher model and training a lightweight student model, as Figure 5 shown, and designing a distillation strategy for the in-vehicle scenario, focusing on retaining the ability to understand driving-related content, and aligning through the attention mechanism to ensure that the lightweight student model focuses on the same key areas as the teacher model. Different quantization strategies are applied to different network layers. For example, the key layer maintains FP16, and the non-key layer is reduced to INT8; and the model size is reduced by implementing weight sharing and sparse representation. In terms of the model structure, for example, a simplified version of Transformer can be designed to reduce the number of attention heads and layers, reduce the computational overhead; and use group convolution to replace the standard convolution to reduce the number of parameters; a confidence-based early termination strategy can also be introduced to end the inference in advance in high-confidence scenarios and speed up the inference speed.
[0058] To make the model more applicable to the scenarios involved in the present invention, for example, vehicle-mounted scenario data can be used for transfer learning to enhance the recognition ability of in-vehicle objects, controls, and states; a vehicle-mounted exclusive vocabulary and semantic representation can also be constructed to optimize the language understanding component, thereby improving the accurate processing ability of driving-related terms and instructions. Through these processes, while maintaining excellent understanding accuracy, the size of the lightweight vision-language understanding model is greatly reduced from the original, the inference speed is increased, and the memory occupancy is reduced, meeting the in-vehicle real-time processing requirements.
[0059] Step S300: Based on the user intention, match it with the learned task rules to obtain a matching result. In steps S100 and S200, it mainly introduced how to obtain the user intention through the user's facial image. Before learning the task rules, the task rules are learned by obtaining the user intention and the user's actual operations; after the task rules are learned, the real-time obtained user intention can be used for matching to obtain a specific matching result. In the present invention, it is precisely through the processing of the user's eye data that the user's fixation point is associated with subsequent operations, establishing the relationship between the user's fixation point and subsequent operations (i.e., the task rules to be learned). After such processing, when the user wants to issue a certain control instruction later, only by looking at the corresponding interaction area can the subsequent automated task be triggered.
[0060] Please refer to Figure 6 , in a specific embodiment of the present invention, the task rules are learned according to steps S310 to S320.
[0061] Step S310: Use the multi-level spatio-temporal graph behavior representation module to learn the user's behavior pattern. The task rules can be, for example, "User A, turn on the air conditioner and lower the temperature. When starting the vehicle at 8-9 am on weekdays, and the ambient temperature is greater than or equal to X °C, and the in-vehicle temperature is greater than or equal to Y °C, and the user wants to adjust the temperature (this condition is the user intention, obtained by processing the user's eye image), confidence level 85%, frequency 15". This is a long-term behavior habit. To learn such task rules, the user's behavior pattern needs to be learned first, and then the behavior patterns with more repeated occurrences are constructed into task rules.
[0062] Please refer to Figure 7 , in a specific embodiment of the present invention, the multi-level spatio-temporal graph behavior representation module includes a perception layer and a behavior layer, and step S310 includes steps S311 and S312.
[0063] Step S311, using the perception layer to obtain the user's operation behavior, the user's interactive behavior within a preset time before the operation behavior, and the vehicle's current state data, the interactive behavior at least includes the user's needs understood based on the user's eye image. When learning the user's behavior pattern, it is based on the user's specific operation behavior, that is, the present invention focuses on each user's operation behavior, and then obtains the user's interactive behavior and the vehicle's current state data within a preset time, providing a basis for the subsequent behavior layer pattern analysis.
[0064] The user's operating behavior may be, for example, turning on the air conditioner, or more specifically adjusting the air conditioner temperature to X°C; the preset time may be, for example, 5 to 60 seconds; the interactive behavior may include, for example, a demand such as "the user wants to adjust the temperature" (this demand is obtained based on the user's action of looking at the air conditioner adjustment area), and the user's instructions to adjust the air conditioner temperature (voice instructions, button instructions, or touch instructions); the vehicle's current status data may include, for example, time information (such as weekdays / holidays, or specific dates, or specific times, or seasons, etc.), location information (such as specific geographic location information, or virtual locations similar to residence / company, etc.) and environmental information (such as weather, or traffic conditions, or ambient temperature, or temperature inside the vehicle, or light intensity, etc.).
[0065] Step S312: Use the behavior layer to aggregate the data obtained by the perception layer to obtain the user's behavior pattern. Through aggregation, a number of user operation behaviors can be refined into specific behavior patterns. For example, the behavior layer can use a time series diagram model to describe the order of occurrence of behaviors and their causal relationships. The behavior layer can also add a weighted attenuation mechanism. For example, recent operation behaviors have a greater influence, and the weight of long-term operation behaviors gradually decreases to ensure that only recent user operation behaviors are focused on.
[0066] Step S320: Generate task rules based on the user's behavior pattern using the ternary graph causal reasoning and task generation module. After obtaining the user's behavior pattern, specific task rules can be constructed based on it.
[0067] See also Figure 8 In a specific embodiment of the present invention, step S320 includes steps S321 to S323.
[0068] Step S321: Use a three-layer spatiotemporal graph structure to construct a triplet, which includes an execution subject, an automated task, and a task execution condition. The construction of a triplet can be understood as knowledge modeling. By constructing a high-precision behavior graph, such as (user A-lower the temperature-hot weather), the system can understand the interactive relationship between user behavior and external context.
[0069] Step S322: Configure the confidence and frequency for each ternary component according to the user's behavior pattern. The purpose of this step is to ensure that the high-confidence mode is preferentially adopted when generating tasks.
[0070] Step S323: Obtain the task rules based on each triple and its corresponding confidence and frequency. The triple corresponds to the automated task that the executor will perform under a certain task execution condition, but only expresses the relationship among the three. In fact, for the task rules, it is also necessary to determine the confidence and frequency of the triple, that is, the confidence and frequency of the automated task that the executor will perform.
[0071] The structure of the ternary graph causal reasoning and task generation module is as Figure 9 shown. In addition to the above basic functions, it can use causal reasoning to replace traditional correlation analysis, and the lack of east system identifies real causal relationships. For example, it will apply the counterfactual intervention framework to simulate the situation of "if the user does not perform a certain behavior, whether the result will be different", avoiding making wrong inferences based solely on data correlation; it can also implement a causal discovery algorithm based on time series, considering the time sequence of user behavior to identify long-term causal relationships; it can also construct a causal Bayesian network to calculate the influence weights of different factors on task generation, helping the system understand complex causal chains.
[0072] When constructing task rules, a common task template library can be predefined, including high-frequency tasks such as environmental control (such as automatically adjusting the air conditioner), information services (such as news broadcasts), and driving assistance (such as route recommendations). A parameterization mechanism can also be designed, and the system dynamically fills task parameters according to the user's personalized habits. For example, "turn on the air conditioner and lower the temperature" in the above-mentioned task rules can be specifically changed to "adjust the air conditioner temperature to X °C".
[0073] In addition to learning task rules according to the user's behavior pattern, the confidence and frequency of the learned task rules can also be adjusted according to the user's response to the recommended tasks. In this way, continuous learning and improvement of task rules can be achieved.
[0074] Please refer to Figure 10 , in a specific embodiment of the present invention, step S300 includes steps S301 to S304.
[0075] It can be understood that the above steps S311, S312, and S321 to S323 mainly introduce how to generate task rules, which correspond to the learning process of task rules. After learning these task rules, subsequent tasks can be executed according to the user's intention. Steps S301 to S304 introduce how to obtain the matching results according to the user's intention.
[0076] Step S301: Obtain the current vehicle state data. The current vehicle state data has been described above, and its main purpose is for the judgment in step S303.
[0077] Step S302: Based on the user intention, traverse the task execution conditions of each triple to obtain several triples containing the user intention. After a lot of triples are constructed through step S320, according to the user intention, first find the triples that can match the user intention. For example, if the user looks at the air conditioner adjustment area and after recognition, the user intention is "the user wants to adjust the air conditioner temperature", then traverse all the triples and filter out the triples containing this intention.
[0078] Step S303: Based on the current vehicle state data, traverse the task execution conditions of several triples containing the user intention to obtain several triples that match the current vehicle state. The purpose of this step is to select the triples that match the current vehicle state data. Suppose two triples are obtained through step S302, which are "User A, turn on the air conditioner and lower the temperature, when starting the vehicle at 8 - 9 am on weekdays, and the ambient temperature is greater than or equal to 30°C, and the interior temperature is greater than or equal to 25°C, and the user wants to adjust the temperature" and "User A, turn on the air conditioner and raise the temperature, when starting the vehicle at 8 - 9 am on weekdays, and the ambient temperature is less than or equal to 10°C, and the interior temperature is less than or equal to 5°C, and the user wants to adjust the temperature". Although both of these two triples contain the same user intention, if the ambient temperature in the obtained current vehicle state data is 32°C and the interior temperature is 26°C, then only the first triple will be retained after step S303.
[0079] Step S304: According to the confidence levels and frequencies of several triples that match the current vehicle state, as well as the preset confidence level threshold and frequency threshold, obtain several matching triples. The reason for introducing the confidence level and frequency is to perform related tasks more intelligently. For example, the confidence level threshold can be set to 60% and the frequency threshold can be set to 5, so that those triples with relatively low confidence levels and / or frequencies can be excluded. And when performing task execution or recommendation later, reasons can be attached, such as "It is recommended to turn on the air conditioner because you have manually adjusted the temperature 7 times recently under similar temperatures". The confidence levels and frequencies of the triples can also support manual adjustment by the user to make it more in line with the user's needs.
[0080] In the above matching process, the same user is taken as an example for illustration. In fact, the judgment of the user identity can also be introduced in the above steps. In a specific embodiment of the present invention, the user identity is also recognized in step S100. Before step S302, the execution subjects can be traversed first to exclude those triples whose execution subjects are not the current user.
[0081] Step S400: Create a task based on the matching result and issue a control instruction. After the processing in the above steps S301 to S304, a matching result can be obtained. The matching result is the task rule that matches the user's intention, and the task rule includes information such as triples, confidence levels, and frequencies. The triples include the execution subject, the automated task, and the task execution conditions. Among them, the execution subject, the task execution conditions, the confidence level, and the frequency are mainly used for matching. The matching result obtained after matching can include only the automated task.
[0082] Please refer to Figure 11 , in a specific embodiment of the present invention, the matching result includes at least the automated task, the confidence level, and the frequency. Step S400 includes: S401: Create an automatic task and an inquiry task according to the automated tasks of several matching triples, as well as the confidence level and frequency of the triples; S402: Issue a control instruction according to the automatic task; S403: Display the inquiry task to the user and issue a control instruction according to the response result of the user to the inquiry task.
[0083] In this embodiment, tasks are executed in different ways through the confidence level and frequency. Here, the confidence level is taken as an example for illustration. For example, if the confidence level of the automated task in a certain task rule is greater than 90%, it can be used as an automatic task and directly executed by the system without user confirmation; the automated tasks in the task rules with a confidence level between 60% and 90% are used as inquiry tasks, that is, the user is asked whether to execute the task through a pop-up window / voice, and then it is judged whether to execute according to the user's response result (touch / voice / nod, etc.). It can be understood that if the user refuses an inquiry task multiple times, the system will reduce its confidence level to ensure that the recommended content meets the user's wishes.
[0084] It should be noted that the step division of the above various methods is only for clear description. When implemented, they can be combined into one step or some steps can be split into multiple steps. As long as they contain the same logical relationship, they are all within the protection scope of this application; adding insignificant modifications to the algorithm or process or introducing insignificant designs, but not changing the core design of its algorithm and process are all within the protection scope of this patent. For example Figure 12 shows the flowchart of another automated task control method. Although the steps in this flowchart are different from the above step descriptions, their overall ideas are the same.
[0085] In a specific embodiment of the present invention, the automated task control method based on eye movement tracking includes the following workflow.
[0086] 1. Initialization stage:
[0087] (1.1) Load the lightweight model and user configuration when the system starts;
[0088] (1.2) Initialize the multi - band camera and eye - tracking module;
[0089] (1.3) Connect to the vehicle bus through CAN / FlexRay / Ethernet protocol to achieve intelligent control of in - vehicle devices.
[0090] 2. User identification and configuration loading:
[0091] (2.1) Ensure accurate identification and loading of the user's specific behavior model and preference settings through three methods: facial recognition, voice authentication, or user login;
[0092] (2.2) Prepare a personalized interaction interface, load parameters such as the user's specific driving habits, entertainment preferences, and temperature control settings, predict the functions to be used based on the historical behavior model, pre - load relevant task modules, and optimize the interface layout through intelligent UI adjustment to ensure that the driver can quickly access the core functions.
[0093] 3. Real - time data collection and processing:
[0094] (3.1) Eye - tracking data capture: Identify the fixation point, blink frequency, and gaze duration, evaluate driving concentration, and continuously capture the user's line - of - sight data through the eye - tracking module;
[0095] (3.2) Combine vehicle status (speed, steering wheel angle), environmental information (temperature, humidity, road conditions outside the vehicle), and user operations (touch - screen interaction, voice commands) to construct a complete user behavior data set for multi - modal data pre - processing and fusion;
[0096] (3.3) Use time - series denoising and outlier detection to ensure data accuracy and reduce the misjudgment rate.
[0097] 4. Behavior understanding and pattern recognition:
[0098] (4.1) The lightweight BLIP model processes visual inputs, identifies the user's line - of - sight focus, infers the current attention object (such as the navigation screen, air - conditioner panel), and understands the user's current behavior;
[0099] (4.2) Through short - term + long - term behavior modeling, continuously optimize the user profile to identify the matching degree between the current behavior and historical patterns, adopt an incremental learning strategy, dynamically update the model, and improve the prediction accuracy;
[0100] (4.3) Calculate the similarity between the current behavior and historical patterns to determine whether the user is repeating a certain task (such as turning up the in - vehicle temperature every morning).
[0101] 5. Causal reasoning and task generation:
[0102] (5.1) Adopt the "user - behavior - context" structure to infer user intentions, such as "temperature drops - driver has not adjusted for a long time - automatic heating up";
[0103] (5.2) Pre - define task templates (air - conditioner adjustment, entertainment recommendation, navigation adjustment), fill in task parameters according to the current context, such as automatically adjusting the temperature to 22°C, and generate potential automated tasks and parameters;
[0104] (5.3) Combine historical data to calculate the confidence level of recommended tasks. High - confidence tasks (execute automatically), medium - confidence tasks (provide recommendations, users can adjust), and low - confidence tasks (not recommended for the time being).
[0105] 6. Resource Scheduling and Task Execution:
[0106] (6.1) The adaptive graph neural network computing module allocates computing resources. High - priority tasks (such as safety warnings) are preferentially allocated GPU computing, and low - priority tasks (such as background data analysis) can be deferred;
[0107] (6.2) Automatically execute high - confidence tasks, such as "when the outside temperature is below 5°C, automatically turn on the seat heating", and users confirm medium - and low - confidence tasks, such as "Do you want to enable the navigation route of the last commute?";
[0108] (6.3) Monitor the task execution situation and detect whether the expected effect is achieved, such as whether the driver further modifies the settings after the temperature adjustment.
[0109] 7. Feedback Collection and Optimization:
[0110] (7.1) Monitor the task execution results and user reactions, such as "After automatically adjusting the temperature, does the user further adjust?";
[0111] (7.2) Collect explicit and implicit feedback and update the causal model;
[0112] (7.3) Adopt incremental learning to optimize the causal reasoning and task recommendation models and improve the prediction accuracy.
[0113] 8. Cloud Synchronization and Sharing:
[0114] (8.1) Allow users to synchronize automated tasks between different vehicles, such as "automatically play a certain radio station when starting the vehicle";
[0115] (8.2) Allow users to share personalized settings with family or friends, such as driving habits and air - conditioner preferences;
[0116] (8.3) Adopt a local - first storage strategy and only upload task data with the user's consent.
[0117] Please refer toFigure 13 , Figure 13 An automated task control device based on eye tracking provided by an embodiment of the present invention includes an image processing unit 101, an intention processing unit 102, a matching unit 103, and an instruction generation unit 104. Among them, the image processing unit 101 is used to obtain the user's eye image and process the eye image to obtain the data of the user's fixation point area; the intention processing unit 102 is used to use the trained lightweight vision-language understanding model to understand the data of the fixation point area and obtain the user's intention; the matching unit 103 is used to match the user's intention with the learned task rules to obtain a matching result; the instruction generation unit 104 is used to create a task according to the matching result and issue a control instruction.
[0118] It should be noted that the task control device in this embodiment is a device corresponding to the above task control method, and the functional modules in the task control device respectively correspond to the corresponding steps in the task control method. The task control device in this embodiment can be implemented in cooperation with the task control method. That is, without conflict, the relevant technical details mentioned in the task control method in the above embodiment can also be applied to the task control device in this embodiment.
[0119] An embodiment of the present invention also provides a vehicle, which includes the above-mentioned automated task control device based on eye tracking.
[0120] Generally speaking, in the present invention, task rules are first learned based on information such as the user's eye image, operation behavior, and vehicle current state data, and then the user's intention is understood based on a certain eye movement made by the user, and by matching with the learned task rules, a suitable task is created and a control instruction is issued to execute. Through core capabilities such as multi-modal perception, causal reasoning, adaptive computing, and task scheduling, an efficient, accurate, and personalized in-vehicle automated experience is achieved. As the user's usage time increases, the system will continuously optimize the behavior model to make task recommendations more intelligent, efficient, and in line with actual needs.
[0121] The above embodiments are only illustrative of the principles and effects of the present invention, and are not used to limit the present invention. Any person familiar with this technology can modify or change the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or changes completed by those with ordinary knowledge in the technical field without departing from the spirit and technical ideas disclosed by the present invention should still be covered by the claims of the present invention.
Claims
1. An automated task control method based on eye movement tracking, characterized in that, including: Obtain the user's eye image, and process the eye image to obtain the user's fixation point area data; Use the trained lightweight vision-language understanding model to understand the fixation point area data and obtain the user's intention; Based on the user's intention, match it with the learned task rules to obtain a matching result; Create a task according to the matching result and issue a control instruction.
2. The automated task control method based on eye movement tracking according to claim 1, wherein, Obtain the user's eye image, and process the eye image to obtain the user's fixation point area data, including: Use an infrared camera to obtain an infrared image of the user's eyes, and use a visible light camera to obtain a visible light image of the user's eyes; Obtain a fused eye image according to the infrared image and the visible light image; Obtain the vehicle's three-axis acceleration and angular velocity data; According to the three-axis acceleration and angular velocity data, and according to the preset mapping relationship between vibration and image movement, process the fused eye image to obtain a jitter-eliminated eye image; Obtain the user's fixation point area data according to the jitter-eliminated eye image.
3. The automated task control method based on eye movement tracking according to claim 2, wherein, Obtain the user's fixation point area data according to the jitter-eliminated eye image, including: Calculate the user's eyeball position and pupil direction vector according to the jitter-eliminated eye image; Project the pupil direction vector into the three-dimensional space model of the key in-vehicle interaction area pre-constructed according to the user's eyeball position to obtain the user's fixation point; Obtain the fixation point area data according to the user's fixation point and the three-dimensional space model. Use a Kalman filter to smooth the user's fixation point, and based on the focus confirmation mechanism of the fixation duration, obtain the user's final fixation point. The lightweight vision-language understanding model is obtained by using the complete BLIP model as the teacher model to train the lightweight student model, and through the attention mechanism alignment, ensure that the lightweight student model focuses on the same key areas as the teacher model.
4. The automated task control method based on eye movement tracking according to claim 1, wherein The task rules are learned according to the following steps: Use the multi-level spatio-temporal graph behavior representation module to learn the user's behavior pattern; According to the user's behavior pattern, use the tripartite graph causal reasoning and task generation module to generate the task rules.
5. The automated task control method based on eye movement tracking according to claim 4, wherein, The multi-level spatio-temporal graph behavior representation module includes a perception layer and a behavior layer; Use the multi-level spatio-temporal graph behavior representation module to learn the user's behavior pattern, including: Use the perception layer to obtain the user's operation behavior, the user's interaction behavior within a preset time before the operation behavior, and the vehicle's current state data, and the interaction behavior at least includes the user's needs understood based on the user's eye image; Use the behavior layer to aggregate the data obtained by the perception layer to obtain the user's behavior pattern. The vehicle's current state data includes time information, position information, and environmental information.
6. The automated task control method based on eye movement tracking according to claim 4, wherein According to the user's behavior pattern, use the tripartite graph causal reasoning and task generation module to generate the task rules, including: Construct a triple using a three-layer spatio-temporal graph structure, and the triple includes an execution subject, an automated task, and task execution conditions; Configure confidence and frequency for each of the ternary components according to the user's behavior pattern; Obtain the task rules according to each of the triples and their corresponding confidence and frequency; 7. The automated task control method based on eye movement tracking according to claim 6, wherein Based on the user intent, match with the learned task rules to obtain a matching result, including: Obtain the current state data of the vehicle; Based on the user intent, traverse the task execution conditions of each of the triples to obtain several triples containing the user intent; Based on the current state data of the vehicle, traverse the task execution conditions of several triples containing the user intent to obtain several triples that match the current state of the vehicle; Obtain several matching triples according to the confidence and frequency of several triples that match the current state of the vehicle, as well as a preset confidence threshold and frequency threshold; 8. The automated task control method based on eye movement tracking according to claim 7, wherein, Create a task and issue a control instruction according to the matching result, including: Create an automatic task and an inquiry task according to the automatic tasks of several matching triples and the confidence and frequency of the triple; Issue the control instruction according to the automatic task; Display the inquiry task to the user and issue the control instruction according to the response result of the user to the inquiry task.
9. An automated task control device based on eye movement tracking, characterized in that, Including: An image processing unit for obtaining an eye image of the user and processing the eye image to obtain the fixation point area data of the user; An intent processing unit for using a trained lightweight vision-language understanding model to understand the fixation point area data to obtain the user intent; A matching unit for matching with the learned task rules based on the user intent to obtain a matching result; An instruction generation unit for creating a task and issuing a control instruction according to the matching result.
10. A vehicle, characterized in that, Including the eye movement tracking-based automatic task control device according to claim 9.