A multimedia interaction control method and system based on multi-modal vision and adaptive line-of-sight stabilization determination

By using multimodal vision and adaptive gaze stabilization determination technology, the problems of poor adaptability and false triggering of gaze interaction in near and far fields are solved, achieving high-precision anti-shake, multi-target locking and low false touch contactless interaction, improving the user experience and reliability of the device.

CN122387307APending Publication Date: 2026-07-14
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Filing Date
2026-04-16
Publication Date
2026-07-14

AI Technical Summary

Technical Problem

Existing gaze interaction technologies have poor adaptability to near and far field scenarios, are easily affected by physiological micro-movements and environmental interference, resulting in a high rate of false triggering and an inability to effectively lock onto the main interaction object. In particular, the reliability and practicality of interaction are insufficient in multi-target scenarios.

Method used

A multimodal vision and adaptive gaze stabilization method is adopted. Through image acquisition, signal processing and interactive control, combined with near and far field adaptive switching, sliding window variance calculation and dynamic anti-shake threshold adjustment, high-precision anti-shake and multi-target stable locking are achieved. Combined with blink confirmation and eye movement trajectory judgment, it can distinguish between unintentional saccades and real interaction intentions.

Benefits of technology

It significantly reduces the probability of false triggering, improves the stability and accuracy of interaction, adapts to complex scenarios, and achieves adaptive, high-precision, and low-latency contactless multimedia control, suitable for public displays, advertising terminals, and vehicle-mounted equipment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122387307A_ABST
    Figure CN122387307A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of computer vision, and discloses a multimedia interaction control method based on multimodal vision and adaptive line-of-sight stability determination, and the operation steps are as follows: step S1: acquiring a video image and performing target detection to extract the distance between a user and a device and the like. The multimedia interaction control method and system based on multimodal vision and adaptive line-of-sight stability determination adaptively switch far and near field line-of-sight estimation modes through a face area, take into account far distance head posture estimation and near distance pupil accurate detection, greatly improve interaction adaptation capability and detection precision under different use distances, adopt joint judgment of a sliding window variance and a convergence trend, combine a dynamic adaptive anti-shake threshold value, can effectively inhibit line-of-sight jitter caused by physiological micro-motion, environmental noise and light change, significantly reduce the false triggering probability, make the gaze determination more stable and reliable, and overall improve the fluency and accuracy of touchless interaction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer vision technology, specifically to a multimedia interactive control method and system based on multimodal vision and adaptive gaze stability determination. Background Technology

[0002] With the popularization of public multimedia displays, smart exhibition halls and in-vehicle interactive devices, contactless eye-tracking interaction has become an important direction for improving user experience. Existing eye-tracking interaction technologies mostly rely on pupil detection at a fixed distance and do not have adaptive designs for near and far field scenarios.

[0003] Meanwhile, in high-value commercial scenarios such as elevator multimedia advertising interaction, shopping mall smart wayfinding screens, and contactless self-service vending machines, the lack of effective anti-shake and action confirmation mechanisms makes the system highly susceptible to interference from human physiological micro-movements, head shaking, and ambient light. This results in the system being unable to distinguish between the user's "unintentional glance (mis-touch)" and "real interaction intent." Furthermore, when multiple targets appear simultaneously, the system cannot effectively lock onto the main interaction object, further reducing the reliability and usability of the interaction. Summary of the Invention

[0004] (a) Technical problems to be solved

[0005] To address the shortcomings of existing technologies, this invention provides a multimedia interactive control method and system based on multimodal vision and adaptive gaze stability determination. It has advantages such as near-field and far-field adaptation, high-precision anti-shake, multi-target stable locking, accurate gaze determination, low false trigger rate, and strong environmental adaptability. It solves the problems of poor near-field and far-field adaptability of traditional gaze interaction, inaccurate recognition at long distances, susceptibility to physiological micro-movements and environmental interference leading to shake and false touches, inability to lock the main interactive object in multi-target scenes, simple and weak gaze determination logic, and unstable interaction in complex scenes.

[0006] (II) Technical Solution

[0007] To achieve the aforementioned objectives of near-field and far-field adaptation, high-precision image stabilization, multi-target stable locking, accurate gaze detection, low false trigger rate, and strong environmental adaptability, this invention provides the following technical solution: a multimedia interactive control method based on multimodal vision and adaptive gaze stabilization detection, comprising image acquisition, signal processing, gaze detection, and interactive control, with the following operation steps:

[0008] Step S1: Acquire video images and perform target detection to extract visual feature information reflecting the distance between the user and the device;

[0009] Step S2: Based on the visual feature information, adaptively switch the gaze mode of far-field head pose estimation or near-field pupil detection;

[0010] Step S3: Filter the line-of-sight sequence and calculate the spatial distribution variance σ² and the variance trend value σDelta σ² using a sliding window.

[0011] Step S4: Determine the valid gaze state through multiple conditions. If the determination is successful, output multimedia interactive control commands through wired or wireless communication.

[0012] Preferably, the extraction of visual feature information reflecting the distance between the user and the device in step S1 specifically includes the calculated face area S_{box}, the distance measurement information obtained by the depth camera, or the pupil distance ratio information; in step S2, when the visual feature information indicates a long distance, the head posture is used to estimate the gaze direction, and when the visual feature information indicates a short distance, pupil detection is used to obtain the precise coordinates of the gaze point.

[0013] Preferably, the method includes a gaze stability determination step: extracting a sequence of gaze points from multiple consecutive frames using a set sliding window, calculating the spatial distribution variance σ² of the gaze points within the window and the variance change trend value Δσ² between adjacent sliding windows; when the gaze point is located in the target interaction area, and σ² is less than a set variance threshold, and Δσ² is less than a preset change threshold, it is determined to be a valid stable gaze; otherwise, it is determined to be an unstable saccade.

[0014] Preferably, the method includes a dynamic anti-shake threshold adjustment step: acquiring visual feature information reflecting the distance between the user and the device, establishing an inverse mapping relationship between the dynamic anti-shake threshold \sigma_{th}^2(t) and the visual feature information; and introducing historical data of gaze jitter as feedback input to adaptively adjust the dynamic anti-shake threshold \sigma_{th}^2(t) in real time for gaze stability determination.

[0015] Preferably, the method includes a multi-target interaction locking step: when multiple target subjects are detected in the field of view, the distance score S_{dist} and stability score S_{stab} of each target are calculated respectively; the distance score and stability score are weighted to obtain a comprehensive score S_{total}; the multiple targets are sorted according to the comprehensive score, the target ranked first is locked as the main interaction object, and the target loss and rematch logic is executed based on the significant decrease in the target score.

[0016] Preferably, the joint determination of valid gaze state in step S4, in addition to being based on the gaze point and variance index, further incorporates the user's blink confirmation, gaze duration, or specific eye movement trajectory as the final triggering condition for instruction execution, in order to distinguish between unintentional touch and genuine interaction intent.

[0017] Preferably, the multimedia interactive control method is applied to elevator multimedia advertising interaction, shopping mall intelligent wayfinding screens, contactless vending machines, and in-vehicle intelligent space scenarios.

[0018] A multimedia interactive control system based on multimodal vision and adaptive gaze stabilization determination includes a hardware device and a communication interface. The hardware device includes at least an image acquisition device, a computing processing unit, and a display terminal. The computing processing unit is used to execute the method steps of any one of claims 1 to 7. The communication interface includes a wired communication interface and / or a wireless communication interface for connecting to external devices. The external devices include elevator controllers or advertising machine control interfaces. After the computing processing unit converts the effective gaze determination result into a control command, it sends it to the external device via the communication interface to trigger the corresponding physical or multimedia action.

[0019] (III) Beneficial Effects

[0020] Compared with the prior art, the present invention provides a multimedia interactive control method and system based on multimodal vision and adaptive gaze stability determination, which has the following beneficial effects:

[0021] 1. This multimedia interactive control method and system based on multimodal vision and adaptive gaze stability determination, by adaptively switching between near and far field gaze estimation modes by the face area, takes into account both long-distance head posture estimation and near-distance pupil accurate detection, which greatly improves the interaction adaptability and detection accuracy at different usage distances. It adopts a joint judgment of sliding window variance and convergence trend, combined with dynamic adaptive anti-shake threshold, which can effectively suppress gaze jitter caused by physiological micro-movements, environmental noise and changes in lighting, significantly reduce the probability of false triggering, make gaze determination more stable and reliable, and improve the smoothness and accuracy of contactless interaction as a whole.

[0022] 2. This multimedia interactive control method and system based on multimodal vision and adaptive line-of-sight stabilization determination can intelligently select the main interactive object in a multi-person environment through a multi-target weighted locking mechanism, avoiding interference and misoperation, and improving the applicability of complex scenarios. The four-fold joint determination logic replaces the traditional single threshold judgment, making the interaction trigger more rigorous and robust. The system as a whole achieves adaptive, high-precision, low-latency, and anti-interference contactless multimedia control, which can be widely used in public displays, advertising terminals, smart spaces and vehicle equipment, and has strong practical value and promotion prospects. Attached Figure Description

[0023] Figure 1 This is a block diagram illustrating the overall system structure and data flow of the present invention;

[0024] Figure 2 This is a flowchart of the eye-tracking interaction control method of the present invention;

[0025] Figure 3 This is a diagram of the functional relationship and logical feedback system for the dynamic threshold adaptive adjustment of the present invention;

[0026] Figure 4 This is a system diagram of multi-target priority selection, main interaction object locking, and target loss re-matching in this invention;

[0027] Figure 5 This is a schematic diagram illustrating the comparison of time-domain waveforms for line-of-sight stability and the determination of false trigger points in this invention.

[0028] Figure 6 This is a flowchart illustrating the smooth switching process between near and far fields with a hysteresis interval according to the present invention. Detailed Implementation

[0029] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0030] Please see Figure 1-6 A multimedia interactive control method based on multimodal vision and adaptive gaze stability determination, the operation steps are as follows:

[0031] Step S1: Acquire video images and perform target detection to extract visual feature information reflecting the distance between the user and the device;

[0032] Step S2: Adaptively switch between far-field head pose estimation and near-field pupil detection gaze modes based on visual feature information;

[0033] Step S3: Filter the line-of-sight sequence and calculate the spatial distribution variance σ² and the variance trend value σDelta σ² using a sliding window.

[0034] Step S4: Determine the valid gaze state through multiple conditions. If the determination is successful, output multimedia interactive control commands through wired or wireless communication.

[0035] In the implementation of the case, the extraction of visual feature information reflecting the distance between the user and the device in step S1 specifically includes the calculated face area S_{box}, the distance measurement information obtained by the depth camera, or the pupil distance ratio information; in step S2, when the visual feature information indicates a long distance, the head posture is used to estimate the direction of gaze, and when the visual feature information indicates a short distance, pupil detection is used to obtain the precise coordinates of the gaze point.

[0036] Among them, the larger the face area Sbox is, the closer the user is to the device, and the smaller the area is, the farther away the user is, so as to achieve accurate differentiation between near and far fields;

[0037] This distance judgment logic can automatically match the optimal line-of-sight detection mode at different usage distances, improving the ability to adapt to all distances.

[0038] In the implementation of the case, the following steps are included to determine the stability of gaze: extract a sequence of gaze points from multiple consecutive frames using a set sliding window, calculate the spatial distribution variance \sigma^2 of the gaze points within the window and the variance change trend value \Delta\sigma^2 between adjacent sliding windows; when the gaze point is located in the target interaction area, and \sigma^2 is less than the set variance threshold, and \Delta\sigma^2 is less than the preset change threshold, it is determined to be a valid stable gaze; otherwise, it is determined to be an unstable saccade.

[0039] Among them, variance σ2 is used to measure the degree of visual focus. The smaller the variance, the more stable the visual focus. The variance trend Δσ2 is used to judge whether the visual tremor tends to converge.

[0040] By using a dual-indicator approach, physiological micro-movements and momentary tremors can be effectively filtered out, thus improving the stability of gaze detection.

[0041] In the implementation of the case, the dynamic anti-shake threshold adjustment step is included: obtaining visual feature information reflecting the distance between the user and the device, establishing an inverse mapping relationship between the dynamic anti-shake threshold \sigma_{th}^2(t) and the visual feature information; and introducing historical data of gaze jitter as feedback input to adaptively adjust the dynamic anti-shake threshold \sigma_{th}^2(t) in real time for gaze stability determination.

[0042] Among them, the threshold is appropriately relaxed as the user is farther away, and the threshold is tightened as the user is closer. At the same time, historical jitter data is combined to suppress environmental noise interference.

[0043] Through a dynamic threshold adaptive mechanism, the image stabilization effect is optimized in real time according to the scene and distance, reducing the false judgment rate.

[0044] In the implementation of the case, the multi-target interaction locking steps are as follows: when multiple target subjects are detected in the field of view, the distance score S_{dist} and stability score S_{stab} of each target are calculated respectively; the distance score and stability score are weighted to obtain the comprehensive score S_{total}; the multiple targets are sorted according to the comprehensive score, and the target ranked first is locked as the main interaction object. The target loss and rematch logic is executed based on the significant decrease in the target score.

[0045] Among them, the closer the target is and the more stable the line of sight is, the higher the overall score, and the more likely it is to be identified as the main interaction object;

[0046] Through weighted scoring and rematching mechanisms, the main user can be accurately identified in multi-user scenarios, avoiding interference and accidental operations.

[0047] In the implementation of the case, the joint determination of the valid gaze state in step S4, in addition to the gaze point and variance index, further combines the user's blink confirmation, gaze duration or specific eye movement trajectory as the final triggering condition for instruction execution, so as to distinguish between unintentional touch and real interaction intention.

[0048] Among them, blink confirmation, dwell time target or specific eye movement trajectory are all clear signals of user active interaction;

[0049] By layering multiple triggering conditions, unintentional scanning and actual operation can be completely distinguished, greatly reducing the probability of accidental triggering in public scenarios.

[0050] In the case implementation, the multimedia interactive control method was applied to elevator multimedia advertising interaction, shopping mall intelligent wayfinding screens, contactless self-service vending machines, and in-vehicle intelligent space scenarios.

[0051] All of the above scenarios are characterized by public contactless interaction, multiple people interfering, and varying distances.

[0052] This method enables safe, stable, and low-mistouch contactless eye-tracking interaction, improving the user experience and hygiene safety of the device.

[0053] A multimedia interactive control system based on multimodal vision and adaptive gaze stabilization determination includes a hardware device and a communication interface. The hardware device includes at least an image acquisition device, a computing processing unit, and a display terminal. The computing processing unit is used to execute the method steps of any one of claims 1 to 7. The communication interface includes a wired communication interface and / or a wireless communication interface for connecting to external devices. The external devices include elevator controllers or advertising machine control interfaces. After the computing processing unit converts the effective gaze determination result into a control command, it sends it to the external device via the communication interface to trigger the corresponding physical or multimedia action.

[0054] Among them, the image acquisition device is used to acquire user visual data, the computing and processing unit is responsible for algorithm operation and data processing, the display terminal is used to present interactive interface and content, the computing and processing unit has real-time image processing and AI computing capabilities, and the communication interface supports multiple protocols such as GPIO, network, and serial port, and is compatible with different terminal devices.

[0055] Through hardware collaboration, a stable operating platform is provided for the entire eye-tracking interaction method, loading and running all the algorithm logic of the control method, completing the entire process from data acquisition to command output, and standardizing the communication method to achieve seamless connection and control between eye-tracking interaction commands and external devices.

[0056] In summary, this multimedia interactive control method and system based on multimodal vision and adaptive gaze stability determination significantly improves the interaction adaptability and detection accuracy at different usage distances by adaptively switching between near and far-field gaze estimation modes based on the face area, taking into account both long-distance head posture estimation and near-distance pupil accuracy detection. By employing a joint judgment of sliding window variance and convergence trend, combined with a dynamic adaptive anti-shake threshold, it can effectively suppress gaze jitter caused by physiological micro-movements, environmental noise, and changes in illumination, significantly reducing the probability of false triggering, making gaze determination more stable and reliable, and improving the overall smoothness and accuracy of contactless interaction.

[0057] Furthermore, through a multi-target weighted locking mechanism, the main interactive object can be intelligently selected in a multi-person environment, avoiding interference and misoperation, and improving the applicability of complex scenarios. The four-fold joint judgment logic replaces the traditional single threshold judgment, making the interaction trigger more rigorous and robust. The system as a whole achieves adaptive, high-precision, low-latency, and anti-interference contactless multimedia control, which can be widely used in public displays, advertising terminals, smart spaces, and vehicle equipment, and has strong practical value and promotion prospects.

[0058] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0059] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A multimedia interactive control method based on multimodal vision and adaptive gaze stability determination, characterized in that: The operation steps are as follows: Step S1: Acquire video images and perform target detection to extract visual feature information reflecting the distance between the user and the device; Step S2: Based on the visual feature information, adaptively switch the gaze mode of far-field head pose estimation or near-field pupil detection; Step S3: Filter the line-of-sight sequence and calculate the spatial distribution variance σ² and the variance trend value σDelta σ² using a sliding window. Step S4: Determine the valid gaze state through multiple conditions. If the determination is successful, output multimedia interactive control commands through wired or wireless communication.

2. The multimedia interactive control method based on multimodal vision and adaptive gaze stability determination according to claim 1, characterized in that: The extraction of visual feature information reflecting the distance between the user and the device in step S1 specifically includes the calculated face area S_{box}, the distance measurement information obtained by the depth camera, or the pupil distance ratio information; in step S2, when the visual feature information indicates a long distance, the head posture is used to estimate the gaze direction, and when the visual feature information indicates a short distance, pupil detection is used to obtain the precise coordinates of the gaze point.

3. The multimedia interactive control method based on multimodal vision and adaptive gaze stability determination according to claim 1, characterized in that: The process includes a gaze stability determination step: extracting a sequence of gaze points from multiple consecutive frames using a set sliding window, calculating the spatial distribution variance \sigma^2 of the gaze points within the window and the variance change trend value \Delta\sigma^2 between adjacent sliding windows; when the gaze point is located in the target interaction area, and \sigma^2 is less than the set variance threshold, and \Delta\sigma^2 is less than the preset change threshold, it is determined to be a valid stable gaze; otherwise, it is determined to be an unstable saccade.

4. The multimedia interactive control method based on multimodal vision and adaptive gaze stability determination according to claim 1, characterized in that: The dynamic stabilization threshold adjustment step includes: acquiring visual feature information reflecting the distance between the user and the device, and establishing an inverse mapping relationship between the dynamic stabilization threshold σ_{th}^2(t) and the visual feature information; Historical data on eye movement is introduced as feedback input, and the dynamic anti-shake threshold σth(t) is adaptively adjusted in real time for eye movement stability determination.

5. The multimedia interactive control method based on multimodal vision and adaptive gaze stability determination according to claim 1, characterized in that: The process includes a multi-target interaction locking step: when multiple target subjects are detected in the field of view, the distance score S_{dist} and stability score S_{stab} of each target are calculated respectively; the distance score and stability score are weighted to obtain a comprehensive score S_{total}; the multiple targets are sorted according to the comprehensive score, and the target ranked first is locked as the main interaction object. The target loss and rematch logic is executed based on the significant decrease in the target score.

6. The multimedia interactive control method based on multimodal vision and adaptive gaze stability determination according to claim 1, characterized in that: The joint determination of valid gaze state in step S4, in addition to being based on the gaze point and variance index, further combines the user's blink confirmation, gaze duration or specific eye movement trajectory as the final trigger condition for instruction execution, so as to distinguish between unintentional touch and real interaction intention.

7. The multimedia interactive control method based on multimodal vision and adaptive gaze stability determination according to claim 1, characterized in that: The multimedia interactive control method is applied to elevator multimedia advertising interaction, shopping mall intelligent wayfinding screens, contactless self-service vending machines, and in-vehicle intelligent space scenarios.

8. A multimedia interactive control system based on multimodal vision and adaptive gaze stabilization determination, characterized in that: The device includes hardware and a communication interface: the hardware includes at least an image acquisition device, a computing processing unit, and a display terminal; the computing processing unit is used to execute the method steps of any one of claims 1 to 7; the communication interface includes a wired communication interface and / or a wireless communication interface for connecting to external devices, the external devices including elevator controllers or advertising machine control interfaces; the computing processing unit converts the valid gaze determination result into a control command and sends it to the external device via the communication interface to trigger the corresponding physical or multimedia action.