An anti-mis-touch interaction method based on medium-free holography
By introducing facial recognition, eye tracking, and spatiotemporal intent determination into non-media holographic interaction, a pre-operation confirmation mechanism is established, which solves the problem of accidental touches in non-media holographic interaction, improves the accuracy and security of the interaction, and ensures the identification of the unique main operator.
Patent Information
- Application Number
- CN202511516901.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-23
- Publication Date
- 2026-03-17
- Estimated Expiration
- 2045-10-23
AI Technical Summary
The lack of a pre-confirmation mechanism based on the user's true intent in medium-free holographic interaction leads to frequent accidental touches, affecting the accuracy of interaction and user experience, and also poses operational security risks.
By introducing a multi-level decision chain that includes facial recognition, eye tracking, spatiotemporal intent determination, and pattern comparison, a pre-operation confirmation mechanism is established. This mechanism includes the construction of semantic interaction entities, determination of the spatial range of faces and eyes, action recognition patterns, spatiotemporal intent determination, and multi-user identity arbitration, ensuring the accuracy and security of the interaction.
It effectively avoids accidental touches caused by natural movements and interference from multiple users, improves the spatial accuracy of interaction and the reliability of timing determination, ensures the identification of the unique main operator, and enhances the reliability and security of media-free holographic interaction.
Smart Images

Figure CN120994073B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of anti-mistouch interaction technology based on media-free holography, and more specifically, to a method for preventing accidental touch interaction based on media-free holography. Background Technology
[0002] In mediumless holographic interaction scenarios, since the image is presented in mid-air, the interaction space is open and lacks physical boundaries, existing technologies usually rely on gestures or actions to enter the camera's field of view and be determined by recognition algorithms to trigger operations.
[0003] However, in actual use, users often naturally turn their heads or make random hand movements simply because they want to browse the content. The body blocking or waving of others when they pass by may also be captured by the system as valid commands.
[0004] When multiple users appear in front of the device at the same time, the system has difficulty distinguishing the identity of the operator in terms of depth and orientation. A wandering or brief pause in the gaze may also be mistaken for an interaction intention. In addition, interference from changes in lighting or imaging noise can cause the system to unintentionally trigger operations.
[0005] These phenomena indicate that existing recognition mechanisms lack prior determination of the operator's intentions and cannot establish effective thresholds for the interaction process at the temporal and spatial levels. As a result, the system frequently experiences accidental touches, which not only affects the accuracy of the interaction and the user experience, but may also bring operational safety risks in critical application scenarios.
[0006] Therefore, the core problem with existing technologies is that mediumless holographic interaction lacks a pre-confirmation mechanism based on the user's true intention, making it difficult to reliably distinguish between arbitrary behavior and actual operation in open and multi-interference environments. Summary of the Invention
[0007] To overcome the aforementioned deficiencies of the prior art, embodiments of the present invention provide a method for preventing accidental touch interaction based on media-free holography. By introducing a pre-operation confirmation mechanism into a multi-level decision chain of face recognition, gaze tracking, spatiotemporal intent determination, and pattern comparison, the method verifies the user's true intent before the interaction is triggered, thereby solving the problems mentioned in the background art.
[0008] To achieve the above objectives, the present invention provides the following technical solution: a method for preventing accidental touch interaction based on media-free holography, comprising:
[0009] S1. A medium-free holographic image is generated by a medium-free holographic imaging device. The medium-free holographic image contains a semantic interaction body, which is composed of multiple interaction voxels or interaction planes.
[0010] S2. A face image is acquired by a camera installed on the medium-free holographic imaging device, and the face position, face orientation, face size and eye distance parameters are identified based on the face image to determine the spatial range of the face in the medium-free holographic image.
[0011] S3. When the face is within the space range, determine the first duration of the face being within the space range. If the first duration is greater than a first preset value, enter the action recognition mode to recognize the user's gesture.
[0012] And / or, when the human eye is within the said spatial range, determine a second duration for which the human eye gaze remains within the said spatial range, and if the second duration is greater than a second preset value, enter the action recognition mode to recognize the user's gesture.
[0013] In a preferred embodiment, determining the second duration for which the human eye's gaze remains within the spatial range when the human eye is within the spatial range includes:
[0014] S31. Acquire human eye images, solve for the pupil center position based on the human eye images and generate a gaze ray, and geometrically align the gaze ray with the semantic interaction body to form a gaze hit probability.
[0015] In a preferred embodiment, the method further includes:
[0016] S4. Input the facial stability, head posture change parameters, blink frequency, gaze trajectory, and hand approach trajectory within a continuous time period as the input sequence to the spatiotemporal intent determination engine, and output the operation intent probability.
[0017] S5. Based on the contrastive learning model, calculate the matching energy between the input sequence and the actual operation mode and non-operation mode respectively, and form an energy difference according to the difference between the two, and use the energy difference as a correction signal for the probability of operation intention.
[0018] In a preferred embodiment, the method further includes S6: when there are two or more candidate faces, a joint probability data association method is used to track all candidate faces, the confidence of the main operator is calculated based on the line-of-sight hit probability, the operation intention probability and the energy difference, and the unique main operator identity is output through the policy network under the constraint of a single main operator.
[0019] In a preferred embodiment, the method further includes S7: performing a weighted calculation based on the line-of-sight hit probability, the operation intention probability, the energy difference, and the main operator's confidence level to obtain a comprehensive score; when the comprehensive score exceeds a preset threshold, triggering a medium-free holographic interactive operation; and completing the final confirmation through gesture recognition or action recognition after triggering.
[0020] In a preferred embodiment, S1, the process of constructing a semantic interactive body in the medium-free holographic image includes the following steps: dividing the three-dimensional spatial region where the medium-free holographic image is located into layers, first dividing the overall space into multiple spatial slices according to the depth direction, and then performing mesh decomposition in each spatial slice to generate multiple sets of interactive voxels or interactive planes.
[0021] For the interactive voxel or interactive plane, the three-dimensional coordinate range of each unit in the device coordinate system is determined based on the projection geometry parameters of the medium-free holographic image, and the spatial position matrix of the unit is solved from the three-dimensional coordinate range; then, the direction vector and intensity distribution of the light rays passing through the unit are calculated based on the light field distribution parameters, the normal vector of the unit is solved from the direction vector, and the geometric boundary parameters of the unit are solved from the intensity distribution.
[0022] The light field distribution parameters include light direction, light intensity distribution, wavelength spectrum, and phase information; the intensity distribution includes the energy density of the light at its spatial location, the brightness gradient, and the amplitude that varies with time.
[0023] The spatial position matrix, normal vector, and geometric boundary parameters are linearly combined according to preset weights to form a spatial description set;
[0024] The spatial description set is used to establish a hierarchical index structure based on the relevance of the interaction intent, and this hierarchical index structure is defined as a semantic interaction body.
[0025] In a preferred embodiment, in S2, a face image is acquired by a camera set on a medium-free holographic imaging device, and face contour feature points are extracted from the face image. The face contour feature points are then used to calculate the face rectangle boundary.
[0026] The coordinate vector of the face position is calculated based on the facial contour feature points, the face size is calculated based on the aspect ratio of the face rectangle boundary, and the interpupillary distance parameter is calculated based on the Euclidean distance between the pupils. Furthermore, a correspondence is established between the two-dimensional coordinates of the measurement points in the face image and the corresponding three-dimensional template coordinates. A projection matrix is calculated to ensure consistency between the two-dimensional and three-dimensional points under the camera imaging geometry, and a rotation component is decomposed from the projection matrix to obtain the face orientation vector. This face orientation vector indicates the orientation of the face in the three-dimensional coordinate system.
[0027] The measurement points include the corners of the eyes, the tip of the nose, and the corners of the mouth; the three-dimensional template coordinates are preset fixed reference positions of the face measurement points in a standard three-dimensional coordinate system.
[0028] The coordinate vector of the face position, face size, eye distance parameter, and face orientation vector are input into the coordinate mapping module to complete the transformation to the spatial coordinate system of the mediumless holographic image, thereby outputting the spatial range of the face in the mediumless holographic image.
[0029] In a preferred embodiment, in S3, under the condition that the range of the face position is determined, an image of the human eye is acquired by a camera, and a set of iris edge points is extracted from the human eye image;
[0030] The iris center position is determined based on the set of iris edge points. The iris center position is used as the pupil center position coordinates, and the straight line direction between the pupil center position coordinates and the camera optical center is defined as the line of sight ray direction vector.
[0031] Perform geometric intersection calculations on the line-of-sight ray direction vector and the spatial description set of the semantic interaction body to solve the intersection point distribution, and form the line-of-sight hit probability based on the correspondence between the intersection point distribution and the interaction unit index.
[0032] In a preferred embodiment, in S4, the spatiotemporal intent determination engine's execution process on the input sequence includes: performing time synchronization and sampling alignment on face stability, head posture change parameters, blink frequency, gaze trajectory, and hand approach trajectory within a continuous time period to construct a fixed-length temporal window; performing normalization and first-order difference operations on each parameter within the temporal window to generate a stability index for face stability, an angular velocity vector for head posture change parameters, a temporal density vector for blink frequency, a velocity and dwell time vector for gaze trajectory, and a velocity and distance vector for hand approach trajectory, respectively; calculating the time delay correlation coefficient between gaze trajectory and hand approach trajectory based on the generated vectors, and establishing a correspondence with the interaction unit index given by the semantic interaction entity to form a spatiotemporal feature matrix; performing spatiotemporal fusion and thresholding statistics based on the spatiotemporal feature matrix to output the probability of operation intent.
[0033] In a preferred embodiment, the process of calculating the energy difference and correcting the probability of operational intent based on a contrastive learning model includes the following steps:
[0034] The input sequence within a continuous time period is input into the contrastive learning model. The contrastive learning model calls the real operation mode branch and the non-operation mode branch respectively, and performs an encoding operation on the input sequence in each branch to transform the input sequence into a feature representation vector with a fixed dimension.
[0035] In the real operation mode branch, the feature representation vector of the input sequence is subjected to vector dot product with the standard feature vector in the real operation mode sample library one by one, and the dot product result is divided by the product of the magnitudes of the two vectors to obtain the similarity score. Then, all similarity scores are accumulated and inverted to form the first matching energy.
[0036] In the non-operation mode branch, the feature representation vector of the input sequence is successively multiplied by the standard feature vector in the non-operation mode sample library, and the dot product result is divided by the product of the magnitudes of the two vectors to obtain the similarity score. Then, all similarity scores are accumulated and inverted to form the second matching energy.
[0037] The first matching energy is subtracted from the second matching energy to obtain the energy difference, which is used to characterize the relative matching strength of the input sequence between the real operating mode and the non-operating mode.
[0038] The energy difference is combined with the operation intent probability output by the spatiotemporal intent determination engine. When the energy difference is greater than zero, the value of the operation intent probability is increased, and when the energy difference is less than zero, the value of the operation intent probability is decreased. The corrected operation intent probability is then output as a correction signal for subsequent comprehensive score calculation.
[0039] The technical effects and advantages of this invention are as follows:
[0040] By introducing a pre-operation judgment mechanism based on facial recognition and eye tracking, the system confirms whether the user has the intention to operate before interaction, thereby effectively avoiding natural movements, passersby passing by, or brief eye lingering being misjudged as valid operations, thus solving the most prominent problem of accidental touch in mediumless holographic interaction.
[0041] By constructing semantic interaction bodies, the mediumless holographic image is divided into interaction voxels or interaction planes, and geometric intersection calculations are performed with the line of sight rays. This allows interaction determination to be based on a precise spatial coordinate system, avoiding ambiguity in line of sight determination caused by boundaryless projection and improving the spatial accuracy of interaction triggering.
[0042] By using a spatiotemporal intent determination engine to fuse and analyze facial stability, head posture, blink frequency, gaze trajectory and hand approach trajectory, and calculate the time delay correlation coefficient, it is possible to distinguish between viewing behavior and actual operation intent, reduce misjudgment caused by single-dimensional parameters, and improve the reliability of timing determination.
[0043] In scenarios where multiple people are present simultaneously, the consistency of candidate face trajectories is maintained through a joint probability data association method. The confidence of the main operator is calculated by combining the probability of eye contact, the probability of operation intention, and the energy difference. Then, identity arbitration is performed through a policy network to ensure that only a unique main operator is output, thus avoiding accidental touches under the interference of multiple users. Attached Figure Description
[0044] Figure 1 This is a schematic diagram illustrating the principle of medium-free holographic imaging in this invention.
[0045] Figure 2 This is a flowchart of the method steps of the present invention. Detailed Implementation
[0046] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0047] Example 1:
[0048] Refer to the instruction manual appendix Figure 1-2 An embodiment of the present invention provides a method for preventing accidental touches based on media-free holography, comprising:
[0049] S1. A medium-free holographic image is generated by a medium-free holographic imaging device. The medium-free holographic image contains a semantic interaction body, which is composed of multiple interaction voxels or interaction planes.
[0050] It should be noted that for S1, the semantic interactive body refers to an interactive reference structure artificially constructed in a mediumless holographic image, which divides the projected image into interactive areas that can be recognized and utilized by the system.
[0051] Interactive voxels refer to dividing three-dimensional space into small cubic units, each unit representing a three-dimensional location point in the image, used to accurately represent the spatial extent of the interactive area; while interactive planes refer to rectangular or polygonal regions divided on a two-dimensional projection plane, used to mark the two-dimensional range where the user may interact; simply put, semantic interactive volumes are spatial index structures composed of a large number of interactive voxels or interactive planes, which transform the originally blurry holographic image interactive area into a set of units that can be geometrically calculated and matched, thereby providing a clear spatial reference for subsequent gaze ray alignment and hit determination;
[0052] S2. A face image is acquired by a camera installed on the medium-free holographic imaging device, and the face position, face orientation, face size and eye distance parameters are identified based on the face image to determine the spatial range of the face in the medium-free holographic image.
[0053] S3. When the face is within the space range, determine the first duration of the face being within the space range. If the first duration is greater than a first preset value, enter the action recognition mode to recognize the user's gesture.
[0054] Furthermore, when the human eye is within the said spatial range, a second duration for the human eye's gaze to remain within the said spatial range is determined. If the second duration exceeds a second preset value, an action recognition mode is entered to recognize the user's gestures.
[0055] The step of determining the second duration of the human eye's gaze within the spatial range when the human eye is within the spatial range includes: acquiring a human eye image, solving for the pupil center position based on the human eye image and generating a gaze ray, and geometrically aligning the gaze ray with the semantic interaction body to form a gaze hit probability.
[0056] In addition to Example 1, Example 2 is also included:
[0057] S4. Input the facial stability, head posture change parameters, blink frequency, gaze trajectory, and hand approach trajectory within a continuous time period as the input sequence to the spatiotemporal intent determination engine, and output the operation intent probability.
[0058] S5. Based on the contrastive learning model, calculate the matching energy between the input sequence and the actual operation mode and non-operation mode respectively, and form an energy difference according to the difference between the two, and use the energy difference as a correction signal for the probability of operation intention.
[0059] Based on Example 2, Example 3 is also included:
[0060] The difference from Example 1 is that it also includes S6: when there are two or more candidate faces, the joint probability data association method is used to track all candidate faces, the confidence of the main operator is calculated based on the line-of-sight hit probability, the operation intention probability and the energy difference, and the unique main operator identity is output through the policy network under the constraint of a single main operator.
[0061] For S6, it should be noted that joint probability data refers to the probability calculation of one-to-one correspondence between multiple face observations captured by the camera and multiple candidate face trajectories established in the system within the same time window, and the joint probability distribution formed by combining all possible observation-trajectory correspondences. This joint probability distribution not only considers the matching probability of a single observation and a single trajectory, but also comprehensively calculates the joint consistency of multiple faces under the overall allocation, in order to reduce erroneous associations in the case of multiple targets.
[0062] When tracking all candidate faces, the position, orientation, and size parameters of the candidate faces are first extracted in each frame of the image and used as observation data to input the joint probabilistic data association method. Then, the observation-trajectory matching probability is established by using the Euclidean distance and pose similarity between the predicted face trajectory state and the current observation data. Finally, based on the joint probability distribution, the allocation scheme with the highest probability is selected from multiple possible allocation results, so as to maintain the identity consistency of each candidate face trajectory between consecutive frames and achieve stable tracking.
[0063] Under the condition that the candidate face is tracked, the system calculates the gaze hit probability, operation intention probability, and energy difference for each candidate face. The three are then input as input vectors to the confidence calculation module. The confidence calculation module obtains the main operator confidence of each candidate face through weighted accumulation and normalization. Subsequently, the confidence vectors of all candidate faces are input to the policy network. Under the constraint of a single main operator, the policy network performs maximum value selection. When the confidence of a candidate face is the highest and exceeds the threshold, the candidate face is output as the unique main operator identity, thereby realizing identity arbitration in the scenario where multiple people exist at the same time.
[0064] The policy network is a pre-trained decision model stored in the system. It is used to select and output a unique main operator identity based on a preset single main operator constraint after inputting the main operator confidence vector.
[0065] It also includes S7: a weighted calculation is performed based on the line-of-sight hit probability, the operation intention probability, the energy difference and the main operator confidence to obtain a comprehensive score. When the comprehensive score exceeds a preset threshold, a medium-free holographic interactive operation is triggered, and the final confirmation is completed through gesture recognition or action recognition after the trigger.
[0066] In S1, the process of constructing a semantic interactive body in the medium-free holographic image includes the following steps: dividing the three-dimensional spatial region where the medium-free holographic image is located into layers, first dividing the overall space into multiple spatial slices according to the depth direction, and then performing mesh decomposition in each spatial slice to generate multiple sets of interactive voxels or interactive planes.
[0067] For the interactive voxel or interactive plane, the three-dimensional coordinate range of each unit in the device coordinate system is determined based on the projection geometry parameters of the medium-free holographic image, and the spatial position matrix of the unit is solved from the three-dimensional coordinate range; then, the direction vector and intensity distribution of the light rays passing through the unit are calculated based on the light field distribution parameters, the normal vector of the unit is solved from the direction vector, and the geometric boundary parameters of the unit are solved from the intensity distribution.
[0068] The light field distribution parameters include light direction, light intensity distribution, wavelength spectrum and phase information, which are used to describe the propagation characteristics of the medium-free holographic image in space; the intensity distribution includes the energy density, brightness gradient and amplitude of the light at the spatial location, which are used to characterize the illumination intensity and variation of the medium-free holographic image in different regions.
[0069] The spatial position matrix, normal vector, and geometric boundary parameters are linearly combined according to preset weights to form a spatial description set for subsequent alignment of line-of-sight rays and semantic interaction objects;
[0070] The spatial description set is used to establish a hierarchical index structure based on the relevance of the interaction intent, and this hierarchical index structure is defined as a semantic interaction body, so as to provide a multi-level alignment benchmark in the matching calculation process between the line of sight ray and the semantic interaction body.
[0071] In S2, a face image is captured by a camera set on a medium-free holographic imaging device, and face contour feature points are extracted from the face image. The face rectangle boundary is calculated using the face contour feature points. The face rectangle boundary refers to the smallest bounding rectangle generated based on the face contour feature points in the detected face area, which is used to limit the two-dimensional range of the face in the image.
[0072] The coordinate vector of the face position is calculated based on the facial contour feature points, the face size is calculated based on the aspect ratio of the face rectangle boundary, and the interpupillary distance parameter is calculated based on the Euclidean distance between the pupils. Furthermore, a correspondence is established between the two-dimensional coordinates of the measurement points in the face image and the corresponding three-dimensional template coordinates. A projection matrix is calculated to ensure consistency between the two-dimensional and three-dimensional points under the camera imaging geometry, and a rotation component is decomposed from the projection matrix to obtain the face orientation vector. This face orientation vector indicates the orientation of the face in the three-dimensional coordinate system.
[0073] The measurement points include the corners of the eyes, the tip of the nose, and the corners of the mouth; the three-dimensional template coordinates are fixed reference positions of the face measurement points in the standard three-dimensional coordinate system, which are used to correspond to the two-dimensional feature points in the image, thereby solving the orientation of the face.
[0074] The coordinate vector of the face position, face size, eye distance parameter, and face orientation vector are input to the coordinate mapping module to complete the transformation to the spatial coordinate system of the medium-free holographic image, thereby outputting the spatial range of the face in the medium-free holographic image. The coordinate mapping module is used to transform the face position coordinate vector, face size, eye distance parameter, and face orientation vector acquired by the camera into the spatial coordinate system of the medium-free holographic image through rotation, translation, and scaling calculations, thereby determining the actual spatial range of the face in the holographic image.
[0075] In S3, under the condition that the range of the face position is determined, the camera captures the human eye image, and extracts the iris edge point set in the human eye image; the iris edge point set refers to the set of continuous pixel coordinates extracted along the junction of the iris and sclera in the captured human eye image, which is used to describe the outer contour of the iris. In practical applications, the iris edge point set can be obtained by performing edge detection and feature extraction algorithms on the human eye image, including calculating the gray-level gradient of the junction area of the iris and sclera and outputting the boundary pixel coordinates;
[0076] The iris center position is determined based on the set of iris edge points. The iris center position is used as the pupil center position coordinates, and the straight line direction between the pupil center position coordinates and the camera optical center is defined as the line of sight ray direction vector.
[0077] Perform geometric intersection calculations on the line-of-sight ray direction vector and the spatial description set of the semantic interaction body to solve the intersection point distribution, and form the line-of-sight hit probability based on the correspondence between the intersection point distribution and the interaction unit index;
[0078] It should be noted that the geometric intersection calculation process first expresses the direction vector of the gaze ray determined by the pupil center position coordinates and the camera optical center as a parametric equation, and then expresses the spatial description set of each interactive unit in the semantic interactive body as a geometric boundary equation. Subsequently, the parametric equation of the gaze ray is substituted into the geometric boundary equation to solve for the coordinates of the intersection point between the ray and the boundary of each interactive unit. After the intersection point coordinates are solved, the distribution of the intersection point coordinates within each interactive unit is further statistically analyzed, and a correspondence between the intersection points and the interactive units is established based on the distribution. Finally, this correspondence is used as the criterion to form the gaze hit probability, which is used to characterize the probability that the user's gaze falls into the interactive unit in the semantic interactive body.
[0079] In S4, the spatiotemporal intent determination engine's execution process for the input sequence includes: performing time synchronization and sampling alignment on face stability, head posture change parameters, blink frequency, gaze trajectory, and hand approach trajectory within a continuous time period to construct a fixed-length temporal window; performing normalization and first-order difference operations on each parameter within the temporal window to generate a stability index for face stability, an angular velocity vector for head posture change parameters, a temporal density vector for blink frequency, a velocity and dwell time vector for gaze trajectory, and a velocity and distance vector for hand approach trajectory; calculating the time delay correlation coefficient between gaze trajectory and hand approach trajectory based on the generated vectors, and establishing a correspondence with the interaction unit index given by the semantic interaction entity to form a spatiotemporal feature matrix; performing spatiotemporal fusion and thresholding statistics based on the spatiotemporal feature matrix to output the probability of operation intent;
[0080] It should be noted that the time delay correlation coefficient refers to the comparison of the time series of the gaze trajectory and the time series of the hand approach trajectory with different time delays within a fixed time window. By calculating the correlation between the two time series at each delay, the maximum correlation value between the two at a certain delay is obtained, and this maximum correlation value is used as the time delay correlation coefficient to characterize the coupling strength between the user's gaze change and the hand approach action in terms of time sequence.
[0081] In addition, in S4, the normalization operation refers to mapping the original values of face stability, head posture change parameters, blink frequency, gaze trajectory, and hand approach trajectory to a unified numerical range through linear transformation in order to eliminate dimensional differences.
[0082] The first-order difference operation in S4 refers to calculating the difference between adjacent values of a normalized sequence over a continuous time period to solve for the rate or increment characteristics of parameter change over time.
[0083] The process of calculating the energy difference and correcting the probability of operational intent based on a contrastive learning model includes the following steps:
[0084] The input sequence within a continuous time period is fed into a contrastive learning model. This model calls both a real operation mode branch and a non-operation mode branch, and in each branch, an encoding operation is performed on the input sequence, transforming it into a feature representation vector with fixed dimensions. The real operation mode branch is a computational path in the contrastive learning model that receives the input sequence and compares it with feature vectors in a real operation mode sample library to calculate the degree of matching between the input sequence and the real operation behavior. The non-operation mode branch is another computational path in the contrastive learning model that receives the same input sequence and compares it with feature vectors in a non-operation mode sample library to calculate the degree of matching between the input sequence and non-operation behavior.
[0085] In the real operation mode branch, the feature representation vector of the input sequence is subjected to vector dot product with the standard feature vector in the real operation mode sample library one by one, and the dot product result is divided by the product of the magnitudes of the two vectors to obtain the similarity score. Then, all similarity scores are accumulated and inverted to form the first matching energy.
[0086] In the non-operation mode branch, the feature representation vector of the input sequence is successively multiplied by the standard feature vector in the non-operation mode sample library, and the dot product result is divided by the product of the magnitudes of the two vectors to obtain the similarity score. Then, all similarity scores are accumulated and inverted to form the second matching energy.
[0087] The first matching energy is subtracted from the second matching energy to obtain the energy difference, which is used to characterize the relative matching strength of the input sequence between the real operating mode and the non-operating mode.
[0088] The energy difference is combined with the operation intent probability output by the spatiotemporal intent determination engine. When the energy difference is greater than zero, the value of the operation intent probability is increased, and when the energy difference is less than zero, the value of the operation intent probability is decreased. The corrected operation intent probability is then output as a correction signal for subsequent comprehensive score calculation.
[0089] It is important to explain that the core problem with existing medium-free holographic interaction is that, because the projected image is suspended in the air and the interactive area is open and lacks physical boundaries, unintentional user actions such as random head movements, passersby walking by, or gesture interference may be incorrectly identified as valid commands, leading to accidental touches. Recognizing this problem, the solution adopts "pre-operation confirmation" as the core strategy to prevent accidental touches. It gradually establishes a complete protection logic by introducing mechanisms such as facial recognition, eye tracking, spatiotemporal behavior determination, comparative learning correction, and multi-user identity arbitration. The establishment of this logic is not a simple linear aggregation of steps, but a collaborative and structured process.
[0090] In terms of implementation steps, the solution first generates a medium-free holographic image using a medium-free holographic imaging device, and constructs a semantic interactive entity within the image, dividing the interactive area into interactive voxels or interactive planes to provide clear spatial references for subsequent gaze and actions. Next, it acquires facial images using a camera, and on this basis, solves the parameters of facial position, facial orientation, facial size, and eye distance. Then, it transforms these parameters into the spatial coordinate system of the medium-free holographic image through a coordinate mapping module, thereby accurately defining the user's interactive space range. After the facial range is determined, it proceeds to the eye image processing stage, extracting the iris edge point set and solving the pupil center position. Combined with the camera's optical center, it generates a gaze ray direction vector, and performs geometric intersection calculations with the semantic interactive entity to form the gaze hit probability. This processing chain ensures that the user's gaze direction must be aligned with the projected interactive area before proceeding to the next step of judgment.
[0091] Subsequently, a spatiotemporal intent determination engine is introduced to synchronize and process facial stability, head posture change parameters, blink frequency, gaze trajectory, and hand approach trajectory over continuous time periods. Through normalization, first-order difference calculation, and time-delay correlation coefficient calculation, these dynamic signals are fused into a spatiotemporal feature matrix. This matrix is used to interpret whether the user has an actual operational intent, such as whether gaze is accompanied by hand movement. The output is the operational intent probability, becoming an important basis for interaction credibility. Based on this, the scheme further introduces a contrastive learning model, performing a two-branch comparison between the input sequence and the actual operation mode and non-operation mode, forming a first matching energy and a second matching energy respectively, and calculating the difference between the two to obtain the energy difference. The energy difference is used as a correction signal on the operational intent probability, enabling the system to clearly distinguish between actual and non-intent behaviors through positive and negative differences, thereby enhancing the robustness of the determination.
[0092] In addition, in multi-person environments, multiple candidate faces are tracked using a joint probability data association method. The probability of eye contact, the probability of operation intent, and the energy difference are input into the confidence calculation module to solve the confidence of the main operator. Then, a policy network is used to output a unique identity under the constraint of a single main operator, which solves the identity confusion problem when multiple people are present at the same time. Finally, the probability of eye contact, the probability of operation intent, the energy difference, and the confidence of the main operator are weighted and calculated to obtain a comprehensive score. When the comprehensive score exceeds a preset threshold, a mediumless holographic interaction operation is triggered. After triggering, a gesture or action is required for final confirmation to achieve a double-layer insurance mechanism.
[0093] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for preventing mistaken touch interaction based on medium-free holography, characterized in that, Comprise: S1, generate a medium-free holographic image by a medium-free holographic imaging device, the medium-free holographic image contains a semantic interactive body, the semantic interactive body is composed of a plurality of interactive voxels or interactive planes; S2, capture a face image by a camera arranged on the medium-free holographic imaging device, and identify a face position, a face orientation, a face size and an eye distance parameter based on the face image, for determining a spatial range of the face in the medium-free holographic image; S3, when the face is in the spatial range, determine a first time length that the face is in the spatial range, and when the first time length is greater than a first preset value, enter an action recognition mode to identify a user gesture; And / or, when a human eye is in the spatial range, determine a second time length that the line of sight of the human eye stays in the spatial range, and when the second time length is greater than a second preset value, enter the action recognition mode to identify the user gesture; S4, input a face stability, a head posture change parameter, a blinking frequency, a line of sight trajectory and a hand approaching trajectory in a continuous period as an input sequence to a space-time intention judgment engine, and output an operation intention probability; S5, calculate a matching energy of the input sequence with a real operation mode and a non-operation mode respectively based on a contrast learning model, form an energy difference according to a difference between the two, and take the energy difference as a correction signal for the operation intention probability. 2.The false touch prevention method based on medium-free hologram according to claim 1, wherein: The determination of the second time length that the line of sight of the human eye stays in the spatial range when the human eye is in the spatial range comprises: Capture an eye image, solve a pupil center position and generate a line of sight ray based on the eye image, geometrically align the line of sight ray with the semantic interactive body, and form a line of sight hit probability.
3. The anti-mis-touch interaction method based on medium-free holography according to claim 2, characterized in that: Further comprising S6: in the case of more than or equal to two candidate faces, track all candidate faces using a joint probability data association method, calculate a main operator confidence based on the line of sight hit probability, the operation intention probability and the energy difference, and output a unique main operator identity under the constraint of a single main operator through a strategy network.
4. The anti-mis-touch interaction method based on medium-free holography according to claim 3, characterized in that: Further comprising S7: perform weighted calculation based on the line of sight hit probability, the operation intention probability, the energy difference and the main operator confidence to obtain a comprehensive score, and trigger a medium-free holographic interaction operation when the comprehensive score exceeds a preset threshold, and complete final confirmation through gesture recognition or action recognition after triggering.
5. The anti-mis-touch interaction method based on medium-free holography according to claim 4, characterized in that: In S1, the process of constructing a semantic interactive body in the medium-free holographic image comprises the following steps: layering and dividing a three-dimensional space region where the medium-free holographic image is located, cutting the whole space into a plurality of space slices according to the depth direction, and then performing grid decomposition in each space slice to generate a plurality of interactive voxels or interactive planes. For the interactive voxel or interactive plane, a three-dimensional coordinate range of each unit in a device coordinate system is determined based on projection geometry parameters of the medium-free holographic image, and a spatial position matrix of the unit is solved from the three-dimensional coordinate range; subsequently, a light ray direction vector and an intensity distribution of a light ray passing through the unit are calculated based on light field distribution parameters, a normal vector of the unit is solved from the direction vector, and geometric boundary parameters of the unit are solved from the intensity distribution; The light field distribution parameters include light ray direction, light intensity distribution, wavelength spectrum, and phase information; and the intensity distribution includes energy density of the light ray on the spatial position, brightness gradient, and amplitude varying with time; The spatial position matrix, the normal vector, and the geometric boundary parameters are linearly combined according to preset weights to form a spatial description set; The spatial description set is established in a hierarchical index structure according to the relevance of the interactive intention, and the hierarchical index structure is defined as a semantic interactive body.
6. The anti-mis-touch interactive method based on medium-free holography according to claim 5, characterized in that: In S2, a camera disposed on the medium-free holographic imaging device is used to capture a face image, and face contour feature points are extracted from the face image; a face rectangular boundary is solved by using the face contour feature points; A coordinate vector of a face position is calculated based on the face contour feature points, a face size is calculated based on an aspect ratio of the face rectangular boundary, an eye distance parameter is calculated based on an Euclidean distance between pupils of two eyes, and on this basis, a two-dimensional coordinate of a measurement point in the face image is corresponded to a corresponding three-dimensional template coordinate; a projection matrix is solved by making the two-dimensional point and the three-dimensional point consistent in camera imaging geometry, and a rotation component is decomposed from the projection matrix, and finally a face orientation vector representing a face orientation is obtained; The face orientation vector is used to indicate a direction of the face in a three-dimensional space coordinate system; The measurement points include a corner point of the eyes, a nose tip point, and a mouth corner point; and the three-dimensional template coordinates are fixed reference positions of the face measurement points in a standard three-dimensional coordinate system; The coordinate vector of the face position, the face size, the eye distance parameter, and the face orientation vector are input into a coordinate mapping module to complete conversion to a medium-free holographic image space coordinate system, so as to output a spatial range of the face in the medium-free holographic image.
7. The anti-mis-touch interactive method based on medium-free holography according to claim 6, characterized in that: In S3, under the condition that the face position range is determined, an eye image is captured by the camera, and an iris edge point set is extracted from the eye image; A center position of the iris is determined based on the iris edge point set, the center position of the iris is taken as a pupil center position coordinate, and a straight line direction between the pupil center position coordinate and an optical center of the camera is defined as a line of sight ray direction vector; The line of sight ray direction vector and the spatial description set of the semantic interactive body are subjected to geometric intersection calculation to solve an intersection point distribution, and a line of sight hit probability is formed based on a corresponding relationship between the intersection point distribution and an interactive unit index.
8. The method of claim 7, wherein the method is based on a medium-free holographic anti-mistouch interaction. In S4, the execution process of the space-time intention judgment engine on the input sequence includes: time synchronization and sampling alignment of the human face stability, head posture change parameter, blink frequency, gaze trajectory and hand proximity trajectory in the continuous time period, and construction of a fixed-length time sequence window; performing normalization and first-order difference operation on each parameter in the time sequence window to generate a stability index of the human face stability, an angular velocity vector of the head posture change parameter, a time density vector of the blink frequency, a speed and dwell time vector of the gaze trajectory, and a speed and distance vector of the hand proximity trajectory; calculating the time delay correlation coefficient of the gaze trajectory and the hand proximity trajectory based on the generated vectors, and establishing a corresponding relationship with the interaction unit index given by the semantic interaction body to form a space-time feature matrix, and performing space-time fusion and thresholding statistics according to the space-time feature matrix to output the operation intention probability.
9. The method of claim 8, wherein the method is based on a medium-free holographic anti-mistouch interaction. The process of calculating the energy difference based on the contrast learning model and correcting the operation intention probability includes the following steps: input the input sequence in the continuous time period into the contrast learning model, call the real operation mode branch and the non-operation mode branch in the contrast learning model respectively, and perform encoding operation on the input sequence in each branch to convert the input sequence into a feature representation vector with fixed dimension; in the real operation mode branch, the feature representation vector of the input sequence is multiplied with the standard feature vector in the real operation mode sample library one by one, and the dot product result is divided by the product of the two vector lengths to obtain the similarity score, and then all the similarity scores are accumulated and inverted to form the first matching energy; in the non-operation mode branch, the feature representation vector of the input sequence is multiplied with the standard feature vector in the non-operation mode sample library one by one, and the dot product result is divided by the product of the two vector lengths to obtain the similarity score, and then all the similarity scores are accumulated and inverted to form the second matching energy; performing subtraction operation on the first matching energy and the second matching energy to obtain the energy difference, which is used to represent the relative matching strength of the input sequence between the real operation mode and the non-operation mode; combine the energy difference with the operation intention probability output by the space-time intention judgment engine, when the energy difference is greater than zero, increase the value of the operation intention probability, when the energy difference is less than zero, decrease the value of the operation intention probability, output the corrected operation intention probability, and use the corrected operation intention probability as a correction signal for subsequent comprehensive score calculation.
Citation Information
Patent Citations
Gesture Interaction method and system based on virtual human
CN108459712A
Methods for two-stage hand gesture input
US20200301513A1