Mistaken touch prevention interaction method based on medium-free holography

By introducing face recognition, eye tracking, and spatiotemporal intent determination into non-media holographic interaction, a multi-level determination chain is constructed, which solves the problem of accidental touch in non-media holographic interaction and achieves higher interaction accuracy and security.

CN120994073AActive Publication Date: 2025-11-21江西像航科技有限公司 +2
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202511516901.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-23
Publication Date
2025-11-21
Estimated Expiration
2045-10-23

AI Technical Summary

Technical Problem

Existing medium-free holographic interaction technology lacks a pre-confirmation mechanism for the user's true intentions, leading to frequent accidental touches in open environments, affecting the accuracy of interaction and user experience, and even posing operational safety risks.

Method used

By introducing facial recognition, eye tracking, and spatiotemporal intent determination, a multi-level determination chain is constructed, including the generation of semantic interaction entities, determination of the spatial range of faces and eyes, action recognition, spatiotemporal intent determination, and multi-user identity arbitration, to ensure the accuracy and security of interactive operations.

Benefits of technology

It effectively avoids accidental touches caused by natural actions and interference from multiple users, improves the spatial accuracy of interaction and the reliability of timing determination, ensures the identification of the unique main operator, and enhances the robustness and security of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120994073A_ABST
    Figure CN120994073A_ABST
Patent Text Reader

Abstract

The invention discloses a false touch prevention interaction method based on medium-free holography, and particularly relates to the field of false touch prevention interaction of medium-free holography, which comprises the following steps: generating a medium-free holographic image through a medium-free holographic imaging device, the semantic interaction body is composed of a plurality of interaction voxels or interaction planes; acquiring a human face image through a camera arranged on the medium-free holographic imaging equipment, and identifying a human face position, a human face orientation, a human face size and an eye distance parameter based on the human face image so as to determine a spatial range of a human face in the medium-free holographic image; when the face is in the spatial range, determining a first duration when the face is in the spatial range; an operation pre-confirmation mechanism is introduced into a multi-level judgment chain of face recognition, sight tracking, space-time intention judgment and mode comparison, so that the real intention of a user is verified before interaction triggering.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of medium-free holographic anti-mistouch interaction, and more particularly to a medium-free holographic anti-mistouch interaction method. BACKGROUND

[0002] In a medium-free holographic interaction scenario, due to the image suspended presentation, open interaction space and lack of physical boundaries, the existing technology usually relies on gestures or actions to enter the camera field of view and is triggered by a recognition algorithm; However, in actual use, users often produce natural head turning or random hand movements just to browse the content, and body occlusion or waving of passers-by may also be captured as valid instructions by the system; When multiple users appear in front of the device at the same time, the system is difficult to distinguish the identity of the operator in depth and orientation, and the wandering or temporary stay of the line of sight may be mistaken as an interaction intention, plus the interference caused by light changes or imaging noise, which will cause the system to trigger operations unintentionally; These phenomena show that the existing recognition mechanism lacks pre-determination of the operator's intention, and cannot establish an effective threshold for the interaction process in the time and space dimensions; as a result, the system frequently triggers false touches, which not only affects the accuracy of the interaction and the user experience, but also may cause operation safety hazards in critical application scenarios; Therefore, the core problem existing in the prior art is that the medium-free holographic interaction lacks a pre-confirmation mechanism based on the real intention of the user, and thus it is difficult to reliably distinguish between random behavior and actual operation in an open and multi-interference environment. SUMMARY

[0003] In order to overcome the above-mentioned defects of the prior art, the embodiments of the present application provide a medium-free holographic anti-mistouch interaction method, which introduces an operation pre-confirmation mechanism in a multi-level determination chain of face recognition, gaze tracking, time and space intention determination and pattern comparison, to verify the real intention of the user before interaction triggering, so as to solve the problems raised in the background art.

[0004] To achieve the above-mentioned purpose, the present application provides the following technical scheme: a medium-free holographic anti-mistouch interaction method, comprising: S1, generating a medium-free holographic image by a medium-free holographic imaging device, the medium-free holographic image containing a semantic interaction body, the semantic interaction body being composed of a plurality of interaction voxels or interaction planes; S2, capturing a face image by a camera arranged on the medium-free holographic imaging device, and identifying a face position, a face orientation, a face size and an eye distance parameter based on the face image, for determining a spatial range of the face in the medium-free holographic image; S3, when the human face is in the space range, determining a first time length that the human face is in the space range, and when the first time length is greater than a first preset value, entering an action recognition mode to recognize a user gesture; and / or, when the human eye is in the space range, determining a second time length that the human eye line of sight stays in the space range, and when the second time length is greater than a second preset value, entering the action recognition mode to recognize the user gesture.

[0005] In a preferred embodiment, the determining, when the human eye is in the space range, a second time length that the human eye line of sight stays in the space range includes: S31, collecting a human eye image, solving a pupil center position based on the human eye image and generating a line of sight ray, geometrically aligning the line of sight ray with the semantic interactive body to form a line of sight hit probability.

[0006] In a preferred embodiment, the method further includes: S4, inputting, as input sequences, a human face stability, a head posture change parameter, a blink frequency, a line of sight trajectory and a hand approaching trajectory in a continuous period to a space-time intention determination engine, and outputting an operation intention probability; S5, respectively calculating, based on a contrast learning model, matching energy of the input sequences and a real operation mode and a non-operation mode, and forming an energy difference according to a difference between the two, and taking the energy difference as a correction signal for the operation intention probability.

[0007] In a preferred embodiment, it further includes S6: when there are greater than or equal to two candidate human faces, using a joint probability data association method to track all candidate human faces, solving a main operator confidence based on the line of sight hit probability, the operation intention probability and the energy difference, and outputting a unique main operator identity under the constraint of a single main operator through a strategy network.

[0008] In a preferred embodiment, it further includes S7: performing weighted calculation based on the line of sight hit probability, the operation intention probability, the energy difference and the main operator confidence to obtain a comprehensive score, and when the comprehensive score exceeds a preset threshold, triggering a medium-free holographic interaction operation, and after triggering, completing final confirmation through gesture recognition or action recognition.

[0009] In a preferred embodiment, in S1, the process of constructing the semantic interactive body in the medium-free holographic image includes the following steps: layering and dividing a three-dimensional space region where the medium-free holographic image is located, first cutting the overall space to form a plurality of space slices according to the depth direction, and then performing grid decomposition in each space slice to generate a plurality of groups of interactive voxels or interactive planes; For the interactive voxel or interactive plane, a three-dimensional coordinate range of each unit in a device coordinate system is determined based on projection geometry parameters of the medium-free holographic image, and a spatial position matrix of the unit is solved from the three-dimensional coordinate range; subsequently, a light ray direction vector and an intensity distribution of the light ray passing through the unit are calculated based on light field distribution parameters, a normal vector of the unit is solved from the direction vector, and geometric boundary parameters of the unit are solved from the intensity distribution; The light field distribution parameters include light ray direction, light intensity distribution, wavelength spectrum, and phase information; and the intensity distribution includes energy density of the light ray on the spatial position, brightness gradient, and amplitude varying with time. The spatial position matrix, the normal vector, and the geometric boundary parameters are linearly combined according to preset weights to form a spatial description set. The spatial description set is established in a hierarchical index structure according to the relevance of the interactive intention, and the hierarchical index structure is defined as a semantic interactive body.

[0010] In a preferred embodiment, in S2, a face image is captured by a camera arranged on the medium-free holographic imaging device, and face contour feature points are extracted in the face image, and a face rectangular boundary is solved using the face contour feature points; A coordinate vector of the face position is calculated based on the face contour feature points, a face size is calculated based on an aspect ratio of the face rectangular boundary, an eye distance parameter is calculated based on an Euclidean distance between pupils of the two eyes, and on this basis, a corresponding relationship between two-dimensional coordinates of a measurement point in the face image and corresponding three-dimensional template coordinates is established, a projection matrix that makes the two-dimensional point and the three-dimensional point consistent under camera imaging geometry is solved, and a rotation component is decomposed from the projection matrix, and finally a face orientation vector representing a face orientation is obtained; the face orientation vector is used to indicate a direction of the face in a three-dimensional space coordinate system; The measurement points include a corner point of the two eyes, a nose tip point, and a mouth corner point; and the three-dimensional template coordinates are fixed reference positions of the face measurement points in a standard three-dimensional coordinate system; The coordinate vector of the face position, the face size, the eye distance parameter, and the face orientation vector are input into a coordinate mapping module to complete conversion to a medium-free holographic image space coordinate system, so as to output a spatial range of the face in the medium-free holographic image.

[0011] In a preferred embodiment, in S3, under the condition that the face position range is determined, a human eye image is captured by the camera, and an iris edge point set is extracted in the human eye image; A center position of the iris is determined based on the iris edge point set, the center position of the iris is taken as a pupil center position coordinate, and a straight line direction between the pupil center position coordinate and an optical center of the camera is defined as a line of sight ray direction vector; The line-of-sight ray direction vector is subjected to a geometric intersection calculation with a spatial description set of the semantic interactive body, a distribution of intersection points is solved, and a line-of-sight hit probability is formed based on a correspondence between the distribution of intersection points and an interactive unit index.

[0012] In a preferred embodiment, in S4, the execution process of the spatio-temporal intention judgment engine on the input sequence includes: performing time synchronization and sampling alignment on the face stability, head posture change parameter, blink frequency, line-of-sight trajectory and hand proximity trajectory in the continuous time period, and constructing a fixed-length time window; performing normalization and first-order difference operation on each parameter in the time window to generate a stability index of face stability, an angular velocity vector of head posture change parameter, a time density vector of blink frequency, a speed and dwell time vector of line-of-sight trajectory, and a speed and distance vector of hand proximity trajectory; calculating the time delay correlation coefficient of the line-of-sight trajectory and the hand proximity trajectory based on the generated vectors, and establishing a correspondence with the interactive unit index given by the semantic interactive body to form a spatio-temporal feature matrix, and performing spatio-temporal fusion and thresholding statistics according to the spatio-temporal feature matrix to output the operation intention probability.

[0013] In a preferred embodiment, the process of calculating the energy difference and correcting the operation intention probability based on the contrast learning model includes the following steps: The input sequence in the continuous time period is input into the contrast learning model, the real operation mode branch and the non-operation mode branch are called in the contrast learning model respectively, and the encoding operation is performed on the input sequence in each branch to convert the input sequence into a feature representation vector with fixed dimensions; In the real operation mode branch, the feature representation vector of the input sequence is subjected to vector dot product calculation with the standard feature vectors in the real operation mode sample library one by one, and the dot product result is divided by the product of the lengths of the two vectors to obtain a similarity score. Then, all the similarity scores are accumulated and inverted to form a first matching energy; In the non-operation mode branch, the feature representation vector of the input sequence is subjected to vector dot product calculation with the standard feature vectors in the non-operation mode sample library one by one, and the dot product result is divided by the product of the lengths of the two vectors to obtain a similarity score. Then, all the similarity scores are accumulated and inverted to form a second matching energy; The first matching energy and the second matching energy are subjected to subtraction operation to obtain an energy difference, which is used to represent the relative matching strength of the input sequence between the real operation mode and the non-operation mode; Combine the energy difference with the operation intention probability output by the space-time intention determination engine, increase the value of the operation intention probability when the energy difference is greater than zero, and decrease the value of the operation intention probability when the energy difference is less than zero, so as to output the corrected operation intention probability, and use the corrected operation intention probability as a correction signal for subsequent comprehensive score calculation.

[0014] Technical effects and advantages of the present application: By introducing the operation pre-determination mechanism on the basis of face recognition and gaze tracking, it is confirmed whether the user has an operation intention before interaction, thereby effectively avoiding natural actions, passers-by or temporary gaze from being misjudged as effective operations, and solving the most prominent mis-touch problem in medium-free holographic interaction. By constructing a semantic interaction body, the medium-free holographic image is divided into interaction voxels or interaction planes, and geometric intersection calculation is performed with the gaze ray, so that the interaction determination is established on the basis of accurate space coordinates, avoiding the ambiguity of gaze determination caused by boundaryless projection, and improving the spatial accuracy of interaction triggering. By fusing and analyzing the face stability, head posture, blink frequency, gaze trajectory and hand proximity trajectory through the space-time intention determination engine, and calculating the time delay correlation coefficient, the watching behavior and the real operation intention can be distinguished, the misjudgment caused by single-dimensional parameters is reduced, and the reliability of time sequence determination is improved. In the scenario where multiple people exist at the same time, the consistency of candidate face trajectories is maintained through the joint probability data association method, the confidence of the main operator is calculated by combining the gaze hit probability, the operation intention probability and the energy difference, and the identity arbitration is performed through the strategy network, so as to ensure that only the unique main operator is output, and the mis-touch under the interference of multiple users is avoided. BRIEF DESCRIPTION OF DRAWINGS

[0015] Fig. 1 It is a schematic diagram of the medium-free holographic imaging principle of the present application.

[0016] Fig. 2 It is a method step flowchart of the present application. DETAILED DESCRIPTION

[0017] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0018] Embodiment 1: Referring to the drawings attached in the specification Figs. 1-2 An embodiment of the present application is a medium-free holographic anti-mis-touch interaction method, which comprises: S1, generating a medium-free holographic image by a medium-free holographic imaging device, the medium-free holographic image containing a semantic interactive body composed of a plurality of interactive voxels or interactive planes; It should be noted that for S1, the semantic interactive body refers to an interactive reference structure artificially constructed in the medium-free holographic image, which divides the projected image into interactive regions that can be recognized and utilized by the system. The interactive voxel refers to the division of three-dimensional space into small cubic units, each unit representing a three-dimensional position point in the image, used to accurately represent the range of the interactive region in space. The interactive plane refers to the rectangular or polygonal region divided in the two-dimensional projection layer, used to mark the two-dimensional range where the user may interact. In short, the semantic interactive body is a spatial indexing structure composed of a large number of interactive voxels or interactive planes, which converts the originally ambiguous holographic image interactive region into a unit set that can be geometrically calculated and matched, thereby providing a clear spatial reference for subsequent line-of-sight ray alignment and hit determination. S2, capturing a face image by a camera arranged on the medium-free holographic imaging device, and identifying a face position, a face orientation, a face size and an eye distance parameter based on the face image, for determining a spatial range of the face in the medium-free holographic image; S3, when the face is in the spatial range, determining a first time length that the face is in the spatial range, and when the first time length is greater than a first preset value, entering an action recognition mode to recognize a user gesture; Further, when the human eye is in the spatial range, a second time length that the line of sight of the human eye stays in the spatial range is determined, and when the second time length is greater than a second preset value, the action recognition mode is entered to recognize the user gesture. The second time length that the line of sight of the human eye stays in the spatial range when the human eye is in the spatial range includes: capturing an eye image, solving a pupil center position based on the eye image and generating a line of sight ray, geometrically aligning the line of sight ray with the semantic interactive body, and forming a line of sight hit probability.

[0019] Based on embodiment 1, embodiment 2 is further included: S4, taking the face stability, head posture change parameter, blinking frequency, line of sight trajectory and hand proximity trajectory in a continuous period as an input sequence, inputting the input sequence into a space-time intention determination engine, and outputting an operation intention probability; S5, calculating the matching energy of the input sequence with the real operation mode and the non-operation mode respectively based on the contrast learning model, and forming an energy difference according to the difference between the two, and taking the energy difference as a correction signal for the operation intention probability.

[0020] On the basis of embodiment 2, embodiment 3 is further included: The difference from embodiment 1 is that S6 is further included: in the presence of more than or equal to two candidate faces, a joint probability data association method is used to track all candidate faces, the main operator confidence is calculated based on the line-of-sight hit probability, the operation intention probability and the energy difference, and the strategy network outputs the unique main operator identity under the constraint of a single main operator; It needs to be explained for S6 that the joint probability data refers to the probability calculation of one-to-one correspondence between the multiple face observations collected by the camera and the multiple candidate face tracks established in the system within the same time window, and the joint probability distribution formed by combining all possible observation-track corresponding relationships; the joint probability distribution not only considers the matching possibility of a single observation and a single track, but also comprehensively calculates the joint consistency of multiple faces under the overall allocation, which is used to reduce the false association in the multi-target case; When tracking all candidate faces, first, the position parameters, orientation parameters and size parameters of the candidate faces in each frame of image are extracted as observation data input into the joint probability data association method; then the Euclidean distance and pose similarity between the predicted face track state and the current observation data are used to establish the observation-track matching probability; finally, the joint probability distribution is used to select the allocation scheme with the maximum probability from multiple possible allocation results, so as to maintain the identity consistency of each candidate face track between consecutive frames and realize stable tracking; Under the condition that the candidate faces are tracked, the system calculates the line-of-sight hit probability, the operation intention probability and the energy difference for each candidate face respectively, and inputs them as an input vector into the confidence calculation module, which obtains the main operator confidence of each candidate face through weighted accumulation and normalization operation; then the confidence vectors of all candidate faces are input into the strategy network, which performs maximum value selection under the constraint of a single main operator, and outputs the candidate face with the maximum confidence and exceeding the threshold as the unique main operator identity, so as to realize identity arbitration in the scene where multiple people exist at the same time; The strategy network is a decision model pre-trained and stored in the system, which is used to select and output the unique main operator identity according to the preset single main operator constraint condition after inputting the main operator confidence vector.

[0021] S7 is further included: based on the line-of-sight hit probability, the operation intention probability, the energy difference and the main operator confidence, a weighted calculation is performed to obtain a comprehensive score, and when the comprehensive score exceeds a preset threshold, a medium-free holographic interaction operation is triggered, and the final confirmation is completed through gesture recognition or action recognition after the trigger.

[0022] In S1, the process of constructing the semantic interactive body in the medium-free holographic image comprises the following steps: layering and dividing the three-dimensional space region where the medium-free holographic image is located, cutting the whole space into multiple space slices according to the depth direction, and then performing grid decomposition in each space slice to generate multiple groups of interactive voxels or interactive planes; For the interactive voxels or interactive planes, the three-dimensional coordinate range of each unit in the device coordinate system is determined based on the projection geometric parameters of the medium-free holographic image, and the spatial position matrix of the unit is solved from the three-dimensional coordinate range; then the light ray direction vector and intensity distribution passing through the unit are calculated based on the light field distribution parameters, the normal vector of the unit is solved from the direction vector, and the geometric boundary parameters of the unit are solved from the intensity distribution; The light field distribution parameters include light ray direction, light intensity distribution, wavelength spectrum and phase information, which are used to describe the propagation characteristics of the medium-free holographic image in space; the intensity distribution includes the energy density, brightness gradient and amplitude changing with time of the light ray at the spatial position, which is used to characterize the light intensity and variation law of the medium-free holographic image in different regions; The spatial position matrix, normal vector and geometric boundary parameters are linearly combined according to the preset weight to form a spatial description set used for subsequent line-of-sight ray and semantic interactive body alignment; The spatial description set is established according to the relevance of the interactive intention to form a hierarchical index structure, and the hierarchical index structure is defined as a semantic interactive body, so as to provide multi-level alignment reference in the matching calculation process of the line-of-sight ray and the semantic interactive body.

[0023] In S2, the camera arranged on the medium-free holographic imaging device is used to capture the face image, and the face contour feature points are extracted from the face image to calculate the face rectangular boundary; the face rectangular boundary refers to the minimum circumscribed rectangle generated based on the face contour feature points in the detected face region, which is used to limit the two-dimensional range of the face in the image; The coordinate vector of the face position is calculated based on the face contour feature points, the face size is calculated based on the aspect ratio of the face rectangular boundary, the eye distance parameter is calculated based on the Euclidean distance between the pupils of the two eyes, and on this basis, the two-dimensional coordinates of the measurement points in the face image are corresponded to the corresponding three-dimensional template coordinates, the projection matrix that makes the two-dimensional points and the three-dimensional points consistent under the camera imaging geometry is calculated, and the rotation component is decomposed from the projection matrix, and finally the face orientation vector representing the face orientation is obtained; the face orientation vector is used to indicate the direction of the face in the three-dimensional space coordinate system; The measuring points include the eye corner points, the nose tip point and the mouth corner points of the two eyes; the three-dimensional template coordinates are fixed reference positions of the preset human face measuring points in a standard three-dimensional coordinate system, which are used to correspond to the two-dimensional feature points in the image, so as to solve the orientation of the human face; The coordinate vector of the human face position, the human face size, the eye distance parameter and the human face orientation vector are input into a coordinate mapping module to complete the conversion to the space coordinate system of the medium-free holographic image, so as to output the spatial range of the human face in the medium-free holographic image; wherein the coordinate mapping module is used to convert the human face position coordinate vector, the human face size, the eye distance parameter and the human face orientation vector obtained by the camera into the space coordinate system of the medium-free holographic image through rotation, translation and scaling calculation, so as to determine the actual spatial range of the human face in the holographic image.

[0024] In S3, under the condition that the human face position range is determined, the human eye image is collected by the camera, and the iris edge point set is extracted in the human eye image; wherein the iris edge point set refers to a set of continuous pixel point coordinates extracted along the junction of the iris and the sclera in the collected human eye image, which is used to describe the external contour of the iris. In actual application, the iris edge point set can be collected by performing edge detection and feature extraction algorithms on the human eye image, including calculating the gray gradient of the junction area of the iris and the sclera and outputting the boundary pixel point coordinates; Based on the iris edge point set, the iris center position is determined, the iris center position is taken as the pupil center position coordinate, and the straight line direction between the pupil center position coordinate and the optical center of the camera is defined as the line of sight ray direction vector; The line of sight ray direction vector and the space description set of the semantic interactive body are calculated to obtain the intersection point distribution, and the line of sight hit probability is formed based on the corresponding relationship between the intersection point distribution and the interactive unit index; It should be noted that the process of geometric intersection calculation first represents the line of sight ray direction vector determined by the pupil center position coordinate and the optical center of the camera as a parametric equation, and represents the space description set of each interactive unit in the semantic interactive body as a geometric boundary equation; then the parametric equation of the line of sight ray is substituted into the geometric boundary equation to solve the intersection point coordinates of the ray and the boundaries of each interactive unit; in the case that the intersection point coordinates are solved, the distribution of the intersection point coordinates in each interactive unit is further counted, and the corresponding relationship between the intersection point and the interactive unit is established according to the distribution; finally, the corresponding relationship is taken as the basis for judgment to form the line of sight hit probability, which is used to represent the probability of the user's line of sight falling into the interactive unit in the semantic interactive body.

[0025] In S4, the execution process of the spatio-temporal intention determination engine on the input sequence includes: time synchronization and sampling alignment of the face stability, head posture change parameter, blink frequency, gaze trajectory and hand proximity trajectory in the continuous time period, and construction of a fixed-length time sequence window; performing normalization and first-order difference operation on each parameter in the time sequence window to generate a stability index of the face stability, an angular velocity vector of the head posture change parameter, a time density vector of the blink frequency, a speed and residence time vector of the gaze trajectory, and a speed and distance vector of the hand proximity trajectory; calculating the time delay correlation coefficient of the gaze trajectory and the hand proximity trajectory based on the generated vectors, and establishing a corresponding relationship with the interaction unit index given by the semantic interaction body to form a spatio-temporal feature matrix, and performing spatio-temporal fusion and thresholding statistics according to the spatio-temporal feature matrix to output the operation intention probability; It should be noted that the time delay correlation coefficient refers to that in a fixed time window, the time sequence of the gaze trajectory and the time sequence of the hand proximity trajectory are respectively compared by shifting according to different time delay amounts, the correlation degree of the two groups of time sequences under each delay amount is calculated, the maximum correlation value of the two under a certain delay is solved, and the maximum correlation value is taken as the time delay correlation coefficient, which is used to represent the coupling strength of the user's gaze change and the hand proximity action in the time sequence. In addition, in S4, the normalization operation refers to mapping the original values of the face stability, head posture change parameter, blink frequency, gaze trajectory and hand proximity trajectory to a unified numerical interval through linear transformation, so as to eliminate the dimensional difference.

[0026] The first-order difference operation in S4 refers to calculating the difference value of adjacent values of the normalized sequence in the continuous time period point by point to solve the speed or incremental feature of the parameter changing with time.

[0027] The process of calculating the energy difference based on the contrast learning model and correcting the operation intention probability includes the following steps: The input sequence in the continuous time period is input into the contrast learning model, and the real operation mode branch and the non-operation mode branch in the contrast learning model are called respectively, and the encoding operation is performed on the input sequence in each branch to convert the input sequence into a feature representation vector with fixed dimension; wherein the real operation mode branch is a calculation path in the contrast learning model, which is used to receive the input sequence and compare with the feature vector in the real operation mode sample library, so as to calculate the matching degree between the input sequence and the real operation behavior; the non-operation mode branch is another calculation path in the contrast learning model, which is used to receive the same input sequence and compare with the feature vector in the non-operation mode sample library, so as to calculate the matching degree between the input sequence and the non-operation behavior; In the real operation mode branch, the feature representation vector of the input sequence is calculated with the standard feature vector in the real operation mode sample library one by one, and the dot product result is divided by the product of the two vector lengths to obtain a similarity score, and all similarity scores are accumulated and inverted to form a first matching energy; In the non-operation mode branch, the feature representation vector of the input sequence is calculated with the standard feature vector in the non-operation mode sample library one by one, and the dot product result is divided by the product of the two vector lengths to obtain a similarity score, and all similarity scores are accumulated and inverted to form a second matching energy; The first matching energy and the second matching energy are subtracted to obtain an energy difference, which is used to represent the relative matching strength of the input sequence between the real operation mode and the non-operation mode; The energy difference is combined with the operation intention probability output by the space-time intention judgment engine, and when the energy difference is greater than zero, the value of the operation intention probability is increased, and when the energy difference is less than zero, the value of the operation intention probability is decreased, so as to output the modified operation intention probability, and the modified operation intention probability is used as a modified signal for subsequent comprehensive score calculation.

[0028] Overall, the core problem of existing medium-free holographic interaction is that the projected image is suspended in the air, the interaction area is open and lacks physical boundaries, and user's unintentional behavior such as random head movement, passerby passing, gesture interference, etc. may be incorrectly recognized as valid instructions, resulting in false touch; after recognizing this problem, the scheme selects "operation pre-confirmation" as the core strategy to prevent false touch, and gradually establishes a complete protection logic by introducing face recognition, gaze tracking, space-time behavior judgment, contrast learning correction, and multi-user identity arbitration mechanisms; The establishment of this logic is not the superposition of single linear links, but a structured process of cooperation; In the implementation step, the scheme first generates a medium-free holographic image through a medium-free holographic imaging device, and constructs a semantic interaction body in the image, divides the interaction area into interaction voxels or interaction planes, so that the subsequent gaze and action have a clear spatial reference; then the face image is collected through the camera, and the face position, face orientation, face size and eye distance parameters are solved based on this, and then they are converted to the spatial coordinate system of the medium-free holographic image through the coordinate mapping module, so as to accurately limit the user's interaction space range; after the face range is determined, the processing link of the eye image is entered, the iris edge point set is extracted and the pupil center position is solved, the gaze ray direction vector is generated combined with the optical center of the camera, and the geometric intersection calculation with the semantic interaction body is performed to form the gaze hit probability; This processing chain ensures that the user's gaze direction must be aligned with the projected interaction area before entering the next step of judgment; Subsequently, a space-time intention determination engine is introduced, and the face stability, head posture change parameter, blink frequency, gaze trajectory and hand approaching trajectory in continuous time periods are synchronized and processed, through normalization, first-order difference and time delay correlation coefficient calculation, these dynamic signals are fused into a space-time feature matrix; the matrix is used to interpret whether the user has actual operation intention, for example, whether the gaze fixation is accompanied by hand movement approaching; the result output is the operation intention probability, which becomes an important basis for interaction credibility; on this basis, the scheme further introduces a contrast learning model, which compares the input sequence with the real operation mode and non-operation mode through double-branch comparison, respectively forming the first matching energy and the second matching energy, and calculating the difference between the two to obtain the energy difference; the energy difference is used as a correction signal for the operation intention probability, so that the system can clearly distinguish between real intention and non-intention behavior through positive and negative differences, thereby enhancing the robustness of the determination; In addition, in a multi-person environment, multiple candidate faces are tracked through a joint probability data association method, the gaze hit probability, operation intention probability and energy difference are input into a confidence solving module to solve the main operator confidence, and then the strategy network outputs a unique identity under the constraint of a single main operator, solving the identity confusion problem when multiple people exist at the same time; finally, the gaze hit probability, operation intention probability, energy difference and main operator confidence are weighted to obtain a comprehensive score, and when the comprehensive score exceeds a preset threshold, a medium-free holographic interaction operation is triggered, and after triggering, a gesture or action is required for final confirmation, to realize a double-layer insurance mechanism.

[0029] The above only describes the preferred embodiments of the present application and is not used to limit the present application, any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A method for preventing mistaken touch interaction based on medium-free holography, characterized in that, Comprise: S1, generating a medium-free holographic image by a medium-free holographic imaging device, the medium-free holographic image containing a semantic interactive body composed of a plurality of interactive voxels or interactive planes; S2, capturing a face image by a camera arranged on the medium-free holographic imaging device, and identifying a face position, a face orientation, a face size and an eye distance parameter based on the face image, for determining a spatial range of the face in the medium-free holographic image; S3, when the face is in the spatial range, determining a first time length that the face is in the spatial range, and when the first time length is greater than a first preset value, entering an action recognition mode to identify a user gesture; And / or, when a human eye is in the spatial range, determining a second time length that the line of sight of the human eye stays in the spatial range, and when the second time length is greater than a second preset value, entering the action recognition mode to identify the user gesture.

2. The false touch prevention method based on medium-free hologram according to claim 1, wherein: The second time length that the line of sight of the human eye stays in the spatial range when the human eye is in the spatial range comprises: Capturing an eye image, solving a pupil center position and generating a line of sight ray based on the eye image, geometrically aligning the line of sight ray with the semantic interactive body to form a line of sight hit probability. 3.The false touch prevention method based on medium-free hologram according to claim 1, wherein: The method further comprises: S4, inputting a face stability, a head posture change parameter, a blink frequency, a line of sight trajectory and a hand proximity trajectory in a continuous period as an input sequence to a space-time intention determination engine, and outputting an operation intention probability; S5, calculating a matching energy of the input sequence with a real operation mode and a non-operation mode respectively based on a contrast learning model, and forming an energy difference according to the difference between the two, and taking the energy difference as a correction signal for the operation intention probability.

4. The anti-mis-touch interaction method based on medium-free holography according to claim 3, characterized in that: Further comprising S6: in the case of more than or equal to two candidate faces, using a joint probability data association method to track all candidate faces, calculating a main operator confidence based on the line of sight hit probability, the operation intention probability and the energy difference, and outputting a unique main operator identity under the constraint of a single main operator through a strategy network.

5. The anti-mis-touch interaction method based on medium-free holography according to claim 4, characterized in that: Further comprising S7: performing weighted calculation based on the line of sight hit probability, the operation intention probability, the energy difference and the main operator confidence to obtain a comprehensive score, and triggering a medium-free holographic interaction operation when the comprehensive score exceeds a preset threshold, and completing final confirmation through gesture recognition or action recognition after triggering.

6. The anti-mis-touch interaction method based on medium-free holography according to claim 5, characterized in that: In S1, the process of constructing a semantic interactive body in the medium-free holographic image comprises the following steps: layering and dividing a three-dimensional space region where the medium-free holographic image is located, cutting the whole space into a plurality of spatial slices according to the depth direction, and then performing grid decomposition in each spatial slice to generate a plurality of interactive voxels or interactive planes; For the interactive voxel or interactive plane, a three-dimensional coordinate range of each unit in a device coordinate system is determined based on projection geometry parameters of the medium-free holographic image, and a spatial position matrix of the unit is solved from the three-dimensional coordinate range; subsequently, a light ray direction vector and an intensity distribution passing through the unit are calculated based on light field distribution parameters, a normal vector of the unit is solved from the direction vector, and geometric boundary parameters of the unit are solved from the intensity distribution; The light field distribution parameters include light ray direction, light intensity distribution, wavelength spectrum and phase information; and the intensity distribution includes energy density of the light ray on the spatial position, brightness gradient and amplitude changing over time. The spatial position matrix, the normal vector and the geometric boundary parameters are linearly combined according to preset weights to form a spatial description set; The spatial description set is established in a hierarchical index structure according to the relevance of the interactive intention, and the hierarchical index structure is defined as a semantic interactive body.

7. The anti-mis-touch interactive method based on medium-free holography according to claim 6, characterized in that: In S2, a camera disposed on the medium-free holographic imaging device is used to capture a face image, and face contour feature points are extracted from the face image; a face rectangular boundary is solved by using the face contour feature points; A coordinate vector of the face position is calculated based on the face contour feature points, a face size is calculated based on an aspect ratio of the face rectangular boundary, an eye distance parameter is calculated based on an Euclidean distance between pupils of two eyes, and a corresponding relationship between two-dimensional coordinates of a measurement point in the face image and corresponding three-dimensional template coordinates is established; a projection matrix is solved by making the two-dimensional point and the three-dimensional point consistent under camera imaging geometry, and a rotation component is decomposed from the projection matrix, so as to finally obtain a face orientation vector representing a face orientation; The face orientation vector is used to indicate a direction of the face in a three-dimensional space coordinate system; The measurement point includes a corner point of the two eyes, a nose tip point and a corner point of the mouth; and the three-dimensional template coordinates are fixed reference positions of the face measurement points in a standard three-dimensional coordinate system; The coordinate vector of the face position, the face size, the eye distance parameter and the face orientation vector are input into a coordinate mapping module to complete conversion to a medium-free holographic image space coordinate system, so as to output a spatial range of the face in the medium-free holographic image.

8. The anti-mis-touch interactive method based on medium-free holography according to claim 7, characterized in that: In S3, under the condition that the face position range is determined, a camera is used to capture an eye image, and an iris edge point set is extracted from the eye image; An iris center position is determined based on the iris edge point set, the iris center position is taken as a pupil center position coordinate, and a straight line direction between the pupil center position coordinate and an optical center of the camera is defined as a line of sight ray direction vector; The line of sight ray direction vector and the spatial description set of the semantic interactive body are subjected to geometric intersection calculation to solve an intersection point distribution, and a line of sight hit probability is formed based on a corresponding relationship between the intersection point distribution and an interactive unit index.

9. The method of claim 8, wherein the method further comprises: In S4, the execution process of the spatio-temporal intention determination engine on the input sequence includes: time synchronization and sampling alignment of the face stability, head posture change parameter, blink frequency, gaze trajectory and hand proximity trajectory in the continuous time period, and construction of a fixed-length time window; performing normalization and first-order difference operation on each parameter in the time window to generate a stability index of face stability, an angular velocity vector of head posture change parameter, a time density vector of blink frequency, a speed and dwell time vector of gaze trajectory, and a speed and distance vector of hand proximity trajectory; calculating the time delay correlation coefficient of the gaze trajectory and the hand proximity trajectory based on the generated vectors, and establishing a corresponding relationship with the interaction unit index given by the semantic interaction body to form a spatio-temporal feature matrix, and performing spatio-temporal fusion and thresholding statistics according to the spatio-temporal feature matrix to output the operation intention probability.

10. The method of claim 9, wherein the method further comprises: The process of calculating the energy difference based on the contrast learning model and correcting the operation intention probability includes the following steps: inputting the input sequence in the continuous time period into the contrast learning model, calling the real operation mode branch and the non-operation mode branch in the contrast learning model respectively, and performing encoding operation on the input sequence in each branch to convert the input sequence into a feature representation vector with fixed dimension; in the real operation mode branch, the feature representation vector of the input sequence and the standard feature vector in the real operation mode sample library are calculated by vector dot product one by one, and the dot product result is divided by the product of the two vector lengths to get the similarity score, then all the similarity scores are accumulated and inverted to form the first matching energy; in the non-operation mode branch, the feature representation vector of the input sequence and the standard feature vector in the non-operation mode sample library are calculated by vector dot product one by one, and the dot product result is divided by the product of the two vector lengths to get the similarity score, then all the similarity scores are accumulated and inverted to form the second matching energy; performing subtraction operation on the first matching energy and the second matching energy to get the energy difference, which is used to represent the relative matching strength of the input sequence between the real operation mode and the non-operation mode; combine the energy difference with the operation intention probability output by the spatio-temporal intention determination engine, when the energy difference is greater than zero, increase the value of the operation intention probability, when the energy difference is less than zero, decrease the value of the operation intention probability, output the corrected operation intention probability as the correction signal for subsequent comprehensive score calculation.

Citation Information

Patent Citations

  • Gesture Interaction method and system based on virtual human

    CN108459712A

  • Man-machine interaction method, device and equipment based on medium-free holographic technology and medium

    CN119002689A

  • Elevator control method and device based on holographic projection, equipment and storage medium

    CN119038335A

  • Holographic intelligent system based on digital human model technology and equipment thereof

    CN119620863A

  • Methods for two-stage hand gesture input

    US20200301513A1