Intelligent air imaging system based on visual adaptation

The smart air imaging system dynamically adjusts focal distance and display height based on user height and gaze analysis, providing precise touch feedback and reducing visual fatigue and misalignment issues.

CN120321380AInactive Publication Date: 2025-07-15SHANGHAI HUANSHENG INTELLIGENT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510450335.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-07-15
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing air imaging technology cannot dynamically adjust the imaging area based on user height and pupil position, resulting in misalignment of the display content and visual fatigue in multiple user scenarios, and the touch feedback is unfriendly, especially for people with poor vision.

Method used

The binocular depth camera and infrared tracking sensor are used to obtain the user's height and pupil center, and combined with the electric sliding lifting component and the DCT-plate lens group, dynamically adjust the imaging height and focus; determine the gaze point through the visual adaptive unit and perform special display; build an LSTM timing trajectory prediction model for touch error compensation; build a priority rating decision model to deal with multi-user conflicts.

Benefits of technology

The dynamic alignment of the imaging area and the user's vision is achieved, which reduces visual fatigue, improves reading efficiency, provides clear tactile feedback, and avoids misoperation and multi-user conflicts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120321380A_ABST
    Figure CN120321380A_ABST
Patent Text Reader

Abstract

The invention relates to an intelligent air imaging system based on visual adaptation, and relates to an air imaging technology. Comprising a hardware support module and an algorithm optimization control module. The hardware supporting module comprises a binocular depth camera, an infrared tracking sensor, a TOF sensor, an electric sliding lifting assembly for bearing a DCT-plate lens group, a corneal curvature meter and an ultrasonic transmitter; the algorithm optimization control module comprises a visual self-adaption unit, a touch control error compensation unit and an exception handling unit. The method can adapt to the height of the user, highlights and / or amplifies the attention area according to the fixation point in the pupil optical axis direction of the user, the reading efficiency is improved, and meanwhile the visual fatigue degree is reduced; the LSTM time sequence model is adopted to pre-judge the touch track of the user, the ultrasonic phase array is combined to dynamically adjust the frequency of the acoustic radiation force field, perceptible tactile feedback is generated, and the phenomenon that the user operates by mistake or touch is not in place is avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to air imaging technology, and specifically to an intelligent air imaging system based on visual adaptation. Background Art

[0002] In recent years, air imaging technology has gradually been applied to fields such as interactive display and medical imaging by generating suspended images in free space through optical elements (such as DCT-plate lens groups). However, the existing technology still has the following core defects:

[0003] Traditional systems rely on fixed focal lengths and display heights, and cannot dynamically adjust the imaging area according to the user's height and pupil position, resulting in misalignment of the displayed content or visual fatigue in multi-user scenarios; air imaging touch feedback mostly uses infrared light curtains or capacitive sensing, which requires the user to actually touch to trigger, there are risks of touch delay and false touch, and it is impossible to predict the operation intention and compensate for errors in advance. In addition, existing air imaging devices cannot perform special displays for the user's gaze area, which is not user-friendly to people with poor eyesight (such as the elderly). Therefore, we provide an intelligent air imaging system based on visual adaptation. Summary of the Invention

[0004] The purpose of the present invention is to provide an intelligent air imaging system based on visual adaptation to solve the above problems.

[0005] The present invention can be achieved through the following technical solutions: an intelligent air imaging system based on visual adaptation, including a hardware support module and an algorithm optimization control module;

[0006] The hardware support module includes a binocular depth camera, an infrared tracking sensor, a TOF sensor, an electric sliding lifting component carrying a DCT-plate lens group, a keratometer, and an array of ultrasonic transmitters;

[0007] The algorithm optimization control module includes a visual adaptation unit, a touch error compensation unit, and an exception handling unit; among them, the visual adaptation unit obtains the user's pupil center and corneal curvature center, determines the optical axis direction and its fixation point in the air imaging area, generates a high-priority area around the fixation point and performs special display processing; the touch error compensation unit constructs an LSTM time series trajectory prediction model to predict the distance between fingertip controls in a certain future time, and drives the ultrasonic transmitter to adjust different frequencies according to the distance size to simulate the fingertip touch feeling; the exception handling unit constructs a scoring decision model to score multiple users according to their priorities, and refreshes regularly. The system responds to the operations of the user with the highest score and blocks the behaviors of other users.

[0008] A further technical improvement of the present invention lies in: the steps of the visual adaptation unit performing special display processing based on the user's height and the position of the fixation point, including:

[0009] Step 1: Obtain the user's height based on a binocular depth camera, and drive the DCT-plate lens group by an electric sliding lifting component to adjust the height to adapt to the height.

[0010] Step 2: Detect the face area based on a face detection algorithm and determine the eye area, and then locate the two-dimensional pixel coordinates of the pupil center in the image.

[0011] Step 3: First convert the two-dimensional pixel coordinates into the camera's three-dimensional coordinates, and then convert the camera's three-dimensional coordinates into the three-dimensional pupil center coordinates in the system's absolute coordinate system.

[0012] Step 4: Obtain the corneal curvature center coordinates and construct an optical axis direction vector in combination with the pupil center coordinates, and determine the intersection point of the vector and the air imaging area, that is, the fixation point.

[0013] Step 5: Determine different forms of high-priority areas according to the content form in the neighborhood where the fixation point is located, perform highlighting display and / or magnification processing on the high-priority areas, and perform grayscale display on the non-high-priority areas at the same time.

[0014] Step 6: Set a trigger condition for the special display method of the high-priority area, and continuously track the user's vision and perform display changes according to the movement of the fixation point.

[0015] A further technical improvement of the present invention is that when constructing the LSTM time series trajectory prediction model, use the finger trajectory data with a time step of 100 ms to predict the position where the finger advances in the next 50 ms. Define the input as the three-dimensional coordinates of the finger with time series, and define the output as the predicted distance between the fingertip and the control.

[0016] A further technical improvement of the present invention is that the priority scoring formula is

[0017] where i represents the number of the person detected to enter the scene, d i is the distance from the detected user's pupil center along the line of sight to the fixation point, t zsi represents the continuous fixation time of the corresponding user on the imaging area, β and γ are the distance weight and the fixation time weight respectively, and β + γ = 1 and β > 0.5 > γ.

[0018] A further technical improvement of the present invention is that when adjusting the height of the DCT-plate lens group in Step 1, its target height is set to H = 0.6h + 300, where H is the target height and h is the user's height.

[0019] A further technical improvement of the present invention lies in: in step three, the process of converting the pixel coordinates of the pupil into three-dimensional pupil center coordinates includes depth information correction and coordinate transformation:

[0020] Construct a three-dimensional camera coordinate system and calculate the parallax depth information Z of the pupil in the camera coordinate system c , that is where f represents the focal length of the camera, b represents the baseline distance of the binocular camera, and d represents the horizontal displacement difference of the pixel coordinates of the pupil center in the left and right images of the binocular camera;

[0021] Obtain the depth information directly measured by the TOF sensor where Δt is the phase difference between the emitted light and the detected reflected light;

[0022] Calculate the finally corrected depth information Z final =(1 - α)·Z c +α·Z TOF ;

[0023] where α is the compensation weight coefficient, and

[0024] Subsequently, obtain the three-dimensional coordinates of the camera according to the corrected depth information, and perform coordinate transformation according to the calibrated external parameter matrix to obtain the system absolute coordinates.

[0025] A further technical improvement of the present invention lies in: the method for determining the contour form of the high-priority area includes:

[0026] The text area generates a rectangular contour by OCR to recognize the text boundary;

[0027] The image area segments the contour by Canny edge detection;

[0028] The interactive control is independently segmented to generate an interactive hot zone.

[0029] A further technical improvement of the present invention lies in: the triggering condition for the special display of the high-priority area is: when the fixation point continuously fixes within a certain neighborhood range of the initial fixation position for a time exceeding the continuous fixation time threshold, trigger the highlight display and / or the enlarged display.

[0030] Compared with the prior art, the present invention has the following beneficial effects:

[0031] 1. Obtain the user's height data through the binocular depth camera, and combine the dynamic adjustment of the height of the DCT-plate lens group by the electric sliding lifting component to always align the imaging plane with the center of the user's visual field. Compared with the traditional fixed-height imaging system, it can match the current user's height and facilitate the user's use;

[0032] 2. Based on the pupil optical axis direction projection and dynamic grid division technology, it preferentially magnifies the fixation point area and grayscales the non-focus area at the same time, so as to quickly help customers focus on the current key information, improve the reading efficiency, and reduce the degree of visual fatigue at the same time;

[0033] 3. Adopt the LSTM time series model to predict the user's touch trajectory, and combine the ultrasonic phased array to dynamically adjust the frequency of the acoustic radiation force field to generate perceivable tactile feedback, so that the user can clearly perceive whether the current finger triggers the corresponding interactive control, and can give the user sufficient feedback to avoid the occurrence of user misoperation or inaccurate touch;

[0034] 4. The present invention also adopts a dynamic priority scoring mechanism to comprehensively consider the user's pupil distance and continuous fixation time, regularly refresh the score and respond to the user with the highest score, so as to avoid the conflict phenomenon caused by multiple users appearing in the scene. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] For the convenience of those skilled in the art to understand, the present invention will be further described below with reference to the accompanying drawings.

[0036] Figure 1 is the system block diagram of the present invention;

[0037] Figure 2 is the schematic diagram of the working process of the visual adaptation unit of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0038] In order to further elaborate on the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the following will describe in detail the specific implementation manners, structures, features and effects of the present invention with reference to the accompanying drawings and preferred embodiments.

[0039] Please refer to Figure 1-2 As shown, an intelligent air imaging system based on visual adaptation includes a hardware support module and an algorithm optimization control module; among them, the hardware support module includes a binocular depth camera, an infrared tracking sensor, a TOF sensor, and an electric sliding lifting component carrying a DCT-plate lens group; the algorithm optimization control module includes a visual adaptation unit, a touch error compensation unit, and an exception handling unit;

[0040] The visual adaptation unit obtains and simulates the pupil visual focus based on the binocular depth camera, infrared tracking sensor, and TOF sensor, and determines the fixation area of the visual focus in the imaging area. Specifically:

[0041] (1) Construct a three-dimensional system coordinate system with a fixed position in the system as the origin. Obtain the user's height through a binocular depth camera, and drive the DCT-plate lens group to perform lifting motion by an electric sliding lifting component, so as to complete the adaptation adjustment for the user's height.

[0042] Specifically, assuming the user's height is h (in millimeters), the target height of the lens group is set to H = 0.6h + 300; drive the electric sliding lifting component according to the gap between the current height and the target height of the lens group, and set the maximum adjustment response time to perform the lifting operation with adaptive speed.

[0043] (2) Obtain the spatial coordinates of the pupil center

[0044] Use a face detection algorithm in the obtained image, such as the CascadeClassifier function in OpenCV, which is a cascade classifier when performing face detection in Opencv, to detect the face area from the image;

[0045] Based on the face area image, perform grayscale processing and Gaussian filtering operations, and then extract face feature points based on the Dlib library, and use 68 feature points to mark the key positions of the face, so as to determine the approximate position of the eye area in the face area image;

[0046] Perform Canny edge detection on the eye area to obtain the contour of the eye area, and perform Hough circle transformation within the contour to locate the pupil, so as to obtain the two-dimensional coordinates of the pupil in the image, that is, pixel coordinates;

[0047] (3) Convert the image two-dimensional coordinates into system three-dimensional coordinates

[0048] Let the image two-dimensional coordinates be (x, y), then the depth information Z of the pupil in the camera coordinate system is calculated according to the binocular disparity c , that is where f represents the camera focal length; b represents the baseline distance of the binocular camera, that is, the distance between the centers of the two cameras in the binocular camera; d represents the horizontal displacement difference of the pixel coordinates of the pupil center in the left and right images of the binocular camera;

[0049] Since the depth error of binocular vision increases with the increase of the distance, the depth information is directly measured by a TOF sensor where Δt is the time difference between the TOF sensor emitting light and detecting the reflected light;

[0050] Correct the disparity depth information according to the directly measured depth information:

[0051] Z final = (1 - α)·Z c + α·Z TOF ;

[0052] where α is a compensation weight coefficient, and

[0053] Taking the optical center of one of the lenses of the binocular camera as the origin and the optical axis direction as the Z-axis to construct a camera three-dimensional coordinate system, and using the internal parameter matrix of the camera (including the focal lengths f x , f y and the principal points c x , c y ), the pixel coordinates are converted into camera plane coordinates (x c , y c ), where

[0054] Combined with the depth information Z final calculate the camera three-dimensional coordinates (X c , Y c , Z c ), where X c = x c ·Z funal , Y c = y c ·Z funal , Z c = Z funal ;

[0055] Let the system absolute coordinate system be (X w , Y w , Z w ), calibrate the external parameter matrix M of the binocular camera, including the rotation matrix R and the translation vector T, Then

[0056] (4) Obtain the eye's line of sight and determine its projection coordinates in the imaging area

[0057] Integrate a corneal topographer in the system, and obtain the corneal curvature center coordinates (X wm , Y wm , Z wm ) in the system absolute coordinate system through the corneal reflection method;

[0058] According to the line connecting the pupil center and the corneal curvature center to form an optical axis direction vector That is

[0059]

[0060] Project the optical axis direction vector onto the plane where the air imaging area is located, so as to obtain the intersection point in this plane. Combine the coordinate set of the air imaging area and the position of the intersection point to determine the two-dimensional coordinates of the intersection point in this plane, and mark the intersection point as the fixation point;

[0061] (5) Process the imaging content within the area where the fixation point is located

[0062] According to the size of the air imaging area, divide the display area into dynamic grids, and mark the grid where the fixation point is located and its adjacent areas as high-priority areas;

[0063] The high-priority areas dynamically adjust their contour shapes according to the content form. Specifically:

[0064] If the high-priority area is a text area, recognize the text boundary based on OCR and generate a rectangular contour;

[0065] If the high-priority area is an image area, segment the contour of the corresponding image through an edge detection algorithm (such as the Canny operator);

[0066] If the high-priority area is an interactive control (such as a button, slider, etc.), independently segment and generate an interactive hot zone;

[0067] Highlight the high-priority areas (such as by indicating the color saturation), and at the same time grayscale the non-high-priority areas;

[0068] Perform content magnification on the high-priority areas: centered on the fixation point, magnify the high-priority areas by 2 to 3 times, and specifically use the bilinear interpolation method in the X and Y directions for magnification processing.

[0069] (6) Set the trigger condition for the special display time of the high-priority areas

[0070] Use a time threshold trigger. Set the continuous fixation time threshold to any value within 300 - 500 ms. When the continuous fixation time of the fixation point within a certain neighborhood range of the initial fixation position exceeds the continuous fixation time threshold, trigger the highlight display and / or magnification display.

[0071] The visual adaptation unit continuously tracks the user's vision with the support of the infrared tracking sensor and makes display changes in response to the movement of the fixation point.

[0072] The touch error compensation unit is used to solve the problem that it is difficult for the user to judge the relative distance between the finger and the virtual interface through vision or feeling, resulting in unsmooth operations. It constructs an LSTM time-series trajectory prediction model to predict the touch intention within the next 100 ms and adjust the tactile compensation in advance;

[0073] Among them, based on the time-series data of a large number of historical finger movement trajectories, with the finger trajectory data at a 100 - ms time step, predict the position where the finger will advance within the next 50 ms, and define the input as the three-dimensional coordinates of the finger with time series (x s , y s , z s) Define the output as the predicted distance D between the fingertip controls;

[0074] According to the predicted distance D between the fingertip controls, adaptively adjust the tactile sensation. The smaller the distance, the stronger the tactile sensation. The generation of the tactile sensation uses a phased array ultrasonic transmitter to form an acoustic radiation force field at the fingertip position, and simulates different intensities of vibration friction by adjusting the frequency, so that the user can actually feel whether they have reached the position of controls such as buttons.

[0075] Moreover, due to the limited number of cameras and sensors, and considering the processing efficiency of the system, when the user is using it, there are situations where multiple people are detected simultaneously, and operation conflicts are likely to occur. The exception handling unit constructs a scoring decision model to score the priorities of multiple users. The scoring formula is:

[0076]

[0077] Among them, i represents the number of the person detected entering the scene, d i is the distance from the center of the pupil of the detected user along the line of sight to the fixation point, t zsi represents the continuous fixation time of the corresponding user on the imaging area, β and γ are the distance weight and fixation time weight respectively, and β + γ = 1 and β > 0.5 > γ;

[0078] Refresh the score every certain time period (such as 0.5s). The system responds to the operation of the user with the highest score and blocks the behaviors of other users to avoid interference conflicts.

[0079] It should be noted that the electric sliding lifting component specifically uses an electric slide table.

[0080] Explanation of related terms:

[0081] The Dlib library is a library containing various machine learning algorithms and has good application support in face detection, face recognition, facial key point detection, etc. In facial key point detection, it predicts the positions of various key points on the face (such as the coordinates of parts like eyes, nose, mouth, etc.) by training a specific model.

[0082] The Hough circle transform is an image processing technology based on the Hough transform for detecting circular contours in images. Its core principle is to map the circles in the image space to the parameter space for voting, and determine the center and radius through the local maximum of the accumulator.

[0083] The corneal reflection method projects an infrared light spot (such as Purkinje spot) onto the corneal surface, uses the geometric distribution characteristics of its reflected light spot, combines with the spherical optical characteristics of the cornea, and constructs a mathematical model to inversely deduce the three-dimensional coordinates of the corneal curvature center. The core assumption of this method is that the anterior surface of the cornea is approximately spherical, and there is a fixed geometric relationship between the position of its reflected light spot and the radius of curvature.

[0084] The above are only the preferred embodiments of the present invention, and do not impose any form of limitation on the present invention. Although the present invention has been disclosed above with the preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some changes or modifications to equivalent embodiments by using the above-disclosed technical content within the scope of the technical solution of the present invention. However, as long as it does not depart from the content of the technical solution of the present invention, any simple modification, equivalent change and modification made to the above embodiments according to the technical essence of the present invention still fall within the scope of the technical solution of the present invention.

Claims

1. An intelligent air imaging system based on visual adaption, characterized in that, It includes a hardware support module and an algorithm optimization control module; The hardware support module includes a binocular depth camera, an infrared tracking sensor, a TOF sensor, an electric sliding lifting component carrying a DCT-plate lens group, a corneal topographer, and an array-distributed ultrasonic transmitter; The algorithm optimization control module includes a visual adaptation unit, a touch error compensation unit, and an exception handling unit; among them, the visual adaptation unit obtains the user's pupil center and corneal curvature center, determines the optical axis direction and its fixation point in the air imaging area, generates a high-priority area around the fixation point and performs special display processing; the touch error compensation unit predicts the distance between fingertip controls in a certain future time by constructing an LSTM time-series trajectory prediction model, and drives the ultrasonic transmitter to adjust different frequencies to simulate fingertip touch according to the distance; the exception handling unit constructs a scoring decision model to score the priorities of multiple users and refreshes regularly, and the system responds to the operations of the user with the highest score and blocks the behaviors of other users.

2. The intelligent air imaging system based on visual adaption according to claim 1, wherein The steps of the visual adaptation unit performing special display processing based on the user's height and the position of the fixation point include: Step 1: Obtain the user's height based on the binocular depth camera, and the electric sliding lifting component drives the DCT-plate lens group to adjust the height to adapt to the height; Step 2: Detect the face area based on the face detection algorithm and determine the eye area, and then locate the two-dimensional pixel coordinates of the pupil center in the image; Step 3: First convert the two-dimensional pixel coordinates into the camera's three-dimensional coordinates, and then convert the camera's three-dimensional coordinates into the three-dimensional pupil center coordinates in the system's absolute coordinate system; Step 4: Obtain the corneal curvature center coordinates and construct an optical axis direction vector in combination with the pupil center coordinates to determine the intersection point of the vector and the air imaging area, that is, the fixation point; Step 5: Determine different forms of high-priority areas according to the content form in the neighborhood where the fixation point is located, perform highlighting display and / or magnification processing on the high-priority areas, and perform grayscale display on the non-high-priority areas at the same time; Step 6: Set a trigger condition for the special display method of the high-priority area, and continuously track the user's vision and perform display changes according to the movement of the fixation point.

3. An intelligent air imaging system based on visual adaption according to claim 1, characterized in that, When constructing the LSTM time-series trajectory prediction model, use the finger trajectory data with a time step of 100 ms to predict the position where the finger advances in the next 50 ms. Define the input as the three-dimensional coordinates of the finger with time series, and define the output as the predicted distance between fingertip controls.

4. An intelligent air imaging system based on visual adaptation according to claim 1, characterized in that, The priority scoring formula is Among them, i represents the number of the person detected to enter the scene, and d i is the distance from the detected pupil center of the user along the line of sight to the fixation point, and t zsi represents the continuous fixation time of the corresponding user on the imaging area. β and γ are the distance weight and the fixation time weight respectively, and β + γ = 1 and β > 0.5 > γ.

5. The intelligent air imaging system based on vision adaptation according to claim 2, wherein When adjusting the height of the DCT-plate lens group in Step 1, its target height is set to H = 0.6h + 300, where H is the target height and h is the user's height.

6. The intelligent air imaging system based on vision adaptation according to claim 2, characterized in that, The process of converting the pixel coordinates of the pupil into the three-dimensional pupil center coordinates in Step 3 includes correcting the depth information and coordinate transformation: Construct a three-dimensional camera coordinate system and calculate the parallax depth information Z of the pupil in the camera coordinate system c , that is where f represents the focal length of the camera, b represents the baseline distance of the binocular camera, and d represents the horizontal displacement difference of the pixel coordinates of the pupil center in the left and right images of the binocular camera; Obtaining depth information directly using a TOF sensor where Δt is the phase difference between the emitted light and the detected reflected light; Calculate the finally corrected depth information Z final = (1 - α)·Z c + α·Z TOF ; where α is a compensation weight coefficient, and Subsequently, obtain the camera's three-dimensional coordinates according to the corrected depth information, and perform coordinate transformation according to the calibrated external parameter matrix to obtain the system's absolute coordinates.

7. The intelligent air imaging system based on vision adaptation according to claim 2, characterized in that The method for determining the contour form of the high-priority area includes: The text area generates a rectangular contour by OCR to recognize the text boundary; The image area is segmented into contours by Canny edge detection; The interactive control is independently segmented to generate an interactive hot zone.

8. An intelligent air imaging system based on visual adaptability according to claim 2, characterized in that, The triggering condition for the special display of the high-priority area is that when the fixation point continuously fixates within a certain neighborhood range of the initial fixation position for a duration exceeding the fixation duration threshold, a highlighted display and / or a magnified display are triggered.