Region-limited virtual touch interaction method and related equipment

By scanning the projection plane boundary and dividing the area, combined with depth perception technology, high-precision user gesture recognition in open and irregular spaces is achieved, solving the problem of low interaction accuracy in the existing technology, and providing a real-time and smooth user experience.

CN120295487AInactive Publication Date: 2025-07-11셴젠 동루 테크놀로지 컴퍼니 리미티드
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510773587.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-11
Publication Date
2025-07-11
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing virtual touch interaction technology is difficult to accurately locate user gestures in an open and irregular space, resulting in low interaction accuracy.

Method used

By combining projection technology and depth perception technology, the projection planes defined by the region are scanned and divided into areas, and the depth camera is used to collect three-dimensional feature points of the user's hand, perform collision detection and visual feedback rendering, so as to achieve accurate capture and response to user's gestures.

Benefits of technology

High-precision user gesture recognition and interaction are achieved in an open and irregular space, improving the freedom and flexibility of interaction, providing real-time, smooth user experience and enhanced interactive reality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120295487A_ABST
    Figure CN120295487A_ABST
Patent Text Reader

Abstract

The invention relates to a region-limited virtual touch interaction method and related equipment, and the method comprises the following steps: carrying out the boundary scanning of a region-limited projection plane, and obtaining a projection region contour coordinate set; performing region division on a projection plane based on the projection region contour coordinate set to obtain interaction region grid data; performing three-dimensional feature point collection on the hand of the user through a depth camera to obtain a hand key point space coordinate sequence; when the hand key point space coordinate sequence is within a depth interaction threshold range of interaction area grid data, collision detection is carried out on the hand key point space coordinate sequence based on the interaction area grid data, and a touch event judgment result is obtained; and performing visual feedback rendering on the interaction area grid data based on the touch event judgment result to obtain interaction state display data, thereby solving the technical problem of low interaction precision caused by difficulty in accurately positioning a gesture of a user in an open and irregular space in an existing solution.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of touch interaction technology, and particularly to a virtual touch interaction method with area limitation and related devices. Background Art

[0002] In current human-computer interaction technologies, the demand for more natural and intuitive interaction methods is increasing. Especially in the fields of public information display, smart home control, and interactive entertainment, users expect to interact with digital content in a way closer to daily life. However, traditional input devices such as touchscreens or mouse keyboards are limited by the existence of physical interfaces, which not only restricts the operation flexibility of users but also makes it difficult to meet the needs of multi-user simultaneous interaction. Therefore, how to break through these limitations and achieve an efficient interaction method without physical contact has become an important research direction.

[0003] Although some virtual interaction systems based on gesture recognition have emerged in the market, they generally have some limitations. For example, most systems require the support of specific hardware, such as wearable devices or dedicated sensors, which increases the deployment cost and technical complexity. In addition, existing solutions often have difficulty accurately positioning users' gestures in an open and irregular space, resulting in low interaction accuracy and poor user experience. These problems have seriously hindered the wide application of virtual touch interaction technology in more scenarios.

[0004] In response to the above challenges, this research proposes an innovative virtual touch interaction method with area limitation. By combining projection technology and depth perception technology, this method achieves precise capture and response to users' gestures without the need for additional wearable devices. More importantly, this method can create a virtual touch interface within an arbitrarily shaped area, greatly improving the freedom and flexibility of interaction. By solving the problems existing in the prior art, this research provides a new idea and implementation solution for future human-computer interaction. Summary of the Invention

[0005] The main object of the present invention is to provide a virtual touch interaction method with area limitation and related devices, which solves the technical problem that existing solutions often have difficulty accurately positioning users' gestures in an open and irregular space, resulting in low interaction accuracy.

[0006] To achieve the above object, the present invention provides a virtual touch interaction method with area limitation, including the following steps: Perform boundary scanning on the projection plane with area limitation to obtain a set of contour coordinates of the projection area; Based on the set of contour coordinates of the projection area, perform area division on the projection plane to obtain interactive area grid data; Collect three-dimensional feature points of the user's hand through a depth camera to obtain a sequence of spatial coordinates of hand key points; When the sequence of spatial coordinates of the hand key points is within the depth interaction threshold range of the interactive area grid data, perform collision detection on the sequence of spatial coordinates of the hand key points based on the interactive area grid data to obtain a touch event determination result; Perform visual feedback rendering on the interactive area grid data based on the touch event determination result to obtain interactive state display data.

[0007] Further, perform boundary scanning on the projection plane defined by the region to obtain a set of contour coordinates of the projection region, including: Emit a raster light spot array to the projection plane defined by the region through a structured light projector, and collect the reflected light intensity distribution of the raster light spot array through a binocular depth camera to obtain depth scattering data of the projection plane; Extract edge contours and perform spatial coordinate mapping on the depth scattering data to obtain a point cloud of the projection region boundary, and perform spatial filtering on the point cloud of the projection region boundary to obtain a boundary feature sequence; Perform polygon fitting and vertex positioning based on the boundary feature sequence to obtain a set of contour coordinates of the projection region; wherein, the set of contour coordinates of the projection region includes the spatial position coordinates of four vertices of the projection plane and the normal vector of the projection plane.

[0008] Further, perform region division on the projection plane based on the set of contour coordinates of the projection region to obtain interactive area grid data, including: Perform homogeneous coordinate transformation on the set of contour coordinates of the projection region to obtain a plane projection transformation matrix, and perform singular value decomposition on the plane projection transformation matrix to obtain main direction feature data of the projection region; Arrange hexagonal grid seed points on the projection plane based on the main direction feature data of the projection region to obtain initial grid layout parameters, and perform Thiessen polygon division on the initial grid layout parameters to obtain regular grid division data; Perform spatial coordinate remapping on the regular grid division data to obtain a set of grid depth parameters, and perform surface interpolation reconstruction on the set of grid depth parameters to obtain grid cell depth attribute data; Perform grid topological relationship construction based on the grid cell depth attribute data to obtain interactive area grid data, wherein the interactive area grid data includes the boundary coordinates of multiple regular hexagonal grid cells, the depth values of the grid cells, and the topological relationships between the grid cells.

[0009] Further, the collecting three-dimensional feature points of the user's hand through a depth camera to obtain a sequence of spatial coordinates of hand key points includes: Collect multi-view depth images of the user's hand through a depth camera to obtain a sequence of depth images of the hand region, and perform structured light coding analysis on the sequence of depth images of the hand region to obtain hand point cloud data; Perform bone joint positioning on the hand point cloud data to obtain a set of candidate hand joint points, and perform motion constraint optimization on the set of candidate hand joint points to obtain hand joint chain structure data, where the hand joint chain structure data includes the connection relationship between joint points, joint mobility parameters, and bone length ratios; Extract gesture features from the hand joint chain structure data to obtain a sequence of hand pose features, and perform temporal smoothing processing on the sequence of hand pose features to obtain hand motion trajectory data; Perform spatial coordinate mapping on the hand motion trajectory data to obtain a sequence of spatial coordinates of hand key points, where the sequence of spatial coordinates of hand key points includes the temporal position data of finger joint points, the center point of the palm, and the reference point of the wrist.

[0010] Further, performing collision detection on the sequence of spatial coordinates of hand key points based on the interactive region grid data to obtain a touch event determination result, including: Perform spatial projection transformation on the sequence of spatial coordinates of hand key points to obtain a set of projected plane mapping points, and perform grid index matching on the set of projected plane mapping points to obtain candidate collision grid data; Perform depth threshold layering on the candidate collision grid data to obtain multiple layers of collision detection domains, and perform spatial overlap analysis on the multiple layers of collision detection domains to obtain collision hierarchy relationship data; Perform temporal collision tracking on the sequence of spatial coordinates of hand key points based on the collision hierarchy relationship data to obtain a sequence of collision events, and perform state transition analysis on the sequence of collision events to obtain collision state transition data; Perform depth gradient calculation on the collision state transition data to obtain a set of depth interaction parameters, and perform touch feature extraction on the set of depth interaction parameters to obtain a group of touch feature vectors; Perform comprehensive determination of touch events based on the group of touch feature vectors to obtain a touch event determination result, where the touch event determination result includes the grid cell index where the collision occurs, the collision duration, and the relative depth value of the collision point.

[0011] Further, the performing depth gradient calculation on the collision state transition data to obtain a set of depth interaction parameters includes: Perform temporal depth sampling on the collision state transition data to obtain a depth change sequence, and perform differential operations on the depth change sequence to obtain depth gradient data, where the depth gradient data includes a depth change rate, a gradient direction vector, and local extreme points; Perform depth distribution statistics based on the depth gradient data to obtain a depth distribution feature set, and perform outlier filtering on the depth distribution feature set to obtain depth statistical parameters; Perform spatial curvature analysis on the depth statistical parameters to obtain curvature feature data, and perform principal curvature extraction on the curvature feature data to obtain surface morphology parameters; Perform contact area estimation based on the surface morphology parameters to obtain contact region parameters, and perform pressure distribution calculation on the contact region parameters to obtain pressure distribution characteristics; Perform feature fusion processing on the pressure distribution characteristics to obtain a depth interaction parameter set, where the depth interaction parameter set includes the contact area size, the pressure distribution pattern, and the depth change trend.

[0012] Furthermore, visually feedback rendering the interactive area grid data based on the touch event determination result to obtain interactive state display data, including: Extract the touch area boundary from the touch event determination result to obtain touch activation range data, and perform light intensity distribution calculation on the touch activation range data to obtain grid brightness mapping data; Perform color space conversion on the touch area corresponding to the projection plane based on the grid brightness mapping data to obtain a grid color parameter set, and perform gradient interpolation processing on the grid color parameter set to obtain color transition data; Perform dynamic ripple superposition on the color transition data to obtain ripple diffusion data, and perform spatio-temporal evolution calculation on the ripple diffusion data to obtain a ripple animation sequence; Perform visual effect synthesis based on the ripple animation sequence to obtain interactive state display data, where the interactive state display data includes the boundary brightness value, color gradient parameters, and dynamic ripple effect parameters of the touch activation area.

[0013] The present invention also provides a region-limited virtual touch interaction device, including: A scanning module for performing boundary scanning on a region-limited projection plane to obtain a set of projection area contour coordinates; A partitioning module for partitioning the projection plane based on the set of projection area contour coordinates to obtain interactive area grid data; An acquisition module for collecting three-dimensional feature points of the user's hand through a depth camera to obtain a sequence of spatial coordinates of hand key points; A detection module, configured to perform collision detection on the hand key-point spatial coordinate sequence based on the interactive area grid data when the hand key-point spatial coordinate sequence is within the depth interaction threshold range of the interactive area grid data, so as to obtain a touch event determination result; A rendering module, configured to perform visual feedback rendering on the interactive area grid data based on the touch event determination result to obtain interactive state display data.

[0014] The present invention further provides a computer device, including a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, the steps of the method described in any one of the above are implemented.

[0015] The present invention further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method described in any one of the above are implemented.

[0016] A virtual touch interaction method with area limitation provided by the present invention includes the following steps: performing boundary scanning on a projected plane with area limitation to obtain a set of contour coordinates of the projected area; dividing the projected plane based on the set of contour coordinates of the projected area to obtain interactive area grid data; collecting three-dimensional feature points of a user's hand through a depth camera to obtain a hand key-point spatial coordinate sequence; when the hand key-point spatial coordinate sequence is within the depth interaction threshold range of the interactive area grid data, performing collision detection on the hand key-point spatial coordinate sequence based on the interactive area grid data to obtain a touch event determination result; performing visual feedback rendering on the interactive area grid data based on the touch event determination result to obtain interactive state display data, which solves the technical problem that existing solutions often have difficulty in accurately positioning a user's gesture in an open and irregular space, resulting in low interaction accuracy, and realizes that when the detected hand key-point spatial coordinate sequence is within the depth interaction threshold range of the interactive area grid data, the system can perform collision detection based on the interactive area grid data and immediately obtain a touch event determination result. This process ensures that the user can obtain a real-time and smooth interaction experience, enhancing the realism and immersion of the interaction. Description of the Drawings

[0017] Figure 1 is a schematic diagram of the steps of a virtual touch interaction method with area limitation according to an embodiment of the present invention; Figure 2 is a structural block diagram of a virtual touch interaction platform with area limitation according to an embodiment of the present invention; Figure 3 is a schematic structural block diagram of a computer device according to an embodiment of the present invention.

[0018] The implementation of the object, functional features, and advantages of the present invention will be further described with reference to the embodiments and the accompanying drawings. Detailed implementation manners

[0019] In order to make the object, technical solutions, and advantages of the present invention more clear and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0020] As Figure 1 shown, Figure 1 is a schematic diagram of the steps of a virtual touch interaction method with area limitation in an embodiment of the present invention; An embodiment of the present invention provides a virtual touch interaction method with area limitation, including the following steps: Step S1, perform boundary scanning on the projection plane with area limitation to obtain a set of contour coordinates of the projection area.

[0021] Specifically, performing boundary scanning on the projection plane with area limitation to obtain a set of contour coordinates of the projection area, the core of this process lies in accurately identifying and digitally processing the physical boundary of the projection plane through technical means. First of all, the projection plane can be a surface of any shape, such as a wall, a desktop, or even a curved screen, and the task of boundary scanning is to determine the actual range of the projection plane and convert it into a set of contour coordinates of the projection area for subsequent processing. To achieve this, high-precision cameras or laser scanning devices can be used, combined with image processing algorithms, to detect the boundary of the projection plane. Specifically, when the projection device projects an image onto the target plane, the scanning device will capture the edge information of the projection area and extract the key point coordinates of the boundary through edge detection algorithms, and these coordinates ultimately form a set of contour coordinates of the projection area. For example, in a smart home control scenario, assume that the user hopes to create a virtual touch interface on the wall to adjust the light brightness or control household appliances. At this time, the system first needs to clarify the specific range of the projection area on the wall, because only within this area can effective gesture interactions be realized. By performing boundary scanning on the projection plane on the wall, the system can obtain a set of contour coordinates of the projection area on the wall. These data not only define the physical boundary of the interaction but also provide a basis for subsequent area division. For example, if the projection area is a rectangle, the scanning result will include the coordinates of the four vertices of the rectangle; if it is an irregular shape, a series of discrete coordinate points describing the boundary will be generated. The accuracy of these coordinate sets directly affects the quality of the generation of subsequent interaction area grid data, thereby ensuring that the user's gesture operations can be recognized and responded to within the correct area. Therefore, this step plays a crucial foundational role in the entire method. It not only ensures the accurate definition of the interaction area but also lays a reliable foundation for subsequent gesture capture and touch event determination.

[0022] Step S2: Based on the set of projection area contour coordinates, divide the projection plane into regions to obtain interactive area grid data.

[0023] Specifically, dividing the projection plane based on the set of projection area contour coordinates to obtain interactive area grid data. The key to this process is to convert the physical boundary information of the projection area into a virtual grid structure for interactive operations. First, the set of projection area contour coordinates has clearly defined the boundary range of the projection plane. Next, it is necessary to perform a logical regional division of the projection plane according to these coordinates. Specifically, the system will use an algorithm to divide the projection area into several regular or irregular small regions, and each small region is assigned a unique identifier and spatial attributes, thus forming the interactive area grid data. The purpose of this grid processing is to provide an accurate spatial reference for subsequent gesture recognition and touch event determination, so that the user's hand movements can be accurately located in specific grid cells. For example, in a smart home control scenario, assume that a projection device projects a rectangular area on the wall and obtains the contour coordinate set of this rectangle through boundary scanning. Next, the system will divide the rectangular area into multiple uniform small grids according to these coordinates. For example, it is divided into a 3×3 nine-square grid layout, and each small grid corresponds to a functional area, such as light adjustment, temperature control, or home appliance switching. These grid data not only define the specific location of each functional area but also provide a basis for subsequent gesture operations. For example, when the user moves their hand into a certain grid area and makes a specific gesture, the system can quickly identify the grid position where the gesture is located and trigger the corresponding functional response. Therefore, this step plays a connecting role in the whole method. It not only converts the physical boundary of the projection area into an operable virtual grid but also lays a spatial foundation for subsequent gesture capture and touch event determination, thus ensuring the efficiency and accuracy of the interaction process.

[0024] Step S3: Collect three-dimensional feature points of the user's hand through a depth camera to obtain a sequence of spatial coordinates of hand key points.

[0025] Specifically, a depth camera is used to collect three-dimensional feature points of the user's hand, obtaining a sequence of spatial coordinates of hand key points. The core of this process lies in leveraging the high-precision perception ability of the depth camera to capture the dynamic changes of the user's hand in three-dimensional space and convert them into a series of processable data. Specifically, the depth camera emits infrared light or other forms of signals, calculates the distance between the target object and the camera by receiving the reflected signals, and thus generates a three-dimensional image containing depth information. On this basis, the system combines computer vision algorithms to identify and locate the key parts of the hand, such as fingertips, knuckles, palm center, etc. The three-dimensional positions of these parts are extracted and recorded in the form of coordinates, ultimately forming a sequence of spatial coordinates of hand key points. For example, in the smart home control scenario, assume the user stands in front of the projection area and hopes to adjust the light brightness or control the home appliance switch through gestures. At this time, the depth camera starts to work and captures the user's hand movements in real time. For example, when the user extends a finger and points to a certain grid area, the depth camera will collect the spatial position of the fingertip and convert it into a set of three-dimensional coordinates; when the user makes a swiping gesture, the camera will continuously record the trajectory of the finger movement, generating a dynamic sequence of spatial coordinates of hand key points. These data can not only reflect the specific form of the user's gesture but also accurately describe the spatial position where the gesture occurs, providing the necessary input for subsequent collision detection and touch event determination. Therefore, this step plays a crucial role in the entire method. It not only achieves the precise capture of hand movements but also lays a technical foundation for the natural interaction between the user and the virtual interface, enabling the user to complete complex operation tasks without touching any physical devices.

[0026] Step S4, when the sequence of spatial coordinates of hand key points is within the depth interaction threshold range of the interactive area grid data, perform collision detection on the sequence of spatial coordinates of hand key points based on the interactive area grid data to obtain a touch event determination result.

[0027] Specifically, when the spatial coordinate sequence of the hand key points is within the depth interaction threshold range of the interactive area grid data, collision detection is performed on the spatial coordinate sequence of the hand key points based on the interactive area grid data to obtain a touch event determination result. This process aims to determine whether the user's gesture triggers a specific interactive operation through precise spatial analysis. First, the system has obtained the spatial coordinate sequence of the key points of the user's hand and defined the interactive area grid data on the projection plane and its corresponding depth interaction threshold range. The depth interaction threshold here refers to the maximum distance range within which hand movements can be recognized to ensure the accuracy of gesture recognition. Once a certain hand key point of the user enters this threshold range, the system will start collision detection, that is, check whether the spatial coordinates of the hand key point intersect or overlap with a certain grid cell. If an intersection is detected, it indicates that the user's gesture intends to perform an operation, and the system takes this as a touch event and generates a corresponding determination result. For example, in the smart home control scenario, assume that the user hopes to adjust the light brightness through gestures. The user stands in front of the pre-set projection area, and the movements of their hand are captured by a depth camera and converted into a spatial coordinate sequence of hand key points. When the user's finger moves above a specific function area (such as the dimming area), the system first determines whether the key point coordinates of the finger are within the depth interaction threshold range of the grid data of this function area. If the condition is met, the system will further perform collision detection to confirm whether the finger "touches" the grid cell of the dimming area. Once a collision is confirmed, the system considers that the user has issued a dimming instruction and adjusts the light brightness according to the specific position and movement direction of the finger. This not only realizes contactless human-computer interaction but also makes the operation more intuitive and convenient, greatly improving the user experience. Therefore, through this precise collision detection mechanism, the system can accurately understand the user's intention and make a timely response.

[0028] Step S5, perform visual feedback rendering on the interactive area grid data based on the touch event determination result to obtain interactive state display data.

[0029] Specifically, based on the touch event determination result, visual feedback rendering is performed on the interactive area grid data to obtain interactive state display data. The core of this process lies in enabling the user to clearly perceive that their gesture operation has been recognized and responded to by the system through real-time visual feedback. Specifically, when the system generates a touch event determination result based on collision detection, it will immediately perform dynamic updates and rendering on the interactive area grid data to visually display the current interactive state to the user. For example, the system can represent that a specific area has been triggered by changing the color, brightness of specific grids within the projection area or adding dynamic effects. This visual feedback not only enhances the user's immersion but also helps the user confirm whether their operation is successful, thereby reducing the possibility of misoperation. For example, in a smart home control scenario, assume that when the user adjusts the light brightness through gestures and their finger moves to a certain grid cell in the dimming area, and the system has determined this operation as a valid touch event. At this time, the system will perform visual feedback rendering on this grid cell based on the touch event determination result. For example, it will change the color of the grid from light gray to bright blue, or add an animated effect of a gradually expanding light circle around the grid to indicate that this area has been activated. At the same time, the system will also dynamically adjust the light brightness according to the specific position of the finger and display the current brightness state through numerical changes or a progress bar on the projection interface. In this way, the user can not only feel the effect of the operation through the actual change of the light but also intuitively understand the correspondence between their gesture and the system response through the visual feedback on the projection interface. Therefore, this step plays a crucial role in the entire method. It not only improves the intuitiveness of the interaction and the user experience but also provides an efficient and natural communication method for contactless human-computer interaction.

[0030] In a specific embodiment, the boundary scanning of the projection plane with area limitation is performed to obtain a set of projection area contour coordinates, including: The structured light projector emits a grid spot array towards the projection plane with area limitation, and the binocular depth camera collects the reflected light intensity distribution of the grid spot array to obtain the depth scattering data of the projection plane; Edge contour extraction and spatial coordinate mapping are performed on the depth scattering data to obtain the projection area boundary point cloud, and spatial filtering is performed on the projection area boundary point cloud to obtain the boundary feature sequence; Based on the boundary feature sequence, polygon fitting and vertex positioning are performed to obtain a set of projection area contour coordinates; wherein, the set of projection area contour coordinates includes the spatial position coordinates of the four vertices of the projection plane and the normal vector of the projection plane.

[0031] Specifically, the process of performing boundary scanning on the projection plane defined by the region to obtain the contour coordinate set of the projection region is achieved through a series of precise technical means. First of all, this process relies on a structured light projector to emit a grid of light spot arrays onto the projection plane defined by the region, and uses a binocular depth camera to collect the reflected light intensity distribution of these grid light spot arrays, thereby obtaining the depth scattering data of the projection plane. Specifically, in the application scenario of smart home control, when a user hopes to create a virtual touch interface on a wall, the structured light projector projects a series of regularly arranged grid light spots onto the wall. After these light spots contact the wall surface, they will generate a specific reflection pattern, and the binocular depth camera is responsible for capturing the light intensity distribution of these reflected light spots, and then generating a data set containing depth information. Since the reflection intensity and position of each light spot carry information about the distance between it and the projection plane, the three-dimensional shape of the projection plane can be initially constructed by analyzing these data. Next, edge contour extraction and spatial coordinate mapping are performed on the depth scattering data to obtain the boundary point cloud of the projection region, and spatial filtering is performed on the boundary point cloud of the projection region to obtain the boundary feature sequence. This means that the system needs to identify the boundary information of the projection region from the original depth scattering data and convert this information into a set of coordinates in three-dimensional space. In this process, the algorithm will first identify those data points representing edge characteristics because they are usually located at the boundary of the projection region. For example, in the above smart home example, when the structured light shines on the wall and is captured by the binocular depth camera, the algorithm can accurately locate the edge of the wall and the position of any possible irregularly shaped parts, such as the position of sockets or decorations. Subsequently, by performing spatial coordinate mapping on these boundary points, the system can establish a point cloud model describing the boundary of the projection region. However, the boundary points directly extracted from the original data may contain noise or redundant information, so spatial filtering needs to be performed on them to remove unnecessary details and retain key boundary features, forming a more concise and accurate boundary feature sequence. Finally, polygon fitting and vertex positioning are performed based on the boundary feature sequence to obtain the contour coordinate set of the projection region. This step involves using geometric principles to approximate the actual boundary of the projection region, simplifying it into a polygon model, and determining the exact positions of each vertex. For the smart home control scenario, this means that the system not only needs to identify the specific area on the wall for creating the virtual touch interface, but also needs to accurately locate the spatial position coordinates and their normal vectors of the four vertices of this area. Through polygon fitting technology, the system can infer the polygon shape closest to the real boundary from the boundary feature sequence and calculate the exact coordinates of each vertex according to this shape. In addition, the determination of the normal vector helps to understand the directionality of the projection plane, which is crucial for the accuracy of subsequent gesture interactions.In summary, through this series of steps, the system can not only accurately define the physical boundaries of the projection area, but also provide a solid foundation for its subsequent functional division, gesture recognition and visual feedback. This process ensures that even in complex or non-standard environments, an efficient and accurate human-computer interaction experience can be achieved.

[0032] In a specific embodiment, the region division of the projection plane based on the projection area contour coordinate set to obtain interactive area grid data includes: Performing homogeneous coordinate transformation on the projection area contour coordinate set to obtain a plane projection transformation matrix, and performing singular value decomposition on the plane projection transformation matrix to obtain main direction feature data of the projection area; Arranging hexagonal grid seed points on the projection plane based on the main direction feature data of the projection area to obtain grid initial layout parameters, and performing Thiessen polygon segmentation on the grid initial layout parameters to obtain regular grid segmentation data; Remapping the spatial coordinates of the regular grid subdivision data to obtain a grid depth parameter set, and reconstructing the grid depth parameter set by surface interpolation to obtain grid unit depth attribute data; A grid topological relationship is constructed based on the grid unit depth attribute data to obtain interactive area grid data, wherein the interactive area grid data includes boundary coordinates of multiple regular hexagonal grid units, depth values ​​of the grid units, and topological relationships between the grid units.

[0033] Specifically, in the process of dividing the projection plane based on the set of contour coordinates of the projection area to obtain the interactive area grid data, it is first necessary to perform homogeneous coordinate transformation on the set of contour coordinates of the projection area to obtain a planar projection transformation matrix, and perform singular value decomposition on this matrix to extract the principal direction feature data of the projection area. The core of this process lies in transforming the original three-dimensional space information into a form suitable for subsequent processing through mathematical transformation. In the application scenario of smart home control, assume that the user hopes to create a virtual touch interface on an irregularly shaped wall area. The system will first use the homogeneous coordinate transformation method to convert the set of contour coordinates of the projection area obtained in the previous step into a representation in a unified coordinate system, and then generate a planar projection transformation matrix. Then, by performing singular value decomposition on this matrix, the system can identify the most representative geometric feature directions within the projection area, and these feature directions are crucial for understanding the layout of the entire area. For example, if there are decorations or sockets on the wall that affect the regularity of the rectangle, singular value decomposition can help the system determine the main axis that best reflects the actual shape and direction. Next, based on the principal direction feature data of the projection area, hexagon grid seed points are arranged on the projection plane to obtain the initial grid layout parameters, and the initial grid layout parameters are segmented by Thiessen polygons to obtain regular grid subdivision data. Here, the system uses the direction information obtained from singular value decomposition to guide the arrangement of hexagon grid seed points. Since hexagons have good filling characteristics and the advantage of uniform distribution, they are widely used in many application scenarios. In the above smart home example, once the main directions are determined, the system can evenly arrange the seed points of the hexagon grid on the projection plane. Then, through Thiessen polygon segmentation technology, the relationship between these seed points can be further refined to form a set of regular grid subdivision data. This method not only ensures that there is no overlap between each hexagon grid cell, but also guarantees the maximization of the entire grid coverage, enabling the user's gesture operations to be accurately captured at any position. Subsequently, spatial coordinate remapping is performed on the regular grid subdivision data to obtain a set of grid depth parameters, and surface interpolation reconstruction is performed on the set of grid depth parameters to obtain the depth attribute data of the grid cells. At this stage, the system needs to remap the previously obtained two-dimensional grid layout back into the three-dimensional space to consider the depth changes at different positions. For example, in a smart home environment, the wall surface may not be completely flat, with some uneven parts. By performing spatial coordinate remapping on the regular grid subdivision data, the system can obtain a set of depth parameters corresponding to each grid cell, and these parameters describe the height difference of the grid cell relative to the reference plane.After that, using the surface interpolation reconstruction technique, more accurate grid cell depth attribute data can be deduced based on these depth parameter sets. This step is particularly crucial for improving the accuracy of gesture recognition because it allows the system to take into account the subtle height variations on the wall surface, thereby more accurately locating the position of the user's gesture. Finally, based on the grid cell depth attribute data, a grid topological relationship is constructed to obtain the interactive area grid data. Among them, the interactive area grid data includes the boundary coordinates of multiple regular hexagonal grid cells, the depth values of the grid cells, and the topological relationships between the grid cells. This means that the system not only needs to define the spatial position and depth attributes of each hexagonal grid cell but also clarify the way they are interconnected. In the smart home control scenario, such a grid structure provides a highly flexible and intuitive interactive interface for users. For example, when a user attempts to adjust the light brightness through gestures, the system can quickly determine the user's intention based on the position and depth information of the specific hexagonal grid cell where the hand movement is located and make corresponding responses. At the same time, the clear topological relationships between the grid cells contribute to improving the efficiency and accuracy of the gesture recognition algorithm because they provide a structured framework that makes gesture tracking and event determination more direct and reliable. In summary, through this series of complex and delicate steps, the system can transform an arbitrarily shaped projection plane into an efficient and accurate virtual touch interface, greatly enhancing the user experience and expanding the possibilities of human-computer interaction.

[0034] In a specific embodiment, the hexagonal grid seed points are arranged on the projection plane based on the projection area main direction feature data to obtain the initial grid layout parameters, including: Perform eigenvector decomposition on the projection area main direction feature data to obtain the main direction component data, and construct an orthogonal basis for the main direction component data to obtain a spatial reference coordinate system; Estimate the grid density of the projection area based on the spatial reference coordinate system to obtain the grid scale parameter, and generate hexagonal cells for the grid scale parameter to obtain a basic grid template; Optimize the spatial arrangement of the basic grid template to obtain a seed point distribution sequence, and perform boundary constraint processing on the seed point distribution sequence to obtain boundary adaptation parameters; Perform local adjustment on the seed points based on the boundary adaptation parameters to obtain a set of grid node coordinates, and construct the connection relationship for the set of grid node coordinates to obtain a grid topological structure; Perform parametric description on the grid topological structure to obtain the initial grid layout parameters, where the initial grid layout parameters include the geometric dimensions of the grid cells, the spatial positions of the grid nodes, and the constraint conditions of the grid boundaries.

[0035] Specifically, the process of arranging hexagonal grid seed points on the projection plane based on the main direction feature data of the projection area to obtain the initial grid layout parameters first involves decomposing the main direction feature data of the projection area into eigenvectors to obtain the main direction component data, and constructing an orthonormal basis for these data to obtain a spatial reference coordinate system. In the application scenario of smart home control, assume that the user hopes to create a virtual touch interface on an irregularly shaped wall. The system first needs to clarify the main direction and features of the wall surface. Through eigenvector decomposition, the system can extract the main direction component data from the previously obtained main direction feature data of the projection area that best represents the geometric characteristics of the area. Then, using these component data to construct a set of orthonormal bases to form a new spatial reference coordinate system. This step is crucial for the subsequent steps because it provides a unified coordinate framework based on the actual geometric shape for the entire grid division process. Estimate the grid density of the projection area based on the spatial reference coordinate system to obtain the grid scale parameters, and generate hexagonal cells for the grid scale parameters to obtain the basic grid template. At this stage, the system needs to determine the appropriate grid density according to the specific size and shape of the projection area. For example, in a smart home environment, if the user hopes to achieve fine gesture control on a large wall, a denser grid of cells may be required to improve the accuracy of gesture recognition; on the contrary, if the wall is small or very high precision is not required, a sparser grid layout can be selected. By carefully analyzing the projection area and combining the actual needs of the user, the system can estimate the most suitable grid scale parameters. Subsequently, use these parameters to generate a series of standard hexagonal cells to form the basic grid template. Hexagons are selected because of their good filling properties. They can not only achieve seamless coverage on the plane but also effectively reduce errors caused by shape. Optimize the spatial arrangement of the basic grid template to obtain the seed point distribution sequence, and perform boundary constraint processing on the seed point distribution sequence to obtain the boundary adaptation parameters. In this process, the system regards each hexagonal cell in the basic grid template as a "seed point" and optimizes its arrangement to ensure that they are evenly distributed throughout the projection area. To achieve this, the algorithm will iteratively adjust the positions of the seed points multiple times until the optimal spatial arrangement scheme is found. At the same time, considering that the projection area may have complex boundary conditions (such as corners, concave and convex parts, etc.), the system also needs to apply boundary constraint processing to the seed point distribution sequence. For example, in the case of smart home control, if there are sockets or decorations on the wall, the system must ensure that the grid does not cross these obstacles but bypasses them for a reasonable arrangement. In this way, a set of boundary adaptation parameters that conform to the actual situation can be obtained, ensuring that the grid structure can not only make full use of the available space but also avoid conflicts with existing physical obstacles.Based on the boundary adaptation parameters, the seed points are locally adjusted to obtain a set of grid node coordinates, and the connection relationships of the set of grid node coordinates are constructed to obtain a grid topology structure. After obtaining the boundary adaptation parameters, the system will further fine-tune the positions of each seed point, especially those nodes close to the boundary, to ensure that their positions meet both geometric requirements and adapt to the actual environmental constraints. For example, in a smart home application, for special structures on the wall, the system may appropriately move some grid nodes to avoid sockets or other obstacles. Once the positions of all nodes are precisely adjusted, the system will start to construct the connection relationships between grid nodes to form a complete grid topology structure. This step not only defines the specific shape of each grid cell but also clarifies how they are interconnected, providing a necessary spatial framework for subsequent gesture interactions. Finally, the grid topology structure is parametrically described to obtain the initial grid layout parameters, where the initial grid layout parameters include the geometric dimensions of the grid cells, the spatial positions of the grid nodes, and the constraint conditions of the grid boundaries. This means that the system needs to integrate all the above information into a format that is easy to understand and use. For example, in a smart home control scenario, the finally generated initial grid layout parameters can not only clearly show the size and shape of each hexagonal grid cell but also accurately mark the exact positions of each grid node and their relationships with the surrounding environment. In addition, it includes detailed constraint conditions on the grid boundaries, which are crucial for ensuring the accuracy and reliability of gesture operations. Through this series of complex but orderly steps, the system can transform an arbitrarily shaped projection plane into a structured and highly operable virtual touch interface, greatly enhancing the user experience and laying a solid foundation for more complex human-computer interaction technologies in the future. This method not only demonstrates the powerful potential of the technology but also reflects its flexibility and practicality in actual applications.

[0036] In a specific embodiment, the three-dimensional feature points of the user's hand are collected by a depth camera to obtain a sequence of spatial coordinates of hand key points, including: The depth images of the user's hand from multiple perspectives are collected by a depth camera to obtain a sequence of depth images of the hand region, and the structured light coding of the sequence of depth images of the hand region is analyzed to obtain hand point cloud data; The bone joints of the hand point cloud data are located to obtain a set of hand joint candidate points, and the set of hand joint candidate points is optimized by motion constraints to obtain hand joint chain structure data, where the hand joint chain structure data includes the connection relationships between joint points, joint mobility parameters, and bone length ratios; The gesture features of the hand joint chain structure data are extracted to obtain a sequence of hand gesture features, and the sequence of hand gesture features is smoothed in time series to obtain hand motion trajectory data; Perform spatial coordinate mapping on the hand movement trajectory data to obtain a sequence of spatial coordinates of hand key points, where the sequence of spatial coordinates of hand key points includes temporal position data of finger joint points, the center point of the palm, and the reference point of the wrist.

[0037] Specifically, the process of collecting three-dimensional feature points of the user's hand through a depth camera to obtain a sequence of spatial coordinates of hand key points first involves using the depth camera to collect depth images of the hand area from multiple perspectives, thereby obtaining a series of depth images of the hand area. In the application scenario of smart home control, when the user hopes to adjust the light brightness or control home appliances through gestures, the depth camera will capture the user's hand movements from different angles and generate a sequence of multiple images containing depth information. These images not only reflect the depth difference between the hand and the background, but also contain rich structural details, providing a basic data source for subsequent processing. Next, the system will perform structured light coding analysis on the obtained sequence of depth images of the hand area. This process involves interpreting the mapping relationship between each pixel point in the depth image and the actual physical distance, and finally extracting the hand point cloud data. Point cloud is a technology representing a discrete point set in three-dimensional space, and each point carries accurate position information in space. Therefore, the hand point cloud data can accurately describe the shape of the user's hand in three-dimensional space. Subsequently, the system needs to perform bone joint positioning on the hand point cloud data to determine the specific positions of each joint of the hand and form a set of hand joint candidate points. This step depends on advanced computer vision algorithms, which can identify the key points in the point cloud that are most likely to represent hand joints. For example, in the smart home control scenario, when the user makes a specific gesture, the system not only needs to identify the positions of key parts such as fingertips, knuckles, and the palm center, but also needs to consider the relative position relationships between these parts. Then, in order to improve the accuracy of joint positioning, the system will optimize the motion constraints of the initially determined set of hand joint candidate points. This means adjusting the positions of these candidate points according to ergonomic principles and the actual range of motion of hand joints to ensure that they conform to the real anatomical structure. Through this process, more accurate hand joint chain structure data can be obtained, which not only contains the connection relationships between joint points, but also covers the range of motion parameters of each joint and the bone length ratio. This is crucial for understanding the complex motion patterns of the hand because it allows the system to more accurately simulate various posture changes of the hand. Then, gesture feature extraction is performed based on the hand joint chain structure data, and a representative sequence of hand gesture features is extracted from it. In a smart home environment, different gestures often correspond to different operation commands. For example, waving the hand can be used to switch TV channels, while making a fist may mean turning up the volume. Therefore, accurately extracting gesture features is particularly crucial for achieving efficient human-computer interaction. To this end, the system will analyze various geometric and dynamic characteristics in the hand joint chain structure data, including finger bending degree, palm orientation, and wrist rotation, etc., and extract a set of feature parameters that can uniquely identify a specific gesture from them. However, due to the influence of sensor noise and environmental factors, the gesture features directly extracted from the original data may have certain fluctuations.To eliminate these interferences, the system also needs to perform temporal smoothing processing on the hand gesture feature sequence, so as to obtain more stable hand movement trajectory data. This step helps to remove unnecessary high-frequency noise and retain the low-frequency movement trend that truly reflects the user's intention. Finally, by performing spatial coordinate mapping on the hand movement trajectory data, the system can obtain the hand key point spatial coordinate sequence. This means converting all the previously obtained information about hand movements into an expression form in a unified spatial coordinate system, so as to facilitate subsequent collision detection and touch event determination. In the application of smart home, specifically, the system not only needs to record the exact positions of finger joints, palm center points, and wrist reference points at different times, but also consider their continuous changes during the entire gesture execution process. In this way, a spatio-temporal data set that comprehensively describes the dynamic characteristics of the user's gesture can be generated, that is, the hand key point spatial coordinate sequence. This set of data not only contains the detailed position information of each key part of the hand, but also reflects the trend of their evolution over time, providing a solid foundation for accurate gesture recognition. To sum up, through the above series of steps, the system can effectively extract the three-dimensional features of the user's hand from the original depth image and convert them into a data format that is easy to understand and process, greatly improving the naturalness and fluency of human-computer interaction. This technical solution not only demonstrates the powerful capabilities of modern computer vision and machine learning, but also opens up a new path for more intelligent and convenient interaction methods in the future.

[0038] In a specific embodiment, the collision detection of the hand key point spatial coordinate sequence based on the interaction area grid data to obtain a touch event determination result includes: Perform spatial projection transformation on the hand key point spatial coordinate sequence to obtain a set of projection plane mapping points, and perform grid index matching on the set of projection plane mapping points to obtain candidate collision grid data; Perform depth threshold layering on the candidate collision grid data to obtain multiple layers of collision detection domains, and perform spatial overlap analysis on the multiple layers of collision detection domains to obtain collision hierarchy relationship data; Perform temporal collision tracking on the hand key point spatial coordinate sequence based on the collision hierarchy relationship data to obtain a collision event sequence, and perform state transition analysis on the collision event sequence to obtain collision state transition data; Perform depth gradient calculation on the collision state transition data to obtain a set of depth interaction parameters, and perform touch feature extraction on the set of depth interaction parameters to obtain a touch feature vector group; Perform comprehensive determination of touch events based on the touch feature vector group to obtain a touch event determination result, where the touch event determination result includes the grid cell index where the collision occurs, the collision duration, and the relative depth value of the collision point.

[0039] Specifically, the process of performing collision detection on the spatial coordinate sequence of hand key points based on the interactive region grid data to obtain a touch event determination result first involves performing a spatial projection transformation on the spatial coordinate sequence of hand key points to obtain a set of projected plane mapping points, and performing grid index matching on these mapping point sets to determine candidate collision grid data. In the application scenario of smart home control, assume that the user hopes to control the light brightness in the room or switch electrical appliances through gestures, and the depth camera has captured the user's hand movements and generated a spatial coordinate sequence of hand key points. Next, the system needs to convert these coordinates in the three-dimensional space into two-dimensional coordinates corresponding to the projection plane in order to identify whether the hand movement occurs within a specific interactive region. For example, when the user's finger points to a virtual button on the wall, the system calculates the specific position of the fingertip on the projection plane through spatial projection transformation and further searches for the corresponding hexagonal grid cell at that position, thus forming a data set containing potential collision grids. Then, perform depth threshold stratification on the candidate collision grid data to obtain multiple layers of collision detection domains, and perform spatial overlap analysis on these detection domains to finally obtain collision hierarchy relationship data. The core of this process lies in considering that there may be different depth levels in the actual operation environment. For example, the user's fingertip may be close to but not fully in contact with the projection plane. Therefore, the system divides the candidate collision grids into multiple levels according to a preset depth threshold, and each level represents a specific depth range. Then, the system will analyze the spatial overlap situation between these different levels to determine whether there is a continuous collision hierarchy. For example, in a smart home environment, if the user tries to adjust the light brightness, their finger may move up and down within a certain range. At this time, the system not only needs to identify the grid cell where the finger is located, but also needs to consider the dynamic changes of the finger at different depth levels to accurately distinguish different types of touch operations such as light touch and press. Perform temporal collision tracking on the spatial coordinate sequence of hand key points based on the collision hierarchy relationship data to obtain a collision event sequence, and perform state transition analysis on these collision event sequences to obtain collision state transition data. This step aims to track the changes of the user's gesture over time to ensure that the system can respond to the user's intention in real time. Specifically, when the user's finger gradually approaches the target grid cell, the system will record the timestamp and position information of each contact attempt, forming a series of ordered collision events. Subsequently, through state transition analysis of these collision events, the system can identify the start, duration, and end stages of the gesture and infer the user's operation intention accordingly. For example, during the process of adjusting the light brightness, if the user's finger stays in a grid cell for a period of time, it may mean that the user hopes to increase the brightness; while a quick swipe may be a signal to switch to the next control option.Next, perform a depth gradient calculation on the collision state transition data to obtain a set of depth interaction parameters, and extract touch characteristics from these parameter sets to generate a touch feature vector group. Here, the system not only needs to consider the position changes of the user's gesture, but also pay attention to the dynamic characteristics of the gesture at different depth levels. For example, in a smart home control scenario, the user may slightly move their finger back and forth to fine-tune the light brightness. To accurately capture such subtle operations, the system calculates the depth change rate in the collision state transition data to obtain a set of parameters describing the depth characteristics of the gesture. These parameters are then used to extract representative touch characteristics, such as the speed and acceleration of finger movement, as well as the dwell time at different depth levels, to form a feature vector group that comprehensively reflects the characteristics of the user's gesture. Finally, based on the touch feature vector group, a comprehensive determination of the touch event is performed to obtain the touch event determination result. This means that the system needs to combine all the collected information, including the grid cell index where the collision occurs, the collision duration, and the relative depth value of the collision point, etc., for a comprehensive analysis and judgment. In a smart home environment, once the user's gesture triggers a specific grid cell, the system will determine the corresponding operation instruction according to the above analysis results. For example, if the user's finger stays on a certain grid cell for a long time and maintains a certain depth pressure, the system may interpret it as a "press" operation and then execute the corresponding function, such as turning on or off a light. On the contrary, if the user's action is more rapid, it may be regarded as an indication of browsing or selecting other functions. In summary, through this series of complex and precise steps, the system not only achieves the accurate capture and parsing of the user's gesture, but also provides the user with a natural and intuitive human-computer interaction method, greatly improving the user experience and the system's response efficiency. This technical solution demonstrates the forefront progress in the fields of modern computer vision and human-computer interaction, and also lays a solid foundation for a more intelligent home control system in the future.

[0040] In a specific embodiment, the performing a depth gradient calculation on the collision state transition data to obtain a set of depth interaction parameters includes: Perform a temporal depth sampling on the collision state transition data to obtain a depth change sequence, and perform a difference operation on the depth change sequence to obtain depth gradient data, where the depth gradient data includes a depth change rate, a gradient direction vector, and local extreme points; Based on the depth gradient data, perform a depth distribution statistics to obtain a depth distribution feature set, and filter out outliers from the depth distribution feature set to obtain depth statistical parameters; Perform a spatial curvature analysis on the depth statistical parameters to obtain curvature feature data, and extract the principal curvature from the curvature feature data to obtain surface form parameters; Estimate the contact area based on the surface morphology parameters to obtain contact region parameters, and calculate the pressure distribution for the contact region parameters to obtain pressure distribution characteristics; Perform feature fusion processing on the pressure distribution characteristics to obtain a set of depth interaction parameters, where the set of depth interaction parameters includes the contact area size, the pressure distribution pattern, and the depth change trend.

[0041] Specifically, the process of calculating the deep gradient of the collision state transition data to obtain the deep interaction parameter set first involves performing temporal deep sampling on the collision state transition data to obtain a deep change sequence, and performing a difference operation on these sequences to obtain deep gradient data. In the application scenario of smart home control, assume that the user hopes to adjust the light brightness or switch the TV channel through gestures, and the system has captured the detailed spatial coordinates of the user's hand movements and their changes over time. To accurately identify the user's operation intention, the system needs to further analyze the dynamic characteristics of these movements in the depth dimension. Through temporal deep sampling, the system can extract the precise deep positions of the gesture at different time points, forming a time series containing multiple deep values. Next, by performing a difference operation on this deep change sequence, the deep difference between adjacent time points can be calculated, and thus deep gradient data including the deep change rate, gradient direction vector, and local extreme points can be obtained. For example, during the process of adjusting the light brightness, the user may slowly move their finger back and forth, and at this time, the deep gradient data of the system will reflect this subtle deep change trend, providing an important basis for subsequent touch feature extraction. Based on the deep gradient data, deep distribution statistics are performed to obtain a deep distribution feature set, and outliers are filtered from these feature sets to obtain deep statistical parameters. This process aims to extract representative feature information from a large amount of raw data. Specifically, the system will statistically analyze various numerical distribution situations in the deep change sequence, such as the average deep value, standard deviation, etc., so as to construct a feature set that comprehensively describes the deep characteristics of the gesture. However, due to some accidental factors in the actual operation process that may cause data fluctuations or incorrect readings, it is also necessary to filter outliers from the generated deep distribution feature set. This step helps to remove those data points that significantly deviate from the normal range, ensuring that the finally obtained deep statistical parameters can truly reflect the user's operation behavior. For example, in a smart home environment, if the user's finger occasionally quickly moves away and then returns to the target area, the system should be able to identify and exclude this short-term deep jump to avoid interfering with the determination of normal touch events. Then, spatial curvature analysis is performed on the deep statistical parameters to obtain curvature feature data, and the principal curvature is extracted from it to obtain surface shape parameters. Spatial curvature analysis is one of the important means to understand the shape of the gesture and its dynamic changes. In the smart home application case, when the user makes a complex three-dimensional gesture, the surface of their hand is not a simple planar structure, but shows certain curvature characteristics. By analyzing the spatial distribution of the deep statistical parameters, the system can calculate the curvature values at each position, and then generate a set of detailed curvature feature data. Subsequently, by extracting the principal curvature from these data, the main shape characteristics of the gesture, such as the degree of finger bending, palm orientation, etc., can be determined, which is particularly crucial for distinguishing different gesture types.For example, when adjusting the light brightness, the user may perform a fine-tuning operation by changing the degree of finger bending. At this time, the system needs to accurately recognize such subtle gesture changes in order to make a correct response. Based on the surface shape parameters, the contact area is estimated to obtain the contact region parameters, and the pressure distribution of these parameters is calculated to obtain the pressure distribution characteristics. In the context of smart home control, understanding the contact situation between the user and the virtual interface is crucial for enhancing the interaction experience. The system first estimates the true contact area between the user's finger and the projection plane according to the surface shape parameters, which not only involves the small-range contact at the fingertip but may also include the participation of part of the palm area. Then, by combining the depth gradient data and the curvature feature data, the system can further simulate the pressure distribution pattern exerted by the finger, thereby obtaining more comprehensive contact region parameters. For example, during the process of adjusting the light brightness, the user may indicate a desire to increase the brightness by increasing the pressing force of the finger. At this time, the system needs to be able to accurately perceive this pressure change and adjust the corresponding output result accordingly. Finally, the pressure distribution characteristics are subjected to feature fusion processing to generate a set of depth interaction parameters. This means integrating all the above information about the depth characteristics of the gesture into a format that is easy to understand and use. In smart home applications, the specific set of depth interaction parameters not only includes static features such as the size of the contact area and the pressure distribution pattern but also covers dynamic information such as the depth change trend. For example, when the user attempts to control the light brightness in the room through gestures, the system not only needs to record the specific area of contact between the finger and the virtual button and the magnitude of the pressure applied but also needs to pay attention to the change in the depth position of the finger during the entire operation process. In this way, the system can comprehensively and accurately capture the user's operation intention and make a timely and accurate response accordingly. In summary, through this series of complex and orderly steps, the system not only realizes a detailed analysis of the depth characteristics of the gesture but also provides the user with a natural and intuitive human-computer interaction method, greatly enhancing the user experience and the response efficiency of the system. This technical solution demonstrates the cutting-edge progress in the fields of modern computer vision and human-computer interaction and also lays a solid foundation for a more intelligent home control system in the future.

[0042] In a specific embodiment, the time-series collision tracking of the hand key point spatial coordinate sequence based on the collision hierarchy relationship data to obtain a collision event sequence includes: Configuring the spatio-temporal sampling interval of the collision hierarchy relationship data to obtain a collision detection frame sequence, and performing time-series registration on the collision detection frame sequence to obtain a collision sampling parameter set; Based on the collision sampling parameter set, performing time-series segmentation on the hand key point spatial coordinate sequence to obtain hand motion segment data, and performing trajectory interpolation reconstruction on the hand motion segment data to obtain continuous motion trajectory data; Perform collision prediction extrapolation on the continuous motion trajectory data to obtain a collision trend feature set, and perform spatial consistency constraint on the collision trend feature set to obtain collision state prediction data; Perform collision event detection based on the collision state prediction data to obtain a collision detection result sequence, and perform temporal correlation analysis on the collision detection result sequence to obtain collision event correlation data; Perform temporal state merging on the collision event correlation data to obtain a collision event sequence, where the collision event sequence includes a collision start time, a duration window, and a collision point trajectory equation.

[0043] Specifically, the process of performing temporal collision tracking on the spatial coordinate sequence of hand key points based on the collision hierarchy relationship data to obtain a collision event sequence first involves configuring the spatio-temporal sampling interval for the collision hierarchy relationship data to obtain a sequence of collision detection frames, and performing temporal registration on these frame sequences to obtain a set of collision sampling parameters. In the application scenario of smart home control, assume that the user wishes to adjust the light brightness or switch the TV channel through gestures. The system has captured the user's hand movements through a depth camera and generated a spatial coordinate sequence of hand key points with multiple depth levels. To accurately track the user's operation process, the system needs to set an appropriate spatio-temporal sampling interval, that is, determine the acquisition time interval and spatial resolution between each frame of image. This step ensures that the system can accurately record the changes in hand movements at different time points, forming a detailed sequence of collision detection frames. Then, through temporal registration of these frame sequences, the system can synchronize the hand position information at different time points, eliminate errors caused by device acquisition delays or environmental factors, and finally generate a unified set of collision sampling parameters. For example, during the process of adjusting the light brightness, the system can ensure that it can capture the subtle forward and backward movements of the user's fingers by setting an appropriate sampling interval, providing a reliable data basis for subsequent collision tracking. Perform temporal segmentation on the spatial coordinate sequence of hand key points based on the set of collision sampling parameters to obtain hand movement segment data, and perform trajectory interpolation and reconstruction on these segment data to obtain continuous motion trajectory data. The core of this process lies in extracting representative motion segments from the original discrete data and connecting them into a smooth continuous trajectory through mathematical methods. Specifically, the system will divide the spatial coordinate sequence of hand key points into several independent motion segments according to the time and spatial information in the set of collision sampling parameters, and each segment represents a specific gesture action within a certain period of time. Then, through trajectory interpolation and reconstruction of these segments, the gaps between actual sampling points can be filled to form a continuous curve describing the complete motion path of the hand. For example, in a smart home environment, when the user attempts to adjust the light brightness through gestures, their fingers may make small up and down fine-tuning movements within a small range. At this time, the system not only needs to identify the specific position of the fingers, but also needs to simulate the smooth motion trajectory of the fingers throughout the process through an interpolation algorithm to more accurately judge the user's operation intention. Perform collision prediction extrapolation on the continuous motion trajectory data to obtain a set of collision trend features, and perform spatial consistency constraints on these feature sets to obtain collision state prediction data. In this step, the system not only focuses on the gesture actions that have occurred, but also tries to predict the future collision possibilities. By analyzing the dynamic characteristics such as speed and acceleration in the continuous motion trajectory data, the system can infer the future motion trend of the hand within a certain period of time, and then generate a set of collision trend features describing potential collision points and their occurrence probabilities.However, considering that there may be various uncertain factors in the actual operating environment, such as slight tremors of the user's hand or external interference, it is necessary to impose spatial consistency constraints on these prediction results. This means combining the current actual position of the hand and known physical constraints to ensure that the predicted collision points conform to the actual situation. For example, in a smart home application, if a user's finger is gradually approaching a virtual button, the system not only needs to calculate the most likely position where the finger will touch, but also consider whether the finger may deviate from the target area due to sudden movements, so as to make a more robust prediction of the collision state. Based on the collision state prediction data, collision event detection is performed to obtain a sequence of collision detection results, and temporal correlation analysis is carried out on these result sequences to obtain collision event correlation data. Once the system has prediction information about potential collision points and their occurrence probabilities, it can start real-time monitoring of the contact situation between hand movements and grid cells within the interaction area. Whenever a new collision event is detected, the system records the corresponding collision detection results, including the time of occurrence of the collision, the duration, and the specific position of the collision point, etc. Then, by performing temporal correlation analysis on these sequences of collision detection results, the system can identify the relationships between individual collision events, such as which collisions occur continuously and which are independent. For example, during the process of adjusting the light brightness, if the user's finger stays in a grid cell for a long time, the system may regard it as a "press" operation; while a quick swipe may be a signal to switch to the next control option. Through this correlation analysis, the system can more accurately understand the user's operation intention and improve the interaction efficiency. Finally, temporal state merging is performed on the collision event correlation data to obtain a sequence of collision events. This means integrating all relevant collision events in chronological order to form a data set that comprehensively describes the entire gesture operation process. The specific sequence of collision events not only contains the basic information of each collision event, such as the start time of the collision, the duration window, and the trajectory equation of the collision point, but also reflects the logical relationships between the events. For example, in a smart home control scenario, when the user adjusts the light brightness through gestures, the system not only needs to record the exact time and position of each finger touch on the virtual button, but also consider the sequence and duration between these touches in order to correctly execute the corresponding functions. In this way, the system can comprehensively and meticulously capture the user's operation behavior and make timely and accurate responses accordingly, greatly enhancing the user experience and the interaction efficiency of the system. This technical solution demonstrates the cutting-edge progress in the fields of modern computer vision and human-computer interaction, and also lays a solid foundation for more intelligent home control systems in the future.

[0044] In a specific embodiment, the visual feedback rendering of the grid data of the interaction area based on the determination result of the touch event to obtain interaction state display data includes: Extract the boundary of the touch area from the determination result of the touch event to obtain touch activation range data, and calculate the light intensity distribution of the touch activation range data to obtain grid brightness mapping data; Perform color space conversion on the touch area corresponding to the projection plane based on the grid brightness mapping data to obtain a set of grid color parameters, and perform gradient interpolation processing on the set of grid color parameters to obtain color transition data; Perform dynamic ripple superposition on the color transition data to obtain ripple diffusion data, and perform spatio-temporal evolution calculation on the ripple diffusion data to obtain a ripple animation sequence; Perform visual effect synthesis based on the ripple animation sequence to obtain interactive state display data, where the interactive state display data includes the boundary brightness value, color gradient parameter, and dynamic ripple effect parameter of the touch activation area.

[0045] Specifically, the process of visually feedback rendering the interactive area grid data based on the touch event determination result to obtain the interactive state display data first involves extracting the touch area boundary from the touch event determination result to obtain the touch activation range data, and calculating the light intensity distribution of these data to obtain the grid brightness mapping data. In the application scenario of smart home control, assume that the user hopes to adjust the light brightness or switch the TV channel through gestures, and the system has already determined the specific touch event based on the spatial coordinate sequence of the hand key points. To provide intuitive feedback to the user, the system needs to clarify which areas have been activated as touch response areas. By analyzing the collision information in the touch event determination result, the system can extract the boundaries of the grid cells that specifically triggered the operation, forming a data set describing the touch activation range. Then, based on this data set, the system further calculates the light intensity distribution on each grid cell, determines its corresponding brightness value, and finally generates a set of grid brightness mapping data for guiding the subsequent rendering process. For example, when adjusting the light brightness, when the user's fingertip touches a specific hexagonal grid cell, the system will increase the brightness of this grid cell, making it stand out from the surrounding environment and providing clear visual feedback to the user. Next, based on the grid brightness mapping data, perform color space conversion on the touch area corresponding to the projection plane to obtain the grid color parameter set, and perform gradient interpolation processing on these parameter sets to obtain the color transition data. This process aims to enhance the visual effect by adjusting the color of the touch area. Specifically, the system first converts the grid brightness mapping data into a color space representation suitable for the display device to generate a set of parameters describing the color characteristics of each grid cell. Then, by performing gradient interpolation processing on these color parameters, smooth transitions between the colors of different grid cells can be achieved, making the entire touch area present a harmonious and coherent visual experience. For example, in a smart home environment, if the user's finger moves from one grid cell to another, the system can make the color change between these two cells appear natural and smooth through color gradient technology, rather than suddenly jumping. This not only improves the visual beauty but also helps the user more intuitively understand their operation trajectory and influence range. Perform dynamic ripple superposition on the color transition data to obtain the ripple diffusion data, and perform spatio-temporal evolution calculation on these data to obtain the ripple animation sequence. In this step, the system introduces a dynamic ripple effect to increase the interest and immediate feedback of the interaction. Specifically, whenever a new touch event is detected, the system will superimpose a ripple pattern starting from the touch point and spreading outward on the corresponding grid cell. To achieve this effect, the system first needs to calculate the initial form of the ripple and its evolution path over time, that is, the ripple diffusion data.Then, by performing spatio-temporal evolution calculations on these diffusion data, the system can simulate the process of how the ripples gradually expand and finally disappear over time, forming a series of ordered ripple animation frames. For example, during the process of adjusting the light brightness, when the user's finger touches a certain virtual button, the system can generate a bright ring at that position and gradually spread outwards as the finger leaves until it covers the entire touch area, providing strong visual feedback to the user. Finally, based on the ripple animation sequence, visual effects synthesis is performed to obtain the interactive state display data. This means integrating all the boundary brightness values, color gradient parameters, and dynamic ripple effect parameters of the touch activation area to form a visual display solution that comprehensively reflects the current interactive state. In the application of smart home, the ultimate goal of the system is to ensure that users can clearly see the results of each of their gesture operations and make the next decision based on this. Therefore, in addition to the basic boundary brightness and color gradient of the touch activation area, the system also adds rich dynamic ripple effects, making each touch accompanied by a unique visual cue. For example, when adjusting the light brightness, if the user hopes to gradually brighten the lights in the room, they can confirm that their operation has been recognized by observing the continuously spreading ripples on the projection plane and adjust the strength and speed of the next gesture according to the changes in the ripples. In this way, the system not only achieves efficient human-computer interaction but also greatly improves the intuitiveness and pleasure of the user experience. In summary, through this series of complex and orderly steps, the system can not only accurately capture the user's operation intention but also ensure that every interaction can be presented in a timely and vivid manner through a rich visual feedback mechanism, laying a solid foundation for a more intelligent and user-friendly home control system in the future. This technical solution demonstrates the forefront progress in the fields of modern computer vision and human-computer interaction and also provides unprecedented convenience and fun for users.

[0046] The above describes the virtual touch interaction method with regional limitation in the embodiments of the present invention. Next, the virtual touch interaction device with regional limitation in the embodiments of the present invention will be described. Please refer to Figure 2 , an embodiment of the virtual touch interaction device with regional limitation in the embodiments of the present invention includes: A scanning module 21, configured to perform boundary scanning on the projection plane with regional limitation to obtain a set of projection area contour coordinates; A dividing module 22, configured to perform regional division on the projection plane based on the set of projection area contour coordinates to obtain interactive area grid data; An acquisition module 23, configured to collect three-dimensional feature points of the user's hand through a depth camera to obtain a sequence of spatial coordinates of hand key points; The detection module 24 is configured to perform collision detection on the spatial coordinate sequence of the hand key points based on the interactive area grid data when the spatial coordinate sequence of the hand key points is within the depth interaction threshold range of the interactive area grid data, so as to obtain a determination result of a touch event; The rendering module 25 is configured to perform visual feedback rendering on the interactive area grid data based on the determination result of the touch event, so as to obtain interactive state display data.

[0047] In this embodiment, for the specific implementation of each unit in the above device embodiment, please refer to the description in the above method embodiment, and details are not described herein again.

[0048] Refer to Figure 3 , an embodiment of the present invention further provides a computer device, and its internal structure may be as Figure 3 shown. The computer device includes a processor, a memory, a display screen, an input device, a network interface, and a database connected through a system bus. Among them, the processor of the computer design is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store the corresponding data in this embodiment. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, the above method is implemented.

[0049] Those skilled in the art can understand that Figure 3 the structure shown in

[0050] is only a block diagram of a partial structure related to the solution of the present invention, and does not constitute a limitation on the computer device to which the solution of the present invention is applied.

[0051] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium provided by the present invention and used in the embodiments can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM, etc.

[0052] It should be noted that in this article, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, device, article or method comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, device, article or method. Without further limitation, an element defined by the statement "comprising one..." does not exclude the existence of additional identical elements in the process, device, article or method comprising such element.

[0053] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structure or equivalent process transformation made by using the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.

Claims

1. A region-limited virtual touch interaction method, characterized in that Including the following steps: Perform boundary scanning on the projection plane defined by the region to obtain a set of contour coordinates of the projection region; Based on the set of contour coordinates of the projection region, divide the projection plane to obtain interactive region grid data; Collect three-dimensional feature points of the user's hand through a depth camera to obtain a sequence of spatial coordinates of hand key points; When the sequence of spatial coordinates of the hand key points is within the depth interaction threshold range of the interactive region grid data, perform collision detection on the sequence of spatial coordinates of the hand key points based on the interactive region grid data to obtain a determination result of a touch event; Based on the determination result of the touch event, perform visual feedback rendering on the interactive region grid data to obtain interactive state display data.

2. The virtual touch interaction method defined by region according to claim 1, wherein The performing boundary scanning on the projection plane defined by the region to obtain a set of contour coordinates of the projection region includes: Emit a raster light spot array to the projection plane defined by the region through a structured light projector, and collect the reflected light intensity distribution of the raster light spot array through a binocular depth camera to obtain depth scattering data of the projection plane; Extract edge contours and perform spatial coordinate mapping on the depth scattering data to obtain a point cloud of the projection region boundary, and perform spatial filtering on the point cloud of the projection region boundary to obtain a boundary feature sequence; Perform polygon fitting and vertex positioning based on the boundary feature sequence to obtain a set of contour coordinates of the projection region; wherein, the set of contour coordinates of the projection region includes the spatial position coordinates of four vertices of the projection plane and the normal vector of the projection plane.

3. The virtual touch interaction method defined by region according to claim 1, wherein, The dividing the projection plane based on the set of contour coordinates of the projection region to obtain interactive region grid data includes: Perform homogeneous coordinate transformation on the set of contour coordinates of the projection region to obtain a plane projection transformation matrix, and perform singular value decomposition on the plane projection transformation matrix to obtain main direction feature data of the projection region; Based on the main direction feature data of the projection region, arrange hexagonal grid seed points on the projection plane to obtain initial grid layout parameters, and perform Thiessen polygon division on the initial grid layout parameters to obtain regular grid division data; Perform spatial coordinate remapping on the regular grid division data to obtain a set of grid depth parameters, and perform surface interpolation reconstruction on the set of grid depth parameters to obtain depth attribute data of grid cells; Based on the depth attribute data of the grid cells, construct a grid topological relationship to obtain interactive region grid data, wherein the interactive region grid data includes the boundary coordinates of multiple regular hexagonal grid cells, the depth values of the grid cells, and the topological relationship between the grid cells.

4. The virtual touch interaction method defined by region according to claim 1, wherein, The collecting three-dimensional feature points of the user's hand through a depth camera to obtain a sequence of spatial coordinates of hand key points includes: Collect multi-view depth images of the user's hand through a depth camera to obtain a sequence of depth images of the hand region, and perform structured light coding analysis on the sequence of depth images of the hand region to obtain hand point cloud data; Perform bone joint localization on the hand point cloud data to obtain a set of candidate hand joint points, and optimize the motion constraints of the set of candidate hand joint points to obtain hand joint chain structure data, where the hand joint chain structure data includes the connection relationships between joint points, joint mobility parameters, and bone length ratios; Extract gesture features from the hand joint chain structure data to obtain a hand pose feature sequence, and perform temporal smoothing processing on the hand pose feature sequence to obtain hand motion trajectory data; Perform spatial coordinate mapping on the hand motion trajectory data to obtain a sequence of hand key point spatial coordinates, where the sequence of hand key point spatial coordinates includes the temporal position data of finger joint points, the palm center point, and the wrist reference point.

5. The virtual touch interaction method limited by a region according to claim 1, wherein Perform collision detection on the sequence of hand key point spatial coordinates based on the interactive area grid data to obtain a touch event determination result, including: Perform spatial projection transformation on the sequence of hand key point spatial coordinates to obtain a set of projected plane mapping points, and perform grid index matching on the set of projected plane mapping points to obtain candidate collision grid data; Perform depth threshold layering on the candidate collision grid data to obtain multiple layers of collision detection domains, and perform spatial overlap analysis on the multiple layers of collision detection domains to obtain collision hierarchy relationship data; Perform temporal collision tracking on the sequence of hand key point spatial coordinates based on the collision hierarchy relationship data to obtain a collision event sequence, and perform state transition analysis on the collision event sequence to obtain collision state transition data; Perform depth gradient calculation on the collision state transition data to obtain a set of depth interaction parameters, and perform touch feature extraction on the set of depth interaction parameters to obtain a touch feature vector group; Perform comprehensive determination of touch events based on the touch feature vector group to obtain a touch event determination result, where the touch event determination result includes the grid cell index where the collision occurs, the collision duration, and the relative depth value of the collision point.

6. The region-defined virtual touch interaction method according to claim 5, characterized in that, The performing depth gradient calculation on the collision state transition data to obtain a set of depth interaction parameters includes: Perform temporal depth sampling on the collision state transition data to obtain a depth change sequence, and perform difference operation on the depth change sequence to obtain depth gradient data, where the depth gradient data includes a depth change rate, a gradient direction vector, and local extreme points; Perform depth distribution statistics based on the depth gradient data to obtain a set of depth distribution features, and perform outlier filtering on the set of depth distribution features to obtain depth statistical parameters; Perform spatial curvature analysis on the depth statistical parameters to obtain curvature feature data, and perform principal curvature extraction on the curvature feature data to obtain surface shape parameters; Perform contact area estimation based on the surface shape parameters to obtain contact area parameters, and perform pressure distribution calculation on the contact area parameters to obtain pressure distribution features; Perform feature fusion processing on the pressure distribution features to obtain a set of depth interaction parameters, where the set of depth interaction parameters includes the contact area size, the pressure distribution pattern, and the depth change trend.

7. The virtual touch interaction method defined by region according to claim 1, characterized in that, Performing visual feedback rendering on the interactive area grid data based on the determination result of the touch event to obtain interactive state display data, including: Extracting the boundary of the touch area from the determination result of the touch event to obtain touch activation range data, and calculating the light intensity distribution of the touch activation range data to obtain grid brightness mapping data; Performing color space conversion on the touch area corresponding to the projection plane based on the grid brightness mapping data to obtain a set of grid color parameters, and performing gradient interpolation processing on the set of grid color parameters to obtain color transition data; Performing dynamic ripple superposition on the color transition data to obtain ripple diffusion data, and performing spatio-temporal evolution calculation on the ripple diffusion data to obtain a ripple animation sequence; Performing visual effect synthesis based on the ripple animation sequence to obtain interactive state display data, wherein the interactive state display data includes the boundary brightness value, color gradient parameters, and dynamic ripple effect parameters of the touch activation area.

8. A region-limited virtual touch interaction device, characterized in that, Including: A scanning module for performing boundary scanning on the projection plane defined by the area to obtain a set of projection area contour coordinates; A partitioning module for partitioning the projection plane based on the set of projection area contour coordinates to obtain interactive area grid data; An acquisition module for collecting three-dimensional feature points of the user's hand through a depth camera to obtain a sequence of spatial coordinates of hand key points; A detection module for performing collision detection on the sequence of spatial coordinates of hand key points based on the interactive area grid data when the sequence of spatial coordinates of hand key points is within the depth interaction threshold range of the interactive area grid data to obtain a determination result of the touch event; A rendering module for performing visual feedback rendering on the interactive area grid data based on the determination result of the touch event to obtain interactive state display data.

9. A computer device, comprising a memory and a processor, wherein a computer program is stored in the memory, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Interactive processing method and system for digital media file

    CN121433554A

  • Dynamic projection interactive calibration method and system for annular science magic space

    CN121685649A

  • A dynamic projection interaction calibration method and system for a ring-shaped science fiction space

    CN121685649B