Somatosensory interaction system based on 3D vision enhancement technology
Through the combination of multimodal sensor arrays and 3D spatiotemporal convolutional networks, an action feature model library is built, and the somatosensory interaction system is optimized using dynamic time regularization algorithms and dual traceability mechanisms, which solves the accuracy and user experience problems of action recognition in complex environments, and achieves high-precision and high-natural user interaction.
Patent Information
- Application Number
- CN202510614413.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-08-15
AI Technical Summary
The existing somatosensory interactive systems have low accuracy in action recognition and poor user experience in complex environments, mainly due to the lack of multi-dimensional dynamic feature fusion of single skeleton joint displacement analysis, making it difficult to distinguish user action signals from environmental noise, and fail to adjust in real time.
A multimodal sensor array is used to collect user skeletal node coordinates, three-dimensional motion trajectories and environmental source data, combine 3D spatiotemporal convolution network to extract multi-scale action features, calculate action matching similarity through dynamic time regularization algorithm, and use environmental analysis and evaluation index dual traceability mechanisms to generate environmental optimization or action optimization suggestions, and combine multimodal feedback mechanism to improve user experience.
It significantly improves the accuracy and robustness of action recognition in complex environments, enhances the naturalness and reality of user interaction, optimizes system performance through adaptive adjustments, reduces interaction delays and improves user experience.
Smart Images

Figure CN120491824A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of visual-body interaction, and in particular to a body interaction system based on 3D vision enhancement technology. Background Art
[0002] As a cutting-edge means of integrating virtual and real life, 3D augmented reality (3D-AR) technology is gradually penetrating multiple fields, with its application in haptic interaction becoming increasingly widespread. Leveraging computer vision and sensor technologies, it allows users to interact naturally with virtual scenes through physical movements, bringing new experiences to industries such as entertainment, education, and healthcare. However, current haptic interactions based on 3D-AR still face numerous challenges: During motion recognition, existing technologies rely on single-point skeletal joint displacement analysis and lack the integration of multi-dimensional dynamic features. This single analysis method results in poor performance in distinguishing similar motions, and ineffective separation of user motion signals from ambient noise in complex environments. Furthermore, the system fails to dynamically adjust based on the real-time environment and user status, resulting in low motion recognition accuracy and a poor user experience in complex environments. In view of the above defects, a solution is now proposed. Summary of the Invention
[0003] The purpose of the present invention is to solve the problems of low accuracy of motion recognition and decreased user experience in existing somatosensory interaction systems in complex environments, and to propose a somatosensory interaction system based on 3D vision enhancement technology.
[0004] The purpose of the present invention can be achieved by the following technical solution: A somatosensory interaction system based on 3D visual enhancement technology, comprising: The motion acquisition module is used to collect the user's skeletal joint coordinates, motion speed, three-dimensional motion trajectory and environmental source data through a multimodal sensor array; The data processing module is used to extract user action data to generate multi-level feature maps, build multi-scale descriptors, and build a human action feature model library based on labeled samples; The action recognition optimization module includes an action matching calculation unit, an analysis unit, and an action failure tracing unit; The action matching calculation unit uses a dynamic time warping algorithm to calculate the optimal path cumulative distance between the real-time action sequence and the standard template, and obtains the action matching similarity SD after normalization. If the action matching similarity SD exceeds the threshold, a virtual interaction event is triggered; otherwise, the action is automatically transmitted to the analysis unit. The analysis unit is divided into an environmental analysis subunit and an action analysis subunit; When the interaction action recognition fails, the environmental analysis subunit analyzes the environmental impact parameters of the interaction space and obtains the environmental analysis evaluation index FXS; When the interactive action recognition fails, the interactive action analysis subunit analyzes the interactive action instruction and obtains the action instruction analysis evaluation index DZC; The action failure tracing unit determines the cause of action recognition failure based on the environment analysis evaluation index FXS and the action command analysis evaluation index DZC, and implements environment optimization suggestions or action optimization suggestions; The somatosensory interaction module is used to generate interactive responses to virtual scenes, optimize the user experience through visual, auditory, and tactile feedback, and perform adaptive adjustments.
[0005] Furthermore, the specific processing process of the data processing module is as follows: Based on the 3D spatiotemporal convolutional network, the coordinates, accelerations, and posture angles of skeletal joints are extracted from point cloud data, and multi-level feature maps are output. Multi-level features are fused to construct multi-scale descriptors. Combined with labeled samples from big data, contrastive learning is used to optimize the feature embedding space, and a human motion feature model library is constructed. Based on this library, a user interaction control model is constructed.
[0006] Furthermore, the specific matching steps of the action matching similarity of the action matching calculation unit are: Load the standard action template from the constructed user interaction control model, and use the dynamic time warping algorithm to calculate the optimal path cumulative distance between the real-time action sequence and the standard template ; Will After normalization, substitute into the formula , get the action matching similarity value SD, where The maximum cumulative distance of the path passing through historical data.
[0007] Furthermore, the specific step 1 of the analysis unit analyzing the interaction environment impact parameter is: The process from when the system is initialized to recognize interactive actions is recorded as the analysis interval. The depth camera is used to collect environmental data of the interactive space within the current analysis interval and calculate the light intensity value GZ, the environmental complexity value FZ, and the volume value of the available interactive space TJ. At the same time, the number of people in the current interactive space is identified to obtain the number of people RS. Establish a three-dimensional coordinate system with the current interactive space identification center point as the reference point, calculate the relative distance JL between the user and the virtual space, and match the preset range to obtain the distance impact value YX; identify and locate the obstructions within the current analysis interval and calculate the obstruction degree value ZD; After normalizing the ambient light intensity value GZ, the environmental complexity value FZ, the available interactive space volume value TJ, and the occlusion degree value ZD, substitute them into the formula Get the environmental analysis index HJS; and These are the influence weight factors of light intensity value, environment complexity value, available interactive space volume value, and occlusion degree value; Combined with the number of people RS and the interactive user distance influence value YX, after normalization, substitute into the formula Get the comprehensive environmental analysis index ZHS; 、 and They represent the environmental analysis reference index, the reference number of people, and the reference value of the distance impact of interactive users in the current interactive environment respectively; and They are the influence weight factors of the environmental analysis index, the number of people and the impact value of the interactive user distance.
[0008] Furthermore, the specific second step of the analysis unit analyzing the interaction environment impact parameters is: The light uniformity index GX is calculated by using the light intensity value within the current analysis interval obtained by the light sensor; the light impact value GY is obtained by matching the preset range; the reflective surface distribution is analyzed using image processing technology to calculate the reflective impact value FG; the number of detected moving objects is marked as YD; Normalize the light impact value GY, the reflection interference value FG, and the number of moving objects YD in the current analysis interval and then enter them into the formula , get the interference environment analysis index GRS; where 、 as well as They are the light impact reference value, the reflection interference reference value and the moving object reference value. and They are the influence weight factors of light impact value, reflection interference value and number of moving objects; After normalizing the environmental comprehensive analysis index ZHS and the interference environment analysis index GRS in the current analysis interval, enter them into the formula , and obtain the environmental analysis and evaluation index FXS; and They represent the comprehensive environmental analysis reference index and interference environment analysis reference index of the current interactive environment respectively. are the impact weight factors of the comprehensive environmental analysis index and the interference environment analysis index respectively.
[0009] Furthermore, the specific steps of the analysis unit analyzing the interaction action impact parameters are as follows: Extract the duration and complexity of the interactive action command, match the preset data package to obtain the action speed estimation DS; identify the action result through the interactive user action command and compare it with the standard trajectory of the possible results to obtain the action trajectory similarity DF; based on the comparison result, after normalization, substitute it into the formula , get the action command analysis evaluation index DZC; where and They represent the number of action trajectory features for interactive action instruction recognition and the number of standard trajectory features for possible results respectively; Indicates the allowable difference between the trajectory characteristics of the interactive action instruction recognition and the standard trajectory characteristics; and They represent the allowed estimation value of motion speed and the allowed value of motion trajectory similarity respectively.
[0010] Furthermore, the specific steps of the action failure tracing unit for action failure tracing are: Based on the comparison of the environment analysis evaluation index FXS and the action command analysis evaluation index DZC with the corresponding thresholds, the cause of failure is determined and optimization is performed; if the environment analysis evaluation index is higher than the threshold and the action command analysis evaluation index is lower than the threshold, it is determined that the current interactive environment has interfered with action recognition, the environment analysis evaluation index is analyzed to identify the interference source, and environmental optimization suggestions are generated. The environmental optimization suggestions include: for occlusion interference, visually prompting the position of the occlusion and guiding the user to adjust it; for reflective interference, dynamically adjusting the camera exposure parameters and gain parameters, and prompting the light source adjustment direction on the interactive interface; if the action command analysis evaluation index DZC is higher than the threshold and the environment analysis evaluation index is lower than the threshold, it is determined that there is a problem with the current user's action command, and action optimization suggestions are output. The action optimization suggestions include dynamic correction suggestions, adjustment of weight coefficients R1, R2, R3, and reference to reinforcement learning mechanisms; If the environment analysis evaluation index FXS and the motion command analysis evaluation index DZC are both higher than the corresponding thresholds, it is determined that there are interference causes in both the interaction environment and the motion command, and the environment optimization suggestions and the motion optimization suggestions are executed simultaneously.
[0011] Furthermore, the somatosensory interaction module includes a virtual mapping unit, a multimodal feedback unit, and an adaptive adjustment unit; The virtual mapping unit maps the user's skeleton coordinates to the virtual space through a homogeneous coordinate transformation matrix, and updates the position offset in real time based on the millimeter wave radar; The multimodal feedback generates visual effects, three-dimensional sound fields, and tactile signals based on the intention recognition results. The visual feedback pre-renders the scene frame based on the optical flow algorithm and dynamically adjusts the brightness and contour highlights; the auditory feedback uses the HRTF algorithm to generate spatialized sound effects and reduces noise based on the complexity of the environment; the tactile feedback synchronizes the action timing through the force feedback glove and dynamically adjusts the gain; The adaptive adjustment unit monitors the recognition action accuracy and delay indicators, triggering the self-diagnosis mechanism to optimize system parameters or guide users to make adjustments.
[0012] Compared with the prior art, the present invention has the following beneficial effects: The present invention uses a multimodal sensor array to collect skeletal joint coordinates, three-dimensional motion trajectories, and environmental source data in real time. It then uses a 3D spatiotemporal convolutional network to extract multi-scale dynamic features, constructing a human motion feature model library that can distinguish motions. It also uses a dynamic time warping algorithm to calculate the matching similarity between real-time motions and standard templates. Through a dual traceability mechanism using the environmental analysis evaluation index (FXS) and the motion command analysis evaluation index (DZC), it identifies environmental interference or motion deviations, and dynamically generates environmental or motion optimization suggestions, effectively improving the accuracy and robustness of motion recognition in complex environments. The present invention maps the user's skeletal coordinates to the virtual space through a homogeneous coordinate transformation matrix, and combines millimeter-wave radar to update the spatial offset in real time to ensure the synchronization of physical and virtual scenes; adopts a multimodal feedback mechanism to achieve coordinated response of vision, hearing and touch; combines a reinforcement learning mechanism to optimize the weights of the motion recognition model, monitors the motion recognition accuracy and delay indicators in real time, and triggers self-diagnosis closed-loop calibration, thereby reducing interaction delays while improving the naturalness of user actions and system adaptability, significantly enhancing the realism of somatosensory interaction and user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] In order to facilitate understanding by those skilled in the art, the present invention will be further described below with reference to the accompanying drawings; Figure 1 This is the overall system block diagram of the present invention. DETAILED DESCRIPTION
[0014] The following will clearly and completely describe the technical solutions of the present invention in conjunction with the embodiments. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts are within the scope of protection of the present invention.
[0015] It should be understood that the terms “include” and “comprising” used in the specification and claims of the present disclosure indicate the presence of described features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or collections thereof.
[0016] It should also be understood that the terminology used in this disclosure is for the purpose of describing specific embodiments only and is not intended to limit the disclosure. As used in this disclosure and the claims, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise. It should further be understood that the term "and / or" as used in this disclosure and the claims refers to any and all possible combinations of one or more of the associated listed items, including and including these combinations.
[0017] like Figure 1 As shown, a somatosensory interaction system based on 3D vision enhancement technology includes a motion acquisition module, a data processing module, a motion recognition optimization module and a somatosensory interaction module.
[0018] The motion acquisition module collects multi-dimensional data of the user through a pre-built multi-modal sensor array to obtain the corresponding user motion data, and records the collected environmental source data; and stores it in a pre-built motion database; The sensor array includes a depth camera, an inertial measurement unit, and a millimeter-wave radar. The depth camera collects motion image data of the user from different angles and captures the user's three-dimensional motion trajectory and skeletal joint coordinates. The inertial measurement unit collects the motion posture of various parts of the user's body, as well as the acceleration and angular velocity during motion. The millimeter-wave radar obtains stereoscopic image data of the user's position and the user's surrounding environment, constructs a three-dimensional spatial model, and obtains corresponding motion data based on the multi-dimensionally collected data. The motion data includes: raising a hand, raising a leg, jumping, and walking. The environmental source data is also recorded simultaneously. The environmental source data includes light intensity and spatial coordinates. The obtained point cloud data, skeletal joint coordinates, and environmental source data are classified and stored in a motion feature database. The data processing module is used to perform feature processing on the collected user motion data and build a corresponding human motion feature model based on a pre-built motion database. The specific construction process includes: extracting the coordinate displacement, acceleration, and posture angle of skeletal joint points from the point cloud data based on a 3D spatiotemporal convolutional network, outputting a multi-level feature map, and building a multi-scale descriptor by fusing the multi-level features through spatial pyramid pooling. In combination with the annotated samples in the big data, contrastive learning is used to optimize the feature embedding space to generate a human motion feature model library that can distinguish motions. Based on the distinguishable motions generated by the human motion feature model library, a corresponding user interaction control model is built, and the motion intention recognition results are output; The action recognition optimization module includes an action matching calculation unit, an analysis unit, and an action failure tracing unit; The action matching calculation unit is used to calculate the action matching similarity between the interactive user's real-time action and the standard action. The specific matching steps of the action matching similarity are: Load the standard action template from the constructed user interaction control model, use the dynamic time warping algorithm, and use the formula , calculate the optimal path cumulative distance between the real-time action sequence and the standard template, where is the feature vector of the i-th frame in the real-time action sequence; is the feature vector of the jth frame in the standard template; is the Euclidean distance between the feature vectors of the two frames; n and m are the number of frames of the real-time sequence and the template respectively, To regularize the path, define boundary constraints and monotonic constraints; For what you have obtained Perform normalization processing to obtain the corresponding action matching similarity value SD, where the mathematical expression formula of normalization processing is: ,in The maximum cumulative distance of the path passing through historical data; The action recognition results are judged based on the preset action matching similarity threshold: when the action matching similarity SD exceeds the preset action matching similarity threshold, the action recognition is judged to be successful, and the system directly maps the action characteristics and action intention to the virtual space, triggering the corresponding interaction event; when the action matching similarity SD is lower than the preset action matching similarity threshold, the action recognition is judged to be unsuccessful, and the system automatically transmits the action to the analysis unit; The analysis unit is used to conduct in-depth analysis of the actions that failed to be recognized and determine the cause of the recognition deviation; the analysis unit includes an environment analysis subunit and an interactive action analysis subunit; When the interaction action recognition fails, the environment analysis subunit analyzes the environmental impact parameters of the interaction space and obtains the environment analysis index HJS. The specific steps are as follows: The process from when the system is initialized to recognize interactive actions is recorded as the analysis interval. The depth camera is used to collect data on the light intensity, environment complexity, and spatial dimension of the interactive space within the current analysis interval. Use the spatial positioning algorithm to process the environmental data collected by the depth camera and calculate the current ambient light intensity value GZ, the environmental complexity value FZ and the available interactive space volume value TJ; at the same time, identify the number of people in the current interactive space and obtain the number of people RS; Establish a three-dimensional coordinate system with the current interactive space identification center point as the reference point; calculate the coordinates and distance value JL of the interactive user relative to the virtual space based on the relative position of the interactive user and the virtual space; match the interactive user's distance value JL with multiple set value ranges to obtain the interactive user's distance influence value YX within the current analysis interval, setting each value range to match a distance influence value; Use depth image analysis technology to identify and locate potential obstructions within the current analysis interval and calculate the occlusion degree value ZD; Substitute the ambient light intensity value GZ, the environmental complexity value FZ, the available interactive space volume value TJ, and the occlusion degree value ZD into the formula , calculate and obtain the environmental analysis index HJS within the current analysis interval; as well as They are the influence weight factors of the light intensity value GZ, the environment complexity value FZ, the available interactive space volume value TJ, and the occlusion degree value ZD; Substitute the environmental analysis index HJS, the number of people RS, and the interactive user distance impact value YX in the current analysis interval into the formula , calculate and obtain the comprehensive environmental analysis index ZHS of the current analysis interval; 、 and Respectively represent the HJS reference environment analysis value, RS reference environment analysis value and YX reference environment analysis value of the current interaction environment; and are the influence weight factors of each parameter respectively; By setting up a preset light sensor, we obtain light intensity values in different directions within the current analysis interval and calculate the light uniformity index GX. We match the current light uniformity index GX with multiple corresponding value ranges to obtain the light impact value GY, with each value range assigned a corresponding light impact value. We use image processing technology to analyze the distribution of reflective surfaces within the current analysis interval and calculate the proportion of reflective surfaces in the overall field of view to obtain the reflective impact value FG. We also assign a reflective interference value to each value range. We also mark the number of moving objects detected within the current analysis interval as YD. Substitute the light impact value GY, the reflection interference value FG and the number of moving objects YD in the current analysis interval into the formula , calculate and obtain the interference environment analysis index GRS of the current analysis interval; 、 as well as They are the light impact reference value, the reflection interference reference value and the moving object reference value. They are the influence weight factors of light impact value, reflection interference value and number of moving objects; Substitute the comprehensive environmental analysis index ZHS and interference environment analysis index GRS in the current analysis interval into the formula , calculate and obtain the environmental analysis assessment index FXS of the current analysis interval; and They represent the comprehensive environmental analysis reference index and interference environment analysis reference index of the current interactive environment respectively. are the impact weight factors of the comprehensive environmental analysis index and the interference environment analysis index respectively; When the interactive action analysis subunit fails to identify the interactive action, it conducts an in-depth analysis of the interactive action instruction. The specific process includes: Extract the duration of the interactive action instruction after feature processing in seconds, and simultaneously obtain the recognition action result corresponding to the current interactive user action instruction, analyze the action complexity, and integrate it with the duration of the interactive user action instruction to obtain the current interactive user action instruction data packet; match the interactive user action instruction data packet in the current analysis interval with the corresponding multiple preset data packets; obtain the interactive user action instruction data packet matching result and the action speed estimation DS; set each preset data packet to correspond to a matching result and action speed estimation DS; the matching results include fast action speed, too fast action speed, normal action speed, slow action speed, and too slow action speed; Compare the interactive user action command recognition result in the current analysis interval with the standard trajectory of the possible results, and calculate the action trajectory similarity DF between the two; Based on the comparison results between the action trajectory characteristics of the interactive user's action instructions in the current analysis interval and the standard trajectory characteristics of the possible results, the above parameters are substituted into the formula , calculate and obtain the interactive user action instruction analysis evaluation index DZC in the current analysis interval; where R1 and R2 represent the number of interactive action instruction recognition action trajectory features and the number of standard trajectory features of possible results respectively; Indicates the allowable difference between the trajectory characteristics of the interactive action instruction recognition and the standard trajectory characteristics; and They represent the allowed estimation value of motion speed and the allowed value of motion trajectory similarity respectively; The action failure tracing unit compares the environmental analysis evaluation index FXS and the action command analysis evaluation index obtained in the current analysis interval with the preset thresholds to determine the cause of the action recognition failure. The specific steps are as follows: When the environment analysis evaluation index FXS is higher than the preset threshold and the action command analysis evaluation index DZC is lower than the preset threshold, it means that the current interaction environment has caused significant interference to action recognition. Then, the following steps are performed: Based on the various parameters of the environmental analysis and evaluation index FXS, the main and secondary environmental interference sources are further analyzed, and the relevant parameters are compared with the preset thresholds of the relevant parameters. The relevant parameters include light intensity value GZ, light uniformity index GX, environmental complexity value FZ, available interactive space volume value TJ, occlusion value ZD, number of people RS, reflective interference value FG and number of moving objects YD; the relative deviation M and relative deviation between the measured value of the relevant parameters and the preset threshold are calculated. ; Compare the relative deviation M with the preset threshold. If the relative deviation M , then the interference can be ignored, if Relative deviation M It is determined to be a secondary interference source; if the relative deviation M , it is determined to be the main interference source; Generates environmental optimization suggestions based on primary and secondary environmental interference sources. These suggestions include: For occlusion interference, visual prompts mark the location of obstructions and guide users to adjust to the system's preset unobstructed interaction area; for reflection interference, dynamically adjusts camera exposure and gain parameters, and prompts the user to adjust the light source direction in the interactive interface; and for light interference, uses an adaptive histogram equalization algorithm based on ambient light sensor data to compensate for uneven lighting and provide recommended values for camera white balance parameters. When the action command analysis evaluation index DZC is higher than the preset threshold and the environment analysis evaluation index is lower than the preset threshold, it means that the current user's action command itself has trajectory deviation or speed problems, resulting in failure. Then, action optimization suggestions are output. The action optimization suggestions include: superimposing the difference between the user's action trajectory and the standard action trajectory through the 3D skeleton model in the virtual space, prompting the user to adjust the action or speed with a color gradient. The color gradient includes red for deviation areas and green for normal areas; and combining voice prompts with virtual coach animations to output dynamic correction suggestions. If the action trajectory similarity DF is lower than the preset threshold and the action speed estimation is higher than the preset threshold, it is determined that the action combination is complex. The complex combination is split into standard actions for step-by-step recognition. Based on the action recognition results, the matching threshold of the user interaction control model is adjusted, and the system algorithm is optimized. Based on historical training data, the feature distribution of highly complex actions is analyzed, and the calculation model of the current action command analysis evaluation index DZC is optimized. The specific optimization process includes adjusting the weight coefficients R1, R2, and R3 of the action trajectory similarity DF and the trajectory feature matching degree; introducing higher-order time series analysis or dynamic feature fusion strategies and adopting reinforcement learning mechanisms to dynamically optimize the action recognition model and improve its adaptability to complex or unconventional actions; When both the environment analysis evaluation index FXS and the motion command analysis evaluation index DZC are higher than the preset threshold, it indicates that there are interference factors in both the interactive environment and the motion command, resulting in motion recognition failure. In this case, both the environment optimization suggestions and the motion optimization suggestions will be implemented simultaneously. The somatosensory interaction module responds to virtual scene interactions and optimizes user experience based on the multi-dimensional data from the motion recognition optimization module. It includes a virtual mapping unit, a multimodal feedback unit, and an adaptive adjustment unit. The virtual mapping unit maps the user interaction control model to the virtual environment and generates a corresponding virtual scene interaction response. The specific process includes: The user's skeletal joint coordinates are mapped to virtual space through a homogeneous coordinate transformation matrix, aligning the position and orientation of the physical and virtual spaces. Millimeter-wave radar is used to update the user's position in real time, adjusting the virtual coordinate offset to ensure that the position of virtual objects is updated synchronously as the user moves. The multimodal feedback unit provides multimodal feedback based on the interactive response of the virtual scene. The multimodal feedback includes visual feedback, auditory feedback, and tactile feedback. Visual feedback is based on user action intention recognition, using Unreal Engine 5 Nanite technology to dynamically generate virtual scenes and overlay interactive effects based on intent tags. An optical flow algorithm is used to predict user action trajectories and pre-render local frames of the scene, using bilinear interpolation to compensate for frame rate fluctuations. The virtual scene brightness is dynamically adjusted based on the light intensity value GZ from the environmental analysis subunit. Furthermore, based on the degree of occlusion, the highlight intensity of virtual object outlines is enhanced to maintain visibility. Auditory feedback is based on user intent recognition results, building a sound effect rule library and establishing a mapping relationship between intent labels and audio signals. This mapping relationship includes binding a short mechanical sound effect to the "grab" intent and matching a continuous wind sound effect to the "wave" intent. The HRTF algorithm combines user head posture data and virtual object spatial coordinates to calculate the three-dimensional sound field spatialization parameters, and dynamically reduces noise based on the environmental complexity value (FZ). Tactile feedback is based on the user's action intention and virtual interaction results, simulating tactile perception through force feedback gloves. The timing of action recognition and tactile signal generation is aligned through a timestamp synchronization mechanism, and the tactile gain is dynamically adjusted based on the occlusion degree value ZD of the environmental analysis subunit. The adaptive adjustment unit monitors system performance indicators in real time and executes a closed-loop optimization process. System performance indicators include motion recognition accuracy and interaction delay. The specific closed-loop optimization process includes: When the motion recognition accuracy or interaction delay is lower than the preset threshold, the self-diagnosis mechanism is triggered, and the system is adjusted based on the cause of the motion recognition failure determined by the analysis results of the motion failure tracing unit; if the environmental analysis index FXS is the cause of the motion recognition failure, optimization is performed based on the environmental optimization suggestions; if DZC is the cause of the motion recognition failure, the weight ty2 of the motion speed estimation DS in the interaction rule is updated through the reinforcement learning mechanism, and the contrast loss function of the intent classification model is iteratively trained based on historical abnormal log data; when the number of automatic optimization calibrations exceeds the preset calibration threshold and the motion recognition accuracy is still lower than the preset threshold, the user guidance protocol is initiated, and environmental adjustment prompts or motion correction instructions are sent to the user. At the same time, user feedback data is collected and injected into the training set, and the embedding space of the multi-scale motion descriptor is updated through the reinforcement learning mechanism.
[0019] The preferred embodiments of the present invention disclosed above are intended only to help illustrate the present invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the present invention to specific embodiments. Obviously, many modifications and variations are possible based on the contents of this specification. These embodiments are selected and described in detail in this specification to better explain the principles and practical applications of the present invention, thereby enabling those skilled in the art to better understand and utilize the present invention. The present invention is limited only by the claims and their full scope and equivalents.
Claims
1. A somatosensory interactive system based on 3D visual enhancement technology, characterized in that: include: The motion acquisition module is used to collect the coordinates of the user's skeletal joints, movement speed, three-dimensional movement trajectory and environmental source data through a multimodal sensor array; The data processing module is used to extract user action data to generate multi-level feature maps, build multi-scale descriptors, and build a human action feature model library based on labeled samples; The action recognition optimization module includes an action matching calculation unit, an analysis unit, and an action failure tracing unit; The action matching calculation unit uses a dynamic time warping algorithm to calculate the optimal path cumulative distance between the real-time action sequence and the standard template, and obtains the action matching similarity SD after normalization. If the action matching similarity SD exceeds the threshold, a virtual interaction event is triggered; otherwise, the action is automatically transmitted to the analysis unit. The analysis unit is divided into an environmental analysis subunit and an action analysis subunit; When the interaction action recognition fails, the environmental analysis subunit analyzes the environmental impact parameters of the interaction space and obtains the environmental analysis evaluation index FXS; When the interactive action recognition fails, the interactive action analysis subunit analyzes the interactive action instruction and obtains the action instruction analysis evaluation index DZC; The action failure tracing unit determines the cause of action recognition failure based on the environment analysis evaluation index FXS and the action command analysis evaluation index DZC, and implements environment optimization suggestions or action optimization suggestions; The somatosensory interaction module is used to generate interactive responses to virtual scenes, optimize the user experience through visual, auditory, and tactile feedback, and perform adaptive adjustments.
2. The somatosensory interactive system based on 3D visual enhancement technology according to claim 1, characterized in that: The specific processing process of the data processing module is as follows: Based on the 3D spatiotemporal convolutional network, the coordinates, accelerations, and posture angles of skeletal joints are extracted from point cloud data, and multi-level feature maps are output. Multi-level features are fused to construct multi-scale descriptors. Combined with labeled samples from big data, contrastive learning is used to optimize the feature embedding space, and a human motion feature model library is constructed. Based on this library, a user interaction control model is constructed.
3. The somatosensory interactive system based on 3D visual enhancement technology according to claim 1, characterized in that: The specific matching steps of the action matching similarity of the action matching calculation unit are: Load the standard action template from the constructed user interaction control model, and use the dynamic time warping algorithm to calculate the optimal path cumulative distance between the real-time action sequence and the standard template ; Will After normalization, substitute into the formula , get the action matching similarity value SD, where The maximum cumulative distance of the path passing through historical data.
4. A somatosensory interactive system based on 3D visual enhancement technology according to claim 1, characterized in that: The specific step 1 of the analysis unit analyzing the interaction environment impact parameters is: The process from when the system is initialized to recognize interactive actions is recorded as the analysis interval. The depth camera is used to collect environmental data of the interactive space within the current analysis interval and calculate the light intensity value GZ, the environmental complexity value FZ, and the volume value of the available interactive space TJ. At the same time, the number of people in the current interactive space is identified to obtain the number of people RS. Establish a three-dimensional coordinate system with the current interactive space identification center point as the reference point, calculate the relative distance JL between the user and the virtual space, and match the preset range to obtain the distance impact value YX; identify and locate the obstructions within the current analysis interval and calculate the obstruction degree value ZD; After normalizing the ambient light intensity value GZ, the environmental complexity value FZ, the available interactive space volume value TJ, and the occlusion degree value ZD, substitute them into the formula Get the environmental analysis index HJS; and These are the influence weight factors of light intensity value, environment complexity value, available interactive space volume value, and occlusion degree value; Combined with the number of people RS and the interactive user distance influence value YX, after normalization, substitute into the formula Get the comprehensive environmental analysis index ZHS; 、 and They represent the environmental analysis reference index, the reference number of people, and the reference value of the distance impact of interactive users in the current interactive environment respectively; and They are the influence weight factors of the environmental analysis index, the number of people and the impact value of the interactive user distance.
5. A somatosensory interactive system based on 3D visual enhancement technology according to claim 4, characterized in that: The specific step 2 of the analysis unit analyzing the interaction environment impact parameters is: The light uniformity index GX is calculated by using the light intensity value within the current analysis interval obtained by the light sensor; the light impact value GY is obtained by matching the preset range; the reflective surface distribution is analyzed using image processing technology to calculate the reflective impact value FG; the number of detected moving objects is marked as YD; Normalize the light impact value GY, the reflection interference value FG, and the number of moving objects YD in the current analysis interval and then enter them into the formula , get the interference environment analysis index GRS; where 、 as well as They are the light impact reference value, the reflection interference reference value and the moving object reference value. and They are the influence weight factors of light impact value, reflection interference value and number of moving objects; After normalizing the environmental comprehensive analysis index ZHS and the interference environment analysis index GRS in the current analysis interval, enter them into the formula , and obtain the environmental analysis and evaluation index FXS; and They represent the comprehensive environmental analysis reference index and interference environment analysis reference index of the current interactive environment respectively. are the impact weight factors of the comprehensive environmental analysis index and the interference environment analysis index respectively.
6. A somatosensory interactive system based on 3D visual enhancement technology according to claim 5, characterized in that: The specific steps of the analysis unit analyzing the interaction action impact parameters are as follows: Extract the duration and complexity of the interactive action command, match the preset data package to obtain the action speed estimation DS; identify the action result through the interactive user action command and compare it with the standard trajectory of the possible results to obtain the action trajectory similarity DF; based on the comparison result, after normalization, substitute it into the formula , get the action command analysis evaluation index DZC; where and They represent the number of action trajectory features for interactive action instruction recognition and the number of standard trajectory features for possible results respectively; Indicates the allowable difference between the trajectory characteristics of the interactive action instruction recognition and the standard trajectory characteristics; and They represent the allowed estimation value of motion speed and the allowed value of motion trajectory similarity respectively.
7. A somatosensory interactive system based on 3D visual enhancement technology according to claim 6, characterized in that: The specific steps of the action failure tracing unit for action failure tracing are: Based on the comparison of the environment analysis evaluation index FXS and the action command analysis evaluation index DZC with the corresponding thresholds, the cause of failure is determined and optimization is performed; if the environment analysis evaluation index is higher than the threshold and the action command analysis evaluation index is lower than the threshold, it is determined that the current interactive environment has interfered with action recognition, the environment analysis evaluation index is analyzed to identify the interference source, and environmental optimization suggestions are generated. The environmental optimization suggestions include: for occlusion interference, visually prompting the position of the occlusion and guiding the user to adjust; for reflection interference, dynamically adjusting the camera exposure parameters and gain parameters, and prompting the light source adjustment direction on the interactive interface; if the action command analysis evaluation index DZC is higher than the threshold and the environment analysis evaluation index is lower than the threshold, it is determined that there is a problem with the current user's action command, and action optimization suggestions are output. The action optimization suggestions include dynamic correction suggestions, adjustment of weight coefficients R1, R2, R3, and reference to reinforcement learning mechanisms; If the environment analysis evaluation index FXS and the motion command analysis evaluation index DZC are both higher than the corresponding thresholds, it is determined that there are interference causes in both the interaction environment and the motion command, and the environment optimization suggestions and the motion optimization suggestions are executed simultaneously.
8. According to claim 1, a somatosensory interactive system based on 3D visual enhancement technology is characterized in that: The somatosensory interaction module includes a virtual mapping unit, a multimodal feedback unit and an adaptive adjustment unit; The virtual mapping unit maps the user's skeleton coordinates to the virtual space through a homogeneous coordinate transformation matrix, and updates the position offset in real time based on the millimeter wave radar; The multimodal feedback generates visual effects, three-dimensional sound fields, and tactile signals based on the intention recognition results. The visual feedback pre-renders the scene frame based on the optical flow algorithm and dynamically adjusts the brightness and contour highlights; the auditory feedback uses the HRTF algorithm to generate spatialized sound effects and reduces noise based on the complexity of the environment; the tactile feedback synchronizes the action timing through the force feedback glove and dynamically adjusts the gain; The adaptive adjustment unit monitors the recognition action accuracy and delay indicators, triggering the self-diagnosis mechanism to optimize system parameters or guide users to make adjustments.