Immersive Arc Display Space Interaction Control Method and System

Through the multi-source fusion of depth cameras and inertial sensor data, combined with wavelet transformation and deep learning technology, it can identify and respond to users' natural actions, and solve the problem that traditional arc display interactive systems are difficult to recognize natural actions, achieving an efficient and stable immersive human-computer interactive experience.

CN119739292BActive Publication Date: 2025-06-24SHENZHEN MEIYAD OPTOELECTRONICS CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510247008.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-06-24
Estimated Expiration
2045-03-04

AI Technical Summary

Technical Problem

Traditional curved display interactive systems are difficult to accurately identify and respond to users' natural movements, resulting in the limitation of the smoothness and nature of the interactive experience.

Method used

The depth image data of the user's continuous actions is collected through the depth camera, and combined with the calculation of the action energy change value and the adaptive threshold segmentation, the effective interactive action segment is identified. Continuous wavelet transformation and multi-scale time-frequency analysis methods are used, combined with deep separation convolutional networks and bidirectional linear attention mechanisms, and action features are extracted. The dual-stage impulse response filtering algorithm combines depth image data and inertial sensor data to obtain the user's real-time spatial coordinates. Human-computer interaction state analysis is performed based on the action recognition results and real-time spatial coordinates, system state prediction data is generated, and interactive control instructions are generated through anti-perturbation model prediction control strategy.

Benefits of technology

It realizes accurate identification of effective interactive action clips, improves the overall performance and user experience of the immersive human-computer interaction system, and enhances the robustness, real-timeness and stability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119739292B_ABST
    Figure CN119739292B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of display space interaction, and discloses an immersive arc display space interaction control method and system. Among them, the method includes: collecting depth image data of continuous user actions through a depth camera set in an arc-shaped display screen, and identifying effective interaction action segments; performing local feature and global feature extraction to output an action recognition result; performing three-dimensional coordinate extraction, and performing two-stage impulse response filtering fusion with the user's inertial sensor data to output real-time spatial coordinates; performing human-computer interaction state analysis based on the action recognition result and real-time spatial coordinates to generate system state prediction data; performing anti-disturbance prediction and dynamic adjustment on the system state prediction data to generate an interaction control instruction for driving the display content of the arc-shaped display screen. The present invention realizes the accurate recognition of effective interaction action segments, and improves the overall performance and user experience of the immersive human-computer interaction system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of display space interaction, and in particular, to an immersive arc display space interaction control method and system. Background Art

[0002] With the rapid development of virtual reality and augmented reality technologies, arc-shaped display screens have been widely used in the field of human-computer interaction due to their unique immersive visual experience and wide field of view. However, traditional arc-shaped display screen interaction systems mainly rely on fixed interaction devices and preset interaction modes, making it difficult to accurately recognize and respond to natural user actions, which severely limits the fluency and naturalness of the interaction experience.

[0003] Currently, although the introduction of devices such as depth cameras and inertial sensors provides a technical basis for user action recognition, there are still many challenges in practical applications: the actions of users in the arc display space have the characteristics of diversity and non-periodicity, resulting in low accuracy of action recognition; secondly, due to the curved surface characteristics of the arc display screen, user position estimation is easily affected by occlusion and environmental interference, affecting the accuracy of interaction positioning; system response delay and control strategies are not optimized enough to ensure the real-time performance and stability of the interaction process. Summary of the Invention

[0004] The present invention provides an immersive arc display space interaction control method and system, which realizes the accurate recognition of effective interaction action segments and improves the overall performance and user experience of the immersive human-computer interaction system.

[0005] In a first aspect, the present invention provides an immersive arc display space interaction control method, and the immersive arc display space interaction control method includes:

[0006] Collecting depth image data of continuous user actions through a depth camera set in the arc display screen, and identifying effective interaction action segments;

[0007] Performing continuous wavelet transform on the effective interaction action segments to generate a multi-scale time-frequency feature map, and performing local feature and global feature extraction to output an action recognition result;

[0008] Extracting three-dimensional coordinates of human skeleton key points in the depth image data, and performing two-stage impulse response filtering fusion with the user's inertial sensor data to output the real-time spatial coordinates of the user in the coordinate system of the arc display screen;

[0009] Based on the action recognition result and the real-time spatial coordinates, performing human-computer interaction state analysis to generate system state prediction data;

[0010] Perform anti-disturbance prediction and dynamic adjustment on the system state prediction data to generate an interactive control instruction for driving the display content of the arc-shaped display screen.

[0011] In a second aspect, the present invention provides an immersive arc-shaped display space interaction control system, and the immersive arc-shaped display space interaction control system includes:

[0012] An acquisition module, configured to collect depth image data of continuous actions of a user through a depth camera disposed in the arc-shaped display screen, and identify valid interactive action segments;

[0013] A feature extraction module, configured to perform continuous wavelet transform on the valid interactive action segments to generate a multi-scale time-frequency feature map, and perform local feature and global feature extraction, and output an action recognition result;

[0014] A filtering and fusion module, configured to extract three-dimensional coordinates of human skeleton key points in the depth image data, and perform two-stage impulse response filtering and fusion with the inertial sensor data of the user, and output the real-time spatial coordinates of the user in the coordinate system of the arc-shaped display screen;

[0015] A state analysis module, configured to perform human-computer interaction state analysis based on the action recognition result and the real-time spatial coordinates to generate system state prediction data;

[0016] A generation module, configured to perform anti-disturbance prediction and dynamic adjustment on the system state prediction data to generate an interactive control instruction for driving the display content of the arc-shaped display screen.

[0017] In the technical solution provided by the present invention, the present invention collects continuous action data of users through a depth camera, combines the calculation of action energy change values and adaptive threshold segmentation, realizes the accurate recognition of effective interaction action segments, and effectively reduces the interference of invalid actions on the system performance. By adopting the continuous wavelet transform and multi-scale time-frequency analysis methods, combining the depth separable convolutional network and the bidirectional linear attention mechanism, the recognition ability of actions with different durations and complexities is improved, and the robustness of the system is enhanced. Through the dual-stage impulse response filtering algorithm, the depth image data and inertial sensor data are fused to realize the high-precision real-time position tracking of the user in the coordinate system of the curved display screen, and the accuracy of spatial positioning is improved. The state equation of human-computer interaction based on the extended state observer is designed, and through the system delay matrix and the state prediction model, the problem of system response delay is effectively solved, and the real-time performance of the interaction process is enhanced. By adopting the anti-disturbance model predictive control strategy, combining the adaptive gain matrix and the fuzzy dynamic compensation, the smooth generation and accurate execution of interaction control commands are realized, and the stability and reliability of the system are improved. The overall solution fully considers the particularity of the curved display space, and through the collaborative optimization of multi-source data fusion and multi-level control strategies, the overall performance and user experience of the immersive human-computer interaction system are significantly improved.

[0018] Other features and advantages of the present invention will be described in the following specification, and, in part, will be obvious from the specification, or will be understood by implementing the present invention. The objectives and other advantages of the present invention are achieved and obtained by the structures specifically pointed out in the specification, claims, and drawings.

[0019] To make the above objectives, features, and advantages of the present invention more obvious and understandable, the following specific preferred embodiments are given, and in conjunction with the accompanying drawings, the detailed description is as follows. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 It is a schematic diagram of an embodiment of the interactive control method for an immersive curved display space in an embodiment of the present invention;

[0021] Figure 2 It is a schematic diagram of an embodiment of the interactive control system for an immersive curved display space in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0022] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0023] As used in the embodiments of the present invention, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally further include other steps or units not listed, or may optionally further include other steps or units inherent to these processes, methods, products or devices.

[0024] For the convenience of understanding this embodiment, first, a detailed introduction is given to an immersive arc display space interaction control method disclosed in the embodiments of the present invention. As Figure 1 shown, the method includes the following steps:

[0025] 101. Collect depth image data of continuous user actions through a depth camera set in the arc-shaped display screen, and identify effective interaction action segments;

[0026] It can be understood that the execution subject of the present invention can be an immersive arc display space interaction control system, or a terminal or a server, and no specific limitation is made here. In the embodiments of the present invention, the server is taken as an example of the execution subject for illustration.

[0027] Specifically, the depth camera in the arc-shaped display screen is used to collect the continuous actions of the user in real time to obtain the initial depth image data. Bilateral filtering and morphological operations are performed on the depth image data. The bilateral filtering technology is used to effectively retain the edge details of the image while smoothing the noise in the area, realizing the preliminary optimization of the depth image. On this basis, morphological operations are introduced to perform erosion and dilation operations on the image, eliminate small-area noise points and enhance the coherence of the main body area, generating denoised depth image data. The depth values of the denoised depth image data are completed and calibrated. The missing depth values are repaired through the interpolation algorithm to ensure the integrity and accuracy of the depth information of each pixel point in the image. At the same time, aiming at the optical distortion caused by the depth camera, the calibration algorithm is used to correct the image coordinates to ensure that the generated preprocessed depth image data has high precision and consistency. On this basis, the preprocessed depth image data is processed. The human body area is extracted through segmentation technology, the background information is removed, and the skeleton features of the human body are extracted to obtain skeleton feature data, which contains the three-dimensional position information of each main joint point of the user. Based on the human body skeleton data, the three-dimensional spatial displacement of the bone joint points between adjacent frames is calculated to obtain the inter-frame action displacement data, which reflects the change trend of the user's actions. To capture the timing characteristics of the actions, the inter-frame action displacement data is processed using the time series differential operation to extract the change rate of the joint point displacement. At the same time, the change amount of the joint point angle is used as another important index, and these features are combined through polynomials to calculate the energy change value of the action. Adaptive threshold segmentation is performed on the action energy change value. By analyzing the distribution characteristics of the action energy change value, the segmentation threshold is dynamically adjusted to determine the optimal segmentation threshold. Based on this threshold, peak detection is performed on the action energy change value, and the starting point of the energy data segment that is continuously greater than the optimal segmentation threshold is marked as the action starting point. This process can accurately identify the moment when the user's action starts, thus avoiding interference caused by noise or misoperations. At the same time, sliding window analysis is performed on the action energy change value, and the end point of the data segment that is continuously less than the optimal segmentation threshold and the window length exceeds the preset duration is marked as the action end point. According to the action starting point and the action end point, the original action sequence in the human body action data is segmented to obtain each independent action segment. The similarity between the segmented actions and the predefined arc-shaped display screen interaction action templates is calculated, and it is judged whether the user's actions belong to valid interaction actions through the feature matching algorithm. The valid interaction action segments are output according to the similarity analysis results.

[0028] 102. Perform continuous wavelet transform on the valid interaction action segments to generate a multi-scale time-frequency feature map, and perform local feature and global feature extraction to output the action recognition result;

[0029] Specifically, perform spectral analysis on the effective interaction action segments to calculate the main frequency characteristics of the action signals. By performing Fourier transform on the signals, obtain their frequency distributions, and determine the main frequency range of the signals through the operation of the autocorrelation function. Based on this main frequency range, select appropriate continuous wavelet basis functions purposefully to ensure the efficient capture of signal characteristics by wavelet transform. Optimize the scale parameters of the wavelet basis functions. Introduce the criterion of minimizing information entropy, where information entropy is used as an index to measure the complexity of the action feature signals at different decomposition scales. By gradually adjusting the scale parameters of the wavelet basis functions and calculating the corresponding information entropy values, select the parameters that can minimize the information entropy to obtain the optimal decomposition scale sequence. Convolve the optimized wavelet basis functions with the effective interaction action segments to achieve wavelet transform and generate the original time-frequency coefficient matrix. Perform singular value decomposition on the original time-frequency coefficient matrix to separate the main components and noise components of the signals and construct the denoised time-frequency coefficient matrix. According to the preset frequency intervals, divide the denoised matrix into subbands, decompose the entire frequency spectrum range into several frequency bands, and extract the data corresponding to each frequency band as independent subband matrices. Perform energy normalization on the subband matrices respectively to eliminate the problem of uneven energy distribution caused by differences in frequency band ranges, making the characteristics of each frequency band more comparable in subsequent calculations to obtain the normalized time-frequency coefficients. Calculate the energy distribution characteristics of each frequency band based on the normalized time-frequency coefficients to generate the time-frequency energy distribution matrix. Process the time-frequency energy distribution matrix through multi-scale fusion technology, combine the feature information at multiple scales, and generate the multi-scale time-frequency feature map. Extract features from the multi-scale time-frequency feature map to obtain local features and global features respectively. The extraction of local features focuses on analyzing the significant changes within a short time in the map, such as the positions of energy peaks and mutation features, while global features capture the overall characteristics of the action signals by statistically analyzing the distribution patterns of the entire map, such as the degree of energy concentration and the proportion of frequency components, etc. By combining local features and global features, achieve the comprehensive recognition and classification of action signals, and finally output the action recognition results.

[0030] The multi-scale time-frequency feature map is input into the depthwise separable convolutional layer. Through pointwise convolution and depthwise convolution decomposition operations, the feature channels and spatial features are gradually separated, effectively reducing the number of model parameters and accelerating the calculation speed, while retaining the feature expression ability. After being processed by the depthwise separable convolution, an initial feature map is obtained. Perform multi-scale spatial pyramid pooling on the initial feature map. Multi-scale spatial pyramid pooling extracts local features at multiple scales by designing pooling kernels of different sizes to capture action patterns in different spatial ranges. In this process, small-scale pooling kernels can extract fine-grained local features, while large-scale pooling kernels can capture a wider feature context, forming a multi-level feature representation. Calculate channel attention for the multi-level local features to enhance the weights of important features. By calculating the importance of each channel, channel weight coefficients are generated, reflecting the contribution of a specific channel in expressing action patterns. According to the channel weight coefficients, a recalibration operation is performed on the feature channels, that is, the weights of the feature channels are adjusted, so that the feature expressions of high-contribution channels are strengthened, while the feature expressions of low-contribution channels are appropriately suppressed, obtaining enhanced local features. The enhanced local feature sequence is input into the bidirectional linear self-attention module. Through the linear projection operations of the query matrix and the key-value matrix, the attention weights of the feature sequence are constructed. The query matrix is used to represent the dependence of the current feature on other features, while the key-value matrix generates attention weights through the inner product operation to quantify the correlation between features. On this basis, the feature sequence is weighted and summed according to the attention weights to obtain the global context feature. The global context feature can capture the long-range dependencies in the feature sequence, enabling the model to not only focus on local features but also understand the overall temporal dynamics of the action pattern. The global context feature and the enhanced local feature are fused into a hybrid feature through a residual connection. Perform feature channel recombination on the hybrid feature. By introducing an adaptive weight network, the importance of each channel's feature is recalculated to obtain the importance coefficient of the feature. Based on the importance coefficient, the hybrid feature is weighted and fused to generate a weighted fusion feature. The weighted fusion feature is input into the fully connected classifier, and the probability distribution of each action category is calculated through the softmax function. The softmax function maps the weighted features to a probability space, so that the probability of each category can intuitively reflect the confidence value of the model for that action category. In this way, the action recognition result is finally output to achieve accurate prediction of action categories.

[0031] 103. Extract the three-dimensional coordinates of the human skeleton key points in the depth image data, and perform two-stage impulse response filtering fusion with the user's inertial sensor data, and output the real-time spatial coordinates of the user in the curved display screen coordinate system;

[0032] Specifically, for human body region segmentation based on depth image data, the human body contour is extracted through an image segmentation algorithm to generate a binary mask representing the human body region, ensuring that subsequent processing only operates on the target region. The human body region mask is subjected to skeleton thinning processing, and the thinning algorithm is used to simplify the human body region into a skeleton contour map with a single-pixel width. Based on the skeleton contour map, key points of the joint region are located by combining human anatomical features, and two-dimensional skeleton key points representing human motion features are extracted. These key points include the main joint positions, such as the head, shoulders, elbows, knees, and ankles, etc., constituting the skeleton framework of the user's actions. According to the extracted two-dimensional skeleton key points and depth image data, the three-dimensional coordinates are calculated through projective transformation. Combining the depth data, the two-dimensional plane coordinates are mapped to the three-dimensional space to generate the initial three-dimensional coordinate data of the skeleton key points. Kinematic constraints are applied to optimize the initial coordinate data. By constructing a kinematic constraint model of human joints, such as joint range of motion and connection length consistency, etc., the optimized three-dimensional coordinate data is more in line with the actual kinematic characteristics. At the same time, the data of the inertial sensors worn by the user are preprocessed to ensure the reliability and accuracy of the data. Zero-drift compensation is performed on the acceleration and angular velocity data collected by the sensors to eliminate the errors caused by sensor drift. At the same time, temperature calibration is carried out to correct the influence of temperature changes on the data. The calibrated acceleration and angular velocity data are integrated to obtain the relative displacement data of the user. To analyze the user's pose information, a pose solution algorithm is used to process the relative displacement data, and the pose angle data of the user is calculated through quaternion operations. On this basis, the coordinate system of the relative displacement data is converted according to the pose angle data, and the displacement data in the local coordinate system is converted into global displacement data, thereby reflecting the position change of the user in the entire interaction space. The three-dimensional coordinate data and the global displacement data are input into the first-stage finite impulse response filter. The first-stage filter smooths the input data through weighted average operations, eliminates high-frequency noise, and retains the motion trend, outputting the initial fusion data. The initial fusion data is input into the second-stage finite impulse response filter for depth fusion processing. The second-stage filter adopts the principle of complementary filtering, and through the complementary characteristics in the frequency domain, it realizes the dynamic weight adjustment of the inertial sensor data and the three-dimensional coordinate data. The low-frequency part is mainly provided by the three-dimensional coordinate data of the depth camera to ensure the global stability of the spatial position, while the high-frequency part is provided by the inertial sensor data to capture the rapid changes of the actions. After complementary filtering operations, the fused real-time spatial coordinates of the user are output. Through the above steps, the real-time spatial coordinates of the user in the coordinate system of the arc-shaped display screen are obtained, reflecting the position and pose of the user.

[0033] 104. Based on the action recognition result and the real-time spatial coordinates, perform human-computer interaction state analysis to generate system state prediction data;

[0034] Specifically, the action recognition results are temporally aligned with the real-time spatial coordinates of the user in the coordinate system of the arc-shaped display to ensure that the action recognition results are consistent with the spatial coordinates in the time dimension, generating the basic data of the system state. After obtaining the basic data of the system state, the operation delay factors of the system itself are considered. The arc-shaped display is affected by data acquisition delay, data transmission delay, and display refresh delay during the interaction process, and these delays are quantitatively calculated. By measuring the response time of each module, a system delay matrix is constructed, including the input delay of the acquisition device, the data transmission delay in network communication, and the refresh delay of the display screen hardware, etc. The delay quantization results are combined with the basic data of the system state for parameter combination operations to establish a state space expression, comprehensively describing the influence of user actions, spatial positions, and system delays on the current interaction state. The state space expression is expanded by Taylor series and linearly approximated to simplify the complex non-linear relationship. Through Taylor expansion, the state space expression is decomposed into computable linear terms and high-order non-linear terms, and the linear part is retained for approximation, ignoring the minor influence of the high-order terms on the system state. The generated human-computer interaction state equation describes the dynamic characteristics of the system in a linear form, facilitating subsequent optimization and solution. The human-computer interaction state equation is subjected to state expansion processing to expand the state variables to enhance the system's adaptability to external disturbances. The expanded state equation is input into the disturbance observation module, and by dynamically monitoring the changes in external disturbance factors, the state expression of the system is corrected. The disturbance observation module can identify the influence of the external environment on the interaction system, such as sudden changes in user behavior or short-term fluctuations in device performance, and integrate these disturbance quantities into the expanded state variables to improve the robustness of the model. To achieve the predictive analysis of the system state, the expanded state variables are substituted into the Gaussian probability distribution model for probability calculation and the generation of state probability distribution parameters. Through the Gaussian distribution, the occurrence probabilities of different states are quantified, and based on these probability values, the future state of the system is inferred. On this basis, combined with the historical state sequence, multiple sets of predicted state values are calculated through a recursive prediction algorithm. The recursive prediction algorithm iterates continuously in the time dimension, using the current state information and historical state characteristics to gradually update the future state, forming a set of multi-step prediction results. Error correction is performed on multiple sets of predicted state values to improve the accuracy of the prediction results. By comparing the deviation between the predicted value and the historical real data, the model parameters are adjusted to correct the calculation error and output more accurate system state prediction data.

[0035] 105. Anti-disturbance prediction and dynamic adjustment are performed on the system state prediction data to generate interaction control instructions for driving the display content of the arc-shaped display.

[0036] Specifically, the system state prediction data is input into the disturbance decoupling module, and the data is decomposed using non - linear separation operations to obtain the internal state sequence and the external disturbance sequence. The internal state sequence represents the dynamic change trend of the system itself, while the external disturbance sequence reflects the influence of external factors on the system. The rolling - horizon optimization calculations are performed on the internal state sequence and the external disturbance sequence respectively. Based on the rolling - optimization idea, a predictive - control cost function is constructed to measure the advantages and disadvantages of different control strategies in the future time domain. The design of the predictive - control cost function needs to comprehensively consider factors such as system performance indicators, energy consumption, and interaction fluency. The quadratic - programming method is used to solve the predictive - control cost function, generating a set of optimized control - input sequences. These control - input sequences ensure that the system can always maintain the best control performance during the dynamic change process by minimizing the cost function. Constraint processing is performed on the optimized control - input sequence. The goal of constraint processing is to ensure that the control - input sequence conforms to the physical and safety boundaries of the actual system. For example, the display range, content - switching speed, and dynamic effects of the arc - shaped display screen all need to meet the preset physical limitations. At the same time, real - time feedback data is introduced into the constraint - processing process, and the control boundaries are corrected in real - time according to the current operating state of the system, generating initial control commands that meet the actual requirements. The initial control commands are input into the adaptive - gain matrix for force allocation. The role of the adaptive - gain matrix is to dynamically adjust the components of the control commands in multiple dimensions according to the requirements of different interaction tasks. For example, in a scenario involving simultaneous interaction of multiple users, different interaction priorities and display resources are allocated to each user, thus achieving reasonable control allocation as a whole. After generating the multi - dimensional control quantities, in order to improve the accuracy and stability of control, fuzzy - dynamic compensation calculations are performed on the multi - dimensional control quantities and the real - time feedback data. Uncertainty factors, such as sudden changes in user actions or interference from the external environment, are processed through a fuzzy - logic system to generate compensation control quantities. The compensation control quantities are mapped and converted into a smooth control sequence to eliminate the mutation points in the control commands and avoid jitter or incoherence during the update of the display content. By optimizing the smoothness of the control signal, the user's interaction experience can be significantly improved. The smooth control sequence is converted into the display parameters of the arc - shaped display screen, including specific display attributes such as content position, size, color, and dynamic special effects. These display parameters are transmitted to the arc - shaped display - screen driving module through an interface, and finally, interaction commands for controlling the display content are output.

[0037] In the embodiments of the present invention, the present invention collects continuous action data of users through a depth camera, and combines the calculation of action energy change values and adaptive threshold segmentation to achieve accurate recognition of effective interaction action segments, effectively reducing the interference of invalid actions on the system performance. By using the continuous wavelet transform and multi-scale time-frequency analysis methods, combined with the depthwise separable convolutional network and the bidirectional linear attention mechanism, the recognition ability of actions with different durations and complexities is improved, and the robustness of the system is enhanced. Through the dual-stage impulse response filtering algorithm, the depth image data and inertial sensor data are fused to achieve high-precision real-time position tracking of the user in the coordinate system of the curved display screen, improving the accuracy of spatial positioning. A human-computer interaction state equation based on an extended state observer is designed, and through the system delay matrix and the state prediction model, the problem of system response delay is effectively solved, enhancing the real-time performance of the interaction process. By adopting the anti-disturbance model predictive control strategy, combined with the adaptive gain matrix and fuzzy dynamic compensation, the smooth generation and accurate execution of interaction control commands are realized, improving the stability and reliability of the system. The overall solution fully considers the particularity of the curved display space, and through the collaborative optimization of multi-source data fusion and multi-level control strategies, significantly improves the overall performance and user experience of the immersive human-computer interaction system.

[0038] In a specific embodiment, the process of executing step 101 may specifically include the following steps:

[0039] Collect the depth image data of the continuous actions of the user through the depth camera set in the curved display screen, perform bilateral filtering and morphological operations on the depth image data to obtain denoised depth image data;

[0040] Perform depth value completion and calibration on the denoised depth image data to obtain preprocessed depth image data, and perform human body region segmentation and skeleton feature extraction on the preprocessed depth image data to obtain human action data;

[0041] Perform three-dimensional space displacement calculation on the bone joint points of adjacent frames in the human action data to obtain inter-frame action displacement data, perform time series differential operation on the inter-frame action displacement data, and perform polynomial combination with the joint point angle change amount to obtain action energy change values;

[0042] Perform adaptive threshold segmentation on the action energy change values to obtain the optimal segmentation threshold, and perform peak detection on the action energy change values according to the optimal segmentation threshold, and mark the starting point of the data segment continuously greater than the optimal segmentation threshold as the action starting point;

[0043] Perform sliding window analysis on the action energy change values, and mark the end point of the data segment continuously less than the optimal segmentation threshold and with a window length greater than the preset duration as the action end point;

[0044] Segment the original action sequence in the human action data according to the action start point and the action end point, calculate the similarity with the predefined interactive action template of the curved display screen, and output the effective interactive action segments.

[0045] Specifically, the depth camera continuously acquires the depth image data of the user's actions in the form of frames, and each frame of data contains the depth value of each pixel point in the scene , where and represent the horizontal and vertical coordinates of the pixel in the image, represents the distance from this pixel point to the depth camera. The original depth image data is affected by noise and environmental interference, and noise reduction processing is performed to extract effective information. In order to remove the random noise in the depth image while retaining the edge details, bilateral filtering is applied to . The formula for bilateral filtering is:

[0046] ;

[0047] where is the filtered depth value, represents the window area, is the spatial distance weight function, which depends on the distance between pixels; is the weight function of the pixel depth difference, which depends on the difference in depth values, is the normalization factor. Bilateral filtering can effectively remove the noise in the flat area while keeping the edges clear. Morphological operations are performed on the filtered depth data, such as erosion and dilation operations, to perform structural adjustment. The erosion operation removes small isolated pixels, and the dilation operation fills the hole areas to generate a smoother noise-reduced depth image data. Depth value completion and calibration are performed on the noise-reduced depth image data. The purpose of completion is to repair the invalid areas caused by occlusion or reflection in the depth sensor, which is achieved through the interpolation method. The specific formula is:

[0048] ;

[0049] where is the completed depth value, is the spatial weight window, is the interpolation weight, which is calculated based on the distance between pixels. The calibration process converts the depth data to the world coordinate system according to the internal parameter matrix and external parameter matrix of the camera to generate the preprocessed depth image data. Human body region segmentation is performed on the preprocessed depth image data. Through the depth threshold method or the semantic segmentation network, the human body region is separated from the background to generate a human body mask. A skeleton extraction algorithm, such as the method based on the medial axis transformation, is applied to the human body mask to generate skeleton feature data. Based on the human anatomy model, the positions of the joint key points are identified in the skeleton feature data , where is the three-dimensional coordinates of the th joint, , , respectively represent its position in the spatial coordinate system. For the extracted human motion data, calculate the three-dimensional displacement of the bone key points between adjacent frames. For the displacement between the th frame and the th frame, it is expressed by the formula:

[0050] ;

[0051] These displacement data reflect the dynamic changes of the user's actions. Perform a temporal differential operation on the inter-frame action displacement data to capture the motion change trend, and at the same time combine the joint angle change to generate the action energy change value using polynomial combination:

[0052] ;

[0053] where is the action energy of the th frame, and are the weight factors of the joint points, is the total number of joint points. For the action energy change value , perform adaptive threshold segmentation. By analyzing the dynamic distribution of the energy values, determine the optimal segmentation threshold , and the specific calculation is:

[0054] ;

[0055] where is the mean value of the energy, is the standard deviation. Use the optimal threshold to detect the peaks of the energy values, and mark the starting point of the segment continuously greater than as the action starting point. At the same time, perform a sliding window analysis on . The end point of the segment with a window length greater than the preset duration and an energy value less than is marked as the action end point. Through the action starting point and end point, segment the original action sequence of the human motion data. Calculate the similarity between the segmented actions and the predefined arc display interaction action templates. For example, use the dynamic time warping algorithm to calculate the similarity between the template and the segmented action :

[0056] ;

[0057] where is the difference value between the two sequences. When the similarity exceeds the set threshold, this segment is marked as a valid interaction action segment and output.

[0058] In a specific embodiment, the process of executing step 102 may specifically include the following steps:

[0059] Perform spectral analysis on the valid interaction action segment, calculate the main frequency characteristics of the action signal, obtain the main frequency range of the signal through autocorrelation function operation, and select the wavelet basis function according to the main frequency range of the signal;

[0060] Optimize the scale parameter of the wavelet basis function, perform multi-scale analysis on the action characteristics through the minimum information entropy criterion, and obtain the optimal decomposition scale sequence;

[0061] Perform convolution operation on the valid interaction action segment and the wavelet basis function to obtain the original time-frequency coefficient matrix, and perform singular value decomposition on the original time-frequency coefficient matrix to construct a noise reduction matrix;

[0062] Divide the noise reduction matrix into sub-band matrices according to the preset frequency interval, obtain multiple sub-band matrices, and perform energy normalization processing on the multiple sub-band matrices respectively to obtain the normalized time-frequency coefficients;

[0063] Calculate the energy distribution characteristics of each frequency band according to the normalized time-frequency coefficients, obtain the time-frequency energy distribution matrix, and perform multi-scale fusion on the time-frequency energy distribution matrix to obtain the multi-scale time-frequency feature map;

[0064] Extract local and global features from the multi-scale time-frequency feature map, and output the action recognition result.

[0065] Specifically, for the valid interaction action segment Perform Fourier transform to obtain the main frequency characteristics of the signal in the frequency domain. The Fourier transform formula is:

[0066] ;

[0067] where is the spectral representation of the signal , is the frequency, is the imaginary unit. By analyzing 's amplitude spectrum, find the main frequency of the signal, that is, the frequency component with the largest amplitude in the spectrum. Use the autocorrelation function to further determine the main frequency range of the signal. The autocorrelation function is defined as:

[0068] ;

[0069] where is the signal At the delay time The correlation under. By searching for The periodic peak of, estimate the main frequency range of the signal . According to the main frequency range of the signal, select the wavelet basis function that matches it, such as the Morlet wavelet, which is defined as:

[0070] ;

[0071] where is the wavelet basis function, is the center frequency. Selecting an appropriate can ensure a good representation of the signal characteristics by the wavelet basis function. After selecting the wavelet basis function, optimize its scale parameter to ensure that the main characteristics of the signal can be captured in the multi-scale analysis. Optimize through the information entropy minimization criterion, and the information entropy is defined as:

[0072] ;

[0073] where is the energy distribution ratio of the signal at the scale. By adjusting the scale parameter of the wavelet transform, find the parameter combination that minimizes to obtain the optimal decomposition scale sequence. Convolve the effective interaction action segment with the selected wavelet basis function to calculate the time-frequency coefficient matrix :

[0074] ;

[0075] where is the scale parameter, is the translation parameter, is the complex conjugate form of the wavelet basis function. The time-frequency coefficient matrix generated by the convolution operation describes the joint distribution of the signal in time and frequency, and is called the original time-frequency coefficient matrix. In order to eliminate the influence of noise on the analysis, perform singular value decomposition (SVD) denoising on the original time-frequency coefficient matrix. The formula of SVD is:

[0076] ;

[0077] where and are orthogonal matrices, is a diagonal matrix containing the singular values of the matrix . By truncating the smaller singular values, construct the denoised matrix to effectively improve the signal-to-noise ratio of the time-frequency characteristics. For the denoised matrix Perform sub-band division according to a preset frequency range. For example, divide the frequency range into low-frequency, medium-frequency, and high-frequency bands to generate multiple frequency sub-matrices. For each sub-matrix perform energy normalization processing and calculate the normalized time-frequency coefficients:

[0078] ;

[0079] where is the normalization coefficient matrix of the th sub-band. According to the normalized time-frequency coefficients calculate the energy distribution characteristics of each frequency band to generate a time-frequency energy distribution matrix :

[0080] ;

[0081] By performing multi-scale fusion on weightedly combine the energy characteristics of different frequency bands to generate a multi-scale time-frequency feature map. After obtaining the multi-scale time-frequency feature map, use a deep learning network to extract local and global features and output the action recognition result. Local feature extraction targets significant changes within a short-time window, such as the position of energy peaks and amplitude mutations, while global feature extraction analyzes the characteristics of action signals from a global perspective by statistically analyzing the distribution patterns of the entire map, such as frequency concentration and energy trajectory trends.

[0082] In a specific embodiment, the process of performing local and global feature extraction on the multi-scale time-frequency feature map and outputting the action recognition result may specifically include the following steps:

[0083] Input the multi-scale time-frequency feature map into a depthwise separable convolutional layer, and use pointwise convolution and depthwise convolution decomposition operations to obtain an initial feature map;

[0084] Perform multi-scale spatial pyramid pooling on the initial feature map to extract multi-level local features through pooling kernels of different scales;

[0085] Perform channel attention calculation on the multi-level local features to obtain channel weight coefficients, and recalibrate the feature channels according to the channel weight coefficients to obtain enhanced local features;

[0086] Input the enhanced local feature sequence into a bidirectional linear self-attention module, and construct attention weights through the linear projection operation of the query matrix and the key-value matrix;

[0087] Weightedly sum the feature sequence according to the attention weights to obtain global context features, and fuse them with the enhanced local features through residual connection to obtain mixed features;

[0088] Recombine the feature channels of the mixed features, calculate the importance coefficients of each feature through the adaptive weight network, obtain the weighted fusion features, and input the weighted fusion features into the fully connected classifier. Calculate the probability distribution of each action category through the softmax function to obtain the action recognition result.

[0089] Specifically, input the multi-scale time-frequency feature map into the depthwise separable convolutional layer. The depthwise separable convolution is divided into two parts: pointwise convolution and depthwise convolution. Its purpose is to reduce the computational complexity while retaining the feature extraction ability. In pointwise convolution, the features of each channel are operated on by a convolution kernel, and the calculation formula is:

[0090] ;

[0091] where is the output feature of channel , is the feature map of the input channel , is 's convolution weight, is the number of channels of the input feature. Depthwise convolution operates independently on each channel, and captures local spatial features through a convolution kernel, and the formula is:

[0092] ;

[0093] where is the range of the convolution kernel, is the weight of the depthwise convolution kernel. Through this separation operation, the initial feature map is generated, which not only reduces the number of parameters but also retains the time-frequency characteristics of the map. Perform multi-scale spatial pyramid pooling on the initial feature map to capture multi-level local features. In pyramid pooling, select pooling kernels of different sizes and perform block pooling operations on the feature map. The pooling result is represented by the formula:

[0094] ;

[0095] where is the size of the pooling area, is the local feature representation after pooling. Through multi-scale pooling, features are extracted at different spatial scales to generate multi-level local feature representations. After multi-level local feature extraction, to enhance the representation ability of the features, channel attention calculation is performed on them. By calculating the importance weight of each channel , recalibrate the features of different channels. The channel attention calculation formula is:

[0096] ;

[0097] where is the global feature aggregation value of channel , and are weight matrices, is the Sigmoid activation function. By calculating the channel weight coefficient , re-weight the feature channels to generate enhanced local features :

[0098] ;

[0099] Input the enhanced local feature sequence into the bidirectional linear self-attention module, which generates attention weights through the linear projection operations of the query matrix , the key matrix and the value matrix . The specific calculation is:

[0100] ;

[0101] where , , , , , are projection matrices, is the dimensionality scaling factor of the key. Through weighted summation of the value matrix , generate the global context feature:

[0102] ;

[0103] The global context feature is fused with the enhanced local feature through a residual connection to generate the hybrid feature :

[0104] ;

[0105] Input the hybrid feature into the adaptive weight network for feature channel recombination. The network calculates the importance coefficient of each feature channel:

[0106] ;

[0107] where is the weight matrix. By applying to , the weighted fusion feature is obtained:

[0108] ;

[0109] Input the weighted fusion feature into the fully connected classifier, and the classifier calculates the probability distribution of each action category through the Softmax function:

[0110] ;

[0111] where is the probability of category , is the prediction score of category , and is the total number of categories. The action recognition result is obtained.

[0112] In a specific embodiment, the process of executing step 103 may specifically include the following steps:

[0113] Based on the depth image data, perform human body region segmentation, extract the human body contour, and obtain the human body region mask;

[0114] Perform skeleton thinning on the human body region mask to obtain the skeleton contour map, and perform key point positioning on the skeleton contour map according to human anatomical features to obtain two-dimensional skeleton key points;

[0115] According to the two-dimensional skeleton key points and the depth image data, perform projection transformation to obtain the initial coordinate data, and perform kinematic constraint optimization on the initial coordinate data to obtain the three-dimensional coordinate data;

[0116] Perform zero drift compensation and temperature calibration on the user's inertial sensor data to obtain the calibrated acceleration and angular velocity data, and perform integral operation on the calibrated acceleration and angular velocity data to obtain the relative displacement data;

[0117] Perform attitude solution on the relative displacement data, obtain the attitude angle data through quaternion operation, and perform coordinate system conversion on the relative displacement data according to the attitude angle data to obtain the global displacement data;

[0118] Input the three-dimensional coordinate data and the global displacement data into the first-stage finite impulse response filter, and obtain the initial fusion data through weighted average operation;

[0119] Input the initial fusion data into the second-stage finite impulse response filter, and output the real-time spatial coordinates of the user in the arc-shaped display screen coordinate system through complementary filtering operation.

[0120] Specifically, based on the depth image data, human body region segmentation is performed, and the depth information of the human body region is extracted through the threshold segmentation method. The depth image data is represented as , where represents the pixel coordinates, represents the depth value of the pixel. By setting the threshold range , the pixel points that meet the conditions are extracted to generate a binary human body region mask :

[0121] ;

[0122] The generated mask is subjected to skeleton thinning processing. The medial axis transformation method is used to perform iterative erosion on , while retaining the topological structure of the region, to generate a skeleton contour map . The goal of skeleton thinning is to reduce the multi-pixel-wide human body region to a single-pixel skeleton structure while retaining the main information of the joints and limbs. Based on the skeleton contour map , key point localization is performed according to human anatomical characteristics. By detecting specific geometric features of the skeleton, such as endpoints and bifurcation points, the positions of the main joints of the human body are determined. Suppose the detected joint points are , where represents the joint point number, and the two-dimensional coordinates of these joint points constitute the skeleton key points of the user. According to the extracted two-dimensional skeleton key points and the depth image data , they are mapped into the three-dimensional space through projection transformation. The internal parameter matrix of the depth camera is , and the specific form is:

[0123] ;

[0124] where is the focal length, is the optical center coordinate. The calculation formula for the three-dimensional point is:

[0125] ;

[0126] Through the above formula, the initial three-dimensional coordinates of each joint point are obtained. To ensure the accuracy of the three-dimensional coordinates, kinematic constraint optimization is performed on them. By setting the length constraint between human joints (where is the predefined length of joints and ), the joint point coordinates are adjusted using a non-linear optimization algorithm to generate optimized three-dimensional coordinate data that conforms to the human body structure. At the same time, for the acceleration and angular velocity collected by the user's inertial sensor Zero-drift compensation and temperature calibration are performed on the data. Through the dynamic bias estimation model , the calibrated acceleration and angular velocity are:

[0127] ;

[0128] where and are the bias values of acceleration and angular velocity respectively. Integral operation is performed on the calibrated data to obtain relative displacement data:

[0129] ;

[0130] where represents relative displacement. The attitude angle data is calculated based on quaternion using the attitude solution method. Let the quaternion be , and the attitude update formula is:

[0131] ;

[0132] where represents quaternion multiplication. By numerically integrating and updating , the user's attitude angle is calculated, and the relative displacement data is transformed into the global coordinate system according to the attitude angle to generate global displacement data . The optimized three-dimensional coordinate data and the global displacement data are input into the first-stage finite impulse response (FIR) filter for preliminary fusion. The output of the FIR filter is:

[0133] ;

[0134] where is the filter weight, is the delayed input data. Through weighted average operation, initial fusion data is generated. The initial fusion data is input into the second-stage FIR filter, and the complementary filtering algorithm is used to fuse the low-frequency three-dimensional coordinate data and the high-frequency inertial sensor data. The formula is:

[0135] ;

[0136] where is the weight factor, and represent the low-frequency and high-frequency components respectively. After two-stage fusion filtering, the real-time spatial coordinates of the user in the arc-shaped display screen coordinate system are output.

[0137] In a specific embodiment, the process of executing step 104 may specifically include the following steps:

[0138] Align the action recognition results with the real-time spatial coordinates of the user in the coordinate system of the arc-shaped display in time series to obtain the basic system state data;

[0139] Quantitatively calculate the data acquisition delay, data transmission delay, and display refresh delay of the arc-shaped display to obtain the system delay matrix, and perform parameter combination operations on the basic system state data and the system delay matrix to obtain the state space expression;

[0140] Perform Taylor series expansion and linear approximation calculation on the state space expression to obtain the human-computer interaction state equation;

[0141] Perform state expansion processing on the human-computer interaction state equation and input it into the disturbance observation module to obtain the expanded state variables;

[0142] Substitute the expanded state variables into the Gaussian probability distribution model for probability calculation to obtain the state probability distribution parameters, and perform recursive prediction operations based on the state probability distribution parameters and the historical state sequence to obtain multiple groups of predicted state values;

[0143] Perform error correction on multiple groups of predicted state values and output the system state prediction data.

[0144] Specifically, for the action recognition results and the real-time spatial coordinates of the user perform time series alignment. Through the linear interpolation method, map the two sets of data to the unified time axis . Assume has timestamps and , has timestamps and , then the aligned data is expressed as:

[0145] ;

[0146] ;

[0147] where and are the action recognition results and spatial coordinate data after alignment. These aligned data together constitute the basic system state data . Quantify the data acquisition delay , data transmission delay and display refresh delay of the arc-shaped display. Measure these delays through time synchronization tools or the round-trip delay of test data packets, and construct the system delay matrix :

[0148] ;

[0149] System status basic data Combined with the delay matrix to generate a state space expression through parameter combination operations :

[0150] ;

[0151] where is the state transition matrix, describing the dynamic characteristics of the system; is the input matrix, is the input signal (such as the user's interaction action). Since the system state expression has non - linear characteristics, for simplifying subsequent calculations, a Taylor series expansion is used to linearly approximate it. For the non - linear function , expanded near the point as:

[0152] ;

[0153] where is 's partial derivative matrix. On this basis, the human - computer interaction state equation is obtained:

[0154] ;

[0155] where is the linearized state transition matrix, is the perturbation matrix, is the random perturbation of the system. To improve the adaptability of the model to complex environments, state expansion processing is performed on the human - computer interaction state equation, adding an expanded state variable , where represents the incremental term of the state, used to capture the rapid changes of the system. The expanded state input perturbation observation module improves the robustness of the system to external disturbances by dynamically estimating 's change characteristics. Substitute the expanded state variable into the Gaussian probability distribution model for probability calculation. Assume obeys a multi - dimensional Gaussian distribution, and its probability density function is:

[0156] ;

[0157] where is the mean vector, is the covariance matrix, is the dimension of the variable. By maximizing this probability density function, the state probability distribution parameters and . Combine the state probability distribution parameters and the historical state sequence , and use a recursive prediction algorithm to generate multiple sets of predicted state values. The formula for recursive prediction is:

[0158] ;

[0159] ;

[0160] where is the predicted state value, is the prediction uncertainty, is the process noise covariance matrix. Error correction is performed on the multiple sets of state values obtained by prediction. By comparing the deviation between the actual state and the predicted state , the model parameters are adjusted to reduce the prediction error. The corrected system state prediction data is output, providing support for the real-time interactive control of the arc-shaped display screen.

[0161] In a specific embodiment, the process of executing step 105 may specifically include the following steps:

[0162] Input the system state prediction data into the disturbance decoupling module, and obtain the internal state sequence and the external disturbance sequence through non-linear separation operation;

[0163] Perform rolling horizon optimization calculation on the internal state sequence and the external disturbance sequence to obtain the predicted control cost function, and solve the predicted control cost function according to the quadratic programming method to obtain the optimized control input sequence;

[0164] Perform constraint processing on the optimized control input sequence, and combine the real-time feedback data to perform display boundary correction to obtain the initial control instruction;

[0165] Input the initial control instruction into the adaptive gain matrix for force control distribution to obtain the multi-dimensional control quantity, and perform fuzzy dynamic compensation calculation on the multi-dimensional control quantity and the feedback data to obtain the compensated control quantity;

[0166] Perform mapping transformation on the compensated control quantity to obtain the smooth control sequence, and convert the smooth control sequence into the arc-shaped display screen display parameters, and output the interactive control instruction for driving the display content of the arc-shaped display screen.

[0167] Specifically, input the system state prediction data into the disturbance decoupling module, and decompose the prediction data into the internal state sequence and the external disturbance sequence through non-linear separation operation. The internal state sequence represents the core dynamic characteristics of the system, while the external disturbance sequence It reflects the influence of the external environment on the system. The non-linear separation operation is achieved by minimizing the following objective function:

[0168] ;

[0169] where is the non-linear mapping function of the system, is the second norm, which is used to measure the decomposition error. Through iterative optimization, the system prediction data is separated into and . For and , the moving horizon optimization calculation is carried out to predict the control performance in the future for a period of time. The core of the moving horizon optimization lies in constructing the predictive control cost function , and its form is:

[0170] ;

[0171] where is the desired system state, is the control input, and are the weight matrices of the state deviation and the control cost respectively, represents the weighted quadratic norm. By solving using the quadratic programming method, the optimized control input sequence is obtained, where:

[0172] ;

[0173] After obtaining the optimized control input sequence, constraint processing is carried out on it to ensure that the generated control instructions conform to the physical boundaries of the actual system. For example, for the display range of the curved display screen, the control input needs to satisfy the following constraints:

[0174] ;

[0175] where and are the upper and lower bounds of the control input. By adding the constraint conditions to the optimization problem, the initial control instructions that meet the physical limitations are generated. The initial control instructions are corrected for the display boundary in combination with the real-time feedback data. The real-time feedback data contains the current display state of the system. By comparing and , the initial control instructions are dynamically adjusted to make them smoother in actual operation. The corrected control instructions are expressed as:

[0176] ;

[0177] wherein is the feedback gain matrix. The corrected control command is input into the adaptive gain matrix to perform multi-dimensional distribution of the control force. The adaptive gain matrix dynamically adjusts the weights of the control components according to the task requirements to generate a multi-dimensional control quantity :

[0178] ;

[0179] wherein elements are adaptively updated according to the state of the system and the task priority. After generating the multi-dimensional control quantity, in order to improve the robustness of the system, for and the feedback data perform fuzzy dynamic compensation calculation. The goal of the compensation is to handle the uncertainties caused by non-linear effects or environmental disturbances. The fuzzy compensator converts the input error into a compensation control quantity :

[0180] ;

[0181] wherein is the fuzzy membership function, is the output value of the compensation rule. The compensation control quantity is subjected to mapping transformation to eliminate mutations and ensure the smoothness of the output. The mapping transformation is implemented through a smoothing function For example:

[0182] ;

[0183] wherein is a common Sigmoid function used to smooth the control quantity. The smoothed control sequence is converted into the display parameters of the arc-shaped display screen, such as the position information, size, color, or special effects of the display content. These display parameters are transmitted to the display device through the drive interface to generate the real-time feedback required for user interaction.

[0184] The above describes the immersive arc-shaped display space interaction control method in the embodiments of the present invention. Next, the immersive arc-shaped display space interaction control system in the embodiments of the present invention will be described. Please refer to Figure 2 One embodiment of the immersive arc-shaped display space interaction control system in the embodiments of the present invention includes:

[0185] An acquisition module 201, configured to collect depth image data of continuous actions of a user through a depth camera disposed in the arc-shaped display screen and identify valid interaction action segments;

[0186] The feature extraction module 202 is used to perform continuous wavelet transform on the effective interaction action segments, generate multi-scale time-frequency feature maps, and perform local feature and global feature extraction, and output the action recognition result;

[0187] The filtering and fusion module 203 is used to extract the three-dimensional coordinates of the human skeleton key points in the depth image data, and perform two-stage impulse response filtering and fusion with the user's inertial sensor data, and output the real-time spatial coordinates of the user in the curved display screen coordinate system;

[0188] The state analysis module 204 is used to perform human-computer interaction state analysis based on the action recognition result and the real-time spatial coordinates, and generate system state prediction data;

[0189] The generation module 205 is used to perform anti-disturbance prediction and dynamic adjustment on the system state prediction data, and generate an interaction control instruction for driving the display content of the curved display screen.

[0190] Through the collaborative cooperation of the above-mentioned various components, the present invention collects the continuous action data of the user through a depth camera, combines the calculation of the action energy change value and the adaptive threshold segmentation, realizes the accurate recognition of the effective interaction action segments, and effectively reduces the interference of invalid actions on the system performance. By using the continuous wavelet transform and multi-scale time-frequency analysis methods, combined with the depth separable convolutional network and the bidirectional linear attention mechanism, the recognition ability of actions with different durations and complexities is improved, and the robustness of the system is enhanced. Through the two-stage impulse response filtering algorithm to fuse the depth image data and the inertial sensor data, the high-precision real-time position tracking of the user in the curved display screen coordinate system is realized, and the accuracy of spatial positioning is improved. The human-computer interaction state equation based on the extended state observer is designed, and through the system delay matrix and the state prediction model, the system response delay problem is effectively solved, and the real-time performance of the interaction process is enhanced. By using the anti-disturbance model predictive control strategy, combined with the adaptive gain matrix and the fuzzy dynamic compensation, the smooth generation and accurate execution of the interaction control instruction are realized, and the stability and reliability of the system are improved. The overall solution fully considers the particularity of the curved display space, and through the collaborative optimization of multi-source data fusion and multi-level control strategies, the overall performance and user experience of the immersive human-computer interaction system are significantly improved.

[0191] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the above-described system, system and unit can refer to the corresponding processes in the foregoing method embodiments, and will not be described herein again.

[0192] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0193] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. An immersive arc display space interactive control method, characterized in that: The method comprises: The depth image data of the user's continuous actions are collected by a depth camera arranged in the curved display screen, and effective interactive action fragments are identified; specifically, the depth image data of the user's continuous actions are collected by a depth camera arranged in the curved display screen, and the depth image data are subjected to bilateral filtering and morphological operations to obtain denoised depth image data; the denoised depth image data are subjected to depth value completion and calibration to obtain pre-processed depth image data, and the pre-processed depth image data are subjected to human body region segmentation and skeleton feature extraction to obtain human body action data; the skeletal joint points of adjacent frames in the human body action data are subjected to three-dimensional spatial displacement calculation to obtain inter-frame action displacement data, and the inter-frame action displacement data are subjected to temporal differentiation operations , and perform a polynomial combination with the angle variation of the joint point to obtain the action energy variation value; perform adaptive threshold segmentation on the action energy variation value to obtain the optimal segmentation threshold, and perform peak detection on the action energy variation value according to the optimal segmentation threshold, and mark the starting point of the data segment that is continuously greater than the optimal segmentation threshold as the action starting point; perform sliding window analysis on the action energy variation value, and mark the end point of the data segment that is continuously less than the optimal segmentation threshold and the window length is greater than the preset time length as the action ending point; according to the action starting point and the action ending point, segment the original action sequence in the human body action data, calculate the similarity with the predefined curved display screen interactive action template, and output the effective interactive action fragment; Performing continuous wavelet transform on the effective interactive action fragments to generate a multi-scale time-frequency feature map, extracting local features and global features, and outputting action recognition results; Extracting three-dimensional coordinates of key points of the human skeleton in the depth image data, and fusing them with the user's inertial sensor data by two-stage impulse response filtering to output the user's real-time spatial coordinates in the coordinate system of the curved display screen; Performing human-computer interaction state analysis based on the action recognition result and the real-time spatial coordinates to generate system state prediction data; Anti-disturbance prediction and dynamic adjustment are performed on the system state prediction data to generate interactive control instructions for driving the curved display screen to display content.

2. The interactive control method of immersive curved display space according to claim 1, characterized in that: The effective interactive action fragments are subjected to continuous wavelet transform to generate a multi-scale time-frequency feature map, and local and global feature extraction is performed to output action recognition results, including: Performing spectrum analysis on the effective interactive action fragments, calculating the main frequency characteristics of the action signal, obtaining the main frequency range of the signal through autocorrelation function calculation, and selecting a wavelet basis function according to the main frequency range of the signal; Optimizing the scale parameters of the wavelet basis function, performing multi-scale analysis on the action features through the information entropy minimization criterion, and obtaining the optimal decomposition scale sequence; Performing a convolution operation on the effective interactive action fragment and the wavelet basis function to obtain an original time-frequency coefficient matrix, and performing singular value decomposition on the original time-frequency coefficient matrix to construct a noise reduction matrix; Dividing the noise reduction matrix into sub-bands according to a preset frequency interval to obtain a plurality of frequency band sub-matrices, and performing energy normalization processing on the plurality of frequency band sub-matrices respectively to obtain normalized time-frequency coefficients; Calculating the energy distribution characteristics of each frequency band according to the normalized time-frequency coefficients to obtain a time-frequency energy distribution matrix, and performing multi-scale fusion on the time-frequency energy distribution matrix to obtain the multi-scale time-frequency feature spectrum; Local features and global features are extracted from the multi-scale time-frequency feature map, and an action recognition result is output.

3. The interactive control method of immersive curved display space according to claim 2, characterized in that: The extracting local features and global features of the multi-scale time-frequency feature map and outputting the action recognition result includes: Input the multi-scale time-frequency feature map into a depth-separable convolutional layer, and use point-wise convolution and depth-wise convolution decomposition operations to obtain an initial feature map; Performing multi-scale spatial pyramid pooling on the initial feature map, and extracting multi-level local features through pooling kernels of different scales; Performing channel attention calculation on the multi-level local features to obtain channel weight coefficients, and recalibrating the feature channels according to the channel weight coefficients to obtain enhanced local features; Input the enhanced local feature sequence into a bidirectional linear self-attention module, and construct an attention weight by performing a linear projection operation of a query matrix and a key value matrix; Performing weighted summation on the feature sequence according to the attention weight to obtain a global context feature, and fusing it with the enhanced local feature through a residual connection to obtain a mixed feature; The mixed features are reorganized into feature channels, the importance coefficient of each feature is calculated through an adaptive weight network to obtain a weighted fusion feature, and the weighted fusion feature is input into a fully connected classifier, and the probability distribution of each action category is calculated through a softmax function to obtain an action recognition result.

4. The interactive control method of immersive curved display space according to claim 3, characterized in that: The three-dimensional coordinates of the key points of the human skeleton in the depth image data are extracted, and two-stage impulse response filtering and fusion are performed with the inertial sensor data of the user to output the real-time spatial coordinates of the user in the coordinate system of the curved display screen, including: Performing human body region segmentation based on the depth image data, extracting human body contours, and obtaining a human body region mask; Performing skeleton refinement processing on the human body region mask to obtain a skeleton outline map, and locating key points of the skeleton outline map according to human anatomical features to obtain two-dimensional skeleton key points; Performing projection transformation on the two-dimensional skeleton key points and the depth image data to obtain initial coordinate data, and performing kinematic constraint optimization on the initial coordinate data to obtain three-dimensional coordinate data; Performing zero drift compensation and temperature calibration on the user's inertial sensor data to obtain calibrated acceleration and angular velocity data, and performing integration operation on the calibrated acceleration and angular velocity data to obtain relative displacement data; Performing attitude calculation on the relative displacement data, obtaining attitude angle data through quaternion calculation, and performing coordinate system conversion on the relative displacement data according to the attitude angle data to obtain global displacement data; Inputting the three-dimensional coordinate data and the global displacement data into a first-stage finite impulse response filter, and obtaining initial fusion data through weighted average operation; The initial fusion data is input into the second-stage finite impulse response filter, and the real-time spatial coordinates of the user in the coordinate system of the arc display screen are output through complementary filtering operation.

5. The interactive control method of immersive curved display space according to claim 4, characterized in that: The human-computer interaction state analysis is performed based on the action recognition result and the real-time spatial coordinates to generate system state prediction data, including: Performing time-series alignment on the action recognition result and the real-time spatial coordinates of the user in the coordinate system of the curved display screen to obtain basic system status data; Quantitatively calculate the data collection delay, data transmission delay and display refresh delay of the curved display screen to obtain a system delay matrix, and perform parameter combination operation on the system state basic data and the system delay matrix to obtain a state space expression; Performing Taylor series expansion and linear approximation calculation on the state space expression to obtain a human-computer interaction state equation; Performing state expansion processing on the human-computer interaction state equation and inputting it into a disturbance observation module to obtain an expanded state variable; Substituting the expanded state variable into a Gaussian probability distribution model for probability calculation to obtain state probability distribution parameters, and performing recursive prediction operations based on the state probability distribution parameters and historical state sequences to obtain multiple groups of predicted state values; Error correction is performed on the multiple groups of predicted state values, and system state prediction data is output.

6. The interactive control method of immersive curved display space according to claim 5, characterized in that: The anti-disturbance prediction and dynamic adjustment of the system state prediction data to generate interactive control instructions for driving the curved display screen to display content includes: Inputting the system state prediction data into a disturbance decoupling module, and obtaining an internal state sequence and an external disturbance sequence through nonlinear separation operation; Performing a rolling time domain optimization calculation on the internal state sequence and the external disturbance sequence to obtain a predictive control cost function, and solving the predictive control cost function according to a quadratic programming method to obtain a control input optimization sequence; Constraining the control input optimization sequence, and performing display boundary correction in combination with real-time feedback data to obtain initial control instructions; Inputting the initial control instruction into an adaptive gain matrix to distribute the control force to obtain a multi-dimensional control quantity, and performing fuzzy dynamic compensation calculation on the multi-dimensional control quantity and feedback data to obtain a compensation control quantity; The compensation control amount is mapped and converted to obtain a smooth control sequence, and the smooth control sequence is converted into a display parameter of an arc display screen, and an interactive control instruction for driving the arc display screen to display content is output.

7. An immersive arc display space interactive control system, characterized in that: A system for executing the immersive curved display space interactive control method according to any one of claims 1 to 6, the system comprising: A collection module, used to collect depth image data of the user's continuous actions through a depth camera set in the curved display screen, and identify effective interactive action segments; A feature extraction module is used to perform continuous wavelet transform on the effective interactive action fragments, generate a multi-scale time-frequency feature map, extract local features and global features, and output action recognition results; A filtering and fusion module, used to extract three-dimensional coordinates of key points of the human skeleton in the depth image data, and perform two-stage impulse response filtering and fusion with the user's inertial sensor data to output the user's real-time spatial coordinates in the coordinate system of the curved display screen; A state analysis module, used to perform human-computer interaction state analysis based on the action recognition result and the real-time spatial coordinates, and generate system state prediction data; A generation module is used to perform anti-disturbance prediction and dynamic adjustment on the system state prediction data, and generate interactive control instructions for driving the curved display screen to display content.

Citation Information

Patent Citations

  • Human-computer interaction method and device

    CN115268619A