Infrared-sensing-based somatosensory interaction control system
By using infrared radiation field analysis and background light-thermal differential technology, the optical attenuation matrix projection and gesture path recognition of the motion-sensing interactive control system were optimized, solving the problems of light interference and misjudgment in complex environments of traditional systems, and realizing high-precision, steady-state motion-sensing interactive control.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANGHAI PAIJIA CULTURE COMMUNICATION CO LTD
- Filing Date
- 2026-04-28
- Publication Date
- 2026-07-21
AI Technical Summary
Traditional motion-sensing interactive control systems are susceptible to light interference in complex environments, leading to misjudgments and command drift. They cannot accurately recognize complex gestures and cannot handle the asynchronous timing issues of hardware rendering delays and perception triggers, resulting in delayed interactive feedback and screen tearing.
The infrared radiation field occlusion analysis module analyzes the energy distribution of the infrared beam, and combined with the background light thermal radiation reference differential stripping, generates the human body occlusion optical attenuation matrix. The spatial coordinates are projected through affine transformation, and combined with topology path comparison and sliding time window filtering, the action-level matching judgment logic is optimized. With the display driver cycle recognition and dynamic response compensation technology, the influence of light noise is eliminated, and the steady state of interactive control is ensured.
It improves the purity and recognition accuracy of human body occlusion signals, enhances the sensitivity of gesture path offset capture, optimizes action-level matching judgment, eliminates accidental touches and improves the steady state of command execution, and ensures a high degree of synchronization between motion control and screen display.
Smart Images

Figure CN122431558A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of human-computer interaction technology, and in particular to a motion-sensing interactive control system based on infrared sensing. Background Technology
[0002] Traditional motion-sensing interactive control systems refer to systems that acquire human motion information through infrared sensing to achieve interactive control. They mainly address the problem of completing interface control without contacting the device. They employ infrared transmitters and receivers arranged in the display area to form a sensing zone. By detecting changes in infrared light occlusion after a person enters the zone, position information is obtained. This is then combined with a pre-set coordinate mapping relationship to map the position of the human hand or limb to control points on the display screen. Simultaneously, by continuously sampling position information, the movement trajectory is calculated and compared with a preset gesture path to determine operations such as clicking, swiping, or hovering. In addition, fixed thresholds are set to determine the action triggering conditions, such as triggering a confirmation operation by having the hand stay in a certain area for a certain period of time. In practical applications, this type of method has problems such as sensitivity to ambient light interference, a large number of false positives due to occlusion, and limited accuracy of motion recognition.
[0003] Traditional motion-sensing interactive control systems rely on simple infrared beam occlusion to determine position, which is highly susceptible to interference from background light and heat radiation in complex exhibition hall environments. This results in a large amount of noise mixed in the sensing signal. The system determines stillness or movement based solely on fixed thresholds, failing to effectively distinguish between ambient light fluctuations and real human movements. This can easily lead to misjudgments and command drift in dynamic scenes with dense crowds. Due to its lack of in-depth analysis of the complex geometric relationship between spatial three-dimensional coordinates and screen pixel planes, it has insufficient accuracy when performing complex gesture recognition. Furthermore, it cannot handle the asynchronous timing issues between hardware rendering delays and sensing triggers, resulting in noticeable lag or screen tearing in interactive feedback. It is difficult to ensure stable operation under high-frequency interactions in large spaces. Summary of the Invention
[0004] The purpose of this invention is to overcome the shortcomings of existing technologies and propose a somatosensory interactive control system based on infrared sensing.
[0005] To achieve the above objectives, the present invention adopts the following technical solution: a motion-sensing interactive control system based on infrared sensing includes: The infrared radiation field occlusion analysis module monitors the energy distribution of the infrared beams produced by the infrared radiation source array of the exhibition booth, tracks the optical attenuation depth of each infrared sensing point under the state of human limb blockage, retrieves the preset exhibition background light and heat radiation benchmark to perform differential stripping calculation, and generates human occlusion optical attenuation matrix. The display screen motion-sensing spatial trajectory prediction module analyzes the relationship between the three-dimensional physical space of the booth and the two-dimensional pixel plane affine transformation of the interactive display screen based on the human body occlusion optical attenuation matrix, projects the spatial coordinates into a set of pixel coordinates, determines the movement vector based on the coordinate deflection direction, and generates the total gesture path offset. The booth control action matching degree determination module calls the total amount of gesture path offset, extracts the standard control gesture topology path of the exhibition area screen, calculates the geometric overlap ratio between the real-time gesture spatial offset trajectory and the standard control gesture topology path, inputs the geometric overlap ratio into the trigger determination threshold logic operator for cross-comparison, and generates the interaction command action-level matching determination coefficient. The interactive screen response timing bias adjustment module obtains the rendering drive cycle corresponding to the current output frame rate of the interactive display screen based on the action-level matching judgment coefficient of the interactive command, identifies the light noise tolerance value caused by the movement of people in the booth environment and performs time delay compensation calculation to generate the interactive drive timing bias of the display screen. The screen motion-sensing command steady-state execution module obtains the screen-spreading interaction driving timing bias, performs smoothing and shaping processing on the motion-sensing command trigger timing, performs sliding time window filtering calculation on the action-level matching judgment coefficient in multiple consecutive rendering cycles, and performs command latching operation when the sliding filter value remains in the action confirmation range, and outputs steady-state motion-sensing interaction control execution command.
[0006] As a further aspect of the present invention, the infrared radiation field blocking analysis module includes: The booth space infrared field construction submodule receives the initial light field distribution pattern around the booth, performs physical blocking monitoring of the infrared beam energy attenuation in the area covered by the light spot, analyzes the infrared energy density on the grid nodes, and generates the initial spatial light intensity mapping array. The background light source interference stripping submodule obtains the pre-stored background light and heat radiation constants of the static ceiling lights and spotlights in the exhibition hall, performs differential stripping operation on the initial spatial light intensity mapping array and the background light and heat radiation constants of the static ceiling lights and spotlights to eliminate the influence of fixed light source illumination in the exhibition hall and obtain the net energy value of the grid. The limb occlusion optical attenuation calculation submodule tracks the spatial drop trend of the net energy value of the grid within a preset infrared scanning cycle. When a sudden drop in energy corresponding to the human body contour area is detected in a continuous grid, the energy attenuation depth and spatial coordinates of the blocked area are calculated to generate a human body occlusion optical attenuation matrix.
[0007] As a further aspect of the present invention, the screen-based motion-sensing spatial trajectory estimation module includes: The screen logical domain mapping submodule calls the human body occlusion optical attenuation matrix to analyze the affine transformation relationship between the three-dimensional physical space coordinate system of the booth and the two-dimensional pixel plane of the interactive display screen. Through matrix multiplication, the three-dimensional physical space coordinates corresponding to the attenuation matrix are projected onto the large screen pixel plane to form a set of pixel action coordinates. The gesture displacement optical slice parsing submodule calculates the physical space coordinate difference between consecutive action frames in the pixel action coordinate set, generates a displacement vector, and logically compares the magnitude of the displacement vector with a preset gesture movement step size threshold. If the magnitude is greater than the preset gesture movement step size threshold, the displacement vector is retained as a valid limb movement vector. The spatial motion continuous geometric deduction submodule performs double integral accumulation processing with respect to time on the effective limb movement vector, splices the displacement spatial increment at the moment of perception, and calculates the total gesture path offset.
[0008] As a further aspect of the present invention, the booth control action matching degree determination module includes: The screen control standard gesture topology comparison submodule retrieves the standard control gesture topology path that matches the current movement trend from the exhibition area interactive control space library based on the total gesture path offset. It then performs point-to-point spatial distance calculation between the physical nodes included in the trajectory tensor and the standard control gesture topology path to determine the trajectory deviation error of each action node. The gesture trajectory space overlap measurement submodule counts the number of action nodes whose trajectory deviation error is less than the preset space tolerance limit, calculates the ratio of the number to the total number of action nodes, obtains the three-dimensional trajectory overlap coefficient, and performs a weighted summation of the three-dimensional trajectory overlap coefficient with the preset trigger judgment threshold to generate the gesture trajectory space matching degree parameter. The trigger threshold cross-judgment submodule prioritizes the preset screen control instruction set according to the magnitude of the gesture trajectory space matching degree parameter, extracts a set of instructions corresponding to the peak value of the coefficient as interactive screen operation instructions to be executed, and generates interactive instruction action-level matching judgment coefficients.
[0009] As a further aspect of the present invention, the interactive screen response timing bias adjustment module includes: The large screen rendering frame rate cycle deconstruction submodule identifies the display driving parameters of the interactive display screen based on the interaction command action level matching judgment coefficient, extracts the single frame rendering time under the current refresh rate, takes the single frame rendering time as the basic processing cycle, and generates the large screen hardware inherent latency benchmark value. The exhibition environment light noise tolerance compensation submodule is based on the inherent delay benchmark value of the large screen hardware. It detects the random brightness fluctuation frequency of the infrared sensing grid in the uninhabited area or non-interactive state, calculates the light noise interference distribution density of people walking within a unit physical area, identifies the exhibition environment correction factor according to the interference distribution density, and generates dynamic response compensation increment. The response delay dynamic bias determination submodule calls the dynamic response compensation increment to be superimposed on the inherent latency benchmark of the large screen hardware, and obtains the screen display interaction driving timing bias by matching the screen buffer driving strategy of the interactive display screen through a lookup table method.
[0010] As a further aspect of the present invention, the screen motion-sensing command steady-state execution module includes: The timing shaping submodule of the control instruction stream calls the timing bias of the screen display interaction driver, manages the enqueue of the interactive screen operation instructions to be executed on the time axis, adjusts the popping rhythm of the instruction queue according to the bias, performs smoothing and anti-shaking processing on the motion control instruction stream, and generates a timing shaping instruction sequence. The motion intent sliding window filtering submodule performs interval windowing mean calculation on the interaction instruction action level matching judgment coefficient corresponding to each instruction in the time-series shaping instruction sequence. It uses a sliding time window with a window length of preset action frames to filter occasional accidental touch actions and outputs the steady-state index of haptic motion. When the steady-state index of the motion control driver exceeds the preset safety anti-mistouch threshold, the screen motion control driver submodule encapsulates the instruction execution identifier and screen frame timestamp, calls the driver communication interface of the underlying main control board of the interactive display screen, sends out the large screen motion control driver signal, and outputs the steady-state motion interaction control execution instruction.
[0011] Compared with the prior art, the advantages and positive effects of the present invention are as follows: In this invention, by introducing infrared radiation field depth analysis and exhibition background light and heat reference differential calculation, the interference of fixed light sources and random light noise on the sensing area is eliminated, enhancing the purity and recognition accuracy of human body occlusion signals. By using affine transformation and time axis double integral derivation, the precise projection of spatial actions to pixel coordinates and continuous tracking of displacement vectors are achieved, enhancing the sensitivity of gesture path offset capture. Combined with topological path geometric overlap comparison and sliding time window filtering mechanism, the action-level matching judgment logic is optimized, eliminating instantaneous false touches and improving the steady state of command execution. With the display driver cycle recognition and dynamic response compensation technology, the inherent hardware delay and timing bias caused by environmental movement light noise are offset, shortening the rendering feedback time and ensuring that the motion control and screen display are highly synchronized. Attached Figure Description
[0012] Figure 1 This is a system flowchart of the present invention. Detailed Implementation
[0013] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0014] Please see Figure 1 An infrared sensing-based motion-sensing interactive control system includes: The infrared radiation field occlusion analysis module monitors the energy distribution of the infrared beams produced by the infrared radiation source array of the exhibition booth, tracks the optical attenuation depth of each infrared sensing point under the state of human limb blockage, retrieves the preset exhibition background light and heat radiation benchmark to perform differential stripping calculation, and generates human occlusion optical attenuation matrix. The display screen motion-sensing spatial trajectory prediction module is based on the human body occlusion optical attenuation matrix. It analyzes the relationship between the three-dimensional physical space of the booth and the two-dimensional pixel plane affine transformation of the interactive display screen, projects the spatial coordinates into a set of pixel coordinates, calculates the Euclidean distance between adjacent nodes and compares it with the preset gesture movement step size threshold, determines the movement vector based on the coordinate deflection direction, and generates the total gesture path offset. The booth control action matching degree judgment module calls the total amount of gesture path offset, extracts the preset standard control gesture topology path of the exhibition area screen, calculates the geometric overlap ratio between the real-time gesture spatial offset trajectory and the standard control gesture topology path, inputs the overlap ratio into the trigger judgment threshold logic operator for cross comparison, and generates the interaction command action-level matching judgment coefficient. The interactive screen response timing bias adjustment module obtains the rendering drive cycle of the interactive display screen at the current output frame rate based on the action-level matching judgment coefficient of the interactive command, calculates the product of the action-level matching judgment coefficient of the interactive command and the preset standard response weight, identifies the light noise tolerance value caused by the movement of people in the booth environment and performs latency compensation calculation to generate the interactive drive timing bias of the display screen. The steady-state execution module for screen-based motion-sensing commands acquires the timing bias of the screen-displaying interaction driver, performs smoothing and shaping processing on the trigger timing of the motion-sensing commands to be executed, performs sliding time window filtering calculations on the action-level matching judgment coefficients in multiple consecutive rendering cycles, and performs command latching operations when the sliding filter value remains within the preset action confirmation range, outputting steady-state motion-sensing interaction control execution commands.
[0015] The infrared radiation field blocking analysis module includes: The booth space infrared field construction submodule receives the initial light field distribution pattern around the booth, performs physical blocking monitoring of the infrared beam energy attenuation in the area covered by the light spot, analyzes the infrared energy density on the grid nodes, and generates the initial spatial light intensity mapping array. The system receives the initial light field distribution around the exhibition booth and acquires temporal data using a 128-channel infrared array sensor deployed on the exhibition hall ceiling. The sensors use 940nm infrared light and capture reflected light signals at a sampling rate of 120Hz, converting the analog electrical signals into a digital matrix via a 12-bit analog-to-digital converter. When monitoring the physical blocking of infrared beam energy attenuation within the light spot coverage area, this submodule allocates a 1920x1080 two-dimensional floating-point array in memory. Each array element corresponds to a 5mm x 5mm grid area in the physical space of the booth. It reads the data stream from the ADC channel, multiplies the received photocurrent value by a preset photoelectric conversion coefficient of 0.85, and analyzes the infrared energy density at the grid nodes. This density data, measured in watts per square meter, is filled into the aforementioned two-dimensional floating-point array to form a spatial energy distribution map, generating an initial spatial light intensity mapping array.
[0016] The background light source interference stripping submodule obtains the pre-stored background light and heat radiation constants of the static ceiling lights and spotlights in the exhibition hall, performs differential stripping operation on the initial spatial light intensity mapping array and the background light and heat radiation constants of the static ceiling lights and spotlights, eliminates the influence of fixed light source illumination in the exhibition hall, and obtains the net energy value of the grid. The process involves acquiring pre-stored background light and thermal radiation constants for static ceiling lights and spotlights in the exhibition hall. This is achieved by reading 3D lighting mapping data of the exhibition hall from a read-only memory via an I2C bus. The exhibition hall environment includes four sets of metal halide ceiling lights and twelve sets of localized focused spotlights. Due to the radiation interference generated by the light sources in the infrared band, light field data is continuously collected for 24 hours under a baseline state of closed exhibition hall and no personnel movement. The infrared radiation baseline value at each grid coordinate is extracted and solidified into a constant matrix. Subsequently, a differential stripping operation is performed between the initial spatial light intensity mapping array and the background light and thermal radiation constants of the static ceiling lights and spotlights. Specifically, each coordinate node in the 1920 x 1080 matrix is traversed, and the energy density value in the initial spatial light intensity mapping array is subtracted from the radiation constant value at the corresponding coordinate in the constant matrix. After the subtraction operation, if the calculation result of a node is less than 0, the value of that node is forcibly reset to zero to avoid overcompensation; if the calculation result is greater than or equal to 0, the difference is retained to eliminate the influence of fixed light source illumination in the exhibition hall and obtain the net energy value of the grid.
[0017] The limb occlusion optical attenuation calculation submodule tracks the spatial drop trend of the net energy value of the grid within a preset infrared scanning cycle. When a sudden drop in energy of the continuous grid that matches the human body contour area is detected, the energy attenuation depth and spatial coordinates of the blocked area are calculated to generate the human body occlusion optical attenuation matrix. Spatial drop trend refers to the rate of change of energy gradient at the same coordinate node in the current infrared scanning cycle and the previous infrared scanning cycle, capturing the vector direction of the continuous decrease in energy value to obtain the spatial drop trend. The energy attenuation depth and spatial coordinates of the blocking region refer to extracting the edge closed curve of the energy drop region, calculating the average energy loss percentage within the area enclosed by the curve, and mapping it to a three-dimensional rectangular coordinate system to determine the centroid position coordinates, thus obtaining the energy attenuation depth and spatial coordinates of the blocking region. The spatial drop trend of the net energy value of the grid is tracked within a preset infrared scanning period, with the period duration set to 8 milliseconds. The spatial drop trend refers to the rate of change of the energy gradient at the same coordinate nodes in the current and previous infrared scanning periods. The energy density value of the grid node with an x-coordinate of 100 and a y-coordinate of 200 in the current frame is extracted, and then subtracted from the energy density value of the node with the same coordinates in the previous frame. The difference is divided by the energy density value of the previous frame. If the calculated rate of change is less than -20%, the node is marked as a decay node. The direction of the energy value decrease is captured by extracting the rate of change of the energy gradient of the eight adjacent nodes around the decay node, finding the adjacent node with the largest decrease in value, and establishing the direction of the line connecting the center of the current node and the center of the adjacent node as the drop vector, thus obtaining the spatial drop trend. When a sudden drop in energy corresponding to the area of a human body contour is detected in a continuous grid, the total number of interconnected decay nodes is counted in a two-dimensional array using a connected component labeling algorithm. If the cumulative physical area of the connected nodes is between 0.15 square meters and 0.45 square meters, it is considered an effective limb occlusion. The energy attenuation depth and spatial coordinates of the blocked area are then calculated. The energy attenuation depth and spatial coordinates of the blocked area refer to the extraction of the closed curve at the edge of the energy drop region. In the binary image composed of connected nodes, morphological dilation and erosion operations are used to remove isolated noise points, and a closed contour composed of an outer ring of mesh nodes is extracted. The average energy loss percentage within the curve's bounding area is calculated by summing the energy attenuation values of all nodes within the contour and dividing by the total number of nodes, resulting in an average energy loss percentage of 45%. This is then mapped to a three-dimensional Cartesian coordinate system to determine the centroid position coordinates. The horizontal and vertical physical coordinate values of all nodes within the contour are summed and divided by the total number of nodes to obtain the planar centroid coordinates. Combined with the fixed infrared sensor height parameter of 3 meters and the principle of projective geometry, the physical height position of the occluder in three-dimensional space is calculated, yielding the energy attenuation depth and spatial coordinates of the blocked area, and generating a human occlusion optical attenuation matrix.
[0018] The screen-based motion-sensing spatial trajectory estimation module includes: The screen logic domain mapping submodule calls the human body occlusion optical attenuation matrix, analyzes the affine transformation relationship between the three-dimensional physical space coordinate system of the booth and the two-dimensional pixel plane of the interactive display screen, and projects the three-dimensional physical coordinates corresponding to the attenuation matrix onto the large screen pixel plane through matrix multiplication to form a set of pixel action coordinates. The optical attenuation matrix for human occlusion is invoked. This matrix stores a set of scattered 3D coordinates in the physical space of the exhibition area caused by the movement of people's hands. The affine transformation relationship between the 3D physical space coordinate system of the exhibition booth and the 2D pixel plane of the interactive display screen is analyzed. The physical space of the exhibition booth is defined as a 3D Cartesian coordinate system with a length of 5000 mm, a width of 5000 mm, and a height of 3000 mm. The resolution of the interactive display screen is 3840 x 2160 pixels. The construction of the affine transformation relationship relies on pre-calibrated intrinsic and extrinsic parameter matrices. The extrinsic parameter matrix contains the translation vector and rotation Euler angle from the origin of the booth coordinate system to the screen plane, while the intrinsic parameter matrix contains the pixel focal length and principal point offset of the screen. The three-dimensional physical coordinates corresponding to the attenuation matrix are projected onto the large screen pixel plane through matrix multiplication. The centroid coordinates of fingertip occlusion in the attenuation matrix are extracted, with a horizontal axis parameter of 2500, a vertical axis parameter of 1500, and a height parameter of 1200. This three-dimensional vector is multiplied by a 3x4 joint mapping matrix of intrinsic and extrinsic parameters to calculate homogeneous coordinates. Then, the horizontal and vertical components of the homogeneous coordinates are divided by the depth scaling factor to complete the perspective division operation, forming a set of pixel motion coordinates.
[0019] The gesture displacement optical slice parsing submodule calculates the physical space coordinate difference between consecutive action frames in the pixel action coordinate set, generates a displacement vector, and logically compares the magnitude of the displacement vector with a preset gesture movement step size threshold. If the magnitude is greater than the preset gesture movement step size threshold, the displacement vector is retained as a valid limb movement vector. The physical space coordinate difference refers to the difference between the physical coordinates of the centroid of the pixel extracted in the current action frame and the corresponding physical coordinates of the centroid of the pixel in the previous adjacent action frame. The physical space coordinate difference between consecutive action frames in the pixel motion coordinate set is calculated by extracting elements from the pixel motion coordinate set sequentially according to frame number. The physical space coordinate difference refers to the difference between the physical coordinates of the centroid of the extracted pixel in the current action frame and the corresponding physical coordinates of the centroid in the adjacent previous action frame. For example, the horizontal coordinate of the centroid of the current 15th frame is 1920 and the vertical coordinate is 1080. The horizontal coordinate of the centroid of the previous action frame (frame 14) is 1900 and the vertical coordinate is 1060. Subtracting the horizontal coordinates and vertical coordinates respectively yields a difference of 20, resulting in the physical space coordinate difference. An array of these differences is used to generate a displacement vector. The magnitude of this displacement vector is logically compared with a preset gesture movement step size threshold. The magnitude of the displacement vector is calculated by adding the squares of the horizontal and vertical differences and then taking the square root, resulting in a magnitude of 28.28 pixels. The preset gesture movement step length threshold is set at 15 pixels. This threshold is derived by analyzing 1000 sets of pixel-level jitter data caused by heartbeat and breathing in a static human body state. Vectors with a magnitude less than 15 pixels are judged as invalid body swaying or ambient light disturbances. When the magnitude exceeds the preset gesture movement step length threshold (28.28 is greater than 15), the displacement vector is retained as a valid limb movement vector.
[0020] The spatial motion continuous geometric deduction submodule performs double integral accumulation processing with respect to time on the effective limb movement vector, splices the displacement spatial increment at the moment of perception, and calculates the total gesture path offset. A double integral accumulation process is performed on the effective limb movement vectors over time. After receiving continuous effective limb movement vectors, the timestamp of each vector's generation is extracted, and the time interval between adjacent timestamps is calculated as the integration step size. The first integration operation divides the displacement vector within the time step by the corresponding time to obtain the instantaneous sliding velocity; the second integration calculates the acceleration variable based on this. The purpose of performing the double integral derivation is to fit a continuous spatial spline curve, eliminating the sense of disjointed movement caused by the discreteness of infrared sensor sampling. The spatial displacement increments at each perceived moment are stitched together, and the fitted displacement increments within each time element are summed end-to-end in the pixel coordinate system to form a continuous pixel line without breaks, thus calculating the total gesture path offset.
[0021] The booth control action matching determination module includes: The screen control standard gesture topology comparison submodule retrieves the standard control gesture topology path that matches the current movement trend from the exhibition area interactive control space library based on the total gesture path offset. It then performs point-to-point spatial distance calculation between the physical nodes included in the trajectory tensor and the standard control gesture topology path to determine the trajectory deviation error of each action node. Based on the total offset of the gesture path, represented as a trajectory tensor containing multiple pixel nodes labeled with a time series, a standard control gesture topology path matching the current movement trend is retrieved from the interactive control space library of the exhibition area. The interactive control space library is stored locally in flash memory and pre-stores standard coordinate sequences for 12 control actions, including "swipe left," "swipe right," "double-tap to zoom," and "rotate counter-clockwise." The principal component of the direction of the current total offset within the start and end time points is extracted. The total lateral displacement of the trajectory reaches -800 pixels, and the total vertical displacement is 50 pixels. The standard topology path of "swipe left" is initially retrieved as a comparison benchmark. Subsequently, point-to-point spatial distance calculations are performed between the physical nodes included in the trajectory tensor and the standard control gesture topology path. A dynamic time warping algorithm is used to measure the spatial geometry. The first to Nth nodes in the trajectory tensor are traversed, and the vertical Euclidean distance from each node to the nearest point on the standard path curve is calculated to determine the trajectory deviation error of each action node.
[0022] The gesture trajectory space overlap measurement submodule counts the number of action nodes whose trajectory deviation error is less than the preset space tolerance limit, calculates the ratio of the number to the total number of action nodes, obtains the three-dimensional trajectory overlap coefficient, and performs a weighted summation with the preset trigger judgment threshold to generate the gesture trajectory space matching degree parameter. The number of motion nodes whose trajectory deviation error is less than a preset spatial tolerance limit (set at 50 pixels) is counted. This tolerance is based on the physical dimensions of the large display screen and the 2000mm standing distance of the operator, calculated to account for the human visual error tolerance. The calculated deviation error array is iterated through, and 120 motion nodes are extracted. Of these, 105 nodes have deviation errors between 0 and 49 pixels; these 105 are recorded as meeting the standard. The ratio of this number to the total number of motion nodes is calculated, and 105 is divided by 120, yielding a decimal of 0.875. This ratio represents the 3D trajectory overlap coefficient. The three-dimensional trajectory overlap coefficient is weighted and summed with the preset trigger judgment threshold. The weight of the three-dimensional trajectory overlap coefficient is set to 0.7. The normalized coefficient of the action speed in the preceding module is introduced as an auxiliary parameter with a weight of 0.3. The speed coefficient is 0.9. The product of 0.875 and 0.7 is added to the product of 0.9 and 0.3. The result is 0.8825, which generates the gesture trajectory spatial matching degree parameter.
[0023] The trigger threshold cross-judgment submodule prioritizes the preset screen control instruction set based on the size of the gesture trajectory space matching degree parameter, extracts a set of instructions corresponding to the peak value of the coefficient as interactive screen operation instructions to be executed, and generates interactive instruction action-level matching judgment coefficients. The set of instructions corresponding to the peak value of the coefficient is extracted as the interactive screen operation instructions to be executed. This means retrieving the set of instructions at the top of the queue in the priority sorting results, determining whether the action-level matching judgment coefficient of the interactive instruction associated with the set of instructions is the largest item in the current calculation cycle, and sending the control signal corresponding to the largest item to the display terminal for execution to obtain the interactive screen operation instructions to be executed. Based on the magnitude of the gesture trajectory spatial matching degree parameter, the screen control instruction set includes page switching and video playback. The generated parameter 0.8825 is compared with the basic trigger requirements of each candidate instruction in the instruction set. If the value exceeds the instruction set baseline, it enters the candidate queue. A set of instructions corresponding to the peak coefficient is extracted as the interactive screen operation instructions to be executed. This refers to the set of instructions at the top of the priority ranking results. When multiple gesture intentions exist, the first one in the ranking list is extracted. It is determined whether the interaction instruction action-level matching judgment coefficient associated with the instruction set is the largest value within the current calculation period. The instruction coefficient value within a 16-millisecond judgment window is retrieved. If 0.8825 is the highest value and leads the second-highest value by a margin of 0.05, the control signal corresponding to the largest value is sent to the display terminal for execution. The control signal is encapsulated in a standard JSON data packet format and sent to obtain the interactive screen operation instructions to be executed, generating the interaction instruction action-level matching judgment coefficient.
[0024] The interactive screen response timing bias adjustment module includes: The large screen rendering frame rate cycle deconstruction submodule identifies the display driving parameters of the interactive display screen based on the interaction command action level matching judgment coefficient, extracts the single frame rendering time under the current refresh rate, takes the single frame rendering time as the basic processing cycle, and generates the large screen hardware inherent latency benchmark value. Based on the action-level matching judgment coefficient of the interactive command, the interactive screen main control board returns the video output mode through the internal communication interface, including the vertical synchronization status, color depth, and panel refresh rate. It reads that the screen is running at a frequency of 60 Hz with a 10-bit color depth. Stripping away the single-frame rendering time at the current refresh rate, and considering the screen refresh rate of 60 Hz, dividing 1 second (1000 milliseconds) by 60, the image rendering time is calculated to be 16.67 milliseconds. Taking into account the logic processing time of the TCON board inside the large screen, an additional fixed clock offset of 2.33 milliseconds is added, using the single-frame rendering time as the basic processing cycle, i.e., 19 milliseconds, to generate the inherent hardware latency baseline value for the large screen.
[0025] The exhibition environment light noise tolerance compensation submodule is based on the inherent delay benchmark value of the large screen hardware. It detects the random brightness fluctuation frequency of the infrared sensing grid in the uninhabited area or non-interactive state, calculates the light noise interference distribution density of people walking within a unit physical area, identifies the exhibition environment correction factor based on the interference distribution density, and generates dynamic response compensation increment. The distribution density of light noise interference caused by personnel movement within a unit physical area refers to obtaining the instantaneous coordinates of non-target interactive personnel within the target monitoring area in the exhibition environment, counting the total number of infrared light intensity change nodes caused by personnel movement within a unit physical projection area, and obtaining the distribution density of light noise interference caused by personnel movement. The exhibition environment correction factor refers to the proportional coefficient used to offset background light noise fluctuations by inputting the distribution density of light noise interference from personnel movement into a preset linear gain compensation function and comparing it with the infrared mapping value in a standard laboratory noise-free environment. Based on the inherent latency baseline of the large-screen hardware, the flickering of lights and the movement of non-target personnel at the exhibition site will generate noise at the edge of the infrared field of view. A 1-meter-wide area around the booth is designated as the background monitoring area. The number of abrupt changes in the energy value of the area grid exceeding the background threshold is sampled over 100 analysis periods to calculate the personnel movement light noise interference distribution density per unit physical area. The personnel movement light noise interference distribution density per unit physical area refers to obtaining the instantaneous coordinates of non-target interactive personnel within the target monitoring area in the exhibition environment. The total number of infrared light intensity change nodes caused by personnel movement within a unit physical projection area is counted. The monitoring area is divided into several 1-square-meter standard units. If 45 infrared brightness-dark reversal nodes appear in a standard unit within 1 second, 45 is defined as the interference density scalar of the standard unit, and the personnel movement light noise interference distribution density is obtained. Based on the interference distribution density, an exhibition environment correction factor is identified. The exhibition environment correction factor refers to inputting the personnel movement light noise interference distribution density into a preset linear gain compensation function. The internal logic of the function is: multiply the density scalar 45 by a preset influence weight of 0.12. By comparing the infrared mapping values under a standard laboratory noise-free environment (the laboratory noise-free constant is 1.0), and adding the constant 1.0 to the density calculation result, the proportional coefficient used to offset background light noise fluctuations is calculated, which is 1.0 plus 5.4, resulting in a correction factor of 6.4 for the exhibition environment. Multiplying the factor by the base processing unit of 1 millisecond generates the dynamic response compensation increment.
[0026] The response delay dynamic bias determination submodule calls the dynamic response compensation increment to be superimposed on the inherent latency benchmark of the large screen hardware, and obtains the screen display interaction drive timing bias by matching the screen buffer drive strategy through the lookup table method. The dynamic response compensation increment is superimposed onto the inherent hardware latency baseline of the large screen. The determined inherent hardware latency baseline value of 19 milliseconds is extracted, and added to it, along with the calculated dynamic response compensation increment of 6.4 milliseconds. The total latency offset is 25.4 milliseconds. A lookup table is used to match the interactive display screen's buffering driving strategy. Internally, a key-value mapping table from latency offset to buffering strategy is established. When the total latency offset falls between 20 and 30 milliseconds, the buffering driving strategy matched in the lookup table is "enable double buffering mechanism and pre-rendering thread". The corresponding timing adjustment parameters are extracted to obtain the interactive display driving timing offset.
[0027] The screen motion control command steady-state execution module includes: The timing shaping submodule of the control instruction stream calls the timing bias of the interactive screen driver, manages the enqueue of the interactive screen operation instructions to be executed on the time axis, adjusts the popping rhythm of the instruction queue according to the bias, performs smoothing and anti-shaking processing on the motion control instruction stream, and generates a timing shaping instruction sequence. The timing offset of the interactive screen display driver is invoked, with the offset set to 25.4 milliseconds. The front-end infrared detection frame rate reaches 120 Hz, while the large screen refresh rate is only 60 Hz, leading to instruction stacking and congestion. Interactive screen operation instructions to be executed are queued and managed on the timeline. A circular instruction buffer based on the first-in, first-out (FIFO) principle is built in memory. Effective control instructions are accompanied by a timestamp and pushed into the buffer in sequence. The popping rhythm of the instruction queue is adjusted according to the offset, and a hardware timer is enabled, with its interrupt period set to half of the timing offset of the interactive screen display driver, i.e., 12.7 milliseconds. Whenever the timer triggers an interrupt, the controller pops an instruction from the top of the buffer. Smoothing and debouncing are performed on the motion-sensing instruction stream. If two consecutive popped instructions are found to be identical page scrolling requests with a timestamp interval of less than 30 milliseconds, the latter instruction is discarded to prevent screen tearing or stuttering caused by repeated calculations in the large screen rendering engine. This mechanism cleans up redundant operations and generates a timing-shaping instruction sequence.
[0028] The motion intent sliding window filtering submodule performs interval windowing mean calculation on the interaction instruction action-level matching judgment coefficient corresponding to each instruction in the timing shaping instruction sequence. It uses a sliding time window with a window length of preset action frames to filter occasional accidental touch actions and outputs the steady-state index of haptic motion. For each instruction in the timing-shaping instruction sequence, the interaction instruction action-level matching judgment coefficient is calculated using an interval windowed mean. To prevent erroneous control triggered by user limb spasms or involuntary waving, a sliding time window with a window length of a preset action frame is used. The width of the time window is rigidly set to 5 action frames. A floating-point one-dimensional array of length 5 is created. When a new instruction enters the processing pipeline, its accompanying matching judgment coefficient is pushed into the head of the array, and the old data at the end of the array is pushed out. The five coefficient values in the array are traversed, the values are added together, and then divided by 5 to obtain the mean of the current window. Occasional accidental touches are filtered out. If the mean falls below the safety baseline of 0.75, the instruction in the last frame of the current time window is determined to be an occasional accidental touch and is blocked. When the average confidence of 5 consecutive frames remains above the safety baseline, the calculated average coefficient is used as the confidence index for passage, and the steady-state index of the haptic action is output.
[0029] When the steady-state index of the motion control driver exceeds the preset safety anti-mistouch threshold, the screen motion control driver module encapsulates the instruction execution identifier and screen frame timestamp, calls the driver communication interface of the underlying main control board of the interactive display screen, sends out the large screen motion control driver signal, and outputs the steady-state motion interaction control execution instruction. When the steady-state metric of the motion sensing exceeds the preset safety threshold for preventing accidental touches (calibrated to 0.82 through real-person testing), the system enters a trigger state when the received steady-state metric reaches 0.89. It then encapsulates the instruction execution identifier and screen frame timestamp to construct a 16-byte data frame. The hexadecimal code representing the operation is filled into the second byte of the data frame as the instruction execution identifier. The current kernel microsecond-level timestamp is obtained, split into four bytes, and filled into bytes 4 through 7 of the data frame. After adding a start character and checksum, the system calls the driver communication interface of the interactive display screen's underlying main control board. The interface uses the RS485 differential serial communication protocol with a baud rate set to 115200. Pulling the transmit enable pin low pushes the constructed 16-byte data stream bit-by-bit in binary form to the communication bus, sending down the large-screen motion sensing control drive signal. Through physical layer isolation and communication downlink mechanisms, the steady-state motion sensing interactive control execution instruction is output.
[0030] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention in any other way. Any person skilled in the art may make changes or modifications to the above-disclosed technical content to create equivalent embodiments that can be applied to other fields. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the protection scope of the present invention.
Claims
1. A motion-sensing interactive control system based on infrared sensing, characterized in that, The system includes: The infrared radiation field occlusion analysis module monitors the energy distribution of the infrared beams produced by the infrared radiation source array of the exhibition booth, tracks the optical attenuation depth of each infrared sensing point under the state of human limb blockage, retrieves the preset exhibition background light and heat radiation benchmark to perform differential stripping calculation, and generates human occlusion optical attenuation matrix. The display screen motion-sensing spatial trajectory prediction module analyzes the relationship between the three-dimensional physical space of the booth and the two-dimensional pixel plane affine transformation of the interactive display screen based on the human body occlusion optical attenuation matrix, projects the spatial coordinates into a set of pixel coordinates, determines the movement vector based on the coordinate deflection direction, and generates the total gesture path offset. The booth control action matching degree determination module calls the total amount of gesture path offset, extracts the standard control gesture topology path of the exhibition area screen, calculates the geometric overlap ratio between the real-time gesture spatial offset trajectory and the standard control gesture topology path, inputs the geometric overlap ratio into the trigger determination threshold logic operator for cross-comparison, and generates the interaction command action-level matching determination coefficient. The interactive screen response timing bias adjustment module obtains the rendering drive cycle corresponding to the current output frame rate of the interactive display screen based on the action-level matching judgment coefficient of the interactive command, identifies the light noise tolerance value caused by the movement of people in the booth environment and performs time delay compensation calculation to generate the interactive drive timing bias of the display screen. The screen motion-sensing command steady-state execution module obtains the screen-spreading interaction driving timing bias, performs smoothing and shaping processing on the motion-sensing command trigger timing, performs sliding time window filtering calculation on the action-level matching judgment coefficient in multiple consecutive rendering cycles, and performs command latching operation when the sliding filter value remains in the action confirmation range, and outputs steady-state motion-sensing interaction control execution command.
2. The infrared sensing-based somatosensory interaction control system according to claim 1, characterized in that, The infrared radiation field blocking analysis module includes: The booth space infrared field construction submodule receives the initial light field distribution pattern around the booth, performs physical blocking monitoring of the infrared beam energy attenuation in the area covered by the light spot, analyzes the infrared energy density on the grid nodes, and generates the initial spatial light intensity mapping array. The background light source interference stripping submodule obtains the pre-stored background light and heat radiation constants of the static ceiling lights and spotlights in the exhibition hall, performs differential stripping operation on the initial spatial light intensity mapping array and the background light and heat radiation constants of the static ceiling lights and spotlights to eliminate the influence of fixed light source illumination in the exhibition hall and obtain the net energy value of the grid. The limb occlusion optical attenuation calculation submodule tracks the spatial drop trend of the net energy value of the grid within a preset infrared scanning cycle. When a sudden drop in energy corresponding to the human body contour area is detected in a continuous grid, the energy attenuation depth and spatial coordinates of the blocked area are calculated to generate a human body occlusion optical attenuation matrix.
3. The infrared sensing-based somatosensory interaction control system according to claim 2, characterized in that, The spatial drop trend refers to the rate of change of energy gradient at the same coordinate nodes in the current infrared scanning cycle and the previous infrared scanning cycle, capturing the vector direction in which the energy value continues to decrease, and thus obtaining the spatial drop trend. The energy attenuation depth and spatial coordinates of the blocking region are obtained by extracting the edge closed curve of the energy drop region, calculating the average energy loss percentage within the area enclosed by the curve, and mapping it to a three-dimensional rectangular coordinate system to determine the centroid position coordinates.
4. The infrared sensing-based somatosensory interaction control system according to claim 2, characterized in that, The display screen motion-sensing spatial trajectory estimation module includes: The screen logical domain mapping submodule calls the human body occlusion optical attenuation matrix to analyze the affine transformation relationship between the three-dimensional physical space coordinate system of the booth and the two-dimensional pixel plane of the interactive display screen. Through matrix multiplication, the three-dimensional physical space coordinates corresponding to the attenuation matrix are projected onto the large screen pixel plane to form a set of pixel action coordinates. The gesture displacement optical slice parsing submodule calculates the physical space coordinate difference between consecutive action frames in the pixel action coordinate set, generates a displacement vector, and logically compares the magnitude of the displacement vector with a preset gesture movement step size threshold. If the magnitude is greater than the preset gesture movement step size threshold, the displacement vector is retained as a valid limb movement vector. The spatial motion continuous geometric deduction submodule performs double integral accumulation processing with respect to time on the effective limb movement vector, splices the displacement spatial increment at the moment of perception, and calculates the total gesture path offset.
5. The infrared sensing-based somatosensory interaction control system according to claim 4, characterized in that, The physical space coordinate difference refers to the difference obtained by extracting the physical coordinates of the centroid of the pixel in the current action frame and subtracting the physical coordinates of the corresponding centroid of the pixel in the adjacent previous action frame.
6. The infrared sensing-based somatosensory interactive control system according to claim 4, characterized in that, The booth control action matching degree determination module includes: The screen control standard gesture topology comparison submodule retrieves the standard control gesture topology path that matches the current movement trend from the exhibition area interactive control space library based on the total gesture path offset. It then performs point-to-point spatial distance calculation between the physical nodes included in the trajectory tensor and the standard control gesture topology path to determine the trajectory deviation error of each action node. The gesture trajectory space overlap measurement submodule counts the number of action nodes whose trajectory deviation error is less than the preset space tolerance limit, calculates the ratio of the number to the total number of action nodes, obtains the three-dimensional trajectory overlap coefficient, and performs a weighted summation of the three-dimensional trajectory overlap coefficient with the preset trigger judgment threshold to generate the gesture trajectory space matching degree parameter. The trigger threshold cross-judgment submodule prioritizes the preset screen control instruction set according to the magnitude of the gesture trajectory space matching degree parameter, extracts a set of instructions corresponding to the peak value of the coefficient as interactive screen operation instructions to be executed, and generates interactive instruction action-level matching judgment coefficients.
7. The infrared sensing-based somatosensory interaction control system according to claim 6, characterized in that, The set of instructions corresponding to the peak value of the intercepted coefficient is used as the interactive screen operation instructions to be executed. This means retrieving the set of instructions at the top of the queue in the priority sorting results, determining whether the interaction instruction action-level matching judgment coefficient associated with the instruction set is the largest item in the current calculation cycle, and sending the control signal corresponding to the largest item to the display terminal for execution to obtain the interactive screen operation instructions to be executed.
8. The infrared sensing-based somatosensory interaction control system according to claim 6, characterized in that, The interactive screen response timing bias adjustment module includes: The large screen rendering frame rate cycle deconstruction submodule identifies the display driving parameters of the interactive display screen based on the interaction command action level matching judgment coefficient, extracts the single frame rendering time under the current refresh rate, takes the single frame rendering time as the basic processing cycle, and generates the large screen hardware inherent latency benchmark value. The exhibition environment light noise tolerance compensation submodule is based on the inherent delay benchmark value of the large screen hardware. It detects the random brightness fluctuation frequency of the infrared sensing grid in the uninhabited area or non-interactive state, calculates the light noise interference distribution density of people walking within a unit physical area, identifies the exhibition environment correction factor according to the interference distribution density, and generates dynamic response compensation increment. The response delay dynamic bias determination submodule calls the dynamic response compensation increment to be superimposed on the inherent latency benchmark of the large screen hardware, and obtains the screen display interaction driving timing bias by matching the screen buffer driving strategy of the interactive display screen through a lookup table method.
9. The infrared sensing-based somatosensory interactive control system according to claim 8, characterized in that, The personnel movement light noise interference distribution density within a unit physical area refers to obtaining the instantaneous coordinates of non-target interactive personnel within the target monitoring area in the exhibition environment, counting the total number of infrared light intensity change nodes caused by personnel movement within a unit physical projection area, and obtaining the personnel movement light noise interference distribution density. The exhibition environment correction factor refers to the proportional coefficient used to offset background light noise fluctuations by inputting the distribution density of light noise interference from personnel movement into a preset linear gain compensation function and comparing it with the infrared mapping value in a standard laboratory noise-free environment.
10. The infrared sensing-based somatosensory interactive control system according to claim 8, characterized in that, The screen motion-sensing command steady-state execution module includes: The timing shaping submodule of the control instruction stream calls the timing bias of the screen display interaction driver, manages the enqueue of the interactive screen operation instructions to be executed on the time axis, adjusts the popping rhythm of the instruction queue according to the bias, performs smoothing and anti-shaking processing on the motion control instruction stream, and generates a timing shaping instruction sequence. The motion intent sliding window filtering submodule performs interval windowing mean calculation on the interaction instruction action level matching judgment coefficient corresponding to each instruction in the time-series shaping instruction sequence. It uses a sliding time window with a window length of preset action frames to filter occasional accidental touch actions and outputs the steady-state index of haptic motion. When the steady-state index of the motion control driver exceeds the preset safety anti-mistouch threshold, the screen motion control driver submodule encapsulates the instruction execution identifier and screen frame timestamp, calls the driver communication interface of the underlying main control board of the interactive display screen, sends out the large screen motion control driver signal, and outputs the steady-state motion interaction control execution instruction.