Intelligent exhibition hall interaction control method and system based on multi-modal data
By using multimodal data processing, a real-time interactive mapping trajectory and intent index value are established, which solves the problem of interactive system failure caused by dynamic interference from multiple people and fluctuations in physiological state in intelligent exhibition halls, and achieves highly reliable interactive control.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- JIANGSU ZHIYUAN MODERN SUPPLY CHAIN CO LTD
- Filing Date
- 2026-03-05
- Publication Date
- 2026-06-09
AI Technical Summary
In the multimodal interaction process of smart exhibition halls, existing technologies cannot effectively cope with the failure of interactive systems caused by dynamic interference from multiple people and fluctuations in users' physiological states, including threshold nonlinear mismatch caused by physiological load offset, lack of intent convergence characteristics, and failure of heterogeneous feature spatiotemporal dimension stripping.
By acquiring multi-source raw data of the interaction space, performing affine transformation of the coordinate system and time stamp synchronization alignment, calculating the power spectral density of the micro-motion high-frequency tremor sequence and the fitting residual of the jerk sequence, generating compensation factors and convergence probabilities, establishing a real-time interactive mapping trajectory, and using the geometric consistency parameters of the skeletal midline vector and the gaze vector to generate intention index values to generate control signals.
It improves the continuity of interactive mapping trajectories, reduces the frequency of erroneous triggering, enhances the system's ability to suppress environmental disturbances, and ensures high reliability for long-term interactive tasks.
Smart Images

Figure CN122172966A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of human-computer interaction and signal processing technology, and in particular to an intelligent exhibition hall interactive control method and system based on multimodal data. Background Technology
[0002] In the multimodal interaction process of intelligent exhibition halls, the skeletal joint points, gaze vectors, and end-effector micro-motion sequences collected by the perception system contain dynamic evolutionary information about user interaction intentions, physiological load, and spatial orientation. Existing interactive control technologies use preset geometric analytical models to correlate limb coordinates with target positions. However, when faced with dynamic interference from multiple users and fluctuations in user physiological states, the interactive system exhibits the following failure mechanisms: The nonlinear threshold mismatch mechanism caused by physiological load offset: When the human body performs directional interactions, muscle tissue generates physiological micro-vibrations in the 6Hz to 12Hz frequency band. When the cumulative interaction time triggers the user to enter a physiological fatigue stage, the physiological load fluctuates, causing a nonlinear shift in the power spectral density of the micro-vibration signal. Since existing decision logic is mostly based on fixed linear mapping thresholds, it cannot provide continuous and matched decision gain adjustment near the physiological load offset point, resulting in a physical mismatch between the expression intensity of the interaction feature and the preset threshold, or frequent switching of the interaction decision logic induced by feature attenuation.
[0003] The mapping mismatch mechanism caused by the lack of intent convergence characteristics: The end trajectory of user interaction actions has a dynamic convergence law. Existing control models usually treat limb pointing as a constant spatial vector, lacking awareness of the dynamic convergence dimension of the interaction intent in the process of approaching the target. In open interaction scenarios, due to the physical superposition interference caused by the movement behavior of multiple people, it is difficult for the output of the preset geometric pointing model in the system to maintain a stable and consistent mapping relationship with the actual convergence state of the user's intent. This leads to false triggering or missed triggering at the critical decision point, causing the overall failure of control accuracy.
[0004] The failure mechanism of separating heterogeneous features from spatiotemporal dimensions: Multimodal sensing data has non-stationary and time-varying characteristics. Low-frequency motion features triggered by large-scale limb movements and high-frequency microtremor features caused by physiological load fluctuations overlap in the temporal and spatial domains. Existing frequency domain analysis methods are limited by the assumption of global stationarity and lack dynamic constraints on the local temporal features of the signal, resulting in feature mismatch in the decoupled physical components, which cannot provide pure physical input for the subsequent construction of dynamic interactive mapping trajectories.
[0005] Based on the above analysis of physical characteristics, due to the feature offset triggered by physiological load fluctuations, the lack of dynamic intention convergence rules, and the nonlinear mismatch characteristics of spatiotemporal phase differences in multi-sensor data in the spatiotemporal dimension, the preset judgment benchmark in the interactive system is difficult to maintain a stable and consistent mapping relationship with the evolution process of the actual behavior state, thus causing the overall failure of the steady-state characteristics of the interactive system. Summary of the Invention
[0006] This invention provides an intelligent exhibition hall interactive control method and system based on multimodal data. In open interactive scenarios, environmental disturbances caused by the physical superposition of multiple people's movement behaviors, as well as the instability of movement characteristics caused by fluctuations in users' physiological load, cause mismatch characteristics in the temporal and spatial dimensions of limb pointing features, motion dynamics features and physiological feedback features. This causes the system's preset judgment benchmark to deviate from the mapping trajectory of the real-time behavior state, resulting in the overall failure of the interactive judgment process.
[0007] In view of the above problems, the present invention provides an intelligent exhibition hall interactive control method based on multimodal data, comprising the following steps: Step a. Obtain multi-source raw data within the interaction space, and extract limb pointing spatial vectors, end-effector velocity sequences, and micro-movement high-frequency tremor sequences; Step b. Perform coordinate system affine transformation and timestamp synchronization alignment on the multi-source raw data; Step c. Calculate the power spectral density of the micro-motion high-frequency tremor sequence, and generate a first compensation factor based on the power spectral density; Step d. Perform a second-order difference operation on the end motion velocity sequence to obtain an acceleration sequence, calculate the fitting residual between the acceleration sequence and the preset decay model, and generate a first convergence probability based on the fitting residual; Step e. Adjust the first convergence probability and the limb pointing space vector with the first compensation factor as a variable to establish a real-time interactive mapping trajectory, and generate an intent index value based on the real-time interactive mapping trajectory; Step f. If the intent index value is greater than the first preset threshold, then generate the first control signal.
[0008] According to the method of claim 1, the limb pointing space vector is obtained by calculating the geometric consistency parameters of the user's skeletal midline vector and the line-of-sight vector.
[0009] Furthermore, the generation process of the first compensation factor in step c includes: Extract the signal components from 6Hz to 12Hz from the micro-motion high-frequency tremor sequence; If the power spectral density amplitude of the signal component is greater than the second preset threshold, the first compensation factor is output according to the power spectral density amplitude.
[0010] Furthermore, the preset attenuation model mentioned in step d is... ,in Let be the convergence coefficient of the action. The first convergence probability is obtained by calculating the fitting residual between the jerk sequence and the preset decay model, and mapping the fitting residual to a normalized interval.
[0011] Furthermore, step e includes a step of determining the adjustment factor by adjusting the intention through a first compensation factor, wherein the first compensation factor is used... This indicates that the intention determination adjustment factor is used. This indicates that the intent determination adjustment factor satisfy: in, Basic regulatory factor, The first preset parameter, This is the second preset parameter. This is the third preset parameter.
[0012] Furthermore, the intent index value is used In other words, the first convergence probability is represented by... This indicates that the intent index value satisfy: in, The angle between the limb pointing spatial vector and the preset pointing center line is defined by the line connecting the geometric center of the candidate interactive target and the origin of the perception unit. It is the order.
[0013] Furthermore, it also includes an anomaly monitoring step: calculating the mismatch distance between the real-time interactive mapping trajectory and the reference trajectory; if the mismatch distance is greater than a third preset threshold, then suppressing the first control signal.
[0014] This invention provides an intelligent exhibition hall interactive control system based on multimodal data, characterized in that it includes: The sensing unit is used to acquire multi-source raw data within the interaction space; The processing unit is used to execute the above method to generate a first control signal; An execution unit is configured to receive the first control signal and execute an interactive response.
[0015] The present invention provides a computer-readable storage medium having a computer program stored thereon, characterized in that the computer program implements the above-described method when executed by a processor.
[0016] The technical solution provided in this application has at least the following technical effects: By extracting the power spectral density of microtremor sequences in the 6Hz to 12Hz frequency band, this scheme constructs a nonlinear adaptive sinking mechanism in which the intent determination adjustment factor dynamically fluctuates with the first compensation factor (physiological load intensity). This mechanism effectively transforms the physiological tremor signal, which is usually regarded as interference noise, into a gain adjustment variable of the system, compensating for the attenuation of interaction feature expression intensity caused by muscle fatigue at the physical level, and improving the continuity of the interaction mapping trajectory under long-term working conditions.
[0017] By nonlinearly coupling the limb pointing spatial vector with the first convergence probability fitted based on the jerk sequence, this scheme reduces the single dependence of the decision logic on instantaneous geometric position features. By applying physical constraints to the jerk features using an exponential decay model, the system can accurately identify and filter non-intentional environmental disturbance signals, significantly reducing the frequency of erroneous actions and making the process of identifying interaction intentions more physically deterministic.
[0018] In the access sovereignty determination phase, geometric consistency constraints are introduced between the skeletal midline vector and the line-of-sight vector. During the operational phase, real-time monitoring of mismatch distance based on the Euclidean L2 norm is implemented, enabling the system's response logic to immediately suppress logical drift caused by sudden environmental changes or sensory occlusion. This closed-loop design minimizes the risk of uncontrolled signal output and enhances the system's ability to shield against abnormal commands in complex dynamic environments.
[0019] The synergistic achievement of the aforementioned technical effects unifies the feature repair process based on physiological compensation and the intent locking process based on dynamic constraints within a nonlinearly coupled framework of intent index value calculation. This enables the system to simultaneously adapt to fluctuations in human physiological load and suppress interference from the physical environment without increasing additional sensing load, thus meeting the high reliability requirements of the control system's steady-state characteristics for high-frequency, long-duration interactive tasks in intelligent exhibition halls. Attached Figure Description
[0020] Figure 1 This is a schematic diagram of the overall process of the intelligent exhibition hall interactive control method based on multimodal data provided in an embodiment of the present invention; Figure 2 This is a structural diagram of an intelligent exhibition hall interactive control system based on multimodal data, provided in an embodiment of the present invention. Detailed Implementation
[0021] The above technical solutions will now be described in detail with reference to the accompanying drawings and specific embodiments to provide a better understanding of them. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. It should be understood that the present invention is not limited to the exemplary embodiments used only to explain the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention. Furthermore, it should be noted that, for ease of description, only the parts related to the present invention are shown in the drawings, not all of them.
[0022] Example
[0023] Please see Figure 2 This invention provides an intelligent exhibition hall interactive control system based on multimodal data. The system includes: a sensing unit for acquiring multi-source raw data in the interactive space; a processing unit for executing the following interactive control method to generate a first control signal; and an execution unit for receiving the first control signal and executing an interactive response.
[0024] Please see Figure 1 This invention provides an intelligent exhibition hall interactive control method based on multimodal data. This method is executed through the aforementioned system and includes the following steps: Step a. Obtain multi-source raw data within the interaction space, and extract limb pointing spatial vectors, end-effector velocity sequences, and micro-movement high-frequency tremor sequences; Step b. Perform coordinate system affine transformation and timestamp synchronization alignment on the multi-source raw data; Step c. Calculate the power spectral density of the micro-motion high-frequency tremor sequence, and generate a first compensation factor based on the power spectral density; Step d. Perform a second-order difference operation on the end motion velocity sequence to obtain an acceleration sequence, calculate the fitting residual between the acceleration sequence and the preset decay model, and generate a first convergence probability based on the fitting residual; Step e. Adjust the first convergence probability and the limb pointing space vector with the first compensation factor as a variable to establish a real-time interactive mapping trajectory, and generate an intent index value based on the real-time interactive mapping trajectory; Step f. If the intent index value is greater than the first preset threshold, then generate the first control signal.
[0025] During system operation, the processing unit also performs anomaly monitoring steps: calculating the mismatch distance between the real-time interactive mapping trajectory and the baseline trajectory; if the mismatch distance is greater than the third preset threshold, then suppressing the first control signal.
[0026] The perception and spatiotemporal normalization phase of multi-source raw data is initiated by the core hub controller. The depth camera in the perception unit captures the spatiotemporal sequence of the user's skeletal joints within the interactive scene, while the infrared tracking device simultaneously captures the instantaneous projection coordinates of the user's pupil center relative to the interactive plane. The processing unit extracts the local pixel region where the interactive end is located from the high frame rate image output by the depth camera, and within a time window when the motion speed of the interactive end is below a preset steady-state threshold (e.g., 0.1 m / s), it uses an optical flow field algorithm to extract the displacement fluctuation data of this local pixel region in the temporal domain. In one embodiment, the processing unit utilizes the oversampling characteristics of a high frame rate (greater than 120 FPS) camera to remove pixel blurring interference caused by macroscopic motion through temporal smoothing filtering. The extraction process of the limb pointing spatial vector is completed by the processing unit through vectorization operations based on the acquired spatiotemporal sequence of skeletal joints. The end-effector motion velocity sequence is generated by the processing unit performing a first-order derivative operation on the displacement data of the interactive end. The micro-motion high-frequency jitter sequence is obtained by the processing unit performing high-pass filtering on the local pixel displacement fluctuation data with a preset cutoff frequency to ensure that the signal components with the retained frequency in the 6Hz to 12Hz band have sufficient signal-to-noise ratio.
[0027] The coordinate system affine transformation and timestamp synchronization alignment process for multi-source raw data is executed by the core hub controller-driven processing unit. The processing unit calls pre-stored affine transformation matrices of each hardware device relative to the global coordinate system of the exhibition hall. These affine transformation matrices are constructed from rotation and translation vectors obtained during the pre-calibration process. The limb pointing spatial vector, the end-effector velocity sequence, and the projected coordinate data output by the infrared tracking device are multiplied with their corresponding affine transformation matrices, mapping the heterogeneous multi-source raw data to a unified three-dimensional geometric reference system of the exhibition hall. The time-dimensional synchronization alignment logic is implemented by the core hub controller using system clock pulses. The processing unit performs linear interpolation compensation based on the timestamps in the data packets uploaded by each sensing device, establishing a millisecond-level correspondence between the limb pointing spatial vector, the end-effector velocity sequence, and the micro-motion high-frequency tremor sequence on the time axis.
[0028] The calculation of geometric consistency parameters between the skeletal midline vector and the gaze vector is initiated by the processing unit after the generation of spatiotemporally synchronized multi-source raw data. The processing unit selects the shoulder center point, elbow joint point, and wrist joint point based on the skeletal joint sequence to construct the limb midline vector. The gaze vector of the user's visual focus point is constructed by the processing unit combining the pupil center coordinates output by the infrared tracking device and head posture parameters. The core hub controller first searches for the closest interactive object to the gaze point in the 3D spatial model of the exhibition hall as a candidate target based on the gaze point coordinates output by the infrared tracking device. If there is no interactive target within a preset radius around the gaze point, the system enters a standby monitoring state and does not initiate subsequent geometric consistency calculations. The preset pointing center line is then determined by connecting the geometric center of the candidate interactive target to the origin of the perception unit. The cosine value of the spatial angle between the limb midline vector and the gaze vector is obtained by the processing unit through vector dot product. The cosine value of the spatial angle participates in the quantification of the user's interactive action pointing characteristics as a geometric consistency parameter. The core hub controller compares the geometric consistency parameter with a preset consistency threshold. If the geometric consistency parameter is greater than the preset consistency threshold, the current interacting entity is determined to have access sovereignty. The limb pointing spatial vector of the entity in the access sovereignty determination state is sent to the adaptive calibration unit as the spatial dimension input for intent index value calculation.
[0029] The quantification of physiological load status and the generation of the first compensation factor are initiated after the spatiotemporally synchronized multi-source raw data output. The processing unit performs filtering on the micro-motion high-frequency tremor sequence using a configured digital bandpass filter operator, with the lower cutoff frequency set to 6Hz and the upper cutoff frequency set to 12Hz. During filtering, frequency components below 6Hz and above 12Hz in the micro-motion high-frequency tremor sequence are suppressed. The power spectral density calculation of the signal components extracted from the micro-motion high-frequency tremor sequence is performed by the processing unit after filtering. The processing unit performs windowing on the extracted signal components and executes a Fast Fourier Transform (FFT) to convert the time-domain displacement fluctuation data into frequency-domain energy distribution data. By squaring the magnitude of the FFT output and integrating it within the frequency range of 6Hz to 12Hz, the processing unit calculates the power spectral density value representing the physiological tremor intensity.
[0030] The generation logic of the first compensation factor is executed by the core hub controller based on the power spectral density value. The core hub controller compares the power spectral density value with a pre-stored second preset threshold. If the power spectral density value is greater than the second preset threshold, the processing unit determines that the interactive subject currently has a physiological load offset. The processing unit generates the first compensation factor based on the magnitude by which the power spectral density value exceeds the second preset threshold, through linear mapping or a preset energy-load lookup table. As a specific implementation data example, the energy-load lookup table adopts a hierarchical ladder mapping structure: when the magnitude of the power spectral density value exceeding the second preset threshold is in the range [0, 10%), the first compensation factor is assigned a value of 0.1; when the magnitude is in the range [10%, 30%), the first compensation factor is assigned a value of 0.3; when the magnitude exceeds 30%, the first compensation factor is assigned a value of 0.6. The output of this first compensation factor provides a real-time variable input for the dynamic adjustment of the subsequent intent determination adjustment factor. It is worth noting that, considering that the extraction of physiological tremor features depends on the quasi-resting state, while the execution of interactive intent is often accompanied by large-amplitude limb movements, the core hub controller is equipped with a state holding register. Once the processing unit calculates the first compensation factor within the quasi-static time window, this value is latched into the state holding register and maintained for a preset valid duration (e.g., 5 seconds) or until the next quasi-static state is detected. During subsequent dynamic interactions, the adaptive calibration unit directly calls the first compensation factor in the state holding register to participate in the real-time calculation of the intention determination adjustment factor.
[0031] The motion dynamics convergence determination and probability calculation process is initiated by the dynamics determination logic in the core hub controller's activation processing unit. The second-order difference operation of the end-effector velocity sequence is performed by the processing unit, which calls the temporal difference operator to perform two consecutive first-order difference operations on the end-effector velocity sequence. Through this second-order difference operation, the end-effector velocity sequence is transformed into an acceleration sequence reflecting the acceleration characteristics of the interactive action. The processing unit uses a least-squares fitting algorithm to fit the numerical trajectory of the acceleration sequence within a preset observation window (e.g., 200ms to 500ms) to a preset decay model. The fitting residual is calculated and output. The fitting residual reflects the geometric deviation between the real-time action acceleration characteristics and the ideal convergence model. Subsequently, the processing unit uses a preset monotonically decreasing mapping function to map the fitting residual to a normalized interval of [0,1], thereby obtaining the first convergence probability. Under this mapping logic, the smaller the fitting residual, the higher the output first convergence probability value, and the clearer the characterization of the interactive subject's action convergence intention.
[0032] The preset attenuation model uses an exponential attenuation function. Describes the physical convergence law of interactive intent as it approaches the goal, in this exponential decay function, The preset motion convergence coefficient (e.g., a value range of 5.0 to 15.0) is calibrated by the system based on the geometric dimensions of the interactive target and the depth resolution of the interactive space. This is a time variable relative to the start point of the sampling time window.
[0033] The calculation of the first convergence probability and the division of the confidence interval are performed by the processing unit based on the fitting residuals. The processing unit converts the fitting residuals into a first convergence probability with a numerical range between 0 and 1 using a mapping function. The processing unit then divides the confidence interval based on the numerical range of the first convergence probability. The calculated first convergence probability is then sent to the adaptive calibration unit as the interaction intent confidence parameter of the dynamic dimension, participating in the subsequent weight reconstruction calculation of the intent index value.
[0034] The mapping trajectory reconstruction and intent index scoring process is initiated by the adaptive calibration unit after the first compensation factor and the first convergence probability are generated. The compensation curve implementation process, where the intent determination adjustment factor changes with the first compensation factor, is completed by the adaptive calibration unit's scheduling processing unit. The adaptive calibration unit uses the first compensation factor as an input variable, according to the formula: The intention is to determine the adjustment factor. In this formula, the first compensation factor is determined by... This indicates that the intention is to determine the regulating factor by express, This represents the baseline adjustment factor (e.g., a value of 0.8). This represents the first preset parameter (with a value range of 0.2 to 0.4). This represents the second preset parameter (with a value range of 1.5 to 3.0). This represents the third preset parameter (used to correct system noise floor, e.g., 0.05). Specific parameter values are calibrated and adjusted based on the resolution and field of view of the exhibition hall cameras. The adjustment factor is determined through the nonlinear mapping of the hyperbolic tangent function. With the first compensation factor The increase shows a monotonically decreasing trend.
[0035] In this formula, the second preset parameter As a sensitivity mapping coefficient, its dimensions are the same as those of the first compensation factor. The dimensions are inverted, thus ensuring the hyperbolic tangent function. The input items are dimensionless numerical values. Additionally, the first compensation factor... Before being substituted into the calculation, a normalization process based on the maximum physiological tremor intensity was performed to lock its value within the [0,1] interval.
[0036] The multimodal feature nonlinear coupling solution process based on the Hill equation is initiated after the intention determination adjustment factor is generated. The adaptive calibration unit acquires the limb pointing space vector and the first convergence probability, and calculates the angle between the limb pointing space vector and the preset pointing center line. The intention index value is calculated by the adaptive calibration unit using the formula: Execution. In this formula, the intent index value is determined by... This indicates that the first convergence probability is given by This indicates that the angle between the limb pointing spatial vector and the preset pointing center line is determined by... express, Represents the order. First convergence probability. With intention judgment moderating factor All values were normalized to the dimensionless range [0,1]. During the calculation, [the following was done / processed]. Take absolute value or limit Within the effective field of view, to ensure the intent index value The monotonic positive positivity. Intent index value The numerical value is controlled by the intention-determining adjustment factor. The real-time fluctuations allow the weight for determining the interaction intent to be nonlinearly reconstructed based on the physiological state reflected by the first compensation factor.
[0037] The real-time construction logic of the dynamic interaction mapping trajectory relies on the continuous output of the intent index value in the time domain. During the generation of the intent index value, the processing unit associates the intent index value with the corresponding limb pointing spatial vector to form a spatiotemporal mapping relationship describing the user's interaction behavior trajectory. If the intent index value is greater than a first preset threshold, the core hub controller determines that the interaction intent is valid and issues an instruction to generate the first control signal. During this process, the processing unit records the changing trends of the intent index value and the limb pointing spatial vector in real time, and the constructed dynamic interaction mapping trajectory serves as the benchmark data for subsequent anomaly monitoring and logic suppression. This trajectory construction method based on deep coupling of multimodal features establishes the logical depth of interaction intent determination, enabling the first control signal output by the system to be based on the technical foundation of the coordinated constraints of physiological state, spatial pointing, and dynamic characteristics.
[0038] The monotonically decreasing relationship between the intent determination adjustment factor and the first compensation factor can be achieved not only through the hyperbolic tangent function, but also through other mathematical functions or logical mappings with monotonically decreasing characteristics. In another implementation, the adaptive calibration unit uses the logistic regression function, i.e., the sigmoid function, instead of the hyperbolic tangent function to calculate the intent determination adjustment factor, making the intent determination adjustment factor smoothly decrease within a preset numerical range as the first compensation factor decreases. In yet another implementation, the adaptive calibration unit uses a piecewise linear function instead of the hyperbolic tangent function, dividing the change path of the intent determination adjustment factor with the first compensation factor into multiple linear compensation intervals with different slopes. Regardless of the specific function form used, a negative correlation adjustment link is established between the first compensation factor and the intent determination adjustment factor.
[0039] Order in the intent index value calculation model The value is adjusted based on the system's trade-off between false trigger rate and response speed; the order is... The change in the value of is an equivalent mathematical transformation of the intent index value calculation model. In one implementation, the order A value of 1 indicates that the intention index value calculation model exhibits basic proportional coupling characteristics. In another implementation, the order... The value is an integer greater than or equal to 2. Regardless of the order. The specific numerical values change while maintaining a consistent nonlinear coupling relationship between the first convergence probability, the limb pointing space vector, and the intent determination adjustment factor. Different order transformations of the intent index value calculation model are all completed under the logical control of the processing unit, achieving nonlinear gain adjustment for interactive intent determination by adjusting the order of the power function.
[0040] The evolution of the intent determination adjustment factor along with the first compensation factor within the extreme numerical range demonstrates the system's ability to repair feature mismatch. When the physiological load shift of the interacting subject approaches a preset maximum value, the value of the first compensation factor is calculated as a scalar value close to 1. The processing unit substitutes the first compensation factor with a value close to 1 into the intent determination adjustment factor calculation formula, resulting in a lower intent determination adjustment factor value output by the formula. At this time, due to the weakening convergence characteristics of the end-effector velocity sequence caused by physiological fatigue, the first convergence probability value output by the processing unit decreases synchronously. Since the decrease in the intent determination adjustment factor value compensates for the decrease in the first convergence probability value, the output result of the intent index value calculation formula remains above the first preset threshold.
[0041] When the intent convergence feature is at a critical state, i.e., the first convergence probability value is low, the system performs logical correction based on the spatial pointing accuracy of the limb pointing space vector. The processing unit extracts the limb pointing space vector and calculates the angle between the limb pointing space vector and the preset pointing center line. When the user points precisely to the interaction target, the value of this angle approaches 0 degrees, and the corresponding cosine function value approaches 1. Even if the first convergence probability output by the dynamic judgment logic is in the weak signal range, under the combined constraints of the low intent judgment adjustment factor after the first compensation factor correction and the high consistency pointing feature of the limb pointing space vector, the intent index value output by the intent index value calculation formula can still reach the strength to trigger the first control signal. This calculation process confirms the technical feasibility of the system using the complementary relationship of multimodal features to repair the inadequacy of single-dimensional feature expression.
[0042] The first control signal is generated by the core hub controller when the intent index value is determined to be greater than a first preset threshold, and then sent to the execution unit to execute the interactive response. The execution unit performs differentiated logic adaptation on the first control signal for different terminal types within the exhibition hall scene. In an implementation where the execution unit includes a large display terminal, the processing unit maps the first control signal to a focus offset command or menu item selection command for the display interface. In an implementation where the execution unit includes a mobile robot terminal, the processing unit converts the first control signal into navigation target coordinates of the interactive target point in the global coordinate system of the exhibition hall. In an implementation where the execution unit includes an intelligent lighting control terminal, the processing unit maps the first control signal to a step change command for lighting brightness or a beam pointing offset parameter. By performing multi-terminal logic transformation on the first control signal, the interactive system achieves the mapping between the first control signal and diverse interactive execution logic.
[0043] The anomaly monitoring step is executed in real-time by the processing unit during the operation of the interactive system. It is used to calculate the mismatch distance between the real-time interactive mapping trajectory and the reference trajectory. The processing unit obtains the mismatch distance by calculating the geometric deviation between the real-time acquired limb pointing spatial vector and the preset reference trajectory vector. This mismatch distance is defined using the L2 norm in Euclidean space. In cases where sudden environmental interference, drastic changes in lighting, or occlusion of the sensing unit causes a non-steady-state shift in the real-time interactive mapping trajectory, the processing unit monitors the fluctuations in the mismatch distance in real-time. If the mismatch distance exceeds a third preset threshold, the core hub controller determines that the current interactive data is in a logical mismatch state and immediately executes an action to suppress the first control signal.
[0044] The physical implementation environment of the intelligent exhibition hall interactive control system based on multimodal data consists of a sensing unit, a processing unit, and an execution unit. The depth camera and infrared tracking device in the sensing unit establish a communication connection with the processing unit via a data bus or network interface. The processing unit includes at least one processor and a computer-readable storage medium coupled to the processor. The computer-readable storage medium stores a computer program that implements the intelligent exhibition hall interactive control method based on multimodal data. When executing the computer program, the processor calls instructions from the memory to complete all computational steps, including multi-source raw data acquisition, spatiotemporal synchronization alignment, generation of the first compensation factor, calculation of the first convergence probability, adjustment of the intent determination adjustment factor, intent index coupling, and generation of the first control signal. After receiving the first control signal from the processing unit, the execution unit drives the display terminal, mobile robot, or controlled lighting equipment to complete the physical-level interactive response.
[0045] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A smart exhibition hall interactive control method based on multimodal data, characterized in that, Includes the following steps: Step a. Obtain multi-source raw data within the interaction space, and extract limb pointing spatial vectors, end-effector velocity sequences, and micro-movement high-frequency tremor sequences; Step b. Perform coordinate system affine transformation and timestamp synchronization alignment on the multi-source raw data; Step c. Calculate the power spectral density of the micro-motion high-frequency tremor sequence, and generate a first compensation factor based on the power spectral density; Step d. Perform a second-order difference operation on the end motion velocity sequence to obtain an acceleration sequence, calculate the fitting residual between the acceleration sequence and the preset decay model, and generate a first convergence probability based on the fitting residual; Step e. Adjust the first convergence probability and the limb pointing space vector with the first compensation factor as a variable to establish a real-time interactive mapping trajectory, and generate an intent index value based on the real-time interactive mapping trajectory; Step f. If the intent index value is greater than the first preset threshold, then generate the first control signal.
2. The method according to claim 1, characterized in that, The limb pointing space vector is obtained by calculating the geometric consistency parameters between the user's skeletal midline vector and the gaze vector.
3. The method according to claim 1, characterized in that, The process of generating the first compensation factor in step c includes: Extract the signal components from 6Hz to 12Hz from the micro-motion high-frequency tremor sequence; If the power spectral density amplitude of the signal component is greater than the second preset threshold, the first compensation factor is output according to the power spectral density amplitude.
4. The method according to claim 1, characterized in that, The preset attenuation model mentioned in step d is ,in Let be the convergence coefficient of the action. The first convergence probability is obtained by calculating the fitting residual between the jerk sequence and the preset decay model, and mapping the fitting residual to a normalized interval.
5. The method according to claim 1, characterized in that, Step e includes the step of determining the adjustment factor by adjusting the intention through a first compensation factor, wherein the first compensation factor is used This indicates that the intention determination adjustment factor is used. This indicates that the intent determination adjustment factor satisfy: in, Basic regulatory factor, The first preset parameter, This is the second preset parameter. This is the third preset parameter.
6. The method according to claim 5, characterized in that, The intent index value is used In other words, the first convergence probability is represented by... This indicates that the intent index value satisfy: in, The angle between the limb pointing spatial vector and the preset pointing center line is defined by the line connecting the geometric center of the candidate interactive target and the origin of the perception unit. It is the order.
7. The method according to claim 1, characterized in that, It also includes an anomaly monitoring step: calculating the mismatch distance between the real-time interactive mapping trajectory and the reference trajectory; if the mismatch distance is greater than a third preset threshold, then suppressing the first control signal.
8. An intelligent exhibition hall interactive control system based on multimodal data, characterized in that, include: The sensing unit is used to acquire multi-source raw data within the interaction space; A processing unit is configured to perform the method according to any one of claims 1 to 7 to generate a first control signal; An execution unit is configured to receive the first control signal and execute an interactive response.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method described in any one of claims 1 to 7.