Human-computer interaction method and system based on intelligent exhibition hall
By synchronously collecting and fusing visual and radar perception data, combined with an accumulation judgment mechanism and digital life image, a human-computer interaction system for the smart exhibition hall is constructed, which solves the problems of unreliable perception and inaccurate decision-making, and improves stability and personalized interactive experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING ZHENGYI QINGSHANG CULTURE MEDIA CO LTD
- Filing Date
- 2026-03-24
- Publication Date
- 2026-06-26
AI Technical Summary
The existing human-computer interaction system in smart exhibition halls is unreliable in perception, inaccurate in decision-making, and lacks learning ability in complex real-world scenarios, resulting in a rigid interactive experience that cannot continuously adapt to audience preferences.
By simultaneously collecting and fusing visual and radar perception data, and through the confidence concept of attribute tags and cumulative judgment mechanism, combined with the interactive behavior of digital life images, a closed loop of perception-decision-execution-evaluation-optimization is constructed to optimize and adjust parameters.
It improves the system's perception capabilities in complex environments, enhances the stability and smoothness of human-computer interaction, provides a highly personalized immersive experience, and continuously improves the system's intelligence level by learning and optimizing interaction strategies.
Smart Images

Figure CN122284816A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of smart exhibition hall and interactive technology, and in particular to human-computer interaction methods and systems based on smart exhibition halls. Background Technology
[0002] As a display platform integrating the Internet of Things, artificial intelligence, and multimedia technologies, smart exhibition halls are evolving from one-way information display to two-way intelligent interaction in their human-computer interaction methods. Existing technologies include solutions that use visual sensors or wearable devices to identify visitors and provide customized explanations or simple interactions accordingly. For example, facial recognition or posture estimation using cameras can trigger the playback of pre-recorded content.
[0003] However, the existing solutions mentioned above have significant shortcomings in their adaptive and long-term optimization capabilities under complex real-world scenarios. Specifically, the systems rely on a single, easily affected by ambient lighting and crowd obstruction, resulting in inaccurate identification of key interactive objects within a group of viewers. Their decision-making logic is largely based on simple rule matching of instantaneous recognition results, making them prone to misjudgments and interaction jitter when faced with recognition noise and dynamic changes in the audience. More importantly, once the system's operating strategy is set, it remains fixed and cannot learn and optimize based on historical data of actual interaction effects. This leads to the interaction strategy failing to continuously align with the audience's true preferences, limiting the level of intelligence. This restricts the depth and appeal of the human-computer interaction experience in smart exhibition halls. Summary of the Invention
[0004] To overcome the above shortcomings, this invention provides a human-computer interaction method and system based on smart exhibition halls, aiming to improve the problems of rigid interactive experience and inability to continuously adapt to audience preferences caused by unreliable perception, inaccurate decision-making and lack of learning ability in existing smart exhibition hall human-computer interaction systems.
[0005] In a first aspect, the present invention provides the following technical solution: a human-computer interaction method based on a smart exhibition hall, comprising the following steps: S1. Within the preset interactive area, simultaneously collect visual perception data and radar perception data of at least one target audience member; S2. Based on the visual perception data and radar perception data, perform data fusion and target tracking to generate and output user status information containing target audience attribute tags and their confidence levels; S3. Receive the user status information and determine the current target interaction mode from multiple pre-stored interaction modes based on the cumulative judgment result of the confidence of the attribute tags within a continuous time window. S4. Invoke the multimedia resources corresponding to the target interaction mode, and drive the digital life image to perform interactive behaviors that match the target interaction mode; S5. During and after the digital life image performs interactive behavior, collect audience feedback data within the interactive area; S6. The user status information, the target interaction mode, and the audience feedback data are associated and stored to form a historical optimization data set; S7. Based on the historical optimized data set, perform parameter optimization and adjustment on the data fusion and target tracking process in step S2 and the cumulative judgment process in step S3.
[0006] Preferably, in step S1, the step of simultaneously acquiring visual perception data and radar perception data of at least one target audience member within a preset interactive area specifically includes: By deploying at least one wide dynamic range camera at different angles in the interactive area, an image sequence containing the target audience is acquired as the visual perception data; At least one millimeter-wave radar deployed in the interactive area is used to collect the outline point cloud and motion trajectory information of the target audience as radar perception data. The image sequence is time-stamped and spatially aligned with the contour point cloud and motion trajectory information to complete the synchronous acquisition.
[0007] Preferably, in step S2, the step of performing data fusion and target tracking based on the visual perception data and radar perception data to generate and output user status information containing target audience attribute tags and their confidence levels specifically includes: The image sequence is associated and matched with the contour point cloud and motion trajectory information to separate, identify and continuously track multiple audience targets within the interaction area, and to determine the main interaction target as the target audience. Based on the analysis of the target audience in the image sequence, the first attribute estimation result and the corresponding first confidence level are obtained; Based on the analysis of the target audience in the contour point cloud and motion trajectory information, the second attribute estimation result and the corresponding second confidence level are obtained; The first attribute estimation result, the second attribute estimation result, the first confidence score, and the second confidence score are weighted and fused to generate user status information containing the target audience attribute labels and their comprehensive confidence scores.
[0008] Preferably, in step S3, the step of receiving the user state information and determining the current target interaction mode from a plurality of pre-stored interaction modes based on the cumulative judgment result of the confidence of the attribute labels within a continuous time window specifically includes: Receive user status information containing the target audience's attribute tags and their overall confidence levels; Based on a preset time window length, a moving average filter is applied to the current overall confidence level and its historical values of the attribute label to obtain the cumulative confidence level evaluation value of the attribute label. The cumulative confidence score is compared with a preset confidence threshold, and the current target interaction mode is determined from the pre-stored multiple interaction modes according to the comparison result and the preset mapping rule. After determining the target interaction mode, the target interaction mode is maintained for at least a preset minimum duration, during which new user state information is ignored from triggering the interaction mode switching.
[0009] Preferably, in step S4, the step of calling multimedia resources corresponding to the target interaction mode and driving the digital life image to perform interactive behaviors matching the target interaction mode specifically includes: Based on the target interaction mode, the corresponding digital life 3D model, dynamic behavior script and narration audio are called from the preloaded multimedia resource library; The digital life 3D model is rendered on the display device of the interactive area, and the model is driven to perform corresponding actions and expressions based on the dynamic behavior script; The audio of the narration is played synchronously, and the lip movements of the digital life 3D model are matched with the audio of the narration to form the interactive behavior.
[0010] Preferably, in step S5, the step of collecting audience feedback data within the interactive area during and after the digital life avatar performs interactive behavior specifically includes: The system uses a wide dynamic range camera to capture images of the audience's facial orientation, dwell time, and crowd density distribution within the interactive area. Millimeter-wave radar is used to collect information on the movement trajectory changes of the audience and the number of people remaining in the area within the interactive area. Based on the facial orientation, dwell time, gathering heat distribution images, movement trajectory changes, and the number of people remaining in the area, quantitative index data reflecting audience participation and interest are extracted and generated as feedback data for the audience group.
[0011] Preferably, the step of extracting and generating quantitative indicator data reflecting audience participation and interest as audience feedback data specifically includes: Based on the facial orientation and the clustering heat distribution image, the proportion of viewers facing the main display area of the digital life image is calculated as the first engagement indicator; Based on the dwell time and the number of people staying in the area, the average dwell time of the audience during the interaction behavior execution cycle is calculated as an interest index. Based on the changes in the movement trajectory, the number of trajectories that approach the core location of the interaction area or perform repetitive interactive actions are identified and counted as a second engagement indicator. The first engagement index, the interest index, and the second engagement index are normalized and weighted to generate the quantitative index data.
[0012] Preferably, in step S6, the step of associating and storing the user status information, the target interaction mode, and the audience feedback data to form a historical optimization data set specifically includes: For each complete human-computer interaction session executed from S1 to S4, generate an associated record; The associated record stores the attribute tags in the user status information that triggered the current session, the determined target interaction mode, and the quantitative indicator data corresponding to the current session. The multiple related records are aggregated in chronological order to form the historical optimization data set.
[0013] Preferably, in step S7, the step of optimizing and adjusting the parameters of the data fusion and target tracking process in step S2 and the cumulative judgment process in step S3 based on the historical optimization data set specifically includes: From the historical optimization data set, inefficient interaction records where the quantitative indicator data is lower than the preset feedback threshold are selected; Analyze the correspondence between the attribute tags stored in the inefficient interaction records and the target interaction mode, and adjust the mapping rules or confidence thresholds from attribute tags to interaction modes in S3. Extract the original perceptual data corresponding to the inefficient interaction records, and retrain or correct the parameters of the analysis model used to generate the first attribute estimation result and / or the second attribute estimation result in S2.
[0014] Secondly, this invention provides the following technical solution: a human-computer interaction system based on a smart exhibition hall, comprising: The perception data acquisition module is used to simultaneously acquire visual perception data and radar perception data of at least one target audience member within a preset interactive area. The state estimation module is used to perform data fusion and target tracking based on the visual perception data and radar perception data, and generate and output user state information containing target audience attribute tags and their confidence levels. The interaction decision module is used to receive the user status information and determine the current target interaction mode from multiple pre-stored interaction modes based on the cumulative judgment result of the confidence of the attribute tags within a continuous time window. The content execution module is used to call multimedia resources corresponding to the target interaction mode and drive the digital life image to perform interactive behaviors that match the target interaction mode. The feedback collection module is used to collect audience feedback data within the interactive area during and after the digital life image performs interactive behavior. The data association module is used to associate and store the user status information, the target interaction mode, and the audience feedback data to form a historical optimization data set. The parameter optimization module is used to optimize and adjust the parameters of the data fusion and target tracking process in the state estimation module and the cumulative judgment process in the interactive decision module based on the historical optimization data set.
[0015] The present invention has the following beneficial effects: 1. In this invention, by synchronously collecting and fusing visual and radar perception data, the inherent defects of a single visual sensor being susceptible to changes in ambient lighting and mutual occlusion by the audience are effectively overcome, providing a more reliable and comprehensive raw data foundation for subsequent processing and improving the system's perception capability in the complex environment of a real exhibition hall.
[0016] 2. In this invention, by introducing the concept of confidence level for attribute labels and a cumulative judgment mechanism within a continuous time window, the system's decision-making does not rely on potentially erroneous instantaneous identification results, but rather on a smooth and prudent judgment based on trends over a period of time. This significantly reduces false triggering and frequent switching of interaction modes caused by instantaneous interference, enhancing the stability and smoothness of the human-computer interaction process.
[0017] 3. In this invention, the most matching target interaction mode is dynamically determined from multiple pre-stored interaction modes based on the result of accumulated confidence judgment, and the digital life image is driven to perform interaction behavior that is completely consistent with it. This enables the interactive content to intelligently match the understanding level and interest preferences of different attribute audiences, thereby providing a highly personalized and contextually coherent immersive experience.
[0018] 4. In this invention, by collecting group feedback data after interaction and associating it with the user status information that triggered the interaction and the interaction mode adopted, a historical optimization data set that can be quantitatively evaluated for interaction effect is constructed, transforming the vague experience into analyzable structured data.
[0019] 5. In this invention, based on historical optimization data sets, the parameters of the front-end perception fusion model and decision rules are optimized and adjusted in a targeted manner, forming a complete closed loop of perception-decision-execution-evaluation-optimization. This enables the system to learn from actual operation, automatically strengthen effective interaction strategies, and correct ineffective matching relationships, thereby allowing the interaction strategies of the entire system to continuously evolve. Attached Figure Description
[0020] Figure 1 This is a flowchart illustrating the human-computer interaction method based on a smart exhibition hall proposed in this invention. Figure 2 This is a schematic diagram of the architecture of the human-computer interaction system based on a smart exhibition hall proposed in this invention. Detailed Implementation
[0021] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0022] Example 1: In the first embodiment of the present invention, the present invention provides a human-computer interaction method based on a smart exhibition hall, such as... Figure 1 As shown, it includes the following steps: S1. Within the preset interactive area, simultaneously collect visual perception data and radar perception data of at least one target audience member.
[0023] Furthermore, in S1, the step of simultaneously collecting visual perception data and radar perception data of at least one target audience member within a preset interactive area specifically includes: At least one wide dynamic range camera is deployed at different angles in the interactive area to capture image sequences containing the target audience as visual perception data. At least one millimeter-wave radar deployed in the interactive area is used to collect the outline point cloud and motion trajectory information of the target audience as radar perception data. The image sequence, contour point cloud, and motion trajectory information are time-stamped and aligned with spatial coordinates to complete synchronous acquisition.
[0024] Specifically, at least two webcams with wide dynamic range capabilities are deployed at the top or upper side of the interactive area. These two cameras cover the entire interactive area from different perspectives, ensuring that at least one camera clearly captures the target audience at any location, thus avoiding occlusion issues from a single perspective. The cameras continuously acquire images at a preset frame rate to form an image sequence. Simultaneously, at least one millimeter-wave radar sensor is deployed near the boundary of the interactive area. This millimeter-wave radar sensor emits frequency-modulated continuous waves and receives their echoes, generating point cloud data through signal processing. Each point in this point cloud data contains information on distance, azimuth, elevation, and radial velocity, thereby reflecting the contour, position, and trajectory of targets within the area. To achieve synchronization between visual perception data and radar perception data, the system employs a hardware-based unified clock source to provide a reference for the timestamps of all sensors. Each image frame and each radar point cloud data frame is marked with a precise timestamp based on this unified clock source upon generation. For spatial coordinate alignment, the system performs a calibration process after deployment. This process uses a calibration board of known size simultaneously within the camera's field of view and the radar's detection range. By identifying the corner coordinates of the calibration board in the image and their corresponding 3D corner coordinates in the radar point cloud, a spatial transformation matrix is calculated to convert the camera image pixel coordinate system to the radar's 3D coordinate system. This transformation matrix is saved and applied in subsequent real-time data streams. Through this timestamp synchronization and spatial coordinate alignment, the system ensures that the image sequence information, contour point cloud, and motion trajectory information of a target at the same time and spatial location can be accurately correlated, thus completing synchronous acquisition. The calculation of the spatial transformation matrix involves the camera imaging model and coordinate transformation. A typical calibration method obtains the transformation parameters by solving the least-squares solution of the following system of equations, where the three-dimensional coordinates of the i-th feature point on the calibration board in the radar coordinate system are: The two-dimensional coordinates in the image pixel coordinate system are A camera model is typically described using an intrinsic parameter matrix K and a rotation / translation matrix [R|t], satisfying: ; in, It is a non-zero scaling factor. This is achieved by collecting multiple corresponding data sets. and The rotation matrix R and translation vector t of the camera relative to the radar coordinate system can be solved, thus establishing the spatial transformation relationship. The coordinates of the millimeter-wave radar itself are predetermined by its installation position and orientation.
[0025] S2. Based on visual perception data and radar perception data, perform data fusion and target tracking to generate and output user status information containing target audience attribute tags and their confidence levels.
[0026] Furthermore, in S2, the steps of data fusion and target tracking based on visual perception data and radar perception data, generating and outputting user state information containing target audience attribute labels and their confidence levels, specifically include: Image sequences are correlated and matched with contour point clouds and motion trajectory information to separate, identify and continuously track multiple audience targets within the interactive area, and identify the main interactive target as the target audience. Based on the analysis of the target audience in the image sequence, the estimation results of the first attribute and the corresponding first confidence level are obtained; Based on the analysis of the target audience in the contour point cloud and motion trajectory information, the second attribute estimation results and the corresponding second confidence scores are obtained; The estimation results of the first attribute and the second attribute, as well as the first confidence level and the second confidence level, are weighted and fused to generate user status information that includes the target audience attribute labels and their comprehensive confidence levels.
[0027] Specifically, the system first performs the following operations on each frame of data: the contour point cloud extracted from the radar perception data is divided into multiple point cloud clusters using a clustering algorithm, with each cluster representing a potential viewer target. The system calculates the 3D centroid coordinates, bounding box size, and average radial velocity of each point cloud cluster as the radar features of that target. Simultaneously, the system processes the image sequences in the visual perception data. Using a deep learning-based target detection model, such as YOLO or Faster R-CNN architecture, the images are analyzed in real time to detect all human bounding boxes in the images. This target detection model is pre-trained using a dataset of human images containing various poses, clothing, and lighting conditions to achieve robust human detection capabilities. Subsequently, the system performs association matching. Using the aligned spatiotemporal coordinates mentioned above, the 3D centroid coordinates of the target in the radar features are projected onto the image pixel plane. The system calculates the Euclidean distance between each projection point and the center point of the human bounding box detected in the image. It then associates the radar target closest to it (within a preset threshold) with the visual target, marking them as the same viewer target and assigning them a unique identifier. For associated targets, the system uses a Kalman filter for continuous tracking. The Kalman filter uses the target's position and velocity as state variables, predicting the state of the current frame from the state of the previous frame and updating it using the actual observations of the current frame. In this way, the system can maintain the identification of each viewer target across consecutive frames, achieving both separation and continuous tracking. The system determines the primary interactive target based on preset judgment logic. This judgment logic includes, but is not limited to: calculating the average distance between each tracked target and a preset core point in the interaction area within a preset time window; and calculating the percentage of frames in which each target's body is facing the main screen within the preset time window. The system comprehensively calculates the distance score and orientation score of each target, identifies the target with the highest weighted total score as the primary interactive target, and marks it as the target audience for this round of interaction. For a identified target audience, the system extracts the corresponding human body region image from the image sequence. This region image is then input into a pre-trained attribute classification model. This attribute classification model is a convolutional neural network (CNN), whose output layer corresponds to different attribute labels, such as "child," "teenager," "adult," and "elderly." The last layer of the model uses the Softmax function, and the value of each element in its output vector represents the probability that the input image belongs to the corresponding category. The system uses this probability value as the first confidence score obtained based on visual analysis and takes the category with the highest probability as the first attribute estimation result. The training process of this CNN uses a dataset of face or full-body images labeled with age stages, and supervised training is performed using the cross-entropy loss function to optimize the network weight parameters. Simultaneously, the system analyzes the radar data of the target audience, including its outline point cloud and historical movement trajectory. The target's height is estimated from the point cloud, calculated as the difference between the maximum and minimum coordinate values of the point cloud cluster in the vertical direction. The system pre-defines typical height ranges for different age groups; for example, those shorter than 1.3 meters are more likely to be children. Based on the degree to which the height estimate falls within the preset range, the system calculates a membership score as a second confidence level based on radar analysis, and uses the main age group corresponding to this range as the second attribute estimation result. Furthermore, the system analyzes the target's movement trajectory information, calculating its average speed and trajectory complexity. For example, children's movement speed may change more rapidly, and their trajectories more irregular. This movement characteristic can serve as an auxiliary criterion in the comprehensive calculation of the second attribute estimation result and the second confidence level. The system fuses attribute estimation results and confidence scores from different modalities. First, it compares the consistency of the first attribute estimation result with the second attribute estimation result. If both point to the same attribute label, this label is directly determined as the final target audience attribute label, and a weighted average method is used to calculate the overall confidence score. The formula for the weighted average method is as follows: ; in, Indicates the overall confidence level. Indicates the first confidence level. This indicates the second confidence level. and These are preset weighting factors for the visual modality and the radar modality, respectively. Their values can be adjusted based on the historical reliability of the sensors, for example, both are preset to 0.5. If the estimation results of the first attribute and the second attribute are inconsistent, the system will compare... and The size of the value determines the attribute label corresponding to the higher confidence level as the final target audience attribute label, and its confidence level is used as the... Finally, the system generates and outputs target audience attribute labels and their overall confidence scores. User status information.
[0028] S3. Receive user status information and determine the current target interaction mode from multiple pre-stored interaction modes based on the cumulative judgment result of the confidence of the attribute labels within a continuous time window.
[0029] Furthermore, in S3, the steps of receiving user state information and determining the current target interaction mode from multiple pre-stored interaction modes based on the cumulative judgment result of the confidence of the attribute labels within a continuous time window specifically include: Receive user status information containing target audience attribute tags and their overall confidence levels; Based on a preset time window length, a moving average filter is applied to the current overall confidence level of the attribute label and its historical values to obtain the cumulative confidence level evaluation value of the attribute label. The cumulative confidence score is compared with the preset confidence threshold, and the current target interaction mode is determined from multiple pre-stored interaction modes based on the comparison result and the preset mapping rule. After determining the target interaction mode, maintain the target interaction mode for at least a preset minimum duration, during which new user state information is ignored from triggering the interaction mode switch.
[0030] Specifically, the system receives user status information. This information includes target audience attribute tags and their corresponding overall confidence scores. The system maintains a fixed-length circular buffer to cache recently received user status information in chronological order. The length of this buffer is the preset time window length. The system extracts historical overall confidence scores belonging to the same target audience attribute tag from the circular buffer and combines them with the latest overall confidence score to form a confidence score sequence. The system performs a moving average filter on this sequence to obtain the cumulative confidence score evaluation value for that attribute tag. The formula for the moving average filter is as follows: ; in, This represents the cumulative confidence score calculated at time t. N represents the number of data points involved in the calculation, which is determined by the preset time window length and the system data update frequency. For example, if the time window length is 2 seconds and the update frequency is 10 Hz, then N = 20. This represents the overall confidence level acquired at time ti. This calculation effectively smooths out any instantaneous fluctuations in the overall confidence level over time. The system will calculate the cumulative confidence score. The system compares the results with a preset confidence threshold. This confidence threshold is a configurable parameter that sets the minimum confidence level required for the system to make a definitive judgment. The system also pre-stores multiple interaction modes, such as "Children's In-Depth Exploration Mode," "Teenager Challenge Mode," "Adult Professional Explanation Mode," and "General Guided Tour Mode." Furthermore, it pre-stores mapping rules from attribute labels to interaction modes. These mapping rules define the interaction modes to be triggered for different attribute labels and their corresponding cumulative confidence evaluation value ranges. The decision logic is as follows: If the cumulative confidence evaluation value corresponding to the current target audience's attribute label is... If the confidence score is greater than or equal to the confidence threshold, the system maps the attribute label to the corresponding specific interaction mode according to the mapping rules, and identifies that mode as the target interaction mode. If the cumulative confidence score... If the value is below the confidence threshold, or the attribute label is "unknown", the system will determine "general navigation mode" as the target interaction mode according to the mapping rules. To prevent frequent switching of interaction modes due to brief visits by viewers or momentary sensor interference, the system starts a minimum duration timer after determining the target interaction mode. Within this minimum duration period, for example, set to 90 seconds, the system will lock the current target interaction mode. During this period, even if newly received user status information, after the aforementioned calculations, indicates a switch to another interaction mode, the system will ignore the switch request and maintain the current interaction mode unchanged until the minimum duration ends. After the minimum duration ends, the system resumes its normal response and decision-making process for new user status information.
[0031] S4. Call the multimedia resources corresponding to the target interaction mode to drive the digital life image to perform interactive behaviors that match the target interaction mode.
[0032] Furthermore, in S4, the specific steps of calling multimedia resources corresponding to the target interaction mode and driving the digital life image to perform interactive behaviors that match the target interaction mode include: Based on the target interaction mode, the corresponding digital life 3D model, dynamic behavior script and narration audio are called from the pre-loaded multimedia resource library; The digital life 3D model is rendered on the display device in the interactive area, and the model is driven to perform corresponding actions and expressions based on dynamic behavior scripts; Synchronize the audio of the narration and match the lip movements of the digital life 3D model with the audio of the narration to create interactive behavior.
[0033] Specifically, the system maintains a pre-loaded multimedia resource library. This library is stored on a local server in the interactive area or in high-speed solid-state storage to ensure fast retrieval. The resource library contains multiple resource packages corresponding one-to-one with different interaction modes. Each resource package includes at least one digital life 3D model file, in FBX or GLTF format, defining the character's mesh, skeleton, materials, and basic textures; a dynamic behavior script file, written in Extensible Markup Language or a specific scripting language, defining a series of action sequences and facial expression changes that the digital life model must perform during the current interaction cycle; and one or more audio files containing narration, whose wording, speaking speed, and tone match the target interaction mode and target audience attribute tags. For example, the audio for the "children" mode uses simpler vocabulary, a slower speaking speed, and a more lively tone. Once the system determines the target interaction mode, it precisely retrieves the corresponding digital life 3D model file, dynamic behavior script file, and narration audio file from the resource library based on the mode identifier and loads them into the runtime memory. The system uses a real-time rendering engine, such as a rendering module built on frameworks like OpenGL, DirectX, Unity, or Unreal Engine. This rendering engine parses and instantiates the loaded digital life 3D model file, then renders the model at a preset position on the screen based on the coordinates and viewpoint parameters of the interactive display area. Simultaneously, the system launches a script parsing and execution module. This module reads and parses dynamic behavior script files, and according to the timestamps and action / expression instructions defined in the script, drives the digital life model to perform corresponding actions and expressions by changing the rotation angles and displacements of the 3D model's skeleton and adjusting the shape or material blending weights of the model's facial bones. This process is procedural, ensuring consistency in behavior each time the same script is executed. The system plays the narration audio file called by the audio playback module. To achieve a realistic interactive effect, the system needs to ensure that the lip movements of the digital life model are precisely synchronized in time with the played audio content. Therefore, the system performs real-time or pre-computed speech analysis while playing the audio. Specifically, the system performs short-time Fourier transform or Mel-frequency cepstral coefficient extraction on the audio stream to analyze its spectral characteristics. Based on the analyzed audio features, the system generates the corresponding lip bone control parameter sequence through a pre-trained lip-shape driven model. The construction and training process of this lip-shape driven model is as follows: A dataset containing a large amount of speech audio and its corresponding lip-shape videos is used. The lip keypoint coordinates of each frame are extracted from the video as training labels, and MFCC features are extracted from the corresponding audio segments as input features. A neural network model, such as a long short-term memory network or a temporal convolutional network, is constructed. The model is trained by minimizing the mean squared error between the model's predicted lip keypoint coordinates and the true coordinates, enabling the model to learn the mapping relationship from audio features to lip movements. After training, the model is deployed in the system. During runtime, the system inputs real-time audio features into the model, and the model outputs lip bone control parameters for each frame. The rendering engine adjusts the pose of the digital life model's lip bones in real time based on these parameters, thereby generating realistic lip movements synchronized with the audio. Through the above resource scheduling, rendering drive and audio-visual synchronization processing, the system ultimately forms a complete interactive behavior that matches the target interaction mode.
[0034] S5. Collect audience feedback data within the interaction area during and after the digital life avatar performs interactive behaviors.
[0035] Furthermore, in S5, the steps for collecting audience feedback data within the interaction area during and after the digital life avatar performs interactive behaviors specifically include: Using a wide dynamic range camera, images of the audience's facial orientation, dwell time, and crowd density distribution within the interactive area are captured. Millimeter-wave radar is used to collect information on changes in the movement trajectory of the audience within the interactive area and the number of people remaining in the area. Based on facial orientation, dwell time, gathering heat distribution images, changes in movement trajectory, and information on the number of people remaining in the area, quantitative indicator data reflecting audience participation and interest are extracted and generated as audience feedback data.
[0036] Furthermore, the steps of extracting and generating quantitative indicators reflecting audience participation and interest as audience feedback data specifically include: Based on images of facial orientation and clustering heat distribution, the proportion of viewers facing the main display area of the digital life image is calculated as the primary engagement indicator. Based on dwell time and the number of people staying in the area, the average dwell time of the audience during the interaction behavior execution cycle is calculated as an interest index. Based on changes in movement trajectories, the number of trajectories that move toward the core location of the interaction area or perform repetitive interactive actions is identified and counted as a second engagement indicator. The first engagement index, interest index, and second engagement index are normalized and weighted to generate quantitative index data.
[0037] Specifically, the system continues to acquire image sequences using the wide dynamic range cameras deployed above. First, the system uses the same object detection model described above to identify the bounding boxes of all spectators in each frame. For each identified spectator, the system further uses a pre-trained facial orientation estimation model, which employs a convolutional neural network architecture. This network takes an image of the upper body region as input, with the image size normalized to 224 pixels high and 224 pixels wide. The network contains the following core modules: Feature extraction module: composed of multiple convolutional layers, batch normalization layers, and modified linear unit activation function layers stacked together, used to extract multi-level feature maps from the input image. Specifically, it can adopt the backbone structure of a ResNet-50 network, removing its final fully connected classification layer; Global average pooling layer: performs global average pooling on the feature map output by the feature extraction module, converting it into a fixed-length feature vector; Fully connected regression module: consists of two consecutive fully connected layers. The first fully connected layer maps the feature vector to a 512-dimensional vector and passes it through the ReLU activation function. The second fully connected layer maps the 512-dimensional vector to a 3-dimensional output vector, which corresponds to the three Euler angles of the head attitude: yaw, pitch and roll. The model is trained using supervised learning, requiring a dataset with real head pose annotations. Public datasets such as 300W-LP or AFLW2000 are used, providing a large number of face images and their corresponding accurate Euler angle annotations. The specific training steps are as follows: Input preparation: Read the images and their corresponding ground truth Euler angles (yaw, pitch, roll) from the training dataset. Loss function definition: Use the mean absolute error as the loss function, calculating the mean of the sum of the absolute errors between the model's predicted Euler angles (yaw^, pitch^, roll^) and the ground truth values. The loss function L is defined as: ; The Adam optimizer is used to minimize the loss function L. The initial learning rate is set to 0.001, and it decays to 0.1 every 20 training epochs. The batch size is set to 64. The total number of training epochs is set to 100. The prepared training data is input into the network, and all weight parameters in the network are iteratively updated using the backpropagation algorithm until the loss function converges or the preset number of training epochs is reached. The trained model is then embedded and integrated into the system. During runtime, the system crops an image of the viewer's upper body (based on a human bounding box), scales it to 224x224 pixels, and inputs it into the model. The model's output is the estimated three Euler angles of the viewer's head in the current frame. Based on business logic, the system primarily uses yaw and pitch angles. By determining whether these two angles are simultaneously within a preset reasonable range, it identifies whether the viewer is "facing the main display area of the digital life image." The system continues to utilize the millimeter-wave radar deployed above to acquire point cloud data. Through clustering and tracking methods, the system continuously tracks the movement trajectories of all viewers within the interactive area. Movement trajectory change refers to the sequence of positions of a tracked target at consecutive points in time. The number of people remaining in the area refers to the total number of tracked targets located within the predefined boundaries of the interactive area at any given time. Based on the raw data collected above, the system calculates the following quantitative indicator: specifically, within a statistical period, the system accumulates the total number of frames among all detected viewers that are determined to be "facing the screen," and divides this by the total number of frames detected for all viewers to obtain the average proportion of viewers facing the screen. The calculation formula is as follows: ; in, This represents the primary engagement metric. T represents the total number of sampled frames within the statistical period. This represents the total number of viewers detected in frame t. ( This is an indicator function that takes a value of 1 when viewer p is determined to be "facing the screen" in frame t, and 0 otherwise. This indicator directly reflects the degree of viewer attention concentration. The system records the time difference between entering and leaving the interactive area for each tracked visitor, which is taken as their dwell time. For visitors still present at the end of the statistical period, their dwell time is calculated up to the current moment. The interest index is the arithmetic mean of the dwell times of all visitors within that statistical period. The calculation formula is as follows: ; Where I represents the interest index, and M represents the total number of unique viewers who appeared in the interaction area during the statistical period. This represents the dwell time of the m-th viewer. This metric reflects the time cost that viewers are willing to pay for this interactive content; The system analyzes the movement trajectory changes of each tracked audience member. When the system detects that an audience member's movement trajectory shows a trend of continuously approaching the core position of the interaction area, and the final position is less than a preset threshold, it is recorded as an "active approach" behavior. In addition, if the system identifies a repetitive specific movement pattern of the audience member's body based on detailed analysis of the radar point cloud, and this movement pattern is consistent with the interactive intent designed for the current interactive behavior, it is recorded as an "interactive action". Within a statistical period, the system counts the number of independent audience trajectory events that occur as "active approach" or "interactive action", and this number is the second engagement index. This indicator reflects the audience's willingness and behavior to actively participate in the interaction; The system normalizes the three indicators mentioned above to eliminate dimensional differences. Normalization uses the minimum-maximum method, based on historical data or preset theoretical maximum values, and can be expressed as: ; in, , and These are the normalized index values, , and This corresponds to the preset maximum reference value; Subsequently, the system calculates the comprehensive quantitative index data Q through weighted summation: ; in, , and The preset weighting coefficients are used, and they satisfy the following conditions: The weighting coefficients can be adjusted according to the feedback dimensions that are expected to be emphasized in different interaction modes. The final Q-value is used as the audience feedback data representing the effect of this interaction.
[0038] S6. Link and store user status information, target interaction patterns and audience feedback data to form a historical optimization data set.
[0039] Furthermore, in S6, the steps of associating and storing user status information, target interaction patterns, and audience feedback data to form a historical optimization dataset specifically include: For each complete human-computer interaction session executed from S1 to S4, generate an associated record; The associated record stores the attribute tags in the user status information that triggered the current session, the determined target interaction mode, and the quantitative indicator data corresponding to the current session. Multiple related records are aggregated in chronological order to form a historical optimization data set.
[0040] Specifically, the system defines a complete human-computer interaction session as follows: from step S1, detecting and initiating a new primary interaction target; through steps S2 and S3, generating a decision; to step S4, completing a full cycle of interaction behavior; and finally, in step S5, collecting the corresponding feedback data for that cycle. Each time the system completes such a human-computer interaction session, it automatically generates a new associated record. Each associated record is assigned a unique session identifier and records the session's start and end timestamps. In this associated record, the system stores the following three types of core data generated in this session in the form of key-value pairs or structured fields: Triggering condition data: stores the core elements of the user status information finally output in step S2. Specifically, it stores the target audience attribute tags for this session. For example, storing the string "child". Additionally, it can selectively store the comprehensive confidence score corresponding to this attribute tag; Execution action data: stores the target interaction mode for this session determined in step S3. For example, storing mode identifiers such as "MODE_CHILD_EXPLORE"; Effect feedback data: stores the quantitative indicator data calculated in step S5 corresponding to the time range of this session. That is, the audience feedback data representing the effect of this interaction, stored as a numerical value Q; The system persistently stores each generated associated record in a database or file in an append-only manner. These chronologically accumulated associated records together constitute a historical optimization dataset, which is a structured dataset in which each record explicitly contains a complete triplet of information: "under what audience attribute conditions," "what interaction strategy was executed," and "what effect was produced." The size of the dataset grows linearly with the system's runtime. The system can create an index for this historical optimization dataset so that subsequent steps in S7 can efficiently retrieve and analyze relevant records based on different query conditions. Through the above steps, the system transforms discrete interactive events into systematic empirical data that can be used for analysis, providing the necessary and structured input for data-driven parameter optimization.
[0041] S7. Based on the historical optimization data set, optimize and adjust the parameters of the data fusion and target tracking process in step S2 and the cumulative judgment process in step S3.
[0042] Furthermore, in S7, the steps of optimizing and adjusting the parameters of the data fusion and target tracking process in step S2 and the cumulative judgment process in step S3 based on the historical optimization data set specifically include: From the historical optimization dataset, inefficient interaction records with quantitative indicators below a preset feedback threshold are selected; Analyze the correspondence between the attribute tags stored in the inefficient interaction records and the target interaction mode, and adjust the mapping rules or confidence thresholds from attribute tags to interaction modes in S3. Extract the original perceptual data corresponding to the inefficient interaction records, and retrain or correct the parameters of the analysis model used in S2 to generate the first attribute estimation result and / or the second attribute estimation result.
[0043] Specifically, the system sets a preset feedback threshold. This threshold represents a lower bound for the quantified indicator data Q. The system periodically queries the historical optimization dataset and filters out all datasets that meet the criteria. The conditions are associated with records, and these records are marked as inefficient interaction records, forming a subset of inefficient records to be analyzed; The system analyzes the data patterns stored in the inefficient record subset, focusing on the correspondence between attribute tags and target interaction patterns. The system statistically analyzes the frequency and average Q-value of each associated target interaction pattern when a specific attribute tag appears in the inefficient record subset. Based on this statistical result, the system makes one or more adjustments to the decision logic in step S3 as follows: Adjusting mapping rules: If it is found that an attribute tag frequently generates inefficient records when mapped to pattern A, but mostly generates efficient records when mapped to pattern B, the system automatically updates the pre-stored mapping rules in step S3, changing the default mapping of that attribute tag from pattern A to pattern B; Adjusting confidence threshold: If it is found that inefficient records often occur when the overall confidence of the attribute tag is in a low range, the system can appropriately increase the confidence threshold required for pattern matching of that attribute tag in step S3. This makes the system more inclined to choose a safe "general navigation mode" rather than a potentially erroneous specific mode when the confidence is insufficient. For each record in the inefficient record subset, the system retrieves the original perceptual data corresponding to that session during the processing in steps S1 and S2 from the log system based on its session identifier and timestamp. Specifically, this includes image sequence fragments collected from the target audience when the session was triggered, along with the corresponding radar contour point cloud and motion trajectory information. The system uses this retrieved original perceptual data and its corresponding, already determined inefficient interaction results to optimize the attribute analysis model in step S2. Specifically, the model is retrained: the system uses the retrieved image sequence fragments as new training samples. For each sample, the training label is not the originally identified attribute label, but a corrected label assigned after reverse inference based on the inefficient result of the interaction or after manual verification. For example, an inefficient interaction labeled "child" might, upon verification, reveal that the audience is actually a "teenager." The system uses these new samples with corrected labels to incrementally train the convolutional neural network classification model that generated the first attribute estimation result in step S2. Incremental training uses a small learning rate to fine-tune the original model parameters to correct judgment biases in certain feature patterns while retaining most of its existing knowledge. The loss function and optimizer used in the training process remain consistent with the original training. Parameter correction: For parameterized methods such as height estimation based on radar data, the system can statistically analyze the actual radar point cloud features of the target audience in inefficient records and compare them with the preset age-height model. If a systematic bias is found, the parameters in the model are corrected. After completing the above optimizations, the updated mapping rules, confidence thresholds, and fine-tuned perception model parameters will be saved. The system can load these new parameters upon the next startup or replace the original parameters during operation via hot updates, thus enabling subsequent human-computer interaction sessions to be executed based on the optimized logic. Through this optimization process, the system can automatically identify and correct perceptual misjudgments and decision-making errors that lead to poor interaction effects, enabling the system to continuously improve itself in real operating environments.
[0044] Example 2: Existing human-computer interaction technologies in smart exhibition halls lack the ability to learn and optimize based on actual interaction effects. This results in fixed interaction strategies failing to continuously adapt to the real preferences of the audience, making it difficult to achieve a long-term, stable, and evolutionary intelligent experience. To address these issues, this invention provides a human-computer interaction system for smart exhibition halls, the structure of which is as follows: Figure 2 As shown. The specific implementation process of this system is as follows: The perception data acquisition module is used to simultaneously acquire visual perception data and radar perception data of at least one target audience member within a preset interactive area. The state estimation module is used to perform data fusion and target tracking based on visual perception data and radar perception data, and to generate and output user state information containing target audience attribute labels and their confidence levels. The interaction decision module is used to receive user status information and determine the current target interaction mode from multiple pre-stored interaction modes based on the cumulative judgment results of the confidence of the attribute labels within a continuous time window. The content execution module is used to call multimedia resources corresponding to the target interaction mode and drive the digital life image to perform interactive behaviors that match the target interaction mode. The feedback collection module is used to collect audience feedback data within the interactive area during and after the digital life avatar performs interactive behaviors. The data association module is used to associate and store user status information, target interaction patterns and audience feedback data to form a historical optimization data set; The parameter optimization module is used to optimize and adjust the parameters of the data fusion and target tracking process in the state estimation module and the cumulative judgment process in the interactive decision-making module based on historical optimization data sets.
[0045] Specifically, the sensing data acquisition module consists of a hardware sensor array and its driving control software. The hardware array includes at least two network cameras with wide dynamic range capabilities and one millimeter-wave radar sensor. The cameras and radar are connected to the same clock source via a hardware synchronization interface to ensure data timestamp synchronization. The driving control software is responsible for controlling the sensors to acquire data at a preset frequency and performing initial spatial coordinate calibration and alignment calculations. The output of this module is a time-synchronized and spatially aligned image sequence, as well as radar point cloud and motion trajectory information. The state estimation module is a software processing unit running on an edge computing device or server. Its input consists of image sequences, radar point clouds, and trajectory information output by the perception data acquisition module. First, the module clusters the radar point cloud to separate different targets and then detects human targets in the image sequence. Subsequently, the module calculates the projection of the radar target's 3D coordinates onto the image and the distance to the detected target, associating and matching the radar target with the visual target, and assigning a unique identifier to each target for continuous tracking. Based on the distance between the target and the core point of the interaction area and the target's orientation, the module determines the primary interactive target. For this primary target, the module uses a pre-trained convolutional neural network model to analyze its image region, outputting a first attribute estimation result and its corresponding first confidence score; simultaneously, it estimates the target's height based on the radar point cloud and analyzes its motion trajectory features, outputting a second attribute estimation result and its corresponding second confidence score. Finally, the module performs a weighted fusion of the two attribute estimation results and confidence scores to generate and output user state information containing the final attribute label and its comprehensive confidence score. The interaction decision module is a software logic unit. It receives user state information output by the state estimation module and caches the comprehensive confidence value corresponding to the attribute label in a fixed-length circular buffer. The module performs a moving average filtering calculation on the historical sequence of comprehensive confidence values within a preset time window to obtain the cumulative confidence evaluation value for that attribute label. The module compares this evaluation value with a preset confidence threshold and, based on pre-stored mapping rules defining interaction modes corresponding to different attribute labels, determines the target interaction mode to be executed. After determining the target interaction mode, the module starts a timer to lock the current mode for a preset minimum duration, ignoring new switching requests. The content execution module consists of a multimedia resource library, a real-time rendering engine, a script parser, and an audio processing unit. The multimedia resource library stores 3D digital life model files, dynamic behavior script files, and narration audio files bound to each interaction mode. This module receives the target interaction mode instructions output by the interaction decision module and calls the corresponding resource files from the resource library. The real-time rendering engine loads the 3D model files and renders them on the display device. The script parser reads and executes the dynamic behavior script files, driving the 3D model to perform actions and expressions according to a predetermined timeline. The audio processing unit synchronously plays the narration audio files and, through a pre-trained lip-syncing driven model, generates lip bone control parameters in real time based on the audio, driving the model's lip movements to synchronize with the speech. The feedback acquisition module reuses the hardware sensors of the perception data acquisition module. This module continuously collects data during and after the content execution module's operation. Using face detection and facial orientation estimation models, the module analyzes whether viewers in the image sequence are facing the main display area and calculates the proportion of viewers facing the screen as the first engagement indicator. The module analyzes viewer dwell time and calculates the average dwell time as an interest indicator. The module also analyzes viewer trajectories tracked by radar and counts the number of trajectories that approach the interactive core area or perform specific interactive actions as a second engagement indicator. Finally, the module normalizes and weights these indicators to generate a comprehensive quantitative indicator data output as viewer group feedback data. The data association module consists of a database management system and data association logic units. This module generates an association record for each complete interactive session executed from perception to content. Each record stores the attribute tags from the user's state information that triggered the session, the target interaction mode determined by the interaction decision module, and the quantitative indicator data generated by the feedback collection module. The module persistently stores multiple such records in chronological order, forming a historical optimization data set. The parameter optimization module is a background analysis and optimization software that runs cyclically. It accesses a historical optimization dataset maintained by the data association module. Based on a preset feedback threshold, the module filters out inefficient interaction records from the dataset whose quantitative indicators are below that threshold. The module analyzes the correspondence between attribute labels and target interaction patterns in these records, and adjusts the mapping rules or confidence thresholds stored in the interaction decision module accordingly. Simultaneously, the module extracts the original image and radar data corresponding to the inefficient records and uses the corrected labels to incrementally train or fine-tune the attribute classification model in the state estimation module. After optimization, the module deploys the updated parameters and rules to the corresponding state estimation and interaction decision modules.
[0046] Finally, it should be noted that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A human-computer interaction method based on smart exhibition halls, characterized in that: Includes the following steps: S1. Within the preset interactive area, simultaneously collect visual perception data and radar perception data of at least one target audience member; S2. Based on the visual perception data and radar perception data, perform data fusion and target tracking to generate and output user status information containing target audience attribute tags and their confidence levels; S3. Receive the user status information and determine the current target interaction mode from multiple pre-stored interaction modes based on the cumulative judgment result of the confidence of the attribute tags within a continuous time window. S4. Invoke the multimedia resources corresponding to the target interaction mode, and drive the digital life image to perform interactive behaviors that match the target interaction mode; S5. During and after the digital life image performs interactive behavior, collect audience feedback data within the interactive area; S6. The user status information, the target interaction mode, and the audience feedback data are associated and stored to form a historical optimization data set; S7. Based on the historical optimized data set, perform parameter optimization and adjustment on the data fusion and target tracking process in step S2 and the cumulative judgment process in step S3.
2. The human-computer interaction method based on a smart exhibition hall according to claim 1, characterized in that, In step S1, the step of simultaneously acquiring visual perception data and radar perception data of at least one target audience member within a preset interactive area specifically includes: By deploying at least one wide dynamic range camera at different angles in the interactive area, an image sequence containing the target audience is acquired as the visual perception data; At least one millimeter-wave radar deployed in the interactive area is used to collect the outline point cloud and motion trajectory information of the target audience as radar perception data. The image sequence is time-stamped and spatially aligned with the contour point cloud and motion trajectory information to complete the synchronous acquisition.
3. The human-computer interaction method based on a smart exhibition hall according to claim 2, characterized in that, In step S2, the step of performing data fusion and target tracking based on the visual perception data and radar perception data, and generating and outputting user status information containing target audience attribute tags and their confidence levels, specifically includes: The image sequence is associated and matched with the contour point cloud and motion trajectory information to separate, identify and continuously track multiple audience targets within the interaction area, and to determine the main interaction target as the target audience. Based on the analysis of the target audience in the image sequence, the first attribute estimation result and the corresponding first confidence level are obtained; Based on the analysis of the target audience in the contour point cloud and motion trajectory information, the second attribute estimation result and the corresponding second confidence level are obtained; The first attribute estimation result, the second attribute estimation result, the first confidence score, and the second confidence score are weighted and fused to generate user status information containing the target audience attribute labels and their comprehensive confidence scores.
4. The human-computer interaction method based on a smart exhibition hall according to claim 3, characterized in that, In step S3, the step of receiving the user state information and determining the current target interaction mode from multiple pre-stored interaction modes based on the cumulative judgment result of the confidence level of the attribute labels within a continuous time window specifically includes: Receive user status information containing the target audience's attribute tags and their overall confidence levels; Based on a preset time window length, a moving average filter is applied to the current overall confidence level and its historical values of the attribute label to obtain the cumulative confidence level evaluation value of the attribute label. The cumulative confidence score is compared with a preset confidence threshold, and the current target interaction mode is determined from the pre-stored multiple interaction modes according to the comparison result and the preset mapping rule. After determining the target interaction mode, the target interaction mode is maintained for at least a preset minimum duration, during which new user state information is ignored from triggering the interaction mode switching.
5. The human-computer interaction method based on a smart exhibition hall according to claim 1, characterized in that, In step S4, the step of calling multimedia resources corresponding to the target interaction mode and driving the digital life image to perform interactive behaviors matching the target interaction mode specifically includes: Based on the target interaction mode, the corresponding digital life 3D model, dynamic behavior script and narration audio are called from the preloaded multimedia resource library; The digital life 3D model is rendered on the display device of the interactive area, and the model is driven to perform corresponding actions and expressions based on the dynamic behavior script; The audio of the narration is played synchronously, and the lip movements of the digital life 3D model are matched with the audio of the narration to form the interactive behavior.
6. The human-computer interaction method based on a smart exhibition hall according to claim 1, characterized in that, In step S5, the step of collecting audience feedback data within the interactive area during and after the digital life avatar performs interactive behavior specifically includes: The system uses a wide dynamic range camera to capture images of the audience's facial orientation, dwell time, and crowd density distribution within the interactive area. Millimeter-wave radar is used to collect information on the movement trajectory changes of the audience and the number of people remaining in the area within the interactive area. Based on the facial orientation, dwell time, gathering heat distribution images, movement trajectory changes, and the number of people remaining in the area, quantitative index data reflecting audience participation and interest are extracted and generated as feedback data for the audience group.
7. The human-computer interaction method based on a smart exhibition hall according to claim 6, characterized in that, The step of extracting and generating quantitative indicator data reflecting audience participation and interest as feedback data for the audience group specifically includes: Based on the facial orientation and the clustering heat distribution image, the proportion of viewers facing the main display area of the digital life image is calculated as the first engagement indicator; Based on the dwell time and the number of people staying in the area, the average dwell time of the audience during the interaction behavior execution cycle is calculated as an interest index. Based on the changes in the movement trajectory, the number of trajectories that approach the core location of the interaction area or perform repetitive interactive actions are identified and counted as a second engagement indicator. The first engagement index, the interest index, and the second engagement index are normalized and weighted to generate the quantitative index data.
8. The human-computer interaction method based on a smart exhibition hall according to claim 7, characterized in that, In step S6, the step of associating and storing the user status information, the target interaction pattern, and the audience feedback data to form a historical optimization data set specifically includes: For each complete human-computer interaction session executed from S1 to S4, generate an associated record; The associated record stores the attribute tags in the user status information that triggered the current session, the determined target interaction mode, and the quantitative indicator data corresponding to the current session. The multiple related records are aggregated in chronological order to form the historical optimization data set.
9. The human-computer interaction method based on a smart exhibition hall according to claim 8, characterized in that, In step S7, the step of optimizing and adjusting the parameters of the data fusion and target tracking process in step S2 and the cumulative judgment process in step S3 based on the historical optimization data set specifically includes: From the historical optimization data set, inefficient interaction records where the quantitative indicator data is lower than the preset feedback threshold are selected; Analyze the correspondence between the attribute tags stored in the inefficient interaction records and the target interaction mode, and adjust the mapping rules or confidence thresholds from attribute tags to interaction modes in S3. Extract the original perceptual data corresponding to the inefficient interaction records, and retrain or correct the parameters of the analysis model used to generate the first attribute estimation result and / or the second attribute estimation result in S2.
10. A human-computer interaction system based on a smart exhibition hall, characterized in that: The human-computer interaction method based on a smart exhibition hall according to any one of claims 1-9 includes: The perception data acquisition module is used to simultaneously acquire visual perception data and radar perception data of at least one target audience member within a preset interactive area. The state estimation module is used to perform data fusion and target tracking based on the visual perception data and radar perception data, and generate and output user state information containing target audience attribute tags and their confidence levels. The interaction decision module is used to receive the user status information and determine the current target interaction mode from multiple pre-stored interaction modes based on the cumulative judgment result of the confidence of the attribute tags within a continuous time window. The content execution module is used to call multimedia resources corresponding to the target interaction mode and drive the digital life image to perform interactive behaviors that match the target interaction mode. The feedback collection module is used to collect audience feedback data within the interactive area during and after the digital life image performs interactive behavior. The data association module is used to associate and store the user status information, the target interaction mode, and the audience feedback data to form a historical optimization data set. The parameter optimization module is used to optimize and adjust the parameters of the data fusion and target tracking process in the state estimation module and the cumulative judgment process in the interactive decision module based on the historical optimization data set.