Interaction Method and System for Intelligent Display Devices Based on Multi-Sensor Fusion
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-19
- Publication Date
- 2026-08-14
AI Technical Summary
[0003]然而在热门展馆的高人流环境中,游客聚集在互动屏前进行隔空翻页操作时,现有技术容易因仅依赖摄像头识别而产生误触发
首先,通过获取手势图像序列、目标距离信息及方向信息,并对所述多源感知数据进行时间对齐处理,使不同来源的感知数据能够在同一时间基准下进行关联,从而保证后续空间位置计算、区域映射及轨迹分析过程中的时间一致性,避免由于不同传感器采集时间差异导致的目标错配、位置偏差或轨迹断裂问题。
Smart Images

Figure CN122239950B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to an interactive method and system for intelligent display devices based on multi-sensor fusion. Background Technology
[0002] Existing intelligent display interaction technologies in exhibition halls and cultural tourism settings mostly employ touchscreens combined with infrared sensors or monocular cameras to achieve human-computer interaction. The system switches displayed content, such as artifact introductions, guide maps, or multimedia playback, by detecting the location of visitors' touches or recognizing simple gestures (such as waving or clicking an air button). While some solutions incorporate voice or visual recognition, there is a lack of deep integration between multiple sensors; they typically operate as independent modules, with simple triggering control from higher-level logic.
[0003] However, in high-traffic environments like popular exhibition halls, when visitors gather in front of interactive screens to perform remote page-turning operations, existing technologies are prone to false triggers due to their reliance solely on camera recognition. When multiple visitors move or converse in front of the screen simultaneously, gestures from non-users in the background may also be recognized as valid input by the system, leading to frequent erroneous interface switching, such as interrupting and redirecting a previously playing cultural and tourism information video. This problem is quite common in actual exhibition halls, stemming from the system's failure to integrate multi-source data such as distance, direction, or individual recognition, making it difficult to accurately identify the actual interactive subject. Summary of the Invention
[0004] The purpose of this invention is to provide an interactive method and system for intelligent display devices based on multi-sensor fusion, aiming to solve the problems mentioned in the background art.
[0005] To solve the above-mentioned technical problems, the technical solution of the present invention is as follows: Firstly, an interaction method for intelligent display devices based on multi-sensor fusion, the method comprising: The system acquires multi-source sensing data collected by a data acquisition device, which includes multiple sensors. These sensors are used to acquire gesture image sequences, target distance information, and direction information. The system also performs time alignment processing on the multi-source sensing data to generate synchronous sensing data. Based on synchronous sensing data, target identification and distance filtering are performed to generate a set of candidate interactive objects within a preset interaction distance range; Based on the target distance and direction information in the synchronous sensing data, the spatial position of each candidate interactive object is calculated in the preset spatial coordinate system, spatial position data is generated, and a three-dimensional interactive volume region corresponding to the display interface is constructed based on the spatial posture information of the display interface, the normal direction of the display interface, and the preset interactive distance range. Based on spatial location data, the set of candidate interactive objects is mapped to a three-dimensional interactive volume region, the spatial occupancy trajectory of each candidate interactive object is calculated, and the corresponding trajectory data is generated. Based on the trajectory data, trajectory continuity analysis and stability assessment are performed to determine the target interaction object and establish the interaction locking state corresponding to the target interaction object and the three-dimensional interaction volume region. In the interactive lock state, the trajectory data of the target interactive object is continuously tracked and continuity detection is performed. When the continuity of trajectory data is interrupted or the current spatial position of the target interactive object deviates from the three-dimensional interactive volume area by more than a preset distance threshold, the interactive lock state is released. After the interaction lock is released, the target interaction object is re-determined based on the time sequence of each candidate interaction object re-entering the three-dimensional interaction volume area, and the consistency between the movement direction of each candidate interaction object and the interaction direction of the display interface. The trajectory data corresponding to the target interactive object is parsed and processed to generate corresponding interactive commands, and the intelligent display device is controlled to perform corresponding operations according to the interactive commands.
[0006] Secondly, an intelligent display device interaction system based on multi-sensor fusion, the system comprising: The data acquisition module is used to acquire multi-source sensing data collected by the acquisition device, which includes multiple sensors. The multiple sensors are used to collect gesture image sequences, target distance information and direction information, and perform time alignment processing on the multi-source sensing data to generate synchronous sensing data. The target recognition and filtering module is used to perform target recognition and distance filtering based on synchronous perception data, and generate a set of candidate interactive objects within a preset interaction distance range; The region construction module is used to calculate the spatial position of each candidate interactive object in a preset spatial coordinate system based on the target distance and direction information in the synchronous sensing data, generate spatial position data, and construct a three-dimensional interactive volume region corresponding to the display interface based on the spatial posture information of the display interface, the normal direction of the display interface, and the preset interactive distance range. The trajectory generation module is used to map the set of candidate interactive objects to a three-dimensional interactive volume area based on spatial location data, calculate the spatial occupancy trajectory of each candidate interactive object, and generate the corresponding trajectory data. The object determination module is used to perform trajectory continuity analysis and stability assessment based on trajectory data, determine the target interactive object, and establish the interactive locking state corresponding to the target interactive object and the three-dimensional interactive volume region. The target tracking release module is used to continuously track the trajectory data of the target interactive object in the interactive lock state and perform continuity detection. When the continuity of the trajectory data is interrupted or the current spatial position of the target interactive object deviates from the three-dimensional interactive volume area by more than a preset distance threshold, the interactive lock state is released. The target re-determination module is used to redetermine the target interactive object after the interaction lock state is released, based on the time order in which each candidate interactive object re-enters the three-dimensional interactive volume area, and the consistency between the movement direction of each candidate interactive object and the interaction direction of the display interface. The interactive control module is used to perform trajectory parsing and processing based on the trajectory data corresponding to the target interactive object, generate corresponding interactive commands, and control the intelligent display device to perform corresponding operations based on the interactive commands.
[0007] The above-described solution of the present invention has at least the following beneficial effects: First, by acquiring gesture image sequences, target distance information, and direction information, and performing time alignment processing on the multi-source sensing data, sensing data from different sources can be correlated under the same time reference, thereby ensuring time consistency in subsequent spatial position calculation, region mapping, and trajectory analysis processes, and avoiding target mismatch, position deviation, or trajectory breakage caused by differences in the acquisition time of different sensors.
[0008] Furthermore, by identifying targets based on synchronous sensing data and filtering the identified targets by distance information, the objects involved in the interaction determination are limited to a preset interaction distance range. This reduces the interference of distant non-interactive personnel, background targets, or irrelevant objects on subsequent spatial positioning and trajectory analysis, and improves the targeting and accuracy of candidate interaction object selection.
[0009] Furthermore, by calculating the spatial position of each candidate interactive object in a preset spatial coordinate system based on the target distance and direction information in the synchronous sensing data, the candidate interactive objects can be expressed in a unified spatial coordinate form, thereby providing an accurate positional basis for subsequent candidate interactive object mapping, trajectory generation, and interaction judgment.
[0010] Furthermore, by constructing a three-dimensional interactive volume region corresponding to the display interface based on the spatial posture information of the display interface, the normal direction of the display interface, and the preset interactive distance range, the interactive judgment region can be matched with the actual installation posture, spatial orientation, and effective interactive distance of the display interface. This extends the interactive judgment from the two-dimensional image level to the three-dimensional spatial constraints corresponding to the display interface, avoiding the misidentification of background personnel actions or actions in non-interactive spaces as interactive operations due to relying solely on image recognition.
[0011] Furthermore, by mapping the set of candidate interactive objects to a three-dimensional interactive volume area based on spatial location data, and generating trajectory data corresponding to each candidate interactive object, the motion process of the candidate interactive object within the corresponding spatial range of the display interface can be continuously recorded. This transforms the interaction judgment from single-frame recognition to a recognition method based on continuous spatial motion trajectory, which is beneficial for distinguishing between real interactive actions, occasional actions, jittery actions, or invalid actions.
[0012] Furthermore, by performing trajectory continuity analysis and stability assessment on trajectory data, the target interaction object can be determined, enabling the selection of the interaction subject to be judged based on the continuity of the trajectory, the stability of spatial changes, and the motion state, rather than relying solely on the target position or detection result at a single moment. This reduces the probability of misselection caused by target intersection, short-term occlusion, or detection fluctuations in multi-person environments.
[0013] Furthermore, by establishing an interaction lock state corresponding to the target interactive object and the three-dimensional interactive volume area, and continuously tracking the trajectory data of the target interactive object in the interaction lock state, the interaction processing can continue to be carried out around the same target interactive object within a certain time range, thereby avoiding interface misoperation caused by frequent switching of candidate objects when multiple people are present at the same time.
[0014] Meanwhile, when the continuity of the trajectory data of the target interactive object is interrupted, or when the current spatial position of the target interactive object deviates from the three-dimensional interactive volume area by more than a preset distance threshold, the interaction lock state is released, so that the system can release the lock in a timely manner according to the actual motion state and spatial position change of the target interactive object, thereby avoiding the target from being continuously identified as the interactive subject after leaving the effective interactive space and causing erroneous control.
[0015] Furthermore, after the interaction lock is released, the target interaction object is re-determined based on the time sequence of each candidate interaction object re-entering the three-dimensional interaction volume area, as well as the consistency between the movement direction of each candidate interaction object and the interaction direction of the display interface. This ensures that the selection of the new interaction subject has both temporal and directional basis, thereby improving the rationality of the interaction subject switching and the continuity of the interaction process.
[0016] Finally, by performing trajectory analysis processing based on the trajectory data corresponding to the target interactive object, corresponding interactive instructions are generated, and the intelligent display device is controlled to perform corresponding operations according to the interactive instructions. This establishes a correspondence between the spatial motion trajectory and the display interface operation, thereby achieving stable, continuous, and interference-resistant intelligent display interactive control in a multi-person environment and reducing the problem of false triggering caused by the actions of non-target personnel. Attached Figure Description
[0017] Figure 1This is a flowchart of an intelligent display device interaction method based on multi-sensor fusion provided in an embodiment of the present invention. Detailed Implementation
[0018] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.
[0019] like Figure 1 As shown, embodiments of the present invention propose an interactive method for intelligent display devices based on multi-sensor fusion, the method comprising: The system acquires multi-source sensing data collected by a data acquisition device, which includes multiple sensors. These sensors are used to acquire gesture image sequences, target distance information, and direction information. The system also performs time alignment processing on the multi-source sensing data to generate synchronous sensing data. Based on synchronous sensing data, target identification and distance filtering are performed to generate a set of candidate interactive objects within a preset interaction distance range; Based on the target distance and direction information in the synchronous sensing data, the spatial position of each candidate interactive object is calculated in the preset spatial coordinate system, spatial position data is generated, and a three-dimensional interactive volume region corresponding to the display interface is constructed based on the spatial posture information of the display interface, the normal direction of the display interface, and the preset interactive distance range. Based on spatial location data, the set of candidate interactive objects is mapped to a three-dimensional interactive volume region, the spatial occupancy trajectory of each candidate interactive object is calculated, and the corresponding trajectory data is generated. Based on the trajectory data, trajectory continuity analysis and stability assessment are performed to determine the target interaction object and establish the interaction locking state corresponding to the target interaction object and the three-dimensional interaction volume region. In the interactive lock state, the trajectory data of the target interactive object is continuously tracked and continuity detection is performed. When the continuity of trajectory data is interrupted or the current spatial position of the target interactive object deviates from the three-dimensional interactive volume area by more than a preset distance threshold, the interactive lock state is released. After the interaction lock is released, the target interaction object is re-determined based on the time sequence of each candidate interaction object re-entering the three-dimensional interaction volume area, and the consistency between the movement direction of each candidate interaction object and the interaction direction of the display interface. The trajectory data corresponding to the target interactive object is parsed and processed to generate corresponding interactive commands, and the intelligent display device is controlled to perform corresponding operations according to the interactive commands.
[0020] In this embodiment of the invention, by acquiring gesture image sequences, target distance information, and direction information, and performing time alignment processing on multi-source sensing data, sensing data from different sources can be correlated under the same time reference, thereby ensuring the time consistency of subsequent spatial position calculation and trajectory change analysis, and avoiding position deviation, target mismatch, or trajectory breakage problems caused by differences in acquisition time.
[0021] After data synchronization is completed, target identification is performed on the synchronized sensing data, and distance filtering is performed in combination with target distance information. This limits the objects involved in the interaction judgment to a preset interaction distance range, thereby reducing the interference of irrelevant objects, background targets, or distant targets on subsequent spatial positioning, area mapping, and trajectory analysis, and allowing subsequent processing to focus on effective interaction objects.
[0022] Based on the target distance and direction information in the synchronous sensing data, the spatial position of each candidate interactive object is calculated in the preset spatial coordinate system, so that the candidate interactive objects can be expressed in a unified spatial coordinate form, thereby providing a stable position basis for subsequent object mapping, trajectory generation and interaction judgment.
[0023] Meanwhile, based on the spatial posture information of the display interface, the normal direction of the display interface, and the preset interaction distance range, a three-dimensional interaction volume region corresponding to the display interface is constructed. This allows the three-dimensional interaction volume region to be determined by the actual spatial orientation and effective interaction distance of the display interface, rather than solely by the instantaneous position distribution of candidate interaction objects. This improves the stability of the interaction region and the accuracy of the spatial correspondence, avoiding the problem of frequent changes in the interaction region or inconsistent judgment benchmarks caused by changes in the position of candidate objects.
[0024] By mapping a set of candidate interactive objects to a three-dimensional interactive volume region and generating trajectory data based on the spatial position changes of the candidate interactive objects within this region, the interactive behavior is transformed from discrete spatial positions into continuous motion trajectories. This allows the interaction process of the candidate interactive objects to be reflected within the corresponding spatial range of the display interface, providing a foundation for subsequent trajectory continuity analysis and stability assessment.
[0025] After the trajectory data is generated, by performing trajectory continuity analysis and stability assessment on the trajectory data, candidate interaction objects that do not meet the continuity condition, have obvious jumps, or have unstable motion states can be screened out, and the target interaction object can be determined from the candidate interaction objects that meet the stability condition. This makes the selection of the interaction subject based on continuous trajectory characteristics and motion stability, rather than based on a single position or instantaneous detection result, thereby reducing the probability of misselection and false triggering.
[0026] Once the target interaction object is determined, an interaction lock state is established between the target interaction object and the corresponding 3D interaction volume region. This allows subsequent interaction processing to continue around the same target interaction object, thereby reducing the problem of frequent switching of interaction subjects caused by target intersection, short-term occlusion, or detection fluctuations in multi-person or multi-target scenarios.
[0027] In the interactive lock state, the trajectory data of the target interactive object is continuously tracked and continuity is detected. When the continuity of the trajectory data is interrupted, or the current spatial position of the target interactive object deviates from the three-dimensional interactive volume area by more than a preset distance threshold, the interactive lock state is released, so that the system can release the lock in time according to the actual state change of the target interactive object, thereby avoiding the target from being continuously identified as the interactive subject after leaving the interaction range and causing erroneous control.
[0028] After the interaction lock is released, the target interaction object is re-determined based on the time sequence of each candidate interaction object re-entering the three-dimensional interaction volume area, as well as the consistency between the movement direction of each candidate interaction object and the interaction direction of the display interface. This ensures that the selection of the new interaction subject has both temporal and directional basis, thereby improving the rationality and continuity of the interaction subject switching.
[0029] Finally, the trajectory data corresponding to the target interactive object is parsed and processed to generate corresponding interactive commands. The intelligent display device is then controlled to perform corresponding operations based on the interactive commands, so that the spatial motion trajectory of the target interactive object can establish a correspondence with the display interface operation of the intelligent display device, thereby realizing intelligent interactive control based on multi-sensor spatial perception and trajectory analysis.
[0030] In a preferred embodiment of the present invention, acquiring multi-source sensing data collected by the acquisition device includes: The image sensor in the control acquisition device acquires a sequence of gesture images within its sensing range; The distance sensor in the control acquisition device collects target distance information between the measured object and the acquisition device within the sensing range; Based on the target position collected by the image sensor, the orientation information collected by the angle sensor, or the orientation data collected by the direction detection sensor, the orientation information of the measured object relative to the acquisition device is obtained. Add sampling time markers and sensor source markers to the gesture image sequence, target distance information, and direction information, respectively; Based on the sampling time marker and sensor source marker, the gesture image sequence, target distance information and direction information are encapsulated to form multi-source perception data; The sampling frequency and sampling range of the image sensor, distance sensor, and orientation detection sensor are set according to the display size of the smart display device, the installation position of the acquisition device, the effective detection range of the sensor, and the user operation distance in the actual interactive scenario.
[0031] In a preferred embodiment of the present invention, time alignment processing is performed on multi-source sensing data to generate synchronized sensing data, including: Obtain the sampling time markers corresponding to the data from each sensor in the multi-source sensing data; The sampling time markers are uniformly processed using a preset synchronization time reference to obtain the sampling time of each sensor data under the same time reference. According to the preset synchronization time interval, the gesture image sequence, target distance information and direction information are divided into time windows; Within the same time window, gesture image frames with the same sampling time or a sampling time difference less than a preset time alignment threshold, target distance information, and direction information are associated and matched. When there is data with different sampling times within the same time window, interpolation, resampling, or nearest neighbor matching is performed based on data from adjacent sampling times. The gesture image frame, target distance information, and direction information after the association matching is completed are combined into synchronous sensing data at the same sampling time. The preset synchronization time reference is determined based on the system clock of the acquisition device, the control clock of the intelligent display device, or the synchronization trigger signal between multiple sensors; the preset time alignment threshold and the preset synchronization time interval are set according to the sampling frequency of each sensor, the data transmission delay, and the system's allowed interactive response delay.
[0032] In a preferred embodiment of the present invention, target identification and distance filtering are performed based on synchronous sensing data to generate a set of candidate interactive objects located within a preset interaction distance range, including: Target detection processing is performed based on the gesture image sequence in the synchronous sensing data to identify human targets, hand targets, or gesture targets in the image; Based on the target detection results, extract the image location, contour features, pose features, or gesture features corresponding to each target; Based on image location, contour features, pose features, or gesture features, each target is classified to determine candidate targets with interactive features. Based on the target distance information in the synchronous sensing data, obtain the distance value corresponding to each candidate target; Compare the distance value corresponding to each candidate target with the preset interaction distance range to determine whether each candidate target is within the preset interaction distance range; Candidate targets whose distance values are within a preset interaction distance range are identified as candidate interaction objects, and a set of candidate interaction objects is generated based on the candidate interaction objects; The preset interaction distance range is set statistically based on the effective sensing distance of the acquisition device, the display size of the intelligent display device, the distribution of normal user operation distance, and the standing range in the actual interaction scenario; the interaction features are obtained by training or statistics based on the morphological features of the human body or hand target, the gesture action features, and actual interaction operation samples.
[0033] In a preferred embodiment of the present invention, in the interactive lock state, the trajectory data of the target interactive object is continuously tracked and continuity detection is performed. When the continuity of the trajectory data is detected to be interrupted or the current spatial position of the target interactive object deviates from the three-dimensional interactive volume area by more than a preset distance threshold, the interactive lock state is released, including: In the interactive lock state, acquire the spatial position data of the target interactive object at continuous sampling times; Based on spatial location data, the trajectory data of the target interactive object is updated to generate an updated trajectory sequence. Based on the updated trajectory sequence, calculate the time interval and spatial displacement change between adjacent trajectory points; The time interval is compared with a preset time threshold, and the change in spatial displacement is compared with a preset displacement change threshold. When the time interval between adjacent trajectory points is greater than a preset time threshold, or the change in spatial displacement between adjacent trajectory points is greater than a preset displacement change threshold, it is determined that the trajectory data of the target interactive object has been interrupted. Simultaneously, based on the spatial relationship between the current spatial position of the target interactive object and the three-dimensional interactive volume region, the deviation distance of the current spatial position relative to the boundary of the three-dimensional interactive volume region is calculated; The deviation distance is compared with a preset distance threshold. When the deviation distance is greater than the preset distance threshold, it is determined that the target interactive object deviates from the three-dimensional interactive volume area. When an interruption in trajectory data is detected, or when the target interactive object deviates from the three-dimensional interactive volume area, the interactive lock mark corresponding to the target interactive object is cleared, and the interactive lock state between the target interactive object and the three-dimensional interactive volume area is released. The preset time threshold is set based on the system sampling frequency, sensor data transmission delay, and allowable short-term trajectory loss time; the preset displacement change threshold is set based on the displacement change distribution of normal user interaction actions, the spatial measurement error of the acquisition device, and the system sampling accuracy; and the preset distance threshold is set based on the boundary range of the three-dimensional interactive volume area, the positioning error of the acquisition device, and the user movement range in the actual interactive scenario.
[0034] In a preferred embodiment of the present invention, after the interaction lock state is released, the target interaction object is re-determined based on the time sequence of each candidate interaction object re-entering the three-dimensional interaction volume region, and the consistency between the movement direction of each candidate interaction object and the interaction direction of the display interface, including: After the interaction lock state is released, the spatial location data of each candidate interaction object in the candidate interaction object set is continuously acquired; Based on the spatial relationship between the spatial location data of each candidate interactive object and the three-dimensional interactive volume region, it is determined whether each candidate interactive object enters the three-dimensional interactive volume region. When a candidate interactive object enters the 3D interactive volume area from outside the 3D interactive volume area, the entry time of the candidate interactive object is recorded. Based on the entry time of each candidate interactive object, the candidate interactive objects that re-enter the 3D interactive volume region are sorted by time. Calculate the motion direction vector of each candidate interactive object based on the spatial position changes of each candidate interactive object before and after entering the three-dimensional interactive volume area; Obtain the interaction direction of the display interface, and calculate the angle between the motion direction vector of each candidate interactive object and the interaction direction of the display interface; The included angle is compared with a preset direction consistency threshold to filter candidate interaction objects that meet the direction consistency condition; From the candidate interactive objects that meet the directional consistency condition, select the candidate interactive object that enters the three-dimensional interactive volume region earliest and re-determine it as the target interactive object; The interaction direction of the display interface is determined based on the normal direction of the display interface, the user's operation direction facing the display interface, or the preset interaction direction; the preset direction consistency threshold is set according to the installation posture of the display interface, the detection angle range of the acquisition device, the direction of common user interaction actions, and the distribution statistics of motion directions in the actual interaction scenario.
[0035] In a preferred embodiment of the present invention, trajectory parsing processing is performed based on the trajectory data corresponding to the target interactive object to generate corresponding interactive instructions, and the intelligent display device is controlled to perform corresponding operations according to the interactive instructions, including: Acquire the trajectory data of the target interactive object when the interaction is locked; Based on the trajectory data, extract the spatial location point sequence of the target interactive object at continuous sampling time; Based on the spatial location point sequence, the trajectory of the target interactive object is divided into trajectory segments to obtain one or more continuous trajectory segments; Based on continuous trajectory segments, calculate the trajectory displacement direction, trajectory displacement amplitude, trajectory duration, trajectory velocity change, and trajectory start and end positions; Based on the trajectory displacement direction, trajectory displacement amplitude, trajectory duration, trajectory velocity change, and trajectory start and end positions, trajectory feature parameters are generated. The trajectory feature parameters are matched with preset interactive action templates to determine the type of interactive action corresponding to the target interactive object; Generate corresponding interaction instructions based on the type of interaction action; According to the interactive instructions, control the intelligent display device to perform corresponding operations, including at least one of page switching, content scrolling, object selection, confirmation operation, cancellation operation, zoom operation, or display content movement operation; The preset interactive action templates are obtained by statistical analysis of trajectory samples of common user operation behaviors in actual interactive scenarios, or by training based on pre-collected gesture trajectory samples. The preset interactive action templates include the trajectory direction range, displacement amplitude range, duration range, speed change range, and start and end position range corresponding to different interactive actions.
[0036] In a preferred embodiment of the present invention, based on the target distance information and direction information in the synchronous sensing data, the spatial position of each candidate interactive object is calculated in a preset spatial coordinate system to generate spatial position data. Then, based on the spatial pose information of the display interface, the normal direction of the display interface, and a preset interactive distance range, a three-dimensional interactive volume region corresponding to the display interface is constructed, including: Based on the fixed installation relationship between the data acquisition device and the intelligent display device, a spatial coordinate system corresponding to the display interface is established with the data acquisition device as the origin. Based on the direction information, determine the direction vector of each candidate interactive object relative to the acquisition device, and based on the target distance information, determine the distance value of each candidate interactive object on the corresponding direction vector; Based on the direction vector and distance value, the spatial position of each candidate interactive object is calculated in the spatial coordinate system to obtain spatial position data under a unified coordinate system; Obtain the spatial orientation information of the display interface in the spatial coordinate system, and determine the normal direction of the display interface based on the spatial orientation information; Based on the normal direction of the display interface and the preset direction offset angle threshold, determine the interaction direction area corresponding to the display interface; Based on the preset interaction distance range, the depth of the interaction direction area is defined along the normal direction of the display interface to construct a three-dimensional interaction volume area corresponding to the display interface.
[0037] In this embodiment of the invention, by establishing a spatial coordinate system corresponding to the display interface with the acquisition device as the origin based on the fixed installation relationship between the acquisition device and the intelligent display device, the spatial position calculation of the candidate interactive objects has a unified reference benchmark, thereby ensuring that different candidate interactive objects can be expressed in the same coordinate system and avoiding position deviations caused by inconsistent coordinate benchmarks.
[0038] Based on this, the direction vector of each candidate interactive object relative to the acquisition device is determined according to the direction information, and the distance value of each candidate interactive object on the corresponding direction vector is determined by combining the target distance information. This allows the direction information and distance information to participate in spatial positioning, thereby enabling the spatial position of each candidate interactive object to be calculated in the spatial coordinate system, providing an accurate positional basis for subsequent area mapping and trajectory generation.
[0039] Meanwhile, by acquiring the spatial orientation information of the display interface in the spatial coordinate system and determining the normal direction of the display interface based on the spatial orientation information, the construction of the three-dimensional interactive volume area can be based on the actual installation posture and spatial orientation of the display interface, thereby ensuring a stable spatial correspondence between the interactive area and the display interface.
[0040] Furthermore, the interaction direction area corresponding to the display interface is determined based on the normal direction of the display interface and the preset direction offset angle threshold, so that the interaction direction area can be limited around the interaction orientation of the display interface, thereby eliminating the spatial range that deviates from the interaction direction of the display interface and reducing the interference of objects or actions in non-target directions on the interaction judgment.
[0041] Finally, based on the preset interaction distance range, the depth of the interaction direction area is limited along the normal direction of the display interface to construct a three-dimensional interaction volume area corresponding to the display interface. This allows the three-dimensional interaction volume area to be constrained by both the direction range and the depth range, thereby forming a stable spatial range that matches the actual interaction space of the display interface. This provides a reliable spatial judgment basis for subsequent candidate interaction object mapping, trajectory generation, and interaction locking.
[0042] In a preferred embodiment of the present invention, based on the fixed installation relationship between the acquisition device and the intelligent display device, a spatial coordinate system corresponding to the display interface is established with the acquisition device as the origin, including: Acquire the installation location and orientation information between the data acquisition device and the intelligent display device; Based on the installation location information and installation posture information, determine the relative positional relationship and relative orientation relationship between the acquisition device and the display interface; The preset reference point of the acquisition device is used as the origin of the spatial coordinate system; Based on the spatial orientation of the display interface, determine the direction of at least one coordinate axis in the spatial coordinate system, so that the direction of at least one coordinate axis corresponds to the normal direction of the display interface; Establish a spatial coordinate system corresponding to the display interface based on the origin and the directions of each coordinate axis; The preset reference point of the acquisition device is determined based on the optical center position of the main sensor, the ranging center position, or the equipment installation reference point in the acquisition device; the installation position information and installation attitude information are obtained through equipment installation calibration, factory calibration, initialization calibration, or manual measurement.
[0043] In a preferred embodiment of the present invention, determining the direction vector of each candidate interactive object relative to the acquisition device based on direction information, and determining the distance value of each candidate interactive object on the corresponding direction vector based on target distance information, includes: Based on the directional information in the synchronous sensing data, obtain the azimuth angle, pitch angle, image coordinate position or sensor orientation data of each candidate interactive object within the sensing range of the acquisition device; Based on the azimuth, pitch, image coordinates or sensor orientation data of each candidate interactive object, calculate the initial direction vector of each candidate interactive object relative to the acquisition device. Based on the attitude relationship between the acquisition device and the spatial coordinate system, the initial direction vector is transformed into the spatial coordinate system to obtain the direction vector of each candidate interactive object relative to the acquisition device. Based on the target distance information in the synchronous sensing data, the distance measurement value of each candidate interactive object relative to the acquisition device is obtained; Based on the candidate interaction object identifier, sampling time or spatial matching relationship, the distance measurement value of each candidate interaction object is associated with the corresponding direction vector; The distance value associated with the direction vector is determined as the distance value of the candidate interactive object on the corresponding direction vector; Specifically, when the orientation information is determined by the image coordinate position, the initial orientation vector is calculated based on the intrinsic parameters of the image sensor, the image coordinate position, and the attitude parameters of the acquisition device; when the orientation information is determined by an angle sensor or orientation detection sensor, the initial orientation vector is calculated based on the azimuth and pitch angles. The attitude relationship between the acquisition device and the spatial coordinate system is obtained through equipment installation calibration or initialization calibration.
[0044] In a preferred embodiment of the present invention, obtaining the spatial orientation information of the display interface in the spatial coordinate system and determining the normal direction of the display interface based on the spatial orientation information includes: Obtain the installation posture information of the display interface in the intelligent display device; Based on the fixed installation relationship between the acquisition device and the intelligent display device, the installation posture information of the display interface is converted into the spatial coordinate system to obtain the spatial posture information of the display interface in the spatial coordinate system. Based on the spatial attitude information, determine the orientation of the plane where the display interface is located in the spatial coordinate system; Calculate the direction vector perpendicular to the plane where the display interface is located, based on the orientation of the plane. Normalize the direction vector perpendicular to the plane where the display interface is located to obtain the normal direction of the display interface; The installation posture information of the display interface is obtained through the factory calibration parameters of the intelligent display device, the equipment installation calibration parameters, the posture sensor measurement parameters, or the initialization calibration parameters; the orientation of the normal direction is determined according to the front orientation of the display interface, the position of the user interaction side, or the relative positional relationship between the acquisition device and the display interface.
[0045] In a preferred embodiment of the present invention, based on spatial location data, a set of candidate interactive objects is mapped to a three-dimensional interactive volume region, the spatial occupancy trajectory of each candidate interactive object is calculated, and corresponding trajectory data is generated, including: Based on the spatial location data, the spatial locations of each candidate interactive object at consecutive sampling times are extracted and arranged in chronological order to generate location sequence data; Based on the location sequence data, the spatial location at each time point is matched with the three-dimensional interactive volume region to determine whether the candidate interactive object is located within the three-dimensional interactive volume region and generate region status marker data. Based on the regional status marker data, continuous spatial locations within the three-dimensional interactive volume region are filtered and correlated in chronological order to generate initial trajectory data; Abnormal location points caused by sampling errors in the initial trajectory data are identified and their positions are corrected to eliminate discrete jumps in the trajectory and generate corresponding trajectory data.
[0046] In this embodiment of the invention, by arranging the spatial location data into a time series, the spatial change relationships of each candidate interactive object at continuous sampling times are preserved, thereby forming a position sequence that reflects the motion process. Based on this, the position sequence is matched with a three-dimensional interactive volume region, marking the spatial locations within that region and distinguishing between valid and invalid interactive locations. Further, continuous spatial locations within the three-dimensional interactive volume region are correlated, transforming discrete position data into continuous trajectory data, thus describing the spatial motion path of the object. Subsequently, abnormal position points in the trajectory caused by sampling errors are identified and corrected, eliminating abrupt changes in the trajectory and ensuring the continuity and stability of the trajectory data, providing a consistent data foundation for subsequent trajectory analysis.
[0047] In a preferred embodiment of the present invention, based on spatial location data, the spatial locations of each candidate interactive object at consecutive sampling times are extracted and arranged in chronological order to generate location sequence data, including: Based on the spatial location data, obtain the three-dimensional coordinate data of each candidate interactive object at different sampling times to generate the original location data; Based on the original location data, the spatial location of the same candidate interactive object at consecutive sampling times is extracted, and the sampling times are marked to generate time-stamped data; Based on the time stamp data, the spatial locations of each candidate interactive object are sorted in chronological order to generate corresponding location sequence data.
[0048] In a preferred embodiment of the present invention, based on the location sequence data, the spatial position at each time point is matched with the three-dimensional interactive volume region to determine whether the candidate interactive object is located within the three-dimensional interactive volume region, and region status marker data is generated, including: Extract the spatial coordinates of each candidate interactive object at each sampling time based on the location sequence data; Based on the spatial boundary of the three-dimensional interactive volume region, the spatial position coordinates are subjected to inclusion relationship determination to determine whether each spatial position is located within the three-dimensional interactive volume region. Based on the judgment results, the spatial position at each sampling time is marked with a state label. Spatial positions located within the three-dimensional interactive volume area are marked as valid states, while spatial positions not located within the three-dimensional interactive volume area are marked as invalid states, thus generating regional state label data.
[0049] In a preferred embodiment of the present invention, based on the region state marker data, continuous spatial locations within the three-dimensional interactive volume region are filtered and correlated in chronological order to generate initial trajectory data, including: Based on the regional status labeling data, spatial locations marked as valid are filtered to generate valid location data; Based on the valid location data, spatial locations that are valid in all consecutive sampling times are extracted to generate a continuous valid location sequence; Based on the continuous valid position sequence, the spatial positions are connected in chronological order to generate initial trajectory data.
[0050] In a preferred embodiment of the present invention, abnormal position points caused by sampling errors in the initial trajectory data are identified, and position correction processing is performed on the abnormal position points to eliminate discrete jumps in the trajectory and generate corresponding trajectory data, including: Based on the initial trajectory data, calculate the displacement change between spatial positions at adjacent sampling times to generate displacement change data; Based on displacement change data, determine whether the displacement between adjacent spatial locations exceeds a preset displacement change threshold, mark spatial locations that exceed the threshold as abnormal location points, and generate abnormal marker data. Based on the anomaly marker data, the anomaly location points are corrected by interpolation or replacement using nearby valid locations to generate corrected location data. Based on the corrected position data, the trajectory is updated to generate continuous and smooth trajectory data.
[0051] In a preferred embodiment of the present invention, trajectory continuity analysis and stability assessment are performed based on trajectory data to determine the target interaction object, and an interaction locking state corresponding to the target interaction object and the three-dimensional interaction volume region is established, including: Based on the time interval and spatial location change of adjacent moments in the trajectory data, the trajectory of each candidate interaction object is segmented to generate multiple trajectory segments. Based on the trajectory segments, select those that meet the conditions of continuous time and continuous spatial change to generate continuous trajectory segment data; Based on the correspondence between each trajectory segment in the continuous trajectory segment data and the candidate interactive objects, the candidate interactive objects with continuous trajectory segments are determined. Based on candidate interactive objects with continuous trajectory segments, calculate the spatial position offset and motion direction change of their continuous trajectory segments, and perform stability calculation on the corresponding candidate interactive objects based on the spatial position offset and motion direction change to generate trajectory stability parameter data. Based on the trajectory stability parameter data, candidate interactive objects with continuous trajectory segments are screened and extracted. Candidate interactive objects whose stability meets the preset stability conditions are extracted and sorted according to the time order of entering the three-dimensional interactive volume region to determine the target interactive object. Based on the trajectory data formed by the target interactive object within the three-dimensional interactive volume area, an interactive locking state is established between the target interactive object and the three-dimensional interactive volume area, and the interactive locking state remains unchanged unless the unlocking condition is met.
[0052] In this embodiment of the invention, the trajectory of each candidate interactive object is segmented according to the time interval and spatial position change of adjacent moments in the trajectory data, so that the trajectory obtained by continuous sampling can be divided into multiple trajectory segments according to the temporal continuity and spatial change, thereby providing a clear analysis unit for subsequent trajectory continuity judgment.
[0053] Based on this, by filtering trajectory segments that meet the conditions of continuous time and continuous spatial change, discontinuous trajectory segments caused by short-term target loss, sampling anomalies, spatial jumps, or detection errors can be eliminated, thereby retaining trajectory data that can truly reflect the continuous motion state of candidate interactive objects.
[0054] Furthermore, based on the correspondence between each trajectory segment and the candidate interactive object in the continuous trajectory segment data, the candidate interactive object with continuous trajectory segments is determined, so that the trajectory continuity judgment result can be associated with the specific candidate interactive object, thereby avoiding the selection of the interactive subject based solely on the instantaneous detection result or a single trajectory point, and improving the reliability of the target interactive object selection.
[0055] By calculating the spatial position offset and motion direction change of the candidate interactive object corresponding to the continuous trajectory segment, and performing stability calculation based on this, the motion stability of the candidate interactive object can be quantitatively expressed, thereby distinguishing stable interactive actions from jitter, drift, sudden changes, or non-interactive movements.
[0056] After obtaining the trajectory stability parameter data, the candidate interactive objects are screened to extract those whose stability meets the preset stability conditions. The objects are then sorted according to their time order of entering the 3D interactive volume region. This process ensures that the determination of the target interactive object takes into account both trajectory stability and entry time order, thereby reducing the probability of misselection caused by target intersection, short-term occlusion, or detection fluctuations in multi-person or multi-target scenarios.
[0057] Finally, based on the trajectory data formed by the target interactive object within the three-dimensional interactive volume area, an interactive locking state is established between the target interactive object and the three-dimensional interactive volume area. The interactive locking state is kept unchanged unless the unlocking condition is met, so that subsequent interactive processing can continue around the same target interactive object, thereby reducing the frequent switching of interactive subjects and improving the continuity, stability and anti-interference ability of the interactive control of the intelligent display device.
[0058] In a preferred embodiment of the present invention, the trajectories of each candidate interactive object are segmented based on the time interval and spatial position change of adjacent moments in the trajectory data to generate multiple trajectory segments, including: Based on the trajectory data, the spatial location and corresponding time information of each candidate interactive object at continuous sampling time are extracted to generate time-series trajectory data; Calculate the time interval and spatial displacement change between adjacent sampling times based on the time-series trajectory data to generate trajectory change data; Based on the trajectory change data, it is determined whether the continuity condition is met between adjacent trajectory points. When the time interval exceeds the preset time threshold or the spatial displacement change exceeds the preset displacement change threshold, the trajectory is segmented to generate multiple trajectory segments. Based on the segmentation results, the trajectories of each candidate interactive object are segmented and marked to generate corresponding trajectory segment data.
[0059] In a preferred embodiment of the present invention, trajectory segments that satisfy the conditions of temporal continuity and spatial change continuity are selected based on the trajectory segments to generate continuous trajectory segment data, including: Based on the trajectory segment data, extract the time span and spatial variation range of each trajectory segment to generate trajectory segment feature data; Based on the trajectory segment feature data, the continuity of each trajectory segment is determined to determine whether the trajectory segment meets the time continuity condition in the time dimension and whether it meets the spatial change continuity condition in the spatial dimension. Based on the judgment results, trajectory segments that simultaneously meet the conditions of temporal continuity and spatial change continuity are selected to generate continuous trajectory segment data. Among them, the time continuity condition and the spatial change continuity condition are set by statistical analysis based on the statistical distribution of trajectory segment characteristics and the characteristics of actual interaction behavior.
[0060] In a preferred embodiment of the present invention, determining candidate interactive objects with continuous trajectory segments based on the correspondence between each trajectory segment in the continuous trajectory segment data and the candidate interactive objects includes: Based on the continuous trajectory segment data, obtain the trajectory segment identifier, start time, end time and the candidate interaction object identifier corresponding to each continuous trajectory segment; Based on the identifier of the candidate interaction object, the continuous trajectory segments belonging to the same candidate interaction object are classified and processed to generate a set of continuous trajectory segments corresponding to each candidate interaction object. Determine whether the set of continuous trajectory segments corresponding to each candidate interaction object is empty, and determine the candidate interaction object whose set of continuous trajectory segments is not empty as the candidate interaction object with continuous trajectory segments; Based on each candidate interaction object with a continuous trajectory segment and its corresponding set of continuous trajectory segments, the association between the candidate interaction object and the continuous trajectory segment is established for subsequent stability calculation.
[0061] In a preferred embodiment of the present invention, based on candidate interactive objects with continuous trajectory segments, the spatial position offset and motion direction change of the continuous trajectory segments are calculated, and stability calculation processing is performed on the corresponding candidate interactive objects based on the spatial position offset and motion direction change to generate trajectory stability parameter data, including: Based on the association between candidate interactive objects with continuous trajectory segments and continuous trajectory segments, extract the spatial location point sequence in the continuous trajectory segment corresponding to each candidate interactive object; Based on the sequence of spatial location points, calculate the spatial displacement between adjacent spatial location points, and calculate the spatial position offset of continuous trajectory segments based on the spatial displacement. Based on the spatial location point sequence, calculate the motion direction vector at adjacent sampling times, and calculate the change in the angle between adjacent motion direction vectors to obtain the change in motion direction of continuous trajectory segments; Calculate the position stability parameter and the direction stability parameter based on the spatial position offset and the change in motion direction, respectively. Based on the position stability parameters and orientation stability parameters, the comprehensive stability of the continuous trajectory segments corresponding to each candidate interaction object is calculated, and the corresponding trajectory stability parameter data is generated. Among them, the spatial position offset can be determined based on the displacement change between adjacent spatial position points, the degree of deviation of the continuous trajectory segment from the fitted trajectory line, or the degree of dispersion of spatial position points within the continuous trajectory segment; the motion direction change can be determined based on the change in the angle between adjacent motion direction vectors.
[0062] In a preferred embodiment of the present invention, candidate interactive objects with continuous trajectory segments are screened based on trajectory stability parameter data to extract candidate interactive objects whose stability meets preset stability conditions. These objects are then sorted according to their time sequence of entry into the three-dimensional interactive volume region to determine the target interactive object, including: Based on the trajectory stability parameter data, obtain the trajectory stability parameters corresponding to each candidate interactive object with continuous trajectory segments; The trajectory stability parameters corresponding to each candidate interaction object are compared with the preset stability conditions, and candidate interaction objects that meet the preset stability conditions are selected to generate a set of stable candidate interaction objects. Based on the trajectory data of each candidate interactive object in the stable candidate interactive object set, determine the time when each candidate interactive object first enters the three-dimensional interactive volume region and generate entry time information. Based on the entry time information, the candidate interaction objects in the stable candidate interaction object set are sorted according to their entry time, and a sorting result is generated. Based on the sorting results, the candidate interactive object that enters the three-dimensional interactive volume region earliest is determined as the target interactive object; The preset stability conditions are set based on the statistical distribution of trajectory stability parameters, system sampling frequency, normal operation trajectory characteristics of candidate interactive objects, and erroneous trigger trajectory characteristics in actual interactive scenarios.
[0063] In a preferred embodiment of the present invention, an interaction lock state is established between the target interaction object and the three-dimensional interaction volume region based on the trajectory data formed by the target interaction object within the three-dimensional interaction volume region, and the interaction lock state is kept unchanged unless the unlocking condition is met, including: Based on the target identifier, trajectory data, and spatial position of the target interactive object within the three-dimensional interactive volume area, establish the binding relationship between the target interactive object and the three-dimensional interactive volume area. Based on the binding relationship, an interaction lock tag is generated, and the target interaction object is identified as the current interaction processing object; In the interactive lock state, continuously acquire the spatial location data of the target interactive object, and update the trajectory data of the target interactive object based on the spatial location data; Based on the updated trajectory data, determine whether the target interactive object meets the unlocking conditions; If the target interaction object does not meet the unlocking condition, maintain the interaction lock flag and continue to use the target interaction object as the current interaction processing object; When the target interactive object meets the unlocking conditions, clear the interactive lock mark and release the binding relationship between the target interactive object and the three-dimensional interactive volume area; The unlocking conditions include the interruption of trajectory data continuity and the spatial deviation condition where the current spatial position of the target interactive object deviates from the three-dimensional interactive volume area by more than a preset distance threshold.
[0064] In a preferred embodiment of the present invention, the spatial position of each candidate interactive object is calculated in a spatial coordinate system based on the direction vector and distance value to obtain spatial position data under a unified coordinate system, including: The unit direction vector of each candidate interactive object in the spatial coordinate system is determined based on the direction vector, and the unit direction vector is decomposed into coordinates to generate component direction data along each coordinate axis. The component direction data is proportionally mapped based on the distance value to establish a correspondence between the unit direction component of each coordinate axis and the distance value, thereby generating component distance data. Based on the component distance data, coordinate synthesis processing is performed on the position of each candidate interactive object in the spatial coordinate system to generate the corresponding three-dimensional coordinate data. For the three-dimensional coordinate data calculated from the sensing data of different sensors, spatial consistency alignment processing is performed on each three-dimensional coordinate data to enable coordinate data from different sources to be uniformly expressed under the same reference coordinate system, generating spatial position data under a unified coordinate system.
[0065] In this embodiment of the invention, the direction vector is decomposed into coordinate components along each coordinate axis, providing basic data for subsequent spatial location calculations. Based on this, the unit direction components are proportionally mapped according to distance values, establishing a correspondence between direction and distance information, thus unifying direction and depth information within the same computational framework. Furthermore, coordinate synthesis is performed on the component distance data to generate three-dimensional coordinate data, enabling the position of candidate interactive objects to be expressed in spatial coordinates. Subsequently, three-dimensional coordinate data from different sensors are aligned for consistency, ensuring that different data sources are expressed under a unified reference coordinate system, thereby guaranteeing the consistency of spatial location data and providing a reliable foundation for subsequent spatial region construction and trajectory analysis.
[0066] In a preferred embodiment of the present invention, the unit direction vector of each candidate interactive object in the spatial coordinate system is determined based on the direction vector, and the unit direction vector is subjected to coordinate decomposition processing to generate component direction data along each coordinate axis, including: Based on the directional information in the synchronous sensing data, obtain the original directional vector of each candidate interactive object relative to the acquisition device; Based on the installation relationship between the acquisition device and the spatial coordinate system, the original direction vector is transformed into the spatial coordinate system to obtain the direction vector of each candidate interactive object in the spatial coordinate system; Normalize the direction vectors to obtain the unit direction vectors corresponding to each candidate interaction object; Based on the definition of the direction of each coordinate axis in the spatial coordinate system, the unit direction vector is projected onto the direction of each coordinate axis to obtain the unit direction component along each coordinate axis; Component direction data is generated based on the unit direction components along each coordinate axis.
[0067] In a preferred embodiment of the present invention, the component direction data is proportionally mapped according to the distance value to establish a correspondence between the unit direction component of each coordinate axis and the distance value, thereby generating component distance data, including: Based on the target distance information, obtain the distance values of each candidate interactive object relative to the acquisition device; Based on the candidate interaction object identifier or timestamp information, the distance value is associated with the component direction data of the corresponding candidate interaction object; Based on the distance value, the unit direction component along each coordinate axis in the component direction data is proportionally mapped to obtain the distance component in each coordinate axis direction. Generate component distance data based on the distance components along each coordinate axis. Among them, the component distance data is used to represent the position components of the candidate interactive object relative to the origin of the spatial coordinate system in each coordinate axis direction.
[0068] In a preferred embodiment of the present invention, for the three-dimensional coordinate data calculated from sensing data from different sensors, spatial consistency alignment processing is performed on each three-dimensional coordinate data, so that coordinate data from different sources can be uniformly expressed under the same reference coordinate system, generating spatial position data under a unified coordinate system, including: Obtain the installation position and orientation of each sensor relative to the acquisition device or intelligent display device, and determine the coordinate transformation relationship between the sensor coordinate system and the unified reference coordinate system for each sensor. Based on the coordinate transformation relationship, the three-dimensional coordinate data calculated from the sensing data of different sensors are transformed into a unified reference coordinate system. Based on the candidate interactive object identifier, sampling time or spatial proximity, the converted 3D coordinate data are correlated and matched to determine the multi-source 3D coordinate data belonging to the same candidate interactive object; Multi-source 3D coordinate data belonging to the same candidate interactive object are fused to obtain the spatial position of the candidate interactive object in a unified reference coordinate system. Based on the spatial position of each candidate interactive object in a unified reference coordinate system, generate spatial position data in a unified coordinate system. The coordinate transformation relationship is determined based on the installation position and orientation of each sensor and the system calibration results. The system calibration results are obtained through equipment installation calibration, initialization calibration, or measurement at known calibration points.
[0069] In a preferred embodiment of the present invention, determining the interaction direction region corresponding to the display interface based on the normal direction of the display interface and a preset direction offset angle threshold includes: The normal direction of the display interface is normalized, and the normalized normal direction is determined as the reference direction. Based on the reference direction and the preset direction offset angle threshold, determine the range of allowed interaction directions centered on the reference direction; Determine the horizontal boundary of the interactive direction area based on the display boundary of the display interface in the spatial coordinate system; Determine the interaction direction area corresponding to the display interface based on the allowed interaction direction range and horizontal boundary; The preset directional offset angle threshold is set according to the installation relationship between the acquisition device and the intelligent display device, the display size of the display interface, and the range of user operation angles in the actual interactive scenario.
[0070] In this embodiment of the invention, by normalizing the normal direction of the display interface and using the normalized normal direction as a reference direction, the determination of the interaction direction area has a unified and stable direction benchmark, thereby avoiding the problem of inconsistent direction judgment caused by differences in the dimensions or length of the normal direction.
[0071] Based on this, according to the reference direction and the preset direction offset angle threshold, the range of allowed interaction directions centered on the reference direction is determined, so that the interaction direction area can expand around the actual interaction direction of the display interface, thereby excluding the space range that deviates from the interaction direction of the display interface and reducing the interference of lateral objects, background targets or non-interactive actions on subsequent interaction judgments.
[0072] Furthermore, the horizontal boundary of the interactive direction area is determined based on the display boundary of the display interface in the spatial coordinate system. This ensures that the interactive direction area is constrained not only by the direction angle but also by the actual display range of the display interface, thereby preventing the interactive direction area from expanding infinitely in the horizontal space and improving the spatial matching degree between the interactive area and the display interface.
[0073] Finally, based on the allowed range of interactive directions and the lateral boundary, the interactive direction area corresponding to the display interface is determined, so that the interactive direction area has both directional constraints and lateral boundary constraints, thereby providing a clear directional and boundary basis for the subsequent depth limitation and construction of the three-dimensional interactive volume area. The preset direction offset angle threshold is set according to the installation relationship between the acquisition device and the intelligent display device, the display size of the display interface, and the range of user operation angles in the actual interaction scenario, so that the allowable interaction direction range can adapt to the actual installation conditions and user operation habits, thereby improving the rationality and applicability of the interaction direction area setting.
[0074] In a preferred embodiment of the present invention, based on a preset interaction distance range, the depth of the interaction direction region is defined along the normal direction of the display interface to construct a three-dimensional interaction volume region corresponding to the display interface, including: Determine the minimum and maximum interaction distances based on the preset interaction distance range; The normal direction of the display interface is used as the spatial depth direction, and the near-end depth boundary corresponding to the minimum interaction distance and the far-end depth boundary corresponding to the maximum interaction distance are determined in the spatial coordinate system. Based on the near-end depth boundary and the far-end depth boundary, the interaction direction region is defined along the spatial depth direction to obtain a depth-limited interaction space; Based on the directional boundary of the interaction direction area, the display boundary of the display interface, the near-end depth boundary, and the far-end depth boundary, a closed boundary constraint is applied to the depth-limited interaction space. The depth-constrained interactive space, after being constrained by closed boundaries, is defined as the three-dimensional interactive volume region corresponding to the display interface.
[0075] In this embodiment of the invention, by determining the minimum and maximum interaction distances according to a preset interaction distance range, the three-dimensional interaction volume region has a clear effective distance range in the spatial depth direction, thereby avoiding objects that are too close or too far away from being misjudged as valid interaction objects.
[0076] By using the normal direction of the display interface as the spatial depth direction, and determining the near-end depth boundary corresponding to the minimum interaction distance and the far-end depth boundary corresponding to the maximum interaction distance in the spatial coordinate system, the depth boundary can be kept consistent with the actual orientation of the display interface, thereby ensuring that the depth limit result matches the interaction space corresponding to the display interface.
[0077] Based on this, the interaction direction region is limited along the spatial depth direction according to the near-end depth boundary and the far-end depth boundary, resulting in a depth-limited interaction space. This limits the interaction direction region to within the effective interaction distance range in the depth dimension, thereby excluding spatial regions that do not meet the interaction distance requirements and reducing the impact of non-effective distance targets on interaction recognition.
[0078] Furthermore, based on the directional boundary of the interaction direction area, the display boundary of the display interface, the near-end depth boundary, and the far-end depth boundary, the depth-limited interaction space is constrained by a closed boundary, so that the interaction space has clear boundaries in the directional, horizontal, and depth dimensions, thereby forming a stable and determinable spatial volume range.
[0079] Finally, the depth-limited interaction space constrained by the closed boundary is determined as the three-dimensional interaction volume region corresponding to the display interface. This allows the three-dimensional interaction volume region to simultaneously reflect the spatial orientation, display range, and effective interaction distance range of the display interface, thereby providing a stable spatial judgment basis for the region mapping, trajectory generation, target locking, and interaction command parsing of candidate interaction objects.
[0080] Embodiments of the present invention also provide an interactive system for an intelligent display device based on multi-sensor fusion, the system comprising: The data acquisition module is used to acquire multi-source sensing data collected by the acquisition device, which includes multiple sensors. The multiple sensors are used to collect gesture image sequences, target distance information and direction information, and perform time alignment processing on the multi-source sensing data to generate synchronous sensing data. The target recognition and filtering module is used to perform target recognition and distance filtering based on synchronous perception data, and generate a set of candidate interactive objects within a preset interaction distance range; The region construction module is used to calculate the spatial position of each candidate interactive object in a preset spatial coordinate system based on the target distance and direction information in the synchronous sensing data, generate spatial position data, and construct a three-dimensional interactive volume region corresponding to the display interface based on the spatial posture information of the display interface, the normal direction of the display interface, and the preset interactive distance range. The trajectory generation module is used to map the set of candidate interactive objects to a three-dimensional interactive volume area based on spatial location data, calculate the spatial occupancy trajectory of each candidate interactive object, and generate the corresponding trajectory data. The object determination module is used to perform trajectory continuity analysis and stability assessment based on trajectory data, determine the target interactive object, and establish the interactive locking state corresponding to the target interactive object and the three-dimensional interactive volume region. The target tracking release module is used to continuously track the trajectory data of the target interactive object in the interactive lock state and perform continuity detection. When the continuity of the trajectory data is interrupted or the current spatial position of the target interactive object deviates from the three-dimensional interactive volume area by more than a preset distance threshold, the interactive lock state is released. The target re-determination module is used to redetermine the target interactive object after the interaction lock state is released, based on the time order in which each candidate interactive object re-enters the three-dimensional interactive volume area, and the consistency between the movement direction of each candidate interactive object and the interaction direction of the display interface. The interactive control module is used to perform trajectory parsing and processing based on the trajectory data corresponding to the target interactive object, generate corresponding interactive commands, and control the intelligent display device to perform corresponding operations based on the interactive commands.
[0081] It should be noted that this system is a system corresponding to the above method. All implementation methods in the above method embodiments are applicable to this embodiment and can achieve the same technical effect.
[0082] Embodiments of the present invention also provide a computing device, including: a processor and a memory storing a computer program, wherein the computer program, when executed by the processor, performs the method described above. All implementations in the above method embodiments are applicable to this embodiment and can achieve the same technical effects.
[0083] Embodiments of the present invention also provide a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to perform the method described above. All implementations in the above method embodiments are applicable to this embodiment and can achieve the same technical effects.
[0084] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. An interactive method for intelligent display devices based on multi-sensor fusion, characterized in that, The method includes: The system acquires multi-source sensing data collected by a data acquisition device, which includes multiple sensors. These sensors are used to acquire gesture image sequences, target distance information, and direction information. The system also performs time alignment processing on the multi-source sensing data to generate synchronous sensing data. Based on synchronous sensing data, target identification and distance filtering are performed to generate a set of candidate interactive objects within a preset interaction distance range; Based on the target distance and direction information in the synchronous sensing data, the spatial position of each candidate interactive object is calculated in the preset spatial coordinate system, spatial position data is generated, and a three-dimensional interactive volume region corresponding to the display interface is constructed based on the spatial posture information of the display interface, the normal direction of the display interface, and the preset interactive distance range. Based on spatial location data, the set of candidate interactive objects is mapped to a three-dimensional interactive volume region, the spatial occupancy trajectory of each candidate interactive object is calculated, and the corresponding trajectory data is generated. Based on the trajectory data, trajectory continuity analysis and stability assessment are performed to determine the target interaction object and establish the interaction locking state corresponding to the target interaction object and the three-dimensional interaction volume region. In the interactive lock state, the trajectory data of the target interactive object is continuously tracked and continuity detection is performed. When the continuity of trajectory data is interrupted or the current spatial position of the target interactive object deviates from the three-dimensional interactive volume area by more than a preset distance threshold, the interactive lock state is released. After the interaction lock is released, the target interaction object is re-determined based on the time sequence of each candidate interaction object re-entering the three-dimensional interaction volume area, and the consistency between the movement direction of each candidate interaction object and the interaction direction of the display interface. The trajectory data corresponding to the target interactive object is parsed and processed to generate corresponding interactive commands, and the intelligent display device is controlled to perform corresponding operations according to the interactive commands. Based on the trajectory data, trajectory continuity analysis and stability assessment are performed to determine the target interaction object and establish the interaction lock state corresponding to the target interaction object and the three-dimensional interaction volume region, including: Based on the time interval and spatial location change of adjacent moments in the trajectory data, the trajectory of each candidate interaction object is segmented to generate multiple trajectory segments. Based on the trajectory segments, select those that meet the conditions of continuous time and continuous spatial change to generate continuous trajectory segment data; Based on the correspondence between each trajectory segment in the continuous trajectory segment data and the candidate interactive objects, the candidate interactive objects with continuous trajectory segments are determined. Based on candidate interactive objects with continuous trajectory segments, calculate the spatial position offset and motion direction change of their continuous trajectory segments, and perform stability calculation on the corresponding candidate interactive objects based on the spatial position offset and motion direction change to generate trajectory stability parameter data. Based on the trajectory stability parameter data, candidate interactive objects with continuous trajectory segments are screened and extracted. Candidate interactive objects whose stability meets the preset stability conditions are extracted and sorted according to the time order of entering the three-dimensional interactive volume region to determine the target interactive object. Based on the trajectory data formed by the target interactive object within the three-dimensional interactive volume area, an interactive locking state is established between the target interactive object and the three-dimensional interactive volume area, and the interactive locking state remains unchanged unless the unlocking condition is met.
2. The interactive method for intelligent display devices based on multi-sensor fusion according to claim 1, characterized in that, Based on the target distance and orientation information in the synchronous sensing data, the spatial position of each candidate interactive object is calculated in a preset spatial coordinate system, generating spatial position data. Then, based on the spatial pose information of the display interface, the normal direction of the display interface, and the preset interactive distance range, a three-dimensional interactive volume region corresponding to the display interface is constructed, including: Based on the fixed installation relationship between the data acquisition device and the intelligent display device, a spatial coordinate system corresponding to the display interface is established with the data acquisition device as the origin. Based on the direction information, determine the direction vector of each candidate interactive object relative to the acquisition device, and based on the target distance information, determine the distance value of each candidate interactive object on the corresponding direction vector; Based on the direction vector and distance value, the spatial position of each candidate interactive object is calculated in the spatial coordinate system to obtain spatial position data under a unified coordinate system; Obtain the spatial orientation information of the display interface in the spatial coordinate system, and determine the normal direction of the display interface based on the spatial orientation information; Based on the normal direction of the display interface and the preset direction offset angle threshold, determine the interaction direction area corresponding to the display interface; Based on the preset interaction distance range, the depth of the interaction direction area is defined along the normal direction of the display interface to construct a three-dimensional interaction volume area corresponding to the display interface.
3. The interactive method for intelligent display devices based on multi-sensor fusion according to claim 1, characterized in that, Based on spatial location data, the set of candidate interactive objects is mapped to a 3D interactive volume region. The spatial occupancy trajectory of each candidate interactive object is calculated, and the corresponding trajectory data is generated, including: Based on the spatial location data, the spatial locations of each candidate interactive object at consecutive sampling times are extracted and arranged in chronological order to generate location sequence data; Based on the location sequence data, the spatial location at each time point is matched with the three-dimensional interactive volume region to determine whether the candidate interactive object is located within the three-dimensional interactive volume region and generate region status marker data. Based on the regional status marker data, continuous spatial locations within the three-dimensional interactive volume region are filtered and correlated in chronological order to generate initial trajectory data; Abnormal location points caused by sampling errors in the initial trajectory data are identified and their positions are corrected to eliminate discrete jumps in the trajectory and generate corresponding trajectory data.
4. The interactive method for intelligent display devices based on multi-sensor fusion according to claim 2, characterized in that, Based on the direction vector and distance value, the spatial position of each candidate interactive object is calculated in the spatial coordinate system, resulting in spatial position data under a unified coordinate system, including: The unit direction vector of each candidate interactive object in the spatial coordinate system is determined based on the direction vector, and the unit direction vector is decomposed into coordinates to generate component direction data along each coordinate axis. The component direction data is proportionally mapped based on the distance value to establish a correspondence between the unit direction component of each coordinate axis and the distance value, thereby generating component distance data. Based on the component distance data, coordinate synthesis processing is performed on the position of each candidate interactive object in the spatial coordinate system to generate the corresponding three-dimensional coordinate data. For the three-dimensional coordinate data calculated from the sensing data of different sensors, spatial consistency alignment processing is performed on each three-dimensional coordinate data to enable coordinate data from different sources to be uniformly expressed under the same reference coordinate system, generating spatial position data under a unified coordinate system.
5. The interactive method for intelligent display devices based on multi-sensor fusion according to claim 2, characterized in that, Based on the normal direction of the display interface and a preset direction offset angle threshold, determine the interaction direction area corresponding to the display interface, including: The normal direction of the display interface is normalized, and the normalized normal direction is determined as the reference direction. Based on the reference direction and the preset direction offset angle threshold, determine the range of allowed interaction directions centered on the reference direction; Determine the horizontal boundary of the interactive direction area based on the display boundary of the display interface in the spatial coordinate system; Determine the interaction direction area corresponding to the display interface based on the allowed interaction direction range and horizontal boundary; The preset directional offset angle threshold is set according to the installation relationship between the acquisition device and the intelligent display device, the display size of the display interface, and the range of user operation angles in the actual interactive scenario.
6. The interactive method for intelligent display devices based on multi-sensor fusion according to claim 2, characterized in that, Based on a preset interaction distance range, the depth of the interaction direction area is defined along the normal direction of the display interface to construct a three-dimensional interaction volume area corresponding to the display interface, including: Determine the minimum and maximum interaction distances based on the preset interaction distance range; The normal direction of the display interface is used as the spatial depth direction, and the near-end depth boundary corresponding to the minimum interaction distance and the far-end depth boundary corresponding to the maximum interaction distance are determined in the spatial coordinate system. Based on the near-end depth boundary and the far-end depth boundary, the interaction direction region is defined along the spatial depth direction to obtain a depth-limited interaction space; Based on the directional boundary of the interaction direction area, the display boundary of the display interface, the near-end depth boundary, and the far-end depth boundary, a closed boundary constraint is applied to the depth-limited interaction space. The depth-constrained interactive space, after being constrained by closed boundaries, is defined as the three-dimensional interactive volume region corresponding to the display interface.
7. An interactive system for an intelligent display device based on multi-sensor fusion, characterized in that, The system, used in the method of any one of claims 1 to 6, comprises: The data acquisition module is used to acquire multi-source sensing data collected by the acquisition device, which includes multiple sensors. The multiple sensors are used to collect gesture image sequences, target distance information and direction information, and perform time alignment processing on the multi-source sensing data to generate synchronous sensing data. The target recognition and filtering module is used to perform target recognition and distance filtering based on synchronous perception data, and generate a set of candidate interactive objects within a preset interaction distance range; The region construction module is used to calculate the spatial position of each candidate interactive object in a preset spatial coordinate system based on the target distance and direction information in the synchronous sensing data, generate spatial position data, and construct a three-dimensional interactive volume region corresponding to the display interface based on the spatial posture information of the display interface, the normal direction of the display interface, and the preset interactive distance range. The trajectory generation module is used to map the set of candidate interactive objects to a three-dimensional interactive volume area based on spatial location data, calculate the spatial occupancy trajectory of each candidate interactive object, and generate the corresponding trajectory data. The object determination module is used to perform trajectory continuity analysis and stability assessment based on trajectory data, determine the target interactive object, and establish the interactive locking state corresponding to the target interactive object and the three-dimensional interactive volume region. The target tracking release module is used to continuously track the trajectory data of the target interactive object in the interactive lock state and perform continuity detection. When the continuity of the trajectory data is interrupted or the current spatial position of the target interactive object deviates from the three-dimensional interactive volume area by more than a preset distance threshold, the interactive lock state is released. The target re-determination module is used to redetermine the target interactive object after the interaction lock state is released, based on the time order in which each candidate interactive object re-enters the three-dimensional interactive volume area, and the consistency between the movement direction of each candidate interactive object and the interaction direction of the display interface. The interactive control module is used to perform trajectory parsing and processing based on the trajectory data corresponding to the target interactive object, generate corresponding interactive commands, and control the intelligent display device to perform corresponding operations based on the interactive commands.
8. A computing device, characterized in that, include: One or more processors; A storage device for storing one or more programs that, when executed by one or more processors, cause the one or more processors to implement the method as described in any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program that, when executed by a processor, implements the method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Gesture recognition and projection fusion non-contact interaction method and device, equipment and medium
CN121236815A
Intelligent interaction system and method based on gesture recognition
CN121811499A