Digital multimedia exhibition hall interactive display control method and system

By using a dynamic visual sensing event camera array and the Loihi 2 neuromorphic chip to build a closed-loop control system in a digital multimedia exhibition hall, the problems of response latency, privacy compliance, and insufficient recognition granularity were solved. This resulted in low-latency, privacy-friendly gesture recognition and closed-loop anti-interference effects, improving the system's engineering reliability and the audience's interactive experience.

CN122507280APending Publication Date: 2026-08-04SHANGHAI LEAPIDEAS MULTIMEDIA SYST CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANGHAI LEAPIDEAS MULTIMEDIA SYST CO LTD
Filing Date
2026-05-13
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

Existing digital multimedia exhibition halls suffer from issues such as response latency, privacy compliance, insufficient recognition granularity, and inadequate closed-loop interference immunity in audience perception and interactive control, making it difficult to achieve control effects that are low-latency, privacy-friendly, fine-grained gesture recognition, and closed-loop interference immunity.

Method used

A closed-loop control system is constructed using a dynamic visual sensing event camera array and the Loihi 2 neuromorphic chip. Through asynchronous event stream processing and directional coherent detection, it identifies gesture categories such as waving, pushing, and selecting, and applies gating suppression after the actuator moves, forming a complete closed-loop control.

Benefits of technology

It significantly reduces audience motion capture latency, improves the selectivity and accuracy of gesture recognition, reduces power consumption, enhances system robustness and the naturalness of audience response, and avoids secondary triggering of non-interactive visual events.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122507280A_ABST
    Figure CN122507280A_ABST
Patent Text Reader

Abstract

This invention relates to the field of digital multimedia technology, and more specifically, to a digital multimedia exhibition hall interactive display control method and system, comprising the following steps: Step 1, an asynchronous event stream is acquired by a dynamic visual sensing event camera array deployed in the exhibition hall, and the propagation direction number of each reserved event is determined based on the adjacency of previously reserved events; Step 2, a primitive neural column array corresponding one-to-one with the interactive primitive unit is constructed on the Loihi 2 neuromorphic chip, and the intent decoding layer decodes the interactive intent binary representing the gesture category and the exhibition item space partition; Step 3, the closed-loop controller encapsulates the interactive intent binary into a control command message and sends it to the multimedia display execution mechanism corresponding to the exhibition item space partition via the fieldbus. This invention achieves significant effects in terms of response latency, privacy compliance, direction recognition selectivity, power consumption, gesture recognition granularity, and closed-loop anti-interference capability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of automatic control technology, specifically relating to a digital multimedia exhibition hall interactive display control method and system. Background Technology

[0002] Digital multimedia exhibition halls are a common core display format in museums, science and technology museums, corporate exhibition centers, brand experience stores, and cultural tourism projects. They typically integrate various multimedia display mechanisms such as projection, directional sound fields, intelligent dimming, and electrochromic glass. The content of the exhibits is driven by the audience's physical movements or positional changes, thereby achieving an immersive and personalized interactive experience. Currently, audience perception and interactive control in digital multimedia exhibition halls have generally developed along two technical routes. The first route is based on traditional frame-based RGB cameras or depth cameras to capture audience images, and then uses deep learning models such as convolutional neural networks for gesture recognition and position localization. Typical examples include open-source frameworks such as OpenPose and MediaPipe Hands, as well as intelligent interaction solutions from various vendors. However, this approach is limited by the frame-based acquisition mechanism, which has limited temporal resolution at a fixed frame rate. It is prone to missed detections and delays for rapid gestures, and the continuous output of image frames that can recognize faces by the sensors faces significant limitations in the increasingly strict privacy compliance environment of exhibition halls. The second approach uses non-visual sensing methods such as piezoelectric floors, millimeter-wave radar, and ultra-wideband positioning tags to estimate audience position or coarse-grained gestures. While this approach is more privacy-friendly, it struggles to accurately distinguish specific gesture categories such as waving, pushing, and selecting, resulting in a mismatch between the recognition granularity and the interactive needs of the exhibits. In the control phase, existing exhibition halls often use centralized industrial control computers with rule engines or finite state machines to schedule exhibits. After the actuators complete their actions, there is usually a lack of reverse gating mechanisms for the front-end event flow. This means that visual disturbances caused by actions such as projection switching and dimming changes can be easily identified as new gestures by sensors, leading to unexpected secondary triggers. In summary, existing technologies have unresolved issues in four aspects: response latency, privacy compliance, recognition granularity, and closed-loop robustness. There is an urgent need for a digital multimedia exhibition hall interactive display control method and system that balances low latency, privacy friendliness, fine-grained gesture recognition, and closed-loop interference resistance. Summary of the Invention

[0003] The main objective of this invention is to provide a digital multimedia exhibition hall interactive display control method and system, which achieves significant results in terms of response latency, privacy compliance, direction recognition selectivity, power consumption, gesture recognition granularity, and closed-loop anti-interference capability. It effectively shields non-interactive visual events caused by the actions of the actuator itself, avoids secondary triggering, and improves the naturalness of the audience's perceived response and the overall engineering reliability of the system.

[0004] To solve the above problems, the technical solution of the present invention is implemented as follows: A method for controlling interactive displays in a digital multimedia exhibition hall includes the following steps: Step 1: The asynchronous event stream is collected by the dynamic visual sensing event camera array deployed in the exhibition hall. After the events are assigned to the interactive primitive units divided in the virtual hexagonal grid coordinate system, they are split according to the polarity identifier and formed into open polarity pulse trains and closed polarity pulse trains after refractory period retention processing. The propagation direction number of each retained event is determined based on the adjacent orientation of the previously retained events. Step 2: Construct a primitive neural column array on the Loihi 2 neuromorphic chip, corresponding one-to-one with the interactive primitive units. Each primitive neural column contains 6 directional coherent detection channels. Each channel receives open-polarity pulses in the same direction and closed-polarity pulses in the opposite direction via open-polarity delay branches and closed-polarity delay branches, respectively, and sends them to the independent input channels of the coherent detection neurons. When all independent input channels receive pulses within a preset coherence window, they emit coherent pulses carrying directional labels. The column-level output neurons output column-level output pulses carrying the winning directional label within a preset integration window. The intent decoding layer decodes the interactive intent representing the gesture category and the spatial partitioning of the exhibit. Figure 2 tuple; Step 3, the closed-loop controller will interact with the intended meaning. Figure 2 The tuple is encapsulated as a control command message and sent to the multimedia display execution mechanism corresponding to the exhibit space partition via the fieldbus. After the multimedia display execution mechanism completes its action, it sends the completion status back to the closed-loop controller. The closed-loop controller applies gating suppression to the asynchronous event stream within a preset stable time and updates the environmental reference of the next decision window, so that the dynamic visual sensing event camera array, Loihi 2 neuromorphic chip and multimedia display execution mechanism form a closed loop.

[0005] Furthermore, the dynamic visual sensing event camera array is deployed on the top of the digital multimedia exhibition hall in a triangular grid manner. After time synchronization and extrinsic parameter calibration, the event timestamp time base of each dynamic visual sensing event camera is aligned with the unified time base of the exhibition hall, and the rectangular pixel coordinates of each dynamic visual sensing event camera are mapped to the unified coordinate system of the exhibition hall through a preset calibration mapping table. A virtual regular hexagonal grid coordinate system is constructed in the unified coordinate system of the exhibition hall and interactive primitive units are divided according to a preset step distance. The interactive primitive unit has 6 adjacent directions, which are denoted as direction 1, direction 2, direction 3, direction 4, direction 5 and direction 6, respectively. Among them, direction 1 and direction 4 are opposite directions, direction 2 and direction 5 are opposite directions, and direction 3 and direction 6 are opposite directions.

[0006] Furthermore, each event in the asynchronous event stream contains pixel coordinates, a timestamp, and a polarity identifier, with the polarity identifier being either open or closed polarity. The pixel coordinates of each event are converted to coordinates in the unified coordinate system of the exhibition hall using a preset calibration mapping table. The event is then assigned to the interactive primitive unit corresponding to the nearest grid center point based on the converted coordinates. Finally, the event streams are split according to the polarity identifier to obtain open polarity event sub-streams and closed polarity event sub-streams. The refractory period retention process is as follows: in the same polarity event sub-stream of the same interactive primitive unit, the earliest event that arrives after a preset refractory period from the timestamp of a retained event is taken as the next retained event, thus forming open polarity pulse trains and closed polarity pulse trains.

[0007] Furthermore, the propagation direction number is determined as follows: One preset neighborhood duration is traced back from the release time of the current reserved event. In the virtual hexagonal grid coordinate system, the nearest prior reserved event of the same polarity is searched among the six primitive positions adjacent to the current reserved event. The adjacent orientation of the prior reserved event relative to the current reserved event is denoted as orientation j. The direction number opposite to orientation j is used as the propagation direction number of the current reserved event. If the number of prior reserved events of the same polarity found within the preset neighborhood duration is 0, the current reserved event is marked as a directionless event and sent to a preset directionless background channel, causing the directionless event to exit subsequent propagation direction-related processing. The source address of a pulse with a propagation direction number is composed of the number of its corresponding interactive primitive unit, its polarity identifier, and the propagation direction number. The release time of the pulse is equal to the timestamp of its corresponding reserved event.

[0008] Furthermore, both the open-polarity delay branch and the closed-polarity delay branch are composed of several relay neurons connected in series. The relay neurons are connected one-to-one by single pulse routing on the Loihi 2 neuromorphic chip. Each pulse generates a fixed unit propagation delay when passing through a relay neuron in the delay branch. The number of relay neurons in the open-polarity delay branch and the closed-polarity delay branch are configured according to a uniform level rule. For the directional coherence detection channel numbered k, k is sequentially selected from 1 to 6. The pulse with propagation direction numbered k in the open-polarity pulse train is sent to the input terminal of the open-polarity delay branch of the directional coherence detection channel numbered k. The pulse with propagation direction numbered k in the closed-polarity pulse train is sent to the input terminal of the closed-polarity delay branch of the directional coherence detection channel numbered k.

[0009] Furthermore, the tail end of the open polarity delay branch is connected to the coherent detection neuron via the first synaptic port, and the tail end of the closed polarity delay branch is connected to the second synaptic port. The first and second synaptic ports are configured as independent input ports on the Loihi 2 neuromorphic chip or distinguished by the graded spike payload field, so that the coherent detection neuron can identify open polarity input and closed polarity input respectively. The firing rule of the coherent detection neuron is as follows: when both the first and second synaptic ports receive pulses within a preset coherence window, a coherent pulse carrying a direction label k is fired, and after firing, the valid reception flags of the first and second synaptic ports within the current preset coherence window are cleared, so that the open polarity pulse and closed polarity pulse of each coherent pair are independently paired.

[0010] Furthermore, the column-level output neuron receives coherent pulses from the six directional coherent detection channels within the primitive neural column. Within a preset integration window, the direction label that first reaches the preset trigger count is taken as the winning direction label. When the preset integration window ends, the column-level output neuron sends a column-level output pulse carrying the winning direction label and the primitive neural column number to the direction channel corresponding to the winning direction label in the intent decoding layer.

[0011] Furthermore, the intent decoding layer includes a set of gesture category neurons, which includes a waving category branch, a pushing / pressing category branch, and a circling category branch. Each category branch consists of a series of sequentially cascaded state detection neurons, time-gated neurons assigned to each level of state detection neurons, resetting inhibition neurons, and a termination trigger neuron at the final level. Each level of state detection neuron is assigned to listen to a preset combination of direction labels and primitive neural column numbers. After the previous level of state detection neuron fires, it drives the time-gated neurons assigned to the next level of state detection neuron to enter the gating open state within a preset inter-level interval. Within the gating open state, when the next level of state detection neuron receives a column-level output pulse that matches its assigned combination of direction labels and primitive neural column numbers, it fires and passes the gating to the next level. If the preset inter-level interval is... When the number of matching column-level output pulses at the end is 0, or when a non-matching column-level output pulse reaches the preset interference count within the listening range of this category branch, the reset inhibition neuron sends a reset pulse to all state detection neurons and all time-gated neurons of this category branch to return it to the initial state; when the final-level state detection neuron fires, it triggers the termination trigger neuron of this category branch to fire a gesture recognition pulse; among them, the inter-level listening direction labels of the waving category branch are arranged alternately in direction 1 and direction 4; the push-press category branch maps the visual axis direction of each exhibit space partition to the push-press listening direction label as the inter-level listening direction label through the local direction mapping table, and the state detection neurons at each level listen to the adjacent primitive nerve column number strings arranged from far to near along the visual axis direction of the exhibit; the inter-level listening direction labels of the circle category branch are arranged cyclically in direction 1 to direction 6.

[0012] Furthermore, the intent decoding layer also includes a set of exhibit partition neurons, gesture inhibition interneurons coupled to the set of gesture category neurons, and partition inhibition interneurons coupled to the set of exhibit partition neurons. Each exhibit partition neuron is bound to one exhibit space partition, receiving column-level output pulses from all primitive neural columns covering its exhibit space partition. When the number of received pulses reaches a preset trigger count within a preset counting window, a partition activation pulse is emitted. The gesture inhibition interneuron receives gesture recognition pulses from all termination trigger neurons in the gesture category neuron set and emits reset pulses to termination trigger neurons, state detection neurons, and time-gated neurons in other category branches except for the branch where the emission source is located. The partition inhibition interneuron receives partition activation pulses from all exhibit partition neurons in the exhibit partition neuron set and emits reset pulses to exhibit partition neurons other than the emission source, causing their count state to return to zero. Within a decision window, the gesture category corresponding to the gesture recognition pulse retained by the gesture inhibition interneuron is bound to the exhibit space partition corresponding to the partition activation pulse retained by the partition inhibition interneuron, thus obtaining the interaction intent of this decision window. Figure 2Tuple; After receiving the completion status, the closed-loop controller applies gating suppression to filter out visual events in the asynchronous event stream that are located within the coverage area of ​​the exhibit space partition and are caused by the action of the multimedia display execution mechanism within a preset stable time period. After the preset stable time period ends, the current multimedia display status of the exhibit space partition is used as the environmental benchmark for the next decision window.

[0013] This invention offers the following advantages: By constructing a complete closed loop through a dynamic visual sensing event camera array, the Loihi 2 neuromorphic chip, and a multimedia display execution mechanism via a closed-loop controller, it achieves outstanding technical results in multiple dimensions. First, by replacing traditional frame-based image acquisition with asynchronous pixel-level brightness change events from the dynamic visual sensing event camera, the capture delay of audience movements reaches the microsecond level, significantly outperforming fixed frame rate schemes. Furthermore, the event stream only carries brightness change information, leaving no image frames that can identify faces, thus achieving privacy-friendly event acquisition while ensuring recognition accuracy. Second, by using a virtual hexagonal grid coordinate system to infer propagation direction numbers based on adjacent orientations, the direction estimation is isotropic in all six directions, avoiding the direction bias introduced by rectangular neighborhoods. Simultaneously, the cross-feeding rule of open-polarity pulses and opposite-direction closed-polarity pulses in the direction coherence detection channel ensures that this channel is only sensitive to genuine gestures with direction reversal characteristics, naturally suppressing pseudo-signals such as unidirectional light flicker and specular reflections, significantly improving recognition selectivity. Third, the Loihi 2 neuromorphic chip directly processes sparse event streams using a pulse-driven approach, resulting in significantly lower power consumption than general-purpose graphics processors, making the dense deployment of large-scale primitive neural column arrays feasible in engineering. Fourth, the intent decoding layer identifies different gesture categories such as waving, pushing, and selecting through cascaded state detection, time gating, and reset suppression mechanisms, and ensures unique interactive intents through a winner-takes-all decision. Figure 2 The tuples provide fine-grained recognition and are resistant to interference. Fifth, after the actuator performs its action, the closed-loop controller applies gating suppression to the affected exhibit space partitions and updates the environmental baseline within a preset stable time. This effectively shields non-interactive visual events caused by the action itself, avoids secondary triggering, and improves the robustness of the closed loop and the naturalness of the audience's perceived response. Attached Figure Description

[0014] Figure 1 A geometrical schematic diagram illustrating the principle of determining the propagation direction numbering in a virtual hexagonal grid coordinate system provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of the timing waveforms of coherent detection neuron firing and column-level output pulse generation provided in an embodiment of the present invention; Figure 3 A schematic diagram of experimental curves illustrating the closed-loop control timing and gating suppression principle provided in an embodiment of the present invention. Detailed Implementation

[0015] A method for controlling interactive displays in a digital multimedia exhibition hall includes the following steps: Step 1: The asynchronous event stream is collected by the dynamic visual sensing event camera array deployed in the exhibition hall. After the events are assigned to the interactive primitive units divided in the virtual hexagonal grid coordinate system, they are split according to the polarity identifier and formed into open polarity pulse trains and closed polarity pulse trains after refractory period retention processing. The propagation direction number of each retained event is determined based on the adjacent orientation of the previously retained events. Step 2: Construct a primitive neural column array on the Loihi 2 neuromorphic chip, corresponding one-to-one with the interactive primitive units. Each primitive neural column contains 6 directional coherent detection channels. Each channel receives open-polarity pulses in the same direction and closed-polarity pulses in the opposite direction via open-polarity delay branches and closed-polarity delay branches, respectively, and sends them to the independent input channels of the coherent detection neurons. When all independent input channels receive pulses within a preset coherence window, they emit coherent pulses carrying directional labels. The column-level output neurons output column-level output pulses carrying the winning directional label within a preset integration window. The intent decoding layer decodes the interactive intent representing the gesture category and the spatial partitioning of the exhibit. Figure 2 tuple; Step 3, the closed-loop controller will interact with the intended meaning. Figure 2 The tuple is encapsulated as a control command message and sent to the multimedia display execution mechanism corresponding to the exhibit space partition via the fieldbus. After the multimedia display execution mechanism completes its action, it sends the completion status back to the closed-loop controller. The closed-loop controller applies gating suppression to the asynchronous event stream within a preset stable time and updates the environmental reference of the next decision window, so that the dynamic visual sensing event camera array, Loihi 2 neuromorphic chip and multimedia display execution mechanism form a closed loop.

[0016] The digital multimedia exhibition hall features multiple dynamic visual sensing event cameras suspended from the ceiling in a triangular grid pattern. A typical deployment density is one camera per 9 to 25 square meters, with a spacing of 4 to 6 meters between adjacent cameras. The camera optical axes are vertically downward or deviate downward from the vertical direction by no more than 15 degrees. The triangular grid ensures that the horizontal distance from any point on the floor within the exhibition hall to the nearest three cameras is approximately equal, resulting in a uniformly distributed field-of-view overlap. Each dynamic visual sensing event camera uses a sensor with asynchronous output capability for brightness changes, such as an event sensor with a resolution of 1280×720, a dynamic range exceeding 120 dB, and a pixel event delay of less than 100 microseconds. The horizontal field of view of the camera lens is between 90 and 120 degrees, ensuring that, with a hall height of 4 to 5 meters, the ground coverage diameter of a single camera reaches more than 6 meters, thereby guaranteeing at least 30% overlap between the fields of view of adjacent cameras within the triangular grid.

[0017] After the physical deployment of the cameras was completed, two pre-processing steps were performed on the dynamic visual sensing event camera array: time synchronization and extrinsic parameter calibration. Time synchronization adopted the IEEE 1588 precise time protocol, enabling boundary clocks at the switch layer to align the event timestamp time bases of all dynamic visual sensing event cameras to the unified time base of the exhibition hall, achieving a synchronization accuracy better than 1 microsecond. If the camera hardware supported external trigger inputs, the event timestamp counters of all cameras were periodically reset via a synchronization trigger line to avoid cumulative drift after long-term operation. Extrinsic parameter calibration employed the scintillation calibration light source method: infrared LEDs driven by square waves from 100 Hz to 1000 Hz were placed at known coordinate points on the exhibition hall floor. Each dynamic visual sensing event camera collected event streams for 2 to 5 minutes and extracted the event cluster centers corresponding to the calibration light sources. Using multiple cluster centers as corresponding points, the extrinsic parameters of each dynamic visual sensing event camera relative to the unified coordinate system of the exhibition hall were solved. Finally, bundle adjustment was used to jointly optimize the pose of all cameras. It is recommended that the calibration results be reviewed every 30 days, and the ceiling of the exhibition hall be redone if there are structural changes.

[0018] After calibration, the rectangular pixel coordinates of each dynamic visual sensing event camera can be mapped to the unified coordinate system of the exhibition hall using a preset calibration mapping table. The unified coordinate system of the exhibition hall takes any fixed reference point on the exhibition hall floor as its origin, with east as the positive x-axis, north as the positive y-axis, and vertically upward as the positive z-axis. The essence of the mapping is to project the pixel ray onto a preset interactive plane based on the camera's intrinsic and extrinsic parameter matrices. This interactive plane is a horizontal plane 1.2 to 1.6 meters above the exhibition hall floor, to match the typical hand-eye coordination height of adult viewers. Let the pixel coordinates of the event be... The coordinates of the exhibition hall after the event is mapped are: The following projection relationship is adopted: ;in , The horizontal and vertical pixel coordinates of the event on the image plane of the dynamic visual sensing event camera. , Let the horizontal and vertical coordinates of the event be defined in the unified coordinate system of the exhibition hall. The dynamic visual sensing event camera uses a pre-calibrated 3x3 homography matrix for the interaction plane, which comprehensively includes camera intrinsic and extrinsic parameters, as well as the interaction plane equations. To improve efficiency, the above equations are not calculated online in the actual system; instead, the mapping result for each pixel is pre-written into a two-dimensional lookup table array, and event attribution is directly performed using... For index reading The lookup table is accurate to the millimeter level.

[0019] A virtual hexagonal grid coordinate system is divided within the unified coordinate system of the exhibition hall according to a preset step distance. Preferably, a flat-topped hexagonal paving is used, and the side length of each hexagon is denoted as . The preset step size, typically 0.10 to 0.20 meters, roughly corresponds to the coverage area of ​​an adult's palm, ensuring appropriate spatial resolution for subsequent directional coherence detection of gesture-level movements. For finer finger-level movements, the step size can be reduced to 0.04 to 0.08 meters; for trunk-level displacements, it can be increased to 0.30 to 0.50 meters. Each regular hexagon constitutes an interactive primitive unit, using axial coordinates... Unique identifier and All values ​​are integers. The transformation from the unified coordinate system of the exhibition hall to the axial coordinate system uses the following formula: ;in , Let the horizontal and vertical coordinates of the aforementioned events be defined in the unified coordinate system of the exhibition hall. The side length of a regular hexagon is... , These represent the axial horizontal and axial vertical coordinates of the event in a virtual hexagonal grid coordinate system. The calculated values ​​are... After rounding and verification to an integer, the interaction primitive corresponding to the nearest grid center point is obtained, and the event is assigned to that interaction primitive. This assignment method has low computational overhead and the assignment result is stable when the event falls near the hexagonal boundary.

[0020] Each interactive primitive unit has 6 adjacent directions. Starting from the geometric center of the interactive primitive unit, the 6 directions pointing from the vertices of the flat-topped regular hexagon to the centers of adjacent interactive primitive units are numbered in counter-clockwise order as direction 1, direction 2, direction 3, direction 4, direction 5, and direction 6. Directions 1 and 4 are opposite directions, directions 2 and 5 are opposite directions, and directions 3 and 6 are opposite directions. There is a fixed correspondence between the direction number and the axial coordinate offset: direction 1 corresponds to... Direction 2 corresponds to Direction 3 corresponds Direction 4 corresponds Direction 5 corresponds to Direction 6 corresponds to In a six-neighbor structure, the geometric distances from the six neighbors to the center are strictly equal, and the directions of motion are isotropic in terms of angle.

[0021] Each event in the asynchronous event stream contains three types of information: pixel coordinates, timestamp, and polarity identifier. The polarity identifier is either "on" or "off." "On" polarity indicates that the logarithmic light intensity at the pixel location has changed positively relative to the time of the last output event, exceeding a threshold. "Off" polarity indicates a negative change exceeding a threshold. After spatial assignment, events are split according to their polarity identifiers: all "on" polarity events within the same interaction primitive unit merge into the "on" polarity event substream of that interaction primitive unit, and "off" polarity events merge into the "off" polarity event substream of that interaction primitive unit. The two event substreams do not mix; the split event substreams maintain their original time sequence according to the event's original timestamp. Polarity splitting is a prerequisite for subsequent directional coherence detection utilizing edge symmetry: a moving bright / dark edge triggers an "on" polarity event at the leading edge and an "off" polarity event at the trailing edge, with the two types of events occurring in pairs.

[0022] Next, refractory period retention processing is applied to the open polarity event sub-stream and the closed polarity event sub-stream respectively. The refractory period retention processing refers to the following: within the same polarity event sub-stream of the same interactive primitive unit, the earliest event arriving after a preset refractory period from the timestamp of a retained event is taken as the next retained event; events falling within the preset refractory period are discarded. The typical value of the preset refractory period is 200 microseconds to 1 millisecond, preferably 500 microseconds; when the exhibition hall lighting is dim and the sensor noise level is high, it can be relaxed to 2 milliseconds; when the application scenario mainly involves rapid gestures and requires higher temporal resolution, it can be shortened to 100 microseconds. When a single pixel of a dynamic visual sensing event camera crosses a bright-dark edge, it usually outputs multiple redundant events continuously due to pixel circuit jitter and differential threshold boundary fluctuations. The event sequence after refractory period retention processing retains the relative temporal relationship of the edge crossing events while keeping the instantaneous density of subsequent pulse trains within a controllable range. The processed open-polarity event sequence and closed-polarity event sequence constitute an open-polarity pulse train and a closed-polarity pulse train, respectively. Each pulse corresponds to a reserved event, and the pulse firing time is equal to the timestamp of the reserved event.

[0023] For each reserved event in the open-polarity pulse train and the closed-polarity pulse train, its propagation direction number is determined. Specifically, the propagation direction number is determined by backtracking one preset neighborhood duration from the current reserved event's emission time, denoted as . The typical value ranges from 3 milliseconds to 15 milliseconds, with 8 milliseconds being the preferred value. The preset neighborhood duration is the time window length for backtracking to find previously retained events. This duration is set to match the side length of the regular hexagon and the expected human movement speed: with a side length of 0.15 meters and an expected gesture speed of 1.5 meters per second, a 10-millisecond neighborhood duration can cover the typical range of gesture speed variations. In the virtual regular hexagonal grid coordinate system, the six primitive positions adjacent to the currently retained event correspond to directions 1 to 6, respectively; the search is conducted within the same polarity event substream at these six primitive positions to find timestamps falling within the interval. All previously retained events are selected, and the one with the largest timestamp is chosen as the most recent previously retained event. This is the timestamp of the currently reserved event, i.e., the time when the currently reserved event was issued. Let the location of the most recent prior reserved event be the position of the currently reserved event. , Taking directions 1 to 6, then it will be related to the orientation. The direction numbers of mutually opposite sides are used as the propagation direction numbers of the currently retained event; for example, if the most recent prior retained event is a neighbor of the current retained event in direction 2, then the propagation direction number of the current retained event is direction 5. This "taking the opposite side" logic reflects the physical process of motion: the prior retained event appears in direction 5. This means that the movement is from the direction The corresponding adjacent interactive primitive unit enters the current interactive primitive unit, therefore the current interactive primitive unit receives the direction. Movement propagating in the opposite direction; through opposite encoding, subsequent stages can complete the alignment and matching of "the direction of arrival of open polarity" and "the direction of departure of closed polarity" in the same directional channel.

[0024] If, within the preset neighborhood timeframe, the number of previously retained events found in the same-polarity event substreams of the six neighboring interactive primitives is zero, it indicates that the currently retained event has no reliable directional reference. In this case, the currently retained event is marked as a non-directional event and sent to the preset non-directional background channel, thus removing it from subsequent propagation direction-related processing. However, it can still be used for event density statistics in the spatial partitioning of this exhibit or for background rate monitoring. This process avoids the direction estimation noise caused by forcibly assigning direction numbers to isolated events, thereby suppressing direction estimation noise. The capacity of the preset non-directional background channel is configured to handle at least 10,000 events per second to accommodate low-speed movements such as visitors pausing or moving slowly in the exhibition hall.

[0025] Finally, the source address of a pulse with a propagation direction number is composed of three parts: the number of its corresponding interactive primitive, the polarity identifier, and the propagation direction number. The number of the interactive primitive is the aforementioned axial coordinate. The integer pairs can be further compressed into a single integer index for easier pulse routing in subsequent stages; the polarity identifier occupies 1 bit, with a value of 0 corresponding to open polarity and a value of 1 corresponding to closed polarity; the propagation direction number occupies 3 bits, with values ​​1 to 6 corresponding to directions 1 to 6 respectively. The source address formed by combining these three parts ensures that each pulse carries complete spatial-polarity-direction information ("where, what polarity, and which direction it came from") when leaving this step, providing the necessary prerequisite for subsequent directional coherent detection. The pulse's emission time is equal to the timestamp of its corresponding reserved event, maintaining the time precision of asynchronous events.

[0026] refer to Figure 1 , Figure 1 The central area features a local grid of 19 closely spaced regular hexagons, laid out in a flat-topped pattern. Each hexagon represents an interaction primitive. The central hexagon is distinguished from the outer hexagons by its filled shape, indicating the interaction primitive currently determining the propagation direction number. Six short arrows radiate outward from the geometric center of the central hexagon, pointing to its six adjacent interaction primitives, corresponding to directions 1 through 6. These six direction labels are placed on a larger-radius circle around the outer edge of the grid to avoid positional conflicts with event markers in the central area. Directions 1 and 4 are opposite each other horizontally, and directions 2 and 5, and 3 and 6 are also paired, reflecting the geometric property of "opposite directions are paired." A circular marker is drawn at the geometric center of the central hexagon, representing the location of the currently preserved event, and its timestamp is recorded as [time stamp value missing]. The other circular marker located at the center of the hexagon in direction 2 represents the location where the previously reserved event occurred, and its timestamp is recorded as follows. ,in This indicates the time of issuance of the prior polarity-preserving event. Indicates the time at which the currently reserved event was issued. and The propagation direction arrow does not exceed the preset neighborhood duration. Starting from the currently retained event, draw a significantly thicker and longer arrow pointing towards the neighbor in direction 5. The geometric meaning of this arrow is: the previously retained event is located adjacent to the current retained event in direction 2, and the motion enters the current interaction primitive unit from direction 2. Therefore, the motion received by the current interaction primitive unit is propagated in the opposite direction of direction 2 (i.e., direction 5). Hence, the propagation direction number of the current retained event is direction 5. Figure 1 The core rule is that the direction of propagation number is marked in text form at the top, and the direction is equal to the opposite side of the azimuth.

[0027] As an optional implementation, the triangular grid can be replaced with a rectangular grid, IEEE 1588 time synchronization can be replaced with a time-sensitive network scheme, the mapping target plane can also be defined as the glass surface of a display case, an interactive wall, or a ground projection surface, the regular hexagonal grid can also adopt a pointed shape, the refractory period retention processing can be replaced with event counting threshold suppression based on a sliding time window, and the preset neighborhood duration can be adaptively adjusted according to the event background rate. Those skilled in the art can flexibly select the above alternatives according to specific scenarios.

[0028] When the interactive display control process of the digital multimedia exhibition hall enters step 2, the Loihi 2 neuromorphic chip receives the open-polarity and closed-polarity pulse trains from step 1 as input, and completes the entire process of conversion from asynchronous pulse stream to structured interactive intent within the chip. The Loihi 2 neuromorphic chip is Intel's second-generation research-oriented neuromorphic processor, integrating approximately 128 neural cores per chip. Each neural core contains approximately 8192 spiking neurons that can be flexibly configured by microinstructions. The chip distributes pulses internally via an asynchronous message bus and supports graded spike message format with a payload field of up to 32 bits. In this embodiment, the algorithm time step of the Loihi 2 neuromorphic chip is set to 100 microseconds, allowing approximately 5 time steps to be processed within a refractory period (0.5 milliseconds), preserving the temporal resolution of events while providing sufficient margin for the chip's timing logic. If the exhibition hall geometry is larger and the primitive neural columns are denser, the algorithm time step can be widened to 200 to 500 microseconds to acquire more neuron resources for expanding channels.

[0029] The digital multimedia exhibition hall is divided into regular hexagons with sides of 0.15 meters. A typical exhibition hall has a floor area of ​​100 to 400 square meters, corresponding to approximately 5100 to 20400 interactive primitive units. Each interactive primitive unit corresponds to one primitive neural column on the Loihi 2 neuromorphic chip. One primitive neural column occupies approximately 100 to 120 neurons. Therefore, a single Loihi 2 neuromorphic chip (approximately 1 million neurons) can support all the primitive neural columns of a medium-sized exhibition hall. For larger exhibition halls or with more channels, the Kapoho Point development board (8 Loihi 2 neuromorphic chips cascaded) or the Hala Point cascade system can be used to synthesize multiple Loihi 2 neuromorphic chips into a unified primitive neural column array through pulse routing between chips. The primitive neural column array uses a two-level indexing method: the first level is the axial coordinate of the interactive primitive unit in the virtual regular hexagonal grid coordinate system. , Let x be the axial and lateral coordinates of this interactive primitive unit. The first level represents the axial and longitudinal coordinates of the interactive primitive unit; the second level represents the Loihi 2 neuromorphic chip number, neural core number, and neuron origin address where the primitive neural column is located. Using a static lookup table array, it can be determined from... Directly reading hardware addresses avoids runtime coordinate resolution, which is especially crucial for pulse streams that run at frequencies of hundreds of thousands of events per second.

[0030] Each primordial neural column internally consists of six directional coherent detection channels and one column-level output neuron. The six directional coherent detection channels are respectively bound to directions 1, 2, 3, 4, 5, and 6. Each channel operates independently, only converging at its terminal via the same column-level output neuron. Each directional coherent detection channel comprises three components: one open-polarity delay branch, one closed-polarity delay branch, and one coherent detection neuron. The open-polarity delay branch and the closed-polarity delay branch are arranged in parallel. The former carries pulses from the open-polarity pulse train whose propagation direction is labeled with the channel's direction tag, while the latter carries pulses from the closed-polarity pulse train whose propagation direction is opposite to the channel's direction tag.

[0031] The delay branch uses a series of relay neurons to construct a delay chain. The chain length can be adjusted to flexibly provide delays ranging from one to hundreds of time steps, offering greater flexibility than the 63 algorithm time steps limit provided by the 6-bit synaptic delay field of the Loihi 2 neuromorphic chip. Each relay neuron is configured with the following microinstructions: a threshold of one synaptic current unit and a leakage conductance of 0. This means that as soon as an input pulse arrives, the relay neuron fires a pulse in the next algorithm time step and returns to zero potential. This "complete pulse forwarding" relay neuron is functionally equivalent to a delay register with a fixed propagation delay equal to one algorithm time step, or 100 microseconds. One delay branch is connected in series. After one relay neuron, the total propagation delay is Each algorithm time step This refers to the number of relay neurons in the delay branch. The number of relay neurons in the open-polarity delay branch and the closed-polarity delay branch is configured according to a uniform level rule, with typical levels being 4, 8, 16, and 32 to cover common gesture speed ranges; within the coherent detection channel in the same direction, the number of relay neurons in the two delay branches... They can be the same or differ by one level; the latter is used to detect directional polarity time-series flow.

[0032] Both the open-polarity pulse train and the closed-polarity pulse train output in step 1 carry a source address composed of the interaction primitive unit number, polarity identifier, and propagation direction number. When an open-polarity pulse arrives at the input interface of the Loihi 2 neuromorphic chip, the chip's pulse routing table processes it according to the source address as follows: first, it locates the corresponding primitive neural column according to the interaction primitive unit number; then, it selects the coherent detection channel with the same number in the primitive neural column according to the propagation direction number; finally, it sends the open-polarity pulse to the input terminal of the open-polarity delay branch of the coherent detection channel in that direction. The processing of the closed-polarity pulse is slightly different, with the core difference being the direction reversal: one propagation direction number is the direction... The polarity pulses are sent into the same primordial nerve column, numbered according to direction. The polarity delay branch input of the directional coherent detection channel in the opposite direction. Taken from directions 1 to 6, the opposite directions are mapped as follows: direction 1 corresponds to direction 4, direction 2 corresponds to direction 5, and direction 3 corresponds to direction 6. The engineering significance of this asymmetric input rule of "open polarity follows the original direction, closed polarity follows the opposite direction" lies in the following: Hand gestures have significant reciprocating motion characteristics. The forward stroke causes a group of interactive primitive units to successively generate open polarity pulses and form an open polarity flow advancing along the main axis. The return stroke causes the same group of interactive primitive units to generate closed polarity pulses and form a closed polarity flow advancing in the opposite direction along the main axis. By connecting the coherent detection channel of the same direction to the open polarity pulses of the same direction and the closed polarity pulses of the opposite direction, this channel provides a method for detecting "open polarity along the direction..." Promoting the polarity of "and" along the direction The system is sensitive to two temporally and spatially opposite pulse flows simultaneously, thus providing highly selective recognition of gestures with directional reversal characteristics, while naturally suppressing pseudo signals with pure unidirectional flow, such as unidirectional continuous motion, specular reflection highlights, and lamp flicker.

[0033] The coherent detection neuron is located at the end of the directional coherent detection channel, receiving two inputs: one from the end of the open polarity delay branch and the other from the end of the closed polarity delay branch. The Loihi 2 neuromorphic chip provides multiple independent input ports for each neuron. In this embodiment, the end of the open polarity delay branch is connected to the first synaptic port of the coherent detection neuron, and the end of the closed polarity delay branch is connected to the second synaptic port. These two ports belong to two different synaptic groups in the Loihi 2 neuromorphic chip's microinstruction configuration and can be identified separately by the chip hardware. If a version of the Loihi 2 neuromorphic chip that supports graded spike payload field pass-through is selected, it can also be distinguished by the preset polarity bit in the payload field. This method can reduce routing table entries between chips in multi-chip cascade scenarios.

[0034] The coherent detection neuron adopts the following behavioral model: it maintains two internal Boolean flags, an open polarity reception flag and an off polarity reception flag, both initially set to 0, and two corresponding window timers. In each algorithm time step, the coherent detection neuron performs the following processing: when the first synaptic port receives at least one pulse in this time step, the open polarity reception flag is set to 1 and the open polarity window timer is started, incrementing from 0; when the second synaptic port receives at least one pulse in this time step, the off polarity reception flag is set to 1 and the off polarity window timer is started; when both reception flags are 1 within the same time step and the corresponding window timer readings do not exceed the preset coherent window length, the coherent detection neuron sends a direction label to the intent decoding layer. coherent pulses, This serves as the direction label for the coherent detection channel in this direction. In the next algorithm time step after transmission, both receive flags are simultaneously cleared, and both window timers are simultaneously reset. This ensures that each coherent pairing's open-polarity pulse and closed-polarity pulse are independently paired; that is, the same open-polarity pulse will not repeatedly pair with subsequent closed-polarity pulses, and vice versa. If any receive flag is set to 1 and its window timer exceeds the preset coherent window length without successful pairing, that flag and its corresponding timer are individually cleared to avoid isolated flags that have been suspended for a long time affecting subsequent pairings. The typical value for the preset coherent window length is 2 to 10 algorithm time steps, i.e., 0.2 milliseconds to 1 millisecond, which coincides with the time difference range between the arrival of the positive and negative leading and trailing edges during the movement edge's traversal of an interactive primitive unit.

[0035] refer to Figure 2 The horizontal axis represents time in milliseconds, and the vertical axis shows several signal channels arranged from top to bottom. The baseline of each channel is a horizontal straight line, and signal pulses are superimposed on the baseline as short vertical lines. The topmost channel is the pulse sequence of an open-polarity pulse train at the input of an open-polarity delay branch. Four short vertical lines represent four open-polarity events entering the branch sequentially. The channel immediately below it is the pulse sequence at the end of the open-polarity delay branch, corresponding one-to-one with the pulses of the topmost channel, but with the overall timestamp shifted by one fixed unit of propagation delay. This fixed unit of propagation delay is equal to the accumulated delay of several relay neurons, representing the propagation process of the pulse in the delay branch. The next channel is the pulse sequence of an off-polarity pulse train at the input of an off-polarity delay branch, and the channel below that is the pulse sequence at the end of the off-polarity delay branch. The delay shift of the off-polarity pulse is independent of the open-polarity pulse, but both are configured using a unified timing rule to ensure that the end pulses are nearly aligned in time. Below the baseline, an additional pre-defined coherence window channel is established, marked by rectangular strips indicating the window activation range of coherent detection neurons in each pairing. The horizontal width of the rectangular strips corresponds to the pre-defined coherence window length, typically as follows: ,in The preset coherence window length is typically between 0.2 ms and 1 ms. Below that is the coherence pulse channel. Each time the pulses at the tails of the open-polarity delay branch and the closed-polarity delay branch fall within the same preset coherence window, the coherence detection neuron fires a coherence pulse. This coherence pulse is represented by a short vertical line on the channel baseline, and its timestamp is slightly later than the pairing completion time. At the bottom is the output channel of the column-level output neuron. This channel is superimposed with a rectangular strip representing the effective range of the preset integration window. The horizontal width of the rectangular strip corresponds to the preset integration window length, denoted as . ,in This indicates the preset integration window length, typically ranging from 5 to 20 milliseconds. At the end of the preset integration window, the column-level output neuron sends a column-level output pulse to the intent decoding layer, displayed as a short vertical line. Figure 2 This multi-channel, vertically parallel layout visually depicts the pulse flow process of "input pulse → delay shift → coherent window alignment → coherent pulse emission → integration window accumulation → column-level output pulse". The column-level output neuron is located at the top of the basic neural column and receives all coherent pulses from the six directional coherent detection channels. The column-level output neuron maintains an independent cumulative counter for each direction label, for a total of six, and also maintains a global integration window timer. At each algorithm time step, the column-level output neuron performs the following processing: upon receiving a pulse carrying a direction label... When the coherent pulse occurs, the direction label The cumulative counters are incremented by 1; the direction label that first reaches the preset trigger count among the 6 cumulative counters is used as the winning direction label for this preset integration window; when the global integration window timer reaches the preset integration window length, this preset integration window is closed, and a column-level output pulse carrying the winning direction label and the primitive neural column number is sent to the direction channel corresponding to the winning direction label in the intent decoding layer. Then, all 6 cumulative counters are simultaneously cleared, the global integration window timer is reset, and this primitive neural column enters the next preset integration window. The typical value of the preset trigger count is 3 to 8; the typical value of the preset integration window length is 50 to 200 algorithm time steps, that is, 5 milliseconds to 20 milliseconds, which matches the resolvable time resolution of human motion.

[0036] The intent decoding layer receives column-level output pulses from all primitive neural columns and performs two tasks in parallel: identifying the global gesture category and locating the spatial partition where the gesture is located. The results of these two tasks are bound to the interaction intent at the end of the decision window. Figure 2 Tuples. The intent decoding layer includes a set of gesture category neurons and a set of exhibit partition neurons. Each of the two sets occupies several neural cores on the Loihi 2 neuromorphic chip. There is no direct pulse coupling between the two sets. Each set has one inhibitory interneuron that maintains the winner-take-all relationship within the set.

[0037] The gesture category neuron set is subdivided into three category branches: waving category branch, pushing / pressing category branch, and circling category branch. Each category branch consists of a pulse sequence detector composed of sequentially cascaded multi-level state detection neurons, typically with four cascade levels. Each level of state detection neuron is equipped with one time-gated neuron; in addition, each category branch has one reset inhibition neuron and one termination trigger neuron at the final level. Each level of state detection neuron is assigned to listen to a preset combination of direction labels and primitive neural column numbers. The combination content is pre-embedded in the pulse routing table of the Loihi 2 neuromorphic chip by the gesture pattern to be recognized by the category branch.

[0038] The inter-level listening direction labels of the waving category branch are arranged alternately as direction 1, direction 4, direction 1, and direction 4. That is, the first level listens for the column-level output pulse of direction 1, the second level listens for the column-level output pulse of direction 4, and so on. Each level of state detection neuron listens for the set of primitive column numbers within the same local primitive group. The typical range of this local primitive group is a regular hexagonal neighborhood with a radius of 3 to 5 interactive primitive units, centered at the waving center point, matching the physical range of the lateral coverage of an adult viewer's palm when waving. The push-press category branch establishes a local direction mapping table for each exhibit space partition. This local direction mapping table is a lookup array of length 6, with indices from direction 1 to direction 6. Only one index position is mapped to the push-press listening direction label, which is the global direction label corresponding to the visual axis direction of the exhibit in that exhibit space partition, denoted as direction. , Taken from directions 1 to 6; the inter-level listening direction labels of the push-by category branches within the exhibition space partition are all set as directions. Each level of state detection neuron listens to the sequence of adjacent primitive neural column numbers arranged from far to near along the visual axis of the exhibit. This local normalization ensures that any pushing or pressing action towards any exhibit can be recognized in the local coordinates of that exhibit, regardless of the exhibit's actual orientation in the exhibition hall. The inter-level listening direction labels of the selected category branch are arranged in a cyclical pattern of direction 1, direction 2, direction 3, direction 4, direction 5, and direction 6, for a total of 6 levels or multiples of 6, to match the 6-step cycle of direction numbers along the six neighborhoods when the audience makes a circular motion. Each level of state detection neuron also listens to the set of primitive neural column numbers within the same local primitive group. The radius of the local primitive group is 2 to 4 interactive primitive units, which corresponds to the circumference radius of a typical adult circular gesture.

[0039] The cascading progression rules within the category branch are the same. Taking the waving category branch as an example: Initially, all state detection neurons in the entire waving category branch are in the initial state, and all time-gated neurons are in the closed state. When the first-level state detection neuron receives a column-level output pulse that matches its assigned direction label and primitive neuron column number combination, the first-level state detection neuron fires a state progression pulse. This pulse sets the time-gated neuron assigned to the second-level state detection neuron to the gated open state. The effective duration of the gated open state is the preset inter-level interval duration, typically ranging from 80 to 2000. The algorithm has 00 time steps, ranging from 8 to 20 milliseconds, which corresponds to the typical time interval between two opposite extreme values ​​within one hand gesture. At the same time, the state propulsion pulse resets the first level to its initial state to avoid repeated responses to the same column-level output pulse. In the gated open state, when the second-level state detection neuron receives a matching column-level output pulse, it also fires a state propulsion pulse and drives the third-level time-gated neuron into the gated open state in the same way. This continues until the final-level state detection neuron fires, triggering the termination trigger neuron of this type of branch to fire a gesture recognition pulse. The resetting inhibition neuron works as follows: All state detection neurons in the current branch synchronously report their firing to the resetting inhibition neuron. Simultaneously, the resetting inhibition neuron monitors the discrepancy pulse counter for the current branch. This counter accumulates the number of column-level output pulses that occur within the gating open state of the current activation level, are within the monitoring range, but do not match the allocation combination of this level. When the discrepancy pulse counter reaches the preset interference count (typically 3 to 10), or when the gating open state timer of any level 1 time-gated neuron reaches the preset inter-level interval duration but the number of matching column-level output pulses is 0, the resetting inhibition neuron fires reset pulses to all state detection neurons and all time-gated neurons in the current branch, returning them to their initial state. The preset interference count makes the current branch insensitive to scattered background noise, while also allowing it to promptly exit the current sequence recognition when clear irrelevant motion occurs.

[0040] Each exhibit partition neuron in the exhibit partition neuron set is bound to one exhibit space partition in the exhibition hall. The exhibit space partitions are pre-divided by the exhibition hall floor plan, and each area corresponds to a specific multimedia exhibit. The binding relationship is written through the pulse routing table of the Loihi 2 neuromorphic chip, meaning that each interactive primitive unit belongs to a unique exhibit space partition. The exhibit partition neuron only receives column-level output pulses emitted by all primitive neural columns covering its own exhibit space partition, and is completely isolated from column-level output pulses from other exhibit space partitions. The exhibit partition neuron adopts a leakage integral firing model, setting one internal counter and one leakage time constant. For each column-level output pulse received, the internal counter increments by 1. After each algorithm time step, the internal counter decays at a preset leakage rate, typically by 1 unit every 10 algorithm time steps, making the internal counter sensitive to short-term dense events while maintaining a low value for sparse background events. When the internal counter reaches the preset trigger count within the preset counting window, the exhibit partition neuron fires one partition activation pulse. The typical value for the preset counting window length is 100 to 500 algorithm time steps, or 10 to 50 milliseconds; the preset trigger count is proportional to the coverage area of ​​the exhibit space partition, with 5 to 10 for small exhibits and 15 to 30 for large exhibits.

[0041] To ensure that objects in only one set of gesture category neurons and one set of presentation partition neurons can be recognized simultaneously, the intent decoding layer introduces two inhibitory interneurons: a gesture inhibition interneuron and a partition inhibition interneuron. The gesture inhibition interneuron receives gesture recognition pulses from all termination trigger neurons in the gesture category neuron set. Upon receiving the first gesture recognition pulse, it immediately sends a reset pulse to the termination trigger neurons, state detection neurons, and time-gated neurons of all other category branches except the one containing the source signal, thus locking the winning category branch and clearing all intermediate states of other branches. The partition inhibition interneuron behaves similarly: it receives partition activation pulses from all presentation partition neurons in the presentation partition neuron set. Upon receiving the first partition activation pulse, it sends a reset pulse to all presentation partition neurons except the one containing the source signal, resetting their count to zero. The engineering value of introducing inhibitory interneurons lies in: reducing the need for... The winner-takes-all structure connected on both sides is compressed to... strip fan entry and The fan-out design significantly saves synaptic resources on the Loihi 2 neuromorphic chip in scenarios where the intent decoding layer involves tens of thousands of exhibit partition neurons.

[0042] The intent decoding layer organizes its output with a preset decision window length as the time granularity. A typical preset decision window length is 500 to 2000 algorithm time steps, or 50 to 200 milliseconds, matching the perceptible threshold of human interaction response latency, ensuring the audience perceives a sufficiently immediate response. Within each decision window, the gesture category corresponding to the gesture recognition pulse retained by the gesture inhibition interneuron is bound to the exhibit space partition corresponding to the partition activation pulse retained by the partition inhibition interneuron. The binding result is the interaction intent for that decision window. Figure 2 The tuple, in binary data structure of "<gesture category, exhibit space partition>", is sent out through the output interface of the Loihi 2 neuromorphic chip. If the gesture inhibition interneuron within a decision window does not retain any gesture recognition impulse, or if the partition inhibition interneuron does not retain any partition activation impulse, this decision window will not output the interaction intention. Figure 2 At the end of the decision window, the intent decoding layer resets the internal counts and gating states of all gesture category neurons and item partition neurons once, so that the next decision window starts computing from a clean initial state.

[0043] As an optional implementation, the six coherent detection neurons can also be merged into one multi-port neuron with six synaptic port pairs; the number of relay neurons can be automatically adjusted by the on-site calibration process; the number of cascaded layers of the gesture category branch can be expanded from 4 levels to 8 levels or more; the partitioning of the display space can be changed to dynamic adjustment; parameters such as preset coherence window length, preset integration window length, preset inter-level interval duration, and preset decision window length can be adjusted according to the scenario in the runtime configuration interface of the Loihi 2 neuromorphic chip. The Loihi 2 neuromorphic chip can be used with a single-chip research and development board, a Kapoho Point 8 chip board, or a Hala Point cascade system.

[0044] After the interactive display control process of the digital multimedia exhibition hall enters step 3, the Loihi 2 neuromorphic chip outputs interactive intentions to the closed-loop controller in cycles based on the decision window. Figure 2 The closed-loop controller, as a unit, performs two tasks: converting the pulse-level semantics output from the neuromorphic chip into industrial control commands recognizable by the actuators in the multimedia display, and feeding the execution results back to the front-end event acquisition stage. The closed-loop controller uses an industrial control computer with hard real-time scheduling capabilities, running a Linux kernel with a PREEMPT_RT patch or an equivalent real-time VxWorks operating system, with system scheduling jitter controlled within 100 microseconds. The closed-loop controller is coupled to the Loihi 2 neuromorphic chip via a PCIe interface or a 10 Gigabit Ethernet interface, and shares the IEEE 1588 precision time protocol time base.

[0045] Interactive meaning Figure 2When the tuple is output from the Loihi 2 neuromorphic chip, it adopts a binary data structure in the form of "<gesture category, exhibit space partition>". The gesture category field occupies 3 bits, with a value of 1 corresponding to a wave, a value of 2 corresponding to a push, and a value of 3 corresponding to a selection; the exhibit space partition field occupies 11 bits, which can address up to 2048 independent exhibit space partitions in the exhibition hall, matching the scale of exhibits that this embodiment can support. The closed-loop controller receives the interaction intent. Figure 2 Following the tuple, the first semantic mapping is completed based on the pre-edited exhibit behavior table: the exhibit behavior table is a lookup array with "<gesture category, exhibit space partition>" as the composite primary key and "<target multimedia display execution mechanism address, target action type, action parameter set, expected execution duration>" as values. The target multimedia display execution mechanism address includes the fieldbus node number of the projection tracking unit, the node number of the directional sound field unit, and the node number of the display case dimming unit. The action parameter set varies depending on the action type. For example, the action parameters of the projection tracking unit include the target pitch angle, target azimuth angle, target focal length, and target brightness level; the action parameters of the directional sound field unit include the target beam azimuth angle, target beam pitch angle, target sound pressure level, and target audio segment number; and the action parameters of the display case dimming unit include the target light transmittance level number and target color temperature level number. The exhibit behavior table allows for updating only the lookup data when adding a new exhibit, while different exhibits can retain their own independent interaction logic.

[0046] After completing the table lookup, the closed-loop controller enters the control command message encapsulation stage. The control command message uses a unified binary frame format, consisting of the following from beginning to end: a 2-byte synchronization header, a 1-byte protocol version number, a 4-byte message sequence number, an 8-byte timestamp, a 2-byte source node address, a 2-byte destination node address, a 1-byte action type, a 2-byte action parameter length, a variable-length action parameter field, and a 2-byte cyclic redundancy check (CRC) code. The message sequence number is incremented and maintained internally by the closed-loop controller to ensure accurate matching of the original command during subsequent status feedback; the timestamp is obtained from the interaction received by the closed-loop controller. Figure 2 The timing of the tuple is accurate to 1 microsecond; the cyclic redundancy check (CRC-16-CCITT) polynomial is used. ,in The formal variable of this polynomial generator is used to detect occasional bit errors during message transmission on the fieldbus. This control command message is sent via the fieldbus interface to the multimedia display actuator corresponding to the exhibition space partition.

[0047] The preferred fieldbus is EtherCAT, which boasts sub-microsecond synchronization accuracy and can complete data exchange between all slave nodes within 1 millisecond. The EtherCAT master station employs a distributed clock mechanism to align the local time base of all slave nodes to the same IEEE 1588 time base as the Loihi 2 neuromorphic chip, ensuring comparability between command issuance times and event timestamps. In smaller-scale scenarios, PROFINET IRT, POWERLINK, or Time-Sensitive Networking (TSN) Ethernet can be used as alternatives to EtherCAT.

[0048] The multimedia display execution mechanism analyzes motion parameters according to the action type and executes physical actions. The projection tracking unit consists of a laser projection mechanism with a dual-axis servo turntable. The servo driver completes the attitude adjustment according to the S-curve trajectory. Typical small angle adjustment takes about 300 milliseconds, and large angle adjustment takes about 800 milliseconds to 1.2 seconds. Image switching is achieved through a double buffering mechanism, with a typical switching time of about 50 milliseconds. The directional sound field unit consists of phased array speakers, which form an acoustic beam pointing to a specific spatial area through a phase delay array, with a response time of less than 10 milliseconds. The display case dimming unit consists of an electrochromic glass panel and adjustable color temperature LED lamps. The switching time from transparent to frosted state of the electrochromic glass is 3 to 8 seconds, and the return time from frosted state to transparent state is 5 to 15 seconds. The color temperature and brightness adjustment of the LED lamps are achieved through PWM drive, with a response time of less than 100 milliseconds.

[0049] Each multimedia display actuator is embedded with a completion status detection unit, which generates a completion status message the instant the physical action is completed. The determination of completion status varies depending on the type of action: For the projection tracking unit, the completion criteria are that the deviation between the servo turntable encoder reading and the target angle is less than 0.1 degrees, the deviation between the brightness feedback value and the target brightness level is less than one level, and the screen switching buffer is completed simultaneously; for the directional sound field unit, the completion criteria are that the deviation between the actual phase of the phase array and the target phase is all less than 2 degrees, and the audio clip playback head reaches the expected position; for the display case dimming unit, the completion criteria are that the deviation between the actual light transmittance of the electrochromic glass and the target light transmittance level is less than 5%, and the deviation between the LED color temperature feedback value and the target level is less than 100K. The format of the completion status message is symmetrical to that of the control command message. The newly added fields are the completion status code (1 byte, 0 indicates successful completion, and a non-zero value indicates an error), the actual execution time (4 bytes, with a precision of 1 microsecond), and the original message sequence number (4 bytes, corresponding to the issued control command message). The completion status message is sent back to the closed-loop controller via the same fieldbus.

[0050] After receiving the completion status message, the closed-loop controller triggers a preset stabilization duration gating suppression process. The preset stabilization duration is used to cover the aftershock phase after the multimedia display actuator completes its action, preventing a large number of non-interactive visual events generated during the action from being misidentified as new gestures by the Loihi 2 neuromorphic chip, triggering unwanted secondary interactions. In this embodiment, different preset stabilization durations are used for the three types of multimedia display actuators: 300 milliseconds for the projection tracking unit, covering the slight echo of the servo turntable and the buffer frame after the image switch; 100 milliseconds for the directional sound field unit, mainly to deal with the micro-vibration reflection of the speaker diaphragm during the phase switching of the phased array; and 2 seconds for the display case dimming unit, matching the chemical layer stabilization time of the electrochromic glass and the thermal equilibrium time of the LED lamps.

[0051] The specific execution method of the preset stable duration gating suppression process is as follows: After receiving the completion status message, the closed-loop controller immediately sends a gating suppression instruction to the event preprocessing path at the front end of the Loihi 2 neuromorphic chip input interface. This instruction contains three fields: gating suppression start time, gating suppression end time, and the set of suppressed exhibit space partition numbers. The gating suppression start time is taken from the timestamp in the completion status message, and the gating suppression end time is taken from that time plus the aforementioned preset stable duration. The event preprocessing path inserts an event filtering logic between the Loihi 2 neuromorphic chip input interface and the asynchronous event stream: During the period from the gating suppression start time to the gating suppression end time, all events falling within the coverage area of ​​the suppressed exhibit space partition are discarded, neither entering the interactive primitive unit ownership process nor entering the open polarity event sub-stream or the closed polarity event sub-stream; after the gating suppression end time arrives, the event filtering logic removes the shielding of the exhibit space partition, and normal event stream processing resumes. It should be noted that gating suppression only takes effect within the designated exhibit space partition and does not affect concurrent interaction in other exhibit space partitions. Therefore, when multiple visitors interact simultaneously in front of different exhibits, the gating suppression of each exhibit space partition is independent of each other and will not delay each other.

[0052] Limiting the gating suppression to specific exhibit space zones, rather than the entire exhibition hall, is a decision based on clear engineering considerations: visual disturbances caused by the physical movements of multimedia display actuators are mainly concentrated near the exhibit. Confining gating suppression to the affected exhibit space zones effectively masks false events while preserving the interactivity of other areas far from the exhibit. Geometrically, the area covered by the suppressed exhibit space zones is determined by the pre-marked polygonal boundaries on the exhibition hall floor plan. Events are then assigned through the attribution process in step 1. Then, the corresponding exhibit space partition number is read directly through a static lookup table array and matched with the set of suppressed exhibit space partition numbers in the gating suppression instruction. If a match is found, the exhibit is discarded; if no match is found, the exhibit is allowed to pass. The entire decision has constant complexity and will not become a bottleneck in the processing of high-speed event streams.

[0053] After the preset stabilization period ends, the closed-loop controller continues to update the environmental baseline for the next decision window. The environmental baseline is a state table maintained in memory by the closed-loop controller, with the exhibit space partition number as the primary key and "<current projection state, current sound field state, current dimming state, current event background rate>" as the value. The current projection state records the most recent target attitude parameters and display content index of the projection tracking unit involved in the exhibit space partition; the current sound field state records the most recent target beam parameters and audio segment index; the current dimming state records the most recent target light transmittance level and color temperature level; the current event background rate is obtained by dividing the event count within the preset observation period (typically 100 milliseconds) by the preset observation period, reflecting the steady-state event density of the exhibit space partition in the new multimedia display state. The specific steps for updating the environmental baseline are as follows: After the preset stabilization period ends, the closed-loop controller counts the asynchronous event stream of the exhibit space partition in the next 100 milliseconds, and divides the count by 0.1 seconds to obtain the new current event background rate; at the same time, the most recent action parameters of the multimedia display actuators involved in the exhibit space partition are written from the exhibit behavior table to the current projection state, current sound field state, and current dimming state; finally, the closed-loop controller sends the updated environmental baseline to the runtime configuration interface of the Loihi 2 neuromorphic chip, and the Loihi 2 neuromorphic chip adjusts several internal parameters accordingly.

[0054] The environmental baseline influences the internal parameters of the Loihi 2 neuromorphic chip as follows: A leaky integral firing type event normalization neuron layer is pre-implemented within the Loihi 2 neuromorphic chip. This layer normalizes events at the entry point of each interactive primitive unit to subtract the impact of the background rate on the firing probability of subsequent coherent detection neurons. The normalization threshold is obtained by linearly mapping the current event background rate; that is, assuming the current event background rate is... The normalization threshold for the event normalization neuron layer is: The following mapping is used: ;in The current event background rate for this exhibit space partition, as obtained in the latest environmental baseline update, is expressed in events per second. The normalized threshold reference value is pre-calibrated for the closed-loop controller, typically taking the value of 1 synaptic current unit; The linear mapping coefficient from the background rate to the threshold is pre-calibrated for the closed-loop controller. Typically, the normalized threshold increases by 0.1 synaptic current units for every 10,000 additional events per second in the background rate. This refers to the normalization threshold of the event-normalized neuron layer within the Loihi 2 neuromorphic chip for this exhibit space partition. This mapping method allows the Loihi 2 neuromorphic chip to automatically raise the recognition threshold when the display content update causes a new steady-state event density after the multimedia display actuator's action, avoiding the generation of too many false gesture recognition pulses under the new high background rate. Conversely, when the exhibit space partition enters a static display state with low event density, the Loihi 2 neuromorphic chip will lower the normalization threshold to restore the recognition sensitivity for slow, slight gestures. The three display states—current projection state, current sound field state, and current dimming state—are not directly converted into neuron parameters of the Loihi 2 neuromorphic chip. Instead, they are written into the "context key" field of the exhibit behavior table, enabling the next decision window to select the appropriate downstream action based on the current display state when generating a new instruction from the table. For example, when the projected image is already on the "exhibit introduction page," a new "waving" gesture is mapped to "turn to the next page" instead of "open the introduction page." This context-based state machine-style behavior mapping upgrades exhibit interaction from a stateless "stimulus-response" model to a multi-step interaction with contextual memory.

[0055] The dynamic visual sensing event camera array, the Loihi 2 neuromorphic chip, and the multimedia display actuator form a complete closed loop through a closed-loop controller. The end-to-end latency of the entire closed loop, from the audience making a gesture to the multimedia display actuator starting to execute the response, is typically no more than 200 milliseconds; the total latency from the completion of the action to the next round of interaction being correctly recognized is typically no more than 500 milliseconds (except for scenarios involving electrochromic glass).

[0056] As an optional implementation, the exhibit behavior table can be expanded into an online editable strategy library; the control command message format can be replaced with OPC UA over TSN; the preset stability duration can be automatically learned by the closed-loop controller using a moving average method; the gating suppression space range can be expanded into the union of multiple adjacent exhibit space partitions under large-scale actions such as multi-projection linkage; the event background rate statistics method can also be changed to an exponentially weighted moving average; and the completion status judgment threshold can be switched from the factory default value to the field calibration value. The multimedia display actuator, fieldbus type, and real-time operating system can all be replaced with other industrial components of equivalent capability.

[0057] refer to Figure 3 The horizontal axis represents time, in milliseconds, with scales ranging from 0 to 1000 at equal intervals; the vertical axis represents asynchronous event flow density, in events per second, with scales ranging from 0 to 6000 at equal intervals. Figure 3The main curve represents the trajectory of the asynchronous event flow density change in a specific exhibit space within a complete interaction cycle. Between 0 and 100 milliseconds, the main curve maintains a low baseline of approximately 50 events per second, with a small fluctuation of about 400 events per second around 60 to 100 milliseconds. This fluctuation corresponds to events triggered by audience gestures; this interval is the first decision window, marked by a rectangle below the curve. The vertical dashed line at 100 milliseconds indicates the moment the closed-loop controller sends a control command message via the fieldbus, marking the start of physical actions by the multimedia display actuators. Between 100 and 400 milliseconds, the main curve rises sharply, forming a peak of approximately 4500 events per second and a secondary peak of approximately 1500 events per second, representing non-interactive visual events such as projection screen switching, dimming changes, electrochromic changes, and mechanical movements of the actuators. The vertical dashed line at 400 milliseconds indicates the moment the multimedia display actuators return their completed state to the closed-loop controller via the fieldbus. The 400-700 millisecond interval is covered by a filled rectangular strip across the entire horizontal axis, representing the gating suppression window applied to the exhibit space partition within a preset stable duration, denoted as _____. ,in This represents the preset stabilization time, typically between 100 milliseconds and 2 seconds. Events falling within this range are directly discarded by the event preprocessing path. The main curve exhibits exponential decay within this range, reflecting the decrease in event density during the aftershock phase. The vertical dashed line at 700 milliseconds indicates the moment when the closed-loop controller completes the environmental baseline update for the next decision window. From this point, the main curve transitions to a new steady-state background level of approximately 95 events per second. This new steady-state background level is slightly higher than the initial baseline 100 milliseconds prior, reflecting the change in steady-state event density caused by changes in the multimedia display state. The 700 to 1000 millisecond range is marked as the next decision window. Figure 3 The measured evolution of asynchronous event flow density within the closed-loop cycle is superimposed on the same time axis along with the three key moments of control command issuance, completion feedback, and environmental baseline update, as well as the preset stable duration gating suppression and decision window boundary. This allows those skilled in the art to clearly grasp the temporal coupling relationship of each link in step 3.

[0058] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for controlling interactive displays in a digital multimedia exhibition hall, characterized in that, Includes the following steps: Step 1: The asynchronous event stream is collected by the dynamic visual sensing event camera array deployed in the exhibition hall. After the events are assigned to the interactive primitive units divided in the virtual hexagonal grid coordinate system, they are split according to the polarity identifier and formed into open polarity pulse trains and closed polarity pulse trains after refractory period retention processing. The propagation direction number of each retained event is determined based on the adjacent orientation of the previously retained events. Step 2: Construct a primitive neural column array on the Loihi 2 neuromorphic chip, corresponding one-to-one with the interactive primitive unit. Each primitive neural column contains 6 directional coherent detection channels. Each channel receives open polarity pulses in the same direction and closed polarity pulses in the opposite direction via open polarity delay branches and closed polarity delay branches, respectively, and sends them to the independent input channel of the coherent detection neuron. When all independent input channels within the preset coherent window receive pulses, they emit coherent pulses carrying directional labels. The column-level output neuron outputs column-level output pulses carrying winning directional labels within the preset integration window. The intent decoding layer decodes the interactive intent binary representing the gesture category and the item space partition. Step 3: The closed-loop controller encapsulates the interaction intent tuple into a control command message and sends it to the multimedia display execution mechanism corresponding to the exhibit space partition via the fieldbus. After the multimedia display execution mechanism completes its action, it sends back the completion status to the closed-loop controller. The closed-loop controller applies gating suppression to the asynchronous event stream within a preset stable duration and updates the environmental reference of the next decision window, so that the dynamic visual sensing event camera array, Loihi 2 neuromorphic chip and multimedia display execution mechanism form a closed loop.

2. The method as described in claim 1, characterized in that, A dynamic visual sensing event camera array is deployed in a triangular grid pattern on the top of the digital multimedia exhibition hall. After time synchronization and external parameter calibration, the event timestamp time base of each dynamic visual sensing event camera is aligned with the unified time base of the exhibition hall. The rectangular pixel coordinates of each dynamic visual sensing event camera are mapped to the unified coordinate system of the exhibition hall through a preset calibration mapping table. A virtual regular hexagonal grid coordinate system is constructed in the unified coordinate system of the exhibition hall and interactive primitive units are divided according to a preset step distance. The interactive primitive unit has 6 adjacent directions, which are denoted as direction 1, direction 2, direction 3, direction 4, direction 5 and direction 6 respectively. Among them, direction 1 and direction 4 are opposite directions, direction 2 and direction 5 are opposite directions, and direction 3 and direction 6 are opposite directions.

3. The method as described in claim 2, characterized in that, Each event in the asynchronous event stream contains pixel coordinates, a timestamp, and a polarity identifier, which can be either on or off polarity. The pixel coordinates of each event are converted to coordinates in the unified coordinate system of the exhibition hall using a preset calibration mapping table. The event is assigned to the interactive primitive unit corresponding to the nearest grid center point after the converted coordinates are used. Then, the event is split into on-polarity event sub-streams and off-polarity event sub-streams according to the polarity identifier. The refractory period retention process is as follows: in the same polarity event sub-stream of the same interactive primitive unit, the earliest event that arrives after a preset refractory period from the timestamp of a retained event is taken as the next retained event, thus forming on-polarity pulse trains and off-polarity pulse trains.

4. The method as described in claim 3, characterized in that, The propagation direction number is determined as follows: Backtracking one preset neighborhood duration from the time the current reserved event is issued, search for the nearest prior reserved event of the same polarity among the six primitive positions adjacent to the current reserved event in the virtual hexagonal grid coordinate system. Record the adjacent orientation of the prior reserved event relative to the current reserved event as orientation j, and use the direction number opposite to orientation j as the propagation direction number of the current reserved event. If the number of prior reserved events of the same polarity found within the preset neighborhood duration is 0, then the current reserved event is marked as a directionless event and sent to the preset directionless background channel, causing the directionless event to exit subsequent propagation direction-related processing. The source address of a pulse with a propagation direction number is composed of the number of its corresponding interactive primitive unit, polarity identifier, and propagation direction number. The pulse's emission time is equal to the timestamp of its corresponding reserved event.

5. The method as described in claim 1, characterized in that, Both the open-polarity delay branch and the closed-polarity delay branch are composed of several relay neurons connected in series. The relay neurons are connected one-to-one by single pulse routing on the Loihi 2 neuromorphic chip. Each pulse generates a fixed unit propagation delay when passing through a relay neuron in the delay branch. The number of relay neurons in the open-polarity delay branch and the closed-polarity delay branch are configured according to a uniform level rule. For the directional coherence detection channel numbered k, k is sequentially selected from 1 to 6. The pulse with propagation direction numbered k in the open-polarity pulse train is sent to the input terminal of the open-polarity delay branch of the directional coherence detection channel numbered k. The pulse with propagation direction numbered k in the closed-polarity pulse train is sent to the input terminal of the closed-polarity delay branch of the directional coherence detection channel numbered k.

6. The method as described in claim 5, characterized in that, The tail end of the open polarity delay branch is connected to the coherent detection neuron via the first synaptic port, and the tail end of the closed polarity delay branch is connected to the second synaptic port. The first and second synaptic ports are configured as independent input ports on the Loihi 2 neuromorphic chip or distinguished by the graded spike payload field, so that the coherent detection neuron can identify open polarity input and closed polarity input respectively. The firing rule of the coherent detection neuron is as follows: when both the first and second synaptic ports receive pulses within a preset coherence window, a coherent pulse carrying a direction label k is fired. After firing, the valid reception flags of the first and second synaptic ports within the current preset coherence window are cleared, so that the open polarity pulse and closed polarity pulse of each coherent pair are independently paired.

7. The method as described in claim 6, characterized in that, The column-level output neuron receives coherent pulses from six coherent detection channels within the elementary nerve column, and within a preset integration window, the direction label that first reaches the preset trigger count is used as the winning direction label. When the preset integration window ends, the column-level output neuron sends a column-level output pulse carrying the winning direction label and the primitive neuron column number to the direction channel corresponding to the winning direction label in the intention decoding layer.

8. The method as described in claim 7, characterized in that, The intent decoding layer includes a set of gesture category neurons, which includes a waving category branch, a push / press category branch, and a selection category branch. Each category branch consists of a series of sequentially cascaded state detection neurons, time-gated neurons assigned to each level of state detection neurons, resetting inhibition neurons, and a termination trigger neuron at the final level. Each level of state detection neuron is assigned to listen to a preset combination of direction labels and primitive neural column numbers. After the previous level of state detection neurons fires, it drives the time-gated neurons assigned to the next level of state detection neurons to enter the gating open state within a preset inter-level interval. In the gating open state, when the next level of state detection neurons receives a column-level output pulse that matches its assigned combination of direction labels and primitive neural column numbers, it fires and passes the gating to the next level. If the preset inter-level interval ends... If the number of matching column-level output pulses is 0, or if a non-matching column-level output pulse reaches the preset interference count within the listening range of this category branch, the reset inhibition neuron sends a reset pulse to all state detection neurons and all time-gated neurons of this category branch to return them to their initial state; when the final-level state detection neuron fires, it triggers the termination trigger neuron of this category branch to fire one gesture recognition pulse; among them, the inter-level listening direction labels of the waving category branch are arranged alternately in direction 1 and direction 4; the push-press category branch maps the visual axis direction of each exhibit space partition to the push-press listening direction label as the inter-level listening direction label through the local direction mapping table, and the state detection neurons at each level listen to the adjacent primitive nerve column number strings arranged from far to near along the visual axis direction of the exhibit; the inter-level listening direction labels of the circle category branch are arranged cyclically from direction 1 to direction 6.

9. The method as described in claim 8, characterized in that, The intent decoding layer also includes a set of presentation partition neurons, gesture inhibition interneurons coupled to the set of gesture category neurons, and partition inhibition interneurons coupled to the set of presentation partition neurons. Each presentation partition neuron is bound to one presentation space partition and receives column-level output pulses from all the basic neural columns covering the presentation space partition. When the number of received pulses reaches a preset trigger count within a preset counting window, a partition activation pulse is emitted. The gesture inhibition interneuron receives gesture recognition pulses from all termination trigger neurons in the gesture category neuron set and emits reset pulses to termination trigger neurons, state detection neurons, and time-gated neurons in other category branches except for the category branch where the emission source is located. The partition inhibition interneuron receives partition activation pulses from all exhibit partition neurons in the exhibit partition neuron set and sends reset pulses to other exhibit partition neurons except the source, causing their counts to return to zero. Within a decision window, the gesture category corresponding to the gesture recognition pulse retained by the gesture inhibition interneuron is bound to the exhibit space partition corresponding to the partition activation pulse retained by the partition inhibition interneuron, thus obtaining the interaction intent tuple for this decision window. After receiving the completion status, the closed-loop controller applies gating inhibition to visual events in the asynchronous event stream that are located within the exhibit space partition coverage area and are caused by the multimedia display execution mechanism within a preset stable duration to filter them out. After the preset stable duration ends, the current multimedia display state of the exhibit space partition is used as the environmental benchmark for the next decision window.