A method and system for interactive image display applied to curtain walls

CN122579402APending Publication Date: 2026-08-14北京建工一建工程建设有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-04-30
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

当前,幕墙照明采用固定光效模式或手动控制,存在互动性不足、智能化程度低的缺点,不能实现与人体行为的精准联动

Benefits of technology

[0009]本申请实施例提供的一种应用于幕墙的影像互动展示方法及系统的有益效果在于:本申请通过摄像装置采集现场视频流并进行人体目标检测,实现基于人体位置的幕墙影像互动,相较于固定播放模式,能够实现人与幕墙之间的实时动态交互,提升幕墙展示的趣味性与沉浸感。通过将人体位置与感兴趣区域匹配判断触发意图,再根据触发-光效映射关系生成泛光照明控制参数,能够准确响应人体行为并驱动照明设备实现动态光效展示,使得幕墙不再是单一显示载体,而是可感知、可互动的智能场景装置,增强了幕墙的展示效果与用户体验。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122579402A_ABST
    Figure CN122579402A_ABST
Patent Text Reader

Abstract

This application provides a method and system for interactive image display applied to curtain walls, belonging to the field of curtain wall lighting control technology. The method includes: acquiring a live video stream through a camera device deployed in the curtain wall area, preprocessing the live video stream and detecting human targets to obtain the positional feature information of the human targets; spatially matching the positional feature information of the human targets with a preset region of interest to determine whether there is a triggering intent and obtaining a trigger determination result; obtaining floodlighting control parameters for controlling the curtain wall lighting equipment based on the trigger determination result and a preset trigger-light effect mapping relationship; generating a control command sequence for each lighting unit of the curtain wall lighting equipment to perform light effect changes based on the floodlighting control parameters; and sending the control command sequence to the controller of the curtain wall lighting equipment to drive the curtain wall lighting equipment to perform dynamic floodlighting display.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of curtain wall lighting control technology, and in particular to an interactive image display method and system applied to curtain walls. Background Technology

[0002] With the advancement of smart city construction, building curtain walls have evolved from simple building envelopes into important carriers for urban nightscape display and human-computer interaction. Curtain wall floodlighting has become a means to enhance building recognition and scene experience. Currently, curtain wall lighting uses fixed lighting effect modes or manual control, which suffers from insufficient interactivity and low intelligence, failing to achieve precise linkage with human behavior. While some existing interactive solutions attempt to detect human presence, they suffer from low positioning accuracy, simple triggering logic, and delayed lighting effect response, and cannot adapt to different ambient lighting and crowd behavior scenarios, leading to false triggers or untimely triggering. Furthermore, the lack of full integration of computer vision and intelligent lighting control technologies prevents dynamic adjustment of lighting effects based on human position and behavioral intentions, thus failing to meet the demands of modern architecture for intelligent, personalized, and interactive displays.

[0003] Therefore, there is an urgent need for an interactive image display method and system for use on curtain walls. Summary of the Invention

[0004] To address the aforementioned technical problems, this application provides a method and system for interactive image display applied to curtain walls.

[0005] A first aspect of this application provides a method for interactive image display applied to a curtain wall, comprising: By deploying camera devices in the curtain wall area to collect on-site video streams, and preprocessing and human target detection of the on-site video streams, the positional feature information of human targets is obtained; The location feature information of the human target is spatially matched with a preset region of interest to determine whether there is a triggering intent, and a triggering determination result is obtained. Based on the trigger determination result and the preset trigger-light effect mapping relationship, the floodlighting control parameters for controlling the curtain wall lighting equipment are obtained; Based on the floodlighting control parameters, a sequence of control instructions is generated for each lighting unit of the curtain wall lighting equipment to perform luminous efficacy changes. The control command sequence is sent to the controller of the curtain wall lighting equipment to drive the device of the curtain wall lighting equipment to perform dynamic floodlighting display.

[0006] A second aspect of this application provides an interactive image display system for use on curtain walls, comprising: The video detection module is used to collect on-site video streams through camera devices deployed in the curtain wall area, and to preprocess and detect human targets in the on-site video streams to obtain the positional feature information of human targets; The intent determination module is used to spatially match the location feature information of the human target with a preset region of interest to determine whether there is a triggering intent and obtain a trigger determination result. The lighting parameter module is used to obtain floodlighting control parameters for controlling the curtain wall lighting equipment based on the trigger determination result and the preset trigger-light effect mapping relationship; The control instruction module is used to generate a sequence of control instructions for each lighting unit of the curtain wall lighting equipment to perform light effect changes based on the floodlighting control parameters. The instruction execution module is used to send the control instruction sequence to the controller of the curtain wall lighting equipment to drive the device of the curtain wall lighting equipment to perform dynamic display of floodlighting.

[0007] A third aspect of this application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor executes the computer program to implement the steps of the above-described interactive image display method applied to a curtain wall.

[0008] A fourth aspect of this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the above-described interactive image display method applied to a curtain wall.

[0009] The beneficial effects of the image interactive display method and system applied to curtain walls provided in this application are as follows: This application uses a camera device to collect on-site video streams and perform human target detection to realize curtain wall image interaction based on human position. Compared with a fixed playback mode, it can realize real-time dynamic interaction between people and the curtain wall, enhancing the fun and immersion of the curtain wall display. By matching the human position with the area of ​​interest to determine the trigger intention, and then generating floodlighting control parameters according to the trigger-light effect mapping relationship, it can accurately respond to human behavior and drive lighting equipment to achieve dynamic light effect display. This makes the curtain wall no longer a single display carrier, but a perceptible and interactive intelligent scene device, enhancing the display effect and user experience of the curtain wall. Attached Figure Description

[0010] Figure 1 A flowchart illustrating an interactive image display method for curtain walls provided in an embodiment of this application; Figure 2 A structural block diagram of an interactive image display system applied to a curtain wall, provided as an embodiment of this application; Figure 3 This is a schematic block diagram of an electronic device provided in an embodiment of this application. Detailed Implementation

[0011] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.

[0012] To make the purpose, technical solution, and advantages of this application clearer, the following will be described in conjunction with the appendix. Figure 1-3 The following is an explanation using specific examples.

[0013] Please refer to Figure 1 , Figure 1 This is a flowchart illustrating an embodiment of an interactive image display method applied to a curtain wall, the method comprising: S101: The video stream is collected by a camera device deployed in the curtain wall area, and the video stream is preprocessed and human target detection is performed to obtain the position feature information of the human target.

[0014] In this embodiment, the curtain wall area refers to the building facade with the curtain wall installed and the interactive space within a certain range in front of it, including the curtain wall body area, the ground interactive area in front of the curtain wall, and the sensing space around the curtain wall. The camera device is an image acquisition device used to capture real-time images of the curtain wall area, including high-definition cameras, infrared cameras, binocular cameras, network cameras, etc., which can be fixedly deployed at suitable locations around the curtain wall to cover the interactive area. The live video stream is a continuous sequence of image frames captured and output in real time by the camera device, including dynamic visual information within the curtain wall area, and transmitted to the processing unit in digital format.

[0015] In this embodiment, preprocessing refers to optimization operations performed on the image frames of the live video stream before human target detection. These operations include, but are not limited to, denoising, cropping, resizing, brightness correction, and distortion correction, to improve detection accuracy. Human target detection is the process of identifying and locating human targets from video frames based on visual algorithms or intelligent models. It can accurately identify human bodies in the image and eliminate non-human interference factors. Human targets are one or more pedestrians detected in the video frame. They are the main entities that trigger the interactive screen and are also the corresponding objects of positional feature information. Positional feature information is a set of relevant data describing the spatial location of the human target, including the human's coordinates in the image, detection confidence, bounding box parameters, etc.

[0016] S102: Spatial match the location feature information of the human target with the preset region of interest to determine whether there is a triggering intent and obtain the triggering determination result.

[0017] In this embodiment, the preset region of interest (ROI) is a specific spatial range pre-defined in the space in front of the curtain wall or in the screen coordinate system to trigger interactive effects on the curtain wall. It is an effective area for determining whether the user has an interactive intention. Spatial matching is a calculation process that compares the spatial position parameters of the human target with the range parameters of the ROI and determines the spatial inclusion relationship to determine whether the human target falls into the trigger area.

[0018] In this embodiment, the triggering intent is the user's subjective behavioral tendency to actively interact with the curtain wall by entering a designated area, staying there, or performing actions. It serves as the basis for the system to initiate interactive lighting effects. The trigger determination result is a conclusive result obtained after spatial matching and intent judgment. It is used to indicate whether the triggering conditions are met and serves as the basis for subsequently controlling the curtain wall lighting equipment to perform interactive displays.

[0019] S103: Based on the trigger determination result and the preset trigger-light effect mapping relationship, obtain the floodlight control parameters for controlling the curtain wall lighting equipment.

[0020] In this embodiment, the trigger-light effect mapping relationship is a pre-set rule that corresponds trigger conditions to corresponding light effects. For example, human presence leads to a warm light gradient, and gesture commands lead to color switching, which serves as the basis for light effect control. Floodlight control parameters are parameters used to drive the curtain wall lighting equipment, including brightness, color, gradient speed, and light effect mode, directly determining the operating state of the lighting equipment. Curtain wall lighting equipment is a lighting device used to achieve interactive floodlight displays on the curtain wall. It consists of multiple independent lighting units, such as LED light strips and floodlights, capable of receiving control commands and executing corresponding light effect changes. Floodlight control parameters are key parameters for controlling the operation of the curtain wall lighting equipment, including brightness, color, and gradient rhythm, ensuring that the light effect display meets preset requirements.

[0021] S104: Based on the floodlighting control parameters, generate a sequence of control instructions for each lighting unit of the curtain wall lighting equipment to perform changes in luminous efficacy.

[0022] In this embodiment, the lighting unit is the smallest independent control unit of the curtain wall lighting equipment, such as a single LED light or LED light group. It can receive instructions independently and complete the adjustment of light effects, serving as the basic execution unit for light effect display. Light effect change refers to the dynamic changes in brightness, color, flicker frequency, and gradation rhythm of the curtain wall lighting equipment under the action of control instructions, used to achieve interactive display effects. The control instruction sequence is a set of instructions composed of multiple consecutive control instructions, arranged according to a preset time sequence. It includes execution information such as the brightness, color, and gradation time of the lighting unit, used to drive each lighting unit to execute light effect changes according to a specified rhythm and sequence, making the light effect display coherent and synchronous.

[0023] S105: Sends a sequence of control commands to the controller of the curtain wall lighting equipment to drive the device of the curtain wall lighting equipment to perform dynamic floodlighting display.

[0024] In this embodiment, the controller is a hardware control unit that receives and parses control commands to drive the curtain wall lighting equipment to perform lighting effect actions. Floodlighting dynamic display is a continuous visual effect of brightness, color, and area changes presented by the curtain wall lighting equipment in space and time according to control commands.

[0025] As can be seen from the above, this application uses a camera device to collect on-site video streams and perform human target detection to achieve interactive screen wall images based on human position. Compared with a fixed playback mode, it can realize real-time dynamic interaction between people and the screen wall, enhancing the fun and immersiveness of the screen wall display. By matching the human position with the area of ​​interest to determine the trigger intention, and then generating floodlighting control parameters according to the trigger-light effect mapping relationship, it can accurately respond to human behavior and drive lighting equipment to achieve dynamic light effect display. This makes the screen wall no longer a single display carrier, but a perceptible and interactive intelligent scene device, enhancing the display effect and user experience of the screen wall.

[0026] In one embodiment of this application, preprocessing and human target detection are performed on the live video stream to obtain the positional feature information of the human target, including: The live video stream is transmitted to the edge computing unit, which operates based on a deep learning-based object detection model. Human targets are detected frame by frame in the live video stream based on the target detection model, and human targets in the picture are identified. Extract the position coordinates, bounding box size, and detection confidence of each human target in the image coordinate system as position feature information.

[0027] In this embodiment, the edge computing unit is computing hardware deployed near the field, possessing independent data processing and model inference capabilities. It is used to process video stream data locally, reducing transmission latency and cloud pressure. The deep learning-based object detection model is an artificial intelligence model trained on deep neural networks, used to automatically identify and locate human targets from images or videos.

[0028] Specifically, the target detection model based on deep neural networks has a hierarchical structure consisting of an input layer, a feature extraction layer, a feature fusion layer, and a detection output layer. These layers are closely interconnected to adapt to the human detection requirements of curtain wall scenarios. The input layer receives single-frame images from the live video stream transmitted by the edge computing unit, performing image denoising, size normalization (to 640×640 pixels), and color gamut standardization. The feature extraction layer uses a deep residual network as its core backbone, comprising four residual modules. Through stacked convolutional layers, batch normalization layers, and the ReLU activation function, it progressively extracts low-level textures, mid-level contours, and high-level semantic features, achieving accurate feature representation of humans at different distances and in different poses. The feature fusion layer uses a feature pyramid network, fusing feature maps of different scales through top-down upsampling and bottom-up feature transfer to address the detection differences between near and far human targets in curtain wall scenarios. The detection output layer includes a classification branch and a regression branch. The classification branch determines whether the target is a human, while the regression branch outputs the bounding box coordinates and size information of the human target, ultimately passing the detection results to the subsequent feature extraction stage.

[0029] The object detection model training process follows a complete workflow of dataset construction, transfer learning, iterative training, and performance verification, enabling the model to adapt to the needs of curtain wall scenarios. First, human image samples are collected from the curtain wall site under different lighting conditions (strong light, weak light, backlight), distances (near, medium, far), and postures (standing, walking, standing). These samples are divided into training, validation, and test sets in an 8:1:1 ratio, and data enhancement operations such as random cropping, flipping, and color gamut dithering are used to expand sample diversity. Second, based on the PyTorch framework, pre-trained ResNet weights are used for transfer learning. The parameters of the feature extraction layer are frozen initially, and only the feature fusion layer and detection output layer are trained to quickly adapt to the curtain wall scene features. Then, all parameters are gradually unfrozen for fine-tuning. During training, 300 iterations are set, with model performance verified every 10 iterations. Training stops when the average accuracy of the validation set shows no improvement for 15 consecutive iterations, and the optimal model weights are saved. The target detection model takes video frame images preprocessed by edge computing units as input and outputs the position coordinates (x, y) of the human target in the image coordinate system, the bounding box size (width, height) and the detection confidence. The output data is directly used as position feature information, providing data support for subsequent spatial matching with the region of interest and triggering intent judgment, and realizing the deep integration of the model with the interactive scene of the curtain wall image.

[0030] In this embodiment, frame-by-frame human target detection involves sequentially performing target detection on each frame of the video stream to achieve continuous, real-time human recognition and tracking. The screen coordinate system is a two-dimensional coordinate system established based on the video frame, used to locate the relative position of the human target within the image. The position coordinates are the horizontal and vertical coordinates of the human target in the screen coordinate system, representing its specific location within the frame. The bounding box size is the width and height of the smallest bounding rectangle enclosing the human target, representing the size and range of the human target within the frame. The detection confidence score is the probability value output by the target detection model representing the current detection result as a real human target, used to indicate the reliability of the detection result.

[0031] As can be seen from the above, this embodiment, by using an edge computing unit equipped with a deep learning object detection model, achieves frame-by-frame human target detection in the live video stream. It can stably identify human targets and extract their position coordinates, bounding boxes, and detection confidence scores under low latency conditions, meeting the real-time requirements of the interactive curtain wall. The edge computing architecture reduces cloud transmission pressure and network latency, improving the system's operational stability in complex outdoor environments. By standardizing the output of position feature information, a unified and reliable data foundation is provided for subsequent spatial matching and trigger judgment, avoiding accidental triggering or delayed response due to unstable detection results. This approach balances detection accuracy and operational efficiency, enabling the interactive curtain wall to maintain good recognition performance even in environments with multiple users, long distances, and complex lighting conditions, thus improving the overall reliability of the interaction.

[0032] In one embodiment of this application, human target detection is performed frame by frame on the live video stream based on a target detection model, and the method further includes: Adjust the detection confidence threshold based on the image quality indicators of the on-site video stream; Obtain the image brightness from the image quality index. If the image brightness is less than the first brightness threshold, lower the detection confidence threshold to the first confidence threshold. When the screen brightness is greater than the second brightness threshold, the detection confidence threshold is raised to the second confidence threshold. Among them, the second brightness threshold is greater than the first brightness threshold, and the second confidence threshold is greater than the first confidence threshold.

[0033] In this embodiment, the image quality index is a quantitative parameter used to represent the imaging quality of the video image, including image brightness, sharpness, contrast, signal-to-noise ratio, etc. The detection confidence threshold is a critical value used by the target detection model to determine whether the detection result is valid; if it is greater than the detection confidence threshold, it is determined to be a valid human target; if it is less than the threshold, it is determined to be an interference or invalid target. Image brightness is a numerical index characterizing the overall brightness of the video image, representing the intensity of ambient light.

[0034] In this embodiment, the first brightness threshold is a set brightness threshold used to determine if the environment is in a low-light state. The second brightness threshold is a set brightness threshold used to determine if the environment is in a strong-light state, and is greater than the first brightness threshold. The first confidence threshold is a smaller detection confidence threshold used in low-light environments. The second confidence threshold is a larger detection confidence threshold used in strong-light environments, and is greater than the first confidence threshold.

[0035] The first brightness threshold is obtained as follows: environmental brightness samples of the curtain wall area under typical lighting scenarios at all times, such as early morning, evening, cloudy days, and night, are collected, with a sample size of no less than 1,000 sets. At the same time, the human target detection accuracy corresponding to each set of samples is recorded. Valid samples are selected with a detection accuracy of no less than 90% as a constraint. The brightness values ​​of the valid samples are statistically analyzed, and the lowest brightness value among them is taken as the first brightness threshold. This ensures that even in low-light scenarios, by reducing the detection confidence threshold, the human target detection accuracy can still meet the design requirements.

[0036] The second brightness threshold is obtained as follows: To ensure the visual experience of the curtain wall light and shadow display, environmental brightness samples of the curtain wall area are collected at different times (early morning, noon, dusk, and late night) and under different weather conditions (sunny, cloudy, and rainy), with a sample size of no less than 1,000 sets; combined with the light and shadow display requirements of the target scene, an effective sample range that can match the light and shadow display effect and the ambient lighting conditions is selected, and the median value within this range is taken as the second brightness threshold, so that the curtain wall light and shadow display can still present the preset visual contrast and sense of layering in complex lighting environments.

[0037] The first confidence threshold was obtained as follows: Taking the human target detection requirements in low-light environments (such as early morning, evening, and cloudy days) in curtain wall scenarios as the core, on-site human detection samples were collected under corresponding lighting conditions. The constraint was to ensure that the human target detection rate was not less than the preset minimum detection rate threshold (e.g., 95%) and to control the false detection rate to be no more than 5%. This was determined through multiple tests and iterative calibrations. This ensured that in low-light scenarios where the screen brightness was less than the first brightness threshold, lowering the detection confidence threshold could effectively improve the detection probability of human targets, avoid missed detections due to insufficient lighting, and control the number of false detections within an acceptable range, thus ensuring the accuracy and reliability of interactive trigger judgment.

[0038] The second confidence threshold was obtained as follows: Combining the actual needs of the interactive curtain wall scenario, and constrained by the accuracy of human target detection and the adaptability to ambient lighting, human detection samples were collected under different lighting conditions (e.g., sunny days and cloudy days). The detection accuracy corresponding to different confidence thresholds was statistically analyzed, and a confidence judgment standard that can balance low false detection rate and high detection rate was selected. This standard was determined after multiple tests and calibrations, taking into account the recognition ability of the target detection model and the response characteristics of the curtain wall lighting equipment. This ensures that in scenarios such as strong light, it can accurately identify human targets, reduce false detections, and avoid missed detections, forming a gradient difference with the first confidence threshold to adapt to the human detection needs under different lighting environments.

[0039] As can be seen from the above, this embodiment improves the robustness of human detection under different lighting conditions by adaptively adjusting the detection confidence threshold based on the screen brightness. Lowering the confidence threshold in low-light conditions avoids missed detections of human bodies; raising the threshold in high-light conditions reduces false detections caused by strong light interference. This dynamic threshold mechanism balances detection sensitivity and accuracy, ensuring stable identification of human targets under varying lighting conditions such as morning / evening, sunny / dark, and indoor / outdoor environments, preventing interactive failures or false triggers due to lighting changes. This adaptive strategy eliminates the need for frequent manual parameter adjustments, improving automation and environmental adaptability, and ensuring stable and reliable operation of the interactive screen image throughout the day.

[0040] In one embodiment of this application, the process of preprocessing and detecting human targets in the live video stream to obtain the location feature information of the human targets further includes: Perform skeletal key point detection on the detected human target and extract the skeletal point coordinate information of the human target; Pose recognition is performed based on skeletal point coordinate information to obtain the orientation information, gesture information, or limb movement information of the human target, which are used as pose feature information. The posture feature information is matched with a preset interactive posture library. If the match is successful, an enhanced trigger signal is generated. When the trigger determination result is a valid trigger, a differentiated control command sequence is generated based on the enhanced trigger signal. The differentiated control command sequence is used to drive the curtain wall lighting equipment to perform differentiated light effects in color change, dynamic frequency change or regional diffusion speed change.

[0041] In this embodiment, skeletal keypoint detection is the process of identifying and locating key skeletal nodes on a human target, such as the head, shoulders, hands, and legs. Skeletal point coordinate information refers to the specific coordinate values ​​of each skeletal keypoint in the image coordinate system, used to characterize the spatial distribution and posture of the human skeleton. Posture recognition is the process of determining the current body posture and direction of movement of a human target based on the spatial distribution relationship of skeletal point coordinate information. Orientation information is the location of the human target in the image, such as front, side, or back, indicating the direction the human is facing. Gesture information is a specific action form composed of skeletal keypoints of the human hand, such as waving, clenching a fist, or raising a hand, used to convey simple interactive commands. Limb movement information is the overall movement formed by the coordinated movement of skeletal keypoints of the human torso and limbs, such as raising a hand, turning around, or waving, used to characterize human interactive behavior. Posture feature information is a set of information composed of human orientation, gestures, and limb movements, used to supplement positional feature information and enrich interaction triggering conditions.

[0042] In this embodiment, the preset interactive posture library is a pre-stored set of standard postures used to trigger differentiated interactions on the curtain wall, including feature data of various preset gestures and body movements. The enhanced trigger signal is generated when the posture feature information successfully matches the interactive posture library, used to identify a clear and personalized interactive intent from the user, distinguishing it from the basic position trigger signal. The differentiated control command sequence is a set of timing control commands generated based on the enhanced trigger signal, different from ordinary trigger commands, used to drive the lighting equipment to present personalized lighting effects. Color change refers to the visual changes such as color switching and gradation presented by the curtain wall lighting equipment under the control commands. Dynamic frequency change refers to the frequency adjustment of the curtain wall lighting equipment's flashing and alternating brightness, used to present dynamic effects with different rhythms. Area diffusion speed change refers to the speed adjustment of the curtain wall light effect's diffusion from the trigger area to the surrounding area, used to present different diffusion visual effects. Differentiated lighting effects are personalized lighting effects that differ from the basic trigger lighting effects, corresponding to different posture interaction intents.

[0043] As can be seen from the above, this embodiment, by adding skeletal key point detection and posture recognition to human body detection, extracts posture features such as orientation, gestures, and limb movements, and matches them with a preset interactive posture library to generate enhanced trigger signals, thus upgrading from position-triggered interaction to posture-based interaction and enriching the interactive forms of the curtain wall. By generating differentiated control commands through posture matching, it is possible to drive lighting to achieve diverse light effect changes such as color, frequency, and diffusion speed, making the interaction more layered and personalized. Compared to simple position-triggered interaction, posture-based interaction is closer to natural human-computer interaction habits, enhancing user participation and experience, while also providing expanded space for customized curtain wall displays and multi-mode interaction, enhancing the technological feel and aesthetic appeal of the device.

[0044] In one embodiment of this application, spatial matching is performed between the location feature information of a human target and a preset region of interest to determine whether a triggering intent exists, thereby obtaining a trigger determination result, including: A preset area on the ground or in the space in front of the curtain wall is taken as the preset region of interest, and a mapping relationship between the ground world coordinate system and the picture coordinate system is established through a perspective transformation matrix. The position coordinates of the human target are transformed to the ground world coordinate system through a perspective transformation matrix to obtain the spatial position information of the human target in real space; Determine whether the spatial location information falls within the preset region of interest; If the target falls into the area, a dwell timer is started to record the continuous dwell time of the human target in the preset region of interest, and to determine whether the continuous dwell time is greater than the preset time threshold. If the duration of continuous stay exceeds the preset time threshold, the presence of a triggering intent is confirmed, and a valid triggering signal is generated as the triggering determination result. If the spatial location information moves out of the preset region of interest, the dwell timer will stop and the continuous dwell time will be reset to zero.

[0045] In this embodiment, the preset area is a designated range for human-computer interaction defined on the ground or in space in front of the curtain wall. The ground world coordinate system is a three-dimensional or two-dimensional real-space coordinate system established based on the actual ground surface of the curtain wall site, used to represent the real position of objects in the physical world. The image coordinate system is a two-dimensional image coordinate system established based on the video image captured by the camera device, used to represent the position of the human target in the image. The perspective transformation matrix is ​​a mathematical transformation matrix used to realize the conversion between the image coordinate system and the ground world coordinate system, used to eliminate shooting perspective distortion. The mapping relationship is the correspondence between image coordinate points and real-world coordinate points established through perspective transformation.

[0046] In this embodiment, the spatial location information is the actual coordinate data of the human target in the ground world coordinate system, representing its actual position in physical space. The dwell timer is a timing module used to count the continuous dwell time of the human target within the region of interest. The continuous dwell time is the cumulative time the human target remains within the region of interest from the moment it enters until it leaves. The preset time threshold is a set minimum dwell time threshold used to determine if the user has an interactive intent. The valid trigger signal is a legitimate trigger command signal generated after the position and dwell time conditions are met, used to activate the interactive lighting effects on the curtain wall.

[0047] As can be seen from the above, this embodiment establishes a mapping between the screen coordinate system and the ground world coordinate system through a perspective transformation matrix, converting the human image coordinates into a real spatial position, achieving accurate spatial matching and judgment, and avoiding misjudgment of the trigger area due to perspective distortion. By judging whether the human body falls into the area of ​​interest and counting the continuous dwell time, it effectively distinguishes between unintentional passing by and intentional interaction behavior, reducing false triggers. Setting a dwell time threshold can improve the rigor of interaction; only users who truly stay and pay attention can trigger changes in lighting effects, ensuring the purposefulness and effectiveness of the curtain wall interaction. This spatial matching mechanism is stable and reliable, suitable for large-area curtain walls and long-distance interaction scenarios, improving the accuracy of interaction judgment and enhancing the overall user experience.

[0048] In one embodiment of this application, the preset time threshold is adjusted according to the distance between the human target and the curtain wall, including: Obtain the vertical distance between the human target and the curtain wall in the spatial location information; If the vertical distance is less than the first distance threshold, the first time threshold is set as the preset time threshold. If the vertical distance is greater than or equal to the first distance threshold and less than the second distance threshold, the second time threshold is set as the preset time threshold. When the vertical distance is greater than or equal to the second distance threshold, a third time threshold is set as the preset time threshold. Among them, the first time threshold is less than the second time threshold, and the second time threshold is less than the third time threshold.

[0049] In this embodiment, the vertical distance is the straight-line distance in the vertical direction between the human target and the curtain wall surface in real space. The first distance threshold is a set near-distance threshold used to distinguish the near-distance range between the human target and the curtain wall. The second distance threshold is a set mid-to-long-distance threshold used to distinguish the mid-to-long-distance range between the human target and the curtain wall. The first time threshold is the triggering time threshold used for the near-distance range. The second time threshold is the triggering time threshold used for the mid-distance range. The third time threshold is the triggering time threshold used for the long-distance range.

[0050] Specifically, the first and second distance thresholds are determined by surveying the effective interactive space in front of the curtain wall on-site, combining the camera device to detect the viewing angle and clarity, statistically analyzing the interaction frequency of the human body at different distances, and selecting the natural dividing points between near-distance and medium-distance, and medium-distance and far-distance, which are respectively used as the first and second distance thresholds.

[0051] The first, second, and third time thresholds are set according to human interaction habits at different distances. Interaction is more convenient at close distances (less than the first distance threshold), so the first time threshold is set; at medium distances (between the first and second distance thresholds), the second time threshold is set; and at long distances (greater than the second distance threshold), a longer stay is required to confirm the interaction intention, so the third time threshold is set, which is consistent with human behavior logic and adaptable to the scenario.

[0052] In this embodiment, the preset time threshold is adaptively adjusted based on the distance between the human target and the curtain wall. The specific process is as follows: the vertical distance between the human target and the curtain wall surface is calculated and obtained from the spatial position information of the human target; if the vertical distance is less than the first distance threshold, it is determined to be a close-range interaction, and the first time threshold is set as the current preset time threshold; if the vertical distance is greater than or equal to the first distance threshold and less than the second distance threshold, it is determined to be a medium-range interaction, and the second time threshold is set as the current preset time threshold; if the vertical distance is greater than or equal to the second distance threshold, it is determined to be a long-range interaction, and the third time threshold is set as the current preset time threshold; wherein, the first time threshold is less than the second time threshold, and the second time threshold is less than the third time threshold, so as to realize the adaptive interaction logic that the closer the distance, the more sensitive the trigger, and the longer the dwell time required, the farther the distance.

[0053] As can be seen from the above, this embodiment dynamically adjusts the dwell time threshold based on the distance between the human body and the curtain wall. The closer the distance, the smaller the threshold; the farther the distance, the larger the threshold, which better aligns with human behavior and visual perception patterns. Users at closer distances have more focused attention and can trigger interactions quickly; users at farther distances require a longer dwell time to confirm their interaction intent, avoiding accidental triggering from a distance. This adjustment strategy is more intelligent and user-friendly, balancing response speed and trigger accuracy, and improving the user's interactive experience at different distances. It also adapts to different curtain wall sizes and different layout scenarios, improving scene adaptability.

[0054] In one embodiment of this application, the device for sending a sequence of control commands to the controller of the curtain wall lighting equipment to drive the curtain wall lighting equipment to perform dynamic floodlighting display further includes: Obtain the equipment configuration information of the curtain wall lighting equipment, including the spatial layout information of the lighting units, equipment response delay parameters, and dimming accuracy parameters; Based on the equipment response delay parameters, the control command sequence is compensated for time axis to generate a first control command sequence that has been time-synchronized. Based on the dimming accuracy parameters, the brightness values ​​in the control command sequence are adapted, and the continuous brightness values ​​are mapped to the discrete dimming levels supported by the lighting unit to generate a second control command sequence. Based on the spatial layout information of the lighting units, the first control command sequence and the second control command sequence are mapped to the control channels corresponding to each lighting unit according to their physical addresses to generate the final control command sequence. The final sequence of control commands is encapsulated into control data packets using a preset communication protocol and sent to the controller of the curtain wall lighting equipment.

[0055] In this embodiment, the device configuration information is a set of parameters describing the hardware characteristics and operational capabilities of the curtain wall lighting equipment, used for command adaptation and synchronous control. The spatial layout information of the lighting units includes the installation position, arrangement, physical number, and area division information of each lighting unit on the curtain wall surface. The device response delay parameter is the time required for the lighting unit to actually execute a change in luminous efficacy from receiving a control command. The dimming accuracy parameter is the minimum brightness adjustment step size that the lighting equipment can achieve, characterizing the fineness of brightness control. Time axis compensation adjusts the timing of command transmission earlier or later based on the device response delay to achieve luminous efficacy synchronization. The first control command sequence is a set of control commands after time synchronization correction and elimination of the impact of device delay.

[0056] In this embodiment, the continuous brightness value is a theoretically calculated continuously changing brightness value, which is a floating-point number. The discrete dimming level is the graded brightness level actually supported by the hardware, which is an adjustable range of finite integer levels. The second control instruction sequence is a set of instructions whose brightness parameters have been adapted to the hardware dimming accuracy and meet the requirements of the discrete level.

[0057] In this embodiment, physical address mapping is the process of binding logical control commands to the actual physical addresses and control channels of the lighting units in a one-to-one correspondence. A control channel is a communication port or drive channel used by the controller to independently drive a single lighting unit or a group of lighting units. The final control command sequence is a complete set of commands that can be directly issued and executed after time correction, brightness adaptation, and address mapping. The preset communication protocol is the data transmission format and interaction specification agreed upon between the controller and the upper-level processing unit. The control data packet is a complete data frame encapsulated according to the communication protocol, including the target address and command content.

[0058] As can be seen from the above, this embodiment uses lighting equipment configuration information to perform time compensation, dimming mapping, and spatial address mapping on control commands, ensuring a precise match between light output and hardware capabilities. Time axis compensation is performed through device response delay to guarantee synchronized operation of multiple lighting units; continuous brightness is mapped to discrete levels based on dimming accuracy, improving hardware compatibility; and spatial layout is mapped to corresponding control channels, achieving precise regional light effect control. Finally, the commands are encapsulated into standard communication data packets for transmission, resulting in smooth, synchronized, and flicker-free light effect display. This mechanism enhances compatibility and stability, making the dynamic display effect of floodlighting more delicate and precise, and improving the overall visual presentation quality of the curtain wall.

[0059] In one embodiment of this application, an interactive image display method applied to a curtain wall further includes: a frame rate adaptive adjustment mechanism for controlling the command sequence. Obtain the command processing capability parameters of the controller of the curtain wall lighting equipment. The command processing capability parameters include the maximum command receiving frequency and the single frame processing time. Based on the required rate of change of light effect in the dynamic display of floodlight illumination, determine the target frame rate of the control command sequence; When the target frame rate is greater than the maximum instruction receiving frequency, the control instruction sequence is time-compressed, and the control instructions of adjacent time slots are merged to generate a frame-down control instruction sequence while maintaining the trend of light effect change. When the target frame rate is less than the product of the maximum command receiving frequency and the preset redundancy coefficient, and the target frame rate is less than the preset critical flicker frequency of the human eye, the control command sequence is time-interpolated, and intermediate control commands are inserted between adjacent time slices to generate a frame-up control command sequence.

[0060] In this embodiment, the frame rate adaptive adjustment mechanism is an adaptive control strategy that automatically adjusts the frame rate of the control command sequence transmission based on hardware processing capabilities and lighting effect display requirements. The command processing capability parameter is a hardware parameter characterizing the data reception and parsing processing capabilities of the curtain wall lighting equipment controller. The maximum command reception frequency is the upper limit of the maximum number of command frames that the controller can stably receive and process per unit time. The single-frame processing time is the time required for the controller to receive, parse, and execute a single frame control command.

[0061] In this embodiment, the required rate of change of light effect refers to the speed at which brightness, color, and area diffusion effects change during the dynamic display of floodlight illumination. The target frame rate is the ideal frame rate for sending the control command sequence to meet the requirements of the light effect display.

[0062] In this embodiment, timing compression is a frame rate reduction method that merges adjacent instructions and reduces the total number of frames without disrupting the overall light effect variation trend. A time slice is an independent time unit corresponding to each frame instruction in the control instruction sequence. The frame rate reduction control instruction sequence is an instruction sequence that has undergone timing compression processing and frame rate reduction to adapt to the upper limit of the controller hardware.

[0063] In this embodiment, the redundancy coefficient is a proportional coefficient set to reserve processing capacity margin to ensure stable system operation. The critical flicker frequency for the human eye is the lowest frequency threshold at which the human eye cannot distinguish light effect flicker and perceives it as a continuous, smooth effect. Timing interpolation is a processing method that inserts transition instructions between adjacent instruction time slices to increase the instruction frame rate and make the light effect smoother. Intermediate state control instructions are transition state instructions calculated based on the instructions of preceding and following frames, used to achieve smooth light effect transitions. The frame-up control instruction sequence is an instruction sequence that, after timing interpolation processing and frame rate increase, results in a smoother light effect.

[0064] As can be seen from the above, this embodiment, by setting an adaptive frame rate adjustment mechanism for control commands, dynamically adjusts the frame rate according to the controller's processing capacity and the requirements of light effect changes, avoiding controller overload due to excessively high frame rates or flickering and stuttering due to excessively low frame rates. When the target frame rate exceeds the hardware limit, timing compression is performed to maintain the light effect trend; when the frame rate is too low, timing interpolation is performed to improve smoothness, making the visual experience comfortable and smooth for the human eye. The adaptive frame rate strategy makes full use of hardware performance, ensuring the continuity and stability of light effect display, improving the visual effect of the curtain wall, while reducing system resource consumption and improving the reliability and service life of the device under long-term operation.

[0065] In one embodiment of this application, the camera device includes at least two non-collinearly arranged camera units; obtaining the positional feature information of the human target includes: Acquire synchronized video frames captured simultaneously by at least two camera units; Human target detection is performed on each synchronized video frame to obtain the image coordinates of the human target from each viewpoint. Based on epipolar geometric constraints, stereo matching of human targets from various perspectives is performed to determine multi-view projection point pairs belonging to the same human target. Based on the multi-view projection point pairs and the internal and external parameters of the camera unit, the spatial position information of the human target in the three-dimensional world coordinate system is calculated by the principle of triangulation. Spatial location information includes three-dimensional coordinate values, which are used to replace or supplement the spatial location information obtained by perspective transformation matrix.

[0066] In this embodiment, the camera unit is an independent image acquisition module constituting the camera device, possessing independent imaging and data output capabilities. Non-collinear arrangement means that at least two camera units are installed in different spatial locations, not on the same straight line, to form a stereoscopic visual baseline. Synchronous video frames are video image frames acquired by different camera units at the same time, ensuring temporal consistency. Image coordinates at each viewpoint are the two-dimensional coordinate positions of the human target in the images captured by different camera units.

[0067] In this embodiment, epipolar geometric constraints are the geometric constraints satisfied by the projection points of the same spatial point under different viewpoints in multi-view visual measurement, used for stereo matching. Stereo matching is the process of finding matching points corresponding to the same spatial object in multi-view images. Multi-view projection point pairs are sets of pixels corresponding to the same human target in three-dimensional space in images from different viewpoints. The intrinsic and extrinsic parameters of the camera unit are the camera's intrinsic parameters (focal length, principal point, distortion coefficients) and extrinsic parameters (rotation matrix, translation vector), used for three-dimensional reconstruction. Triangulation is a measurement method that uses multiple known viewpoints and observation directions to calculate the three-dimensional coordinates of a target point through geometric trigonometric relationships. The three-dimensional world coordinate system is a three-dimensional rectangular coordinate system describing the actual physical space of the curtain wall site, used for precise positioning of human targets. The three-dimensional coordinate values ​​are the X, Y, and Z axis coordinates of the human target in the three-dimensional world coordinate system, representing its precise spatial position.

[0068] As can be seen from the above, this embodiment achieves stereoscopic vision positioning by employing at least two non-collinear camera units. It calculates the three-dimensional spatial position of the human body through epipolar geometric constraints and triangulation, resulting in higher positioning accuracy and anti-interference capabilities compared to monocular perspective transformation. Three-dimensional coordinates can more accurately represent the true position, distance, and movement state of the human body, avoiding positional deviations caused by shooting angles and lens distortion. Multi-view synchronous detection and stereo matching improve recognition stability in multi-person scenarios, providing a solid data foundation for high-precision interaction, area segmentation, and group behavior analysis, enabling the interactive wall to be applied to more demanding and complex large-scale public scenarios.

[0069] In one embodiment of this application, determining whether a triggering intent exists and obtaining a triggering determination result further includes: Obtain the continuous motion trajectory of the human target in the three-dimensional world coordinate system, and perform time series modeling of the continuous motion trajectory; Based on Kalman filtering or linear prediction algorithms, a short-term predicted trajectory of a human target is generated, and the predicted arrival time of the human target entering a preset region of interest is calculated. Based on the predicted arrival time, the timing of the dwell timer is dynamically adjusted so that the dwell timer enters a pre-start state before the human target enters the region of interest.

[0070] In this embodiment, the continuous motion trajectory is a sequence of spatial position points of the human target that changes continuously over time in a three-dimensional world coordinate system, used to characterize the human's movement path. Time series modeling involves mathematically modeling the position data of the human target over time to describe its motion patterns and trends. Kalman filtering is an optimal state estimation algorithm used to predict the target's true position and movement trend from noisy motion data. The linear prediction algorithm is a prediction method based on historical position data, extrapolating the target's future position through linear fitting. The short-term predicted trajectory is the movement path of the human target in the near future, predicted based on historical motion trajectories. The predicted arrival time is the time point at which the human target is expected to enter the region of interest, calculated based on motion prediction.

[0071] Specifically, the linear prediction algorithm consists of four layers: a data input layer, a feature processing layer, a linear prediction layer, and a result output layer. The input layer receives continuous motion trajectory data of the human target in a three-dimensional world coordinate system, performing data denoising, outlier removal, and time series alignment preprocessing. The feature processing layer extracts temporal features from the preprocessed trajectory data, focusing on core features such as timestamps, changes in three-dimensional coordinates, and motion speed, constructing standardized temporal feature vectors. The linear prediction layer, the core layer of the algorithm, uses a linear regression model to construct a trajectory prediction function, achieving short-term trajectory prediction by fitting the linear relationship between historical trajectory features and time. The output layer converts the prediction results into a format, outputting the short-term predicted trajectory of the human target and the predicted arrival time upon entering the preset region of interest, directly connecting to the logic for adjusting the start timing of the dwell timer.

[0072] The training process of the linear prediction algorithm is as follows: First, continuous motion trajectory data of human figures in a three-dimensional world coordinate system are collected at the curtain wall site. This includes trajectory samples with different motion speeds (uniform speed, acceleration, deceleration) and different motion directions (approaching, moving away from, parallel to the curtain wall). A total of 500 complete trajectory data sets are collected and divided into training and test sets in a 7:3 ratio. Second, the training set data is preprocessed to remove noise, outlier removal, and time alignment, constructing a time-series feature vector. The least squares method is used to train a linear regression model, fitting the linear relationship between historical trajectory coordinates and time, and iteratively optimizing the model coefficients. During training, the mean squared error (MSE) between the predicted and actual coordinates is used as the loss function. When the loss function value is less than 1e... -4 Furthermore, if there is no decrease after 10 consecutive iterations, training is stopped and the optimal model parameters are saved. The model input is the continuous motion trajectory data of the human target in the three-dimensional world coordinate system (including timestamps, X / Y / Z three-dimensional coordinates), and the output is the short-term predicted trajectory of the human body (continuous predicted coordinate sequence) and the predicted arrival time when entering the region of interest. The output data is directly used to dynamically adjust the timing of the dwell time timer, realizing deep integration with the curtain wall trigger judgment logic and improving the efficiency of interactive response.

[0073] In this embodiment, the dwell timer is a timing module used to count the continuous dwell time of a human target within the region of interest. The start time is the point at which the dwell timer begins timing. The pre-start state is a preparatory working state that the timer enters before the human target actually enters the area, used to improve the trigger response speed.

[0074] As can be seen from the above, this embodiment, by modeling human motion trajectories and using Kalman filtering to predict short-term trajectories, can predict in advance the moment when a human enters the region of interest, achieving pre-start and predictive interaction. Dynamically adjusting the start timing of the dwell timer makes the response smoother, reduces user waiting time, and improves the smoothness of the interaction. The trajectory prediction mechanism can effectively reduce latency and achieve instant response, making it suitable for fast-moving crowds and interactive screen scenarios with high smoothness requirements, thus improving the overall interactive experience and technological feel.

[0075] In one embodiment of this application, an interactive image display method applied to a curtain wall further includes: Construct a state machine that includes "entry-stay-departure" states to record the dwell trajectory patterns of human targets within a preset region of interest; Obtain the velocity vector sequence of the human target within the region of interest, and calculate the rate of velocity change. When the velocity vector of a human target is detected to exhibit a typical interactive trajectory pattern of deceleration-pause-acceleration, a pre-trigger signal is generated; When the velocity vector of a human target is detected to be passing through a non-interactive trajectory at a constant speed, the trigger determination is blocked, and no valid trigger signal is generated even if the continuous dwell time exceeds the preset time threshold.

[0076] In this embodiment, the state machine is a behavioral logic model composed of a finite number of states and transition rules between states, used to describe the behavioral changes of a human target within a region of interest. The enter-stay-leave state is one of the three basic states in the state machine used to characterize the interaction between the human target and the region of interest, corresponding to the behavioral stages of entering the region, staying within the region, and leaving the region, respectively. The dwell trajectory pattern is the behavioral characteristic pattern formed by the human target's movement path, dwell time, and changes in movement state within the region of interest.

[0077] In this embodiment, the velocity vector sequence is a data set of the magnitude and direction of human target motion speed arranged in chronological order. The rate of change of velocity is the magnitude of change in the human target's motion speed per unit time, used to characterize the acceleration and deceleration characteristics of the motion. The typical interactive trajectory pattern is the motion characteristic exhibited when a user actively interacts with the curtain wall, characterized by a behavior pattern of decelerating upon entry, pausing within the area, and then accelerating away. The pre-trigger signal is a preliminary trigger signal generated after recognizing typical interactive behavior, used to improve the timeliness and accuracy of subsequent trigger responses.

[0078] In this embodiment, the non-interactive trajectory mode refers to the motion characteristics exhibited when a user merely passes through an area without any intention to interact. This is characterized by moving directly through the area of ​​interest at a constant speed, without any noticeable pauses or deceleration. The trigger filtering process filters out behaviors that meet the required dwell time but lack interactive intent, rejecting the generation of valid trigger signals to avoid false triggers. A valid trigger signal is a legitimate trigger signal generated after determining that the user has a genuine intention to interact, used to initiate interactive lighting effects on the curtain wall.

[0079] As can be seen from the above, this embodiment distinguishes between interactive trajectories and passage trajectories through state machine and velocity vector analysis, identifies typical interaction patterns of deceleration-pause-acceleration, and shields against false triggers caused by simple passage behavior. Trajectory pattern judgment enhances the accuracy of trigger intent recognition, avoids frequent false actions in densely populated areas, and improves system reliability. This makes the curtain wall interaction more intelligent and more closely matches real user intentions, reduces invalid lighting effect switching, lowers system energy consumption, and enhances the audience's interactive experience, making the interaction more natural and targeted.

[0080] In one embodiment of this application, an interactive image display method applied to a curtain wall further includes: Integrate ambient light sensors into curtain wall lighting equipment to collect real-time ambient illuminance on the curtain wall surface; Using ambient illuminance as a feedforward compensation variable, the target brightness value and gradient curve parameters in the floodlight control parameters are corrected so that the light output of the curtain wall lighting equipment maintains the preset visual contrast under different ambient lighting conditions. Obtain real-time status feedback information of each lighting unit of the curtain wall lighting equipment. The real-time status feedback information includes the current brightness value and the progress of light effect change. The real-time status feedback information is timestamped with the trigger judgment result. When a human target is detected to have left the area of ​​interest but the progress of the light effect change has not been completed, a fast reset command is generated to force the curtain wall lighting equipment to enter the dimming reset process in advance.

[0081] In this embodiment, the ambient light sensor is integrated into the curtain wall lighting equipment and is a sensing device used to detect the ambient light intensity in real time. Ambient illuminance is a physical quantity that characterizes the overall brightness of the environment surrounding the curtain wall surface. The feedforward compensation variable uses ambient illuminance as a pre-intervention correction parameter to dynamically adjust the lighting output. The target brightness value is the ideal output brightness value set in the floodlighting control parameters. The gradient curve parameter describes the rate of change and curve shape of the light effect from one brightness / color to another. Visual contrast is the degree of brightness difference between the curtain wall lighting effect and the ambient background, used to ensure that interactive effects are clearly visible.

[0082] In this embodiment, the real-time status feedback information is the working status data transmitted back by the lighting unit in real time, used for closed-loop control and progress monitoring. The current brightness value is the actual brightness value currently output by the lighting unit. The light effect change progress is the proportion or stage of the current light effect animation that has been executed. Timestamp alignment is the time synchronization matching between the trigger judgment time and the light effect status feedback time. The quick reset command is a control command used to forcibly interrupt the current light effect and quickly restore it to the default lighting state. The gradual dimming reset process is the reset procedure where the lighting device smoothly reduces from the current brightness to the standby brightness.

[0083] As can be seen from the above, this embodiment achieves brightness feedforward compensation through an ambient light sensor, ensuring stable visual contrast of the curtain wall's lighting effect under various conditions, including daytime, nighttime, and changes in lighting, thus improving display consistency. Real-time status feedback and timestamp alignment enable closed-loop control of the lighting effect's progress, allowing for rapid resetting of lighting when people leave prematurely, avoiding ineffective lighting and resource waste. Environmental adaptation and the closed-loop status mechanism enhance the system's intelligence level, ensuring stable and controllable curtain wall display effects, reducing energy consumption, and extending the lifespan of lighting equipment.

[0084] In one embodiment of this application, an interactive image display method applied to a curtain wall further includes: Calculate the dense optical flow field frame by frame in the on-site video stream to obtain the motion vector distribution of pixels in the area in front of the curtain wall; Spatial clustering is performed based on the magnitude of the motion vector to divide the region into high-activity areas and static areas of interest. Dynamically designate high-activity areas as temporary regions of interest, and prioritize responding to triggering intents of human targets located within these temporary regions of interest; The priority of light effect response within the stationary focus area is reduced, and when a human target within the stationary focus area generates a triggering intent, a low brightness or slow gradient mode is used to respond.

[0085] In this embodiment, the dense optical flow field is a vector field calculated from all pixels in the video stream, representing the inter-frame pixel displacement, used to describe the overall motion information of the scene. Motion vector distribution is the overall spatial distribution of the magnitude and direction of the motion vectors of each pixel within the area in front of the screen. The amplitude of the motion vector is the numerical value of its length, representing the speed of the pixel's movement. Spatial clustering is the process of dividing spatially adjacent pixels into regions with different motion characteristics based on the similarity of their motion vector amplitudes. High-activity regions are spatial regions with large motion vector amplitudes and frequent, intense movement of people or objects. Static areas of interest are spatial regions with small motion vector amplitudes and where people are stationary or moving slowly.

[0086] In this embodiment, the temporary region of interest (ROI) is a dynamically adjusted area based on real-time motion activity, used to prioritize triggering interactions. Trigger intent refers to the behavioral tendency of a human target to interact with the screen through position, stillness, or movement. Light effect response priority is the order of response to triggered behaviors within different areas and the weight of resource allocation. Low brightness mode is a light effect output mode with brightness lower than the standard trigger effect. Slow gradient mode is a gradient display mode with a smoother light effect change speed and a longer transition time.

[0087] As can be seen from the above, this embodiment calculates the motion vector distribution through a dense optical flow field, dividing the area into high-activity regions and static areas of interest, thus achieving intelligent priority response. High-activity regions receive priority response and enhanced lighting effects; static regions receive lower priority and softer lighting effects, resulting in more rational allocation of system resources. Dynamic regions of interest can automatically adjust with crowd movement, improving the smoothness and relevance of interaction, avoiding rigid responses in fixed areas, and enhancing the intelligence and scene adaptability of the interactive wall.

[0088] In one embodiment of this application, an interactive image display method applied to a curtain wall further includes: Spatial clustering of multiple human targets within a high-activity area, and counting the number of targets within each cluster; When the number of targets exceeds the preset group response threshold, the lighting effect mode is automatically switched to group response mode. Group response modes include ripple diffusion light effects, regional synchronous flashing light effects, or light effect pattern generation that matches the number of targets.

[0089] In this embodiment, spatial clustering is a process of grouping multiple targets that are close together and clustered together into the same cluster based on their spatial distribution. The number of targets within a cluster is the total number of independent human targets included in the same spatial cluster area. The group response threshold is a set minimum number of people required to determine whether to trigger a group interactive lighting effect. Specifically, the group response threshold is obtained through statistical analysis of multiple scenarios on the curtain wall. In high-activity areas, interactive behavior samples are collected when different numbers of people gather. The minimum reasonable number of people required to trigger a group lighting effect when people gather is calculated. Combined with the lighting effect coverage of the curtain wall lighting equipment and the needs of the interactive scenario, a critical number of people that can distinguish between single-person interaction and group interaction is determined as the group response threshold. This ensures that the group lighting effect is not mistakenly triggered due to too few people, and that the mode can be switched in time when many people gather to adapt to the actual interactive scenario.

[0090] In this embodiment, the light effect mode is a fixed combination of visual effects formed by the operation of the curtain wall lighting equipment according to specific rules. The group response mode is a group light effect display mode specifically designed for multi-person interactive scenarios, distinct from single-person triggering. The ripple diffusion light effect is a dynamic light effect that spreads outward layer by layer from the cluster center, resembling water ripples. The area synchronous flashing light effect is a group light effect in which the lighting units corresponding to the crowd gathering area synchronously flash alternately between bright and dark. The target quantity matching light effect pattern generation is a customized light effect pattern that generates a corresponding number of light points, light strips, or graphics based on the number of human targets within the cluster.

[0091] As can be seen from the above, this embodiment spatially clusters and counts the number of people in high-activity areas. When the number exceeds the group response threshold, it automatically switches to group response mode, achieving intelligent switching between individual and group interaction. Light effects such as ripple diffusion, synchronized flashing, and people-matching patterns allow the curtain wall to adapt to scenes with large crowds, enhancing the atmosphere of public spaces. This enhances the scene scalability and visual appeal of the curtain wall interaction, enabling the device to serve both individual experiences and create group visual effects, thereby increasing its commercial and landscape value.

[0092] In one embodiment of this application, an interactive image display method applied to a curtain wall further includes: During the process of recording the continuous dwell time of a human target within a preset region of interest, the dwell time difference between adjacent time slices is calculated. When the dwell time difference value shows a positive jump and exceeds the preset rate threshold, it is determined that the human target has begun to intentionally linger, and a preload command is immediately generated. The preload command is sent to the controller of the curtain wall lighting equipment to drive the lighting unit into the preheating state. The preheating state is to increase the brightness of the lighting unit to the intermediate brightness level between the standby brightness and the target brightness, and complete the response delay precharge of the drive circuit. After the formal trigger judgment is completed, a complete brightness change command is generated, driving the lighting unit to quickly rise from the preheating state to the target brightness.

[0093] In this embodiment, continuous dwell time is the cumulative time a human target spends continuously within a preset region of interest. A time slice is the smallest unit of time used to segment and statistically analyze dwell time. The dwell time difference is the change in dwell time between two adjacent time slices, used to determine abrupt changes in dwell behavior.

[0094] In this embodiment, a positive jump is a sudden increase in the dwell time difference, showing a significant rise. The rate threshold is the critical rate of change used to determine whether the dwelling behavior has changed from unintentional passing through to intentional dwelling. Intentional dwelling is a behavioral state in which a human target actively stays in the area with the intention to interact. The preload command is a preparatory command issued in advance before the formal triggering, used to prepare for the light effect output.

[0095] The rate threshold is obtained through on-site measurement and scene adaptation calibration. Based on the behavior pattern of human beings staying in the area of ​​interest, samples of the rate of change of the dwell time from passing through to intentional dwelling are collected in different scenarios. The critical rate at which the difference in dwell time shows a significant positive jump is statistically determined. Combined with the response delay of lighting equipment and the pre-charging requirements of the drive circuit, the minimum rate value that can accurately determine the intentional dwelling of human beings is used as the rate threshold. This enables accurate identification of intentional dwelling behavior, avoids false triggering of preload commands, and ensures a smooth connection between the preheating state and subsequent triggering.

[0096] In this embodiment, the preheating state before triggering is the preparatory working state that the lighting unit enters before formal triggering, falling between standby and output. Standby brightness is the base brightness maintained by the curtain wall lighting equipment when there is no interactive trigger. Target brightness is the set brightness value that the lighting unit needs to achieve after interactive triggering. Intermediate brightness levels are transitional brightness levels between standby brightness and target brightness. The driving circuit is a hardware circuit module used to drive the lighting unit to achieve brightness and color adjustment. Response delay pre-charge is the pre-power-on and voltage-stabilizing pre-charge operation of the driving circuit to eliminate circuit response delay. The complete brightness change command is the control command used to increase the brightness from the preheating state to the target brightness after formal triggering.

[0097] As can be seen from the above, this embodiment detects intentional user lingering behavior by differentiating the dwell time and sends a preload command in advance, causing the lighting unit to enter a preheating intermediate state, shortening the response time after formal triggering, and achieving a near-instantaneous leap in luminous efficiency. The preload mechanism effectively offsets hardware response latency, improves the immediacy and smoothness of interaction, and makes the user experience closer to zero-latency interaction. At the same time, the preheating state does not affect normal standby power consumption, balancing the needs of fast response and energy saving, and improving the smoothness of curtain wall interaction.

[0098] In one embodiment of this application, an interactive image display method applied to a curtain wall further includes: Establish a time window constraint between the preload command and the formal trigger determination. If no formal trigger determination is received within the preset time window after the preload command is sent, control the lighting unit to exit the waiting-to-trigger preheating state and restore it to standby brightness. The preset time window is dynamically adjusted based on the average dwell time of the human target in the region of interest.

[0099] In this embodiment, the time window constraint is a rule that limits the effective time range between the preload command and the formal trigger determination, used to avoid invalid warm-up. The preset time window is the maximum time interval allowed from the preload state to the formal trigger. The formal trigger determination is the final confirmed valid trigger determination result after all conditions such as dwell time, posture, and trajectory are met. Exiting the preheating state means the lighting unit ends the pre-brightness and circuit pre-charge states, returning to the standby mode without interaction. Standby brightness is the normal low brightness state of the curtain wall lighting equipment when there is no interaction and no warm-up. The average dwell time is the average dwell time of multiple human targets within the region of interest within a statistical period.

[0100] As can be seen from the above, this embodiment sets a time window constraint for the preload command. If the command is not triggered within the time limit, it automatically exits the preheating process and returns to standby mode, avoiding energy waste and light pollution caused by ineffective preheating. The time window is dynamically adjusted based on the average dwell time, improving the adaptability of the strategy. This mechanism ensures both high response speed and strict reliability and energy efficiency, avoids light pollution and equipment damage caused by erroneous preheating, improves the long-term stable operation capability of the curtain wall, and makes the intelligent preheating mechanism safer and more practical.

[0101] Corresponding to the image interactive display method applied to curtain walls in the above embodiments, Figure 2 This is a structural block diagram of an interactive video display system applied to a curtain wall, according to an embodiment of this application. For ease of explanation, only the parts relevant to the embodiment of this application are shown. References Figure 2 The interactive video display system 20 applied to the curtain wall includes: a video detection module 21, an intent judgment module 22, a lighting parameter module 23, a control command module 24, and a command execution module 25.

[0102] Among them, the video detection module 21 is used to collect on-site video streams through camera devices deployed in the curtain wall area, and to preprocess and detect human targets in the on-site video streams to obtain the position feature information of human targets; The intent determination module 22 is used to spatially match the location feature information of the human target with the preset region of interest to determine whether there is a triggering intent and obtain the trigger determination result. The lighting parameter module 23 is used to obtain the floodlighting control parameters for controlling the curtain wall lighting equipment based on the trigger determination result and the preset trigger-light effect mapping relationship; Control instruction module 24 is used to generate a sequence of control instructions for each lighting unit of the curtain wall lighting equipment to perform luminous efficacy changes based on the floodlighting control parameters; The instruction execution module 25 is used to send a sequence of control instructions to the controller of the curtain wall lighting equipment, and drive the device of the curtain wall lighting equipment to perform dynamic display of floodlighting.

[0103] See Figure 3 , Figure 3 This is a schematic block diagram of an electronic device provided according to an embodiment of this application. Figure 3 The electronic device 300 in this embodiment may include one or more processors 301, one or more input devices 302, one or more output devices 303, and one or more memories 304. The processors 301, input devices 302, output devices 303, and memories 304 communicate with each other via a communication bus 305. The memories 304 store computer programs, including program instructions. The processors 301 execute the program instructions stored in the memories 304. Specifically, the processors 301 are configured to invoke the program instructions to perform the functions of the modules in the aforementioned device embodiments, for example... Figure 2 The functions of the video detection module 21, intent judgment module 22, lighting parameter module 23, control command module 24, and command execution module 25 shown are illustrated.

[0104] It should be understood that, in the embodiments of this application, the processor 301 may be a central processing unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.

[0105] Input device 302 may include a touchpad, a fingerprint sensor (for collecting the user's fingerprint information and fingerprint orientation information), a microphone, etc., and output device 303 may include a display (LCD, etc.), a speaker, etc.

[0106] The memory 304 may include read-only memory and random access memory, and provides instructions and data to the processor 301. A portion of the memory 304 may also include non-volatile random access memory. For example, the memory 304 may also store device type information.

[0107] In specific implementations, the processor 301, input device 302, and output device 303 described in the embodiments of this application can execute the implementation methods described in any embodiment of the interactive image display method for curtain walls provided in the embodiments of this application, or they can execute the implementation methods of the electronic devices described in the embodiments of this application, which will not be repeated here.

[0108] In another embodiment of this application, a computer-readable storage medium is provided. This computer-readable storage medium stores a computer program, which includes program instructions. When executed by a processor, the program instructions implement all or part of the processes in the methods described above. Alternatively, the computer program can instruct related hardware to complete the process. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include any entity or device capable of carrying computer program code, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc.

[0109] The computer-readable storage medium can be an internal storage unit of the electronic device in any of the foregoing embodiments, such as a hard disk or memory of the electronic device. The computer-readable storage medium can also be an external storage device of the electronic device, such as a plug-in hard disk, smart media card (SMC), secure digital card (SD), flash card, etc., equipped on the electronic device. Furthermore, the computer-readable storage medium can include both internal and external storage units of the electronic device. The computer-readable storage medium is used to store computer programs and other programs and data required by the electronic device. The computer-readable storage medium can also be used to temporarily store data that has been output or will be output.

[0110] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this application.

[0111] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the electronic devices and units described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0112] In the several embodiments provided in this application, it should be understood that the disclosed electronic devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces or units, or it may be an electrical, mechanical, or other form of connection.

[0113] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of the embodiments of this application, depending on actual needs.

[0114] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0115] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for interactive image display applied to curtain walls, characterized in that, include: By deploying camera devices in the curtain wall area to collect on-site video streams, and preprocessing and human target detection of the on-site video streams, the positional feature information of human targets is obtained; The location feature information of the human target is spatially matched with a preset region of interest to determine whether there is a triggering intent, and a triggering determination result is obtained. Based on the trigger determination result and the preset trigger-light effect mapping relationship, the floodlighting control parameters for controlling the curtain wall lighting equipment are obtained; Based on the floodlighting control parameters, a sequence of control instructions is generated for each lighting unit of the curtain wall lighting equipment to perform luminous efficacy changes. The control command sequence is sent to the controller of the curtain wall lighting equipment to drive the device of the curtain wall lighting equipment to perform dynamic floodlighting display.

2. The image interactive display method applied to curtain walls according to claim 1, characterized in that, The preprocessing and human target detection of the live video stream to obtain the positional feature information of the human target includes: The live video stream is transmitted to an edge computing unit, wherein the edge computing unit operates based on a deep learning-based target detection model. Based on the target detection model, human target detection is performed frame by frame on the live video stream to identify human targets in the scene; The position coordinates, bounding box size, and detection confidence of each human target in the image coordinate system are extracted as the position feature information.

3. The image interactive display method applied to curtain walls according to claim 2, characterized in that, The step of performing human target detection frame by frame on the live video stream based on the target detection model further includes: The detection confidence threshold is adjusted based on the image quality index of the live video stream. Obtain the image brightness from the image quality index. If the image brightness is less than the first brightness threshold, reduce the detection confidence threshold to the first confidence threshold. When the brightness of the image is greater than the second brightness threshold, the determination threshold of the detection confidence is raised to the second confidence threshold. Wherein, the second brightness threshold is greater than the first brightness threshold, and the second confidence threshold is greater than the first confidence threshold.

4. The image interactive display method applied to curtain walls according to claim 1, characterized in that, The step of preprocessing and human target detection of the live video stream to obtain the positional feature information of the human target also includes: Perform skeletal key point detection on the detected human target and extract the skeletal point coordinate information of the human target; Based on the skeletal point coordinate information, posture recognition is performed to obtain the orientation information, gesture information, or limb movement information of the human target, which are used as posture feature information. The posture feature information is matched with a preset interactive posture library. If the match is successful, an enhanced trigger signal is generated. When the trigger determination result is a valid trigger, a differentiated control instruction sequence is generated based on the enhanced trigger signal. The differentiated control instruction sequence is used to drive the curtain wall lighting equipment to perform differentiated light effects in color change, dynamic frequency change or regional diffusion speed change.

5. The image interactive display method applied to curtain walls according to claim 1, characterized in that, The step of spatially matching the location feature information of the human target with a preset region of interest to determine whether there is a triggering intent and obtain a trigger determination result includes: A preset area on the ground or in space in front of the curtain wall is taken as the preset region of interest, and a mapping relationship between the ground world coordinate system and the screen coordinate system is established through a perspective transformation matrix. The position coordinates of the human target are transformed to the ground world coordinate system through the perspective transformation matrix to obtain the spatial position information of the human target in real space; Determine whether the spatial location information falls within the preset region of interest; If the target falls into the area, a dwell timer is started to record the continuous dwell time of the human target in the preset region of interest, and it is determined whether the continuous dwell time is greater than a preset time threshold. If the continuous dwell time exceeds the preset time threshold, then a triggering intention is confirmed, and a valid triggering signal is generated as the triggering determination result. If the spatial location information moves out of the preset region of interest, the dwell timer is stopped and the continuous dwell time is reset to zero.

6. The image interactive display method applied to curtain walls according to claim 5, characterized in that, The preset time threshold is adjusted based on the distance between the human target and the curtain wall, including: Obtain the vertical distance between the human target and the curtain wall in the spatial location information; If the vertical distance is less than the first distance threshold, the first time threshold is set as the preset time threshold; If the vertical distance is greater than or equal to the first distance threshold and less than the second distance threshold, the second time threshold is set as the preset time threshold. When the vertical distance is greater than or equal to the second distance threshold, a third time threshold is set as the preset time threshold. Wherein, the first time threshold is less than the second time threshold, and the second time threshold is less than the third time threshold.

7. The image interactive display method applied to curtain walls according to claim 1, characterized in that, The method of sending the control command sequence to the controller of the curtain wall lighting equipment to drive the curtain wall lighting equipment to perform dynamic floodlighting display also includes: Obtain the equipment configuration information of the curtain wall lighting equipment, which includes the spatial layout information of the lighting units, equipment response delay parameters, and dimming accuracy parameters; Based on the device response delay parameters, the control command sequence is time-axis compensated to generate a first control command sequence after time synchronization correction; Based on the dimming accuracy parameters, the brightness values ​​in the control command sequence are adapted, and the continuous brightness values ​​are mapped to the discrete dimming levels supported by the lighting unit to generate a second control command sequence. Based on the spatial layout information of the lighting units, the first control command sequence and the second control command sequence are mapped to the control channels corresponding to each lighting unit according to their physical addresses to generate the final control command sequence. The final control command sequence is encapsulated into a control data packet using a preset communication protocol and sent to the controller of the curtain wall lighting equipment.

8. The image interactive display method applied to a curtain wall according to claim 7, characterized in that, Also includes: The frame rate adaptive adjustment mechanism of the control command sequence: Obtain the instruction processing capability parameters of the controller of the curtain wall lighting equipment, the instruction processing capability parameters including the maximum instruction receiving frequency and the single frame processing time; Based on the required rate of change of light effect in the dynamic display of floodlight illumination, the target frame rate of the control command sequence is determined; When the target frame rate is greater than the maximum instruction receiving frequency, the control instruction sequence is time-compressed, and the control instructions of adjacent time slots are merged to generate a frame-down control instruction sequence while maintaining the trend of light effect change. When the target frame rate is less than the product of the maximum instruction receiving frequency and the preset redundancy coefficient, and the target frame rate is less than the preset critical flicker frequency of the human eye, the control instruction sequence is time-interpolated, and intermediate state control instructions are inserted between adjacent time slices to generate a frame-up control instruction sequence.

9. The image interactive display method applied to curtain walls according to claim 1, characterized in that, The camera device includes at least two non-collinearly arranged camera units; obtaining the positional feature information of the human target includes: Acquire synchronized video frames captured simultaneously by the at least two camera units; Human target detection is performed on each of the synchronized video frames to obtain the image coordinates of the human target from each viewpoint; Based on epipolar geometric constraints, stereo matching is performed on human targets under each viewpoint to determine multi-view projection point pairs belonging to the same human target. Based on the multi-view projection point pairs and the internal and external parameters of the camera unit, the spatial position information of the human target in the three-dimensional world coordinate system is calculated by the principle of triangulation. The spatial location information includes three-dimensional coordinate values, which are used to replace or supplement the spatial location information obtained by perspective transformation matrix conversion.

10. An interactive image display system applied to curtain walls, characterized in that, include: The video detection module is used to collect on-site video streams through camera devices deployed in the curtain wall area, and to preprocess and detect human targets in the on-site video streams to obtain the positional feature information of human targets; The intent determination module is used to spatially match the location feature information of the human target with a preset region of interest to determine whether there is a triggering intent and obtain a trigger determination result. The lighting parameter module is used to obtain floodlighting control parameters for controlling the curtain wall lighting equipment based on the trigger determination result and the preset trigger-light effect mapping relationship; The control instruction module is used to generate a sequence of control instructions for each lighting unit of the curtain wall lighting equipment to perform light effect changes based on the floodlighting control parameters. The instruction execution module is used to send the control instruction sequence to the controller of the curtain wall lighting equipment to drive the device of the curtain wall lighting equipment to perform dynamic display of floodlighting.