Infrared light shadow interactive system and method

By using an infrared light and shadow interaction system, combined with image input devices and infrared interaction devices, the system enables multi-layer synthesis and dynamic visual interaction of user-drawn outline images and static ancient painting materials. This solves the problem of static image generation in existing technologies and enhances the degree of interaction freedom and visual expressiveness.

CN120762573BActive Publication Date: 2025-11-25SOUTHERN UNIVERSITY OF SCIENCE AND TECHNOLOGY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511293653.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-11
Publication Date
2025-11-25
Estimated Expiration
2045-09-11

AI Technical Summary

Technical Problem

Existing infrared light and shadow interactive technology lacks personalization capabilities in image generation, resulting in static content and insufficient freedom in image generation, interactive precision, and display hierarchy.

Method used

By using an infrared light and shadow interaction system, combined with image input devices, an interactive main platform, and infrared interactive devices, the system enables multi-layer synthesis of user-drawn outline images and static ancient painting materials. Furthermore, it utilizes infrared laser point trajectories to control the positional changes of the masked composite image within the multi-layered images, thereby achieving dynamic visual interaction.

Benefits of technology

A user-input-driven intelligent image processing system was built, realizing a closed loop from user creation to image processing output, which enhances the degree of freedom of interaction, visual expressiveness and cultural immersion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120762573B_ABST
    Figure CN120762573B_ABST
Patent Text Reader

Abstract

The application relates to the field of image processing, and discloses an infrared light and shadow interactive system and method. The system comprises an image input device, an interactive main platform, an infrared interactive device and a video output device. The image input device is used for acquiring a contour image drawn by a user. The interactive main platform is used for pre-processing the contour image to obtain a mask combined image, performing layer synthesis on the mask combined image and a preset static ancient painting material, and obtaining a multi-layer image. The infrared interactive device is used for capturing a laser point trajectory of a user operation point. The video output device is used for displaying the multi-layer image, binding coordinate information of the laser point with a coordinate position of the mask combined image, and controlling a position of the mask combined image in the multi-layer image based on the laser point coordinate information. The application constructs an interactive system with user input, intelligent processing and dynamic display feedback, and realizes a full-process closed loop from user creation, image processing and immersive interaction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing, and more particularly to an infrared light and shadow interaction system and method. Background Technology

[0002] In existing technologies, when using infrared laser point recognition to achieve wall projection interaction, non-contact control capabilities are available. However, the displayed content is mostly preset static resources, lacking personalized generation capabilities and unable to be deeply integrated with AI image generation mechanisms. Consequently, light and shadow interaction technology has significant shortcomings in terms of image generation freedom, infrared interaction accuracy, and display level. Specifically, light and shadow interaction suffers from the problem of static image display. Summary of the Invention

[0003] To address the problem of static image display in existing light and shadow interaction technologies, this application provides an infrared light and shadow interaction system and method.

[0004] In a first aspect, this application provides an infrared light and shadow interaction system, comprising:

[0005] Image input device for acquiring a user-drawn outline image;

[0006] The interactive main platform is used to receive the outline image sent by the image input device, preprocess the outline image to obtain a masked composite image, and composite the masked composite image with the preset static ancient painting material to obtain a multi-layer image; the masked composite image and the static ancient painting material in the multi-layer image are displayed as single layers, and each layer is independent of the others.

[0007] An infrared interactive device is used to capture the laser point trajectory of a user's operation point and extract the laser point coordinate information in the laser point trajectory;

[0008] A video output device is used to display the multi-layer image, and binds the laser point coordinate information with the coordinate position of the mask combination image in the multi-layer image, so as to control the position change of the mask combination image in the multi-layer image based on the laser point coordinate information, thereby realizing dynamic visual interaction.

[0009] In an optional embodiment, the infrared interactive device includes an infrared laser lifting rod and an infrared camera; wherein the infrared camera is fixedly installed at the target position of the video output device so that the infrared signal acquisition range of the infrared camera covers the boundary area of ​​the video output device;

[0010] The infrared camera is used to capture the trajectory of the laser point emitted by the infrared laser lifting rod in real time when the user moves the infrared laser lifting rod.

[0011] In an optional implementation, the step of preprocessing the contour image to obtain the masked composite image includes:

[0012] The preset image model is invoked to stylize the outline image, resulting in a lantern image with the target style;

[0013] Convert the lantern image with the target style into a transparent background image;

[0014] The transparent background image is subjected to image masking processing to obtain a masked composite image.

[0015] In an optional implementation, the process of generating the graph-generated model includes:

[0016] Acquire and preprocess a set of image samples containing elements of the target style colored lights;

[0017] Each image in the image sample set is annotated in both Chinese and English, and a custom trigger word is embedded in the annotation text. The custom trigger word is bound to the style weight of the graph-generated graph model to be generated. The custom trigger word is used to quickly trigger the graph-generated graph model to call the style features of the target style.

[0018] The Stable Diffusion model structure was trained using the LoRA algorithm in conjunction with the image sample set to obtain a graph-to-graph model with the ability to transfer elements of the target style lanterns.

[0019] In an optional implementation, acquiring and preprocessing the image sample set containing the target style colored light elements includes:

[0020] Multiple ancient painting images containing elements of the target style lanterns were processed by image matting and AI synthesis. The resulting matted sample images and AI sample images were used as the image sample set for model training.

[0021] The Real-ESRGAN deep learning super-resolution algorithm is used to uniformly enlarge each sample image in the image sample set to a preset pixel size;

[0022] A preset edge enhancement algorithm is used to enhance the contours and details of each magnified sample image, and a color equalization algorithm is used to restore the color of each sample image.

[0023] In an optional implementation, the main interactive platform is further used for:

[0024] The performance of the trained graph-generated graph model is validated, specifically including:

[0025] Obtain the lantern structure outline drawn by the user, and use the image-generated image mode to take the lantern structure outline as the input image to trigger the image-generated image model to generate a lantern image with lantern elements of the target style based on the custom trigger words and the style weights;

[0026] The lantern image is compared with the lantern image in the ancient painting corresponding to the image-generated image mode to obtain the performance verification result of the image-generated image model.

[0027] In an optional implementation, the step of performing image masking processing on the transparent background image to obtain a masked composite image includes:

[0028] The transparent background image is overlaid onto a preset Circle component using the Transform component to obtain a mask image;

[0029] The masked image is multiplied and synthesized with preset AI dynamic video material to obtain a masked composite image.

[0030] In an optional implementation, acquiring the laser point trajectory of the user's operation point and extracting the laser point coordinate information from the laser point trajectory includes:

[0031] A real-time video stream containing the laser point trajectory of the user's operation points is acquired, and the image in the video stream is cropped to obtain the target processing area; the target processing area is an image area that includes the screen area of ​​the video output device.

[0032] The target processing area is blurred to enhance the brightness features of the laser points within the target processing area. Threshold segmentation is then performed on the blurred target processing area to separate the laser points from the target processing area, resulting in laser point image data containing only the position information of the laser points.

[0033] The laser point image data is converted into geometric contour data to obtain the geometric edge information of the laser point;

[0034] The geometric contour data is converted into channel data to obtain the coordinate information of the laser point from the geometric edge information. The coordinate information is then normalized to map the coordinate information from the original pixel coordinate system to the standard coordinate system, thus obtaining the laser point coordinate information.

[0035] In an optional implementation, controlling the position of the contour image in the multi-layer image based on the laser point coordinate information includes:

[0036] The horizontal and vertical coordinates in the laser point coordinate information are mapped to the displacement parameters of the contour image in the multi-layer image, so as to achieve synchronization between the change of the laser point coordinate position and the change of the contour image coordinate position in the multi-layer image.

[0037] Secondly, this application provides an infrared light and shadow interaction method, applied to the aforementioned infrared light and shadow interaction system, the method comprising:

[0038] Obtain the outline image drawn by the user;

[0039] After preprocessing the outline image, a masked composite image is obtained. The masked composite image is then combined with a preset static ancient painting material to obtain a multi-layer image. The masked composite image and the static ancient painting material in the multi-layer image are displayed as single layers, and each layer is independent of the others.

[0040] Obtain the laser point trajectory of the user's operation point, and extract the laser point coordinate information from the laser point trajectory;

[0041] The laser point coordinate information is bound to the coordinate position of the masked combined image in the multi-layer image, so as to control the position change of the masked combined image in the multi-layer image based on the laser point coordinate information, thereby realizing dynamic visual interaction.

[0042] The embodiments of this application have the following beneficial effects:

[0043] This application provides an infrared light and shadow interactive system. Through the deep integration of intelligent image processing and infrared laser interaction, an interactive system with user input drive, intelligent image processing, and dynamic display feedback is constructed. Combined with layer synthesis and dynamic control technology, it realizes real-time dynamic interaction between images and videos, and constructs a closed-loop system integrating input-image generation-interaction-display. It breaks through the limitations of traditional technologies in terms of content generation freedom, interaction accuracy, and display level, and realizes a closed loop of the entire process from user creation to image processing output and then to immersive interaction. At the same time, it also improves the degree of interaction freedom, visual expressiveness, and cultural immersion. Attached Figure Description

[0044] To more clearly illustrate the technical solutions of this application, the accompanying drawings used in the embodiments will be briefly described below. It should be understood that the following drawings only show some embodiments of this application and therefore should not be considered as a limitation on the scope of protection of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0045] Figure 1A schematic diagram of an infrared light and shadow interaction system in an embodiment of this application is shown;

[0046] Figure 2 A flowchart of an infrared light and shadow interaction method in an embodiment of this application is shown;

[0047] Figure 3 A schematic diagram of the interactive interface of the image input device in an embodiment of this application is shown;

[0048] Figure 4 This illustration shows a flowchart of the graph model generation process in an embodiment of this application.

[0049] Figure 5 A schematic diagram of a sample image containing target style colored light elements is shown in an embodiment of this application;

[0050] Figure 6 This application shows a schematic diagram of an image of a static ancient painting material in an embodiment of the present application;

[0051] Figure 7 A static schematic diagram of an image frame within a multi-layer image is shown in an embodiment of this application. Detailed Implementation

[0052] The technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments.

[0053] The components of the embodiments of this application described and illustrated in the accompanying drawings can be arranged and designed in a variety of different configurations. Therefore, the following detailed description of the embodiments of this application provided in the drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.

[0054] In the following, the terms “comprising,” “having,” and their cognates, which may be used in various embodiments of this application, are intended only to indicate a particular feature, number, step, operation, element, component, or combination thereof, and should not be construed as excluding, firstly, the presence of one or more other features, numbers, steps, operations, elements, components, or combinations thereof, or adding the possibility of one or more features, numbers, steps, operations, elements, components, or combinations thereof.

[0055] Furthermore, the terms "first," "second," and "third" are used only to distinguish descriptions and should not be interpreted as indicating or implying relative importance.

[0056] Unless otherwise specified, all terms used herein (including technical and scientific terms) shall have the same meaning as commonly understood by one of ordinary skill in the art to which the various embodiments of this application pertain. Terms (such as those defined in commonly used dictionaries) shall be interpreted as having the same meaning as in their contextual meaning in the relevant technical field and shall not be construed as having an idealized or overly formal meaning, unless clearly defined in the various embodiments of this application.

[0057] The following detailed description of some embodiments of this application is provided in conjunction with the accompanying drawings. Unless otherwise specified, the following embodiments and features can be combined with each other.

[0058] This application provides an infrared light and shadow interaction system, wherein, as shown in the embodiments, Figure 1 As shown, the system includes an image input device 110, an interactive main platform 120, an infrared interactive device 130, and a video output device 140.

[0059] In this embodiment, the outline image of any object is drawn on the image input device 110. After a series of preset processing on the outline image, the interactive main platform 120 obtains a multi-layer image with multiple layers superimposed and each layer independent of the others. The multi-layer image includes a mask layer. The multi-layer image is displayed on the video output device 140. During this process, the infrared laser output by the infrared interactive device 130 falls on the image display area of ​​the video output device 140. By binding the coordinates of the laser point in the image display area with the coordinate position of the mask combination image in the multi-layer image, the position change of the mask combination image in the multi-layer image can be indirectly controlled by the change of the coordinate position of the infrared laser, thereby realizing dynamic interaction of light and shadow.

[0060] The outline image of any object described in this embodiment is only used as an example of a lantern.

[0061] For reference, the image input device 110 is used to acquire the outline image of any object drawn by the user (the arbitrary object is a lantern, and the outline image mentioned here and below refers to the outline image of the lantern); wherein, the image input device 110 includes, but is not limited to, electronic devices with electronic drawing functions such as graphics tablets, iPads, and mobile phones.

[0062] That is, the user draws the outline of the target object in the drawing area of ​​the image input device 110, so that the image input device 110 obtains the electronic outline image from the drawing area and sends the outline image to the interactive main platform 120.

[0063] Optionally, the image input device 110 and the interactive main platform 120 may not establish direct communication, but rather achieve data interaction based on an intermediate device; that is, the image input device 110 may act as an intermediate machine for receiving images, transmitting the outline image identified in the drawing area to the host computer, and the host computer may then transmit the outline image to the interactive main platform 120.

[0064] It should be noted that the communication architecture and communication lines between the image input device 110 and the interactive main platform 120 can be configured according to actual needs, as long as they can achieve dynamic interaction of light and shadow in conjunction with other parts of the interactive system.

[0065] The main interactive platform 120 receives contour images, preprocesses them to obtain masked composite images, and then combines these masked composite images with preset static ancient painting materials to obtain multi-layered images. The masked composite image includes a mask layer and a bottom layer, with the preprocessed contour image embedded within the mask layer.

[0066] Specifically, the interactive main platform 120 first preprocesses the outline image to generate a specific masked composite image, and then composites it with preset static ancient painting materials to obtain a multi-layer image. In this multi-layer image, the masked composite image and the static ancient painting materials are displayed as single layers, and each layer is independent of the others. In other words, editing the content of any layer in the multi-layer image will not affect the content of other layers, and the visibility of the content of each layer in the multi-layer image is also independent.

[0067] In one example, the interactive main platform 120 includes a LoRA image generation module and a mask compositing module. The LoRA image generation module is used to parse the contour image to form dynamically identifiable graphic boundary data, and then call a locally pre-deployed image-to-image model to perform a series of processes such as stylization to obtain a transparent background image with the target style. The mask compositing module is used to perform image masking processing on the transparent background image, and then to perform layer compositing with the preset static ancient painting material to obtain a multi-layer image.

[0068] Optionally, the main interactive platform 120 is a programming development platform with visual interactive creation capabilities. The main interactive platform 120 can be independently developed and designed or can call an existing development platform. This embodiment does not limit this; for example, the main interactive platform 120 can be the TouchDesigner platform, etc.

[0069] Furthermore, the infrared interactive device 130 is used to capture the laser point trajectory of the user's operation point and extract the laser point coordinate information in the laser point trajectory.

[0070] In one embodiment, the infrared interactive device 130 includes an infrared laser lift rod and an infrared camera; wherein, the infrared camera is fixedly installed at the target position of the video output device 140 so that the infrared signal acquisition range of the infrared camera covers the boundary area of ​​the video output device 140; the infrared camera is used to capture the trajectory of the laser point emitted by the infrared laser lift rod in real time when the user moves the infrared laser lift rod.

[0071] It is understood that the user interacts with the multi-layered image displayed on the video output device 140 by moving the infrared laser lever. The infrared laser lever emits infrared laser during the movement, and the infrared laser falls on the display screen (i.e., the image display area) of the video output device 140 to form a laser dot trajectory. This laser dot trajectory represents the user's laser pointing operation. The infrared camera is used to capture the laser dot trajectory of the infrared laser output by the infrared laser lever in real time.

[0072] It is important to note that this infrared laser positioning system has a clear and controllable signal source, ensuring accurate identification unaffected by background or obstructions. Furthermore, the fixedly installed infrared camera can capture the laser point at high frequency, achieving millisecond-level response. Consequently, the laser beam trajectory emitted from the infrared laser pole is clear and stable, allowing for precise motion control. The physical infrared laser pole enhances the user's tactile feedback and the atmosphere of the exhibition, increasing its ceremonial and participatory nature. The form of the infrared laser pole is not limited; in practical scenarios, it can be replaced with devices or apparatuses capable of emitting infrared lasers, such as laser pointers, as needed.

[0073] Furthermore, the video output device 140 is used to display multi-layer images, and binds the laser point coordinate information with the coordinate position of the mask combination image in the multi-layer images. Based on the laser point coordinate information, it controls the position change of the mask combination image in the multi-layer images to achieve dynamic visual interaction. Then, the user controls the position or rotation angle of the mask combination image in the multi-layer images by controlling the laser displacement emitted by the infrared laser lifting rod.

[0074] Optionally, the video output device 140 includes electronic devices such as a display screen or projector that can directly receive video signals and display images; for example, the video output device 140 can be a MAXHUB large screen.

[0075] For example, this embodiment uses TouchDesigner to build the main interactive platform 120, which is supplemented by hardware and software modules such as drawing electronic devices, infrared laser lifting rods, infrared cameras, and MAXHUB large screens to work together to realize a complete closed-loop process from user input to dynamic interactive output.

[0076] As an optional implementation, the system uses a sensing device instead of the infrared interaction device 130 to recognize the user's gestures. Then, during the display of multi-layered images on the video output device 140, the user can control the position of the masked composite image within the multi-layered images through gestures, achieving dynamic visual interaction. The sensing device can be an electronic device with video recording or gesture recognition capabilities, such as a camera.

[0077] This application embodiment uses the infrared light and shadow interaction system to process the outline image drawn by the user to obtain a masked combined image, and then combines it with static ancient painting materials to obtain a multi-layer image. The position of the masked combined image in the combined multi-layer image is controlled by the laser point trajectory of the user operation point captured by infrared, thereby realizing the real-time interactive control of dynamic visual effects by the user through infrared laser.

[0078] Please refer to Figure 2 This application also provides an infrared light and shadow interaction method, applied to the aforementioned infrared light and shadow interaction system, wherein, as... Figure 2 As shown, the method includes the following steps:

[0079] S210, Obtain the outline image drawn by the user.

[0080] S220: After preprocessing the outline image, a masked composite image is obtained. The masked composite image is then combined with a preset static ancient painting material to obtain a multi-layer image.

[0081] In this embodiment, the user manually draws a contour image through the image input device 110, and then the interactive main platform 120 preprocesses the contour image, including stylization processing, background transparency processing, image masking processing, etc.

[0082] As an optional implementation, the interactive main platform 120 analyzes the contour image using a preset image analysis algorithm before stylizing it. Optionally, during the analysis process, this embodiment can also use a contour detection algorithm to extract contour information from the analyzed contour image data, then select the main contour that best matches the lantern contour, and extract boundary coordinate data from the main contour. The specific operations and algorithms used in this image analysis process can be set according to actual needs, and this embodiment is not limited to any particular method.

[0083] Furthermore, the interactive main platform 120 calls the preset image-to-image model to stylize the outline image and then convert it into a transparent background image.

[0084] In one example, the open-source ComfyUI-To-TD plugin interface is built into the main interactive platform 120 to connect to the ComfyUI graphical generation platform, thereby calling the functionality of the ComfyUI graphical generation platform's local graph-generated image model (such as Stable Diffusion). Specifically, this embodiment calls the locally pre-deployed graph-generated image model to convert the outline image into a transparent background image (such as a PNG format image) with target style colored light elements.

[0085] In another example, the interactive main platform 120 can build the program framework of the ComfyUI graphical generation platform, enabling it to directly possess the image processing function of the ComfyUI graphical generation platform. Then, the interactive main platform 120 can directly call the graph model within the ComfyUI graphical generation platform framework to convert the outline image into the corresponding transparent background image.

[0086] It should be understood that this embodiment does not impose too many restrictions on the relationship or configuration between the interactive main platform 120 and the ComfyUI graphical generation platform, as long as it can combine the corresponding functions to achieve the relevant image processing described herein.

[0087] Optionally, in real-world scenarios, such as Figure 3 As shown, the user hand-draws an outline image in the image input device 110, and then the interactive main platform 120 calls the image-generated model to perform a series of processes on it (such as stylization and background transparency processing) to obtain a transparent background image with the target style. Then, the transparent background image can be synchronously displayed to the user through the image input device 110. During this process, the interactive main platform 120 also performs image masking and layer compositing processing on the transparent background image with the target style.

[0088] In a real-world scenario, the user can redraw the outline image multiple times using the image input device 110. For example, the image input device 110 recognizes the outline image within the target drawing frame on the display interface. Only after receiving the response information generated when the user triggers the "Complete" button or component on the display interface, the device saves the outline image within the current target drawing frame, uses it as the target outline image for subsequent processing, and transmits it to the interactive main platform 120.

[0089] If the image input device 110 receives a response message generated when the user triggers the "redraw" button or component on the display interface during the user's drawing process, it will clear the outline image within the current target drawing frame.

[0090] After the interactive main platform 120 processes the outline image transmitted by the image input device 110, it outputs a transparent background image with the target style to the image input device 110. Then, the image input device 110 transmits the transparent background image to the target display box on the display interface for display.

[0091] In one embodiment, the interactive main platform 120 may also transmit the masked composite image to the image input device 110 after converting the outline image into a masked composite image, so that the image input device 110 displays it within a target display frame on the display interface. That is, the object displayed within the target display frame on the display interface of the image input device 110 can be a transparent background image with a target style or a masked composite image, which is not limited in this embodiment.

[0092] As an optional implementation, the image-generated image model can also be mounted in the image input device 110. When the image input device 110 recognizes the image in the drawing area, it can call the local image-generated image model for processing, and then transmit the processed transparent background image with the target style to the interactive main platform 120 for further processing (i.e., image masking processing and layer compositing processing).

[0093] The graph-generating model can be pre-trained; in one embodiment, such as Figure 4 As shown, the generation process of this graph-generated model includes the following steps:

[0094] S410, acquire and preprocess an image sample set containing elements of the target style colored lights.

[0095] Multiple painting images containing lantern elements of the target style are processed by image matting and AI synthesis. The resulting matted sample images and AI sample images are used as an image sample set for model training. This image sample set contains multiple matted sample images and multiple AI sample images, and the number of images in the image sample set is not limited here. The specific lantern elements of the target style are set according to actual needs and are not limited thereto. This embodiment takes Bianjing lantern elements as an example. Figure 5 As shown, the painting image is an ancient painting image that includes elements of ancient Bianjing lanterns.

[0096] Furthermore, ancient paintings containing elements of ancient Bianjing lanterns, such as works with ancient court themes, were collected from image resource libraries. The necessary image data was obtained from these ancient paintings in high-resolution (at least 2K) digital format to ensure image quality and processability. Subsequently, through image annotation and segmentation processes, image editing tools (such as Photoshop) were used to precisely extract the lantern areas from the image data. After removing the background, images with a uniform white background (such as PNG format images) were generated. Then, a uniform cropping ratio and composition method were adopted to ensure that each white background image maintained target centrality and clear composition, thereby constructing an image sample set of ancient Bianjing lanterns with stylistic purity and data consistency.

[0097] In one example, since the original ancient paintings of Bianjing lanterns generally suffer from low resolution, blurred lines, and color degradation, the image sample set can be preprocessed before being used as training samples to train the model. This preprocessing systematically optimizes the image quality and improves the accuracy of subsequent model training. This preprocessing can be implemented through the AI ​​image processing module within the main interactive platform 120; it includes image enlargement, image enhancement, and image equalization, among other operations, and is not limited to these.

[0098] For example, the Real-ESRGAN deep learning super-resolution algorithm is used to enlarge each sample image in the image sample set to a preset pixel size to meet the size requirements of subsequent model training. This preset pixel size is set according to actual needs and is not limited thereto. For example, the preset pixel size is 1024×1024 pixels. Next, a preset edge enhancement algorithm is used to enhance the contours and details of each enlarged sample image to strengthen the pattern contours and structural details, improving the image's process reproduction accuracy. Then, a color equalization algorithm is used to perform color restoration processing on each sample image to automatically restore the image's hue, saturation, and contrast. Furthermore, this embodiment can also perform a series of image enhancement processes (such as rotation, mirroring, scaling, etc.) on the images after the above preprocessing to effectively improve the diversity of training samples and the model's style generalization ability.

[0099] All sample images obtained after the above processing are uniformly centered, with transparent backgrounds and a resolution of 1024×1024, to ensure style consistency and training stability.

[0100] S420 performs bilingual (Chinese and English) annotations on each image in the image sample set and embeds custom trigger words in the annotation text. The custom trigger words are bound to the style weights of the graph-to-graph model to be generated. The custom trigger words are used to quickly trigger the graph-to-graph model to call the style features of the target style.

[0101] For each sample image in the image sample set, perform Prompt keyword annotation, specifically using both Chinese and English for annotation to enhance the cross-semantic generation ability of the model. For example, the English annotation text can be: "deng, a lantern, Song Dynasty, imperial lantern, silk texture, ancient Chinese craft style, white background"; the Chinese annotation text is: "灯笼, 一个灯笼, 古代, 宫廷灯笼, 丝绸质感, 中国古代工艺风格, 白色背景". Among them, all images form one-to-one corresponding training samples with the Prompt keyword annotation text.

[0102] Next, to further improve the consistency of style activation and the inference stability of the model, a custom trigger word (such as "deng") is uniformly embedded in all Prompt keyword annotation texts. This custom keyword will be bound to the style weight of the image-to-image model during the subsequent model training process and serve as the semantic anchor point for invoking the image-to-image model in the inference stage. Furthermore, it can be used to quickly activate the specific style features of the "ancient Bianjing colored lantern style" during the process of the image-to-image model converting the user's hand-drawn contour image into a transparent background image with the ancient Bianjing colored lantern style, thereby enhancing the style recognition ability and invocation efficiency of the image-to-image model.

[0103] S430, Use the LoRA algorithm to train the Stable Diffusion model structure in combination with the image sample set to obtain an image-to-image model with the ability to transfer the target style colored lantern elements.

[0104] This embodiment is to construct an efficient, low-resource-consuming, and highly adaptable image-to-image model with style transfer ability. Specifically, based on the Stable Diffusion model structure (such as Stable Diffusion XL (SDXL) 1.0) as the model backbone architecture, the LoRA (Low-Rank Adaptation) algorithm is used for fine-tuning modeling to achieve high-fidelity transfer and image generation of the ancient Bianjing colored lantern style.

[0105] As a further example, the model training process was completed using the Kohya-SS open-source trainer, with the core parameters set as follows: Alpha value 8, learning rate 1e-4, 10 training epochs, batch size 2, AdamW optimizer selected, and fp16 half-precision mixed training enabled to save GPU memory and accelerate computational efficiency. Furthermore, the model training employed a caption+image paired input mechanism, and a weight decoupling optimization strategy was implemented between the Text Encoder and UNet modules in the model structure. This ensured that the LoRA parameters only applied to key sub-modules, thereby accurately adapting to the target style while retaining the general capabilities of the large model, ensuring high fidelity in image details, composition, and texture. The final model was exported in .safetensors format, containing only weight difference parameters, with a file size controlled within 160MB. It supports flexible deployment and multi-style fusion, making it suitable for the ComfyUI graphical generation platform, enabling efficient dynamic generation tasks with multiple styles coexisting.

[0106] In one embodiment, this embodiment also involves performance verification of the trained graph-to-graph model; specifically, after the interactive main platform 120 obtains the lantern structure outline drawn by the user through the image input device 110, it adopts the graph-to-graph mode, uses the lantern structure outline as the input image, and triggers the graph-to-graph model to generate a lantern image with lantern elements of the target style based on custom trigger words and style weights; the lantern image is compared with the lantern image in the ancient painting image corresponding to the graph-to-graph mode to obtain the performance verification result of the graph-to-graph model.

[0107] Demonstratively, after the image-to-image model is trained, it is loaded into the ComfyUI graphical generation platform, and its stylistic expressiveness is verified through an image generation workflow. In the verification process, an image-to-image mode is used, allowing users to manually draw the lantern structure outline via a contour sketching panel (or image input device 110), serving as input to guide generation. This enables the image-to-image model to combine the Prompt trigger word "deng" and style weight information to generate a complete image of ancient Bianjing lanterns. This process verifies the image-to-image model's adaptability and restoration ability to the target style under structural control, providing a stable style output foundation for the infrared light and shadow interaction system. Furthermore, the transparent background image output by the image-to-image model is compared with lantern patterns in original ancient paintings, evaluating dimensions including color reproduction, detail retention, structural symmetry, and stylistic consistency to ensure the model's output image has a high degree of cultural reproduction capability. Based on this, the model output is further integrated with the main interactive platform 120, enabling real-time transmission and retrieval of the model's output image.

[0108] Furthermore, the interactive main platform 120 converts the user's hand-drawn outline image into a transparent background image with the target style by calling the graph-to-graph model trained above. Then, it performs image masking on the transparent background image and composites it with preset static ancient painting materials to obtain a multi-layer image. Among them, the multi-layer image is a dynamic video.

[0109] Specifically, a transparent background image with the target style is overlaid onto a Circle component within the ComfyUI graphical generation platform or the ComfyUI graphical generation platform embedded in the main interactive platform 120 using a Transform component to obtain a mask image. The mask image is then multiplied and synthesized with preset AI dynamic video material to obtain a masked composite image. Finally, the masked composite image is synthesized with preset static ancient painting material to obtain a multi-layered image. The Circle component is used to draw circles and supports setting attributes such as circle size, fill color, and border color.

[0110] For example, the preset static ancient painting material can be pre-processed using AI technology, such as image restoration; the preset AI dynamic video material is a dynamic scroll video constructed using AI technology based on static ancient painting material containing lantern festival scenes (i.e., static ancient painting material). The video content includes typical scenes such as streets, bridges, and lantern festivals generated based on AI image restoration and material enhancement. Figure 6 As shown, the "Along the River During the Qingming Festival" scroll is used as an example to illustrate this static ancient painting material. The specific processing procedures involving AI technology in the static ancient painting material and the AI ​​dynamic video material can be set and processed according to actual needs; this embodiment does not limit this.

[0111] Furthermore, pre-made static ancient painting materials (i.e., static ancient painting materials) and AI dynamic video materials are first imported through the two Movie File In TOP nodes in the main interactive platform 120. The two materials are then initially combined using the Composite TOP node to provide a foundation for subsequent image masking processing.

[0112] Then, add a corresponding Circle component to the main interactive platform 120. The transparent background image with the target style output by the graph-generated image model is overlaid on the Circle component using the Transform component in an "over" manner. Increase the edge feathering value and brightness in the built-in parameters of the Circle component, and vertically and horizontally center the transparent background image and the Circle component. Combine them into one using the Transform component to output the mask image.

[0113] To define the masking area, after connecting the masking image created with the Transform component in the above steps to the AI ​​dynamic video footage, set the output parameter "Combine with Input" to "Set Resolution Only" to ensure that the output image only contains the resolution information of the masking area and does not introduce other image content.

[0114] The process of connecting the mask image to the AI ​​dynamic video footage involves using the Multiply TOP component to combine the mask image and the AI ​​dynamic video footage, performing multiplication and synthesis to obtain a masked composite image. Since the RGB values ​​of the transparent background area in the mask image are all 0, while the RGB values ​​of the circular area and the outline of the colored lights in the transparent background image are all 1, after multiplying it with the AI ​​dynamic video footage, the RGB values ​​of the transparent area remain 0, and the circular area and the outline of the colored lights are preserved in the video content, achieving the effect of displaying the video only in the circular area.

[0115] Then, the Composite TOP node is used to composite the processed Circle video masking effect (i.e., the masked composite image) with the static ancient painting material to form the final output image (i.e., the multi-layer image). In other words, the static ancient painting material is overlaid on the AI ​​dynamic video material in the Composite component to complete the infrared flashlight masking effect.

[0116] In short, the interactive main platform 120 can achieve multi-layer compositing. This multi-layer image includes multiple layers, specifically a static ancient painting layer, an AI dynamic video layer, and a user-manipulated mask layer. The lantern image in this mask layer is a masked composite image with a target style bound to the Circle component. For example... Figure 7 As shown, you can obtain an image with multiple layers superimposed (i.e., a multi-layer image).

[0117] Clearly, the final image displayed by the video output device 140 employs a dual-layer overlay mechanism. The upper layer is a static ancient painting layer, and the lower layer is an AI-generated dynamic video material. The user-controlled Circle component area (i.e., the mask layer) acts as a mask window, displaying the underlying dynamic content at its location, while the non-masked areas remain static. In other words, in the image displayed by the video output device 140, the masked composite image and the static ancient painting material are each displayed as a single layer, and each layer is independent of the others. This results in only the masked composite image portion displaying as dynamic video in the image, while the rest remains static.

[0118] Understandably, based on this infrared light and shadow interactive system, users can first draw the structural outline of the lantern, then generate a corresponding transparent background image with elements of ancient Bianjing lantern style through an image-generated model. This image is then combined with AI dynamic video through masking, and finally composited with static ancient painting materials to obtain a multi-layered image. This multi-layered image is then simultaneously displayed on the video output device 140, realizing a complete process from structural input to style generation to multimedia presentation. This system forms an end-to-end cultural content generation chain from lantern structural outline construction to real-time interactive display of image style generation, and can be widely applied in diverse application scenarios such as assisting in lantern art creation and digital interpretation of intangible cultural heritage elements.

[0119] S230: Obtain the laser point trajectory of the user's operation point and extract the laser point coordinate information from the laser point trajectory.

[0120] In this embodiment, the interactive main platform 120 captures the laser point trajectory of the user's operation point, and then binds the coordinate information of the laser point trajectory with the coordinate information of the lantern image in the multi-layer image displayed in the video output device 140, so as to realize interactive control of the lantern image.

[0121] As a further example, firstly, a real-time video stream containing the laser point trajectory of the user's operation point is acquired, and the image in the video stream is cropped to obtain the target processing area; wherein, the target processing area is an image area containing the screen area of ​​the video output device 140.

[0122] Next, the target processing area is blurred to enhance the brightness characteristics of the laser points within the target processing area. Then, threshold segmentation is performed on the blurred target processing area to separate the laser points from the target processing area, resulting in laser point image data containing only the position information of the laser points.

[0123] Then, the laser point image data is converted into geometric contour data to obtain the geometric edge information of the laser point; the geometric contour data is converted into channel data to obtain the coordinate information of the laser point from the geometric edge information, and the coordinate information is normalized to map the coordinate information from the original pixel coordinate system to the standard coordinate system to obtain the laser point coordinate information.

[0124] It is understandable that, based on the above-mentioned structure of the infrared light and shadow interaction system, in a real-world scenario, a user holds an infrared laser handle with infrared laser function and performs laser pointing operations on the video output device 140. An infrared camera deployed at a fixed position on the video output device 140 captures in real time the trajectory of the infrared laser point emitted by the infrared laser handle, which only appears within the screen range of the video output device 140.

[0125] In this process, an infrared camera is connected to the main interactive platform 120 via the Video Device In component to acquire a real-time video stream. To ensure processing efficiency and accuracy, the Crop component is used to crop the image in the real-time video stream, limiting the processing area to capture only the screen area of ​​the video output device 140. Subsequently, Blur TOP is applied to blur the image within the limited processing area to enhance the brightness characteristics of the laser points, and Threshold TOP is used for thresholding to separate the laser points from the background, resulting in an image containing only the laser point position information (i.e., laser point image data).

[0126] The image containing only laser point information is converted into geometric contour data using Trace SOP to extract the edge information of the laser points. Using SOP to CHOP, the geometric contour data is converted into channel data to obtain the horizontal and vertical coordinates (i.e., X (tx) and Y (ty) coordinates) of the laser points. The X and Y channels are selected separately using Select CHOP, and the horizontal and vertical coordinate data are normalized using Math CHOP, mapping them from the original pixel coordinate system to a normalized range of 0 to 1 to match the coordinate system of subsequent graphics. Furthermore, the normalized coordinate data can be cached using Null CHOP to ensure data stability and accessibility.

[0127] S240 binds the laser point coordinate information with the coordinate position of the masked composite image in the multi-layer image, so as to control the position change of the masked composite image in the multi-layer image based on the laser point coordinate information, and realize dynamic visual interaction.

[0128] In this embodiment, the video output device 140 maps the horizontal and vertical coordinates in the laser point coordinate information to the displacement parameters of the mask combination image in the multi-layer image, so as to bind the laser point coordinate information with the coordinate position of the mask combination image in the multi-layer image, thereby realizing the synchronization of the change in the laser point coordinate position and the change in the coordinate position of the mask combination image in the multi-layer image.

[0129] As a further example, the X and Y coordinates of the laser point after normalization in the aforementioned steps are mapped onto the displacement parameters of the mask window, respectively, to achieve real-time synchronization between the laser point position and the mask image position, allowing users to directly control the movement direction and position of the mask image in the multi-layer image through physical laser actions.

[0130] In other words, by identifying the coordinates of the laser point and binding them to the coordinates of the mask image within the mask composite image, when a change in the position of the laser point is detected, the mask image is controlled to move to the position pointed to by the laser point in the multi-layer image, thus achieving synchronous movement of the laser point and the mask image. As a result, when the user holds and moves the infrared laser lifting rod, the mask image in the multi-layer image can move with the lifting rod, completing the light and shadow interaction.

[0131] In another embodiment, the video output device 140 may only perform image display and infrared signal capture functions, while the interactive main platform 120 binds the coordinates of the laser point to the coordinates of the mask image in the masked composite image. Then, when a change in the position of the laser point is detected, the mask image is correspondingly controlled to move to the laser point's pointing position in the multi-layer image, achieving synchronous movement of the laser point and the mask image. The control effect of this process is displayed through the video output device 140. For example, the video output device 140 captures the laser point trajectory or position of the user's operation point via infrared, and then sends the laser point trajectory or position to the interactive main platform 120 in real time. The interactive main platform 120 identifies the coordinates of the laser point within the display area of ​​the video output device 140 based on the received laser point trajectory or position, and then binds the laser point coordinates to the coordinates of the mask image in the masked composite image. Based on the received laser point position, when the position of the laser point changes, the video output device 140 correspondingly controls the mask image to move to the laser point's pointing position in the multi-layer image.

[0132] That is, both the video output device 140 and the interactive main platform 120 can realize the binding process of the laser point position coordinates and the position coordinates of the mask image. The specific device executing this process is not limited here.

[0133] In real-world scenarios, users can control the movement of lantern images within multiple image feeds on the video output device 140 by moving the infrared laser lever, thereby achieving a long-distance light and shadow interaction effect between the infrared laser and the lantern images.

[0134] It is understandable that users input images by hand-drawing lantern outlines. The interactive system then uses an image-based model to stylize and generate a transparent background image with the style of ancient Bianjing lanterns. This image is then masked and layered to produce a multi-layered image. An infrared laser device enables real-time control of the masked image's movement and direction within the multi-layered image. Finally, the main interactive platform 120 achieves multi-layered compositing and dynamic visual interaction, thus realizing a closed-loop process from user input to image generation and dynamic display. This enhances the freedom of interaction, visual expressiveness, and cultural immersion. Users experience an immersive cultural journey from "hand-drawing lanterns" to "laser-controlled lantern movement" to "illuminating a dynamic nighttime scene," constructing an innovative fusion of traditional aesthetics and digital image generation.

[0135] Among them, by integrating infrared laser interaction technology with a stable diffusion image generation model with low-rank adaptation (LoRA) fine-tuning, an end-to-end system process was realized, from user hand-drawn input to personalized image generation, and then to multi-layer dynamic synthesis and display. This breakthrough overcomes the limitations of existing technologies in terms of interaction depth, freedom of image content generation, and dynamic visual feedback, and achieves higher freedom of content creation, a more immersive visual experience, and more accurate physical interaction response.

[0136] Furthermore, compared to traditional infrared interactive systems that can only trigger preset image content, users can only perform simple response operations (such as zooming in, replacing, etc.) within limited projected content, lacking the ability to create content based on the user's subjective intentions. This application's embodiment introduces a user-defined drawing input and real-time image model generation mechanism at the interactive front end, allowing users to input lantern outline structure information through freehand drawing, combined with a Prompt keyword triggering mechanism, to achieve proactive definition and precise guidance of image content. User input is not limited to geometric outlines but can also include semantic tags, participating in the content control process before image generation, greatly enhancing the creative expression space and controllability of the interaction, possessing high personalization, unpredictability, and creative freedom; through this "creation-generation-fusion" mechanism, a content closed loop from user input to image output is achieved, significantly improving the system's generation dynamism and interactive openness.

[0137] Secondly, in terms of image control methods, compared to existing technologies that only use prompts as input parameters and lack user behavior dimensions in their image control paths, this application combines the structural guidance mechanism of ControlNet to generate artistic images that combine ancient aesthetics with modern graphic quality based on the lantern outline drawn by the user, taking into account both the flexibility of hand-drawing and the quality stability of AI generation.

[0138] Furthermore, traditional image generation systems are mostly limited to single-image display or offline output, lacking deep integration with real-time interactive devices, especially in dynamic multi-layer compositing. This application achieves real-time binding and layer compositing of generated images and dynamic video materials through an interactive main platform 120, constructing a three-layer structure of static ancient painting - interactive mask window - underlying dynamic video, thus forming a composite visual effect of interactive triggering - local dynamic visibility - overall static control, enhancing the correlation between visual rhythm and user feedback. This compositing logic enables the dynamic introduction of generated content into fixed scenes. This control method not only improves the spatial accuracy of interaction but also enhances the user's subjective sense of control through the binding between physical behavior and image movement, giving the display effect immersive, responsive, and multi-dimensional characteristics, with significant advantages in cultural exhibitions, public art, and interactive entertainment scenarios.

[0139] Furthermore, in terms of visual content construction, this application uses AI modeling and physical rendering methods to digitally reconstruct static ancient painting scenes to obtain AI dynamic video materials. These AI dynamic video materials not only restore spatial structures such as Bianjing street scenes, water systems, and pavilions, but also realistically reconstruct lantern styles, lighting logic, and environmental dynamic effects based on material properties. Details such as the swaying of lanterns, the flickering of candlelight, and the refraction of water ripples enhance the realism and artistry of cultural expression, enabling the entire system to achieve a high degree of integration between cultural reproduction and technological realization. This solves the problems of static images lacking dynamic expression and interactive content lacking cultural accumulation in existing technologies.

[0140] Based on this, this application has made systematic improvements in multiple dimensions, which not only enhances the interactive capabilities and content customization of AI image generation, but also strengthens the real-time linkage between user participation and image visual output, ultimately achieving significant technical improvements in cultural expression effects, interactive experience methods, and visual presentation quality.

[0141] This application also provides a computer device, exemplary of which includes a processor and a memory, wherein the memory stores a computer program, and the processor executes the computer device to perform the above-described infrared light and shadow interaction method by running the computer program.

[0142] The processor can be an integrated circuit chip with signal processing capabilities. The processor can be a general-purpose processor, including at least one of a Central Processing Unit (CPU), Graphics Processing Unit (GPU), Network Processor (NP), Digital Signal Processor (DSP), Application-Specific Integrated Circuit (ASIC), Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The general-purpose processor can be a microprocessor or any conventional processor, capable of implementing or executing the methods, steps, and logic block diagrams disclosed in the embodiments of this application.

[0143] The memory can be, but is not limited to, Random Access Memory (RAM), Read Only Memory (ROM), Programmable Read-Only Memory (PROM), Erasable Programmable Read-Only Memory (EPROM), Electrically Erasable Programmable Read-Only Memory (EEPROM), etc. The memory is used to store computer programs, and the processor can execute the computer programs accordingly after receiving execution instructions.

[0144] This application also provides a computer storage medium for storing the computer program used in the aforementioned computer device. The computer storage medium can be a readable storage medium, a non-volatile storage medium, or a volatile storage medium. For example, the computer storage medium may include, but is not limited to, various media capable of storing program code, such as a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0145] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can also be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the architecture, functionality, and operation of possible implementations of apparatus, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that, in alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0146] In addition, the functional modules or units in the various embodiments of this application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.

[0147] If the aforementioned functions are implemented as software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a smartphone, personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application.

[0148] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.

Claims

1. An infrared light and shadow interactive system, characterized in that, include: Image input device for acquiring a user-drawn outline image; The interactive main platform is used to receive the outline image sent by the image input device, preprocess the outline image to obtain a masked composite image, and composite the masked composite image with the preset static ancient painting material to obtain a multi-layer image; the masked composite image and the static ancient painting material in the multi-layer image are displayed as single layers, and each layer is independent of the others. An infrared interactive device is used to capture the laser point trajectory of a user's operation point and extract the laser point coordinate information in the laser point trajectory; A video output device is used to display the multi-layer image, and binds the laser point coordinate information with the coordinate position of the mask combination image in the multi-layer image, so as to control the position change of the mask combination image in the multi-layer image based on the laser point coordinate information, thereby realizing dynamic visual interaction. The step of capturing the laser point trajectory of the user's operation point and extracting the laser point coordinate information from the laser point trajectory includes: A real-time video stream containing the laser point trajectory of the user's operation points is acquired, and the image in the video stream is cropped to obtain the target processing area; the target processing area is an image area that includes the screen area of ​​the video output device. The target processing area is blurred to enhance the brightness features of the laser points within the target processing area. Threshold segmentation is then performed on the blurred target processing area to separate the laser points from the target processing area, resulting in laser point image data containing only the position information of the laser points. The laser point image data is converted into geometric contour data to obtain the geometric edge information of the laser point; The geometric contour data is converted into channel data to obtain the coordinate information of the laser point from the geometric edge information. The coordinate information is then normalized to map the coordinate information from the original pixel coordinate system to the standard coordinate system, thus obtaining the laser point coordinate information.

2. The infrared light and shadow interactive system according to claim 1, characterized in that, The infrared interactive device includes an infrared laser lifting rod and an infrared camera; wherein, the infrared camera is fixedly installed at the target position of the video output device so that the infrared signal acquisition range of the infrared camera covers the boundary area of ​​the video output device; The infrared camera is used to capture the trajectory of the laser point emitted by the infrared laser lifting rod in real time when the user moves the infrared laser lifting rod.

3. The infrared light and shadow interactive system according to claim 1, characterized in that, The process of preprocessing the contour image to obtain the masked composite image includes: The preset image model is invoked to stylize the outline image, resulting in a lantern image with the target style; Convert the lantern image with the target style into a transparent background image; The transparent background image is subjected to image masking processing to obtain a masked composite image.

4. The infrared light and shadow interactive system according to claim 3, characterized in that, The process of generating the graph-generated model includes: Acquire and preprocess a set of image samples containing elements of the target style colored lights; Each image in the image sample set is annotated in both Chinese and English, and a custom trigger word is embedded in the annotation text. The custom trigger word is bound to the style weight of the graph-generated graph model to be generated. The custom trigger word is used to quickly trigger the graph-generated graph model to call the style features of the target style. The Stable Diffusion model structure was trained using the LoRA algorithm in conjunction with the image sample set to obtain a graph-to-graph model with the ability to transfer elements of the target style lanterns.

5. The infrared light and shadow interactive system according to claim 4, characterized in that, The acquisition and preprocessing of the image sample set containing the target style colored light elements includes: Multiple ancient painting images containing elements of the target style lanterns were processed by image matting and AI synthesis. The resulting matted sample images and AI sample images were used as the image sample set for model training. The Real-ESRGAN deep learning super-resolution algorithm is used to uniformly enlarge each sample image in the image sample set to a preset pixel size; A preset edge enhancement algorithm is used to enhance the contours and details of each magnified sample image, and a color equalization algorithm is used to restore the color of each sample image.

6. The infrared light and shadow interactive system according to claim 4 or 5, characterized in that, The main interactive platform is also used for: The performance of the trained graph-generated graph model is validated, specifically including: Obtain the lantern structure outline drawn by the user, and use the image-generated image mode to take the lantern structure outline as the input image to trigger the image-generated image model to generate a lantern image with lantern elements of the target style based on the custom trigger words and the style weights; The lantern image is compared with the lantern image in the ancient painting corresponding to the image-generated image mode to obtain the performance verification result of the image-generated image model.

7. The infrared light and shadow interactive system according to claim 3, characterized in that, The step of performing image masking processing on the transparent background image to obtain a masked composite image includes: The transparent background image is overlaid onto a preset Circle component using the Transform component to obtain a mask image; The masked image is multiplied and synthesized with preset AI dynamic video material to obtain a masked composite image.

8. The infrared light and shadow interactive system according to claim 3, characterized in that, The method of controlling the position change of the masked composite image in the multi-layer image based on the laser point coordinate information includes: The horizontal and vertical coordinates in the laser point coordinate information are mapped to the displacement parameters of the masked combined image in the multi-layer image, so as to achieve synchronization between the change of the laser point coordinate position and the change of the coordinate position of the masked combined image in the multi-layer image.

9. An infrared light and shadow interaction method, characterized in that, The method, applied to the infrared light and shadow interactive system as described in any one of claims 1-8, comprises: Obtain the outline image drawn by the user; After preprocessing the outline image, a masked composite image is obtained. The masked composite image is then combined with a preset static ancient painting material to obtain a multi-layer image. The masked composite image and the static ancient painting material in the multi-layer image are displayed as single layers, and each layer is independent of the others. Obtain the laser point trajectory of the user's operation point, and extract the laser point coordinate information from the laser point trajectory; The laser point coordinate information is bound to the coordinate position of the masked combined image in the multi-layer image, so as to control the position change of the masked combined image in the multi-layer image based on the laser point coordinate information, thereby realizing dynamic visual interaction. The step of acquiring the laser point trajectory of the user's operation point and extracting the laser point coordinate information from the laser point trajectory includes: A real-time video stream containing the laser point trajectory of the user's operation points is acquired, and the image in the video stream is cropped to obtain the target processing area; the target processing area is an image area that includes the screen area of ​​the video output device. The target processing area is blurred to enhance the brightness features of the laser points within the target processing area. Threshold segmentation is then performed on the blurred target processing area to separate the laser points from the target processing area, resulting in laser point image data containing only the position information of the laser points. The laser point image data is converted into geometric contour data to obtain the geometric edge information of the laser point; The geometric contour data is converted into channel data to obtain the coordinate information of the laser point from the geometric edge information. The coordinate information is then normalized to map the coordinate information from the original pixel coordinate system to the standard coordinate system, thus obtaining the laser point coordinate information.

Citation Information

Patent Citations

  • A large-screen interaction system based on laser radar positioning

    CN109828695A

  • Hybrid interactive virtual environment system and method of use thereof

    WO2024199422A1