A cultural interaction device and method based on artificial intelligence and multi-modal perception
By using a cultural interactive device based on artificial intelligence and multimodal perception, combined with rotating mechanical devices and multimodal perception technology, real-time linkage between user input and dynamic content is achieved, solving the problem of insufficient dynamic adjustment in existing technologies and enhancing the user's immersive cultural experience and interactivity.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SOUTHERN UNIVERSITY OF SCIENCE AND TECHNOLOGY
- Filing Date
- 2025-02-28
- Publication Date
- 2026-05-01
AI Technical Summary
Existing cultural interactive devices are difficult to dynamically adjust through user input, resulting in an inability to meet diverse interactive needs and insufficient real-time performance and immersion.
The cultural interaction device, based on artificial intelligence and multimodal perception, captures users' visual data and generates and displays personalized dynamic text and image information through the collaborative work of multimodal perception units, generative content processing units and physical display units, and achieves real-time linkage with rotating mechanical devices.
It enhances the user's immersive cultural interaction experience, strengthens the fun and interactivity of traditional culture, reduces the complexity of interaction, provides a natural interaction method, and enhances the user's sense of participation and immersion.
Smart Images

Figure CN120143981B_ABST
Abstract
Description
A cultural interaction device and method based on artificial intelligence and multimodal perception Technical Field
[0001] This application relates to the field of human-computer interaction technology, and in particular to a cultural interaction device and method based on artificial intelligence and multimodal perception. Background Technology
[0002] While numerous digital-based methods of cultural dissemination exist, these interactive systems often present high learning curves. Complex interfaces can overload users' cognitive load, limiting natural user participation and diminishing the immersive and engaging experience of cultural transmission. The revolving lantern, a classic handicraft, cleverly combines structure and light and shadow art. Driven by hot airflow, it rotates a paddlewheel, projecting dynamic images onto the lantern's surface, creating flowing visual art. Its structure includes key components such as a paper lantern, paper-cutting, and a paddlewheel, relying on the hot airflow generated by the candlelight to display the dynamic images. This traditional installation not only possesses profound cultural significance but also provides unique inspiration for modern light and shadow art and mechanical device design. In recent years, with the application of digital technology, more and more innovative installations have attempted to combine traditional cultural elements with modern technology, expanding their application value in the fields of art display and dissemination. However, these studies largely focus on replicating appearance or craftsmanship, lacking systematic exploration of the deep integration of physical intelligent interaction and cultural experience.
[0003] Digital craftsmanship, by combining digital materials, electronic components, and computer-aided design techniques, has given traditional handicrafts new creative possibilities, such as parametric metal art based on growth logic and interactive silverware. This approach not only enriches the sensory experience of the material carrier but also stimulates emotional resonance and cultural cognition through the dynamic expression of cultural symbols. For example, triggering the sounds of artisans at work enhances cultural imagination, and textile logic gates stimulate thinking about alternatives to computational technologies. In recent years, multimodal interactive systems have developed rapidly in the fields of artificial intelligence and digital technology. These systems rely on computer vision and deep learning algorithms to capture user input (such as facial expressions and gestures) and generate matching digital content. However, most research focuses only on the single domain of image or text generation, lacking deep interaction between physical entities and digital content. For example, some studies have attempted to project user-generated images onto rotating surfaces, but due to insufficient mapping rigidity and real-time feedback, the user experience is stiff and lacks fluidity. Furthermore, dynamic mechanical devices are usually only used as passive display mediums, making it difficult to achieve real-time interaction between user input and content generation, thus weakening the emotional expression and immersive experience of the interaction.
[0004] Existing intelligent interactive designs incorporating cultural elements also have limitations, often remaining at the static presentation of traditional symbols and failing to delve deeper into cultural connotations and interactive potential. For example, dynamic video installations based on "carousels" mostly adopt a rotating display mode with fixed content, making it difficult to achieve dynamic adjustments through user input. This fails to meet diverse interactive needs and cannot convey deeper cultural values and emotional experiences.
[0005] Therefore, existing technologies still need to be improved and developed. Summary of the Invention
[0006] The main purpose of this application is to provide a cultural interaction device and method based on artificial intelligence and multimodal perception, which aims to solve the problem that existing cultural interaction devices are difficult to dynamically adjust through user input, thus failing to meet diverse interaction needs.
[0007] A first aspect of this application provides a cultural interaction device based on artificial intelligence and multimodal perception. The device includes a multimodal perception unit, a generative content processing unit, a physical display unit, and a control system. The multimodal perception unit is connected to the generative content processing unit, which is connected to the physical display unit via the control system. The multimodal perception unit acquires visual data from a user and sends it to the generative content processing unit. The generative content processing unit processes the visual data to generate dynamic text-image information for the user and sends it to the physical display unit via the control system. The physical display unit displays the dynamic text-image information to the user.
[0008] Optionally, in one embodiment of this application, the multimodal sensing unit includes a lantern shell and a camera, the camera being disposed on the lantern shell; the generative content processing unit includes a data processing module, a digital content generation module, and a content storage and update module connected in sequence; the physical display unit includes a screen controller and a flexible screen connected in sequence, the content storage and update module being connected to the screen controller; the control system includes a power management module, a data transmission module, and a motor control module, the power management module being connected to the camera, the screen controller, the flexible screen, and the motor control module respectively, the camera being connected to the data processing module through the data transmission module, and the motor control module being used to control the rotation of the flexible screen.
[0009] Optionally, in one embodiment of this application, the physical display unit further includes a rotating mechanism, the flexible screen is connected to the rotating mechanism, the motor control module is connected to the rotating mechanism, and the rotating mechanism is used to drive the flexible screen to rotate.
[0010] Optionally, in one embodiment of this application, the rotating mechanism includes a connecting plate and a DC motor, a gear set, and a curved screen bracket disposed on the connecting plate. The connecting plate is disposed on the lantern housing, the DC motor is connected to the gear set, and the flexible screen is disposed in the curved screen bracket.
[0011] Optionally, in one embodiment of this application, multiple flexible screens are provided, and the multiple flexible screens are arranged around the lantern shell to display information to the user.
[0012] Optionally, in one embodiment of this application, the motor control module is a drive board, the drive board is disposed on the connection board, and the DC motor is connected to the drive board.
[0013] A second aspect of this application also provides an interaction method for a cultural interaction device based on artificial intelligence and multimodal perception, as described in any of the above-mentioned solutions. The interaction method includes: the multimodal perception unit acquiring visual data from a user and sending it to the generative content processing unit; the generative content processing unit processing the visual data to generate dynamic text image information from the user and sending it to the entity display unit through the control system; and the entity display unit displaying the dynamic text image information to the user.
[0014] Optionally, in one embodiment of this application, the visual data includes posture features and facial features; the multimodal perception unit acquires the user's visual data specifically by: responding to the user's trigger action, the camera collects the user's posture features and facial features after the user completes the trigger action.
[0015] Optionally, in one embodiment of this application, the dynamic text image information includes cultural poems and artistic images; the generative content processing unit processes the visual data to generate the user's dynamic text image information, specifically including: the data processing module receiving the user's posture features and appearance features, and preprocessing the posture features and appearance features to obtain key features; the digital content generation module generating the cultural poems and artistic images that match the user based on the key features.
[0016] Optionally, in one embodiment of this application, the entity display unit displays the dynamic text and image information to the user, specifically: the screen controller displays the cultural poem and the artistic image on the flexible screen to the user.
[0017] Beneficial effects: This application provides a cultural interaction device and method based on artificial intelligence and multimodal perception. Through the cooperation of a multimodal perception unit, a generative content processing unit and a physical display unit, this application realizes the capture, processing and analysis of user input features, as well as the generation and dynamic display of personalized digital content, providing users with an immersive cultural interaction experience. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 is an exploded view of the overall structure of a preferred embodiment of the cultural interaction device based on artificial intelligence and multimodal perception according to this application;
[0020] Figure 2 is a structural diagram of the rotating mechanism in a preferred embodiment of the cultural interaction device based on artificial intelligence and multimodal perception of this application;
[0021] Figure 3 is a circuit diagram of a preferred embodiment of the cultural interaction device based on artificial intelligence and multimodal perception according to this application;
[0022] Figure 4 is a flowchart of a preferred embodiment of the interaction method of the cultural interaction device based on artificial intelligence and multimodal perception of this application;
[0023] Figure 5 is an interaction flowchart in a preferred embodiment of the interaction method of the cultural interaction device based on artificial intelligence and multimodal perception of this application;
[0024] Figure 6 is the interaction system flow in a preferred embodiment of the cultural interaction device based on artificial intelligence and multimodal perception of this application;
[0025] Figure 7 shows a multi-agent AI workflow in a preferred embodiment of the cultural interaction device based on artificial intelligence and multimodal perception according to this application.
[0026] Explanation of reference numerals in the attached figures:
[0027] 1. Insulation layer; 2. Slip ring; 3. Fixing block; 4. Curved screen bracket; 5. Clamp; 6. Ball bearing; 7. Connecting plate; 8. Coupling; 9. Large gear; 10. Small gear; 11. DC motor; 12. PCB board; 13. Lantern shell; 14. Display frame.
[0028] The accompanying drawings have illustrated specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to specific embodiments. Detailed Implementation
[0029] To make the objectives, technical solutions, and effects of this application clearer and more explicit, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. The described embodiments are only possible technical implementations of this application and not all possible implementations. Based on the embodiments in this application, those skilled in the art can obtain other embodiments without creative effort, and these embodiments are also within the protection scope of this application.
[0030] Existing systems also suffer from significant shortcomings in real-time performance and immersion. Many mechanical motion-based interactive devices have failed to address the latency issues between input capture and output response, resulting in a lack of consistency between generated content and physical dynamics. While some research has attempted to optimize response speed using more efficient algorithms, no breakthroughs have been achieved in the deep integration of dynamic devices and generated content. In short, existing technologies have significant limitations in terms of the depth, flexibility, and emotional experience of combining dynamic physical entities with intelligent interactive systems. Therefore, exploring the real-time linkage between dynamic mechanical structures and user input, and expanding the interactive capabilities of traditional cultural elements through intelligent technologies, has become an important research direction for intelligent cultural interactive devices. Multimodal interactive systems have gradually become an important research direction in the fields of artificial intelligence and digital technology in recent years.
[0031] The following describes the terms used in the embodiments of this application:
[0032] LLMs: Large Language Models;
[0033] FOC: Field-Oriented Control;
[0034] HDMI: High-Definition Multimedia Interface;
[0035] OpenCV: Open Source Computer Vision Library;
[0036] ControlNet: A posture control module or technology for posture recognition and posture control;
[0037] ZhipuAI: A generative AI model;
[0038] Stable Diffusion: An image generation model or technique used to generate high-quality artistic images;
[0039] Remove BG: An image background removal technique or tool used to optimize image backgrounds.
[0040] This interactive device, by integrating dynamic capture of user input, multimodal generation algorithms, and high-precision physical output, provides a new technological path for real-time linkage and immersive interaction. This fusion not only more naturally combines user input with dynamic output but also endows the device with higher emotional expression capabilities and cultural transmission potential, laying an important foundation for innovation in the field of human-computer interaction. By introducing more intelligent technologies, the interactive device optimizes the interaction method, making it more intuitive and natural, and reducing the psychological pressure on users during operation. Combined with generative artificial intelligence, the device can not only dynamically generate personalized artistic content but also support co-creation between users and the device, thereby significantly enhancing participation and emotional connection. Furthermore, the physical interactive design retains the classic viewing form of the traditional "carousel," achieving an innovative fusion between modern digital technology and traditional culture.
[0041] The following description, with reference to the accompanying drawings, describes a cultural interaction device and method based on artificial intelligence and multimodal perception, according to embodiments of this application. Addressing the problem in the aforementioned related technologies where interactive devices struggle to dynamically adjust based on user input, thus failing to meet diverse interactive needs, this application provides a cultural interaction device based on artificial intelligence and multimodal perception. In this method, the cooperation of a multimodal perception unit, a generative content processing unit, and a physical display unit enables the capture, processing, and analysis of user input features, as well as the generation and dynamic display of personalized digital content, providing users with an immersive cultural interaction experience. This solves the technical problem in the related technologies where interactive devices struggle to dynamically adjust based on user input, thus failing to meet diverse interactive needs.
[0042] The technical solutions of this application will be described in detail below with specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.
[0043] As shown in Figures 1 and 2, this application embodiment provides a cultural interaction device based on artificial intelligence and multimodal perception. The device includes a multimodal perception unit, a generative content processing unit, a physical display unit, and a control system. The multimodal perception unit is connected to the generative content processing unit, which is connected to the physical display unit through the control system. The multimodal perception unit acquires the user's visual data and sends it to the generative content processing unit. The generative content processing unit processes the visual data to generate dynamic text image information for the user and sends it to the physical display unit through the control system. The physical display unit displays the dynamic text image information to the user.
[0044] This application addresses the shortcomings of existing technologies in human-computer interaction, deep integration of cultural dissemination, and digital transmission of intangible cultural heritage. Inspired by the structure of the traditional intangible cultural heritage "revolving lantern," it combines a cylindrical flexible LED screen and a rotating mechanical device. Through the collaborative work of a multimodal perception unit, a generative content processing unit, and a physical display unit, it achieves real-time interaction between user input and dynamic digital content. By capturing the user's body posture and facial features using posture recognition and deep learning algorithms, and combining semantic extraction and Prompt optimization techniques, it generates highly personalized cultural content, including artistic poems and paintings with their cultural interpretations. The generated content is dynamically presented through a high-precision LED screen and a rotating mechanical device, which not only optimizes the quality of content generation and the relevance of user input but also significantly enhances the immersiveness of the interactive experience and the modern expression of cultural dissemination.
[0045] It should be noted that this application is based on multimodal intelligent agent technology. By integrating computer vision algorithms and large language models (LLMs), it extracts features from visual data such as user input posture and appearance to generate highly matching poetry and image content. Through the collaborative work of Prompt optimization technology and deep learning models, it ensures that the generated content is highly correlated with user characteristics in terms of semantics and style. The generated content is automatically identified by AI technology, and then separated from the background layer for seamless connection and dynamic adaptation on a flexible display screen. The interactive process of this application's cultural interactive device ensures smooth interaction and accurate feedback through multiple user data collections and real-time updates of generated content. By collecting user posture information in stages and dynamically displaying it in conjunction with the generated poetry and painting content, users can obtain a personalized and contextualized visual and auditory interactive experience in a short time. Furthermore, the design of this cultural interactive device combines the rotating dynamic structure of traditional lanterns and the shape features of the "pavilion" in garden architecture, integrating cultural imagery into the dynamic device's expression to enhance the user's immersive experience and emotional expression.
[0046] Understandably, this application addresses the limitations of existing technologies in terms of the linkage between dynamic physical devices and user interaction, the matching of personalized generated content, and emotional expression. By integrating visual data such as user-input postures and expressions with the dynamic display of generated content, it explores a possible path to enhance the human-computer interaction experience. Through the combination of intelligent technology and dynamic display, it achieves the following technical effects: enhanced cultural experience: optimizing the user's immersive experience through multimodal interaction technology, enhancing the fun and interactivity of traditional cultural dissemination; empowering creative participation: combining generative AI models, enabling users to directly participate in content creation, enhancing the personalization of cultural activities; natural interaction and immersive experience: by optimizing the linkage mechanism between user input and device output, it significantly reduces the complexity of interaction, providing a more natural interaction method, enhancing the user's immersion and enjoyment; wide application scenarios: suitable for exhibition spaces, public cultural venues, and educational environments, promoting the innovative dissemination and popularization of intangible cultural heritage.
[0047] In one embodiment of this application, as shown in Figures 1 and 2, the multimodal sensing unit includes a lantern shell 13 and a camera, the camera being disposed on the lantern shell 13; the generative content processing unit includes a data processing module, a digital content generation module, and a content storage and update module connected in sequence; the physical display unit includes a screen controller and a flexible screen connected in sequence, the content storage and update module being connected to the screen controller; the control system includes a power management module, a data transmission module, and a motor control module, the power management module being connected to the camera, the screen controller, the flexible screen, and the motor control module respectively, the camera being connected to the data processing module through the data transmission module, and the motor control module being used to control the rotation of the flexible screen.
[0048] Specifically, cameras are used to capture user input features such as posture and appearance; sensors (such as infrared sensors and accelerometers) are used to assist in capturing user movement or location information. Real-time multimodal information such as user posture and expression is captured through cameras or sensors, providing a data foundation for subsequent content generation. The data processing module receives and preprocesses data from the multimodal perception unit. AI models (such as OpenCV visual algorithms and large language models LLMs) analyze the processed data and generate personalized digital content (such as poetry, paintings, and their cultural interpretations). The content storage and update module stores the generated digital content and updates or adjusts it as needed. By analyzing and processing the data captured by the multimodal perception unit, personalized digital content matching user characteristics is generated, and this generated content is stored and managed to ensure its accuracy and real-time performance during display. Flexible LED screens are used to dynamically display the generated digital content, and the screen controller is responsible for controlling the screen's display content and synchronization. The digital content generated by the generative content processing unit is presented to the user in a visual form, achieving an immersive experience. Synchronous display of content across multiple screens is ensured to improve visual expressiveness. The power management module is responsible for converting external power into the voltage required by the internal components of the device. The data transmission module is responsible for data communication and transmission between the components. The motor control module (such as the FOC drive board) is used to control the rotation effect of the device to ensure the smoothness of the dynamic display; to provide a stable power supply to ensure the normal operation of the internal components of the device, to achieve efficient data transmission and synchronization, to ensure the collaborative work between the components, to control the rotation effect of the device, and to enhance the immersiveness and visual impact of the dynamic display.
[0049] In one embodiment of this application, the physical display unit further includes a rotating mechanism, the flexible screen is connected to the rotating mechanism, the motor control module is connected to the rotating mechanism, and the rotating mechanism drives the flexible screen to rotate.
[0050] In one embodiment of this application, as shown in FIG2, the rotating mechanism includes a connecting plate 7 and a DC motor 11, a gear set, and a curved screen bracket 4 disposed on the connecting plate 7. The connecting plate 7 is disposed on the lantern shell 13, the DC motor 11 is connected to the gear set, and the flexible screen is disposed in the curved screen bracket 4.
[0051] In one embodiment of this application, multiple flexible screens are provided, and the multiple flexible screens are arranged around the lantern shell 12 to display information to the user.
[0052] In one embodiment of this application, as shown in FIG2, the motor control module is a drive board, the drive board is disposed on the connecting plate 7, and the DC motor 11 is connected to the drive board.
[0053] Specifically, a self-designed Field-Oriented Control (FOC) driver board, in conjunction with a brushless DC motor 11, is used to achieve precise control of the rotating frame. A flexible LED screen serves as the content display medium, dynamically presenting images and text content generated based on user input, synchronized in real time with the mechanical movement.
[0054] As shown in Figure 1, the screen display uses six flexible LED screens (i.e., flexible screens) to simulate the traditional "carousel" effect. Its screen design and display functions integrate high-precision hardware layout and advanced synchronization technology. The screen layout supports flexible bending to fit the cylindrical frame and is fixed magnetically. The main frame uses insulating materials to ensure operational safety and structural stability. The display content is divided and adapted to the six screens by the screen controller. Dynamic content, including traditional patterns, poems, and animated characters, is developed using the TouchDesigner engine to convey rich cultural and visual information. Multi-screen synchronization technology achieves precise multi-screen output through the controller, ensuring smooth and continuous display of content during rotation, further optimizing visual performance.
[0055] As shown in Figure 2, the rotation effect is achieved as follows: To realize the rotation display effect, a self-designed Field-Oriented Control (FOC) driver board is used, which controls the dynamic response of the rotation mechanism through a brushless DC motor 11. The FOC driver board has advanced dynamic response adjustment function, which can optimize the motor speed and torque parameters according to the load conditions, ensuring operating efficiency and stability under different usage scenarios. The onboard SPI interface supports communication with the TouchDesigner engine, realizing real-time adjustment and optimization of key parameters such as motor speed and torque. The mechanical part adopts a precision gear transmission structure. The motor is coupled to the frame's large gear 9 through the small gear 10, driving the smooth rotation of the display frame 14. This rotation drive module provides high-precision rotation control through FOC technology, significantly improving the system's real-time performance and smoothness, and providing a reliable technical guarantee for the stable presentation of dynamic display content.
[0056] Understandably, current devices capture user posture and facial features using WiFi cameras and computer vision algorithms (such as ControlNet). Possible alternatives would be to use depth sensors (such as Kinect), LiDAR, or other biometric devices (such as heart rate monitoring or EEG brainwave detection) to enhance the diversity of user input.
[0057] Further, referring to Figures 1 and 2, the lantern shell 13 includes a display frame 14 and a cover. The rotating mechanism is located in the cover and the display frame 14. The curved screen bracket 4 is connected to the connecting plate 7. The connecting plate 7 is set inside the lantern shell 13, and the curved screen bracket 4 is connected to the display frame 14. The connecting plate 7 is provided with clamps 5 on the top and bottom. The connecting plate 7 is rotatably connected to the rotating shaft through ball bearings 6. The rotating shaft is fixedly connected to the curved screen bracket 4. The curved screen bracket 4 is connected with a fixing block 3, a slip ring 2, and an insulating layer 1. The insulating layer 1 is connected to the cover. A coupling 8 is set below the connecting plate 7. The coupling 8 is connected to the large gear 9 in the gear set. Below the connecting plate 7 is a PCB board 12 (i.e., a drive board). The PCB board 12 is connected to the DC motor 11. The DC motor 11 is connected to the small gear 10 in the gear set. The small gear 10 is meshed with the large gear 9.
[0058] In this embodiment, as shown in Figure 3, the circuit connection design is as follows: The internal circuit design of the device adopts a modular structure, combined with efficient power management and real-time data transmission strategies to ensure stable system operation and collaborative performance. The power management module converts 220V AC to 12V DC through a transformer to provide stable power to the host and screen controller; subsequently, a secondary transformer further reduces the 12V to 5V to power the LED flexible screen, thereby meeting the voltage requirements of different components. The data transmission module connects to each LED screen through the screen controller, uses data cables for signal transmission, and achieves high-speed communication with the host through an HDMI interface to ensure synchronous display of content across multiple screens. The camera power supply and communication module is driven by an independent 12V battery and establishes a data transmission link with the host via WiFi to achieve real-time acquisition and efficient transmission of attitude data. The optimized circuit design effectively improves the real-time performance of data transmission and the stability of system operation, providing reliable technical support for the collaborative work of multiple modules.
[0059] This application utilizes a multimodal artificial intelligence model to capture user input features such as posture and facial expressions, generating personalized digital content (such as poems, paintings, and their cultural interpretations), which is then presented immersively through a dynamic display device. The system adopts a modular design, including multi-layered collaborative optimization of mechanical structure, circuit design, generation algorithms, and user interaction processes, ensuring efficient operation of the device and an immersive user experience.
[0060] Based on the above embodiments, this application also provides an interaction method for the cultural interaction device based on artificial intelligence and multimodal perception, as shown in FIG4. The interaction method includes the following steps:
[0061] In step S101, the multimodal perception unit acquires the user's visual data and sends it to the generative content processing unit;
[0062] In one possible implementation, the visual data includes posture features and facial features; in response to the user's trigger action, the camera acquires the user's posture features and facial features after the user completes the trigger action.
[0063] In step S102, the generative content processing unit processes the visual data to generate the user's dynamic text image information, and sends it to the physical display unit through the control system.
[0064] In one possible implementation, the dynamic text image information includes cultural poems and artistic images; the data processing module receives the user's posture features and appearance features, and preprocesses the posture features and appearance features to obtain key features; the digital content generation module generates the cultural poems and artistic images that match the user based on the key features.
[0065] In step S103, the entity display unit displays the dynamic text image information to the user.
[0066] In one possible implementation, the screen controller displays the cultural poems and artistic images to the user on the flexible screen.
[0067] Specifically, the process involves several steps: User Triggering and Initialization: The user triggers the device via a method such as a proximity sensor or a button. The device responds and begins the initialization process, including power-on and system self-test. Visual Data Acquisition: A top-mounted camera captures visual data such as the user's posture and facial expressions. This data is transmitted in real-time to the processing unit via WiFi or other communication methods. Data Preprocessing and Feature Extraction: Visual algorithms such as OpenCV are used to preprocess the user's image, including face detection and posture recognition, to extract key features such as facial features and body postures, providing a foundation for subsequent content generation. Personalized Content Generation: Based on the extracted features, Large Language Models (LLMs) are used to generate poems that match the user's features. Simultaneously, image generation models (such as Stable Diffusion) are used to generate artistic images related to the user's posture. Content Optimization and Synchronization: The generated poems and images are optimized to ensure semantic consistency and visual expressiveness. The optimized content is then synchronized to the flexible LED screen for display. Dynamic Display and Interactive Feedback: The LED screen dynamically displays the generated poems and images, simulating a "carousel" effect. Users experience the direct impact of their input on the generated content through visual and behavioral feedback. Data recording and analysis: The system stores all generated text and image content as an Excel file, providing a complete record for subsequent data analysis and optimization; Interactive experience end and output: After the interactive experience ends, the user can receive a printed photo containing the AI-generated work and background story, and the device enters standby mode, waiting for the next trigger.
[0068] This application's human-computer interaction method draws inspiration from the traditional "carousel," integrating a cylindrical flexible LED screen with a rotating mechanical structure to achieve dynamic aesthetics while ensuring structural stability. By capturing the user's posture in real time and interacting naturally with the device, the intuitiveness and smoothness of human-computer interaction are optimized, effectively reducing the user's cognitive load and significantly improving the immersion and user experience quality during the interaction process. The system integrates a high-precision WiFi camera and deep learning algorithms (such as ControlNet), dynamically generating highly personalized artistic poems and cultural interpretation texts through posture recognition and semantic extraction technologies. Combined with an intelligent rotation drive module, it retains the user viewing process of the traditional "carousel" format, enhancing the cultural experience and technological adaptability of the interaction. Generative AI technology optimization integrates multiple generative AI models (such as ZhipuAI and Stable Diffusion) to achieve cross-semantic to visual content generation. Through Prompt optimization and keyword extraction, it ensures that the generated content is highly relevant to user characteristics, while Remove BG technology is used to improve the background clarity and overall expressiveness of the generated images.
[0069] As shown in Figure 5, a user-centric approach combined with multimodal perception technology enables an immersive cultural interaction experience. When a user triggers the device, a prompt guides them to adjust their posture, while a top camera captures their posture and facial expressions. Through real-time posture recognition and semantic extraction technology, the system transforms user characteristics into personalized digital content, including poems, paintings, and their cultural interpretations. This content is dynamically presented on a flexible LED screen, allowing users to experience the direct impact of their input on the generated content through visual and behavioral feedback. After the interactive experience, users can collect a printed photo, which includes the AI-generated artwork and its background story. This process, through real-time processing of visual data and the display of generative content, transforms the user from a "viewer" into a "creator," enhancing participation and immersion.
[0070] As shown in Figure 6, this application is developed using the Python programming language and integrates OpenCV vision algorithms and Large Language Models (LLMs) to support user-generated and created poems and images. The system deeply explores the path of AI model integration and optimizes the collaborative work between different large language models. Natural human-computer interaction is achieved through a pose recognition input device. The system can capture image and trajectory information from the input device and utilize computer vision algorithms for user face detection and pose recognition. The user's facial features and body posture directly affect the semantic content of the generated poems and the image composition. The system employs a deep learning-based keyword extraction algorithm and Prompt optimization technology to ensure that the generated poems accurately reflect the user's facial features. Simultaneously, through the pose control module in ControlNet, the generated character image can highly match the user's body posture, thereby achieving accurate human-computer interactive creation. The workflow includes three key stages to achieve dynamic processing of visual data and personalized presentation of generated content. First, in the initial data acquisition stage, the system identifies the user and captures the first image through a WiFi camera, using semantic analysis and keyword extraction technology to provide a foundation for content generation. The system then enters the content generation and update phase. Every 30 seconds, it collects the user's new posture and dynamically generates poetic and pictorial content that matches the user's characteristics through real-time calculations, ensuring that the generated content reflects the user's interactive changes in real time. Finally, the system stores all generated text and image content in an Excel file, providing a complete record for subsequent data analysis and optimization. This design of multiple data collections and real-time generation significantly enhances the interactive linkage between user input and generated content, while improving the immersiveness and responsiveness of the personalized creation experience.
[0071] As shown in Figure 7, the system's multimodal AI workflow consists of seven highly efficient and collaborative functional modules that work together to achieve a complete process of personalized content generation and dynamic display. First, the image acquisition module uses a WiFi camera to capture user input data, including posture and facial expression information. Next, the posture recognition module extracts key posture features based on ControlNet technology, ensuring data accuracy and real-time performance. The keyword extraction module extracts key semantic information from images using deep learning algorithms, providing a precise semantic foundation for content generation. The Prompt optimization module further optimizes the input instructions of the generation model, improving the semantic consistency and accuracy of the generated content. In the generative AI module, ZhipuAI and BaiduAI models are integrated to generate ancient-style poems, showcasing text content with high cultural characteristics. The image generation module uses StableDiffusion to generate artistic images and optimizes the image background using Remove BG technology to enhance visual expressiveness. Finally, the display module transmits the generated text and images to a flexible LED screen, achieving dynamic content presentation. The modular design architecture significantly improves the system's flexibility and adaptability, enabling generated content to accurately match user characteristics while optimizing the efficiency and reliability of visual data processing.
[0072] Understandably, while this application draws inspiration from the "carousel," other dynamic display forms of traditional culture may also be incorporated. For example, the dynamic fluttering of traditional kites or the dynamic projection in shadow puppetry could be combined with generative AI as an alternative design. The system currently integrates generative AI models such as ZhipuAI and Stable Diffusion for content generation. Other alternatives may employ different generative models (such as OpenAI's DALLE and MidJourney) or be based on self-developed generative algorithms to achieve semantic-to-content transformation. Multimodal input expansion: The device primarily collects data through user posture and facial expressions. Possible variations include adding voice interaction, touch input, or environmental awareness (such as temperature and light) as additional interaction dimensions, thereby expanding the diversity of user input. Adjustment of content display and user participation modes: The current device achieves an immersive user experience through dynamic generation and real-time display. Alternative designs may combine preset content with user input, moving the content generation portion to the cloud or a separate computing platform, thus reducing the device's reliance on computing power. Based on the extension of different cultural elements: This application is inspired by the "carousel". The alternative design may utilize other dynamic cultural elements, such as the dynamic display of ancient mechanical clocks and the rotating scenery of traditional opera stages, to generate content in combination with user input, thereby avoiding the core characteristics of the "carousel" form.
[0073] In the description of this application, unless otherwise expressly specified and limited, the terms "installation," "connection," "linking," "fixing," etc., should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral part; they can refer to a mechanical connection, an electrical connection, or a connection that allows communication between them; they can refer to a direct connection or an indirect connection through an intermediate medium; they can refer to the internal communication between two components or the interaction between two components. Those skilled in the art can understand the specific meaning of the above terms in this application according to the specific circumstances.
[0074] In the description of this application, it should be understood that the terms "center", "longitudinal", "lateral", "length", "width", "thickness", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise", "axial", "radial", "circumferential", etc., indicating the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings, are only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of this application.
[0075] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified.
[0076] It should be noted that, in this application, unless otherwise expressly specified and limited, "above" or "below" the second feature can mean that the first feature is in direct contact with the second feature, or that the first feature is in indirect contact with the second feature through an intermediate medium. Furthermore, "above," "on top of," and "over" the second feature can mean that the first feature is directly above or diagonally above the second feature, or simply indicates that the first feature is at a higher horizontal level than the second feature. "Below," "below," and "under" the second feature can mean that the first feature is directly below or diagonally below the second feature, or simply indicates that the first feature is at a lower horizontal level than the second feature.
[0077] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a particular order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of the application described herein can be implemented, for example, in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0078] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "illustrative embodiment," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this disclosure. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0079] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.
Claims
1. A cultural interaction device based on artificial intelligence and multimodal perception, characterized in that, The cultural interaction device based on artificial intelligence and multimodal perception includes a multimodal perception unit, a generative content processing unit, a physical display unit, and a control system. The multimodal perception unit is connected to the generative content processing unit, and the generative content processing unit is connected to the physical display unit through the control system. The multimodal perception unit acquires the user's visual data and sends it to the generative content processing unit. The generative content processing unit processes the visual data to generate dynamic text image information for the user and sends it to the physical display unit through the control system. The physical display unit displays the dynamic text image information to the user. The multimodal sensing unit includes a lantern shell and a camera, with the camera mounted on the lantern shell. The generative content processing unit includes a data processing module, a digital content generation module, and a content storage and update module connected in sequence. The physical display unit includes a screen controller and a flexible screen connected in sequence, with the content storage and update module connected to the screen controller. The control system includes a power management module, a data transmission module, and a motor control module. The power management module is connected to the camera, the screen controller, the flexible screen, and the motor control module. The camera is connected to the data processing module through the data transmission module. The motor control module is used to control the rotation of the flexible screen.
2. The cultural interaction device based on artificial intelligence and multimodal perception according to claim 1, characterized in that, The physical display unit also includes a rotating mechanism, the flexible screen is connected to the rotating mechanism, the motor control module is connected to the rotating mechanism, and the rotating mechanism is used to drive the flexible screen to rotate.
3. The cultural interaction device based on artificial intelligence and multimodal perception according to claim 2, characterized in that, The rotating mechanism includes a connecting plate and a DC motor, a gear set, and a curved screen bracket disposed on the connecting plate. The connecting plate is disposed on the lantern shell, the DC motor is connected to the gear set, and the flexible screen is disposed in the curved screen bracket.
4. The cultural interaction device based on artificial intelligence and multimodal perception according to claim 3, characterized in that, Multiple flexible screens are provided, and the multiple flexible screens are arranged around the lantern shell to display information to the user.
5. The cultural interaction device based on artificial intelligence and multimodal perception according to claim 3, characterized in that, The motor control module is a drive board, which is mounted on the connection board, and the DC motor is connected to the drive board.
6. An interaction method based on the cultural interaction device based on artificial intelligence and multimodal perception as described in any one of claims 1 to 5, characterized in that, The interaction method includes: the multimodal perception unit acquiring the user's visual data and sending it to the generative content processing unit; the generative content processing unit processing the visual data to generate dynamic text image information for the user and sending it to the entity display unit through the control system; and the entity display unit displaying the dynamic text image information to the user.
7. The interaction method of the cultural interaction device based on artificial intelligence and multimodal perception according to claim 6, characterized in that, The visual data includes posture features and facial features; the multimodal perception unit acquires the user's visual data by: responding to the user's trigger action and collecting the user's posture features and facial features after the user completes the trigger action through the camera.
8. The interaction method of the cultural interaction device based on artificial intelligence and multimodal perception according to claim 7, characterized in that, The dynamic text image information includes cultural poems and artistic images; the generative content processing unit processes the visual data to generate the user's dynamic text image information, specifically including: the data processing module receives the user's posture features and appearance features, and preprocesses the posture features and appearance features to obtain key features; the digital content generation module generates the cultural poems and artistic images that match the user based on the key features.
9. The interaction method of the cultural interaction device based on artificial intelligence and multimodal perception according to claim 8, characterized in that, The physical display unit displays the dynamic text and image information to the user, specifically: the screen controller displays the cultural poems and artistic images to the user on the flexible screen.
Citation Information
Patent Citations
Smart home scene understanding and interaction method and system based on multi-modal fusion
CN119398159A