Cultural interaction device and method based on artificial intelligence and multi-modal perception
By integrating multimodal perception unit, generative content processing unit and physical display unit in the cultural interaction device, the problem of difficult user input in the prior art is solved, and the dynamic generation and immersive experience of personalized content are realized.
Patent Information
- Application Number
- CN202510233077.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-02-28
AI Technical Summary
Existing cultural interaction devices are difficult to achieve dynamic adjustment through user input, resulting in the inability to meet diverse interaction needs, which weakens the immersion and interest of cultural communication.
Using a cultural interaction device based on artificial intelligence and multimodal perception, through the cooperation of multimodal perception unit, generative content processing unit and physical display unit, the capture, processing and analysis of user input features, as well as the generation and dynamic display of personalized digital content.
Real-time capture of user input features and dynamic generation of personalized content are realized, which improves users' immersive cultural interaction experience, meets diverse interactive needs, and enhances the modern expression of cultural communication.
Smart Images

Figure CN120143981A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of human-computer interaction technology, and in particular to a cultural interaction device and method based on artificial intelligence and multimodal perception. Background Art
[0002] Although there are many cultural communication methods based on digital technology, these interactive systems often have high learning thresholds, and complex operation interfaces easily cause a large cognitive load on users, limiting the possibility of natural user participation, and also weakening the immersion and fun of cultural communication. As a classic handicraft, the revolving lantern cleverly combines structure and light and shadow art. The impeller is driven by hot air flow to rotate, and dynamic images are projected on the surface of the lantern to form a flowing image art. Its structure includes key components such as paper lanterns, paper cuttings and impellers, and the dynamic image display is completed by the hot air flow generated by candlelight. This traditional device not only has profound cultural implications, but also provides unique inspiration for modern light and shadow art and mechanical device design. In recent years, with the application of digital technology, more and more innovative devices have tried to combine traditional cultural elements with modern technology to expand their application value in the field of art display and communication. However, these studies mostly focus on appearance or craft reproduction, and there is still a lack of systematic exploration of the deep integration of physical intelligent interaction and cultural experience.
[0003] Digital crafts have given traditional crafts new creative possibilities by combining digital materials, electronic components and computer-aided design technology, such as parametric metal art based on growth logic and silverware with interactive functions. This approach not only enriches the sensory experience of the material carrier, but also stimulates emotional resonance and cultural cognition through the dynamic expression of cultural symbols. For example, cultural imagination is enhanced by triggering the working sound of craftsmen, and thinking about alternative computing technologies is stimulated by textile logic gates. In recent years, multimodal interactive systems have developed rapidly in the fields of artificial intelligence and digital technology. Such systems rely on computer vision and deep learning algorithms to capture user input (such as expressions and gestures) and generate matching digital content. However, most studies only focus on a single field of image or text generation, lacking the deep linkage between physical entities and digital content. For example, some studies have tried to project user-generated images onto a rotating surface, but due to the lack of mapping rigidity and real-time feedback, the user experience is stiff and lacks fluency. In addition, dynamic mechanical devices are usually only used as passive display media, making it difficult to achieve real-time linkage between user input and content generation, weakening the emotional expression and immersive experience of interaction.
[0004] Existing intelligent interaction designs incorporating cultural elements also have limitations. They mostly remain at the static presentation of traditional symbols and fail to deeply explore cultural connotations and interaction potential. For example, dynamic imaging devices based on "Chinese lanterns" mostly adopt a rotating display mode with fixed content, making it difficult to achieve dynamic adjustment through user input, unable to meet diverse interaction needs, and unable to convey deep cultural values and emotional experiences.
[0005] Therefore, the existing technologies still need to be improved and developed. Summary of the Invention
[0006] The main objective of this application is to provide a cultural interaction device and method based on artificial intelligence and multimodal perception, aiming to solve the problem in the existing technologies that cultural interaction devices are difficult to achieve dynamic adjustment through user input, resulting in the inability to meet diverse interaction needs.
[0007] In the first aspect of the embodiments of this application, a cultural interaction device based on artificial intelligence and multimodal perception is provided. The cultural interaction device based on artificial intelligence and multimodal perception includes a multimodal perception unit, a generative content processing unit, an entity display unit, and a control system. The multimodal perception unit is connected to the generative content processing unit, and the generative content processing unit is connected to the entity display unit through the control system. The multimodal perception unit is configured to obtain visual data of the user and send it to the generative content processing unit. The generative content processing unit is configured to process the visual data, generate dynamic text image information of the user, and send it to the entity display unit through the control system. The entity display unit is configured to display the dynamic text image information to the user.
[0008] Optionally, in an embodiment of this application, the multimodal perception unit includes a lantern housing and a camera, and the camera is disposed on the lantern housing. The generative content processing unit includes a data processing module, a digital content generation module, and a content storage and update module that are connected in sequence. The entity display unit includes a screen controller and a flexible screen that are connected, and the content storage and update module is connected to the screen controller. The control system includes a power management module, a data transmission module, and a motor control module. The power management module is respectively connected to the camera, the screen controller, the flexible screen, and the motor control module. The camera is connected to the data processing module through the data transmission module, and the motor control module is configured to control the rotation of the flexible screen.
[0009] Optionally, in an embodiment of this application, the entity display unit further includes a rotating mechanism, the flexible screen is connected to the rotating mechanism, the motor control module is connected to the rotating mechanism, and the rotating mechanism is configured to drive the flexible screen to rotate.
[0010] Optionally, in an embodiment of the present application, the rotation mechanism includes a connecting plate, a DC motor, a gear set, and a curved screen bracket arranged on the connecting plate. The connecting plate is arranged on the lantern housing. The DC motor is connected to the gear set, and the flexible screen is arranged in the curved screen bracket.
[0011] Optionally, in an embodiment of the present application, a plurality of flexible screens are provided, and the plurality of flexible screens are arranged around the lantern housing to display information to the user.
[0012] Optionally, in an embodiment of the present application, the motor control module is a drive board. The drive board is arranged on the connecting plate, and the DC motor is connected to the drive board.
[0013] A second aspect of the embodiments of the present application further provides an interaction method for a cultural interaction device based on artificial intelligence and multi-modal perception according to any one of the above solutions. The interaction method includes: the multi-modal perception unit acquires visual data of the user and sends it to the generative content processing unit; the generative content processing unit processes the visual data to generate dynamic text image information of the user, and sends it to the entity display unit through the control system; the entity display unit displays the dynamic text image information to the user.
[0014] Optionally, in an embodiment of the present application, the visual data includes pose features and personal appearance features; the multi-modal perception unit acquires visual data of the user, specifically: in response to a trigger action of the user, the camera captures the pose features and personal appearance features of the user after completing the trigger action.
[0015] Optionally, in an embodiment of the present application, the dynamic text image information includes cultural poems and artistic images; the generative content processing unit processes the visual data to generate dynamic text image information of the user, specifically including: the data processing module receives the pose features and personal appearance features of the user, and preprocesses the pose features and personal appearance features to obtain key features; the digital content generation module generates the cultural poems and artistic images matching the user according to the key features.
[0016] Optionally, in an embodiment of the present application, the entity display unit displays the dynamic text image information to the user, specifically: the screen controller displays the cultural poems and artistic images on the flexible screen to the user.
[0017] Beneficial effects: The present application provides a cultural interaction device and method based on artificial intelligence and multimodal perception. Through the cooperation of the multimodal perception unit, the generative content processing unit, and the entity display unit, the present application realizes the capture, processing, and analysis of user input features, as well as the generation and dynamic display of personalized digital content, providing users with an immersive cultural interaction experience. Description of the Drawings
[0018] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0019] Figure 1 Exploded view of the overall structure of the preferred embodiment of the cultural interaction device based on artificial intelligence and multimodal perception of the present application;
[0020] Figure 2 Structure diagram of the rotating mechanism in the preferred embodiment of the cultural interaction device based on artificial intelligence and multimodal perception of the present application;
[0021] Figure 3 Circuit diagram in the preferred embodiment of the cultural interaction device based on artificial intelligence and multimodal perception of the present application;
[0022] Figure 4 Flowchart of the preferred embodiment of the interaction method of the cultural interaction device based on artificial intelligence and multimodal perception of the present application;
[0023] Figure 5 Interaction flowchart in the preferred embodiment of the interaction method of the cultural interaction device based on artificial intelligence and multimodal perception of the present application;
[0024] Figure 6 Interaction system process in the preferred embodiment of the cultural interaction device based on artificial intelligence and multimodal perception of the present application;
[0025] Figure 7 Workflow of multi-agent AI in the preferred embodiment of the cultural interaction device based on artificial intelligence and multimodal perception of the present application.
[0026] Description of the reference numerals:
[0027] 1. Insulation layer; 2. Slip ring; 3. Fixed block; 4. Curved screen bracket; 5. Fixture; 6. Ball bearing; 7. Connecting plate; 8. Coupling; 9. Large gear; 10. Small gear; 11. DC motor; 12. PCB board; 13. Lantern housing; 14. Display frame.
[0028] Through the above-mentioned drawings, specific embodiments of the present application have been shown, and will be described in more detail hereinafter. These drawings and written descriptions are not intended to limit the scope of the concept of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by reference to specific embodiments. Detailed Description of the Embodiments
[0029] To make the objectives, technical solutions and effects of the present application clearer and more definite, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. The described embodiments are only possible technical implementations of the present application, not all possible implementations. Based on the embodiments in the present application, those skilled in the art can fully combine the embodiments of the present application to obtain other embodiments without creative work, and these embodiments are also within the protection scope of the present application.
[0030] Existing systems also have obvious deficiencies in terms of real-time performance and immersion. Many interactive devices based on mechanical motion fail to solve the problem of delay in input capture and output response, and the linkage between generated content and physical dynamics lacks consistency. Although some studies have tried to use more efficient algorithms to optimize the response speed, no breakthrough has been made in the deep integration of dynamic devices and generated content. In short, the existing technologies have significant deficiencies in the depth, flexibility and emotional experience of the combination of dynamic physical entities and intelligent interaction systems. Therefore, exploring the real-time linkage between dynamic mechanical structures and user inputs and expanding the interaction capabilities of traditional cultural elements through intelligent technologies have become important research directions for intelligent cultural interaction devices. Multimodal interaction systems have gradually become an important research direction in the fields of artificial intelligence and digital technologies in recent years.
[0031] The following introduces the terms involved in the embodiments of the present application:
[0032] LLMs: Large Language Models;
[0033] FOC: Field-Oriented Control;
[0034] HDMI: High-Definition Multimedia Interface;
[0035] OpenCV: Open Source Computer Vision Library;
[0036] ControlNet: A pose control module or technology for pose recognition and pose control;
[0037] ZhipuAI: A generative AI model;
[0038] Stable Diffusion: An image generation model or technology for generating high-quality artistic images;
[0039] Remove BG: An image background removal technology or tool for optimizing image backgrounds.
[0040] The interactive device of this application provides a new technical path for real-time linkage and immersive interaction by integrating dynamic capture of user input, multimodal generation algorithms, and high-precision physical output. This integration not only combines user input and dynamic output more naturally but also endows the device with higher emotional expression ability and cultural transmission potential, laying an important foundation for innovation in the field of human-computer interaction. By introducing more intelligent technologies, the interactive device optimizes the interaction method, making it more intuitive and natural, and reducing the psychological pressure of users during operation. Combined with generative artificial intelligence, the device can not only dynamically generate personalized artistic content but also support co-creation between users and the device, thus significantly enhancing the sense of participation and emotional connection. In addition, the physical interaction design retains the classic viewing form of the traditional "Chinese lantern", achieving an innovative integration between modern digital technology and traditional culture.
[0041] The following describes the cultural interaction device and method based on artificial intelligence and multimodal perception according to the embodiments of this application with reference to the accompanying drawings. Aiming at the problem that the interactive device in the above-mentioned related technologies is difficult to achieve dynamic adjustment through user input, resulting in the inability to meet diverse interaction needs, this application provides a cultural interaction device based on artificial intelligence and multimodal perception. In this method, through the cooperation of the multimodal perception unit, the generative content processing unit, and the physical display unit, the capture, processing, and analysis of user input features, as well as the generation and dynamic display of personalized digital content, are realized, providing users with an immersive cultural interaction experience. Thus, the technical problem that the interactive device in the related technologies is difficult to achieve dynamic adjustment through user input, resulting in the inability to meet diverse interaction needs, is solved.
[0042] The following uses specific embodiments to elaborate on the technical solutions of this application in detail. These several specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.
[0043] Such as Figure 1 And Figure 2As shown, an embodiment of the present application provides a cultural interaction device based on artificial intelligence and multimodal perception, and the cultural interaction device based on artificial intelligence and multimodal perception includes a multimodal perception unit, a generative content processing unit, a physical display unit and a control system, the multimodal perception unit is connected to the generative content processing unit, and the generative content processing unit is connected to the physical display unit through the control system; the multimodal perception unit is used to obtain the user's visual data and send it to the generative content processing unit; the generative content processing unit is used to process the visual data, generate dynamic text image information of the user, and send it to the physical display unit through the control system; the physical display unit is used to display the dynamic text image information to the user.
[0044] In order to solve the shortcomings of the existing technology in terms of human-computer interaction experience, deep integration of cultural communication, and digital inheritance of intangible cultural heritage, this application is inspired by the structure of the traditional intangible cultural heritage "revolving lantern", combined with a cylindrical LED flexible screen and a rotating mechanical device, through the collaborative work of a multimodal perception unit, a generative content processing unit, and a physical display unit, to achieve real-time linkage between user input and dynamic digital content. The user's body posture and character appearance features are captured through posture recognition and deep learning algorithms, and semantic extraction and Prompt (prompt or guidance) optimization technology are combined to generate highly personalized cultural content, including art poems and paintings and their cultural interpretation texts. The generated content is dynamically presented through a high-precision LED screen and a mechanical rotating device, which not only optimizes the quality of content generation and the correlation between user input, but also significantly enhances the immersion of the interactive experience and the modern expression of cultural communication.
[0045] It should be noted that this application is based on multimodal intelligent agent technology, and by integrating computer vision algorithms and large language models (LLMs), features are extracted from user input visual data such as posture and appearance to generate highly matched poetry and image content. Through the collaborative work of Prompt optimization technology and deep learning models, it is ensured that the generated content is highly relevant to user characteristics in terms of semantics and style. The generated content automatically identifies the foreground layer through AI technology, and then separates it from the background layer for seamless connection and dynamic adaptation on the flexible display screen. The interactive process of the cultural interaction device of this application ensures the smoothness of interaction and the accuracy of feedback through multiple user data collection and real-time update of generated content. By collecting user posture information in stages and combining the generated poetry and painting content for dynamic display, users can obtain personalized and contextualized visual and auditory interactive experience in a short time. And the design of the cultural interaction device combines the rotating dynamic structure of the traditional revolving lantern and the modeling characteristics of the "pavilion" of the garden building, and integrates cultural imagery into the expression form of the dynamic device to enhance the user's immersive experience and emotional expression.
[0046] It is understandable that, in view of the limitations of the prior art in terms of the linkage between dynamic physical devices and user interactions, the matching of personalized generated content, and emotional expression, this application explores a possible path to enhance the human-computer interaction experience by integrating visual data such as the gestures and expressions input by users with the dynamic display of generated content. Through the combination of intelligent technology and dynamic display, it has the following technical effects: enhancing cultural experience, optimizing the immersive experience of users through multi-modal interaction technology, and enhancing the interest and interactivity of traditional culture dissemination; endowing a sense of creative participation, combining with generative AI models to enable users to directly participate in content creation and enhance the personalization of cultural activities; natural interaction and immersive experience: by optimizing the linkage mechanism between user input and device output, significantly reducing the complexity of interaction, providing a more natural interaction method, and enhancing the user's sense of immersion and interest; wide range of application scenarios, applicable to exhibition spaces, public cultural venues, and educational environments, and capable of promoting the innovative dissemination and popularization of intangible cultural heritage.
[0047] In one embodiment of this application, as Figure 1 and Figure 2 shown, the multi-modal perception unit includes a lantern housing 13 and a camera, and the camera is disposed on the lantern housing 13; the generative content processing unit includes a data processing module, a digital content generation module, and a content storage and update module connected in sequence; the entity display unit includes a screen controller and a flexible screen connected to each other, and the content storage and update module is connected to the screen controller; the control system includes a power management module, a data transmission module, and a motor control module. The power management module is respectively connected to the camera, the screen controller, the flexible screen, and the motor control module. The camera is connected to the data processing module through the data transmission module, and the motor control module is used to control the rotation of the flexible screen.
[0048] Specifically, the camera is used to capture input features such as the user's posture and appearance; sensors (such as infrared sensors, accelerometers, etc.) are used to assist in capturing the user's movement or position information. The camera or sensors capture the user's real-time posture, expression and other multimodal information, providing a data basis for subsequent content generation. The data processing module is responsible for receiving and preprocessing the data from the multimodal perception unit. The AI model (such as OpenCV vision algorithm, large language model LLMs, etc.) is used to analyze the processed data and generate personalized digital content (such as poems, paintings and their cultural interpretations). The content storage and update module stores the generated digital content and updates or adjusts it as needed; by analyzing and processing the data captured by the multimodal perception unit, personalized digital content matching the user's characteristics is generated, and these generated contents are stored and managed to ensure their accuracy and real-time performance during display. The flexible LED screen is used to dynamically display the generated digital content, and the screen controller is responsible for controlling the display content and synchronization of the screen; the digital content generated by the generative content processing unit is presented visually to the user to achieve an immersive experience. Ensure the synchronous display of multi-screen content and improve visual expressiveness. The power management module is responsible for converting the external power supply into the voltages required by each component inside the device. The data transmission module is responsible for data communication and transmission between components. The motor control module (such as an FOC driver board) is used to control the rotation effect of the device to ensure the smoothness of dynamic display; provide a stable power supply to ensure the normal operation of each component inside the device, achieve efficient data transmission and synchronization, ensure the collaborative work between components, control the rotation effect of the device, and enhance the immersion and visual impact of dynamic display.
[0049] In an embodiment of the present application, the entity display unit further includes a rotating mechanism, the flexible screen is connected to the rotating mechanism, the motor control module is connected to the rotating mechanism, and the rotating mechanism drives the flexible screen to rotate.
[0050] In an embodiment of the present application, as Figure 2 shown, the rotating mechanism includes a connecting plate 7 and a DC motor 11, a gear set and a curved screen support 4 arranged on the connecting plate 7. The connecting plate 7 is arranged on the lantern housing 13. The DC motor 11 is connected to the gear set, and the flexible screen is arranged in the curved screen support 4.
[0051] In an embodiment of the present application, a plurality of flexible screens are provided, and the plurality of flexible screens are arranged around the lantern housing 12 to display information to the user.
[0052] In an embodiment of the present application, as Figure 2 shown, the motor control module is a driver board, the driver board is arranged on the connecting plate 7, and the DC motor 11 is connected to the driver board.
[0053] Specifically, an independently designed Field-Oriented Control (FOC) driver board is adopted, cooperating with the brushless DC motor 11 to achieve precise control of the rotating frame. The flexible LED screen serves as the content display medium, dynamically presenting the image and text content generated based on user input, and being synchronized with the mechanical movement in real time.
[0054] As Figure 1 shown, in the screen display, 6 flexible LED screens (i.e., flexible screens) are used to simulate the traditional "carousel" effect. The screen design and display function integrate high-precision hardware layout and advanced synchronization technology. The screen layout supports flexible bending to fit the cylindrical outer frame and is fixed by magnetic attraction. The main frame is made of insulating material to ensure operational safety and structural stability. The display content is segmented by the screen controller and adapted to 6 screens. The TouchDesigner engine is used to develop dynamic content, including traditional patterns, poems, and animated characters, to convey rich cultural visual information. The multi-screen synchronization technology achieves precise multi-screen output through the controller, ensuring that the content is displayed smoothly and continuously during rotation, and further optimizing the visual performance.
[0055] As Figure 2 shown, the rotation effect is achieved as follows. To achieve the rotation display effect, an independently designed Field-Oriented Control (FOC) driver board is adopted, and the dynamic response of the rotating mechanism is controlled by the brushless DC motor 11. The FOC driver board has an advanced dynamic response adjustment function, which can optimize the motor speed and torque parameters according to the load conditions to ensure the operation efficiency and stability in different usage scenarios. The on-board SPI interface supports communication with the TouchDesigner engine to realize real-time adjustment and optimization of key parameters such as motor speed and torque. The mechanical part adopts a precision gear transmission structure. The motor is coupled with the large gear 9 of the frame through the small gear 10 to drive the smooth rotation of the display frame 14. This rotation drive module provides high-precision rotation control through FOC technology, significantly improving the real-time performance and fluency of the system, and providing a reliable technical guarantee for the stable presentation of dynamic display content.
[0056] It can be understood that the current device captures the user's posture and personal appearance features through a WiFi camera and computer vision algorithms (such as ControlNet). A possible alternative is to use depth sensors (such as Kinect), lidar (LIDAR), or other biometric devices (such as heart rate monitoring or EEG brain wave detection) to enhance the diversity of user input.
[0057] Furthermore, referring to Figure 1 and Figure 2, the lantern housing 13 includes a display frame 14 and a cover. The rotating mechanism is within the cover and the display frame 14. The curved screen support 4 is connected to the connecting plate 7. The connecting plate 7 is arranged inside the lantern housing 13. The curved screen support 4 is connected to the display frame 14. Clamps 5 are arranged above and below the connecting plate 7. The connecting plate 7 is rotatably connected to a rotating shaft through a ball bearing 6. The rotating shaft is fixedly connected to the curved screen support 4. A fixed block 3, a slip ring 2, and an insulating layer 1 are connected to the curved screen support 4. The insulating layer 1 is connected to the cover. A coupling 8 is arranged below the connecting plate 7. The coupling 8 is connected to a large gear 9 in a gear set. There is also a PCB board 12 (i.e., the driving board) below the connecting plate 7. The PCB board 12 is connected to a DC motor 11. The DC motor 11 is connected to a small gear 10 in the gear set. The small gear 10 is meshed with the large gear 9.
[0058] In this embodiment, as Figure 3 shown, the circuit connection design is as follows: The internal circuit design of the device adopts a modular structure, combining efficient power management and real-time data transmission strategies to ensure the stable operation and collaborative performance of the system. The power management module converts 220V AC power into 12V DC power through a transformer to provide stable power supply for the host and the screen controller. Subsequently, the 12V is further reduced to 5V through a secondary transformer to supply power to the LED flexible screen, thus meeting the voltage requirements of different components. The data transmission module is connected to each LED screen through the screen controller, uses data lines for signal transmission, and communicates with the host at high speed through an HDMI interface to ensure the synchronous display of multi-screen content. The camera power supply and communication module is driven by an independent 12V battery, establishes a data transmission link with the host through WiFi, and realizes the real-time acquisition and efficient transmission of attitude data. The optimized circuit design effectively improves the real-time performance of data transmission and the stability of system operation, providing a reliable technical guarantee for the collaborative work of multiple modules.
[0059] This application captures input features such as the user's posture and expression through a multi-modal artificial intelligence model, generates personalized digital content (such as poems, paintings, and their cultural interpretations), and presents it immersively through a dynamic display device. The system adopts a modular design, including multi-layer collaborative optimization of mechanical structure, circuit design, generation algorithm, and user interaction process, ensuring the efficient operation of the device and the immersion of the user experience.
[0060] Based on the above embodiments, this application also provides an interaction method for the cultural interaction device based on artificial intelligence and multi-modal perception, as Figure 4 shown. The interaction method includes the following steps:
[0061] In step S101, the multi-modal perception unit acquires the visual data of the user and sends it to the generative content processing unit;
[0062] In a possible implementation, the visual data includes pose features and personal appearance features; in response to the triggering action of the user, the camera captures the pose features and personal appearance features of the user after the triggering action is completed.
[0063] In step S102, the generative content processing unit processes the visual data, generates the dynamic text image information of the user, and sends it to the entity display unit through the control system;
[0064] In a possible implementation, the dynamic text image information includes cultural poems and artistic images; the data processing module receives the pose features and personal appearance features of the user, and preprocesses the pose features and personal appearance features to obtain key features; the digital content generation module generates the cultural poems and artistic images that match the user according to the key features.
[0065] In step S103, the entity display unit displays the dynamic text image information to the user.
[0066] In a possible implementation, the screen controller displays the cultural poems and artistic images on the flexible screen to the user.
[0067] Specifically, for user triggering and initialization, the user triggers the device through a certain method (such as a proximity sensor, button, etc.). The device responds to the trigger and starts the initialization process, including power-on startup, system self-check, etc.; for visual data collection, the top camera captures visual data such as the user's posture and expression, and the data is transmitted to the processing unit in real time via WiFi or other communication methods; for data preprocessing and feature extraction, visual algorithms such as OpenCV are used to preprocess the user image, such as face detection, posture recognition, etc., and key features such as facial features and limb postures are extracted to provide a basis for subsequent content generation; for personalized content generation, based on the extracted features, large language models (LLMs) are used to generate poems that match the user's features, and at the same time, image generation models (such as Stable Diffusion) are used to generate artistic images related to the user's posture; for content optimization and synchronization, the generated poems and images are optimized to ensure semantic consistency and visual expressiveness, and the optimized content is synchronized to the flexible LED screen for display; for dynamic display and interactive feedback, the LED screen dynamically displays the generated poems and images, simulating the "lantern slide" effect, and the user can feel the direct impact of their own input on the generated content through visual and behavioral feedback. For data recording and analysis, the system stores all the generated text and image content as an Excel file to provide a complete record for subsequent data analysis and optimization; for the end of the interactive experience and output, after the interactive experience ends, the user can receive a printed photo containing the AI-generated works and background stories, and the device enters the standby state, waiting for the next trigger.
[0068] The human-computer interaction method of this application is inspired by the traditional "lantern slide", integrating a cylindrical flexible LED screen and a rotating mechanical structure to achieve the performance of dynamic beauty while ensuring the structural stability. By real-time capturing the user's body posture and forming a natural interaction with the device, the intuitiveness and fluency of human-computer interaction are optimized, effectively reducing the user's cognitive load, and significantly enhancing the immersion and user experience quality during the interaction process. The system integrates a high-precision WiFi camera and deep learning algorithms (such as ControlNet), and through posture recognition and semantic extraction technologies, dynamically generates highly personalized artistic poems, paintings and cultural interpretation texts. Combining with the intelligent rotation drive module, it retains the user viewing process in the form of the traditional "lantern slide", strengthening the cultural experience and technical adaptability of the interaction. For the optimization of generative AI technology, the system integrates multiple generative AI models (such as ZhipuAI and Stable Diffusion) to achieve cross-semantic to visual content generation. Through Prompt optimization and keyword extraction, it ensures that the generated content is highly relevant to the user's features, and at the same time uses the Remove BG technology to improve the background clarity and overall expressiveness of the generated images.
[0069] Such as Figure 5As shown, user-centered and combined with multi-modal perception technology, an immersive cultural interaction experience is achieved. When the user triggers the device, a prompt sound guides the user to adjust their body posture, and the top camera captures the user's posture and expression information. Through real-time posture recognition and semantic extraction technology, the system transforms user characteristics into personalized digital content, including poems, paintings, and their cultural interpretation texts. These contents are dynamically presented on the LED flexible screen, and the user can feel the direct impact of their own input on the generated content through visual and behavioral feedback. After the interaction experience ends, the user can receive printed photos, which include AI-generated works and background stories. This process, through the real-time processing of visual data and the display of generative content, transforms the user from an "audience" to a "creator", enhancing the sense of participation and immersion.
[0070] As Figure 6 shown, this application is developed using the Python programming language, integrating OpenCV vision algorithms and large language models (LLMs) to support users in generating and creating poems and images. The system deeply explores the path of AI model integration and optimizes the collaborative work between different large language models. Through a posture recognition input device for natural human-computer interaction, the system can capture the image and trajectory information of the input device and use computer vision algorithms for user face detection and posture recognition. The user's facial features and body postures will directly affect the semantic content of poem generation and the image composition. The system adopts a deep learning-based keyword extraction algorithm and Prompt optimization technology to ensure that the generated poems accurately reflect the user's facial features. At the same time, through the pose control module in ControlNet, the generated human image can highly match the user's body posture, thus achieving precise human-computer interactive creation. The workflow includes three key stages to achieve the dynamic processing of visual data and the personalized presentation of generated content. First, in the initial data collection stage, the system identifies the user through a WiFi camera and captures the first image, using semantic analysis and keyword extraction technology to provide a basis for content generation. Subsequently, it enters the content generation and update stage, where the system captures the user's new posture every 30 seconds and dynamically generates poem and painting content that matches the user's characteristics in real-time, ensuring that the generated content can reflect the user's interaction changes in real-time. Finally, the system stores all the generated text and image content as an Excel file, providing a complete record for subsequent data analysis and optimization. This design of multiple data collections and real-time generation significantly enhances the interaction linkage between user input and generated content, while improving the immersion and response accuracy of the personalized creation experience.
[0071] As Figure 7As shown, the system's multi-modal AI workflow consists of seven highly collaborative functional modules, jointly realizing the complete process of personalized content generation and dynamic display. First, through the image acquisition module, a WiFi camera is used to capture user input data, including gesture and expression information. Then, the gesture recognition module extracts the user's key gesture features based on ControlNet technology to ensure the accuracy and real-time nature of the data. The keyword extraction module extracts the key semantic information in the image through deep learning algorithms, providing an accurate semantic basis for content generation. The Prompt optimization module further optimizes the input instructions of the generation model to improve the semantic consistency and accuracy of the generated content. In the generative AI module, ZhipuAI and BaiduAI models are integrated to generate ancient-style poems, presenting text content with high cultural characteristics. The image generation module uses StableDiffusion to generate artistic images and optimizes the image background through Remove BG technology to enhance visual expressiveness. Finally, the display module transmits the generated text and images to the flexible LED screen to achieve dynamic presentation of the content. The modular design architecture significantly improves the flexibility and adaptability of the system, enabling the generated content to accurately match user characteristics while optimizing the efficiency and reliability of visual data processing.
[0072] It can be understood that although this application is inspired by the "lantern slide", other forms of dynamic display of traditional culture may also be introduced. For example, the dynamic fluttering form of traditional kites or the combination of dynamic projection in "shadow puppetry" and generative AI can be used as alternative designs. The system currently integrates generative AI models such as ZhipuAI and Stable Diffusion for content generation. Other alternative solutions may adopt different generation models (such as DALLE of OpenAI, MidJourney) or generative algorithms based on independent research and development to achieve the transformation from semantics to content. Expansion of multi-modal input: The device mainly collects data through the user's gestures and expressions. Possible variant solutions include adding voice interaction, touch input, or environmental perception (such as temperature, light) as additional interaction dimensions to expand the diversity of user input. Adjustment of content display and user participation mode: The current device achieves an immersive experience for users through dynamic generation and real-time display. Alternative designs may combine preset content with user input, moving the content generation part to the cloud or an independent computing platform, thereby reducing the device's dependence on computing power. Extension based on different cultural elements: This application is inspired by the "lantern slide". Alternative designs may utilize other dynamic cultural elements, such as the dynamic display of ancient mechanical clocks, the rotating scenery on the traditional opera stage, etc., combined with user input for content generation, thus avoiding the core features of the "lantern slide" form.
[0073] In the description of this application, unless otherwise clearly specified and defined, terms such as "installation", "connection", "linkage", "fixation" shall be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or integrated; it may be a mechanical connection, an electrical connection, or capable of communicating with each other; it may be directly connected, or indirectly connected through an intermediate medium, and may be the communication inside two components or the interaction relationship between two components. For those of ordinary skill in the art, the specific meanings of the above terms in this application can be understood according to specific circumstances.
[0074] In the description of this application, it should be understood that the orientation or positional relationship indicated by terms such as "center", "longitudinal", "transverse", "length", "width", "thickness", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise", "axial", "radial", "circumferential", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing this application and simplifying the description, rather than indicating or implying that the device or component referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation to this application.
[0075] In addition, the terms "first" and "second" are only used for descriptive purposes and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of this application, the meaning of "a plurality" is at least two, such as two, three, etc., unless otherwise clearly and specifically defined.
[0076] It should be noted that in this application, unless otherwise clearly specified and defined, the first feature may be in direct contact with the second feature "on" or "under" the second feature, or the first and second features may be indirectly in contact through an intermediate medium. Moreover, the first feature being "above", "over" and "on" the second feature may mean that the first feature is directly above or obliquely above the second feature, or merely indicates that the first feature has a higher horizontal height than the second feature. The first feature being "under", "beneath" and "under" the second feature may mean that the first feature is directly below or obliquely below the second feature, or merely indicates that the first feature has a lower horizontal height than the second feature.
[0077] In the description of the specification, claims, and the above-mentioned drawings of this application, the terms "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of this application described here, for example, can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that comprises a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0078] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "schematic embodiments", "examples", "specific examples", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present disclosure. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0079] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application and are not intended to limit them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.
Claims
1. A cultural interaction device based on artificial intelligence and multimodal perception, characterized in that: The cultural interaction device based on artificial intelligence and multimodal perception includes a multimodal perception unit, a generative content processing unit, a physical display unit and a control system, wherein the multimodal perception unit is connected to the generative content processing unit, and the generative content processing unit is connected to the physical display unit through the control system; The multimodal perception unit is used to obtain visual data of the user and send it to the generative content processing unit; The generative content processing unit is used to process the visual data, generate dynamic text image information of the user, and send it to the physical display unit through the control system; The physical display unit is used to display the dynamic text image information to the user.
2. The cultural interaction device based on artificial intelligence and multimodal perception according to claim 1, characterized in that: The multimodal sensing unit comprises a lantern housing and a camera, wherein the camera is arranged on the lantern housing; The generative content processing unit comprises a data processing module, a digital content generation module and a content storage and updating module connected in sequence; The physical display unit includes a screen controller and a flexible screen connected to each other, and the content storage and update module is connected to the screen controller; The control system includes a power management module, a data transmission module and a motor control module. The power management module is respectively connected to the camera, the screen controller, the flexible screen and the motor control module. The camera is connected to the data processing module through the data transmission module. The motor control module is used to control the rotation of the flexible screen.
3. The cultural interaction device based on artificial intelligence and multimodal perception according to claim 2, characterized in that: The physical display unit further includes a rotating mechanism, the flexible screen is connected to the rotating mechanism, the motor control module is connected to the rotating mechanism, and the rotating mechanism is used to drive the flexible screen to rotate.
4. The cultural interaction device based on artificial intelligence and multimodal perception according to claim 3 is characterized in that: The rotating mechanism includes a connecting plate and a DC motor, a gear set and a curved screen bracket arranged on the connecting plate. The connecting plate is arranged on the lantern housing, the DC motor is connected to the gear set, and the flexible screen is arranged in the curved screen bracket.
5. The cultural interaction device based on artificial intelligence and multimodal perception according to claim 4 is characterized in that: There are multiple flexible screens, and the multiple flexible screens are arranged around the lantern shell to display information to the user.
6. The cultural interaction device based on artificial intelligence and multimodal perception according to claim 4, characterized in that: The motor control module is a driving board, the driving board is arranged on the connecting board, and the DC motor is connected to the driving board.
7. An interactive method based on the cultural interaction device based on artificial intelligence and multimodal perception according to any one of claims 1 to 6, characterized in that: The interaction method comprises: The multimodal perception unit acquires visual data of the user and sends it to the generative content processing unit; The generative content processing unit processes the visual data to generate dynamic text image information of the user, and sends it to the physical display unit through the control system; The physical display unit displays the dynamic text image information to the user.
8. The interactive method of the cultural interaction device based on artificial intelligence and multimodal perception according to claim 7, characterized in that: The visual data includes posture features and character appearance features; The multimodal perception unit obtains the user's visual data, specifically: In response to the triggering action of the user, the posture characteristics and appearance characteristics of the user after completing the triggering action are collected through a camera.
9. The interactive method of the cultural interaction device based on artificial intelligence and multimodal perception according to claim 8, characterized in that: The dynamic text image information includes cultural poems and artistic images; The generative content processing unit processes the visual data to generate dynamic text image information of the user, specifically including: The data processing module receives the posture features and the appearance features of the user, and pre-processes the posture features and the appearance features of the user to obtain key features; The digital content generation module generates the cultural poem and the artistic image matching the user according to the key features.
10. The interactive method of the cultural interactive device based on artificial intelligence and multimodal perception according to claim 9, characterized in that: The entity display unit displays the dynamic text image information to the user, specifically: The screen controller displays the cultural poem and the artistic image to the user on the flexible screen.
Citation Information
Patent Citations
Robot interaction control method and device, storage medium and robot
CN115213884A
Man-machine interaction system and method based on multi-modal large language model
CN117762257A
Multi-modal AI brain construction method based on vehicle-mounted OS, terminal and storage medium
CN118790271A
Intelligent question number large-screen display system and method based on multi-modal large model
CN119088820A
Wearable immersive multi-modal feedback man-machine interaction system and method
CN119200850A
Cited By
Digital display system and method for non-perpetual cultural heritage
CN121187451A