Digital media art interactive experience system

By building AI-driven multimodal perception and dynamic content generation modules, immersive and personalized interaction of digital media art design is realized, user participation and cultural adaptability are improved, and problems of insufficient interaction and content solidification in traditional design are solved.

CN120560508AInactive Publication Date: 2025-08-29ZHOUKOU NORMAL UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510667384.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-22
Publication Date
2025-08-29
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional digital media art design lacks in-depth interaction and personalized experience, lack of user behavior understanding, it is difficult to meet the needs of global exhibitions, and the content presentation method is solidified, making it difficult to meet the diverse aesthetic needs.

Method used

A multimodal perception module based on AI is built to perceive user behavior in real time through positioning and motion capture technology, combining dynamic content generation modules and intelligent recommendation modules to realize immersive, personalized and adaptive interaction between users and digital art content.

Benefits of technology

It improves the interactivity and sense of participation of artistic works, meets users' personalized needs, enhances cultural adaptability, realizes accurate behavioral semantic analysis and emotional recognition, and solves the problems of one-way output and solidified content in traditional systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120560508A_ABST
    Figure CN120560508A_ABST
Patent Text Reader

Abstract

The invention discloses a digital media art interactive experience system, which relates to the technical field of digital interaction and comprises a multi-mode sensing module, a dynamic content generation module, an interactive experience module and an intelligent recommendation module. The multi-mode sensing module is used for carrying out real-time positioning and micro-motion capture on a user; the dynamic content generation module is used for constructing an artistic work parameterized model library, deconstructing color, form and motion track elements into adjustable parameters, and training an interaction strategy network through reinforcement learning, so that the system autonomously optimizes the content according to an audience behavior mode; the interactive experience module is used for superposing and projecting the digital media art content, the generated dynamic content and the corresponding projection substrate to a physical space; and the intelligent recommendation module is used for constructing a user knowledge graph, dynamically adjusting art symbols and guide prompts, and recommending interested digital media art for the user. According to the application, the user participation degree can be improved, diversified dynamic interaction can be realized, and the cultural adaptability is relatively high.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of digital interactive technology, and in particular to a digital media art interactive experience system. Background Art

[0002] Digital media art design is an artistic design activity that utilizes computer technology to digitally process, store, transmit, and display various media information. This design approach transcends the limitations of time and space, allowing artworks to be presented to users in richer and more diverse forms. Traditional digital media art relies on a one-way output model, with users participating only as passive recipients. This lacks deep interaction and personalized experience, resulting in relatively insufficient interactivity, participation, and emotional resonance in artworks. Furthermore, content presentation methods are often rigid and monolithic, making it difficult to meet the increasingly diverse aesthetic needs of contemporary audiences. Furthermore, existing digital interactive systems generally rely on single sensor technologies, which are limited in areas such as semantic analysis of user behavior and emotion recognition. Furthermore, they lack cultural adaptability, making it difficult to meet the demands of global exhibitions. Summary of the Invention

[0003] In light of this, this application provides a digital media art interactive experience system that addresses the one-way output model of traditional digital media art, where users participate only as passive recipients, lacking deep interaction and personalized experience. This results in relatively low interactivity, engagement, and emotional resonance in artworks. Furthermore, the content presentation format is rigid, lacks a deep understanding of user behavior, and is culturally inadequate, making it difficult to meet the technical requirements of global exhibitions.

[0004] To achieve the above objectives, the present invention provides the following technical solution: a digital media art interactive experience system. This system constructs a dynamic digital art space that can perceive user behavior and environmental data, enabling immersive, personalized, and adaptive interaction between users and digital art content. By leveraging AI (artificial intelligence) generation and virtual interaction technology, this system breaks through the traditional one-way output model of digital media art and establishes a closed-loop system of "perception-understanding-response-evolution." The system primarily consists of four core components: a multimodal perception module, a dynamic content generation module, an interactive experience module, and an intelligent recommendation module. The multimodal perception module integrates positioning and motion capture technologies to provide real-time spatial localization of users and capture precise micro-movements. Sensed user position, posture, and motion information are transmitted to the dynamic content generation module via a data transmission protocol. The dynamic content generation module constructs a parametric model library for artworks and, using structured data processing methods, deeply deconstructs artistic elements of digital media art, such as color, form, and motion trajectory, into a flexibly adjustable parameter system. Based on reinforcement learning theory, an interactive strategy network is trained, enabling the system to autonomously optimize dynamic content generation strategies based on audience behavior and feedback, achieving adaptive content iteration. The interactive experience module establishes a two-way data exchange channel with the dynamic content generation module. Using a high-precision projection device, digital media art content and a dynamically generated projection base are projected into the physical space. Simultaneously, augmented reality (AR) technology is used to overlay virtual content, generated in real time and precisely synchronized with the physical projection, onto the projection base. Furthermore, this module provides real-time feedback on audio and tactile information, leveraging multi-channel perception technology to create an immersive interactive experience for users. The intelligent recommendation module constructs a personalized knowledge graph based on user behavior data and, combined with a cultural dimension analysis model, dynamically adjusts the presentation of artistic symbols and guidance strategies. Furthermore, collaborative filtering and deep learning algorithms are employed to identify user interests and preferences, providing precise recommendations of digital media art works and related content that align with their interests, enhancing their artistic experience and engagement.

[0005] Furthermore, the multimodal perception module includes a user verification unit, a positioning unit and a motion recognition unit, wherein the user verification unit uses a deep learning-based facial feature extraction model to extract features from the facial image of the user when entering the venue obtained by the facial acquisition device, encrypts the extracted facial features, and establishes a temporary user profile based on the encrypted facial features.

[0006] The facial collection device is generally installed at the entrance of the exhibition hall. When the user enters, a temporary user profile can be established through simple facial recognition, and the profile will be automatically destroyed after 48 hours of storage. The facial collection device includes a camera, an image recognizer, a communication module, and a power supply module. The image recognizer uses a deep learning-based facial feature extraction model to extract features from the facial image collected by the camera, and uploads it to the cloud server for user profile registration and information recording.

[0007] The positioning unit is used to establish the three-dimensional spatial coordinate system of the exhibition hall through the scanning data of the exhibition hall, and to construct a real-time coordinate mapping relationship between the physical space and digital content, as well as to obtain the user's current physical space coordinate position by detecting the user's three-dimensional position in real time based on the IMU module.

[0008] Before positioning, it is necessary to obtain a three-dimensional scanning model of the exhibition hall based on the laser scanning device, and select a certain point as the origin based on the three-dimensional scanning model to establish a three-dimensional space coordinate system. According to the preset coordinate conversion relationship, the coordinates of each pixel point of the digital media art work are converted, and the real-time position of the user is obtained based on the IMU module on the AR device worn by the user, and the user's current physical space coordinate position is calculated.

[0009] The motion recognition unit is used to detect the user's basic body movements using the skeleton key point detection method, and to track the user's hand joints using the OpenPose deep learning model.

[0010] Furthermore, the multimodal perception module also includes a behavior semantic parsing unit, which is connected to the action recognition unit. The behavior semantic parsing unit is used to build a behavior pattern analysis model based on the LSTM network and the Transformer network, analyze the behavior target in combination with the scene context, and convert the user action into an art adjustment parameter. The behavior pattern analysis model includes a time series modeling unit and an interaction intention prediction unit. The time series modeling unit is used to process the action sequence S using a bidirectional LSTM network and output the behavior classification probability distribution P. The interaction intention prediction unit is used to establish a behavior-intention mapping based on the Transformer model. ,in, is the contribution of actions at different time steps to the final intention, is a learnable parameter matrix, and d is the feature dimension. LSTM captures local temporal patterns, while Transformer captures global dependencies, improving the accuracy of behavior analysis.

[0011] The structure of the bidirectional LSTM network includes an input layer, a forward LSTM, a backward LSTM, a hidden state, and a fully connected layer. The hidden layer dimension of each LSTM unit is h (128), and the output forward hidden state is and the backward hidden state , the final representation of each time step is , the input data is an action sequence , where each time step contains user action features, such as joint coordinates, velocity, acceleration, etc., and the dimension is the number of time steps Feature dimension, the fully connected layer is used to transform Mapped to the behavioral category space, , Apply Softmax to each time step and output the behavior classification probability distribution .

[0012] The structure of the Transformer model is the fusion of the input layer, the self-attention layer, the linear layer and the fully connected layer. The input layer is used to transform the temporal feature matrix of the output of the bidirectional LSTM into Environmental parameters corresponding to the scene context are concatenated to form enhanced features N. The self-attention layer is used to calculate the query, key, and value. The linear fusion layer concatenates the outputs of multiple single-head attentions. Through linear fusion, a fully connected layer maps the attention outputs to the intent space. Finally, the temporal features H output by the LSTM are concatenated with the intent score Intent to obtain the fused features F. Regression is then used to generate the parameters that drive the dynamic adjustment of the artwork.

[0013] Furthermore, the dynamic content generation module includes a parameter structure unit and a behavior-driven evolution unit; The parameter structure unit is used to analyze the characteristics of the original art material according to the color, shape, and motion trajectory elements using a color rhythm encoder, a morphological topology analyzer, and motion dynamics modeling, obtain the adjustable parameters of the corresponding parameter space, and build an art parameter model library; The behavior-driven evolution unit is used to generate the user's real-time behavior feature vector and ambient context parameters Generate state representation through generative adversarial network GAN and output parameter increment through policy network , then the rendering engine applies the new parameters to generate content and according to the reward function: Update the policy network, where U() is the real-time behavior matching, S() is the cosine similarity of parameter changes, and C() is the aesthetic term. The pre-trained aesthetic evaluation model is used for scoring, and the parameter change speed constraint is set to ensure a continuous transition of artistic styles. The constraint formula is: .

[0014] Furthermore, the dynamic content generation module also includes a multi-person interaction management unit, which is used for a dynamic resource allocation method based on Q-Learning, giving priority to responding to highly engaged users, and using the Shapley value algorithm to fairly allocate control rights when conflicts in interactive actions occur.

[0015] Furthermore, the interactive experience module includes a projection control unit, an AR interaction control unit, and a feedback adjustment unit. The projection control unit is used to project the digital media art content and the generated dynamic content projection substrate into the physical space through the projection device; The AR interaction control unit is used to overlay the real-time generated virtual content synchronized with the physical projection onto the projection base through the AR device through a timestamp synchronization protocol, and provide real-time feedback of audio and tactile information; The feedback adjustment unit is connected to the AR interaction control unit. The feedback adjustment unit is used to dynamically adjust the tactile feedback intensity according to the movement amplitude. The adjustment formula is: ,in, is the tactile feedback intensity, is the scaling factor, is the sum of the squares of the weights of all action components, which is used to reflect the overall amplitude of the action.

[0016] Furthermore, the intelligent recommendation module includes a knowledge graph construction unit, an adjustment prompt unit, and a user recommendation unit; The knowledge graph construction unit is used to take the artistic elements, users, eye movement patterns, and cultural dimensions in the original art materials as core entities and define the relationships between the entities. The constructed knowledge graph uses the Neo4j graph database to implement triple storage and update management, and the TransE algorithm is used to implement embedding representation learning. The adjustment prompt unit is connected to the knowledge graph construction unit. The adjustment prompt unit is used to associate cultural dimensions with artistic symbols and establish a mapping relationship according to the formula: , in, is the cultural dimension coefficient dynamically adjusted by Q-learning, s is the Sigmoid function, which compresses the weight to [0,1]. is the symbol-culture association matrix, where k is the number of artistic symbols and d is the number of cultural dimensions. Symbol weights are dynamically calculated, and each coefficient is dynamically adjusted through Q-learning. Fuzzy logic is also used to handle multi-dimensional cultural feature conflicts. It is also used to automatically increase the weight of the cultural openness dimension when users repeatedly gaze at cross-cultural conflicting elements, and to adjust the salience of artistic symbols based on the distribution of gaze hotspots. The user recommendation unit is connected to the knowledge graph construction unit. The user recommendation unit includes an eye tracking unit, an attention calculation unit and a hybrid recommendation unit. The eye tracking unit is used to obtain the eye movement trajectory and gaze hot zone, and synchronize the projection screen coordinate mapping, establish the correspondence between the gaze point and the pixel of the digital art canvas, and automatically identify the artistic elements in the painting through the image segmentation algorithm. The attention calculation unit is connected to the eye tracking unit. The attention calculation unit is used to model the interest of artistic elements, count the cumulative gaze time and the number of return views of each artistic element, and construct the attention weight vector , in, is the duration of looking at element i, is the number of replays, For the total gaze duration, an LSTM temporal model is constructed to process the gaze path sequence composed of color, composition, and symbols, capturing the deep interest pattern and outputting the temporal attention weight. a, combined with temporal attention weight Calculate the attention score of each art element i using the following formula: , in, is the cumulative gaze duration of element i, is the number of times the user looks back at i, The temporal attention weight output by the LSTM temporal model is used to reflect the importance of elements in the browsing path, and then the sensitivity of the attention weight is adjusted according to the user's cultural parameter UAI. , where β is a learnable parameter that controls the intensity of cultural influence; the hybrid recommendation unit is connected to the attention calculation unit, and the hybrid recommendation unit uses the attention gating mechanism to dynamically fuse the traditional recommendation score with the attention weight. The formula is as follows: ,in, is the collaborative filtering score based on user graph similarity, is the content similarity score, γ is the dynamic gating coefficient, which is determined by the confidence of eye movement data, ,in, is the total fixation duration of the current session, is the historical average gaze duration. When the user's concentration is higher than the average level, γ is reduced to enhance the content weight. Then, the DDPG algorithm is used to combine the real-time feedback of clicks and stays with the attention signal to perform reinforcement learning to adjust the recommendation ranking and obtain the recommendation list. The formula is as follows: ,in, To balance the coefficient, the KL divergence term penalizes the deviation of the recommendation results from the user's long-term interests to avoid over-catering to instantaneous attention. Finally, the digital media art of interest to the user is recommended based on the obtained recommendation list.

[0017] Furthermore, the intelligent recommendation module also includes a user selection unit, which is connected to the user recommendation unit. The user selection unit is used to select the corresponding digital media art work according to the user's selection trigger signal, and send the selection information to the interactive experience module for execution.

[0018] It can be seen from the above technical solution that the advantages of the present invention are: This application utilizes a multimodal perception module and a dynamic content generation module to locate users in real time, capture and identify micro-movements, and dynamically adjust content based on interactive actions, transforming users from passive recipients into active participants in artworks, enhancing interactivity and engagement. The intelligent recommendation module constructs a user knowledge graph and, in conjunction with a cultural dimension model, dynamically adjusts artistic symbols and guidance prompts to recommend content of interest to users, meeting their personalized needs and enhancing cultural adaptability. Furthermore, the dynamic content generation module constructs a parametric model library of artworks, deconstructing multiple elements into adjustable parameters. This allows the system to autonomously optimize content based on audience behavior patterns, addressing the issue of rigid, monotonous content presentation and meeting the diverse aesthetic needs of contemporary audiences. Furthermore, it enables precise behavioral semantic analysis and emotion recognition, overcoming the limitations of existing systems that rely on single sensor technology. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The drawings that constitute a part of this application are used to provide further understanding of this application. The illustrative embodiments of this application and their descriptions are used to explain this application and do not constitute improper limitations on this application.

[0020] Figure 1 This is a schematic diagram of the composition structure of this application.

[0021] Figure 2 Schematic diagram of the behavior semantic recognition process of this embodiment.

[0022] Figure 3 Schematic diagram of the dynamic content evolution process of this embodiment.

[0023] Figure 4 Schematic diagram of the interaction process of this embodiment. DETAILED DESCRIPTION

[0024] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail in conjunction with the embodiments and drawings. Here, the illustrative embodiments of this application and their descriptions are used to explain this application, but are not intended to limit this application.

[0025] Digital media art is an art form that combines modern technologies such as multimedia, network technology, and computer graphics. It not only covers traditional visual art forms such as painting, sculpture, and photography, but also includes a variety of expression methods such as sound, animation, and virtual reality. Traditional digital media art is a one-way output model, and the user as the "audience" lacks artistic experience. Participation is low, and the content presentation method is rigid. There is also a lack of in-depth understanding of user behavior and insufficient cultural adaptability, making it difficult to meet the needs of global exhibitions. Figures 1 to 4 In the illustrated embodiment, the digital media art interactive experience system was developed based on a customized system framework. The system specifically includes a multimodal perception module, a dynamic content generation module, an interactive experience module, and an intelligent recommendation module. The multimodal perception module integrates a multi-source environmental sensor array. Through collaborative perception technology using the AR device's inertial measurement unit and computer vision, it achieves millimeter-level positioning of the user's spatial coordinates and high-precision capture of gestures and other body language, transmitting real-time perception data to the dynamic content generation module. Based on parametric modeling theory, the dynamic content generation module constructs a library of artistic element models covering color space, geometric topology, and motion dynamics. Using a feature decoupling algorithm, it converts visual elements into a set of adjustable parameters. This module utilizes a deep reinforcement learning framework, using user interaction data as training samples to iteratively optimize an attention-based interaction strategy network, enabling dynamic and adaptive adjustment of the artistic content generation strategy. The interactive experience module communicates bidirectionally with the dynamic content generation module, utilizing multi-channel projection fusion technology to project digital art content and dynamically generated visual substrates into physical space. Furthermore, the module aligns the spatiotemporal alignment of AR virtual content with the physical projection, integrating spatial audio rendering and haptic feedback devices to create a multimodal immersive interactive scene. The intelligent recommendation module leverages knowledge graph construction technology to integrate user behavior data and cultural background information, constructing a multidimensional semantic network encompassing user interests, preferences, and cultural cognitive characteristics. Based on a theoretical model of cultural dimensions, it dynamically adjusts the semantic mapping of artistic symbols and interaction guidance strategies. Furthermore, through a hybrid recommendation algorithm of collaborative filtering and deep learning, personalized digital media art content recommendations are achieved.

[0026] Specifically, the multimodal perception module includes a user verification unit, a positioning unit and a motion recognition unit; the user verification unit is used to use a facial feature extraction model based on the ResNet-50 model to extract features from the facial image of the user when entering the venue obtained by the face acquisition device, output a 512-dimensional feature vector, encrypt the output facial feature vector using a homomorphic encryption method, and establish a temporary user profile based on the encrypted facial features; the positioning unit is used to establish a three-dimensional spatial coordinate system of the exhibition hall through the scanning data of the exhibition hall, and construct a real-time coordinate mapping relationship between the physical space and the digital content, and obtain the user's current physical space coordinate position by detecting the user's three-dimensional position in real time according to the IMU module; the motion recognition unit is used to detect the user's basic body movements using a skeleton key point detection method, and track the user's hand joints using the OpenPose deep learning model.

[0027] In this embodiment, the facial acquisition device is generally installed at the entrance of the exhibition hall. When a user enters, a temporary user profile can be established through simple facial recognition, and the profile is automatically destroyed after being stored for 48 hours. The facial acquisition device includes a camera, an image recognizer, a communication module, and a power supply module. The image recognizer extracts the facial features of the user upon entry through the ResNet-50 model, outputs a 512-dimensional feature vector, and uploads it to the cloud server for user profile registration and information recording. Before positioning, it is necessary to obtain a three-dimensional scanning model of the exhibition hall using a laser scanning device, and select a point as the origin based on the three-dimensional scanning model to establish a three-dimensional spatial coordinate system. The coordinates of each pixel point of the digital media art work are converted according to the preset coordinate conversion relationship, and the real-time position of the user is obtained based on the IMU module on the AR device worn by the user, and the user's current physical space coordinate position is calculated.

[0028] The multimodal perception module also includes a behavior semantic parsing unit, such as Figure 2 As shown in the behavior semantic recognition process diagram, the behavior semantic parsing unit is connected to the action recognition unit. The behavior semantic parsing unit is used to build a behavior pattern analysis model based on the LSTM network and the Transformer network, analyze the behavior target in combination with the scene context, and convert the user action into artistic adjustment parameters. The behavior pattern analysis model includes a time series modeling unit and an interaction intention prediction unit. The time series modeling unit is used to process the action sequence S using a bidirectional LSTM network and output the behavior classification probability distribution P. The interaction intention prediction unit is used to establish a behavior-intention mapping based on the Transformer model. ,in, is the contribution of actions at different time steps to the final intention, is a learnable parameter matrix, and d is the feature dimension. LSTM captures local temporal patterns, while Transformer captures global dependencies, improving the accuracy of behavior analysis.

[0029] In this embodiment, the structure of the bidirectional LSTM network includes an input layer, a forward LSTM, a backward LSTM, a hidden state, and a fully connected layer. The hidden layer dimension of each LSTM unit is h, and the output forward hidden state is and the backward hidden state , the final representation of each time step is , the input data is an action sequence , where each time step Contains user action features, such as joint coordinates, velocity, acceleration, etc., and the dimension is the number of time steps Feature dimension, the fully connected layer is used to transform Mapped to the behavioral category space, , Apply Softmax to each time step and output the behavior classification probability distribution The structure of the Transformer model is the fusion of the input layer, the self-attention layer, the linear layer and the fully connected layer. The input layer is used to transform the temporal feature matrix of the output of the bidirectional LSTM into The environmental parameters corresponding to the scene context are spliced ​​to form the enhanced feature N. The self-attention layer is used to calculate the query, key, and value. The linear layer fusion layer splices multiple single-head attention outputs. Through linear layer fusion, the fully connected layer maps the attention output to the intent space.

[0030] For example, the user continuously performs the action of “raising hand → pointing → grabbing”, with feature dimension d=30; then time series modeling: bidirectional LSTM output H∈ (h=128), the action classification layer determines the pointing action as "selection" (with a probability of 0.92). Intent prediction is then performed: the Transformer self-attention calculation indicates the intention is to "adjust the parameters of the selected object." Finally, the parameters are generated: the regression layer outputs the parameters that drive the dynamic adjustment of the artwork.

[0031] The dynamic content generation module is used to build a parametric model library for artworks, deconstructing the elements of color, form, and motion trajectory into adjustable parameters, and training the interactive strategy network through reinforcement learning, so that the system can autonomously optimize the content according to the audience's behavior patterns. The dynamic content generation module includes a parameter structure unit and a behavior-driven evolution unit; Figure 3The dynamic content evolution flow chart shown in the figure shows a parameter structure unit that uses a color rhythm encoder, a morphological topology analyzer, and motion dynamics modeling to analyze the original artistic material based on its color, morphology, and motion trajectory elements, obtaining adjustable parameters in the corresponding parameter space and constructing a parameterized art model library. The color rhythm encoder extracts the hue standard deviation Hstd, saturation gradient Sgrad, and lightness entropy Ventropy from the HSV space; the morphological topology analyzer constructs morphological feature vectors; and the motion dynamics modeling analyzes the trajectory frequency domain characteristics based on the Hamilton equation, outputting the dominant frequency and chaos index, thereby constructing a parameterized model library.

[0032] The behavior-driven evolution unit is used to generate the user's real-time behavior feature vector and ambient context parameters Generate state representation through the generative adversarial network GAN, and output parameter increments through the policy network using the TD3 algorithm , then the rendering engine applies the new parameters to generate content and according to the reward function: Update the policy network, where U() is the real-time behavior matching, S() is the cosine similarity of parameter changes, and C() is the aesthetic term. The pre-trained aesthetic evaluation model is used for scoring, and the parameter change speed constraint is set to ensure a continuous transition of artistic styles. The constraint formula is: .

[0033] In this embodiment, the dynamic content generation module also includes a multi-person interaction management unit, which is used to implement a dynamic resource allocation method based on Q-Learning, give priority to responding to highly engaged users, and use the Shapley value algorithm to fairly allocate control rights when conflicts in interactive actions occur.

[0034] In this embodiment, the interactive experience module is used to project the projection base of the digital media art content and the generated dynamic content into the physical space through the projection device, and to superimpose the virtual content generated in real time and synchronized with the physical projection onto the projection base through the AR device, and to provide real-time feedback on the audio and tactile information. Specifically, the interactive experience module includes a projection control unit, an AR interaction control unit, and a feedback adjustment unit. The projection control unit is used to project the projection base of the digital media art content and the generated dynamic content into the physical space through the projection device; the AR interaction control unit is used to superimpose the virtual content generated in real time and synchronized with the physical projection onto the projection base through the AR device through the timestamp synchronization protocol, and to provide real-time feedback on the audio and tactile information; the feedback adjustment unit is connected to the AR interaction control unit, and the feedback adjustment unit is used to dynamically adjust the tactile feedback intensity according to the amplitude of the movement. The adjustment formula is: ,in, is the tactile feedback intensity, is the scaling factor, is the sum of the squares of the weights of all action components, which is used to reflect the overall amplitude of the action.

[0035] After the user creates a temporary profile, the intelligent recommendation module loads their historical behavior and initial values ​​of cultural dimensions, transmits gaze data through the eye-tracking camera on the AR device, analyzes the gaze data, adjusts the weight of the corresponding cultural or artistic symbols based on the analysis structure, and recommends related digital media art works.

[0036] Specifically, the intelligent recommendation module includes a knowledge graph construction unit, an adjustment prompt unit, and a user recommendation unit. The knowledge graph construction unit is used to identify artistic elements, users, eye movement patterns, and cultural dimensions in the original art materials as core entities, define relationships between these entities, and implement triple storage and update management within the constructed knowledge graph using a Neo4j graph database. Artistic elements include color, composition, and symbols; users include demographic attributes such as user ID, region, and age, as well as behavioral labels; eye movement patterns include gaze duration, number of return glances, and scan paths, such as "gaze-jump-return"; and cultural dimensions include quantitative values ​​for six dimensions based on the Hofstede model (power distance, individualism, etc.). Relationships are defined as: user-[preference]-artistic element; artistic element-[cultural affiliation]-cultural dimension; user-[trigger]-eye movement pattern. Triple storage formats consist of a head entity, a relationship, and a tail entity, such as user A, preference, and ink painting style. Embedding representation learning is implemented using the TransE algorithm, capturing the implicit association between "cultural similarity → similar artistic style."

[0037] The adjustment prompt unit is connected to the knowledge graph construction unit. The adjustment prompt unit is used to associate cultural dimensions with artistic symbols and establish a mapping relationship according to the formula: ,in, is the cultural dimension coefficient dynamically adjusted by Q-learning, s is the Sigmoid function, which compresses the weight to [0,1]. is the symbol-culture association matrix, k is the number of artistic symbols, and d is the number of cultural dimensions. The symbol weights are dynamically calculated, and each coefficient is dynamically adjusted through Q-learning. Fuzzy logic is used to handle multi-dimensional cultural feature contradictions. It is also used to automatically increase the weight of the cultural openness dimension when the user repeatedly gazes at cross-cultural conflict elements, and adjust the significance of artistic symbols according to the distribution of gaze hotspots.

[0038] The user recommendation unit is connected to the knowledge graph construction unit. The user recommendation unit includes an eye tracking unit, an attention calculation unit and a hybrid recommendation unit. The eye tracking unit is used to obtain the eye movement trajectory and gaze hot zone, and synchronize the projection screen coordinate mapping, establish the correspondence between the gaze point and the pixel of the digital art canvas, and automatically identify the artistic elements in the painting through the image segmentation algorithm. The attention calculation unit is connected to the eye tracking unit. The attention calculation unit is used to model the interest of artistic elements, count the cumulative gaze time and the number of return views of each artistic element, and construct the attention weight vector , in, is the duration of looking at element i, is the number of replays, For the total gaze duration, an LSTM temporal model is constructed to process the gaze path sequence composed of color, composition, and symbols, capturing the deep interest pattern and outputting the temporal attention weight. a, combined with temporal attention weight Calculate the attention score of each art element i using the following formula: , in, is the cumulative gaze duration of element i, is the number of times the user looks back at i, The temporal attention weight output by the LSTM temporal model is used to reflect the importance of elements in the browsing path, and then the sensitivity of the attention weight is adjusted according to the user's cultural parameter UAI. , where β is a learnable parameter that controls the intensity of cultural influence; the hybrid recommendation unit is connected to the attention calculation unit, and the hybrid recommendation unit uses the attention gating mechanism to dynamically fuse the traditional recommendation score with the attention weight. The formula is as follows: ,in, is the collaborative filtering score based on user graph similarity, is the content similarity score, γ is the dynamic gating coefficient, which is determined by the confidence of eye movement data, ,in, is the total fixation duration of the current session, is the historical average gaze duration. When the user's concentration is higher than the average level, γ is reduced to enhance the content weight. Then, the DDPG algorithm is used to combine the real-time feedback of clicks and stays with the attention signal to perform reinforcement learning to adjust the recommendation ranking and obtain the recommendation list. The formula is as follows: ,in, To balance the coefficient, the KL divergence term penalizes the deviation of the recommendation results from the user's long-term interests to avoid over-catering to instantaneous attention. Finally, the digital media art of interest to the user is recommended based on the obtained recommendation list.

[0039] In this embodiment, the intelligent recommendation module also includes a user selection unit, which is connected to the user recommendation unit. The user selection unit is used to select the corresponding digital media art work according to the user's selection trigger signal, and send the selection information to the interactive experience module for projection adjustment.

[0040] like Figure 4 As shown, the interaction process of this embodiment is as follows: Upon entering, a temporary user profile is created through facial recognition. Inertial sensors capture initial motion patterns and generate basic interaction parameters. A multimodal perception module receives motion commands and adjusts the trajectory of the art particles through behavioral semantic analysis. The AR device generates an augmented reality layer synchronized with the physical projection in real time. A cultural adaptation system automatically switches the guide language and recommends content based on language and cultural dimensions and user behavior. This invention can significantly enhance user engagement and satisfaction with digital media artworks, thereby promoting a deeper understanding and emotional resonance of the culture.

[0041] The above description is merely a preferred embodiment of the present application and is not intended to limit the present application. Those skilled in the art will appreciate that various modifications and variations of the present embodiment are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present application shall be included within the scope of protection of the present application.

Claims

1. A digital media art interactive experience system, characterized by: include: Multimodal perception module, dynamic content generation module, interactive experience module and intelligent recommendation module; The multimodal perception module is used to perform real-time positioning and micro-motion capture of the user, and send the identification information to the dynamic content generation module; The dynamic content generation module is used to build a parametric model library of artworks, deconstructing the elements of color, form, and motion trajectory into adjustable parameters, and train the interactive strategy network through reinforcement learning, so that the system can autonomously optimize the content according to the audience's behavior patterns; The interactive experience module is connected to the dynamic content generation module, and is used to project the digital media art content and the projection substrate corresponding to the generated dynamic content into the physical space through a projection device, and to superimpose the virtual content generated in real time and synchronized with the physical projection onto the projection substrate through an AR device, and to provide real-time feedback of audio and tactile information; The intelligent recommendation module is connected to the interactive experience module. The intelligent recommendation module is used to build a user knowledge graph, dynamically adjust artistic symbols and guidance prompts in combination with a cultural dimension model, and recommend digital media art of interest to users.

2. The digital media art interactive experience system according to claim 1, characterized in that: The multimodal perception module includes a user verification unit, a positioning unit and a motion recognition unit; The user verification unit is used to extract features from the facial image of the user entering the venue acquired by the facial acquisition device using a deep learning-based facial feature extraction model, encrypt the extracted facial features, and create a temporary user profile based on the encrypted facial features; The positioning unit is used to establish the exhibition hall's three-dimensional spatial coordinate system through the exhibition hall's scanning data, and to build a real-time coordinate mapping relationship between the physical space and the digital content, and to obtain the user's current physical space coordinate position by detecting the user's three-dimensional position in real time based on the IMU module; The motion recognition unit is used to detect the user's basic body movements using the skeleton key point detection method, and to track the user's hand joints using the OpenPose deep learning model.

3. The digital media art interactive experience system according to claim 2, characterized in that: The multimodal perception module also includes a behavior semantic parsing unit, which is connected to the action recognition unit. The behavior semantic parsing unit is used to construct a behavior pattern analysis model based on the LSTM network and the Transformer network, analyze the behavior target in combination with the scene context, and convert the user action into an artistic adjustment parameter. The behavior pattern analysis model includes a time series modeling unit and an interaction intention prediction unit. The time series modeling unit is used to process the action sequence S using a bidirectional LSTM network and output a behavior classification probability distribution P. The interaction intention prediction unit is used to establish a behavior-intention mapping based on the Transformer model. , where d is the feature dimension.

4. The digital media art interactive experience system according to claim 3, characterized in that: The dynamic content generation module includes a parameter structure unit and a behavior driven evolution unit; The parameter structure unit is used to analyze the characteristics of the original art material according to the color, shape, and motion trajectory elements using a color rhythm encoder, a morphological topology analyzer, and motion dynamics modeling, obtain adjustable parameters in the corresponding parameter space, and construct an art parameter model library; The behavior driven evolution unit is used to generate a real-time behavior feature vector of the user. and ambient context parameters Generate state representation through generative adversarial network GAN and output parameter increment through policy network , then the rendering engine applies the new parameters to generate content and according to the reward function: Update the policy network, where U() is the real-time behavior matching, S() is the cosine similarity of parameter changes, and C() is the aesthetic term. The pre-trained aesthetic evaluation model is used for scoring, and the parameter change speed constraint is set to ensure a continuous transition of artistic styles. The constraint formula is: .

5. The digital media art interactive experience system according to claim 4, characterized in that: The dynamic content generation module also includes a multi-person interaction management unit, which is used to implement a dynamic resource allocation method based on Q-Learning, give priority to responding to users with high participation, and use a Shapley value algorithm to fairly allocate control rights when conflicting interactive actions occur.

6. The digital media art interactive experience system according to claim 1, characterized in that: The interactive experience module includes a projection control unit, an AR interaction control unit, and a feedback adjustment unit. The projection control unit is used to project the digital media art content and the generated dynamic content projection substrate into the physical space through a projection device; The AR interaction control unit is used to overlay the virtual content generated in real time and synchronized with the physical projection onto the projection base through the AR device through a timestamp synchronization protocol, and provide real-time feedback of audio and tactile information; The feedback adjustment unit is connected to the AR interaction control unit and is used to dynamically adjust the tactile feedback intensity according to the action amplitude. The adjustment formula is: ,in, is the tactile feedback intensity, is the scaling factor, is the sum of the squares of the weights of all action components, which is used to reflect the overall amplitude of the action.

7. The digital media art interactive experience system according to claim 1, characterized in that: The intelligent recommendation module includes a knowledge graph construction unit, an adjustment prompt unit and a user recommendation unit; The knowledge graph construction unit is used to take the artistic elements, users, eye movement patterns and cultural dimensions in the original art materials as core entities, and define the relationship between the entities. The constructed knowledge graph uses the Neo4j graph database to implement triple storage and update management, and uses the TransE algorithm to implement embedded representation learning; The adjustment prompt unit is connected to the knowledge graph construction unit, and is used to associate cultural dimensions with artistic symbols and establish a mapping relationship according to the formula: , in, is the cultural dimension coefficient dynamically adjusted by Q-learning, s is the Sigmoid function, which compresses the weight to [0,1]. is the symbol-culture association matrix, where k is the number of artistic symbols and d is the number of cultural dimensions. Symbol weights are dynamically calculated, and each coefficient is dynamically adjusted through Q-learning. Fuzzy logic is also used to handle multi-dimensional cultural feature conflicts. It is also used to automatically increase the weight of the cultural openness dimension when users repeatedly gaze at cross-cultural conflicting elements, and to adjust the salience of artistic symbols based on the distribution of gaze hotspots. The user recommendation unit is connected to the knowledge graph construction unit, and the user recommendation unit includes an eye tracking unit, an attention calculation unit and a hybrid recommendation unit. The eye tracking unit is used to obtain the eye movement trajectory and the gaze hot zone, and synchronize the projection screen coordinate mapping, establish the correspondence between the gaze point and the digital art canvas pixel, and automatically identify the art elements in the painting through the image segmentation algorithm. The attention calculation unit is connected to the eye tracking unit and is used to perform art element interest modeling, count the cumulative gaze time and the number of return glances for each art element, and construct the attention weight vector , in, is the duration of looking at element i, is the number of replays, For the total gaze duration, an LSTM temporal model is constructed to process the gaze path sequence composed of color, composition, and symbols, capturing the deep interest pattern and outputting the temporal attention weight. a, combined with temporal attention weight Calculate the attention score of each art element i using the following formula: , in, is the cumulative gaze duration of element i, is the number of times the user looks back at i, The temporal attention weight output by the LSTM temporal model is used to reflect the importance of elements in the browsing path, and then the sensitivity of the attention weight is adjusted according to the user's cultural parameter UAI. , where β is a learnable parameter that controls the intensity of cultural influence; the hybrid recommendation unit is connected to the attention calculation unit, and the hybrid recommendation unit uses the attention gating mechanism to dynamically fuse the traditional recommendation score with the attention weight. The formula is as follows: ,in, is the collaborative filtering score based on user graph similarity, is the content similarity score, γ is the dynamic gating coefficient, which is determined by the confidence of eye movement data, ,in, is the total fixation duration of the current session, is the historical average gaze duration. When the user's concentration is higher than the average level, γ is reduced to enhance the content weight. Then, the DDPG algorithm is used to combine the real-time feedback of clicks and stays with the attention signal to perform reinforcement learning to adjust the recommendation ranking and obtain the recommendation list. The formula is as follows: ,in, To balance the coefficient, the KL divergence term penalizes the deviation of the recommendation results from the user's long-term interests to avoid over-catering to instantaneous attention. Finally, the digital media art of interest to the user is recommended based on the obtained recommendation list.

8. The digital media art interactive experience system according to claim 7, characterized in that: The intelligent recommendation module also includes a user selection unit, which is connected to the user recommendation unit. The user selection unit is used to select a corresponding digital media art work according to a user's selection trigger signal and send the selection information to the interactive experience module for projection adjustment.

Citation Information

Cited By

  • Interactive processing method and system for digital media file

    CN121433554A