Edge-computing-based real-time scene experience system and method for creative and cultural products
By using edge computing technology, multimodal experience data is collected synchronously and parallel emotion recognition is performed. Combined with user digital profiles, a matching sequence of cultural and creative content interaction is generated, which solves the problem of insufficient real-time emotional interaction continuity in existing technologies and achieves efficient emotion-preference coupling content adaptation.
Patent Information
- Application Number
- CN202511374418.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-25
- Publication Date
- 2025-10-31
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In existing real-time experience solutions for cultural and creative products, the latency of processing multimodal data in the cloud is too high, the accuracy of single-modal emotion recognition is low, and the push of cultural and creative content does not take into account the user's historical behavior preferences, resulting in insufficient continuity of real-time emotional interaction.
Based on edge computing, multimodal experience data streams (facial expressions, voice signals, and body movements) are collected synchronously. Multi-threaded parallel emotional state recognition is performed through the frame buffer queue on the edge server. Combined with the user's digital profile, a sequence of cultural and creative content interaction matching the user's emotional state is generated, and real-time rendering and push are performed through high-performance rendering nodes in the cloud.
It achieves low-latency emotion recognition and content response, improves the accuracy and comprehensiveness of emotion recognition, ensures the content adaptation and rendering stream edge adaptation transmission of emotion-preference coupling in the real-time scenario-based experience of cultural and creative products, and guarantees the continuity of real-time emotional interaction.
Smart Images

Figure CN120873296A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of distributed resource management technology, and more specifically, to a real-time scenario-based experience system and method for cultural and creative products based on edge computing. Background Technology
[0002] With the deep integration of the cultural and creative industries and digital technologies, cultural and creative products are transforming from traditional static displays to real-time, scenario-based, and emotionally interactive experiences. Users' demands for personalized adaptation, real-time response speed, and emotional resonance in cultural and creative experiences have significantly increased.
[0003] Existing real-time experience solutions for cultural and creative products mostly adopt a model of "centralized cloud processing + single-modal emotion recognition + fixed content push". That is, the cloud server receives and processes the user's single-modal data (such as facial expressions), generates cultural and creative content based on preset rules, and then pushes the rendered content directly to the terminal. However, this model has three major technical problems: First, when processing multimodal data in the cloud, the long network transmission link leads to excessive data processing latency, which cannot meet the needs of real-time interaction. Second, single-modal emotion recognition relies on only one dimension of data, which is prone to incomplete information and low accuracy in judging emotional state, making it difficult to accurately capture the user's true emotions. Third, the push of cultural and creative content does not take into account the user's historical behavior preferences, but is only generated according to a general template, resulting in insufficient adaptability to the user's real-time emotional state. Moreover, the rendering stream is directly pushed from the cloud, which is susceptible to network fluctuations, resulting in stuttering or frame order disorder. Therefore, how to achieve content adaptation that couples emotion and preference and edge adaptation of the rendering stream in the real-time scenario-based experience of cultural and creative products to ensure the continuity of real-time emotional interaction has become a challenge for the industry. Summary of the Invention
[0004] This application provides a real-time contextualized experience system and method for cultural and creative products based on edge computing. It can realize content adaptation and rendering stream edge adaptation transmission with emotion-preference coupling in the real-time contextualized experience of cultural and creative products to ensure the continuity of real-time emotional interaction.
[0005] Firstly, this application provides a method for real-time contextualized experience of cultural and creative products based on edge computing, achieving low-latency emotion recognition and content response when target users engage in real-time contextualized experience of cultural and creative products. The method includes: Simultaneously collect facial expressions, voice signals, and body movements of target users during real-time scenario-based experiences with cultural and creative products, thereby obtaining multimodal experience data streams of target users; The frame buffer queue built on the edge server performs multi-threaded parallel emotional state recognition on the multimodal experience data stream to obtain all emotional feature vectors of the target user in the process of experiencing cultural and creative products, and then constructs a dynamic evolution trajectory of the target user's emotional state changes based on all emotional feature vectors. The behavioral preference features of target users' cultural and creative product experiences are extracted from the pre-constructed digital user profiles. Then, based on the dynamic evolution trajectory and the behavioral preference features, a sequence of cultural and creative content interactions that matches the target user's emotional state is generated. The interactive sequence of cultural and creative content is rendered in real time by a high-performance rendering node in the cloud and the rendering stream is pushed to the interactive terminal through an edge server, thereby enabling the cultural and creative product content experience response to interact emotionally with the target user in the interactive terminal.
[0006] In some embodiments, the simultaneous collection of facial expressions, voice signals, and body movements of target users during real-time contextualized experiences with cultural and creative products, thereby obtaining a multimodal experience data stream of the target users, specifically includes: By synchronously capturing the facial expressions and body movements of target users during real-time scenario-based experiences of cultural and creative products using a visual sensor array, a sequence of facial expression video streams and body movement depth maps is obtained. The voice signal stream is obtained by collecting the voice signal of the target user during the real-time scenario-based experience of cultural and creative products through a high-fidelity microphone array. The facial expression video stream, the motion depth map sequence, and the speech signal stream are filtered and aligned to generate a multimodal experience data stream for the target user.
[0007] In some embodiments, the multi-threaded parallel emotional state recognition of the multimodal experience data stream is performed based on a frame buffer queue built on an edge server to obtain all emotional feature vectors of the target user during the experience of cultural and creative products, specifically including: Create a frame buffer queue and a multimodal emotion recognition thread pool on the edge server side; The multimodal experience data stream is time-sliced and classified and cached into a frame buffer queue according to facial expression frames, voice feature frames, and body motion frames; For each time slice, the recognition thread in the multimodal emotion recognition thread pool performs emotion recognition on the facial expression frames, speech feature frames and body motion frames in the time slice of the frame buffer queue, and outputs facial sub-vectors, speech sub-vectors and body motion sub-vectors. The facial sub-vector, the voice sub-vector, and the tactile sub-vector are fused to generate the target user's emotional feature vector in each time slice, thereby obtaining the target user's emotional feature vector in all time slices during the experience of cultural and creative products.
[0008] In some embodiments, constructing a dynamic evolution trajectory of the target user's emotional state changes based on all emotional feature vectors specifically includes: Temporal calibration is performed on all sentiment feature vectors to obtain a temporal calibration set of sentiment features; Based on the emotional feature time-series calibration set, establish the emotional state coordinates of the target user at each time point during the experience of cultural and creative products; Trend fitting is performed on the emotional state coordinates at all time points to generate a dynamic evolution trajectory of the target user's emotional state changes.
[0009] In some embodiments, extracting behavioral preference features of target users' experiences with cultural and creative products from pre-built user digital profiles specifically includes: Call all user digital profiles from the pre-built cultural and creative behavior profile library; Based on all user digital profiles, target user experience and historical behavioral data of cultural and creative products are used to identify the target user experience. The core preferences of the target users are extracted from the historical behavioral data, thereby obtaining the behavioral preference characteristics of the target users' experience with cultural and creative products.
[0010] In some embodiments, generating a cultural and creative content interaction sequence that matches the target user's emotional state based on the dynamic evolution trajectory and the behavioral preference features specifically includes: From the dynamic evolution trajectory, core preferences in the synchronously associated behavioral preference features of emotional mutation nodes are screened, thereby determining all emotional guidance anchors of the target user in the process of experiencing cultural and creative products; Map all emotionally driven anchors to the experiential content of cultural and creative products, and output a set of suitable candidate cultural and creative content. The candidate cultural and creative content set is sorted according to the emotional change trend of the dynamic evolution trajectory to generate a cultural and creative content interaction sequence that matches the emotional state of the target user.
[0011] In some embodiments, rendering the cultural and creative content interaction sequence in real time using a high-performance rendering node in the cloud and pushing the rendering stream to the interactive terminal via an edge server specifically includes: The interaction sequence of the cultural and creative content is analyzed to obtain multiple rendering sub-tasks with a unified time base. The dynamic rendering optimization engine is started to render all rendering subtasks and the rendered result frames are passed to the frame buffer synchronization queue for timing calibration, resulting in a timing-coherent rendering stream. The rendering stream is pushed to the target user's interactive terminal via an edge server.
[0012] In some embodiments, the content experience response of cultural and creative products that engages in emotional interaction with target users in an interactive terminal specifically includes: After receiving the sequentially rendered stream, the edge server calls the intelligent encoding adaptation module to encode the data according to the type of interactive terminal. The edge-terminal timing phase-locked mechanism ensures that the timing deviation between the dynamically adjusted content displayed on the interactive terminal and the user's emotional feedback is less than a preset deviation threshold.
[0013] In some embodiments, the multimodal emotion recognition thread pool includes a facial expression recognition thread, a voice emotion recognition thread, and a body motion recognition thread.
[0014] Secondly, this application provides a real-time scenario-based experience system for cultural and creative products based on edge computing, including: The data acquisition module is used to simultaneously collect facial expressions, voice signals and body movements of target users when they experience cultural and creative products in a real-time contextualized manner, thereby obtaining a multimodal experience data stream of the target users; The processing module is used to perform multi-threaded parallel emotional state recognition on the multimodal experience data stream based on the frame buffer queue built on the edge server, to obtain all emotional feature vectors of the target user in the process of experiencing cultural and creative products, and then to construct a dynamic evolution trajectory of the target user's emotional state changes based on all emotional feature vectors. The processing module is used to extract behavioral preference features of target users’ cultural and creative product experiences from the pre-constructed user digital profile, and then generate a cultural and creative content interaction sequence that matches the target user’s emotional state based on the dynamic evolution trajectory and the behavioral preference features. The execution module is used to render the cultural and creative content interaction sequence in real time through a high-performance rendering node in the cloud and push the rendering stream to the interactive terminal through an edge server, thereby enabling the cultural and creative product content experience response to interact emotionally with the target user in the interactive terminal.
[0015] The technical solutions provided by the embodiments disclosed in this application have the following beneficial effects: The edge computing-based real-time scenario-based experience system and method for cultural and creative products provided in this application first synchronously collects facial expressions, voice signals, and body movements of the target user during the real-time scenario-based experience of cultural and creative products, thereby obtaining a multimodal experience data stream of the target user; based on a frame buffer queue built on the edge server, the multimodal experience data stream is subjected to multi-threaded parallel emotional state recognition to obtain all emotional feature vectors of the target user during the experience of cultural and creative products, and then a dynamic evolution trajectory of the target user's emotional state changes is constructed based on all emotional feature vectors; behavioral preference features of the target user's experience of cultural and creative products are extracted from the pre-built user digital profile, and then a cultural and creative content interaction sequence matching the target user's emotional state is generated based on the dynamic evolution trajectory and the behavioral preference features; the cultural and creative content interaction sequence is rendered in real time through a high-performance rendering node in the cloud and the rendering stream is pushed to the interactive terminal through the edge server, thereby enabling the cultural and creative product content experience response that interacts with the target user emotionally in the interactive terminal.
[0016] Therefore, this application utilizes high-performance cloud rendering nodes to render the cultural and creative content interaction sequence in real time and pushes the rendering stream to the interactive terminal via an edge server, thereby enabling the cultural and creative product content experience response to interact emotionally with the target user on the interactive terminal. First, the multimodal experience data stream is determined to be a fused data sequence recording the facial expressions, voice signals, and body movements of the target user during real-time scenario-based experience of the cultural and creative product. The determination of the multimodal experience data stream involves simultaneously collecting the target user's facial expression video stream (i.e., visual dimension), voice signal stream (i.e., auditory dimension), and body movement depth map sequence (i.e., movement dimension), and then filtering, aligning, and fusing them into a unified data sequence. This provides a three-dimensional emotional feature input of "visual-auditory-movement," thereby improving the accuracy and comprehensiveness of emotion recognition through multi-dimensional data fusion, thus avoiding the limitations of single-modal data. Then, the dynamic evolution trajectory is determined to obtain a curve showing the emotional changes of the target user throughout the entire cultural and creative experience. The determination of the dynamic evolution trajectory involves time-series calibration, emotional state coordinate establishment, and trend simulation of the full-time emotional feature vector obtained through multi-threaded parallel recognition on the edge server. The generated emotional change curve upgrades "static emotional judgment" to "dynamic emotional tracking," clearly reflecting the user's emotional tendencies at different stages of the experience (such as increased arousal during haptic interaction and fluctuations in pleasure during artifact tours), and even accurately pinpointing moments of emotional abrupt change (such as a sudden increase in advantage when a certain artifact's details are displayed), fundamentally solving the technical pain point of "inability to capture dynamic emotional changes." Finally, determining the interaction sequence of cultural and creative content yields a sequence of cultural and creative experience content generated based on the target user's emotional change trends and behavioral preferences during the experience of cultural and creative products. The determination of the cultural and creative content interaction sequence is based on both dynamic evolution trajectory (real-time emotions) and behavioral preference characteristics (historical needs) to generate a candidate content set, and sorting it according to emotional change trends improves the adaptability of cultural experience content with user emotions and preferences, fundamentally solving the technical pain point of existing solutions that only generate cultural and creative content based on preset general templates without combining the user's real-time emotional state and historical behavioral preferences, resulting in insufficient content adaptability. In summary, based on the above solutions, emotional-preference coupled content adaptation and rendering stream edge adaptation transmission can be achieved in the real-time scenario-based experience of cultural and creative products to ensure the continuity of real-time emotional interaction. Attached Figure Description
[0017] Figure 1 This is an exemplary flowchart of a real-time contextualized experience method for cultural and creative products based on edge computing, as shown in some embodiments of this application. Figure 2 This is an exemplary flowchart illustrating the determination of sentiment feature vectors according to some embodiments of this application; Figure 3This is a flowchart illustrating the operation of determining the interactive sequence of cultural and creative content according to some embodiments of this application; Figure 4 This is a schematic diagram of the structure of a real-time scenario-based experience system for cultural and creative products based on edge computing, as shown in some embodiments of this application. Figure 5 This is an internal structural diagram of a computer device that implements a real-time scenario-based experience method for cultural and creative products based on edge computing, according to some embodiments of this application. Detailed Implementation
[0018] To better understand the technical solution of this application, the technical solution of this application will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0019] refer to Figure 1 The figure is an exemplary flowchart of a real-time contextualized experience method for cultural and creative products based on edge computing, according to some embodiments of this application. The real-time contextualized experience method for cultural and creative products based on edge computing mainly includes the following steps: In step 101, facial expressions, voice signals and body movements of the target user are collected simultaneously during the real-time scenario-based experience of cultural and creative products, thereby obtaining the multimodal experience data stream of the target user.
[0020] In some embodiments, the simultaneous acquisition of facial expressions, voice signals, and body movements of target users during real-time contextualized experiences with cultural and creative products, thereby obtaining a multimodal experience data stream of the target users, can be achieved through the following steps: By synchronously capturing the facial expressions and body movements of target users during real-time scenario-based experiences of cultural and creative products using a visual sensor array, a sequence of facial expression video streams and body movement depth maps is obtained. The voice signal stream is obtained by collecting the voice signal of the target user during the real-time scenario-based experience of cultural and creative products through a high-fidelity microphone array. The facial expression video stream, the motion depth map sequence, and the speech signal stream are filtered and aligned to generate a multimodal experience data stream for the target user.
[0021] In practical implementation, the facial expressions and body movements of the target user during real-time scenario-based experience of cultural and creative products are simultaneously captured by a visual sensor array. The resulting facial expression video stream and body movement depth map sequence can be achieved using a two-camera visual sensor array. One high dynamic range camera focuses on the target user's facial area to acquire facial images and extract facial feature points (such as eye corner coordinates, mouth corner coordinates, etc.) to capture micro-expressions. Existing video coding optimization techniques (such as intra-frame compression algorithms) are used to compress the acquired facial image frames to generate a facial expression video stream. The other high dynamic range camera focuses on the target user's limb area to acquire body movement images and extracts the three-dimensional coordinates of the target user's core limb joints (such as elbow, knee, hip joints, etc.) based on structured light depth perception technology, thereby outputting the body movement depth map sequence. The system generates a motion depth map sequence, frame by frame. Simultaneously, local clock synchronization technology at edge nodes adds a unified timestamp to each frame of the facial expression video stream and the motion depth map sequence, ensuring spatiotemporal alignment of the two types of data. The facial expression video stream is a dynamic image sequence capturing micro-expression changes in the face of a target user during a real-time, scenario-based experience with cultural and creative products. This provides visual evidence of facial emotional characteristics, accurately capturing subtle expressions such as smiles and frowns, providing intuitive and rich visual features for emotion recognition. The motion depth map sequence is a depth image sequence recording the three-dimensional spatial position changes of the target user's limb movements during a real-time, scenario-based experience with cultural and creative products. This motion depth map sequence reflects the emotional tendency of the target user's limb movements, supplementing emotional expression with information such as movement amplitude and rate, thereby enhancing the dynamism and realism of emotion recognition.
[0022] In specific implementation, the voice signal of the target user during real-time scenario-based experience of cultural and creative products is collected through a high-fidelity microphone array. The resulting voice signal stream can be achieved in the following way: a 4-unit cardioid high-fidelity microphone array (with a sampling rate of 48 kHz) deployed around the cultural and creative experience interactive terminal (which can be arranged in a regular quadrilateral pattern with a spacing of 50 cm) can be used to collect the voice signal of the target user during real-time scenario-based experience of cultural and creative products. A beamforming algorithm is then used to suppress environmental noise (such as background noise in the exhibition hall) and enhance the voice signal of the target user to generate a voice signal stream. The voice signal stream is a sequence of user voice sound wave signals during real-time scenario-based experience of cultural and creative products. This voice signal stream can provide emotional information contained in the voice, reflecting the user's emotional state through features such as tone and speech rate, and working in conjunction with visual and motion data to improve the completeness of emotion recognition.
[0023] In specific implementation, the facial expression video stream, the motion depth map sequence, and the speech signal stream are filtered and aligned to generate the multimodal experience data stream for the target user. This can be achieved in the following way: First, the facial expression video stream can be Gaussian filtered (e.g., window size 3×3, standard deviation 1.2) to remove image noise. The motion depth map sequence can be Kalman filtered (state equation is "current joint coordinate = previous moment coordinate + action speed × time interval", observation equation is "observed coordinate = true coordinate + Gaussian noise") to smooth the action trajectory. The speech signal stream can be filtered using a short-time Fourier transform algorithm to remove environmental noise. Then, based on a unified timestamp, the filtered facial expression video stream, motion depth map sequence, and speech signal stream can be imported into an edge server and aligned using a queue scheduling algorithm (e.g., arranged in ascending order of timestamp, discarding data frames with timestamp deviations exceeding 5 milliseconds) to generate the multimodal experience data stream for the target user.
[0024] It should be noted that in this application, the multimodal experience data stream is a fusion data sequence that records the facial expressions, voice signals and body movements of the target user when experiencing cultural and creative products in a real-time contextualized manner. This multimodal experience data stream can provide comprehensive input for emotion state recognition, and improve the accuracy and comprehensiveness of emotion recognition through multi-dimensional data fusion, thereby avoiding the limitations of single-modal data.
[0025] In step 102, the multi-threaded parallel emotional state recognition of the multimodal experience data stream is performed based on the frame buffer queue built on the edge server to obtain all emotional feature vectors of the target user in the process of experiencing cultural and creative products, and then the dynamic evolution trajectory of the target user's emotional state changes is constructed based on all emotional feature vectors.
[0026] In some embodiments, reference Figure 2 The diagram is an exemplary flowchart illustrating the determination of emotional feature vectors according to some embodiments of this application. In this application, the multi-threaded parallel emotional state recognition of the multimodal experience data stream, based on a frame buffer queue constructed on an edge server, to obtain all emotional feature vectors of the target user during the experience of cultural and creative products, can be achieved through the following steps: In step 1021, a frame buffer queue and a multimodal emotion recognition thread pool are created on the edge server side; In step 1022, the multimodal experience data stream is time-sliced and classified and cached into a frame buffer queue according to facial expression frames, voice feature frames and body motion frames; In step 1023, for each time slice, the recognition thread in the multimodal emotion recognition thread pool performs emotion recognition on the facial expression frames, speech feature frames and body motion frames in the time slice of the frame buffer queue, and outputs facial sub-vectors, speech sub-vectors and body motion sub-vectors. In step 1024, the facial sub-vector, the voice sub-vector, and the tactile sub-vector are fused to generate the target user's emotional feature vector in time slices, thereby obtaining the target user's emotional feature vector in all time slices during the experience of cultural and creative products.
[0027] It should be noted that in this application, the frame buffer queue is an ordered storage structure for caching multimodal data frames on the edge server side. This frame buffer queue can realize the temporary storage and ordered scheduling of multimodal data, avoid data congestion and ensure timing consistency, and provide stable input for parallel recognition. The multimodal emotion recognition thread pool is a collection of threads at the edge end that process different modalities of emotion recognition, including three types of dedicated recognition threads: facial expression recognition thread, voice emotion recognition thread, and body motion recognition thread. The multimodal emotion recognition thread pool can execute facial, voice, and body motion emotion analysis tasks in parallel, thereby improving the efficiency of emotion recognition and meeting real-time requirements.
[0028] In specific implementation, the creation of frame buffer queues and multimodal emotion recognition thread pools on the edge server can be achieved in the following way: A multimodal adaptation frame buffer queue can be built on the edge server, using the programming language C++ to encapsulate a thread-safe queue structure, setting three modal partitions: face, voice, and body sensing. The maximum cache capacity of each partition is configured according to "100 frames / second of single-modal data" to avoid data overflow. At the same time, a multimodal emotion recognition thread pool is created and configured with three types of dedicated recognition threads. Among them, the facial expression recognition thread can call a pre-trained network based on the pleasure-arousal-dominance emotion dimension model for emotion recognition; the voice emotion recognition thread can be equipped with a voice intonation spectrum analysis algorithm (such as the Mel spectrum dynamic time warping algorithm) for emotion recognition; and the body sensing action recognition thread can integrate preset action-emotion mapping rules for emotion recognition. The multimodal emotion recognition thread pool achieves efficient resource utilization through a dynamic scheduling mechanism (such as prioritizing tasks for idle threads).
[0029] In specific implementation, the multimodal experience data stream can be time-sliced and classified into facial expression frames, voice feature frames, and motion-sensing frames and cached in a frame buffer queue in the following way: First, the multimodal experience data stream can be time-sliced in 10-millisecond units, and a unique timestamp can be added to each time slice using the local clock of the edge server to ensure data temporal continuity; then, the multimodal experience data within each time slice is split, namely: the facial expression video stream within the time slice is converted into facial expression frames according to the 68-point facial feature encoding rule (i.e., extracting key coordinates such as the corners of the eyes and mouth); the voice signal stream within the time slice is used to extract 13-dimensional emotional tone features through Mel-frequency cepstral coefficients to generate voice feature frames; and the motion-sensing depth map sequence within the time slice is parsed into a joint three-dimensional coordinate sequence (such as elbow and knee joint coordinates) to form motion-sensing frames; finally, the three types of data frames are classified according to modality type. Do not cache the corresponding partition in the frame buffer queue. During caching, a queue overflow detection mechanism (i.e., triggering data priority processing when the partition cache reaches 90%) can be used to avoid multimodal data accumulation. Among them, the facial expression frame is a static image frame containing the facial expression features of the target user when experiencing cultural and creative products in a real-time contextualized manner. As the basic unit of facial emotion analysis, the facial expression frame can capture the micro-expression details of the target user, thereby improving the accuracy of facial emotion recognition. The voice feature frame is an audio segment that carries the emotional feature information of the target user's voice when experiencing cultural and creative products in a real-time contextualized manner. The voice feature frame can reduce redundant information interference by focusing on key features such as tone and speech rate. The body motion frame is a three-dimensional data frame that records the instantaneous state of the target user's limbs when experiencing cultural and creative products in a real-time contextualized manner. The body motion frame can reflect the emotional tendency of the limb movements through quantitative data such as joint coordinates, thereby making the movement emotion analysis more accurate.
[0030] In specific implementation, the recognition threads in the multimodal emotion recognition thread pool perform emotion recognition on facial expression frames, speech feature frames, and haptic motion frames in the time slice of the frame buffer queue, and output facial sub-vectors, speech sub-vectors, and haptic motion sub-vectors. This can be achieved by using a thread scheduling algorithm (threads can be allocated according to the priority order of "face → speech → haptic motion") to call the three types of dedicated recognition threads in the multimodal emotion recognition thread pool for parallel recognition. Among them, the facial expression recognition thread can analyze the displacement and deformation of 68 facial feature points in the facial expression frame through a convolutional neural network, and output a 3D facial sub-vector containing pleasantness, arousal, and dominance; the speech emotion recognition thread can analyze the intonation fluctuation (i.e., fundamental frequency change rate) and speech rate (i.e., syllables per second) in the speech feature frame through a speech intonation spectrum analysis algorithm, and output a 2D speech sub-vector containing intonation fluctuation and speech rate; the haptic motion recognition thread... The system can calculate the amplitude and rate of joint movements in motion frames and output a 1D motion-sensation sub-vector containing excitation level through preset motion-emotion mapping rules (such as large-amplitude limb movements corresponding to high excitation level). The facial sub-vector represents the facial emotions of the target user during a real-time, scenario-based experience of cultural and creative products. This facial sub-vector quantifies the target user's facial emotional state, providing emotional data support at the facial dimension. The voice sub-vector represents the voice emotional state of the target user during a real-time, scenario-based experience of cultural and creative products. This voice sub-vector quantifies the emotional attributes in the target user's voice, supplementing non-visual emotional information and enhancing the comprehensiveness of recognition. The motion-sensation sub-vector represents the action-emotion state of the target user during a real-time, scenario-based experience of cultural and creative products. This motion-sensation sub-vector quantifies the emotions contained in the target user's limb movements, improving the dynamism of emotion recognition through dynamic motion features.
[0031] It should be noted that, in this application, the emotional feature vector is a multi-dimensional feature vector representing the facial emotional state, voice emotional state, and action emotional state of the target user during a real-time scenario-based experience of cultural and creative products. This emotional feature vector can comprehensively reflect the emotional state of the target user and integrate multi-modal advantages to provide accurate emotional basis for subsequent content response. In specific implementation, the facial sub-vector, the voice sub-vector, and the body sensation sub-vector are fused to generate the target user's emotional feature vector under the time slice. This can be achieved in the following way: a weighted fusion algorithm can be used to fuse the facial sub-vector, voice sub-vector, and body sensation sub-vector under the same time slice to generate the target user's emotional feature vector under the time slice; that is, the facial sub-vector, voice sub-vector, and body sensation sub-vector can be weighted and summed according to the weight allocation rule of "facial 0.4, voice 0.3, body sensation 0.3" to generate a 6-dimensional emotional feature vector containing pleasure, arousal, dominance, tone intensity, speech rate intensity, and excitement. In other embodiments, other weight allocation rules can also be used, which are not specifically limited here.
[0032] In some embodiments, constructing a dynamic evolution trajectory of the target user's emotional state changes based on all emotional feature vectors can be achieved through the following steps: Temporal calibration is performed on all sentiment feature vectors to obtain a temporal calibration set of sentiment features; Based on the emotional feature time-series calibration set, establish the emotional state coordinates of the target user at each time point during the experience of cultural and creative products; Trend fitting is performed on the emotional state coordinates at all time points to generate a dynamic evolution trajectory of the target user's emotional state changes.
[0033] In specific implementation, temporal calibration of all emotional feature vectors to obtain an emotional feature temporal calibration set can be achieved in the following way: First, based on the acquisition time of multimodal data in the frame buffer queue, all emotional feature vectors are sorted in ascending order according to the experience timeline of cultural and creative products. Then, combined with existing anomaly filtering technology (such as Gaussian filtering), all sorted emotional feature vectors are checked and filtered frame by frame to obtain the emotional feature temporal calibration set. In the anomaly filtering process, an anomaly threshold can be set for any emotional dimension value of an emotional feature vector fluctuating by more than 50% (such as a sudden drop in pleasure caused by blurred facial expression frames). The emotional feature temporal calibration set is a continuous and valid dataset formed after temporally sorting all emotional feature vectors and removing invalid vectors caused by sensor noise. This emotional feature temporal calibration set serves as a time benchmark for unified emotional data, which can eliminate data temporal disorder and noise interference, and ensure the temporal consistency and data reliability of subsequent emotional state analysis.
[0034] In specific implementation, establishing the emotional state coordinates of the target user at each time point during the experience of cultural and creative products based on the emotional feature time-series calibration set can be achieved in the following way: For the emotional feature vector of each time slice in the emotional feature time-series calibration set, the emotional feature vector can be decomposed into three dimensions: facial emotional intensity (i.e., the average of pleasure, arousal, and dominance) corresponding to the facial sub-vector micro-expression intensity, voice emotional intensity (i.e., the average of tone intensity and speech rate intensity) corresponding to the voice sub-vector intonation fluctuation and speech rate, and action emotional intensity (i.e., excitement) corresponding to the body sensory sub-vector motion amplitude. Then, a coordinate mapping algorithm is used to determine the coordinates of each dimension in the emotional feature vector. The emotional state coordinates of each time slice are calculated by assigning dimensional weights (e.g., 0.4 for face, 0.3 for voice, and 0.3 for body sensation) to the values of each dimension. Through the above steps, the emotional state coordinates of each time slice in the emotional feature time-series calibration set can be obtained. This yields the emotional state coordinates of the target user at each time point during the experience of cultural and creative products. The emotional state coordinates are three-dimensional quantitative indicators that display the emotional state of the target user at each time point during the cultural and creative experience. These emotional state coordinates can accurately locate the emotional state of the target user at each time point during the cultural and creative experience, providing concrete and calculable data support for subsequent trend fitting.
[0035] In practice, trend fitting of the emotional state coordinates at all time points to generate a dynamic evolution trajectory of the target user's emotional state changes can be achieved in the following way: First, the emotional state coordinates at all time points can be smoothly connected with a sliding window of 500 milliseconds. Then, the emotional state transition probability of the emotional state coordinates at all time points (such as the transition pattern from "low arousal-calm" to "high arousal-pleasure") can be captured by existing time-series state inference models (such as the transition pattern from "low arousal-calm" to "high arousal-pleasure") and emotional mutation nodes can be marked simultaneously (such as the time point where the fluctuation of any dimension in the three-dimensional index of emotional state coordinates exceeds 30% is the corresponding key node of cultural and creative content interaction). Thus, a dynamic evolution trajectory containing emotional change trends and mutation markings can be output.
[0036] It should be noted that in this application, the dynamic evolution trajectory is a curve showing the emotional changes of the target user throughout the entire cultural and creative experience. This dynamic evolution trajectory can clearly reflect the fluctuation trend of the target user's emotions as they interact with the cultural and creative content, providing accurate emotional basis for the subsequent generation of matching cultural and creative content interaction sequences.
[0037] In step 103, behavioral preference features of the target user's experience with cultural and creative products are extracted from the pre-constructed user digital profile, and then a sequence of cultural and creative content interaction that matches the target user's emotional state is generated based on the dynamic evolution trajectory and the behavioral preference features.
[0038] In some embodiments, extracting behavioral preference features of target users' experiences with cultural and creative products from a pre-built digital user profile can be achieved through the following steps: Call all user digital profiles from the pre-built cultural and creative behavior profile library; Based on all user digital profiles, target user experience and historical behavioral data of cultural and creative products are used to identify the target user experience. The core preferences of the target users are extracted from the historical behavioral data, thereby obtaining the behavioral preference characteristics of the target users' experience with cultural and creative products.
[0039] It should be noted that, in this application, the cultural and creative behavior profile library is a structured database that integrates frame buffer queue time-series indexing technology to store multimodal behavioral data during the user's cultural and creative experience. The multimodal behavioral data includes facial interaction frames (such as micro-expression dynamic images), voice interaction logs (including tone fluctuation features), and body-sensing operation trajectories (i.e., joint three-dimensional coordinate sequences). Each user digital profile entry is associated with a unique device number as an identifier. This cultural and creative behavior profile library can provide complete and low-latency raw data support for subsequent extraction of user preferences, avoiding the latency problem of cloud data retrieval. At the same time, the frame buffer queue time-series calibration ensures data continuity and meets the need for rapid location of user behavior records at the edge.
[0040] In specific implementation, the following method can be used to call all user digital profiles in the pre-built cultural and creative behavior profile library: the edge profile retrieval interface can be started, a user identifier set (such as device number) can be input, and the full set of user digital profiles can be quickly matched through the frame buffer queue time sequence index, and all user digital profiles can be output. This avoids the network latency of cloud retrieval, ensures the efficiency of subsequent target user data location, and conforms to the technical characteristics of low latency processing of edge computing. The user digital profile is a personalized data set that integrates user basic information and historical interaction records. This user digital profile depicts the user's unique behavior and emotional tendencies in the cultural and creative scenario, which can be used to accurately locate the user's core needs and provide a concrete user feature basis for subsequent extraction of behavioral preference features.
[0041] In specific implementation, the historical behavior data of the target user's experience with cultural and creative products based on all user digital profiles can be achieved in the following way: a hash algorithm can be used to quickly match the target user profile entries in all user digital profiles by the unique identifier of the target user; then, based on the data collection time recorded in the frame buffer queue, valid historical behavior data within the past 6 months is filtered out (such as excluding accidental touch records with a single interaction duration of <30s), and the historical behavior data of the target user's experience with cultural and creative products is output; wherein, the historical behavior data is the valid cultural and creative interaction records of the target user within the past 6 months selected from the cultural and creative behavior profile library, including interaction time, cultural and creative product type, operation actions (such as gesture swiping, voice commands) and emotional feedback data (such as emotional state vectors). This historical behavior data provides the original behavioral basis for preference extraction, which can ensure the accuracy and reliability of subsequent preference analysis.
[0042] In specific implementation, the core preferences of target users are extracted from the historical behavioral data to obtain the behavioral preference characteristics of target users' cultural and creative product experience. This can be achieved in the following way: First, three core dimensions can be set according to the behavioral preference quantification system derived from the pleasure-awakening-dominance emotional dimension model, including: cultural and creative type preference (such as digital exhibitions and interactive games), interaction modality preference (such as haptic, voice and touch screen), and dwell scene preference (such as exhibition hall and rest area). Then, for each core dimension, the historical records of various behaviors under the core dimensions in the historical behavioral data, including interaction frequency (i.e., number of clicks), dwell time (i.e., duration of each experience), and repeated operations (i.e., number of visits to the same type of cultural and creative product), can be weighted according to the formula "interaction frequency 0.4 + dwell time 0.3 + repeated operations 0.3" to calculate the preference score. The preference score of behavior under each core dimension can be obtained through the above steps. Finally, the behaviors corresponding to high preference scores are summarized as the core preferences of each core dimension through feature clustering algorithms (such as K-clustering), and the vector composed of all core preferences is used as the behavioral preference characteristics of target users' cultural and creative product experience.
[0043] It should be noted that in this application, the behavioral preference feature is a feature vector extracted from historical behavioral data to represent the core preference dimensions of the target user (such as preference for cultural and creative types, preference for interaction modalities, and preference for scene dwell). This behavioral preference feature clarifies the target user's demand tendency in the cultural and creative experience and can provide accurate preference basis for subsequent matching of emotionally responsive cultural and creative content, which is in line with the design logic of emotionally interactive cultural and creative products.
[0044] In some embodiments, reference Figure 3The figure is a flowchart illustrating the operation of determining the interaction sequence of cultural and creative content according to some embodiments of this application. In this application, generating the interaction sequence of cultural and creative content that matches the emotional state of the target user based on the dynamic evolution trajectory and the behavioral preference features can be achieved by the following steps: From the dynamic evolution trajectory, core preferences in the synchronously associated behavioral preference features of emotional mutation nodes are screened, thereby determining all emotional guidance anchors of the target user in the process of experiencing cultural and creative products; Map all emotionally driven anchors to the experiential content of cultural and creative products, and output a set of suitable candidate cultural and creative content. The candidate cultural and creative content set is sorted according to the emotional change trend of the dynamic evolution trajectory to generate a cultural and creative content interaction sequence that matches the emotional state of the target user.
[0045] It should be noted that in this application, the emotion-oriented anchor point is an emotional node with preference annotation during the target user's experience with cultural and creative products. This emotion-oriented anchor point can be used to establish a relationship between "emotional mutation" and "user preference", breaking the limitation of content matching that relies solely on emotional data, so that subsequent cultural and creative content recommendations not only fit the emotional state but also match the user's inherent preferences, thereby improving the accuracy of content adaptation.
[0046] In specific implementation, the process of selecting emotional mutation nodes from the dynamic evolution trajectory and synchronously associating them with core preferences in behavioral preference features to determine all emotional guidance anchors for the target user during the cultural and creative product experience can be achieved as follows: Extract all emotional mutation nodes from the dynamic evolution trajectory and synchronously call existing behavioral preference-emotion association algorithms (such as cosine similarity association algorithms) to associate and match emotional mutation nodes with core preferences in behavioral preference features (such as high-frequency interactive cultural relic tour modalities and long-stay time motion-sensing game scenarios). Output a vector containing emotional mutation timestamps and associated core preference labels as emotional guidance anchors during the cultural and creative product experience, thereby obtaining the emotional guidance anchors for the target user during the cultural and creative product experience. All emotionally oriented anchor points; among them, the behavior preference-emotion association algorithm can match the emotional mutation node with the core preferences in the behavior preference features by calculating the cosine similarity between the emotional state coordinates of the emotional mutation node and the emotional state vectors of each core preference in the historical behavior data (e.g., a cosine similarity greater than 0.8 is considered an association), and output the associated emotional mutation timestamp and associated core preference label as emotionally oriented anchor points in the cultural and creative product experience process; the emotional mutation node is the vector that locates the moment when the target user's emotions change suddenly due to the interaction with cultural and creative content. This emotional mutation node can be used to eliminate false emotional fluctuations caused by sensor noise, provide accurate time anchor points for subsequent matching and adaptation of cultural and creative content, and avoid invalid content push.
[0047] In practice, all emotionally driven anchor points are mapped to the experiential content of cultural and creative products. The output of a set of suitable candidate cultural and creative content can be achieved in the following way: Frame buffer queue temporal indexing technology can be used to search for experiential content in the system's built-in cultural and creative content library according to the core preference tags in all emotionally driven anchor points, based on the type of cultural and creative product (such as in-depth cultural relic tours, interactive puzzle-solving, immersive historical narratives, etc.). All the retrieved cultural and creative content is then combined into a set of candidate cultural and creative content. The set of candidate cultural and creative content is a set of suitable content filtered according to the preferences of the target user when experiencing cultural and creative products. This set of candidate cultural and creative content can provide the target user with multi-dimensional suitable content options, avoiding the limitations of single content. At the same time, the low latency characteristics of the frame buffer queue ensure the speed of content retrieval, meeting the real-time requirements of edge processing.
[0048] In a specific implementation, the candidate cultural and creative content set can be sorted according to the emotional change trend of the dynamic evolution trajectory to generate a cultural and creative content interaction sequence that matches the emotional state of the target user. This can be achieved in the following way: the experience content in the candidate cultural and creative content set can be sorted in chronological order according to the emotional change timestamp of the corresponding emotional guidance anchor point, and the sorted content sequence can be used as the cultural and creative content interaction sequence that matches the emotional state of the target user.
[0049] It should be noted that in this application, the cultural and creative content interaction sequence is a sequence of cultural and creative experience content generated according to the emotional change trend and behavioral preferences of the target user in the process of experiencing cultural and creative products. This cultural and creative content interaction sequence can make the push of cultural and creative content conform to the user's emotional flow pattern, ensure the continuity of user experience, avoid the experience fragmentation caused by the conflict between emotion and content, and fit the core design logic of emotional interactive cultural and creative products.
[0050] In step 104, the cultural and creative content interaction sequence is rendered in real time through a high-performance rendering node in the cloud and the rendering stream is pushed to the interactive terminal through an edge server, thereby enabling the cultural and creative product content experience response to interact emotionally with the target user in the interactive terminal.
[0051] In some embodiments, the real-time rendering of the cultural and creative content interaction sequence using a high-performance rendering node in the cloud and the push of the rendering stream to the interactive terminal via an edge server can be achieved through the following steps: The interaction sequence of the cultural and creative content is analyzed to obtain multiple rendering sub-tasks with a unified time base. The dynamic rendering optimization engine is started to render all rendering subtasks and the rendered result frames are passed to the frame buffer synchronization queue for timing calibration, resulting in a timing-coherent rendering stream. The rendering stream is pushed to the target user's interactive terminal via an edge server.
[0052] In specific implementation, the task parsing of the cultural and creative content interaction sequence to obtain multiple rendering sub-tasks with unified time bases can be achieved in the following way: the multimodal content parsing engine can be called to decompose the cultural and creative content interaction sequence to obtain the rendering sub-tasks corresponding to each module. For example, for the cultural relics module, the number of model vertices (e.g., 500,000 vertices) and texture resolution (e.g., 2K) can be extracted; for the motion sensing module, the computing power requirements of the physics engine (e.g., 1000 collision detections / second) and frame rate (e.g., 60 frames per second) can be extracted. At the same time, a unified format timestamp can be bound to each rendering sub-task through a time-series anchoring algorithm to ensure that the time base deviation of all sub-tasks is <1 millisecond, thereby outputting multiple rendering sub-tasks with unified time bases. The rendering sub-task is an independent rendering unit with timestamps and rendering parameters decomposed from the cultural and creative content interaction sequence. This rendering sub-task allows the dynamic rendering optimization engine to process according to the content type, thereby achieving load balancing of rendering tasks (i.e., high-complexity sub-tasks are allocated to high-performance nodes), and the unified time base avoids subsequent time sequence chaos, meets the low-latency rendering requirements of the edge, and lays the foundation for smooth emotional interaction.
[0053] In specific implementation, the dynamic rendering optimization engine is started to render all rendering sub-tasks and the resulting frames after rendering are passed to the frame buffer synchronization queue for timing calibration. The resulting time-coherent rendering stream can be achieved in the following way: The dynamic rendering optimization engine can be started to adapt the rendering technology according to the type of each rendering sub-task. For example, the cultural relics module can use ray tracing rendering to improve the realism of materials, and the motion sensing module can use a real-time physics engine to simulate limb collision feedback. The rendering time of a single frame is controlled to be less than 15 milliseconds by emotional frame priority scheduling (i.e., the frame priority associated with the emotional mutation node is set to the highest). The rendered result frames are then passed to the frame buffer synchronization queue and sorted in ascending order of timestamp. Out-of-order frames with timestamp deviation > 2 milliseconds and duplicate frames with content similarity > 95% are removed, and a time-coherent rendering stream is output. The rendering stream is a coherent set of frames that provide time-stable cultural and creative content data after being rendered by the dynamic rendering optimization engine. This rendering stream can ensure that the interactive terminal display is smooth and the frame order is coherent, avoiding the impact of content discontinuity on the user's emotional interaction experience, which meets the core requirements of real-time performance and smoothness for emotional interactive cultural and creative products in Shang Zhong'an's literature.
[0054] In specific implementation, the rendering stream can be pushed to the target user's interactive terminal via the edge server in the following way: the rendering stream can be pushed to the terminal through the edge-terminal short path transmission protocol combined with a dynamic transmission adjustment mechanism (such as reducing the bit rate by 10% when the network packet loss rate is >5%).
[0055] In some embodiments, the content experience response of cultural and creative products that engages in emotional interaction with the target user in an interactive terminal can be achieved through the following steps: After receiving the sequentially rendered stream, the edge server calls the intelligent encoding adaptation module to encode the data according to the type of interactive terminal. The edge-terminal timing phase-locked mechanism ensures that the timing deviation between the dynamically adjusted content displayed on the interactive terminal and the user's emotional feedback is less than a preset deviation threshold.
[0056] In specific implementation, after receiving the sequentially rendered stream, the edge server calls the intelligent encoding adaptation module to encode according to the interactive terminal type. This can be achieved in the following way: First, the target interactive terminal attributes (such as virtual reality glasses, interactive large screens, and touch tablets) can be obtained through the terminal type identification interface. Then, the encoding scheme is matched based on the terminal bandwidth and display capabilities. This includes: virtual reality glasses with low bandwidth requirements use low bitrate encoding (such as a bitrate of 2-4 megabits per second and a low encoding complexity level), interactive large screens with high image quality requirements use high bitrate encoding (such as a bitrate of 6-8 megabits per second and a high encoding complexity level), and touch tablets with balanced requirements use medium bitrate encoding (such as a bitrate of 4-6 megabits per second). During the encoding process, intra-frame compression optimization is enabled to ensure that the encoding latency is <10 milliseconds.
[0057] In specific implementation, the timing deviation between the dynamically adjusted content displayed on the interactive terminal and the user's emotional feedback is less than a preset deviation threshold can be ensured through an edge-terminal timing phase-locked loop mechanism. This can be achieved in the following way: First, the local clocks of the edge server and the interactive terminal are calibrated using a network time protocol to ensure that the time reference deviation between the two is less than 1 millisecond. Then, the terminal collects user emotional feedback data (such as facial micro-expressions and body movements) in real time and adds a 1-millisecond precision timestamp. The edge server synchronously obtains the timestamp of the rendering stream display and compares the two types of timestamps in real time using a timing deviation calculation formula (such as timing deviation = rendering stream display timestamp - emotional feedback timestamp) to ensure that the timing deviation is less than a preset deviation threshold (such as 3 milliseconds). If the timing deviation is greater than or equal to the preset deviation threshold, the edge server can dynamically correct it through a push rhythm adjustment algorithm (such as pushing the next frame of the rendering stream 5 milliseconds in advance).
[0058] Furthermore, in another aspect of this application, in some embodiments, this application provides a real-time contextualized experience system for cultural and creative products based on edge computing, see reference. Figure 4 The figure is a schematic diagram of the structure of a real-time contextualized experience system for cultural and creative products based on edge computing, according to some embodiments of this application. The real-time contextualized experience system 400 for cultural and creative products based on edge computing includes: a data acquisition module 401, a processing module 402, and an execution module 403, which are described below: The acquisition module 401 in this application is mainly used to synchronously acquire the facial expressions, voice signals and body movements of the target user when experiencing cultural and creative products in a real-time contextualized manner, so as to obtain the multimodal experience data stream of the target user. Processing module 402, in this application, is mainly used to perform multi-threaded parallel emotional state recognition on the multimodal experience data stream based on the frame buffer queue built on the edge server, to obtain all emotional feature vectors of the target user in the process of experiencing cultural and creative products, and then to construct the dynamic evolution trajectory of the target user's emotional state change based on all emotional feature vectors. It should be noted that the processing module 402 in this application is also used to extract the behavioral preference features of the target user’s cultural and creative product experience from the pre-constructed user digital profile, and then generate a cultural and creative content interaction sequence that matches the target user’s emotional state based on the dynamic evolution trajectory and the behavioral preference features. The execution module 403 in this application is mainly used to render the cultural and creative content interaction sequence in real time through a high-performance rendering node in the cloud and push the rendering stream to the interactive terminal through an edge server, thereby enabling the cultural and creative product content experience response to interact emotionally with the target user in the interactive terminal.
[0059] The modules in the aforementioned edge computing-based real-time scenario-based experience system for cultural and creative products can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the corresponding operations of each module.
[0060] In another embodiment, this application provides a computer device, which may be a server, and its internal structure diagram may be as follows. Figure 5 As shown, the computer device includes a processor, memory, and a network interface connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The database stores data related to a real-time contextualized experience method for cultural and creative products based on edge computing. The network interface communicates with external terminals via a network connection. When the computer program is executed by the processor, it implements a real-time contextualized experience method for cultural and creative products based on edge computing.
[0061] Those skilled in the art will understand that Figure 5 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0062] In one embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above embodiment of the method for real-time contextualized experience of cultural and creative products based on edge computing.
[0063] In one embodiment, a computer-readable storage medium is provided storing a computer program that, when executed by a processor, implements the steps described in the above embodiment of the method for real-time contextualized experience of cultural and creative products based on edge computing.
[0064] In one embodiment, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the steps described in the embodiment of the real-time contextualized experience method for cultural and creative products based on edge computing.
[0065] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the methods described above. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical storage, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0066] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0067] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.
Claims
1. A method for real-time contextualized experience of cultural and creative products based on edge computing, characterized in that, The method includes the following steps: Simultaneously collect facial expressions, voice signals, and body movements of target users during real-time scenario-based experiences with cultural and creative products, thereby obtaining multimodal experience data streams of target users; The frame buffer queue built on the edge server performs multi-threaded parallel emotional state recognition on the multimodal experience data stream to obtain all emotional feature vectors of the target user in the process of experiencing cultural and creative products, and then constructs a dynamic evolution trajectory of the target user's emotional state changes based on all emotional feature vectors. The behavioral preference features of target users' cultural and creative product experiences are extracted from the pre-constructed digital user profiles. Then, based on the dynamic evolution trajectory and the behavioral preference features, a sequence of cultural and creative content interactions that matches the target user's emotional state is generated. The interactive sequence of cultural and creative content is rendered in real time by a high-performance rendering node in the cloud and the rendering stream is pushed to the interactive terminal through an edge server, thereby enabling the cultural and creative product content experience response to interact emotionally with the target user in the interactive terminal.
2. The method as described in claim 1, characterized in that, Simultaneously collect facial expressions, voice signals, and body movements of target users during real-time, scenario-based experiences with cultural and creative products, thereby obtaining a multimodal experience data stream of the target users, specifically including: By synchronously capturing the facial expressions and body movements of target users during real-time scenario-based experiences of cultural and creative products using a visual sensor array, a sequence of facial expression video streams and body movement depth maps is obtained. The voice signal stream is obtained by collecting the voice signal of the target user during the real-time scenario-based experience of cultural and creative products through a high-fidelity microphone array. The facial expression video stream, the motion depth map sequence, and the speech signal stream are filtered and aligned to generate a multimodal experience data stream for the target user.
3. The method as described in claim 1, characterized in that, The frame buffer queue built on the edge server performs multi-threaded parallel emotional state recognition on the multimodal experience data stream to obtain all emotional feature vectors of the target user during the experience of cultural and creative products, specifically including: Create a frame buffer queue and a multimodal emotion recognition thread pool on the edge server side; The multimodal experience data stream is time-sliced and classified and cached into a frame buffer queue according to facial expression frames, voice feature frames, and body motion frames; For each time slice, the recognition thread in the multimodal emotion recognition thread pool performs emotion recognition on the facial expression frames, speech feature frames and body motion frames in the time slice of the frame buffer queue, and outputs facial sub-vectors, speech sub-vectors and body motion sub-vectors. The facial sub-vector, the voice sub-vector, and the tactile sub-vector are fused to generate the target user's emotional feature vector in each time slice, thereby obtaining the target user's emotional feature vector in all time slices during the experience of cultural and creative products.
4. The method as described in claim 1, characterized in that, Constructing a dynamic evolution trajectory of the target user's emotional state changes based on all emotional feature vectors specifically includes: Temporal calibration is performed on all sentiment feature vectors to obtain a temporal calibration set of sentiment features; Based on the emotional feature time-series calibration set, establish the emotional state coordinates of the target user at each time point during the experience of cultural and creative products; Trend fitting is performed on the emotional state coordinates at all time points to generate a dynamic evolution trajectory of the target user's emotional state changes.
5. The method as described in claim 1, characterized in that, Specifically, extracting behavioral preference characteristics of target users' experiences with cultural and creative products from pre-constructed digital user profiles includes: Call all user digital profiles from the pre-built cultural and creative behavior profile library; Based on all user digital profiles, target user experience and historical behavioral data of cultural and creative products are used to identify the target user experience. The core preferences of the target users are extracted from the historical behavioral data, thereby obtaining the behavioral preference characteristics of the target users' experience with cultural and creative products.
6. The method as described in claim 1, characterized in that, The generation of cultural and creative content interaction sequences that match the target user's emotional state based on the dynamic evolution trajectory and behavioral preference characteristics specifically includes: From the dynamic evolution trajectory, core preferences in the synchronously associated behavioral preference features of emotional mutation nodes are screened, thereby determining all emotional guidance anchors of the target user in the process of experiencing cultural and creative products; Map all emotionally driven anchors to the experiential content of cultural and creative products, and output a set of suitable candidate cultural and creative content. The candidate cultural and creative content set is sorted according to the emotional change trend of the dynamic evolution trajectory to generate a cultural and creative content interaction sequence that matches the emotional state of the target user.
7. The method as described in claim 1, characterized in that, Specifically, the real-time rendering of the cultural and creative content interaction sequence using high-performance rendering nodes in the cloud and the push of the rendering stream to the interactive terminal via an edge server include: The interaction sequence of the cultural and creative content is analyzed to obtain multiple rendering sub-tasks with a unified time base. The dynamic rendering optimization engine is started to render all rendering subtasks and the rendered result frames are passed to the frame buffer synchronization queue for timing calibration, resulting in a timing-coherent rendering stream. The rendering stream is pushed to the target user's interactive terminal via an edge server.
8. The method as described in claim 1, characterized in that, The content experience response of cultural and creative products that engages in emotional interaction with target users in interactive terminals specifically includes: After receiving the sequentially rendered stream, the edge server calls the intelligent encoding adaptation module to encode the data according to the type of interactive terminal. The edge-terminal timing phase-locked mechanism ensures that the timing deviation between the dynamically adjusted content displayed on the interactive terminal and the user's emotional feedback is less than a preset deviation threshold.
9. The method as described in claim 3, characterized in that, The multimodal emotion recognition thread pool includes a facial expression recognition thread, a voice emotion recognition thread, and a body motion recognition thread.
10. A real-time scenario-based experience system for cultural and creative products based on edge computing, characterized in that, include: The data acquisition module is used to simultaneously collect facial expressions, voice signals and body movements of target users when they experience cultural and creative products in a real-time contextualized manner, thereby obtaining a multimodal experience data stream of the target users; The processing module is used to perform multi-threaded parallel emotional state recognition on the multimodal experience data stream based on the frame buffer queue built on the edge server, to obtain all emotional feature vectors of the target user in the process of experiencing cultural and creative products, and then to construct a dynamic evolution trajectory of the target user's emotional state changes based on all emotional feature vectors. The processing module is used to extract behavioral preference features of target users’ cultural and creative product experiences from the pre-constructed user digital profile, and then generate a cultural and creative content interaction sequence that matches the target user’s emotional state based on the dynamic evolution trajectory and the behavioral preference features. The execution module is used to render the cultural and creative content interaction sequence in real time through a high-performance rendering node in the cloud and push the rendering stream to the interactive terminal through an edge server, thereby enabling the cultural and creative product content experience response to interact emotionally with the target user in the interactive terminal.