Visual immersive interactive display method and device for multi-modal data of text and blog scene

By collecting multimodal data and building user profiles, the problem of user differences not being considered in traditional cultural heritage scenarios has been solved, and personalized content recommendations and immersive experiences have been improved.

CN121478124APending Publication Date: 2026-02-06SHENZHEN COMMSCOPE INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511641961.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-11
Publication Date
2026-02-06

AI Technical Summary

Technical Problem

The digital interactive display technology in traditional cultural heritage sites has failed to fully consider the differences among users, making it difficult for novice users to understand professional content, while experienced users cannot obtain in-depth information, resulting in a lack of personalization and immersion.

Method used

By using multimodal data collection and authenticity verification mechanisms, a multi-dimensional user profile is constructed. Combining static attributes and dynamic behavior sequences, a personalized interactive content recommendation strategy is generated, including an interest index and a duration preference model, to achieve adaptive matching of content, format, and time.

Benefits of technology

It significantly improves the depth and accuracy of user modeling, optimizes information transmission efficiency, and enhances the personalization of interactive displays and the user's immersive experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121478124A_ABST
    Figure CN121478124A_ABST
Patent Text Reader

Abstract

The invention discloses a visual immersive interactive display method and device for multi-modal data of a cultural relic scene, and relates to the technical field of digital display. The method comprises the following steps: acquiring basic attribute information of a user, collecting a figure image for authenticity verification, and determining the user passing the verification as a target user; collecting the staying duration, question data, active operation behavior data and skipped content classification data of the target user in the interaction process, and extracting a real-time interestingness index; constructing a label based on the basic attribute and the behavior information, and generating a multi-dimensional user portrait; determining recommendation content, information density and presentation form of the personalized interaction content according to the multi-dimensional portrait, and generating recommendation duration in combination with a user duration preference model and a real-time interestingness index; and fusing multi-modal data to realize personalized interactive display through an immersive interactive terminal. According to the method and the device, the content, form and duration all-directional self-adaptive literature and blog interaction experience is realized, and the immersion of the user is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of digital display technology, specifically to a method and apparatus for immersive interactive display of multimodal data visualization in cultural heritage settings. Background Technology

[0002] Traditional digital interactive display technologies in cultural heritage sites typically provide uniform content to all visitors based on pre-set display formats. These technologies deliver information through fixed-duration video playback, standardized text and image descriptions, or virtual tours with pre-defined paths. They fail to adequately consider the differences in cognitive levels, interests, and information processing abilities among different visitors. This results in novice users struggling to understand specialized content, while experienced users are unable to access in-depth information, thus limiting the effectiveness of interaction and the immersive experience of the visit.

[0003] Existing technologies have introduced some solutions to filter content by introducing user tags, but they mostly focus on static attributes or single behavioral dimensions, failing to integrate real-time behavioral sequences and visual feedback data. In particular, they lack dynamic and fine-grained control over the duration of content display, resulting in insufficient personalization, weak adaptability, and limited immersive experience. Summary of the Invention

[0004] This invention addresses the technical problems of insufficient personalization of content display, single-dimensional user profile construction, and lack of dynamic control over content display duration in existing cultural heritage scenarios by providing a multimodal data visualization immersive interactive display method and device for cultural heritage scenarios.

[0005] The technical solution of the present invention to solve the above-mentioned technical problems is as follows: In a first aspect, the present invention provides an immersive interactive display method for multimodal data visualization in cultural heritage scenarios, including: After a user enters the immersive interactive system of the cultural heritage scene, the system obtains the user's basic attribute information. Collect the user's image information, verify the authenticity of the basic attribute information, and define the user who passes the authenticity verification as the target user; Real-time collection of the target user's behavioral information during the immersive interaction process, including dwell time, question data, active operation behavior data, and skipped content category data, and extraction of the target user's real-time interest index from the person image information; Based on the target user's basic attribute information and behavioral information, static and dynamic tags are constructed and fused to generate a multi-dimensional profile of the target user. Based on the multi-dimensional profile, the recommended content, information density, and presentation format of personalized interactive content are determined, and the recommended duration of personalized interactive content is generated based on the user time preference model and the real-time interest index. After integrating the corresponding multimodal data, interactive display is achieved through an immersive interactive terminal.

[0006] Secondly, the present invention provides an immersive interactive display device for multimodal data visualization in cultural heritage scenarios, comprising: The information acquisition module is used to acquire the user's basic attribute information after the user enters the immersive interactive system of the cultural heritage scene; The verification module is used to collect the user's image information, verify the authenticity of the basic attribute information, and define users who pass the authenticity verification as target users. The behavior collection and analysis module is used to collect the target user's behavior information in real time during the immersive interaction process. The behavior information includes dwell time, question data, active operation behavior data and skipped content classification data, and extracts the target user's real-time interest index from the person image information. The user profile building module is used to build static and dynamic tags based on the target user's basic attribute information and behavioral information, and then merge them to generate a multi-dimensional profile of the target user. The recommendation strategy generation module is used to determine the recommended content, information density and presentation format of personalized interactive content based on the multi-dimensional profile, and to generate the recommended duration of personalized interactive content based on the user duration preference model and the real-time interest index. The interactive display module is used to integrate the corresponding multimodal data and then achieve interactive display through an immersive interactive terminal.

[0007] The beneficial effects of this invention are: Compared to existing technologies, this invention first ensures the authenticity and comprehensiveness of the data relied upon for user profile construction through a multimodal data collection and authenticity verification mechanism, laying a reliable data foundation for subsequent personalized recommendations. Secondly, by integrating static attributes and dynamic behavioral sequences to construct multi-dimensional user profiles, it can more accurately depict users' cognitive levels, interest preferences, and content rejection tendencies, significantly improving the depth and accuracy of user modeling. Thirdly, based on multi-dimensional profiles and real-time interest indices, it dynamically generates personalized interactive content recommendation durations, achieving comprehensive adaptive matching of content display strategies across content, form, and time dimensions, effectively optimizing information delivery efficiency. Finally, through a closed-loop process from data collection and profile construction to strategy generation and execution, it significantly enhances the personalization of interactive displays and the user's immersive experience in cultural heritage scenarios. Attached Figure Description

[0008] Figure 1 A flowchart illustrating the immersive interactive display method for multimodal data visualization in cultural heritage scenarios provided by this invention; Figure 2This is a schematic diagram of the structure of the immersive interactive display device for multimodal data visualization in cultural heritage scenarios provided by the present invention.

[0009] In the attached diagram, the components represented by each number are as follows: Information acquisition module 11, verification module 12, behavior collection and analysis module 13, user profile construction module 14, recommendation strategy generation module 15, and interactive display module 16. Detailed Implementation

[0010] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0011] In the description of this invention, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the stated features. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.

[0012] In the description of this invention, the term "for example" is used to mean "used as an example, illustration, or description." Any embodiment described as "for example" in this invention is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is provided to enable any person skilled in the art to make and use the invention. Details are set forth in the following description for purposes of explanation. It should be understood that those skilled in the art will recognize that the invention can be made without using these specific details. In other instances, well-known structures and processes will not be described in detail to avoid obscuring the description of the invention with unnecessary detail. Therefore, the invention is not intended to be limited to the embodiments shown, but is consistent with the broadest scope of the principles and features disclosed herein.

[0013] Example 1, as Figure 1 As shown, this embodiment of the invention provides an immersive interactive display method for multimodal data visualization in cultural heritage scenarios, including: S10: After a user enters the immersive interactive system of the cultural heritage scene, the system obtains the user's basic attribute information; First, once a user enters the immersive interactive system within the cultural heritage scene, the system obtains the user's basic attribute information.

[0014] Specifically, the basic attribute information includes at least: the user's self-declared identity attributes, the level of knowledge in the cultural heritage field determined through cultural heritage cognition association, the user's actively selected interactive language preferences, and the user's declared information receiving assistance needs.

[0015] Specifically, this basic attribute information includes user characteristic data across multiple dimensions. Among these, the user-declared identity attributes include basic identity characteristics such as age range, occupation category, and educational background, obtained through direct input from the user and the interactive interface. The knowledge level in the cultural heritage field is obtained through a correlation assessment based on cultural heritage cognition. This assessment process automatically classifies the user's knowledge level by presenting standardized test questions in the cultural heritage field and determining the degree of matching between the user's answers and a pre-set knowledge graph. For example, the knowledge level can be divided into three levels: beginner, intermediate, and advanced.

[0016] Interactive language preferences are obtained through user active selection. Immersive interactive systems offer multiple language options, and users choose the interactive language according to their own needs. This interactive language preference will directly affect the language presentation of all subsequent interactive content.

[0017] Information reception assistance needs are obtained through user submissions. These needs specifically refer to special auxiliary functions required by users during information reception, including visual assistance (such as font magnification) and auditory assistance (such as sign language interpretation, synchronized subtitles, and volume amplification). Users can choose to enable or disable specific auxiliary functions according to their own needs.

[0018] All of the above basic attribute information is stored in the local immersive interactive system via encrypted transmission and serves as the basic data source for building user profiles in subsequent processing. This information acquisition process ensures the integrity and accuracy of the initial data required for subsequent personalized recommendations.

[0019] S20: Collect user image information, verify the authenticity of basic attribute information, and define users who pass the authenticity verification as target users; After obtaining the user's basic attribute information, it is further necessary to collect the user's image information. Since unintentional misreporting or intentional falsification may occur during the user's self-declaration of basic attribute information, this directly affects the accuracy of subsequent user profile construction and the effectiveness of personalized recommendations. Therefore, by collecting image information, biometric data strongly correlated with the user's identity can be obtained. Comparing this biometric data with pre-stored biometric templates corresponding to the user's declared identity attributes can objectively verify the authenticity of the declared information. This verification step effectively ensures the reliability of the basic attribute information and prevents user profile deviations due to inaccurate information.

[0020] First, the system collects the user's image information in real time using image acquisition devices deployed within the immersive interactive environment of the cultural heritage site. These devices automatically activate when the user enters the interactive area, capturing high-resolution image data containing the user's facial features.

[0021] Then, the authenticity of the basic attribute information is verified based on the image information of the person.

[0022] Specifically, the authenticity of basic attribute information is verified, including: The user's image information is matched with the biometric information corresponding to the identity attribute in the basic attribute information, and the matching degree is calculated. If the matching degree is greater than or equal to the preset matching degree threshold, the authenticity verification is deemed to have passed; if the matching degree is less than the preset matching degree threshold, the authenticity verification is deemed to have failed, and a secondary verification process is triggered.

[0023] Specifically, the authenticity verification process is achieved by collecting the user's image information and cross-referencing it with the declared basic attribute information.

[0024] First, biometric information corresponding to the identity attributes is extracted from the basic attribute information declared by the user. Specifically, this biometric information is retrieved from a pre-established biometric database and is a pre-stored biometric information template uniquely bound to that identity attribute. The biometric database is an encrypted database constructed by collecting facial image data from registered users and standardizing it using feature extraction algorithms. It is obtained through the biometric collection and encoding process completed during the initial user registration phase, and each identity identifier corresponds to a unique biometric information template in the database. Second, the real-time collected human image information is processed through feature extraction algorithms to generate the current biometric information.

[0025] Then, the real-time extracted biometric information is compared with the biometric information corresponding to the identity attributes to calculate the matching degree. This calculation process outputs a quantified matching degree value, which is used to objectively evaluate the consistency between the two types of feature data. For example, the cosine similarity calculation method is used, which calculates the cosine of the angle between two feature vectors in the vector space to obtain a matching degree value between 0 and 1. This matching degree value objectively reflects the spatial distribution consistency of the two biometric feature vectors; the higher the value, the stronger the feature consistency.

[0026] The calculated matching score is compared with a preset matching score threshold. This threshold is a decision threshold pre-set based on the security level requirements of the actual application scenario. Its value is determined by optimizing the false acceptance rate and false rejection rate on a large-scale test dataset. For example, in a cultural heritage scenario requiring high security, this threshold can be set to 0.85. When the calculated matching score is greater than or equal to the preset matching score threshold, the user's identity verification is considered successful. At this point, the user is officially defined as the target user and authorized to use subsequent personalized interaction functions.

[0027] When the matching score is less than the preset matching score threshold, the authenticity verification fails. In this case, a secondary verification process will be automatically triggered. This process may include supplementary verification methods such as re-collecting image information, manual verification, or using alternative verification methods. For example, the user may be prompted to read the chip information in their ID card, and the obtained identity data will be cross-compared with the basic attribute information declared by the user.

[0028] S30: Real-time collection of target user behavior information during immersive interaction, including dwell time, question data, active operation behavior data, and skipped content category data, and extraction of the target user's real-time interest index from the person image information; After verifying the user's identity and identifying the user as the target user, continuous multi-dimensional behavioral data monitoring and recording of the target user during the immersive interaction process is initiated.

[0029] Specifically, it collects real-time behavioral information of target users during immersive interaction, including dwell time, question data, active operation data, and skipping content category data, and extracts the target user's real-time interest index from the user's image information, including: Record the entry and exit timestamps of target users on each immersive interactive page, and calculate the dwell time; Collect question data input by target users through the interactive system, including voice or text question data; Record the active operation data of the target users, including at least the click to view specific cultural relics content, the selection and switching of content modules, and the interactive commands to the cultural relics visualization model; Identify skip actions triggered by the target user, categorize the skip actions by content type, and use this as skip content classification data; By integrating dwell time, question data, active behavior data, and skipped content category data, behavioral information of the target user is formed. Using computer vision algorithms, facial expressions and body postures of target users are identified from human image information and quantified into a real-time interest index according to preset scoring rules.

[0030] Specifically, the collection of behavioral information covers four main aspects. First, by recording the entry and exit timestamps of users in various interactive pages or display units, the dwell time is accurately calculated. Second, user-inputted questions via voice or text are collected, preserving their original semantic content. Third, user-initiated actions are recorded, including clicking on specific cultural relic images and text, switching between different content modules, and rotating or scaling visualized cultural relic models. Fourth, explicitly triggered skip actions by users are identified and categorized, establishing skip content classification data based on the knowledge domain or media type of the skipped content. These four types of data, after integration and processing, together constitute behavioral information reflecting user interaction preferences.

[0031] Simultaneously, image acquisition devices are deployed to continuously acquire user images, and computer vision algorithms are used to analyze these images in real time. Specifically, this computer vision algorithm is based on deep learning-based object detection and pose estimation methods. First, a face detection algorithm based on convolutional neural networks is used to locate the face region in the user image information, while a human pose estimation algorithm extracts the body contour and key skeletal points. Then, facial expression features are identified based on facial landmark detection technology, including geometric features such as the curvature of the corners of the mouth, the shape of the eyebrows, and the state of the eyes; limb pose features are calculated by analyzing the three-dimensional coordinates of human landmarks, including kinematic parameters such as the head's turning angle and the forward tilt angle of the body's principal axis relative to the vertical direction.

[0032] Secondly, based on preset quantitative scoring rules, the aforementioned visual features are converted into a numerical real-time interest index. This real-time interest index can dynamically reflect the user's level of focus and interest during the interaction process. The higher the value, the more active the user's facial expressions and the closer their body posture is to the interactive interface, indicating a stronger cognitive engagement and content interest intensity; the lower the value, the more indifferent the user's facial expressions or avoidance posture, indicating a state of distraction or lack of interest.

[0033] For example, the quantification process from visual features to a real-time interest index is implemented through a pre-trained regression model. This regression model takes standardized feature vectors extracted from facial expressions and body postures as input and outputs a real-time interest index representing the user's current level of interest. This real-time interest index is then processed by a linear scaling layer and finally mapped to a pre-defined target numerical range, such as 0.8 to 1.2, which facilitates subsequent calculations. This regression model is trained on a large amount of labeled user interaction video data, ensuring the accuracy and consistency of the evaluation results.

[0034] In summary, by simultaneously collecting behavioral information and real-time interest indices, we can obtain dual feedback from users' explicit operational behaviors and implicit emotional responses, providing comprehensive and multi-dimensional data support for the subsequent construction of dynamic user profiles.

[0035] S40: Construct static and dynamic tags based on the target user's basic attribute and behavioral information, and merge them to generate a multi-dimensional profile of the target user; Specifically, static and dynamic tags are constructed based on the target user's basic attribute and behavioral information, and then integrated to generate a multi-dimensional profile of the target user, including: Based on the mapping of the target user's basic attribute information, static tags are generated, including identity attribute tags, knowledge level tags, interaction language preference tags, and information reception assistance requirement tags; Dynamic tags are constructed based on the behavioral information of target users. These dynamic tags include interest tags, focus tags, preferred content tags, and content rejection tags. According to the preset fusion rules, static tags and dynamic tags are integrated to form a multi-dimensional profile that includes user attribute characteristics, cognitive level characteristics, interest preference characteristics, and content rejection characteristics.

[0036] First, static tags are generated based on the target user's basic attribute information. This process maps the basic attribute information into a standardized tag system: identity attribute tags record the user's demographic characteristics such as age and occupation; knowledge level tags identify the user's cognitive level in the cultural heritage field; interaction language preference tags record the user's selected interaction language category; and information reception assistance needs tags mark the special assistance functions required by the user.

[0037] Secondly, dynamic tags are constructed based on the behavioral information of the target users.

[0038] Specifically, dynamic tags are constructed based on the target user's behavioral information, including: Content with a dwell time greater than or equal to a preset dwell time threshold is marked as content of interest. Content of interest is then clustered into similar categories to form interest preference tags. The system uses a pre-trained natural language processing model in the cultural heritage field to analyze the question data, extract the demand keywords and count their frequency, select high-frequency demand keywords that meet the preset question frequency threshold, and perform semantic clustering to form focus direction tags. Active operation content is extracted from active operation behavior data and its frequency is counted. High-frequency active operation content that meets the preset operation frequency threshold is selected, and similar clusters are formed to create preferred content tags. Skipped content category data is counted by content type, and content types with a skip count greater than or equal to a preset skip count threshold are used as content rejection tendency tags. Integrate interest-oriented tags, focus-oriented tags, preferred content tags, and content aversion tags to form dynamic tags.

[0039] First, interest-based tags are constructed by analyzing dwell time data. The dwell time of users in each content unit is compared with a preset dwell time threshold, and content that reaches or exceeds the threshold is marked as content of interest. Then, based on the content classification standards of the cultural heritage knowledge system, content of interest is clustered into similar groups (e.g., K-means clustering, DBSCAN clustering), forming interest-based tags reflecting areas of sustained user attention. The dwell time threshold is set comprehensively based on the experience of experts in the cultural heritage field and the results of large-scale user behavior data analysis. Its value must simultaneously consider the inherent information content of the content unit and the average reading speed of users. For example, for exhibits with medium information density in the ceramics category, the threshold can be set to 30 seconds; while for detailed appreciation units with high information density in the calligraphy and painting category, the threshold is correspondingly increased to 90 seconds.

[0040] Secondly, focus tags are constructed by analyzing question data. A natural language processing model trained on a corpus in the cultural heritage field is used to semantically analyze users' voice or text questions, extracting core demand keywords. By statistically analyzing the frequency of keyword occurrences, high-frequency demand keywords that reach a preset question frequency threshold are selected. Finally, these keywords are clustered using semantic similarity calculations to form focus tags representing users' deeper knowledge needs. The preset question frequency threshold is set based on the total number of user questions and keyword distribution characteristics in a single interactive session, aiming to filter out statistically significant core needs. For example, in a standard interactive session lasting 30 minutes, if a keyword appears at least 3 times in all user questions, it is determined that the keyword has reached the preset question frequency threshold.

[0041] Secondly, preference content tags are constructed by analyzing proactive user behavior. The content of the user's actions (clicks, switches, model operations, etc.) is extracted from these records, and the frequency of each action is statistically analyzed. High-frequency actions that reach a preset frequency threshold are selected and clustered according to content type (e.g., K-means clustering, DBSCAN clustering), forming preference content tags that reflect user interaction preferences. The preset frequency threshold is set based on a combination of the percentage of total proactive actions during a user session and the absolute frequency, aiming to identify preference behaviors with significant statistical bias. For example, if a user performs 50 proactive actions in a single session, and the interaction with a specific 3D model of an artifact reaches 10% of the total actions (5 times), then that content is considered to have reached the preset frequency threshold.

[0042] Finally, content rejection bias tags are constructed by analyzing skip operation data. The content skipped by users is categorized and statistically analyzed by type, and the number of skips for each category is calculated. Content types with skip counts reaching or exceeding a preset skip count threshold are directly labeled as content rejection bias tags to identify content areas that users are not interested in. The preset skip count threshold is set based on the total number of content types a user encounters during the current session and statistical significance requirements, aiming to exclude the influence of accidental skipping behavior. For example, among 10 knowledge categories encountered by a user, if the number of skips for a certain category reaches 20% of the total number of categories encountered (i.e., 2 times), then that content type is considered to have reached the preset skip count threshold.

[0043] Ultimately, by integrating tag data from four dimensions—interest tendency tags, focus direction tags, preferred content tags, and content rejection tendency tags—a dynamic tag set that comprehensively reflects users' interests, preferences, and behavioral characteristics is formed, providing a precise basis for subsequent personalized recommendations.

[0044] For example, a fusion rule based on feature vector concatenation and attention weight allocation can be adopted. Specifically, firstly, static and dynamic labels are encoded into feature vectors of fixed dimensions; then, the weight distribution of the dynamic label vector relative to the static label vector is calculated through an attention mechanism, where the static label is used as the basic feature vector and the dynamic label is used as the context feature vector; finally, the weighted dynamic feature vector is concatenated with the static basic vector and fused through a fully connected layer for dimensionality reduction.

[0045] The resulting multi-dimensional profile includes user attribute features, cognitive level features, interest preference features, and content rejection features. This multi-dimensional profile can structurally represent the user's overall state and real-time tendencies, providing a precise data foundation for subsequent personalized interactions.

[0046] S50: Based on multi-dimensional profiles, determine the recommended content, information density, and presentation format of personalized interactive content, and generate the recommended duration of personalized interactive content based on user time preference model and real-time interest index; Specifically, based on multi-dimensional profiles, the recommended content, information density, and presentation format of personalized interactive content are determined, including: Based on the interest and preference features and content rejection features in the multi-dimensional profile, personalized interactive content recommendations are generated from the preset content library; Information density of personalized interactive content generated based on cognitive level features in multi-dimensional profiles; Personalized interactive content is generated based on user attribute features from multi-dimensional profiles.

[0047] First, content items matching the user's interests are selected from a pre-set content library based on their interests and preferences. Content exclusion features are then used to filter the initial selection, removing content that falls into categories the user dislikes. This results in a final set of recommended content that aligns with the user's interests while avoiding their excluded areas. The pre-set content library is a collection of digital resources built through systematic collection and structured processing of multimodal data in the cultural heritage field. Obtained through authorization from professional institutions, integration of academic resources, and digital acquisition and processing, it includes various types of content materials such as high-resolution images of cultural relics, 3D models, historical documents, audio lectures by experts, research paper abstracts, and educational videos, all indexed and organized according to the cultural heritage professional classification system.

[0048] Secondly, information density is determined. Information density refers to the amount of knowledge and information complexity contained in a unit of interactive content, and its determination is related to cognitive level characteristics. Specifically, when the cognitive level characteristics identify the user as a basic cognitive level, a low information density mode is automatically selected, corresponding to the use of basic concept introductions and concise knowledge points; when an intermediate cognitive level is identified, a medium information density mode is adopted, incorporating an appropriate amount of background knowledge and related explanations; if identified as an advanced cognitive level, a high information density mode is activated, providing in-depth academic analysis, professional terminology, and multi-faceted perspectives.

[0049] Next, the presentation format is determined, and its configuration is driven by user attribute characteristics. Based on the interaction language preferences recorded in the attribute characteristics, content materials in the corresponding language are called up; according to the information reception assistance needs, the corresponding auxiliary function modules are automatically activated, such as generating synchronized subtitles, sign language VR, enlarged text and images, and voice broadcasts for users with hearing or visual impairments; at the same time, combined with the age information in the identity attributes, the interaction design and visual style are adapted to different age groups, such as presenting in the form of animations and popular science videos based on the user's age.

[0050] Furthermore, the recommended duration of personalized interactive content is generated based on a user time preference model and a real-time interest index, including: Obtain a pre-trained user time preference model; Input the average dwell time associated with the target user's interest tags, and output the predicted recommendation time of personalized interactive content; The predicted recommendation duration is adjusted according to the scenario to obtain the recommendation duration of personalized interactive content.

[0051] First, a pre-trained user time-consuming preference model is invoked. This model is a prediction model trained using a regression algorithm based on massive amounts of historical user interaction data in cultural heritage scenarios. It can establish a mapping relationship between user interest characteristics and ideal content consumption time.

[0052] The average dwell time associated with the target user's interest tags is used as a key input feature and fed into the user time preference model. By analyzing the user's typical dwell patterns across different interest categories and combining this with the general behavioral patterns of their user group, the model outputs a predicted recommended dwell time for the current personalized interactive content. This predicted value reflects a baseline dwell time suggestion that adapts to the user's interest preferences in typical scenarios.

[0053] Furthermore, the predicted recommendation duration is adjusted to adapt to different scenarios. Since the predicted recommendation duration is a static prediction value generated based on users' historical behavior data, it reflects the typical duration requirements under users' long-term interests and preferences, and cannot respond to dynamic changes in the current interaction context in real time. Therefore, it is necessary to adjust the predicted recommendation duration adaptively by introducing a real-time interest index as a dynamic adjustment factor to capture users' instantaneous focus and emotional feedback, ultimately obtaining an accurate recommendation duration that reflects users' long-term preferences and adapts to real-time interaction states.

[0054] Specifically, the predicted recommendation duration is adjusted for specific scenarios to obtain the recommendation duration for personalized interactive content, including: The average effective reception time of each personalized interactive content type in the cultural heritage scenario is obtained as the baseline duration for the scenario. The average effective reception time of user groups whose multi-dimensional profiles of the target users meet the preset similarity threshold for each personalized interactive content type is obtained as the group reference time. The weighted average of the predicted recommendation duration, the baseline duration of the scenario, and the reference duration of the group is calculated, and then multiplied by the real-time interest index to obtain the recommendation duration of personalized interactive content.

[0055] First, the baseline duration of the scene is obtained. This baseline duration is obtained by statistically analyzing the average effective reception time of a large number of users for various content types in cultural heritage scenarios. It represents an ideal duration benchmark that is generally applicable to specific content types. For example, for content such as the display of 3D models of ceramics, the baseline duration of the scene may be statistically determined to be 90 seconds.

[0056] Simultaneously, a group reference duration is obtained. Based on the multi-dimensional profile of the target user, similar user groups with profile similarity reaching a preset similarity threshold are retrieved from the historical user database. The average effective reception time of this similar user group for each content type is calculated. The final determined group reference duration reflects the interaction habits of groups with similar characteristics to the target user on similar content, providing a group behavior reference for personalized recommendations. The historical user database is constructed by continuously collecting and storing user behavior data generated in the immersive interactive system of cultural heritage scenarios. This database records users' basic attributes, complete interaction behavior sequences, and corresponding content metadata. The preset similarity threshold is set according to the balance between recommendation accuracy and computational efficiency. Its specific value is optimized after analyzing the distribution characteristics of multi-dimensional profile feature vectors. For example, in scenarios where cosine similarity is used to calculate profile feature vectors, this threshold is usually set to 0.75. When the similarity calculation result of the profile feature vectors of the target user and historical users is greater than or equal to 0.75, the historical user is included in the similar user group for calculating the group reference duration.

[0057] Finally, the predicted recommendation duration, the scenario baseline duration, and the group reference duration are weighted and averaged according to preset weights. The allocation of each weight is based on prior knowledge of its importance to the final recommendation duration or determined through experimental optimization; for example, the weight of the predicted recommendation duration can be set to 0.5, the scenario baseline duration to 0.3, and the group reference duration to 0.2. This weighted calculation process effectively integrates information from individual preferences, scenario characteristics, and group behavior. Then, the weighted average is multiplied by the real-time interest index, which acts as an adjustment factor; a value greater than 1 indicates an upward adjustment of the duration, while a value less than 1 indicates a downward reduction. Through this calculation process, the final output recommendation duration retains the stability of the baseline prediction while possessing dynamic flexibility to adapt to real-time interaction states.

[0058] S60: After integrating the corresponding multimodal data, interactive display is achieved through an immersive interactive terminal.

[0059] Specifically, the system retrieves multimodal data resources corresponding to the recommended content from a pre-defined content library, including high-resolution images of cultural relics, 3D models, audio explanations, documentary materials, and video content. Based on the determined information density, the text and audio content undergo intelligent summarization or detailed expansion processing to ensure that the amount of information accurately matches the user's cognitive level. Simultaneously, based on the generated presentation parameters, the visual style of the interactive interface, the language and speed of the voice broadcast, and the font size are configured in real time.

[0060] Finally, the integrated and processed multimodal data stream is sent to the immersive interactive terminal. This terminal automatically selects the optimal presentation medium based on the content characteristics; for example, it projects 3D artifact models onto augmented reality glasses for stereoscopic display, displays historical background information via a high-definition touchscreen with appropriate fonts and layouts, and plays narration in the corresponding language through spatial audio equipment. The entire interactive display process strictly adheres to the generated recommended duration, initiating a smooth content switching mechanism at the end of the preset time, thus constructing a highly personalized immersive cultural heritage interactive experience in terms of content, form, and duration.

[0061] In summary, the embodiments of this application have at least the following technical effects: Compared to existing technologies, this invention effectively improves the data quality and reliability of user profile construction by introducing multimodal data fusion and authenticity verification mechanisms. The multi-dimensional profile construction method based on static attributes and dynamic behavioral sequences can comprehensively depict users' cognitive characteristics and interest preferences, significantly enhancing the accuracy of personalized recommendations. By combining a user time preference model with a real-time interest index, dynamic and refined control of content display time is achieved, solving the problem of poor user experience caused by fixed display time in traditional methods. The resulting complete technical solution from data collection to personalized display achieves significant improvements in content adaptability, interactive immersion, and information transmission efficiency, providing an innovative solution for digital interaction in cultural heritage scenarios.

[0062] Example 2, as Figure 2 As shown, based on the same inventive concept as the immersive interactive display method for multimodal data visualization in cultural heritage scenes provided in Embodiment 1, this embodiment of the invention also provides an immersive interactive display device for multimodal data visualization in cultural heritage scenes, including: The information acquisition module 11 is used to acquire the user's basic attribute information after the user enters the immersive interactive system of the cultural heritage scene; The verification module 12 is used to collect the user's image information, verify the authenticity of the basic attribute information, and define the user who passes the authenticity verification as the target user. The behavior collection and analysis module 13 is used to collect the behavior information of the target user in the immersive interaction process in real time. The behavior information includes dwell time, question data, active operation behavior data and skipped content classification data, and extracts the real-time interest index of the target user from the image information of the person. User profile building module 14 is used to build static and dynamic tags based on the basic attribute information and behavioral information of the target user, and merge them to generate a multi-dimensional profile of the target user. The recommendation strategy generation module 15 is used to determine the recommended content, information density and presentation format of personalized interactive content based on multi-dimensional profiles, and to generate the recommended duration of personalized interactive content based on user duration preference model and real-time interest index. The interactive display module 16 is used to achieve interactive display through an immersive interactive terminal after fusing the corresponding multimodal data.

[0063] The information acquisition module 11 is specifically used for: After a user enters the immersive interactive system of the cultural heritage scene, the system obtains the user's basic attribute information, which includes at least: the user's self-declared identity attributes, the level of knowledge in the cultural heritage field determined through cultural heritage cognition association, the user's actively selected interactive language preferences, and the user's declared information receiving assistance needs.

[0064] Specifically, the verification module 12 is used for: Verify the authenticity of basic attribute information, including: The user's image information is matched with the biometric information corresponding to the identity attribute in the basic attribute information, and the matching degree is calculated. If the matching degree is greater than or equal to the preset matching degree threshold, the authenticity verification is deemed to have passed; if the matching degree is less than the preset matching degree threshold, the authenticity verification is deemed to have failed, and a secondary verification process is triggered.

[0065] Among them, the behavior acquisition and analysis module 13 is specifically used for: Real-time data collection of target user behavior during immersive interaction, including dwell time, question data, active operation data, and skipped content category data, and extraction of the target user's real-time interest index from the user's image information, including: Record the entry and exit timestamps of target users on each immersive interactive page, and calculate the dwell time; Collect question data input by target users through the interactive system, including voice or text question data; Record the active operation data of the target users, including at least the click to view specific cultural relics content, the selection and switching of content modules, and the interactive commands to the cultural relics visualization model; Identify skip actions triggered by the target user, categorize the skip actions by content type, and use this as skip content classification data; By integrating dwell time, question data, active behavior data, and skipped content category data, behavioral information of the target user is formed. Using computer vision algorithms, facial expressions and body postures of target users are identified from human image information and quantified into a real-time interest index according to preset scoring rules.

[0066] The user profile building module 14 is specifically used for: Static and dynamic tags are constructed based on the target user's basic attribute and behavioral information, and then merged to generate a multi-dimensional profile of the target user, including: Based on the mapping of the target user's basic attribute information, static tags are generated, including identity attribute tags, knowledge level tags, interaction language preference tags, and information reception assistance requirement tags; Dynamic tags are constructed based on the behavioral information of target users. These dynamic tags include interest tags, focus tags, preferred content tags, and content rejection tags. According to the preset fusion rules, static tags and dynamic tags are integrated to form a multi-dimensional profile that includes user attribute characteristics, cognitive level characteristics, interest preference characteristics, and content rejection characteristics.

[0067] Among them, dynamic tags are constructed based on the target user's behavioral information, including: Content with a dwell time greater than or equal to a preset dwell time threshold is marked as content of interest. Content of interest is then clustered into similar categories to form interest preference tags. The system uses a pre-trained natural language processing model in the cultural heritage field to analyze the question data, extract the demand keywords and count their frequency, select high-frequency demand keywords that meet the preset question frequency threshold, and perform semantic clustering to form focus direction tags. Active operation content is extracted from active operation behavior data and its frequency is counted. High-frequency active operation content that meets the preset operation frequency threshold is selected, and similar clusters are formed to create preferred content tags. Skipped content category data is counted by content type, and content types with a skip count greater than or equal to a preset skip count threshold are used as content rejection tendency tags. Integrate interest-oriented tags, focus-oriented tags, preferred content tags, and content aversion tags to form dynamic tags.

[0068] The recommendation strategy generation module 15 is specifically used for: Based on multi-dimensional profiles, the recommended content, information density, and presentation format of personalized interactive content are determined, including: Based on the interest and preference features and content rejection features in the multi-dimensional profile, personalized interactive content recommendations are generated from the preset content library; Information density of personalized interactive content generated based on cognitive level features in multi-dimensional profiles; Personalized interactive content is generated based on user attribute features from multi-dimensional profiles.

[0069] Furthermore, the recommended duration of personalized interactive content is generated based on a user time preference model and a real-time interest index, including: Obtain a pre-trained user time preference model; Input the average dwell time associated with the target user's interest tags, and output the predicted recommendation time of personalized interactive content; The predicted recommendation duration is adjusted according to the scenario to obtain the recommendation duration of personalized interactive content.

[0070] Specifically, the predicted recommendation duration is adjusted for specific scenarios to obtain the recommendation duration for personalized interactive content, including: The average effective reception time of each personalized interactive content type in the cultural heritage scenario is obtained as the baseline duration for the scenario. The average effective reception time of user groups whose multi-dimensional profiles of the target users meet the preset similarity threshold for each personalized interactive content type is obtained as the group reference time. The weighted average of the predicted recommendation duration, the baseline duration of the scenario, and the reference duration of the group is calculated, and then multiplied by the real-time interest index to obtain the recommendation duration of personalized interactive content.

[0071] The interactive display module 16 is specifically used for: After integrating the corresponding multimodal data, interactive display is achieved through an immersive interactive terminal.

[0072] It should be noted that the order of the embodiments described above is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, the above description focuses on specific embodiments of this specification. Additionally, the processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired results. In some implementations, multitasking and parallel processing are possible or may be advantageous.

[0073] The above description is only a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

[0074] This specification and accompanying drawings are merely illustrative examples of this application and are intended to cover any and all modifications, variations, combinations, or equivalents within the scope of this application. Clearly, those skilled in the art can make various alterations and modifications to this application without departing from its scope. Therefore, if such modifications and modifications fall within the scope of this application and its equivalents, this application intends to include such modifications and modifications.

Claims

1. A method for immersive interactive display of multimodal data visualization in cultural heritage scenarios, characterized in that: include: After a user enters the immersive interactive system of the cultural heritage scene, the system obtains the user's basic attribute information. Collect the user's image information, verify the authenticity of the basic attribute information, and define the user who passes the authenticity verification as the target user; Real-time collection of the target user's behavioral information during the immersive interaction process, including dwell time, question data, active operation behavior data, and skipped content category data, and extraction of the target user's real-time interest index from the person image information; Based on the target user's basic attribute information and behavioral information, static and dynamic tags are constructed and fused to generate a multi-dimensional profile of the target user. Based on the multi-dimensional profile, the recommended content, information density, and presentation format of personalized interactive content are determined, and the recommended duration of personalized interactive content is generated based on the user time preference model and the real-time interest index. After integrating the corresponding multimodal data, interactive display is achieved through an immersive interactive terminal.

2. The immersive interactive display method for multimodal data visualization in cultural heritage scenarios according to claim 1, characterized in that, The basic attribute information includes at least: the user's self-declared identity attributes, the knowledge level in the cultural heritage field determined through cultural heritage cognition association, the user's actively selected interactive language preferences, and the user's declared information receiving assistance needs.

3. The immersive interactive display method for multimodal data visualization in cultural heritage scenes according to claim 1, characterized in that, The authenticity of the basic attribute information is verified, including: The user's image information is matched with the biometric information corresponding to the identity attribute in the basic attribute information, and the matching degree is calculated. If the matching degree is greater than or equal to the preset matching degree threshold, the authenticity verification is deemed to have passed; if the matching degree is less than the preset matching degree threshold, the authenticity verification is deemed to have failed, and a secondary verification process is triggered.

4. The immersive interactive display method for multimodal data visualization in cultural heritage scenarios according to claim 1, characterized in that, Real-time collection of behavioral information of the target user during immersive interaction, including dwell time, question data, active operation behavior data, and skipped content category data, and extraction of the target user's real-time interest index from the user's image information, including: Record the entry and exit timestamps of the target user on each immersive interactive page, and calculate the dwell time; Collect question data input by the target user through the interactive system, wherein the question data includes voice question information or text question information; Record the active operation behavior data of the target user, wherein the active operation behavior data includes at least the click to view specific cultural relic content, the selection and switching of content modules, and the interactive commands of the cultural relic visualization model; Identify the skip operation content actively triggered by the target user, classify the skip operation content according to content type, and use it as skip content classification data; By integrating dwell time, question data, active behavior data, and skipped content category data, behavioral information of the target user is formed. Using computer vision algorithms, the facial expressions and body postures of the target user are identified from the image information of the person, and quantified into a real-time interest index according to a preset scoring rule.

5. The immersive interactive display method for multimodal data visualization in cultural heritage scenes according to claim 1, characterized in that, Based on the target user's basic attribute and behavioral information, static and dynamic tags are constructed and fused to generate a multi-dimensional profile of the target user, including: Based on the target user's basic attribute information mapping, static tags are generated that include identity attribute tags, knowledge level tags, interaction language preference tags, and information reception assistance requirement tags; Dynamic tags are constructed based on the behavioral information of the target users, wherein the dynamic tags include interest tendency tags, focus direction tags, preferred content tags, and content rejection tendency tags; According to the preset fusion rules, the static tags and dynamic tags are integrated to form a multi-dimensional profile that includes user attribute features, cognitive level features, interest preference features, and content rejection features.

6. The immersive interactive display method for multimodal data visualization in cultural heritage scenes according to claim 5, characterized in that, Dynamic tags are constructed based on the target user's behavioral information, including: Content with a dwell time greater than or equal to a preset dwell time threshold is marked as content of interest. The content of interest is then clustered into similar categories to form interest preference tags. The question data is analyzed using a pre-trained natural language processing model in the cultural heritage field, the demand keywords are extracted and their frequency is counted, high-frequency demand keywords that meet the preset question frequency threshold are selected, and semantic clustering is performed to form attention direction tags. Active operation content is extracted from active operation behavior data and its frequency is counted. High-frequency active operation content that meets the preset operation frequency threshold is selected, and similar clusters are formed to create preferred content tags. Skipped content category data is counted by content type, and content types with a skip count greater than or equal to a preset skip count threshold are used as content rejection tendency tags. The aforementioned interest-oriented tags, focus-oriented tags, preferred content tags, and content rejection tags are integrated to form the dynamic tags.

7. The immersive interactive display method for multimodal data visualization in cultural heritage scenarios according to claim 1, characterized in that, Based on the multi-dimensional profile, the recommended content, information density, and presentation format of personalized interactive content are determined, including: Based on the interest preference features and content rejection features in the multi-dimensional profile, personalized interactive content recommendations are generated from the preset content library; Information density of personalized interactive content generated based on cognitive level features in the multi-dimensional profile; The presentation format of personalized interactive content is generated based on the user attribute features in the multi-dimensional profile.

8. The immersive interactive display method for multimodal data visualization in cultural heritage scenes according to claim 7, characterized in that, The recommended duration of personalized interactive content is generated based on the user time preference model and the real-time interest index, including: Obtain a pre-trained user time preference model; Input the average dwell time associated with the target user's interest tags, and output the predicted recommendation time of personalized interactive content; The predicted recommendation duration is adjusted for scene adaptation to obtain the recommendation duration of personalized interactive content.

9. The immersive interactive display method for multimodal data visualization in cultural heritage scenes according to claim 8, characterized in that, The predicted recommendation duration is adjusted for scene adaptation to obtain the recommendation duration of personalized interactive content, including: The average effective reception time of each personalized interactive content type in the cultural heritage scenario is obtained as the baseline duration for the scenario. The average effective reception time of user groups whose multi-dimensional profiles of the target users meet the preset similarity threshold for each personalized interactive content type is obtained as the group reference time. The weighted average of the predicted recommendation duration, the scene baseline duration, and the group reference duration is calculated and then multiplied by the real-time interest index to obtain the recommendation duration of personalized interactive content.

10. An immersive interactive display device for multimodal data visualization in cultural heritage scenarios, characterized in that: The method for performing the immersive interactive display of multimodal data visualization in cultural heritage scenes according to any one of claims 1-9 includes: The information acquisition module is used to acquire the user's basic attribute information after the user enters the immersive interactive system of the cultural heritage scene; The verification module is used to collect the user's image information, verify the authenticity of the basic attribute information, and define users who pass the authenticity verification as target users. The behavior collection and analysis module is used to collect the target user's behavior information in real time during the immersive interaction process. The behavior information includes dwell time, question data, active operation behavior data and skipped content classification data, and extracts the target user's real-time interest index from the person image information. The user profile building module is used to build static and dynamic tags based on the target user's basic attribute information and behavioral information, and then merge them to generate a multi-dimensional profile of the target user. The recommendation strategy generation module is used to determine the recommended content, information density and presentation format of personalized interactive content based on the multi-dimensional profile, and to generate the recommended duration of personalized interactive content based on the user duration preference model and the real-time interest index. The interactive display module is used to integrate the corresponding multimodal data and then achieve interactive display through an immersive interactive terminal.