A virtual reality interaction method and system for displaying classical paintings
By constructing a three-dimensional spatial model and semantic label matching, combined with user behavior analysis, the problem of inconsistent cultural context and artistic atmosphere in the display of famous paintings in virtual reality is solved, and immersive and personalized display of famous paintings is realized, which enhances users' artistic cognition and interactive experience.
Patent Information
- Application Number
- CN202510855280.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2045-06-25
AI Technical Summary
The existing virtual reality famous painting display system does not fully consider the cultural context and artistic atmosphere of classical famous paintings, resulting in the interactive mode that may interrupt or misunderstand the historical situation of the painting's expression.
Through multi-dimensional high-definition image acquisition, three-dimensional spatial model is constructed, cultural context information is extracted, and semantic classification and labeling is performed. Real-time matching is combined with user behavior data, interactive feedback is generated consistent with the painting style, and typical interaction paths and misunderstanding areas are identified through clustering analysis, and context labels and response logic are dynamically adjusted.
It realizes immersive and semantic-driven virtual reality interaction of classical paintings, improves users' cognitive accuracy and cultural experience quality of classical art, and ensures the consistency of interactive feedback and artistic style.
Smart Images

Figure CN120428866B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of virtual reality interaction technology, and in particular to a virtual reality interaction method and system applied to the display of classical paintings. Background Art
[0002] Virtual reality interaction applied to the display of classical masterpieces utilizes virtual reality (VR) technology to present classic paintings in an immersive and interactive manner. Through VR, viewers can immerse themselves in the depicted scenes and even interact with the elements within them, such as zooming in on details, hearing backstory stories, or watching a simulation of the artist's creative process. This technology not only enhances the exhibition experience but also facilitates a deeper understanding of the historical context and artistic value of the works.
[0003] The existing technology has the following shortcomings:
[0004] In existing VR systems for displaying masterpieces, viewers often interact with the paintings through gestures, gaze, or voice (e.g., clicking on a character to trigger a narration, focusing on a specific element to trigger an animation). However, these interaction methods often rely on common user interface logic (e.g., web page clicks, game-like task flows), failing to fully consider the cultural context and artistic atmosphere of classical masterpieces. This can result in the historical context of the paintings being interrupted or misinterpreted. Summary of the Invention
[0005] The purpose of the present invention is to provide a virtual reality interaction method and system for displaying classical paintings, so as to overcome the shortcomings of the background technology.
[0006] To achieve the above-mentioned object, the present invention provides the following technical solution: a virtual reality interaction method for displaying classical paintings, comprising:
[0007] Capture multi-dimensional high-definition images of target classical paintings and construct a three-dimensional spatial model based on the compositional elements of the paintings;
[0008] Extract cultural context information related to the paintings and perform semantic classification and labeling based on element dimensions;
[0009] Obtain behavioral data related to the user's gaze point, movement posture, and duration of stay, and perform real-time matching based on contextual tags;
[0010] Generate interactive feedback information consistent with the current painting style based on user behavior data and matched cultural context tags;
[0011] Cluster analysis is performed based on the interaction behavior data of several users to identify typical interaction paths and areas with high incidence of misunderstandings. Context labels and response logic are dynamically adjusted to achieve personalized recommendations.
[0012] Preferably, multi-dimensional high-definition image acquisition of the target classical painting includes: using a ring light source device to achieve uniform illumination, using structured light scanning or laser ranging to collect surface texture information of the painting; collecting shallow parallax image sequences from the front of the painting and multiple micro-angles such as top, bottom, left, right, front and back; fusing image data in different bands, including visible light, infrared and ultraviolet imaging, to obtain information on the underlying drawing and pigment layers.
[0013] Preferably, a three-dimensional space model based on composition elements is constructed: the visual elements in the picture, including characters, background, symbols and light sources, are extracted using an image semantic segmentation algorithm; spatial depth relationships are established based on the composition features of perspective lines, main viewing angles and symmetry axes; a depth map is generated using structured light or a multi-view stereo algorithm, and a three-dimensional point cloud is constructed based on the depth map; a three-dimensional mesh model is generated through point cloud reconstruction or mesh generation technology, and the original high-definition image is mapped to the model surface as a texture map.
[0014] Preferably, matching user behavior data with context tags in real time includes:
[0015] Bind a unique ID and a set of semantic tags to each composition element;
[0016] When the user's gaze or gesture falls within the range of a certain composition element, the semantic tag corresponding to the element is triggered;
[0017] The matching mechanism is based on the behavior-semantic mapping rules, combining the user behavior type, action frequency, and viewing duration to perform similarity weighted calculation;
[0018] If the behavior matching degree exceeds the set threshold, a trigger event is generated.
[0019] Preferably, cluster analysis is performed based on the interactive behavior data of several users, including: encoding user behavior data into high-dimensional feature vectors, including gaze heat, behavior frequency, label preference and path sequence, clustering the behavior data, and extracting typical interaction paths and misunderstanding behavior areas in each cluster; and using the results to update the user portrait library and optimize the context recommendation logic.
[0020] Preferably, after analyzing the wandering characteristics of the user's behavior trajectory in the misunderstanding area, a perspective switching anomaly value is generated. The generation method is as follows: the perspective data generated by each user in the VR system includes: : The i-th timestamp; for every two consecutive time points and , calculate the rate of change of its viewing angle , the expression is: ; In a time window w, count the directional fluctuation value of the user's perspective change , the expression is: ;in, is the viewing angle vector in the current window, is the average viewing angle direction in the window, n is the total number of time points; the obtained viewing angle change rate Directional fluctuation value with viewing angle change Normalization processing is performed, and the normalized perspective change rate and the perspective change direction fluctuation value are weighted averaged and calculated to obtain the perspective switching anomaly value; when the perspective switching anomaly value exceeds the set threshold, the picture area is determined to be a perspective wandering abnormal area.
[0021] Preferably, a semantic deviation value is generated after comparing and analyzing the semantic labels bound to the abnormal view wandering area with the user's feedback behavior. The generation method is:
[0022] Build or access an existing art knowledge graph, uniquely encode all nodes, and attach category attributes to them, converting the user's interactive behavior in virtual reality into an intention node u: Use a graph traversal algorithm to find the shortest semantic path length between nodes u and t in the graph. , the expression is: ; Set the ideal understanding depth of each target label t to Indicates the normal depth of the path from the knowledge root node to the label t in the ideal understanding path, and defines the semantic deviation value as: ; Where SDS is the semantic deviation value.
[0023] Preferably, the perspective switching anomaly values and semantic deviation values are converted into comprehensive feature vectors, and the comprehensive feature vectors are used as inputs of the machine learning model. The machine learning model uses each group of comprehensive feature vectors to predict the comprehensive misunderstanding analysis coefficient value labels in each region as the prediction target, and uses minimizing the sum of the prediction errors of the comprehensive misunderstanding analysis coefficient value labels in all regions as the training target. The machine learning model is trained until the sum of the prediction errors reaches convergence, and the model training is stopped. The comprehensive misunderstanding analysis coefficient value in each region is determined according to the model output results, wherein the machine learning model is a polynomial regression model.
[0024] Preferably, the obtained comprehensive misunderstanding analysis coefficient value in each area is compared with a gradient standard threshold, the gradient standard threshold includes a first standard threshold and a second standard threshold, and the first standard threshold is less than the second standard threshold, and the comprehensive misunderstanding analysis coefficient value in each area is compared with the first standard threshold and the second standard threshold respectively;
[0025] If the comprehensive misunderstanding analysis coefficient value in each area is greater than the second standard threshold, it is marked as a high-risk area and requires optimization of labels, feedback or guidance content;
[0026] If the comprehensive misunderstanding analysis coefficient value in each area is greater than or equal to the first standard threshold and less than or equal to the second standard threshold, it is marked as a medium-risk area and supplementary explanation or voice guidance is introduced;
[0027] If the comprehensive misunderstanding analysis coefficient value in each area is less than the first standard threshold, it is marked as a low-risk area and used as an example area.
[0028] The present invention also provides a virtual reality interactive system for displaying classical paintings, comprising an image acquisition module, a context information management module, a behavior perception module, an interactive feedback module, and a semantic optimization module;
[0029] Image acquisition module: This module collects multi-dimensional high-definition images of target classical paintings and constructs a three-dimensional spatial model based on the compositional elements of the paintings;
[0030] Contextual information management module: extracts cultural contextual information related to paintings and performs semantic classification and labeling based on element dimensions;
[0031] Behavior perception module: obtains behavioral data related to the user's gaze point, movement posture, and dwell time, and combines it with contextual tags for real-time matching;
[0032] Interactive feedback module: Generates interactive feedback information consistent with the current painting style based on user behavior data and matched cultural context labels;
[0033] Semantic Optimization Module: Performs cluster analysis based on the interaction behavior data of several users, identifies typical interaction paths and areas with high incidence of misunderstandings, and dynamically adjusts context labels and response logic to achieve personalized recommendations.
[0034] In the above technical solution, the technical effects and advantages provided by the present invention are:
[0035] 1. This invention integrates high-definition image acquisition, 3D modeling, cultural context semantic tagging, and user behavior perception technologies to provide an immersive, semantically driven method for interacting with classical paintings in virtual reality. Compared to existing display methods based on general interaction logic, this invention not only achieves high-fidelity restoration of the painting's compositional structure and artistic style, but also establishes a behavioral-semantic mapping mechanism based on cultural understanding, imbuing interactive feedback with greater artistic consistency and cultural depth.
[0036] 2. This invention incorporates user behavior cluster analysis and machine learning prediction models to identify typical user paths and areas of high misunderstanding. It also dynamically optimizes contextual labels and response logic, enabling personalized recommendations and adaptive system adjustments. Overall, this approach significantly improves users' cognitive accuracy, interactive immersion, and cultural experience of classical art. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments described in the present invention. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.
[0038] Figure 1 This is a mind map of the method of the present invention.
[0039] Figure 2 This is a mind map of the system modules of the present invention. DETAILED DESCRIPTION
[0040] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0041] Example 1, please refer to Figure 1 As shown, the virtual reality interaction method for displaying classical paintings described in this embodiment includes:
[0042] Capture multi-dimensional high-definition images of target classical paintings and construct a three-dimensional spatial model based on the compositional elements of the paintings;
[0043] Extract cultural context information related to the paintings and perform semantic classification and labeling based on element dimensions;
[0044] Obtain behavioral data related to the user's gaze point, movement posture, and duration of stay, and perform real-time matching based on contextual tags;
[0045] Generate interactive feedback information consistent with the current painting style based on user behavior data and matched cultural context tags;
[0046] Cluster analysis is performed based on the interaction behavior data of several users to identify typical interaction paths and areas with high incidence of misunderstandings. Context labels and response logic are dynamically adjusted to achieve personalized recommendations.
[0047] A ring light fixture is used to control lighting uniformity, avoiding reflections, shadows, or color distortion. Reflected images from different time periods or wavelengths are synthesized to capture pigment details and textures (e.g., cracks and brushstrokes).
[0048] Based on a fixed frame, the image is captured from multiple micro-angles (e.g., up, down, left, right, front, and back) to form a shallow parallax image sequence for subsequent depth of field analysis. Laser scanning or structured light acquisition technology is used to record the image's relief (e.g., relief painting) and canvas structure information.
[0049] Introducing multispectral imaging such as infrared and ultraviolet, we can deeply restore the original painting, covering layers and color-degraded areas, and enhance the comprehensive understanding of the painting structure.
[0050] Use image semantic segmentation algorithms (such as U-Net and DeepLabV3+) to identify the main compositional elements in the image: people, background, objects, light source direction, etc. Label the elements in layers and mark their relative size and position on the 2D canvas.
[0051] Use a combination of deep learning and geometric analysis to identify basic artistic structures such as the main perspective, vanishing point, perspective lines, composition axes (such as the axis of symmetry and the rule of thirds) of the picture.
[0052] Build a spatial perspective framework based on visual guide lines to provide a geometric foundation for three-dimensional models.
[0053] Combining images captured from multiple angles, a depth map is generated using structured light reconstruction or parallax calculation algorithms (such as SFM and MVS). For purely two-dimensional works, depth estimation is generated through vanishing point and composition logic simulation, creating a "stage-like" 3D model.
[0054] Utilize point cloud reconstruction or mesh generation techniques (such as Poisson Surface Reconstruction) to reconstruct the 3D structures in the image, such as buildings, human figures, and background layers. Layer the high-definition image as a texture map and map it onto the 3D model surface, preserving the original image details.
[0055] Once constructed, the 3D spatial model is imported into a VR engine (such as Unity or Unreal Engine) to adapt to changes in perspective caused by user head movements or interactive actions, providing a dynamic visual response. During the model construction process, artistic style constraints are introduced to prevent the original aesthetic of the image from being disrupted by a "heaviness" or "dislocation" of the model. For example, the consistency of the perspective ratio, light and shadow direction, and light and dark relationships in the original work is maintained; for stylized compositions, corresponding geometric constraint templates are introduced.
[0056] Acquire multi-source semantic data related to paintings, including but not limited to: painters’ autobiography or letters, art history literature and journal reviews, museum commentaries, exhibition scripts, expert review articles and cultural research materials, authoritative encyclopedias and art knowledge bases.
[0057] Natural language processing (NLP) tools are used to perform word segmentation, entity recognition, and syntactic analysis on the text to extract key semantic units related to the composition of the picture, the historical context of the scene background, and the meaning of the image symbols.
[0058] Divide the interactive or visually focused objects in the picture into several composition element dimensions, such as: character elements, background elements, symbolic elements, and visual guide elements.
[0059] Each compositional element is assigned a number of cultural contextual labels, including historical labels (time, location, and event associations), symbolic labels, stylistic labels (school of painting, perspective type, and color expression), and emotional labels (e.g., solemnity, warmth, mystery, and anxiety). All labels are uniformly encoded using a knowledge graph structure or multi-label embedding matrix to support subsequent rapid query, matching, and semantic reasoning.
[0060] To adapt to audiences from different cultural backgrounds, the system allows the label library to be expanded into multiple languages and cultural backgrounds. Through audience feedback or expert proofreading mechanisms, labels are continuously updated and revised to form a learnable context model.
[0061] Through virtual reality devices, users' behavioral characteristics (such as gaze, gestures, body movements, etc.) are accurately identified, and these behavioral data are matched with the established cultural context label system in real time to trigger appropriate interactive feedback.
[0062] The built-in infrared camera of the VR headset is used to collect the user's pupil position and movement path; gaze trajectory data and heat maps are generated to determine the user's gaze focus in real time; the accuracy is preferentially controlled within the ±1° viewing angle range to trigger the gaze time threshold judgment.
[0063] With the help of handle sensors, hand tracking (Leap Motion, Ultraleap) or full-body motion capture devices, user motion data is collected: finger clicks, swipes, pointing; body tilts, steps, avoidance movements; posture angle changes (pitch, yaw, roll).
[0064] The system determines the user's potential interest based on the time they stay in front of a specific area or element (e.g., >3 seconds). Combined with spatial positioning information (e.g., Collider + Time in Unity), it can determine whether there is "interaction intent." The system can prioritize the same user's behavior in different areas and archive their behavior trajectories.
[0065] Each composition element (such as a person or symbol) is bound to a unique ID and a set of semantic tags in the virtual space; when the gaze point or hand movement is within the boundary of the element (such as a 3D bounding box or a sight ray hit), an index match is triggered.
[0066] In this invention, the "behavior-semantic mapping" mechanism serves as the core intermediary, matching users' specific behavior patterns with an established semantic tagging system in real time to identify their interaction intent and drive the generation of subsequent feedback content. This mechanism is based on the logical mapping relationship between multiple common behavior types and semantic elements in paintings, specifically including the following typical scenarios:
[0067] Matching gaze behavior with character semantics: When a user gazes at a character in a painting through a VR headset and the gaze time exceeds a preset threshold (for example, more than 2 seconds), the system will determine that the user has a specific attention intention towards the character.
[0068] Matching pointing behavior with symbolic semantics: When a user selectively points to an object in the image, such as a white dove, a scroll, or a key, using hand gestures or a controller, the system recognizes this as a request to interact with a specific symbol. If the element's label indicates it symbolizes "peace," "wisdom," or "power," the system can guide the user to relevant cultural context, interpret the meaning of artistic symbolism, or trigger a display module with supplementary historical information.
[0069] Matching dwell behavior with the semantics of the background scene: When a user dwells for an extended period on a particular background area, such as a church arch or a cityscape, the system determines their exploration intent based on their spatial location and dwell time. If the area is labeled "Gothic Architecture" or "Renaissance Urban Environment," the system will automatically present an analysis of the architectural style, an introduction to the social context of the time, or other information modules related to the semantics of the area.
[0070] Intelligent judgment based on the fusion of multiple behaviors: In some cases, users may generate multiple behavioral signals simultaneously, such as gazing at a person's face while pointing at an object they hold. In these situations, the system comprehensively assesses the intensity of the behavior, the weight of the target element, and the frequency of interaction, prioritizing semantic paths that are more consistent with the primary viewpoint. For example, if the person is a sage holding a scroll symbolizing knowledge, the system will prioritize "character identity interpretation" and subsequently supplement it with "symbol interpretation," ensuring coherence in the interaction logic and in-depth expression of cultural context.
[0071] The matching mechanism selects the optimal context path response by judging the similarity weight between behavioral features and label attributes.
[0072] When multiple user behaviors (such as gazing + pointing) occur simultaneously, the system uses a weighted matching strategy to determine the interaction direction based on the behavior priority and contextual environment; preventing false triggering or semantic misinterpretation, and improving interaction stability.
[0073] All behavior recognition and matching processing are run locally on the client using a low-latency module, with response times kept at <50ms. Matching results are fed back to the interaction engine (such as the Unity Event System) in real time to trigger composite visual, auditory, or tactile feedback. All interaction data is uploaded to the backend database for user portrait construction and context system optimization.
[0074] Based on real-time user behavior and matched cultural contextual tags, interactive feedback is dynamically generated, aligning with the style and narrative logic of classic paintings. This immersive multimodal feedback mechanism enhances the user's perceptual richness and contextual immersion in the VR scene, while ensuring that the feedback does not undermine the original artistic style and historical context, achieving a balance between "technical presentation and artistic language."
[0075] The system first obtains the following information combination through the "behavior recognition" and "semantic label matching" completed in the previous stage:
[0076] The screen element that the user is currently looking at or interacting with; the cultural context label to which the element belongs; the user behavior type and its intensity (such as dwell time, gesture pointing, or dynamic approach); and the artistic style characteristics of the current scene. These information together constitute the input trigger conditions for the feedback content generation module.
[0077] Based on the historical background or artistic analysis content associated with semantic tags, the corresponding explanation text in the voice library is called; a speech synthesis style that fits the style of the painting is adopted, such as using a low and solemn tone when displaying religious paintings, and a soft and gentle voice in the Rococo style; the explanation trigger method can be that the user stares still, makes an active choice, or the system intelligently identifies their area of interest; supports segmented explanations and structured narratives (character → scene → symbol → artistic style), so that users can gain a coherent knowledge experience through immersion.
[0078] Non-intrusive visual highlighting is performed on interactive target elements, such as using soft apertures and low-contrast strokes to gently highlight them from the picture. The style of the light effect animation is consistent with the overall tone of the painting, such as the Gothic style uses sharp light and shadow cuts, and Impressionist works use soft fluctuating halos. The picture texture is not directly modified to ensure the visual integrity of the original work and enhance the effect of guiding user attention.
[0079] Based on the era, scene, and cultural labels of the image, historical soundscapes or situational sound effects are used for spatial arrangement. The clamor of a market and the sound of horse hooves can be restored against the backdrop of a medieval street scene. The soundscape uses spatial audio rendering, and the position of the sound source changes as the user moves, enhancing the sense of presence. The soundscape can be triggered based on regional residence or behavioral intention.
[0080] Static elements with strong symbolic meaning in the painting are non-destructively and lightweightly animated: the animation style follows the original color and brushstroke style, using a "pseudo-painting style" rendering algorithm to avoid a sense of disconnection from the original painting; the animation's performance rhythm is controlled within the principles of "slow, respectful, and guiding", mainly presenting auxiliary information without causing visual interference.
[0081] To prevent interactive feedback from destroying the original artistic style of classical paintings, the system introduces the following style adaptation mechanism:
[0082] All feedback content (voice, lighting effects, soundscapes, animations) is bound to the style attributes of the current painting. Before the feedback content is called, it is subjected to a "style consistency check" to ensure that the visual language, timbre, mood, and rhythm control are consistent with the tonality of the painting. An artistic style template library is set up as a reference for the style rendering of the feedback content.
[0083] Leveraging user behavioral data within the VR art display system and using machine learning methods like cluster analysis, we automatically identify high-frequency interaction paths (i.e., common viewing methods) and areas prone to cognitive bias or misunderstanding. This analysis can be used to optimize subsequent interactive guidance design, adjust feedback content, and supplement explanatory information, thereby enhancing the system's intelligent adaptability and the consistency of the viewing experience.
[0084] Collect behavior logs of multiple users in a virtual reality environment, mainly including the following data fields:
[0085] User ID (anonymized), timestamp sequence, gaze point coordinates (Gaze X, Y, Z), gesture and operation events (click, point, zoom, etc.), dwell time (residence time in each screen area), behavior path (action sequence arranged in chronological order), element ID matching semantic tags, data is stored in a structured format (such as JSON or CSV) as the basis for subsequent analysis.
[0086] User behavior data is converted into a vector feature space for clustering. Each user is represented as a set of vectors, such as gaze point heat distribution (a two-dimensional grid distribution), interaction frequency vector (the number of times different elements are triggered), time series path embedding (processed using LSTM or dynamic time warping (DTW)), and semantic dimension distribution (the tag categories that the user pays the most attention to). Vector features can be combined into sparse matrices or high-dimensional embeddings for training clustering models.
[0087] Select appropriate clustering methods (such as K-Means, DBSCAN, Gaussian Mixture Model) to cluster the behavior vectors and obtain the following results:
[0088] Several high-frequency interaction pattern clusters (e.g., rapid skimming, deep exploration, symbol preference, etc.);
[0089] Average path, average dwell time, and distribution of attention semantic tags of users in each mode;
[0090] The visualization results are presented in the form of heat maps, path overlay maps or 3D model playback.
[0091] The specific steps to identify typical interaction paths include:
[0092] Step 1: Perform time series normalization on the user behavior trajectories within each cluster to make the paths comparable in space and time; extract highly overlapping path segments, i.e., path intervals that are repeatedly visited or gazed upon by multiple users, and define them as typical path nodes; construct a path map and connect these nodes to form the audience's subjective path network.
[0093] Step 2: Classify typical paths according to their semantic preference, interaction density, and start-end structure. For example: start with the protagonist → gradually explore the background → linger on the symbolic objects; browse from the periphery of the scene → quickly enter the core character → quickly exit; use this to build interaction guidance strategies, such as recommended routes, visual guide lines, or interaction entry optimization.
[0094] If multiple users frequently trigger inconsistent behaviors in a certain area, it is marked as a misinterpretation area. Such behaviors include: long gaze time without further interaction; excessive clicks without successful feedback (indicating that the user expects certain content but the system does not provide it); and chaotic and wandering behavior trajectories (such as frequent switching of perspectives and returning to repeated areas).
[0095] After analyzing the wandering characteristics of the user's behavior trajectory in the misunderstanding area, the perspective switching anomaly is generated. The generation method is as follows:
[0096] The perspective data generated by each user in the VR system includes:
[0097] : the i-th timestamp;
[0098] : The viewing direction at that moment, expressed as azimuth (horizontal rotation) and pitch (vertical rotation) , the unit is degree.
[0099] : The location coordinates at that moment, used to assist in judging mobility and spatial wandering behavior.
[0100] For every two consecutive time points and , calculate the rate of change of its viewing angle , the expression is: ; Within a time window w (e.g. 5 seconds), count the directional fluctuation values of the user's perspective change , the expression is: ;in, is the viewing angle vector in the current window, is the average viewing direction within the window, and n is the total number of time points.
[0101] The obtained viewing angle change rate Directional fluctuation value with viewing angle change Normalization processing is performed, and the normalized view angle change rate and the view angle change direction fluctuation value are weighted averaged and summed to obtain the view angle switching anomaly value.
[0102] When the perspective switching abnormal value exceeds the set threshold (for example, set to above the 75th percentile value based on historical behavior data), the screen area is determined to be a perspective wandering abnormal area.
[0103] Compare the semantic labels associated with unusual areas of camera wandering with user feedback to identify semantic deviations. For example, if a user frequently clicks on a background area, indicating they have questions about the background, but the system doesn't associate a semantic label with it or the feedback is overly simplistic.
[0104] After comparing and analyzing the semantic labels bound to the abnormal view wandering area with the user's feedback behavior, a semantic deviation value is generated. The generation method is as follows:
[0105] Constructing or accessing an existing art knowledge graph requires a hierarchical structure; all nodes are uniquely encoded (such as IDs) and accompanied by category attributes (characters, scenes, symbols, etc.); and can be represented using RDF / OWL standards or graph databases (such as Neo4j).
[0106] The user's interactive behavior in virtual reality is converted into an intention node u: if the user looks at and attempts to interact with a "pigeon" multiple times, the system can infer that their behavioral goal is "peace"; behavioral intention is generated through keywords, gaze objects, semantic history matching, etc.
[0107] Use a graph traversal algorithm (such as Dijkstra) to find the shortest semantic path length between nodes u and t in the graph , the expression is: ; If u=t, the deviation is 0 (completely consistent); if u and t are far apart, it means that the understanding deviation is large.
[0108] The ideal understanding depth of each target label t is predefined by the system or specified by experts. ; represents the normal depth of the path from the general knowledge root node to the label t in the ideal understanding path; can be regarded as the standard semantic cognitive distance.
[0109] Combining the above two paths, the semantic deviation value is defined as: Where SDS is the semantic deviation value. If SDS is equal to 1, it indicates that the user's understanding path is the same as the ideal path. If SDS is greater than 1, it indicates that the user's behavioral intention deviates from the expected semantic path, which may lead to misunderstanding. If SDS is less than 1, it may indicate that the user has a deeper semantic connection (which can be used as an expert mode feature).
[0110] The perspective switching anomalies and semantic deviation values are converted into comprehensive feature vectors, and the comprehensive feature vectors are used as the input of the machine learning model. The machine learning model uses each set of comprehensive feature vectors to predict the comprehensive misunderstanding analysis coefficient value label in each region as the prediction target, and takes minimizing the sum of the prediction errors of the comprehensive misunderstanding analysis coefficient value labels in all regions as the training target. The machine learning model is trained until the sum of the prediction errors reaches convergence, and the model training is stopped. The comprehensive misunderstanding analysis coefficient value in each region is determined according to the model output results, wherein the machine learning model is a polynomial regression model.
[0111] Comparing the obtained comprehensive misunderstanding analysis coefficient value in each area with the gradient standard threshold, the gradient standard threshold includes a first standard threshold and a second standard threshold, and the first standard threshold is less than the second standard threshold, and comparing the comprehensive misunderstanding analysis coefficient value in each area with the first standard threshold and the second standard threshold respectively;
[0112] If the comprehensive misunderstanding analysis coefficient value in each area is greater than the second standard threshold, it is marked as a high-risk area. The misunderstanding risk in this area is extremely high, and the label, feedback, or guidance content needs to be optimized;
[0113] If the comprehensive misunderstanding analysis coefficient value in each area is greater than or equal to the first standard threshold and less than or equal to the second standard threshold, it is marked as a medium-risk area, indicating possible cognitive bias, and it is recommended to introduce supplementary explanations or voice guidance;
[0114] If the comprehensive misunderstanding analysis coefficient value in each area is less than the first standard threshold, it is marked as a low-risk area, where users have a relatively consistent understanding and can be used as an interaction template or example area.
[0115] All areas are summarized by risk level; a heat map visualization is rendered, displaying the image in a color-coded manner, for example: red: high risk; orange: medium risk; green: low risk. The heat map can be overlaid on the VR scene as a background analysis tool or as a reference view for content optimization.
[0116] Collect and analyze historical user behavior characteristics, including: gaze frequency and path, dwell time distribution, gesture interaction records, and semantic tag preferences (semantic dimensions that users trigger more frequently);
[0117] Use unsupervised clustering algorithms (such as K-Means, DBSCAN, and HDBSCAN) to cluster users and form user behavior portrait model groups, such as visually oriented users, symbol-preferred users, fast-browsing users, and deep-exploration users.
[0118] The original contextual labels for each composition element are optimized based on the following factors:
[0119] A certain label appears repeatedly in a high-frequency misinterpretation area → triggering label semantic refinement or reassociation;
[0120] Multiple users focus on a certain image area, but the original label is unresponsive → The system recommends adding a new label dimension;
[0121] The user feedback path deviates significantly from the tag trigger content → triggering tag weight adjustment or hierarchical re-arrangement;
[0122] The tag adjustment method can be: tag confidence weighting (dynamically adjust the weight according to user trigger frequency and accuracy).
[0123] The interactive response content corresponding to each context label (such as voice explanation, light effect prompts, and animation feedback) is adjusted as follows:
[0124] Add pre-guidance to the response logic for areas with high misunderstanding, such as "background introduction" or "cultural tips";
[0125] Adjust the response order based on user profile preferences (e.g., prioritize conclusions for fast-paced users);
[0126] Response materials can switch styles based on user type (e.g. concise explanation / detailed interpretation / focused on visual prompts);
[0127] The response logic is dynamically loaded in an event-driven manner, processed in a modular manner, and supports on-demand switching.
[0128] Based on the user's current behavioral characteristics and historical profile, the system provides real-time recommendations: interactive entry points within specific paintings (to guide users to focus on areas they are more likely to understand or be interested in); semantic hierarchy for priority feedback (such as religious symbolism vs. character emotions); and supplementary content (such as additional historical background, the painter's life, a symbol dictionary, etc.). The recommendation mechanism can use heuristic strategies or simple classifiers (such as Decision Tree), running in real time, lightweight and efficient.
[0129] Example 2, please refer to Figure 2As shown, the virtual reality interactive system for displaying classical paintings described in this embodiment includes an image acquisition module, a context information management module, a behavior perception module, an interactive feedback module, and a semantic optimization module;
[0130] Image acquisition module: This module collects multi-dimensional high-definition images of target classical paintings and constructs a three-dimensional spatial model based on the compositional elements of the paintings;
[0131] Contextual information management module: extracts cultural contextual information related to paintings and performs semantic classification and labeling based on element dimensions;
[0132] Behavior perception module: obtains behavioral data related to the user's gaze point, movement posture, and dwell time, and combines it with contextual tags for real-time matching;
[0133] Interactive feedback module: Generates interactive feedback information consistent with the current painting style based on user behavior data and matched cultural context labels;
[0134] Semantic Optimization Module: Performs cluster analysis based on the interaction behavior data of several users, identifies typical interaction paths and areas with high incidence of misunderstandings, and dynamically adjusts context labels and response logic to achieve personalized recommendations.
[0135] The above formulas are all dimensionless and numerical calculations. The formulas are obtained by collecting a large amount of data and performing software simulation to obtain the most recent real situation. The preset parameters in the formulas are set by technicians in this field according to actual conditions.
[0136] It should be understood that the term "and / or" as used herein simply describes a relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A alone, A and B together, or B alone. A and B can be singular or plural. Furthermore, the character " / " as used herein generally indicates an "or" relationship between the associated objects, but it may also indicate an "and / or" relationship. For specific understanding, please refer to the context.
[0137] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0138] The above is only a specific implementation method of the present application, but the scope of protection of the present application is not limited thereto. Any technician familiar with this technical field can easily think of changes or replacements within the technical scope disclosed in this application, which should be covered by the scope of protection of the present application.
Claims
1. A virtual reality interaction method for displaying classical paintings, characterized by: include: Capture multi-dimensional high-definition images of target classical paintings and construct a three-dimensional spatial model based on the compositional elements of the paintings; Extract cultural context information related to the paintings and perform semantic classification and labeling based on element dimensions; Obtain behavioral data related to the user's gaze point, movement posture, and duration of stay, and perform real-time matching based on contextual tags; Generate interactive feedback information consistent with the current painting style based on user behavior data and matched cultural context tags; Cluster analysis is performed based on the interaction behavior data of several users to identify typical interaction paths and areas with high misunderstandings. Context labels and response logic are dynamically adjusted to achieve personalized recommendations. Specifically, this involves performing cluster analysis based on the interactive behavior data of several users, including encoding user behavior data into high-dimensional feature vectors, including gaze intensity, behavior frequency, tag preference, and path sequence, clustering the behavior data, and extracting typical interaction paths and misunderstood behavior areas in each cluster; using the results to update the user portrait library and optimize the contextual recommendation logic; After analyzing the wandering characteristics of the user's behavior trajectory in the misunderstanding area, the perspective switching anomaly value is generated. The generation method is as follows: the perspective data generated by each user in the VR system includes: : The i-th timestamp; for every two consecutive time points and , calculate the rate of change of its viewing angle , the expression is: Where, is the azimuth, is the pitch angle; within a time window w, the direction fluctuation value of the user's viewing angle change is counted , the expression is: ;in, is the viewing angle vector in the current window, is the average viewing angle direction in the window, n is the total number of time points; the obtained viewing angle change rate Directional fluctuation value with viewing angle change Normalization is performed, and the normalized view angle change rate and the view angle change direction fluctuation value are weighted averaged and calculated to obtain the view angle switching abnormality value; when the view angle switching abnormality value exceeds the set threshold, the image area is determined to be a view angle wandering abnormal area; The semantic labels bound to the abnormal view wandering area are compared and analyzed with the user's feedback behavior to generate a semantic deviation value. The generation method is as follows: construct or access an existing art knowledge graph, uniquely encode all nodes and attach category attributes, and convert the user's interactive behavior in virtual reality into an intention node u: use a graph traversal algorithm to find the shortest semantic path length between nodes u and t in the graph. , the expression is: ; Set the ideal understanding depth of each target label t to : ; Indicates the normal depth of the path from the knowledge root node to the label t in the ideal understanding path, and defines the semantic deviation value as: ; Where SDS is the semantic deviation value; The perspective switching anomalies and semantic deviation values are converted into comprehensive feature vectors, and the comprehensive feature vectors are used as inputs of the machine learning model. The machine learning model uses each set of comprehensive feature vectors to predict the comprehensive misunderstanding analysis coefficient value labels in each region as the prediction target, and takes minimizing the sum of the prediction errors of the comprehensive misunderstanding analysis coefficient value labels in all regions as the training target. The machine learning model is trained until the sum of the prediction errors reaches convergence, and the model training is stopped. The comprehensive misunderstanding analysis coefficient value in each region is determined based on the model output results, wherein the machine learning model is a polynomial regression model; Comparing the obtained comprehensive misunderstanding analysis coefficient value in each area with the gradient standard threshold, the gradient standard threshold includes a first standard threshold and a second standard threshold, and the first standard threshold is less than the second standard threshold, and comparing the comprehensive misunderstanding analysis coefficient value in each area with the first standard threshold and the second standard threshold respectively; If the comprehensive misunderstanding analysis coefficient value in each area is greater than the second standard threshold, it is marked as a high-risk area and requires optimization of labels, feedback or guidance content; If the comprehensive misunderstanding analysis coefficient value in each area is greater than or equal to the first standard threshold and less than or equal to the second standard threshold, it is marked as a medium-risk area and supplementary explanation or voice guidance is introduced; If the comprehensive misunderstanding analysis coefficient value in each area is less than the first standard threshold, it is marked as a low-risk area and used as an example area.
2. The virtual reality interaction method for displaying classical paintings according to claim 1, characterized in that: Multi-dimensional high-definition image acquisition of target classical paintings includes: using a ring light source device to achieve uniform illumination, using structured light scanning or laser ranging to collect surface texture information of the painting; collecting shallow parallax image sequences from the front of the painting and multiple micro-angles such as top, bottom, left, right, front and back; fusing image data in different bands, including visible light, infrared and ultraviolet imaging, to obtain information on the underlying drawing and pigment layers.
3. The virtual reality interaction method for displaying classical paintings according to claim 2, characterized in that: Construct a three-dimensional spatial model based on compositional elements: Use image semantic segmentation algorithms to extract visual elements in the picture, including people, background, symbols, and light sources; establish spatial depth relationships based on perspective lines, main viewing angles, and symmetry axis compositional features; use structured light or multi-view stereo algorithms to generate a depth map, and construct a three-dimensional point cloud based on the depth map; generate a three-dimensional mesh model through point cloud reconstruction or mesh generation technology, and map the original high-definition image as a texture map to the model surface.
4. The virtual reality interaction method for displaying classical paintings according to claim 1, characterized in that: Real-time matching of user behavior data with contextual tags includes: Bind a unique ID and a set of semantic tags to each composition element; When the user's gaze or gesture falls within the range of a certain composition element, the semantic tag corresponding to the element is triggered; The matching mechanism is based on the behavior-semantic mapping rules, combining the user behavior type, action frequency, and viewing duration to perform similarity weighted calculation; If the behavior matching degree exceeds the set threshold, a trigger event is generated.
5. A virtual reality interactive system for displaying classical paintings, for implementing the virtual reality interactive method for displaying classical paintings as claimed in any one of claims 1 to 4, characterized in that: It includes image acquisition module, context information management module, behavior perception module, interactive feedback module and semantic optimization module; Image acquisition module: This module collects multi-dimensional high-definition images of target classical paintings and constructs a three-dimensional spatial model based on the compositional elements of the paintings; Contextual information management module: extracts cultural contextual information related to paintings and performs semantic classification and labeling based on element dimensions; Behavior perception module: obtains behavioral data related to the user's gaze point, movement posture, and dwell time, and combines it with contextual tags for real-time matching; Interactive feedback module: Generates interactive feedback information consistent with the current painting style based on user behavior data and matched cultural context labels; Semantic Optimization Module: Performs cluster analysis based on the interaction behavior data of several users, identifies typical interaction paths and areas with high incidence of misunderstandings, and dynamically adjusts context labels and response logic to achieve personalized recommendations.
Citation Information
Patent Citations
Virtual reality system and interaction method applied to art painting and calligraphy exhibition
CN118131911A
Immersive exhibition hall intelligent guide display method and system based on user behaviors
CN120182488A