A system and method for generating interactive digital media displays and immersive experiences.
By acquiring virtual-real fusion information, determining gaze deviation constraints and element rendering visual effects, seamless fusion and pose connection interpolation complementarity are performed to generate recommended visual positions and display delay tolerance. This solves the problems of virtual-real element misalignment and interaction accuracy, and improves the interactive efficiency and immersive experience quality of digital media interactive display.
Patent Information
- Application Number
- CN202511609615.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-05
- Publication Date
- 2026-03-06
- Estimated Expiration
- 2045-11-05
AI Technical Summary
In the current generation of interactive digital media displays and immersive experiences, the integration of virtual and real information is not systematically collected, resulting in spatial misalignment of virtual and real elements, inconsistent lighting, visual discontinuity of interactive elements, and the inability to meet user needs in terms of interaction accuracy and experience smoothness.
By acquiring virtual-real fusion information, determining gaze deviation constraints and element rendering visual effects, seamless fusion is performed to form a normal contribution ratio. Pose connection interpolation is obtained and complemented with gaze deviation constraints to generate recommended visual positions and display delay tolerance, and feedback response closed-loop adjustment is performed.
It enhances the naturalness and visual comfort of virtual-real integration, improves the visual coordination of interactive elements, improves the smoothness and real-time response of cross-scene experience, and enhances the integrity and stability of the immersive experience of virtual-real integrated display.
Smart Images

Figure CN121070506B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of virtual reality technology, and more specifically, to a system and method for generating interactive digital media displays and immersive experiences. Background Technology
[0002] Virtual reality (VR) refers to a virtual three-dimensional space constructed using computer technology that provides users with an immersive experience. It is the core technological foundation for interactive display and immersive experience generation in digital media. Through head-mounted displays, sensors, and other devices, it renders and outputs virtual scenes in real time, while simultaneously capturing user body movements, head rotations, and other behaviors to achieve real-time interaction between the user and the virtual environment. By constructing virtual scenes that closely resemble the physical laws of the real world, it allows users to experience a sense of immersion, feeling as if they are "in a virtual scene," through visual, auditory, and even tactile perception. Because VR needs to meet requirements such as high frame rates, low latency, and precise interaction, its implementation requires balancing image realism with device performance. Excessive latency or large interaction deviations can easily lead to user dizziness and a disconnected experience. Therefore, optimizing scene rendering efficiency and interaction accuracy is crucial to enhancing VR's ability as a carrier for digital media display and experience.
[0003] However, in existing digital media interactive displays and immersive experience generation, there is a lack of systematic collection of virtual-real fusion information for fitted scenes and the absence of gaze deviation constraints. The lack of quantitative basis for seamless integration of element rendering visual effects leads to disordered visual recommendation positions. During cross-scene interaction, pose transitions lack interpolation processing, and intent judgment does not incorporate pose and gaze constraints. Furthermore, feedback responses are not linked to visual recommendation positions and display latency tolerance, resulting in spatial misalignment of virtual and real elements, inconsistent lighting, visual discontinuities in interactive elements, and low core interaction reach. Consequently, the virtual-real synergy, interaction accuracy, and experience smoothness of digital media interactive displays fail to meet user needs. Therefore, how to implement feedback response closed-loop adjustment for the construction of virtual-real fusion constraints in digital media interactive displays and immersive experience generation scenarios to improve the interactive efficiency of digital media interactive displays is a challenge facing the industry. Summary of the Invention
[0004] This application provides a system and method for generating interactive digital media displays and immersive experiences. It can perform feedback response closed-loop adjustment of the construction of virtual-real fusion constraints in the context of interactive digital media displays and immersive experiences, so as to improve the interactive efficiency of interactive digital media displays.
[0005] In a first aspect, this application provides a method for generating interactive digital media displays and immersive experiences, the method comprising the following steps:
[0006] Acquire virtual-real fusion information corresponding to each fitted scene during interactive digital media display, and determine the gaze deviation constraint when digital media is displayed in the scene separation plane based on all virtual-real fusion information.
[0007] Determine the element rendering visual effects in the interactive information during digital media interactive display, seamlessly integrate the element rendering visual effects to form the normal contribution ratio corresponding to the interactive display resources, and then determine the recommended visual position when the interactive information is generated based on the normal contribution ratio;
[0008] The pose connection interpolation value of the user when experiencing cross-scene interaction is obtained, and the pose connection interpolation value and the gaze deviation constraint are complemented by intent to obtain the intent adaptation logic of the interaction permission when rendering the user experience screen. Then, the display delay tolerance when the contextualized interaction response is generated is determined by the intent adaptation logic.
[0009] The system provides feedback and response to the virtual-real fusion display scene based on the recommended visual position and the display delay tolerance, and generates interactive display and immersive experience results of digital media.
[0010] In this embodiment, the virtual-real fusion information refers to a set of data describing the collaborative adaptation of real scenes and virtual elements in the dimensions of spatial location, lighting conditions, and occlusion relationships in interactive digital media displays.
[0011] In this embodiment, determining the gaze deviation constraint when digital media is displayed in the scene separation plane based on all virtual-real fusion information specifically includes:
[0012] Establish a coordinate mapping relationship between the virtual scene and the physical space;
[0013] Based on the virtual-real fusion information, determine the user viewpoint data when digital media is displayed in the scene separation plane;
[0014] The gaze deviation constraint when digital media is displayed in the scene separation plane is determined by the user viewpoint data and the coordinate mapping relationship.
[0015] In this embodiment, determining the rendering visual effects of elements in the interactive information during digital media interactive display specifically includes:
[0016] Real-time capture of user input data during the interaction process yields dynamic parameters of the interactive elements;
[0017] The dynamic parameters are matched with a predefined visual effects rule library to generate element-level rendering instructions;
[0018] The element-level rendering instructions determine the element rendering visual effects in the interactive information during digital media interactive display.
[0019] In this embodiment, the element rendering visual effect refers to the visual effect presented by the interactive element according to the rendering instructions, including information on materials, lighting, animation and multi-sensory feedback.
[0020] In this embodiment, the normal contribution ratio refers to the proportion of value of a single interactive display resource to the overall interactive experience.
[0021] In this embodiment, determining the recommended viewpoint for interactive information generation based on the normal contribution ratio specifically includes:
[0022] The visual saliency weights of each visual element are quantified and ranked according to the normal contribution ratio.
[0023] By integrating the aforementioned visual saliency weights with real-time user gaze focus data, attention distribution hotspots in the visual space are determined;
[0024] The recommended viewpoints for generating interactive information are determined by the attention distribution hotspots.
[0025] In this embodiment, the cross-scene interactive connection refers to the process of transitioning interactive behaviors when a user switches between different digital media scenarios.
[0026] In this embodiment, determining the display delay tolerance when generating a contextualized interactive response by the intent adaptation logic specifically includes:
[0027] The latency sensitivity coefficient for generating contextualized interactive responses is determined based on the intent adaptation logic.
[0028] Determine the latency tolerance baseline for contextualized interaction response generation in the current environment;
[0029] The display latency tolerance for generating contextualized interactive responses is determined based on the latency sensitivity coefficient and the latency tolerance baseline.
[0030] Secondly, this application provides a digital media interactive display and immersive experience generation system, used to execute a digital media interactive display and immersive experience generation method, the generation system comprising:
[0031] The information acquisition module is used to acquire virtual-real fusion information corresponding to each fitted scene during the interactive display of digital media, and to determine the gaze deviation constraint when the digital media is displayed in the scene separation plane based on all the virtual-real fusion information.
[0032] The seamless integration module is used to determine the element rendering visual effects in the interactive information during the interactive display of digital media, seamlessly integrate the element rendering visual effects to form the normal contribution ratio corresponding to the interactive display resources, and then determine the recommended visual position when the interactive information is generated based on the normal contribution ratio.
[0033] The intent complementation module is used to obtain the pose connection interpolation value when the user experiences cross-scene interaction connection, and to perform intent complementation between the pose connection interpolation value and the gaze deviation constraint to obtain the intent adaptation logic of the interaction permission when rendering the user experience screen. Then, the intent adaptation logic determines the display delay tolerance when the contextualized interaction response is generated.
[0034] The results generation module is used to provide feedback and response to the virtual-real fusion display scene based on the recommended visual position and the display delay tolerance, and to generate interactive display and immersive experience results of digital media.
[0035] The technical solutions provided by the embodiments disclosed in this application have the following beneficial effects:
[0036] The system acquires virtual-real fusion information corresponding to each fitted scene during interactive digital media display, and determines the gaze deviation constraint when digital media is displayed in the scene separation plane based on all virtual-real fusion information; it determines the element rendering visual effects in the interactive information during interactive digital media display, seamlessly integrates the element rendering visual effects to form the normal contribution ratio corresponding to the interactive display resources, and then determines the recommended visual position when generating interactive information based on the normal contribution ratio; it acquires the pose connection interpolation when the user experiences cross-scene interaction connection, performs intention complementarity between the pose connection interpolation and the gaze deviation constraint, obtains the intention adaptation logic of the interaction permission when rendering the user experience screen, and then determines the display delay tolerance when generating contextualized interactive response based on the intention adaptation logic; it provides feedback response to the virtual-real fusion display scene based on the recommended visual position and the display delay tolerance, and generates interactive display and immersive experience results of digital media.
[0037] Therefore, this application demonstrates that by acquiring virtual-real fusion information for each fitted scene and determining gaze deviation constraints, it can solve the problems of spatial misalignment of virtual and real elements, lighting incoordination, and user visual fatigue in existing technologies, thereby improving the naturalness and visual comfort of virtual-real fusion. By determining the rendering visual effects of elements and seamlessly integrating them to form a normal contribution ratio to determine the recommended visual position, it can compensate for the defects of disjointed element visual effects and low core interaction reach, enhancing the visual coordination and efficiency of interactive elements. By acquiring pose connection interpolation and forming intention complementarity with gaze deviation constraints, it can obtain intention adaptation logic and display delay tolerance, thereby improving the dizziness and uncontrolled response delay caused by cross-scene pose abrupt changes, and enhancing the smoothness and real-time response of cross-scene experience. By providing feedback and generating results based on the recommended visual position and display delay tolerance, it can solve the problems of disordered feedback and disjointed immersive perception, improve the experience integrity and stability of virtual-real fusion display, and comprehensively optimize the interactive display and immersive experience effects of digital media.
[0038] In summary, the technical solution adopted in this application can provide feedback response closed-loop adjustment for the construction of virtual-real fusion constraints in the context of digital media interactive display and immersive experience generation, thereby improving the interactive efficiency of digital media interactive display. Attached Figure Description
[0039] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only for this embodiment of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0040] Figure 1 This is an exemplary flowchart of a method for generating interactive digital media displays and immersive experiences according to this application;
[0041] Figure 2 This is a flowchart illustrating the determination of the normal contribution ratio provided in this application;
[0042] Figure 3 This is a flowchart illustrating the intent-based adaptation logic provided in this application;
[0043] Figure 4 This is a module structure diagram of a digital media interactive display and immersive experience generation system provided in this application. Detailed Implementation
[0044] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0045] This application provides a system and method for generating interactive digital media displays and immersive experiences. The core of this system involves acquiring virtual-real fusion information corresponding to various fitted scenes during interactive digital media displays; determining gaze deviation constraints when digital media is displayed in scene separation planes based on all virtual-real fusion information; determining the element rendering visual effects in the interactive information during interactive digital media displays; seamlessly merging the element rendering visual effects to form a normal contribution ratio corresponding to interactive display resources; and then determining the recommended visual position when generating interactive information based on the normal contribution ratio; acquiring the pose connection interpolation when the user experiences cross-scene interactive connections; performing intention complementarity between the pose connection interpolation and the gaze deviation constraints to obtain the intention adaptation logic of interactive permissions when rendering user experience screens; and then determining the display delay tolerance when generating contextualized interactive responses based on the intention adaptation logic; and providing feedback responses to the virtual-real fusion display scene based on the recommended visual position and the display delay tolerance, thereby generating interactive digital media displays and immersive experience results.
[0046] Example 1: To better understand the above technical solution, the following will provide a detailed description of the technical solution in conjunction with the accompanying drawings and specific implementation methods. (Refer to...) Figure 1 As shown, this figure is an exemplary flowchart of a digital media interactive display and immersive experience generation method according to this embodiment of the application. The generation method includes the following steps:
[0047] In step S1, the virtual-real fusion information corresponding to each fitted scene during the interactive display of digital media is obtained, and the gaze deviation constraint when the digital media is displayed in the scene separation plane is determined based on all the virtual-real fusion information.
[0048] In specific implementation, the virtual-real fusion information corresponding to each fitted scene during interactive digital media display can be obtained in the following way: First, a depth camera (such as Kinect) and a high-definition camera are used to capture 360° images of the real area of the fitted scene, collecting data on spatial structure (such as wall height and table size), object distribution (such as furniture position), and lighting conditions (such as light source direction and brightness). Then, through synchronous positioning and mapping technology, feature points of the real scene (such as wall corners and table corners) are extracted to construct a 3D map. The 3D coordinates of virtual elements are aligned with the coordinates of the 3D map of the real scene to determine spatial coordination data. Subsequently, a light sensor is used to collect the light source parameters of the real scene, and the shadow angle and brightness adjustment coefficient of the virtual elements are calculated by combining the lighting estimation algorithm to obtain lighting coordination data. Finally, the depth of real objects is detected by a depth sensor, and the order of virtual-real occlusion is determined by collision detection technology to obtain occlusion coordination data. This occlusion coordination data is used as the virtual-real fusion information corresponding to each fitted scene during interactive digital media display.
[0049] It should be noted that, in this application, virtual-real fusion information refers to a data set that describes the collaborative adaptation of real scenes and virtual elements in the dimensions of spatial location, lighting conditions, and occlusion relationships in interactive digital media displays.
[0050] In this embodiment, determining the gaze deviation constraint when digital media is displayed in the scene separation plane based on all virtual-real fusion information can be achieved through the following steps:
[0051] Establish a coordinate mapping relationship between the virtual scene and the physical space;
[0052] Based on the virtual-real fusion information, determine the user viewpoint data when digital media is displayed in the scene separation plane;
[0053] The gaze deviation constraint when digital media is displayed in the scene separation plane is determined by the user viewpoint data and the coordinate mapping relationship.
[0054] In practice, firstly, a 3D point cloud map is constructed by scanning the physical space using a depth camera and LiDAR, establishing a physical coordinate system with a fixed metal marker as the origin. A matching virtual scene is then built in Unity, establishing a virtual coordinate system with the corresponding virtual marker as the origin. AprilTag QR codes are placed, and the coordinates of the physical markers are calculated using a high-definition camera and intrinsic parameters. The least squares method is used to calculate the transformation matrix between the two coordinate systems, completing the mapping. The mapping result is used as the coordinate mapping relationship between the virtual scene and the physical space. Next, a fitted scene is built based on the virtual-real fusion information, digital media is projected, and lighting is matched. An eye-tracking system is deployed, and after user calibration, the user watches for 5 minutes, during which the system collects pupil coordinates and eye angles. The 3D coordinates of the gaze point are calculated using the coordinate mapping, and outlier data is removed using the 3σ criterion to obtain the user's viewpoint data when the digital media is displayed in the scene separation plane. Finally, convert the gaze point coordinates to a virtual coordinate system, using the center of the digital media as the reference point; calculate the spatial distance and angular offset, and take the 95th percentile as the initial value based on expert experience or experiments; adjust with reference to comfort data (if the angle exceeds 5°, adjust to 5°), and determine the gaze deviation constraints when the digital media is displayed in the scene separation plane in combination with the display requirements.
[0055] It should be noted that, in this application, coordinate mapping relationship refers to the correspondence rule between the virtual scene coordinate system and the physical space coordinate system; user viewpoint data refers to the dataset that records the coordinates, angle and duration of the user's gaze point when gazing at digital media; gaze deviation constraint refers to the maximum allowable offset range between the gaze point and the digital media reference point; scene separation surface refers to the physical or visual boundary that divides the real scene and the virtual scene in the interactive display of digital media.
[0056] In step S2, the element rendering visual effects in the interactive information during digital media interactive display are determined, and the element rendering visual effects are seamlessly integrated to form the normal contribution ratio corresponding to the interactive display resources. Then, the recommended visual position when the interactive information is generated is determined by the normal contribution ratio.
[0057] In this embodiment, determining the rendering visual effects of elements in the interactive information during digital media interactive display can be achieved through the following steps:
[0058] Real-time capture of user input data during the interaction process yields dynamic parameters of the interactive elements;
[0059] The dynamic parameters are matched with a predefined visual effects rule library to generate element-level rendering instructions;
[0060] The element-level rendering instructions determine the element rendering visual effects in the interactive information during digital media interactive display.
[0061] In practical implementation, firstly, devices can be deployed according to the interaction scenario: mobile devices use a 120Hz capacitive touch sensor to capture touch position, pressure, and duration; VR scenarios use a six-DOF controller to collect spatial position, rotation angle, and button signals; desktop devices combine a mouse sensor and keyboard module to capture click position and swipe speed. Kalman filtering is used to remove data noise, and pressure, speed, and other data are normalized to the 0-1 range. Parameters are mapped according to element type: for button types, the offset is calculated by comparing the touch position with the initial position, and the pressure value is mapped to a transparency coefficient; for slider types, the swipe speed is mapped to the color gradient speed, obtaining the dynamic parameters of the interactive elements. Then, a rule base can be built according to element type and interactive behavior classification: for button click behavior, pressure 0.3-0.6 and offset ≤5 pixels correspond to matte material, shadow +20%, and 0.2-second scaling animation; pressure >0.6 corresponds to metallic material, shadow +30%, and vibration animation. The rules are stored in JSON. A hash table is used to index and locate the element and behavior subclass, and a binary search is used to match the parameter range, finding the visual effect configuration within 10ms. The configuration is converted into the rendering engine format: Unity engine instructions, including element IDs, material IDs, lighting parameters, and animation frame sequences, are sorted by material replacement > animation playback > vibration feedback to generate element-level rendering instructions. Finally, the parsing module converts the instructions into engine parameters: In Unreal Engine 5, material instructions extract metallicity and reflectivity and assign them to material instances; lighting instructions adjust light source intensity and direction; animation instructions determine the frame sequence, playback speed, and loop count. The real-time rendering pipeline is invoked: material replacement uses instance dynamic update functionality, lighting adjustments use linear interpolation with a 0.1-second transition to avoid abrupt changes, animations are played through the Blueprint system, and the device vibration module is synchronously invoked to trigger vibration at keyframes. The frame buffer is used to detect the rendered image, compare it with the visual effects configuration, and resend instructions if there are discrepancies to determine the rendering visual effects of elements in the interactive information during digital media interactive displays.
[0062] It should be noted that, in this application, the dynamic parameters of interactive elements refer to the quantitative data describing the real-time state of interactive elements obtained after processing user input data; the predefined visual effect rule library refers to the database that stores the mapping relationship between "dynamic parameters and visual effect configuration"; element-level rendering instructions refer to the set of instructions that guide the rendering engine to execute visual effects; user input data refers to the raw data of operation generated when the user interacts with digital media; element rendering visual effects refer to the visual effects presented by interactive elements according to rendering instructions, including information on materials, lighting, animation, and multi-sensory feedback.
[0063] Preferably, in this embodiment, the rendering visual effects of the elements are seamlessly integrated to form a normal contribution ratio corresponding to interactive display resources, with reference to... Figure 2 As shown in the figure, this is a flowchart illustrating the determination of the normal contribution ratio in some embodiments of this application. In this embodiment, the determination of the normal contribution ratio can be achieved by the following steps:
[0064] In step S21, the rendering resource consumption of each visual element is analyzed based on the rendering visual effects of the elements.
[0065] In step S22, the interactive attribute characteristics of visual elements during hybrid rendering are determined based on the total rendering resource consumption.
[0066] In step S23, the visual representation description corresponding to the interactive display resource is determined based on the interactive attribute features;
[0067] In step S24, the normal contribution ratio corresponding to the interactive display resource is determined based on the visual representation description.
[0068] In practice, the first step is to invoke the resource monitoring module of the digital media rendering engine (such as Unity or Unreal Engine) to enable real-time data collection. For the rendering effect of each visual element, three core metrics are collected: first, computational resource consumption, such as the number of polygons drawn per frame and the shader calculation time (in milliseconds); second, storage resource consumption, such as the memory size occupied by the element's texture map (in MB) and the storage capacity of the material file; and third, bandwidth resource consumption (if it's for cross-device display), such as the transmission rate of the element's rendering data (in Mbps). Following the "single element - single effect" correspondence, the collected raw data is converted according to a unified standard (e.g., converting shader time to "average time per frame" and texture memory to "peak usage"), ultimately forming a statistical table of rendering resource consumption for each visual element, thus obtaining the rendering resource consumption of each visual element. Next, the visual elements are categorized, distinguishing between interactive elements (such as buttons and virtual props) and decorative elements (such as background textures and static lighting). For interactive elements, their rendering resource consumption is analyzed in conjunction with interaction behavior data: if an element has high rendering resource consumption (e.g., shader time > 5ms) and user click frequency > 10 times / minute, its interaction attribute characteristics are defined as "high resource consumption - high frequency interaction - core function type"; if an element has low resource consumption (e.g., texture memory < 2MB) and is only used to guide the eye (no direct click operation), it is defined as "low resource consumption - low frequency interaction - auxiliary guidance type". For decorative elements, characteristics are defined only based on resource consumption (e.g., "resource consumption - no interaction - atmosphere creation"), ultimately forming the interaction attribute characteristics of visual elements in hybrid rendering. Then, a mapping rule base of "interactive attribute features - visual representation" is established: if the feature is "high resource consumption - high frequency interaction - core function type", the corresponding visual representation description is "using high saturation colors + real-time dynamic lighting and shadows, triggering obvious scaling animation (scaling ratio 1.1x) during interaction, and maintaining the top layer of the screen hierarchy"; if the feature is "low resource consumption - low frequency interaction - auxiliary guidance type", the description is "using low saturation colors (contrast with background ≥ 2:1), no complex animation, only slight brightening when hovering, and the screen hierarchy is below the core elements". Combining the display scene style of digital media (such as the lifestyle style of AR shopping, the industrial style of VR training), the descriptive language is adjusted according to the scene (such as "soft gloss" for lifestyle scenes, and "metallic texture" for industrial style), finally forming the visual representation description of each interactive display resource. Finally, a mapping relationship was established between visual representation description and contribution dimensions: "high saturation + high frequency interaction" in the description corresponds to "interaction frequency contribution" (weight 0.3), "core functional type + upper layer of screen" corresponds to "functional core degree contribution" (weight 0.4), and "real-time dynamic light and shadow" corresponds to "visual attractiveness contribution" (weight 0.3).For each interactive display resource, scores are assigned across three dimensions based on its description (out of 100). For example, a resource described as "high saturation + high-frequency interaction + core functionality" would receive 90 points for interaction frequency, 100 points for core functionality, and 80 points for visual appeal. The resource's overall score is then calculated using weighted averages (90 × 0.3 + 100 × 0.4 + 80 × 0.3 = 91 points). This overall score is then divided by the sum of the overall scores of all interactive resources to obtain the resource's normal contribution ratio (e.g., if the total score is 455, the contribution ratio is 91 ÷ 455 = 20%).
[0069] It should be noted that, in this application, rendering resource consumption refers to a set of numerical indicators that quantify the computational and storage resources occupied by each visual element during the rendering process; interactive attribute characteristics refer to a set of characteristics of the relevance, functionality, and response efficiency of visual elements in a mixed rendering scene with the user; visual representation description refers to transforming abstract interactive attribute characteristics into concrete and practical visual presentation language; and normal contribution ratio refers to measuring the proportion of value of a single interactive display resource to the overall interactive experience.
[0070] In this embodiment, determining the recommended viewpoint for interactive information generation based on the normal contribution ratio can be achieved through the following steps:
[0071] The visual saliency weights of each visual element are quantified and ranked according to the normal contribution ratio.
[0072] By integrating the aforementioned visual saliency weights with real-time user gaze focus data, attention distribution hotspots in the visual space are determined;
[0073] The recommended viewpoints for generating interactive information are determined by the attention distribution hotspots.
[0074] In practice, the process begins by collecting the normal contribution ratio data of all interactive visual elements and sorting them from highest to lowest value. Linear normalization is then applied: the highest contribution ratio is set to 1.0, and the relative values of other elements are calculated as "their own contribution ratio ÷ the highest contribution ratio," serving as the initial visual saliency weight. Decorative elements (without a normal contribution ratio) are assigned a fixed weight of 0.1-0.3 based on their rendering resource consumption (e.g., less than 10% of total resources). This forms the final visual saliency weight for each visual element. Next, an eye-tracking device (such as the Tobii eye tracker) is deployed to collect the user's gaze points in the 3D coordinates of the visual space in real time, with a sampling frequency of 60Hz, recording continuously for 5 minutes. The visual space is divided into 10cm × 10cm grid units, and the number (frequency) of gaze points within each grid is counted. The visual saliency weight is then fused with the grid gaze frequency: the heat value of each grid = (total weights of elements within the grid × 0.6) + (gaze frequency × 0.4). A heatmap generation algorithm (such as Gaussian blur interpolation) is used to visualize heat values. Grid areas with heat values > 80% are identified as attention distribution hotspots in the visual space. Finally, the continuous grid areas with the highest heat values are selected from these attention distribution hotspots, prioritizing the center of each hotspot (the grid with the highest heat value). Based on the size of the interactive information (e.g., a button diameter of 5cm), an adaptation area is defined at the center, ensuring no obstructions (avoiding other high-weight elements) and maintaining a safe distance (≥ 10cm) from scene separation surfaces. Multiple hotspots are sorted by heat value, and the top three areas are selected as candidate recommendation positions. Finally, the area with the highest matching degree to the current interactive scene (e.g., in a shopping scene, the area closest to the product in the hotspot) is selected as the visual recommendation position.
[0075] It should be noted that, in this application, visual saliency weight refers to a quantitative indicator that measures the ability of a visual element to attract the user's attention in the visual space; real-time user gaze focus data refers to dynamic data collected in real time by an eye-tracking device, recording the three-dimensional coordinates and dwell time of the user's gaze point in the visual space; attention distribution hot zone refers to the visual distribution of the degree of user attention concentration in the visual space; and visual recommendation position refers to the information delivery position in the visual space that is suitable for placing interactive information.
[0076] In step S3, the pose connection interpolation value of the user when experiencing cross-scene interaction is obtained, and the pose connection interpolation value and the gaze deviation constraint are complemented by intent to obtain the intent adaptation logic of the interaction permission when rendering the user experience screen. Then, the display delay tolerance when the contextualized interaction response is generated is determined by the intent adaptation logic.
[0077] In practice, obtaining the pose transition interpolation value for users experiencing cross-scene interaction can be achieved as follows: First, inertial measurement unit (IMU) sensors are worn on the user's head, hands, and waist, with a sampling frequency of 100Hz, to collect position coordinates, angles, and acceleration data in real time. When the user switches between scenes, the initial and target pose data for 1 second before and after the switch are extracted. A cubic spline interpolation algorithm is used to generate 20 transition data points between the two, with an interval of 0.05 seconds, so that the changes in position and angle are distributed according to a smooth curve, forming the pose transition interpolation value for the user experiencing cross-scene interaction.
[0078] It should be noted that, in this application, "experience cross-scene interaction connection" refers to the process of transitioning interactive behaviors when a user switches between different digital media scenarios; "pose connection interpolation" refers to the parameters for the smooth transition of body posture when a user crosses scenarios.
[0079] Preferably, in this embodiment, the pose interpolation and the gaze deviation constraint are complemented by intent to obtain the intent adaptation logic of interaction permissions when rendering the user experience screen, referring to... Figure 3 As shown in the figure, this is a flowchart illustrating the intent adaptation logic in some embodiments of this application. In this embodiment, the intent adaptation logic can be implemented using the following steps:
[0080] In step S31, the pose connection interpolation is mapped to the gaze deviation constraint to obtain the visual coordination sequence;
[0081] In step S32, the intent compensation range of the interaction permission when rendering the user experience screen is determined according to the visual coordination sequence.
[0082] In step S33, a dynamic allocation strategy for interaction permissions in the user experience screen is determined based on the intent compensation range;
[0083] In step S34, the intent adaptation logic of the interaction permissions when rendering the user experience screen is determined according to the dynamic allocation strategy.
[0084] In practice, the process begins by establishing a coordinate mapping between the virtual scene and physical space, converting the position coordinates and angle data (physical space) from pose transition interpolation into virtual scene coordinate system data. Transitional data (20 data points, 0.05-second intervals) from pose transition interpolation is extracted frame-by-frame. Combined with the gaze coordinates of the corresponding frames collected in real-time by the eye-tracking device, it is determined whether the gaze point in each frame meets the gaze deviation constraints (e.g., spatial distance offset ≤ 3cm, angle offset ≤ 5°). The "pose parameters – gaze coordinates – compliance" data are arranged chronologically to generate a visual alignment sequence containing 20 sets of data. Next, abnormal frames in the visual alignment sequence are analyzed: the start and end times of consecutive abnormal frames are statistically analyzed (e.g., frames 5-8, corresponding to 0.2-0.35 seconds), determining the time compensation interval; the maximum offset between the gaze point and the constraint range in the abnormal frame is calculated (e.g., spatial offset 4cm, angle offset 6°), and combined with the pose transition direction (e.g., head tilt), the spatial compensation interval is determined (e.g., extending the allowable gaze range by 2cm in the user's tilt direction). The time interval (0.2–0.35 seconds) and spatial interval (extended by 2 cm in the forward tilt direction) are integrated, and the corresponding pose parameters within this interval are labeled to form the intent compensation interval for interactive permissions when rendering the user experience screen. Then, the intent compensation interval is graded according to "compensation intensity": if the gaze offset within the interval is ≤5cm and the duration is ≤0.5 seconds, it is defined as "mild compensation"; if the offset is >5cm or the duration is >0.5 seconds, it is defined as "severe compensation". Within the mild compensation interval, core interactive permissions (such as click and zoom) are prioritized, while secondary permissions (such as rotation and deletion) are restricted. Within the severe compensation interval, only basic guidance permissions (such as highlighting interactive areas) are enabled, core permissions are suspended, and pose guidance is triggered (such as displaying a "adjust head angle" prompt on the screen). Within the non-compensation interval, all interactive permissions are enabled as usual, forming a dynamic allocation strategy for interactive permissions in the user experience screen. Finally, the dynamic allocation strategy is broken down into conditional logic recognizable by the rendering engine: When the system detects that the current scene is in the "light compensation zone" (determined by time and space parameters), logic 1 is triggered—calling the permission rendering module to highlight core interactive elements (such as glowing button borders) and graying out secondary elements; when in the "heavy compensation zone," logic 2 is triggered—rendering guidance text and icons, calling the permission blocking interface, and disabling core interactive responses; when in the "non-compensation zone," logic 3 is triggered—restoring normal rendering of all elements and opening all interactive response interfaces. This logic code is embedded into the rendering engine's frame update process (such as Unity's Update function) to ensure that the corresponding logic is judged and executed in real time each frame, forming the intent adaptation logic for interactive permissions when rendering the user experience screen.
[0085] It should be noted that, in this application, the visual alignment sequence refers to the time-ordered "pose-gaze" matching data chain formed by associating user pose transition data with gaze range constraints; the intent compensation interval refers to the time and space range within which intent judgment deviations need to be compensated when the gaze point deviates from the constraints in the visual alignment sequence; the dynamic allocation strategy refers to a rule system formulated based on the characteristics of the intent compensation interval to adjust the openness of interaction permissions according to scene changes; interaction permissions refer to the scope of permissions that allow users to perform specific operations, obtain content, or adjust parameters on the rendered experience screen in cross-scene interactions; and the intent adaptation logic refers to the executable logic that transforms the dynamic allocation strategy into an executable logic that guides the rendering engine to adjust the presentation and response of interaction permissions.
[0086] In this embodiment, determining the display delay tolerance when generating a contextualized interactive response using the intent adaptation logic can be achieved through the following steps:
[0087] The latency sensitivity coefficient for generating contextualized interactive responses is determined based on the intent adaptation logic.
[0088] Determine the latency tolerance baseline for contextualized interaction response generation in the current environment;
[0089] The display latency tolerance for generating contextualized interactive responses is determined based on the latency sensitivity coefficient and the latency tolerance baseline.
[0090] In practical implementation, firstly, the intent adaptation logic is broken down into three core scenarios: the "non-compensated interval" corresponds to normal interactions (such as clicking the confirmation button), the "lightly compensated interval" corresponds to low-deviation interactions (such as scaling operations under slight offset), and the "heavily compensated interval" corresponds to high-deviation guidance (such as posture adjustment prompts). Basic sensitivity coefficients are set for these three scenarios: for normal interactions, users have a strong perception of latency, so the basic coefficient is set to 1.2; for low-deviation interactions, the perception is moderate, so the coefficient is set to 0.9; and for high-deviation guidance, the perception is weak, so the coefficient is set to 0.6. Combining the normal contribution ratio of interactive elements within the scene (e.g., if the core element's contribution ratio is 30%, the coefficient increases by 0.2), the latency sensitivity coefficient for generating contextualized interactive responses is finally calculated (e.g., the normal interaction coefficient for core elements = 1.2 + 0.2 = 1.4). Then, a latency testing environment is built to simulate the currently used hardware (such as AR glasses, PC host) and rendering engine (such as Unity) configuration. For the three scenarios in the intent adaptation logic, 100 response latency tests were conducted for each: In the "non-compensated interval," a button click was triggered, and the time from the operation to the screen feedback was recorded; in the "slightly compensated interval," a scaling operation was triggered, and the time from parameter change to screen rendering was recorded; in the "heavily compensated interval," a guided prompt was triggered, and the time from the instruction being issued to the prompt being displayed was recorded. The average value of the test data for each scenario was taken (e.g., the average latency for normal interaction is 15ms), and this average value is the latency tolerance baseline for contextualized interaction responses in the current environment. Finally, the calculation logic of "tolerance = baseline × sensitivity coefficient" can be used: if the latency tolerance baseline for the "non-compensated interval" is 15ms and the sensitivity coefficient is 1.4, then the tolerance = 15 × 1.4 = 21ms; for the "slightly compensated interval," the baseline is 20ms and the coefficient is 0.9, so the tolerance = 20 × 0.9 = 18ms; for the "heavily compensated interval," the baseline is 25ms and the coefficient is 0.6, so the tolerance = 25 × 0.6 = 15ms. Referring to industry visual comfort standards (such as delays exceeding 20ms that can easily cause dizziness), the calculation results were fine-tuned (such as reducing the normal interaction tolerance from 21ms to 20ms), and the final display delay tolerance for generating contextualized interactive responses was determined (such as 20ms for normal, 18ms for mild, and 15ms for severe).
[0091] It should be noted that in this application, the latency sensitivity coefficient refers to an indicator that quantifies the user's sensitivity to the perceived latency of the interaction response under different intent adaptation scenarios; the latency tolerance baseline refers to the baseline value of the interaction response latency that the system can stably support under the current hardware and software environment; and the display latency tolerance refers to the maximum latency value of the contextualized interaction response that the user can accept under the current environment and scenario sensitivity.
[0092] In step S4, the virtual-real fusion display scene is responded to based on the recommended visual position and the display delay tolerance, and the interactive display and immersive experience results of digital media are generated.
[0093] In specific implementation, the feedback response to the virtual-real fusion display scene based on the recommended visual position and the display latency tolerance, and the generation of interactive display and immersive experience results of digital media, can be achieved in the following way: First, user interaction signals (such as clicking on a recommended visual position element) are received in real time. The trigger position is first checked to see if it is within the recommended visual position; if so, the response process is initiated. During the response, the display latency tolerance data is called, and the corresponding tolerance (e.g., core interaction ≤ 20ms) is matched according to the interaction type (core / non-core). Resources are dynamically scheduled through the rendering engine: priority is given to rendering recommended visual position elements, and texture precision is reduced for non-critical elements to ensure that the response latency is within the tolerance. Feedback forms include simultaneous triggering of visual (element highlighting / animation), tactile (device vibration), and auditory (prompt sounds). Subsequently, the virtual-real fusion scene and interactive feedback mechanism are integrated to develop a runnable digital media application (such as AR shopping, VR education). After optimization through user testing (collecting interaction completion time, accidental touch rate, and subjective experience rating), interactive display content and an immersive experience result without latency or virtual-real disconnect are generated.
[0094] It should be noted that, in this application, the virtual-real fusion display scenario refers to a scenario that integrates real physical space with virtual digital elements for users to have an interactive experience.
[0095] Therefore, this application demonstrates that by acquiring virtual-real fusion information for each fitted scene and determining gaze deviation constraints, it can solve the problems of spatial misalignment of virtual and real elements, lighting incoordination, and user visual fatigue in existing technologies, thereby improving the naturalness and visual comfort of virtual-real fusion. By determining the rendering visual effects of elements and seamlessly integrating them to form a normal contribution ratio to determine the recommended visual position, it can compensate for the defects of disjointed element visual effects and low core interaction reach, enhancing the visual coordination and efficiency of interactive elements. By acquiring pose connection interpolation and forming intention complementarity with gaze deviation constraints, it can obtain intention adaptation logic and display delay tolerance, thereby improving the dizziness and uncontrolled response delay caused by cross-scene pose abrupt changes, and enhancing the smoothness and real-time response of cross-scene experience. By providing feedback and generating results based on the recommended visual position and display delay tolerance, it can solve the problems of disordered feedback and disjointed immersive perception, improve the experience integrity and stability of virtual-real fusion display, and comprehensively optimize the interactive display and immersive experience effects of digital media.
[0096] In summary, the technical solution adopted in this application can provide feedback response closed-loop adjustment for the construction of virtual-real fusion constraints in the context of digital media interactive display and immersive experience generation, thereby improving the interactive efficiency of digital media interactive display.
[0097] Example 2: This application provides a digital media interactive display and immersive experience generation system, referencing... Figure 4As shown, this figure is a modular structure diagram of a digital media interactive display and immersive experience generation system according to this embodiment of the present application. The generation system includes:
[0098] The information acquisition module 100 is used to acquire virtual-real fusion information corresponding to each fitted scene during the interactive display of digital media, and to determine the gaze deviation constraint when the digital media is displayed in the scene separation plane based on all the virtual-real fusion information.
[0099] The seamless integration module 200 is used to determine the element rendering visual effects in the interactive information during the interactive display of digital media, seamlessly integrate the element rendering visual effects to form the normal contribution ratio corresponding to the interactive display resources, and then determine the recommended visual position when the interactive information is generated based on the normal contribution ratio.
[0100] The intent complementation module 300 is used to obtain the pose connection interpolation value when the user experiences cross-scene interaction connection, and to perform intent complementation between the pose connection interpolation value and the gaze deviation constraint to obtain the intent adaptation logic of the interaction permission when rendering the user experience screen, and then the intent adaptation logic determines the display delay tolerance when the contextualized interaction response is generated.
[0101] The result generation module 400 is used to provide feedback and response to the virtual-real fusion display scene based on the recommended visual position and the display delay tolerance, and to generate interactive display and immersive experience results of digital media.
[0102] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0103] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, including read-only memory (ROM), random access memory (RAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), one-time programmable read-only memory (OTPROM), electrically-Erasable Programmable Read-Only Memory (EEPROM), compactdisc read-only memory (CD-ROM) or other optical disc storage, disk storage, magnetic tape storage, or any other computer-readable medium capable of carrying or storing data.
[0104] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
Claims
1. A method for digital media interactive presentation and immersive experience generation, characterized by, The generation method comprises the following steps: Obtain the virtual-real fusion information corresponding to each fitting scene in the digital media interactive display, and determine the gaze deviation constraint of the digital media in the scene separation surface according to all the virtual-real fusion information, wherein the virtual-real fusion information refers to a data set describing the spatial position, lighting condition and occlusion relationship dimension coordination adaptation of the real scene and the virtual element in the digital media interactive display; Determine the element rendering visual effect in the interactive information of the digital media interactive display, seamlessly integrate the element rendering visual effect, form the normal contribution ratio corresponding to the interactive display resource, and then determine the view recommendation position of the interactive information generation according to the normal contribution ratio, wherein the normal contribution ratio refers to the value proportion of a single interactive display resource to the overall interactive experience, and the view recommendation position refers to the information touch position suitable for placing the interactive information in the view space; Obtain the pose connection interpolation of the user in the experience of cross-scene interaction connection, complement the intention of the pose connection interpolation and the gaze deviation constraint, obtain the intention adaptation logic of the interactive authority in the rendering of the user experience picture, and then determine the display delay tolerance of the situational interaction response generation according to the intention adaptation logic; The intention adaptation logic of the interactive authority in the rendering of the user experience picture is obtained by complementing the intention of the pose connection interpolation and the gaze deviation constraint, and comprises: mapping the pose connection interpolation into the gaze deviation constraint to obtain a visual coordination sequence; determining the intention compensation interval of the interactive authority in the rendering of the user experience picture according to the visual coordination sequence; determining the dynamic allocation strategy of the interactive authority in the user experience picture according to the intention compensation interval; and determining the intention adaptation logic of the interactive authority in the rendering of the user experience picture according to the dynamic allocation strategy, wherein the visual coordination sequence refers to the matching data chain of the pose and the gaze point sorted by time after associating the user pose transition data with the gaze range constraint; the intention compensation interval refers to the time and space range that needs to compensate the intention judgment deviation for the case that the gaze point deviates from the constraint in the visual coordination sequence; the dynamic allocation strategy refers to the rule system for adjusting the opening degree of the interactive authority according to the characteristics of the intention compensation interval; the interactive authority refers to the authority range that allows the user to perform specific operations, obtain content or adjust parameters on the rendered experience picture in cross-scene interaction; and the intention adaptation logic refers to the executable logic for converting the dynamic allocation strategy into and guiding the rendering engine to adjust the interactive authority presentation and response; According to the view recommendation position and the display delay tolerance, the virtual-real fusion display scene is fed back and responded, and the interactive display and immersive experience result of the digital media is generated.
2. The method of claim 1, wherein, The gaze deviation constraint of the digital media in the scene separation surface is determined according to all the virtual-real fusion information, and comprises: Establishing a coordinate mapping relationship between the virtual scene and the physical space; Determining the user viewpoint data of the digital media in the scene separation surface according to the virtual-real fusion information; The gaze deviation constraint of the digital media when displayed in the scene separation surface is determined according to the user viewpoint data and the coordinate mapping relationship.
3. The method of claim 1, wherein, The element rendering visual effect in the interactive information when the digital media is interactively displayed specifically includes: Real-time capturing of user input data in the interactive process to obtain dynamic parameters of the interactive element; Matching the dynamic parameters with a predefined visual effect rule library to generate element-level rendering instructions; Determining the element rendering visual effect in the interactive information when the digital media is interactively displayed according to the element-level rendering instructions.
4. The method of claim 1, wherein, The element rendering visual effect refers to the visual effect of the interactive element presented according to the rendering instructions, including material, light and shadow, animation, and multi-sensory feedback information.
5. The method of claim 1, wherein, The visual scene recommendation position when the interactive information is generated is specifically determined according to the normal contribution ratio, which includes: Quantifying and sorting the visual saliency weight of each visual element according to the normal contribution ratio; Fusing the visual saliency weight and the real-time gaze focus data of the user to determine the attention distribution hot area in the visual scene space; Determining the visual scene recommendation position when the interactive information is generated through the attention distribution hot area.
6. The method of claim 1, wherein, The experience cross-scene interaction transition refers to the process of completing the transition of the interactive behavior when the user switches between different digital media scenes.
7. The method of claim 1, wherein, The display delay tolerance when the situational interactive response is generated is specifically determined according to the intention adaptation logic, which includes: Determining the delay sensitivity coefficient when the situational interactive response is generated according to the intention adaptation logic; Determining the delay tolerance baseline of the generation of the situational interactive response in the current environment; Determining the display delay tolerance when the situational interactive response is generated according to the delay sensitivity coefficient and the delay tolerance baseline.
8. A digital media interactive presentation and immersive experience generation system for performing a digital media interactive presentation and immersive experience generation method according to any one of claims 1 to 7, characterized by, The generation system includes: An information acquisition module for acquiring virtual-real fusion information corresponding to each fitting scene when the digital media is interactively displayed, and determining the gaze deviation constraint of the digital media when displayed in the scene separation surface according to all virtual-real fusion information; A seamless fusion module for determining the element rendering visual effect in the interactive information when the digital media is interactively displayed, seamlessly fusing the element rendering visual effect, forming a normal contribution ratio corresponding to the interactive display resource, and then determining the visual scene recommendation position when the interactive information is generated according to the normal contribution ratio; An intention complementary module for obtaining the pose transition interpolation when the user experiences the cross-scene interaction transition, complementing the intention of the pose transition interpolation and the gaze deviation constraint, obtaining the intention adaptation logic of the interactive authority when the user experience picture is rendered, and then determining the display delay tolerance when the situational interactive response is generated according to the intention adaptation logic; An achievement generation module for feeding back to the virtual-real fusion display scene according to the visual scene recommendation position and the display delay tolerance, and generating the interactive display and immersive experience achievement of the digital media.
Citation Information
Patent Citations
Attitude synchronization interaction exhibition hall display system based on XR technology
CN115202485A
Digital exhibition hall panoramic roaming construction method based on virtual-real fusion
CN120876791A