AI and image detection-based art design virtual reality content synchronous display system
By extracting artistic design features through multispectral image acquisition and dynamic semantic segmentation networks, combined with generative adversarial networks and distributed synchronization engines, high-precision restoration and dynamic interactive response of artistic entities are achieved, solving the problems of existing systems in detail restoration and device adaptation, and improving user experience.
Patent Information
- Application Number
- CN202510625724.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-15
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2045-05-15
AI Technical Summary
Existing digital display systems for art design have difficulty capturing multi-dimensional features such as surface reflection, material layers, and micro-topological structures of paintings. They also lack the ability to interact with user interactions, resulting in errors in the virtual model's restoration of authentic artistic styles. Furthermore, they are difficult to adapt to the display characteristics of different terminals, affecting user immersion and interactive experience.
A multispectral image acquisition device is combined with a dynamic semantic segmentation network to extract artistic design features, and a generative adversarial network is used to generate virtual reality content that matches the terminal display characteristics. User interaction operations are monitored in real time for incremental content rendering and multi-terminal synchronous push.
It achieves high-precision restoration of artistic entities and dynamic interactive response, solves the problems of existing systems in detail fidelity and device adaptation, ensures low-latency multi-terminal synchronous push, and enhances user immersion and interactive experience.
Smart Images

Figure CN120707782A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of virtual reality display technology, and in particular to an art design virtual reality content synchronous display system based on AI and image detection. Background Art
[0002] With the rapid development of virtual reality (VR) technology, the art and design field has gradually explored the migration of visual elements such as two-dimensional paintings, three-dimensional modeling, and material textures into virtual space for immersive display. However, existing digital display systems for art and design generally have the following technical bottlenecks:
[0003] First, traditional image acquisition methods primarily rely on RGB three-channel information, making it difficult to capture multi-dimensional features such as surface reflections, material layers, and microscopic topological structures. This results in significant errors in the virtual model's ability to reproduce authentic artistic styles and a lack of fidelity in detailing textures. Second, existing content generation mechanisms mostly utilize static texture mapping or pre-rendered models, lacking the ability to integrate with user interactions. This results in significant latency and rigid responses in terms of dynamic response, tactile feedback, and personalized adjustments.
[0004] In addition, in the content push scenario for multi-terminal VR devices, the current system generally adopts a unified rendering content broadcast or a simple resolution compression method, which is difficult to adapt to the display characteristics of different terminals (such as refresh rate, field of view, color gamut coverage), and easily causes image tearing, rendering distortion and synchronization lag, seriously affecting the user's immersion and interactive experience. Summary of the Invention
[0005] The present invention provides an art design virtual reality content synchronous display system based on AI and image detection, which can integrate high-dimensional image detection, AI feature modeling and multi-terminal adaptive art design virtual reality content display system to achieve true restoration of art entities, interactively driven content updates and low-latency multi-device synchronous push.
[0006] Art design virtual reality content synchronous display system based on AI and image detection, including:
[0007] A feature acquisition unit acquires surface texture data of an art design object through a multispectral image acquisition device and extracts at least three art design feature sets using a dynamic semantic segmentation network. The art design feature sets include brushstroke trajectory features, material reflectivity features, and spatial topological relationship features.
[0008] A parameter matching unit inputs the art design feature set into a pre-trained generative adversarial network model, synchronously obtains display characteristic parameters of the target VR terminal through device parameter analysis, and generates a virtual reality content frame sequence that matches the display characteristic parameters, wherein the virtual reality content frame sequence includes a dynamic light and shadow effect layer and an interactive response event marker;
[0009] Virtual reality generation unit: monitors the spatial posture data generated by user interaction operations in real time, triggers feature update instructions based on the interaction response event marker, feeds back the updated art design feature set to the generative adversarial network model for incremental content rendering, and pushes the adapted virtual reality content to multiple terminals through a distributed synchronization engine.
[0010] Optionally, in the feature acquisition unit, based on a multispectral image acquisition device, a preset wavelength range is used to scan the surface of the art design entity object three times in an orthogonal polarization mode to obtain visible light reflection spectrum, near-infrared absorption spectrum and short-wave infrared scattering spectrum respectively.
[0011] Optionally, the feature acquisition unit also includes constructing a two-stream dynamic semantic segmentation network, wherein the first sub-network uses a dilated convolution kernel to extract the spatiotemporal continuity features of the brush stroke trajectory, and tracks the brush stroke pressure change curve through a long short-term memory module; the second sub-network uses a spectral feature fusion layer to weightedly splice the channel dimensions of the three groups of atlases, and outputs a pixel-level feature map including the material reflectance gradient value.
[0012] Optionally, a three-dimensional graph attention mechanism is deployed at the end of the dual-stream dynamic semantic segmentation network to generate a non-Euclidean feature vector representing the spatial topological relationship by establishing a curvature change correlation matrix of adjacent pixels, and finally the spatiotemporal continuity features, pixel-level feature maps and non-Euclidean feature vectors are combined into the artistic design feature set.
[0013] Optionally, the generative adversarial network model in the parameter matching unit adopts a multimodal conditional generative adversarial network, including a generator G and a discriminator D.
[0014] Optionally, the input end of the generator G is connected to three parallel channels of the art design feature set:
[0015] The first channel maps the stroke trajectory features to latent space vectors and generates the basic geometric grid through the spatiotemporal encoder;
[0016] The second channel inputs the material reflectivity characteristics into the physical rendering pipeline and combines it with the Monte Carlo ray tracing algorithm to generate a dynamic light and shadow effect layer;
[0017] The third channel receives the display characteristic parameter set output by device parameter analysis, including the screen refresh rate, color gamut coverage, and field of view curvature radius of the target VR terminal, and generates device-related content rendering constraints through the parameter adaptation layer.
[0018] Optionally, a device perception verification mechanism is deployed in the discriminator D, specifically including:
[0019] Color gamut matching verification: Calculate the ΔE2000 color difference between the generated content and the P3 color gamut of the target device;
[0020] Motion blur prediction: Predicts the image smear index based on screen refresh rate and eye tracking data;
[0021] Geometric distortion correction: Dynamically adjust the vertex shader's deformation compensation parameters based on the field of view curvature radius.
[0022] Optionally, the method further includes embedding an interactive response event marker in the content frame sequence through an event marker generator, wherein the event marker includes:
[0023] Tactile feedback trigger area: the pressure response level is divided according to the material reflectivity gradient value;
[0024] Light-shadow interaction sensitive area: Marks the dynamic shadow update priority based on ray tracing results;
[0025] Device adaptation identifier: records the hash value of the rendering constraints for verification by the distributed synchronization engine.
[0026] Optionally, the virtual reality generation unit specifically includes:
[0027] a. Acquire six-degree-of-freedom pose data by fusing a nine-axis inertial measurement unit with an infrared optical tracking system, use a filter to eliminate jitter noise, and generate a purified pose stream including a timestamp;
[0028] b. Build an event marker driven engine to trigger feature update operations based on user interaction trigger areas:
[0029] c uses differential encoding technology in the incremental rendering stage:
[0030] The feature hash comparison algorithm is used to identify the areas that need to be updated, and only the feature subset with hash value differences greater than 15% is regenerated;
[0031] Use a progressive generation strategy: generate a low-resolution base mesh in the first frame, and gradually improve it to the target resolution through residual connections in subsequent frames;
[0032] d When the distributed synchronization engine performs multi-terminal push:
[0033] Attach a device fingerprint identifier to each content frame, including the screen color temperature calibration value and rendering pipeline delay compensation coefficient;
[0034] Adopting a layered synchronization protocol: the base layer transmits geometric topology data, and the enhanced layer transmits light and shadow effects data in a streaming manner;
[0035] When version differences between terminals are detected, a content merging algorithm based on the minimum spanning tree is triggered.
[0036] Optionally, the event marker driving engine execution logic is:
[0037] Parse the haptic feedback trigger area boundary coordinates in the interaction response event marker;
[0038] When the operation trajectory in the purified pose stream intersects with the trigger area:
[0039] Calculate the feature set update amount ΔF according to the pressure-displacement curve;
[0040] Generate an update instruction packet with a priority tag, where the priority is determined by the heat value Hv of the interaction area.
[0041] Beneficial effects of the present invention:
[0042] This invention combines multispectral image acquisition with a dual-stream dynamic semantic segmentation network to heterogeneously decouple and jointly model brushstroke trajectories, material reflectivity, and spatial topological relationships. In particular, it significantly improves the signal-to-noise ratio of feature extraction with the help of orthogonal polarization imaging under conditions of high reflective interference. The non-Euclidean space vectors generated by the graph attention mechanism can fully restore the three-dimensional configuration of complex artistic textures, solving the problem of traditional RGB image extraction lacking material physical semantics, and providing a more accurate digital restoration foundation for virtual artwork modeling.
[0043] The present invention constructs a multimodal conditional generative adversarial network, deeply integrates the artistic feature set with the VR terminal display parameters, introduces three types of discrimination mechanisms: color gamut matching verification, motion blur prediction, and geometric distortion correction, and realizes fine-grained adaptive adjustment of content frames. At the same time, with the help of a tactile response event-driven engine and an LSTM structure, the feature update amount of the trigger pressure change in the user's real-time interaction trajectory is estimated, ensuring that the incremental content rendering is dynamically consistent with the user's operation, and solving the problems of large response delay and poor device adaptation in the existing system.
[0044] This invention proposes a differential coding mechanism based on hash comparison and residual connection, combined with a progressive resolution improvement strategy, to perform target-level rendering only on sub-regions where significant changes have occurred in the content, thereby reducing bandwidth and computing power consumption; and combines device fingerprint identification with a minimum spanning tree content merging algorithm to construct a multi-terminal push system that takes into account both low latency and consistency, effectively solving problems such as difficult synchronization and slow merging in multi-device heterogeneous rendering scenarios, and ensuring the comprehensive optimization of virtual reality experience in terms of performance, quality and latency. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only for the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0046] Figure 1 Schematic diagram of system functional units according to an embodiment of the present invention;
[0047] Figure 2 Schematic diagram of the device perception verification mechanism according to an embodiment of the present invention. DETAILED DESCRIPTION
[0048] The present invention is described in detail below with reference to the accompanying drawings and specific embodiments. It is also noted that, to provide a more detailed description, the following embodiments are best and preferred embodiments, and those skilled in the art may employ alternative methods for implementing certain known technologies. Furthermore, the accompanying drawings are intended only to provide a more detailed description of the embodiments and are not intended to limit the present invention.
[0049] like Figure 1-Figure 2 As shown, the art design virtual reality content synchronous display system based on AI and image detection includes:
[0050] Feature acquisition unit: acquires surface texture data of an art design object through a multispectral image acquisition device, and uses a dynamic semantic segmentation network to extract at least three art design feature sets, including brushstroke trajectory features, material reflectivity features, and spatial topological relationship features;
[0051] Parameter matching unit: This unit inputs the art design feature set into a pre-trained generative adversarial network model, synchronously obtains the display characteristic parameters of the target VR terminal through device parameter analysis, and generates a virtual reality content frame sequence that matches the display characteristic parameters. The virtual reality content frame sequence includes a dynamic light and shadow effect layer and interactive response event markers.
[0052] Virtual reality generation unit: monitors the spatial posture data generated by user interaction operations in real time, triggers feature update instructions based on interaction response event markers, feeds the updated art design feature set back to the generative adversarial network model for incremental content rendering, and pushes adapted virtual reality content to multiple terminals through a distributed synchronization engine.
[0053] The feature acquisition unit specifically includes:
[0054] Multispectral image collaborative acquisition subunit: Using a multispectral imaging module with a wavelength range of 380-2500nm, the surface of the art design object is scanned three times in orthogonal polarization mode to obtain the following atlas data:
[0055] Visible light reflectance spectrum: λ1∈[400nm,700nm];
[0056] Near-infrared absorption spectrum: λ2∈[701nm,1100nm];
[0057] Short-wave infrared scattering spectrum: λ3∈[1101nm,2500nm];
[0058] Among them, orthogonal polarization is used to suppress mirror reflection interference and improve the signal-to-noise ratio by 60dB; the spectral resolution accuracy is controlled at ±5nm, which is used to analyze sub-surface material structure.
[0059] Semantic segmentation network construction subunit: Construct a two-stream neural network structure consisting of two sub-networks:
[0060] The first sub-network is used to extract the spatiotemporal continuity features of the brush stroke trajectory. The network structure is configured as follows:
[0061] The convolution module uses a dilated convolution kernel, and the expansion coefficient is set to:
[0062] d∈{2,4,8}, stride=2, receptive field=7×7, where d is the dilation coefficient of the dilated convolution kernel, stride is the step size of the convolution kernel, and receptive field is the receptive field size;
[0063] Subsequently, an LSTM module is added for dynamic tracking, and the configuration is as follows:
[0064] Hidden layer dimension = 256, time window length = 15 frames;
[0065] The second sub-network processes the spectral data from the three sets of atlases, uses the spectral feature fusion layer to perform channel splicing and weighting, and outputs a pixel-level feature map containing the material reflectance gradient. The fusion calculation method is:
[0066] Among them, Ffused Represents the pixel-level material feature map (reflectance gradient map) after weighted fusion of three sets of spectral maps. is the feature map corresponding to the i-th band spectrum, w i Represents the weight factor dynamically adjusted by the learnable parameter matrix, satisfying ∑w i =1.
[0067] Relationship modeling and feature set combination subunit: A three-dimensional graph attention mechanism is introduced at the end of the two-stream network to generate non-Euclidean space feature vectors:
[0068] Establish the curvature change correlation matrix between pixels:
[0069] Among them, K(i,j) represents the Gaussian curvature of the i,j pixel point, z(x,y) represents the grayscale height or reflectivity value of the image at the pixel, represents the second-order partial derivative of z with respect to the x direction, that is, the rate of change of the concave and convex along the horizontal direction, represents the second-order partial derivative of z with respect to y, that is, the rate of change of concavity and convexity along the vertical direction, and (i, j) represents the pixel coordinates of the i-th row and j-th column in the image;
[0070] The graph attention mechanism is introduced, and 8 parallel attention heads are used for non-Euclidean encoding. Finally, the three types of features are combined to form a complete art design feature set:
[0071]
[0072] in, Represents the spatiotemporal continuity characteristics of the brushstroke trajectory, Represents the pixel-level material reflectivity gradient feature map, It represents the non-Euclidean space topology vector obtained by hyperbolic geometric space mapping, and the mapping curvature radius R = 0.8.
[0073] The LSTM module is used to model the temporal changes of the brush stroke trajectory, especially the trend of the brush stroke pressure over time. The following is the state calculation process for each time step t:
[0074] 1. Input gate: controls the current input x t Whether the unit status is written:
[0075] i t =σ(W i x t +U i h t-1 +b i );
[0076] 2. Forget gate: controls the state c of the previous step t-1Whether to retain: f t =σ(W f x t +U f h t-1 +b f );
[0077] 3. Output gate: controls the cell state c t Which parts of are used for the current output:
[0078] o t =σ(W o x t +U o h t-1 +b o );
[0079] 4. Candidate state: Generate candidate memory content for the current time step:
[0080]
[0081] 5. Unit state update: Update the current memory state according to the forget gate and input gate:
[0082]
[0083] 6. Current output: Combine the output gate and the current memory state to calculate the current hidden state:
[0084] h t =o t ⊙tanh(c t );
[0085] Above: x t is the current input vector, the current input x t is the stroke image feature vector extracted from the tth frame, including the spatial position information of the stroke path and the corresponding pressure value encoding, which is used to reflect the local spatiotemporal state of the drawing action in this frame. t-1 is the output of the previous time step, c t-1 is the cell state of the previous time step, i t ,f t ,o t They are input gate, forget gate, and output gate respectively. is the candidate unit state at the current time step, σ(·) is the sigmoid activation function, tanh(·) is the hyperbolic tangent activation function, and ⊙ represents the element-wise product; W * ,U * ,b * are the trainable weight matrices and bias parameters.
[0086] The LSTM module acts on the time series input {x1,x2,…,x n}, by recursively calculating the output sequence {h1,h2,…,h n}, used to represent the pressure change trend during continuous drawing and embedded into the feature set as a high-order dynamic feature of the brush stroke trajectory middle.
[0087] The parameter matching unit specifically includes:
[0088] Multimodal Conditional Generative Adversarial Network Construction Subunit: Construct a Multimodal Conditional Generative Adversarial Network (MC-GAN), whose generator G receives three types of parallel inputs from the art design feature set:
[0089] First channel: Set the stroke trajectory feature Mapped to a latent space vector, the basic geometric grid is generated through the spatiotemporal encoder: ; Among them, CapsEnc(·) represents the encoder using capsule network, and the capsule network dimension is set to [8, 16, 32];
[0090] Second channel: Material reflectivity characteristics Input to the physical rendering pipeline, and use the Monte Carlo ray tracing algorithm to generate dynamic lighting and shadow effect layers:
[0091] N = 256; where I(x, y) is the brightness value of the final composite image at the pixel point (x, y), L i (x, y) represents the lighting result of the i-th path, N is the number of sampling points per pixel, and the depth is set to 8, that is, after each light reflection, the path is terminated based on probability to avoid infinite bounces of all paths and reduce the amount of calculation. The setting of "depth of 8" means that the light is allowed to bounce up to 8 times, and the path exceeding this depth will be forcibly terminated;
[0092] The third channel: receiving the target device parameter set Includes the following:
[0093] Among them, f VR is the screen refresh rate, range: 90–144, C color Indicates color gamut coverage (DCI-P3), R fov is the field of view curvature radius, range: 500-1500mm;
[0094] The parameter adaptation layer adjusts the resolution of the output image and the interpolation calculation error ∈ res Control is: ∈ res <0.5 pixels,∈ resIndicates the interpolation deviation between the content image and the device resolution.
[0095] Device-aware verification subunit of the discriminator: In the adversarial network discriminator D, the device adaptation-aware verification mechanism is integrated:
[0096] Color gamut matching verification: Calculated based on the CIEΔE2000 color difference formula, the matching standard is: ΔE 00 <2.3, calculation parameters follow: 2° observation angle, standard light source D65 conditions, ΔE 00 CIE 2000 color difference value, used to measure the color deviation between the generated image and the standard color gamut (P3) of the target device;
[0097] Motion blur prediction: Predicting the smear index of generated content in VR devices Expressed as:
[0098] Where τ is the response time constant, t is the response time, v max is the maximum eye angular velocity, f VR is the refresh rate, and the streaking threshold is I<5%.
[0099] Geometric distortion correction: Based on the field of view curvature radius R fov ,The vertex shader is used to perform deformation compensation, and the distortion parameters used include: k1∈[-0.15,0.15], p1, p2∈[-0.05,0.05], where k1 is the radial distortion coefficient, and p1, p2 are the tangential distortion coefficients.
[0100] Interactive response event marker embedding subunit: The response region of the output frame sequence is embedded through the event marker generator, which includes:
[0101] Haptic feedback trigger area: based on the material reflectivity gradient value Classification of tactile response levels:
[0102] F p ∈[0.1,5]N / mm 2 , F p It is the tactile feedback pressure response value, reflecting the tactile feedback strength required in different areas;
[0103] The classification is based on the Hertz contact model with a mesh accuracy of 0.1 mm2.
[0104] Light-shadow interaction sensitive area: When the shadow edge change rate satisfies: ΔS / Δt>15% / frame, it is marked as a high-priority update area. ΔS is the change in pixel area of the shadow edge, indicating the change in shadow shape between two frames. Δt is the time interval between adjacent frames, which is equal to 1.
[0105] Device adaptation identifier: Generate a hash value from rendering constraints for distributed synchronization verification:
[0106] ; where hash generation and verification delay T verify Controlled at: T verify <0.3ms.
[0107] The virtual reality generation unit specifically includes:
[0108] (a) Six-DOF pose fusion and purification: A nine-axis inertial measurement unit (IMU) is fused with an infrared optical tracking system to obtain six-DOF pose data (position x, y, z and attitude angles φ, θ, ψ). The extended Kalman filter (EKF) is then used to eliminate noise and generate a purified pose stream.
[0109] Pose solution frequency: f pose ≥120Hz;
[0110] Attitude angle error: ∈ θ <0.1°;
[0111] Spatial positioning standard deviation: σ x,y,z ≤0.3mm;
[0112] The Kalman filter noise covariance matrix is set as:
[0113] Q=diag([0.01,0.01,0.01,0.001,0.001,0.001]);
[0114] The IMU sensor parameter range is:
[0115] Acceleration range: ±16g; gyroscope range: ±2000° / s; magnetometer range: ±4900μT.
[0116] (b) Event marker driven engine: triggers feature update operations based on user interaction trigger areas:
[0117] Calculation of feature set update amount: When the operation trajectory intersects with the tactile feedback area, the calculated update amount is:
[0118] ΔF=K p ∫(F(t)-F0)dt; where ΔF is the characteristic adjustment amount to be injected, K p is the material stiffness coefficient (canvas: 0.8, watercolor paper: 0.3), F(t) is the pressure at the current interaction moment, and F0 is the trigger threshold pressure.
[0119] The interaction heat value is calculated as: Among them, H v is the heat value of the interaction area, N interactis the number of interactions within a unit time window, is the pressure gradient (intensity of regional pressure distribution change), T is the time window, sliding mechanism, length is 500ms, step size is 100ms, α = 0.6, β = 0.4 represent weight coefficients.
[0120] (c) Incremental rendering and differential encoding: Rendering optimization uses differential encoding and progressive generation strategies:
[0121] Hash comparison strategy: Use CityHash64 to compare the feature subset of each frame. If the hash difference ΔH>15%, the region is regenerated.
[0122] Progressive residual generation:
[0123] The first frame output basic grid size: R0 = 256 × 256;
[0124] Subsequent level-by-level residual upsampling: R i =2 i 256, i = 1, 2, ..., 3 (up to 2048 × 2048), each upsampling stage introduces a deformable convolution kernel to adapt to local structure changes.
[0125] (d) Distributed synchronization and content push strategy: Multi-terminal synchronization relies on device identification and layered data flow structure:
[0126] Frame attaching device fingerprint: Among them, T white is the screen color temperature calibration value, ranging from 2000-10000K, λ delay The rendering pipeline delay compensation coefficient is in the range of [0.8, 1.2]. The device fingerprint is encrypted and transmitted using HMAC-SHA1.
[0127] Layered Synchronization Protocol:
[0128] The base layer transmits geometric topology data, with a frame data volume of ≤5KB / frame;
[0129] Enhanced layer transmission light and shadow effect stream: bit rate ∈ [8,50] Mbps;
[0130] Content merging mechanism: When content version differences between terminals are detected, a content merging algorithm based on the minimum spanning tree (MST) is used:
[0131] Network delay weight w i,j Set as:
[0132] Merge processing delay τ merge Controlled by: τ merge <10ms.
[0133] The present invention encompasses any alternatives, modifications, equivalents, and solutions that fall within the spirit and scope of the present invention. To provide a thorough understanding of the present invention, specific details are described in detail below in connection with the preferred embodiments of the present invention, but those skilled in the art will be able to fully understand the present invention without these detailed descriptions. Furthermore, to avoid unnecessary confusion regarding the essence of the present invention, well-known methods, processes, procedures, components, and circuits have not been described in detail.
[0134] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as within the scope of protection of the present invention.
Claims
1. Art design virtual reality content synchronous display system based on AI and image detection, characterized by: include: A feature acquisition unit acquires surface texture data of an art design object through a multispectral image acquisition device and extracts at least three art design feature sets using a dynamic semantic segmentation network. The art design feature sets include brushstroke trajectory features, material reflectivity features, and spatial topological relationship features. A parameter matching unit inputs the art design feature set into a pre-trained generative adversarial network model, synchronously obtains display characteristic parameters of the target VR terminal through device parameter analysis, and generates a virtual reality content frame sequence that matches the display characteristic parameters, wherein the virtual reality content frame sequence includes a dynamic light and shadow effect layer and an interactive response event marker; Virtual reality generation unit: monitors the spatial posture data generated by user interaction operations in real time, triggers feature update instructions based on the interaction response event marker, feeds back the updated art design feature set to the generative adversarial network model for incremental content rendering, and pushes the adapted virtual reality content to multiple terminals through a distributed synchronization engine.
2. The art design virtual reality content synchronous display system based on AI and image detection according to claim 1 is characterized in that: In the feature acquisition unit, based on a multispectral image acquisition device, a preset wavelength range is used to scan the surface of the art design entity object three times in an orthogonal polarization mode to obtain a visible light reflection spectrum, a near-infrared absorption spectrum and a short-wave infrared scattering spectrum respectively.
3. The art design virtual reality content synchronous display system based on AI and image detection according to claim 2 is characterized in that: The feature acquisition unit also includes constructing a two-stream dynamic semantic segmentation network, in which the first sub-network uses a dilated convolution kernel to extract the spatiotemporal continuity characteristics of the brush stroke trajectory and tracks the brush stroke pressure change curve through a long short-term memory module; the second sub-network uses a spectral feature fusion layer to weightedly splice the channel dimensions of the three groups of atlases and output a pixel-level feature map including the material reflectance gradient value.
4. The art design virtual reality content synchronous display system based on AI and image detection according to claim 3 is characterized in that: A three-dimensional graph attention mechanism is deployed at the end of the dual-stream dynamic semantic segmentation network. By establishing a curvature change correlation matrix of adjacent pixels, a non-Euclidean feature vector representing the spatial topological relationship is generated. Finally, the spatiotemporal continuity features, pixel-level feature maps and non-Euclidean feature vectors are combined into the artistic design feature set.
5. The art design virtual reality content synchronous display system based on AI and image detection according to claim 1 is characterized in that: The generative adversarial network model in the parameter matching unit adopts a multimodal conditional generative adversarial network, including a generator G and a discriminator D.
6. The art design virtual reality content synchronous display system based on AI and image detection according to claim 5 is characterized in that: The input of the generator G is connected to three parallel channels of the art design feature set: The first channel maps the stroke trajectory features to latent space vectors and generates the basic geometric grid through the spatiotemporal encoder; The second channel inputs the material reflectivity characteristics into the physical rendering pipeline and combines it with the Monte Carlo ray tracing algorithm to generate a dynamic light and shadow effect layer; The third channel receives the display characteristic parameter set output by device parameter analysis, including the screen refresh rate, color gamut coverage, and field of view curvature radius of the target VR terminal, and generates device-related content rendering constraints through the parameter adaptation layer.
7. The art design virtual reality content synchronous display system based on AI and image detection according to claim 5 is characterized in that: The discriminator D is deployed with a device perception verification mechanism, which specifically includes: Color gamut matching verification: Calculate the ΔE2000 color difference between the generated content and the P3 color gamut of the target device; Motion blur prediction: Predicts the image smear index based on screen refresh rate and eye tracking data; Geometric distortion correction: Dynamically adjust the vertex shader's deformation compensation parameters based on the field of view curvature radius.
8. The art design virtual reality content synchronous display system based on AI and image detection according to claim 7 is characterized in that: The method further includes embedding an interactive response event marker in the content frame sequence by an event marker generator, wherein the event marker includes: Tactile feedback trigger area: the pressure response level is divided according to the material reflectivity gradient value; Light-shadow interaction sensitive area: Marks the dynamic shadow update priority based on ray tracing results; Device adaptation identifier: records the hash value of the rendering constraints for verification by the distributed synchronization engine.
9. The art design virtual reality content synchronous display system based on AI and image detection according to claim 1 is characterized in that: The virtual reality generation unit specifically includes: a. Acquire six-degree-of-freedom pose data by fusing a nine-axis inertial measurement unit with an infrared optical tracking system, use a filter to eliminate jitter noise, and generate a purified pose stream including a timestamp; b. Build an event marker driven engine to trigger feature update operations based on user interaction trigger areas: c uses differential encoding technology in the incremental rendering stage: The feature hash comparison algorithm is used to identify the areas that need to be updated, and only the feature subset with hash value differences greater than 15% is regenerated; Use a progressive generation strategy: generate a low-resolution base mesh in the first frame, and gradually improve it to the target resolution through residual connections in subsequent frames; d When the distributed synchronization engine performs multi-terminal push: Attach a device fingerprint identifier to each content frame, including the screen color temperature calibration value and rendering pipeline delay compensation coefficient; Adopting a layered synchronization protocol: the base layer transmits geometric topology data, and the enhanced layer transmits light and shadow effects data in a streaming manner; When version differences between terminals are detected, a content merging algorithm based on the minimum spanning tree is triggered.
10. The art design virtual reality content synchronous display system based on AI and image detection according to claim 9 is characterized in that: The execution logic of the event marker driving engine is as follows: parsing the tactile feedback triggering area boundary coordinates in the interactive response event marker; When the operation trajectory in the purified pose stream intersects with the trigger area: Calculate the feature set update amount ΔF according to the pressure-displacement curve; Generate an update instruction packet with a priority tag, where the priority is determined by the heat value Hv of the interaction area.
Citation Information
Patent Citations
Apparatus for implementing a sequence transform neural network for transforming an input sequence and learning method using the same
KR1020240027347A
Systems and methods for augmented reality art creation
US20180276882A1
Method and system for augmented imaging using multispectral information
WO2020025696A1