Traditional art 3D content conversion method based on artificial intelligence
By constructing a cross-scene semantic anchor point map and a spatiotemporal constraint baseline, combined with millisecond-level preheating frame calibration and texture bounce mechanism, the problem of semantic tag mismatch in traditional art 3D content conversion is solved, achieving visual consistency and interaction stability in multiple scenes, and improving the artistic fidelity and user experience of virtual exhibitions.
Patent Information
- Application Number
- CN202511419848.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-30
- Publication Date
- 2026-02-24
AI Technical Summary
In virtual exhibitions and other multi-scene switching environments, the lack of a stable constraint mechanism for semantic tags in the conversion of traditional art 3D content leads to the incorrect projection of character feature tags onto the background material layer, and the embedding of background details into the foreground, causing visual disorder and affecting the accuracy of artistic reproduction and the stability of the interactive experience.
We construct a cross-scene semantic anchor point map and a spatiotemporal constraint baseline, calibrate semantic anchor points through millisecond-level warm-up frames, perform hierarchical binding and mapping snapshot tracing, use texture bounce mechanism to correct deviations, and establish an adaptive adversarial verification loop in combination with user interaction feedback to ensure semantic stability.
It achieves visual consistency and semantic integrity of traditional art 3D content across multiple scene transitions, and enhances the artistic fidelity of virtual exhibitions and the stability of immersive user interaction.
Smart Images

Figure CN121563752A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, specifically to a method for converting traditional art content into 3D based on artificial intelligence. Background Technology
[0002] The transformation of traditional art into 3D content based on artificial intelligence refers to using AI technology to intelligently analyze and extract the two-dimensional image features, brushstroke styles, color layers, and artistic semantics inherent in traditional artworks (such as traditional Chinese paintings, oil paintings, calligraphy, and sculptures) through deep learning, computer vision, and generative modeling methods. These elements are then mapped, restored, or extended into a three-dimensional digital space to generate three-dimensional content with the aesthetic appeal of traditional art. This process not only preserves the artistic charm and cultural connotations of the original work but also endows it with new spatial expressions, enabling immersive reproduction and innovative applications in scenarios such as virtual reality, augmented reality, digital exhibition halls, and cultural heritage preservation.
[0003] Existing technologies suffer from the following shortcomings: In current technologies, the transformation of traditional 3D art content based on artificial intelligence often lacks a stable constraint mechanism for semantic tags in multi-scene switching environments such as virtual exhibitions. When scenes rapidly switch between different exhibition halls, backgrounds, or interactive states, the semantic tags carried by traditional elements are prone to cross-layer mismatch during instantaneous mapping. Feature tags that should belong to the main subject are incorrectly projected onto the background material layer, while details that should remain in the foreground are embedded in the environmental texture, resulting in significant visual disorder. This problem not only undermines the semantic integrity of traditional artworks in three-dimensional space but also causes strong visual confusion for viewers during immersive experiences, severely impacting the artistic fidelity and interactive experience stability of virtual exhibitions.
[0004] The information disclosed in the background section is only intended to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention
[0005] The purpose of this invention is to provide a method for converting traditional art content into 3D based on artificial intelligence, so as to solve the problems in the background art mentioned above.
[0006] To achieve the above objectives, the present invention provides the following technical solution: a method for converting traditional art 3D content based on artificial intelligence, comprising the following steps: Construct a cross-scene semantic anchor point map to generate verifiable semantic anchor point signatures for character features and background materials in traditional art images, and simultaneously establish a spatiotemporal constraint baseline based on spatial coordinates and time series as a reference for subsequent semantic mapping. Before scene switching, a millisecond-level warm-up frame is loaded based on the spatiotemporal constraint baseline, and temporal calibration is performed on the semantic anchor signature to output a phase reference for controlling the subsequent mapping process; Based on phase reference, hierarchical binding of semantic features is performed, binding character features to the foreground layer and background material to the environment layer, and generating a mapping snapshot of the current semantic state. The mapping snapshot supports state traceability. A cross-frame tolerance corridor is established based on the mapping snapshot. Deviation detection is performed by comparing the actual position of the semantic target with the snapshot record frame by frame. If the detected deviation exceeds the tolerance range, the texture bounce mechanism is triggered to correct the mapping offset and maintain the stability of the layered binding. For semantic anomalies remaining after bias correction, semantic consistency adjudication is performed based on residual analysis. Mismatch source points are identified through multi-view comparison and skeleton topological constraints, and local semantic remapping is performed after identification. After local semantic remapping is completed, an adaptive adversarial verification loop driven by human-computer resonance is constructed based on the user's gaze trajectory and interactive gestures. Semantic weights are adjusted in real time, and the label state is dynamically controlled through the label freezing and unfreezing mechanism to form a semantically stable closed loop in the cross-scene switching process.
[0007] Preferably, the steps for establishing a spatiotemporal constraint baseline based on spatial coordinates and time series are as follows: Frame by frame, traditional art images are analyzed to extract significant expressive areas of human features and background materials, and human features and background materials are semantically labeled and semantically encoded respectively. Based on spatial location, boundary contour, feature texture, color level and semantic encoding, semantic anchor signatures of human features and background material are generated, and the semantic anchor signatures are bound to the spatial location of the image. By combining semantic anchor signatures with time series and three-dimensional spatial coordinates, the positional offset, morphological changes and spatial trajectory of semantic anchors in the frame sequence are constructed to form a spatiotemporal constraint baseline. The semantic anchor graph, semantic anchor signature, and spatiotemporal constraint baseline are uniformly organized to establish a semantic anchor graph set for multi-scenario loading, which is used for semantic label consistency mapping and spatial position correction during scene switching.
[0008] Preferably, the phase reference output step for controlling the subsequent mapping process is as follows: High frame rate image sequences with a loading time range of 15 to 40 milliseconds are used as warm-up frames to extract semantic anchor point state data of character features and background materials; The constructed spatiotemporal constraint baseline is used to perform frame-level comparison of each semantic anchor point in the preheating frame, and to perform temporal series calibration of positional deviation, morphological changes and texture direction. Based on the motion trend, edge frequency, texture gradient and spatial displacement of each anchor point, state phase parameters are generated, and a stable phase reference for each anchor point in the time window is output. Based on the stable phase reference, the spatial location, layer number, depth sorting and label weight of the semantic anchor in the new scene are initialized and configured to complete the mapping preparation before scene switching.
[0009] Preferably, the steps for generating the mapping snapshot of the current semantic state are as follows: Based on the dynamic trajectory and spatial behavior trend of semantic anchors in the warm-up frame, and combined with phase reference values, layer positioning preparation is carried out to distinguish between foreground layer and environment layer to bind candidate anchors. Based on the layer positioning results, perform a layered binding operation of semantic anchor points, bind the character features to the foreground layer, bind the background material to the environment layer, and set the layer number and layer sort number; Based on the bound semantic anchor state, generate a semantic state mapping snapshot containing 3D coordinates, layer number, sort number, texture index, phase reference value and timestamp; The current mapping snapshot is combined with the previous frame snapshot to construct a time-series snapshot chain, recording the state change trajectory of each anchor point, forming a dynamic data chain that supports semantic state tracing.
[0010] Preferably, the phase reference value in the semantic state mapping snapshot is used to control the rendering priority of the anchor point in the layer, and the layer sort number of each anchor point is dynamically adjusted according to the stability of the phase reference value, so as to enhance the visual hierarchy stability of character features and background materials during multi-scene switching.
[0011] Preferably, if the detection deviation exceeds the tolerance range, the steps to trigger the texture bounce mechanism to correct the mapping offset and maintain the stability of the layered binding are as follows: Based on the three-dimensional spatial position and phase reference value of the semantic anchor point in the previous frame's mapping snapshot, a spatial tolerance corridor with defined three-axis boundaries is constructed. Extract the actual 3D position of the semantic anchor point in the current frame and compare it with the tolerance corridor axis by axis difference to determine whether there is a deviation exceeding the limit. For anchor points with excessive deviation, the bounce path is calculated based on three-frame snapshot data, and a texture bounce correction operation is performed. The corrected anchor state is written to the current frame mapping snapshot and updated to the cross-frame tracing chain to achieve dynamic stability of the semantic state.
[0012] Preferably, the local semantic remapping execution steps are as follows: Based on the anchor point state after texture bounce correction, perform residual analysis to determine whether the differences between the three-dimensional coordinates, texture binding area and phase reference value exceed the stability threshold. For semantic anchors that are determined to be abnormal, the semantic anchors that constitute mismatch source points are identified by comparing multi-view image features and calculating the angle between the skeleton topology. Perform local semantic remapping operations on the spatial location, texture attribution, and layer priority of the identified mismatch source points to generate a new mapping snapshot; The remapped anchor state is written into the traceability chain and local consistency verification is performed to ensure structural integrity and semantic binding continuity.
[0013] Preferably, after local semantic remapping is completed, the following steps are performed to construct a semantically stable closed loop: Collect user eye movements and interactive gestures to identify the main interaction focus area; Calculate the weight level of semantic anchor points based on the main interactive focus area, and establish a semantic response structure; For semantic anchors whose weight factors exceed a set threshold, a freeze operation is performed, and the freeze timestamp and unfreeze conditions are set. Unfreeze the system when there is a gaze shift, insufficient gaze duration, or a change in interaction action. Record the change in frozen state information into the semantic tag management chain to achieve continuity and stability control of semantic behavior.
[0014] The technical effects and advantages provided by the present invention in the above technical solution are as follows: This invention achieves precise correspondence between semantic tags and 3D spatial structures by constructing verifiable semantic anchor signatures and spatiotemporal constraint baselines. The introduction of millisecond-level warm-up frames and the generation of phase references ensure temporal coherence and responsiveness during semantic mapping. Layered semantic feature binding and a mapping snapshot tracing mechanism ensure logical isolation and visual consistency between character features and background materials. Cross-frame tolerance corridors and texture bounce technology enable highly sensitive correction of dynamic offsets. Residual analysis and skeleton topological constraints achieve accurate identification of mismatched source points and local semantic remapping, effectively repairing visual disturbances. Finally, by combining a resonance feedback mechanism of user gaze trajectory and interactive gestures, an adaptive adversarial verification closed loop is established, dynamically adjusting semantic weights and achieving stable inheritance and seamless transition of semantic tags across multiple scenarios through a freeze-thaw mechanism. The overall solution not only enhances the artistic fidelity and spatial expressiveness of traditional 3D art conversion but also significantly improves the stability and continuity of immersive user interaction in virtual exhibitions. Attached Figure Description
[0015] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this invention. For those skilled in the art, other drawings can be obtained based on these drawings.
[0016] Figure 1 This is a flowchart of the method for converting traditional art 3D content based on artificial intelligence, according to the present invention. Detailed Implementation
[0017] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, they are provided so that the description of this disclosure will be more complete and fully convey the concept of the exemplary embodiments to those skilled in the art.
[0018] This invention provides, for example Figure 1 The illustrated method for converting traditional art content into 3D, based on artificial intelligence, includes the following steps: Construct a cross-scene semantic anchor point map to generate verifiable semantic anchor point signatures for character features and background materials in traditional art images, and simultaneously establish a spatiotemporal constraint baseline based on spatial coordinates and time series as a reference for subsequent semantic mapping. By establishing a semantic anchor point map and combining spatial coordinates with time series information to construct a spatiotemporal constraint baseline, a stable reference basis is provided for subsequent semantic mapping. This step is divided into stages: This approach involves frame-by-frame analysis of traditional art images to extract salient regions of human features and background textures, and semantic annotation of each visual element. Specifically, a deep convolutional neural network is used to perform boundary recognition and feature extraction on compositional elements in the image, separating the subject from the background at the pixel level. Human feature extraction includes not only structural contours but also multi-dimensional visual aspects such as facial expressions, body postures, clothing textures, and movement trends. Background texture extraction encompasses spatial elements such as landscape layout, decorative brushstrokes, architectural patterns, and the spatial relationships of objects. Each extraction operation is performed using a trained model and semantically categorized based on a labeled dataset. After categorization, a unique semantic code is assigned to each identified semantic category, and this code is bound to a specific region in the image, establishing a preliminary correspondence between semantic elements and spatial locations.
[0019] After completing the semantic label encoding and binding, a verifiable semantic anchor signature is constructed. This signature is a joint description of the spatial location, boundary contour, feature texture, color level, and semantic classification of each visual element in the image, forming a unique representation structure with multi-dimensional attributes. Taking human features as an example, its semantic anchor signature will include the starting coordinates of the image where the person is located, the coordinate set of the shape contour, the spatial distribution of facial feature points, the histogram features of the color layer, and the corresponding semantic encoding combination; for background materials, its anchor signature includes texture direction vector, regional average color value, stroke thickness variation, layer stacking order, and spatial depth estimation data. By uniformly encapsulating the above information into a structured semantic anchor signature, it is possible to ensure accurate identification of each semantic element in subsequent 3D reconstruction and scene switching, and to effectively track and verify them. This semantic anchor signature differs from the semantic binding methods based on fixed labels or global features in existing technologies. Its innovation lies in introducing spatial density description and content hierarchy separation, which has stronger stability and recognizability, and can maintain the label unchanged in dynamic scenes.
[0020] After constructing the semantic anchor signatures, a complete spatiotemporal constraint baseline is further established by combining the spatial coordinates of the anchors and their frame positions in the image sequence. This baseline, with time frames as the main axis, records the positional offset, morphological changes, and semantic stability of each semantic anchor signature in consecutive frames, forming a temporal path based on inter-frame continuity. Simultaneously, the two-dimensional position of each semantic anchor in the image is mapped to three-dimensional spatial coordinates, generating a complete spatial reference trajectory. Taking human features as an example, the spatial trajectory description includes the coordinate change paths of key points such as the head, shoulders, and limbs in each frame; the background material is represented by a spatial curve from the edge of the image to the visual vanishing point through depth estimation and perspective correction. This spatiotemporal constraint baseline plays a core reference role in this invention, not only for precise alignment during subsequent semantic mapping but also as a criterion for evaluating semantic stability in dynamic scenes. Unlike existing technologies that only use static labels or rule-based mapping, this invention achieves the binding of multi-dimensional spatiotemporal features through frame-level tracking and spatial synchronization mechanisms, thereby enhancing semantic continuity and robustness.
[0021] After constructing the spatiotemporal constraint baseline, the semantic anchor point map, its corresponding anchor point signatures, and their trajectory information within the spatiotemporal constraint baseline are uniformly organized to establish a cross-scenario semantic anchor point map set for multi-scenario applications. This map set is organized by semantic categories, with each category of semantic anchor points corresponding to its signature data and spatiotemporal trajectory set, allowing for rapid indexing and retrieval during different scene loading. During scene switching, this map set provides standardized anchor point templates to achieve consistent semantic recognition mapping, avoiding misalignment of character features or background materials due to scene changes. For example, when switching from exhibition hall A to exhibition hall B, the system first searches the map for the semantic anchor point signature that best matches the current content captured in the current frame, compares it with the baseline trajectory, and then corrects the spatial mapping position and temporal phase of the semantic tags.
[0022] Before scene switching, a millisecond-level warm-up frame is loaded based on the spatiotemporal constraint baseline, and temporal calibration is performed on the semantic anchor signature to output a phase reference for controlling the subsequent mapping process; To achieve temporal alignment and stable mapping of semantic anchor points during scene transitions in traditional 3D content conversion, based on the constructed spatiotemporal constraint baseline and semantic anchor point map, short-time high-frequency image sequences are used for anchor point preheating before scene transitions are triggered. Combined with dynamic trajectories, precise calibration of the anchor point temporal phase is achieved. The specific implementation includes the following steps: A high-frame-rate image sequence with a time window ranging from 15 to 40 milliseconds is loaded as a warm-up frame sequence before scene switching. The image content of the warm-up frames must be derived from a continuous frame sequence of the currently active scene, and data should not be retrieved from future target scenes to avoid semantic drift. The acquisition frequency of each frame is set to more than 60 frames per second to ensure that semantic detail changes are completely recorded at the millisecond level. During image preprocessing, pixel-level brightness equalization, color channel decoupling, edge enhancement, and noise removal are performed on each frame to ensure subsequent recognition accuracy. After loading, semantic anchor point extraction is performed on each frame. Taking the features of a person in a traditional art scene as an example, the extracted content includes five specific visual features: head outline, eye gaze direction, mouth opening / closing state, hair ornament patterns, and the position of clothing fold edges; while for background materials, the direction lines of mountain and rock textures, the grayscale changes of distant mountains, the diffusion radius of light spots on the water surface, the perspective angle of window frames, and the light and shadow transition boundaries are extracted. The above information serves as the preliminary anchor point state data for the corresponding frame for subsequent calibration.
[0023] After the preheating frame sequence is loaded, the semantic anchors in each frame are compared and calibrated using the spatiotemporal constraint baseline built during the image analysis phase. The specific operation includes three sub-steps: First, the timestamp, position coordinates, edge morphology point set, texture descriptor, and color gradient vector contained in the anchor signature are extracted and compared one-to-one with historical data of the same category of semantic anchors in the spatiotemporal constraint baseline. Then, the average positional deviation, morphological contour deviation, and texture direction angle deviation of each anchor in the frame relative to its historical trajectory are calculated. Taking the human eye anchor as an example, if the angle of change in gaze direction in five consecutive preheating frames differs from historical data by more than 5 degrees, an angle correction algorithm is triggered to adjust the angle value to be consistent with the direction of the most recent time frame; if the mouth opening / closing state fluctuates by more than 15 pixels in consecutive frames, boundary smoothing is performed. Finally, the trajectories of all anchors within the preheating frame time period are dynamically aligned to achieve frame-level synchronization between their temporal sequence state and the change patterns in historical trajectories, thus realizing the temporal behavior trend correction of semantic anchors.
[0024] After completing the temporal calibration of anchor point signatures, phase state parameters are extracted for each semantic anchor point in the aligned preheating frames, and a phase reference is generated to control subsequent mapping. This phase reference is not a static timestamp, but a comprehensive expression of the anchor point's change frequency, structural stability, spatial drift amplitude, and semantic fluctuation value over a short period of time. Specifically, the implementation is as follows: First, the mean and variance of the changes in key point positions for each anchor point within the preheating time window are statistically analyzed to determine its movement trend; second, Fourier analysis is performed on the edge morphology point set to extract its dominant frequency features; third, the average slope and maximum gradient jump value of the anchor point's texture gradient transformation in consecutive frames are evaluated; fourth, the displacement amplitude of the anchor point along the depth axis (Z-axis) in three-dimensional space is recorded. After all parameters are extracted, a weighted aggregation operation is performed to output a stable phase value representing the state center of the anchor point within this time window. For example, if a person's head feature anchor point maintains a constant gaze direction, minimal boundary contour changes, and stable facial texture across consecutive frames, its phase reference will be set as the primary calibration point at the midpoint of the frame time interval. Conversely, if the water surface texture in the background exhibits a periodic diffusion trend, its phase reference will fall on the frame with the least perturbation during this diffusion process. Through this phase extraction mechanism based on behavioral feature evolution, each semantic anchor point obtains a mapping control reference that highly matches its semantic evolution behavior.
[0025] After the phase reference is constructed, the initialization preparation process for semantic anchors is immediately initiated to ensure consistent state continuity and mapping during scene transitions. First, based on the generated phase reference, the spatial initialization parameters of the current semantic anchor in the upcoming new scene are adjusted, including position coordinates, binding layer number, depth sorting order, and foreground / background occlusion weights. Taking a character's head anchor as an example, based on its stable gaze direction and 3D depth position in the current pre-warm-up frame, it is mapped to a position in the new scene with the same viewing angle and consistent distance from the camera. If the anchor is located at the character's foreground boundary, the layer binding order is set to the highest level to prevent it from being covered by background textures. Second, the mapping intensity weight of the semantic anchor is initialized, assigning it a high, medium, or low priority label based on its semantic stability index in the pre-warm-up frame. Third, the dynamic backtracking permissions for anchors in subsequent mappings are pre-set. For example, if the texture direction of a leaf texture anchor in the background material shows a left-right drift trend in the pre-warm-up frame, the system will record its most frequent direction as a backtracking reference direction for subsequent position correction. Finally, the initialization state of all anchor points is verified for consistency to ensure that upon entering the target scene, anchor points do not mutate, misalign, or become invalid due to viewpoint shifts, texture blending, or rendering delays. After this initialization, all anchor points are in a state of temporal and spatial synchronization, providing a stable semantic structure foundation for scene transitions.
[0026] Based on phase reference, hierarchical binding of semantic features is performed, binding character features to the foreground layer and background material to the environment layer, and generating a mapping snapshot of the current semantic state. The mapping snapshot supports state traceability. To ensure the visual stability and semantic consistency of traditional 3D art content during virtual scene transitions, after constructing the semantic anchor point phase reference, it is necessary to further perform a structured, layered binding operation of semantic features based on this phase reference. Simultaneously, a mapping snapshot of the current state is generated, and a traceable data chain is constructed. The specific implementation process includes the following steps: Based on the dynamic trajectory features and spatial behavior trends exhibited by semantic anchors in the preheating frames, layer positioning preparation is performed for all semantic anchors using phase reference values. The goal of this layer positioning preparation is to assign a clear layer candidate level to each anchor, thereby providing a benchmark for subsequent binding operations. The operation process is as follows: For human feature anchors, such as head contour anchors, facial feature anchors, limb contour anchors, and clothing pattern anchors, the three-dimensional spatial coordinate change path of each anchor is first extracted for three consecutive frames in the preheating frame sequence. Combined with the stability score provided by the phase reference, the depth axis displacement range is quantified. When the displacement range is less than a set threshold (e.g., 15 spatial pixels), and it remains within the forward viewing area between 0° and 15° of the main viewing angle in each frame, it is determined to have a high probability of being assigned to the foreground. Further judgment is made based on texture sharpness: if the average texture gradient change of the human feature anchor in the preheating frames exceeds a set brightness slope threshold (e.g., 0.85), it is confirmed as a foreground binding candidate. Similarly, for background material anchor points, such as mountain boundary anchor points, ground brick seam anchor points, window lattice structure anchor points, and sky background anchor points, if the slope change of their texture features in consecutive frames is less than a preset stable value (e.g., 0.2), and they are always located in the outer area of the main viewpoint (between 15° and 45°), accompanied by depth information that is consistently greater than a set value (e.g., 100 spatial pixels), then they are classified as candidate anchor points for the environment layer. This layer positioning, based on multi-dimensional quantitative indicators such as image texture clarity, depth spatial position, viewpoint direction, and temporal phase stability, ensures that the anchor point classification is based on clear and reliable principles.
[0027] After completing the anchor point layer positioning preparation, immediately perform the layered binding operation of semantic anchor points. Clearly bind character feature anchor points to the foreground layer and background material anchor points to the environment layer, assigning a unique binding number and layer priority index value to each anchor point. The specific operation process is as follows: For foreground layer binding, based on phase reference values, head-related anchor points (including eyebrow and eye boundaries, nose bridge center point, and corners of the mouth) are set as the highest priority anchor points in the layer, ensuring they are rendered first in the scene and closest to the viewpoint. Secondary body anchor points such as shoulders, torso, arms, and legs are arranged sequentially according to their 3D positions, constructing a binding order from near to far and from center to outward. The binding data for each anchor point includes five items: 3D coordinate vector, texture reference index, lighting reflection coefficient, binding layer number, and layer position index. The environment layer binding operation also follows the principle of from center to outward and from clear to blurry; for example, wall anchor points are bound to the second layer of the environment layer, ground anchor points to the third layer, and sky anchor points to the fourth layer. Once the binding is complete, all semantic anchor points have been divided into layers based on phase references to ensure clear separation and stable boundaries between the character and the background during the switching process.
[0028] Based on the binding results, a semantic state mapping snapshot of the current frame is constructed, and state data is encapsulated for each bound anchor point. This mapping snapshot has a seven-dimensional data structure: a unique anchor point identifier, three-dimensional spatial coordinates, binding layer number, layer internal sort number, texture sample index, current phase reference value, and binding timestamp. The snapshot generation operation is performed based on the current image frame, with the foreground layer and environment layer executed independently. For example, in the process of generating a snapshot of a person's feature anchor point, for the anchor point in the person's facial region, the X, Y, and Z values of the nose tip coordinate point in three-dimensional space (e.g., 125, 48, 37), its belonging layer number is the first foreground layer, its sort number is the second foreground layer, its texture sample index points to the 24th texture block, its phase reference value is 12.6 degrees, and its timestamp marks the current frame number as Frame_1050; this is a complete snapshot node. Similarly, in the environment layer, for example, the anchor point of the window pattern, its snapshot record coordinates are (88, 26, 198), layer number is the 2nd layer of the environment layer, sequence number is the 4th, texture sample is tile number 51, phase reference is 21.4 degrees, and timestamp is Frame_1050. All snapshot data is centrally written to the current frame mapping dataset in a structured manner, realizing a complete record of the layer and anchor point states.
[0029] To ensure the traceability of future states, after the mapping snapshot is generated, an anchor point snapshot chain is constructed based on the temporal sequence relationship between the current snapshot and the previous frame snapshot, and the range of inter-frame variation differences is defined. This snapshot chain connects the snapshot nodes of the same anchor point in each frame in chronological order, forming a dynamic recording trajectory of the anchor point's state changes in consecutive frames. For example, if the Z-axis coordinate of the character's head anchor point is 38 in Frame_1049 and 37 in Frame_1050, the displacement is -1 spatial unit. If this displacement value increases unidirectionally for three consecutive frames without any texture or layer changes, its movement trend is determined to be a stable forward movement, and no correction is triggered. Conversely, if the Z-axis of the anchor point suddenly changes to 60 in the next frame Frame_1051, exceeding the threshold range (set to 15), it is recorded as an abnormal change, and a texture bounce mechanism is subsequently triggered for repair. Furthermore, each snapshot chain not only records coordinate information changes but also includes the history of layer number changes, phase reference fluctuation amplitude, and texture matching similarity value, providing a precise basis for subsequent mismatch tracking, semantic drift detection, and layer repair. By establishing a complete snapshot chain, it is ensured that the semantic state of all anchor points can be accurately recovered, restored, and compared at any frame time point, thereby significantly enhancing controllability and semantic continuity during multi-scene switching.
[0030] A cross-frame tolerance corridor is established based on the mapping snapshot. Deviation detection is performed by comparing the actual position of the semantic target with the snapshot record frame by frame. If the detected deviation exceeds the tolerance range, the texture bounce mechanism is triggered to correct the mapping offset and maintain the stability of the layered binding. To ensure the spatial stability of semantic anchor points in traditional 3D art content during continuous frame rendering and scene transitions, a cross-frame tolerance corridor needs to be constructed based on the semantic position and attribute data after generating the mapping snapshot. This involves detecting deviations frame by frame and executing texture bounce when deviations exceed the tolerance range, thus achieving continuous and stable control of the layered binding structure. This process can be divided into the following steps: Based on the spatial location data of each semantic anchor point in the previous frame's mapping snapshot, a tolerance corridor is established for that anchor point in the current frame. The tolerance corridor is the allowed range of movement for the anchor point in three-dimensional space around its original position; its structure is a three-axis spatial volume region. Specifically, it is constructed by extracting the three-dimensional coordinates of feature-type anchor points (such as the nose tip anchor point) from the previous frame snapshot, for example, (124, 67, 39). The X-direction is set to ±5 units, the Y-direction to ±4 units, and the Z-direction to ±3 units. A cubic space with boundary lengths of 10, 8, and 6 units is constructed centered on this anchor point. The tolerance value is determined by the anchor point's phase reference stability parameter in the previous frame. For example, when its phase value fluctuates less than 2°, the tolerance range is compressed to 70% of the standard value; if the phase fluctuation is greater than 5°, it is expanded to 130% of the standard value. After this step, each anchor point has a set of permissible spatial movement regions generated based on its historical movement trends and semantic stability, which are used as bounding boxes for subsequent deviation detection.
[0031] After the current frame is rendered, the actual 3D position data of each semantic anchor point in the frame is extracted, and the difference between the anchor point and the corresponding anchor point position in the previous frame's snapshot is calculated on each axis to determine whether the current frame position is within the tolerance corridor. For example, in the current frame, if the coordinates of the anchor point at the left corner of the head are (131, 70, 41), then its displacement relative to the previous frame position (124, 67, 39) is X-axis +7, Y-axis +3, and Z-axis +2. This vector difference is then compared with the tolerance boundary constructed by the anchor point. If the X-axis direction is set to ±5, and the actual deviation is +7, it is considered to exceed the tolerance limit. Further analysis is performed on the anchor point's layer binding status and texture attachment area in the current frame, comparing it with the recorded data in the previous frame. If it is found that the bound layer has changed from the first level of the foreground layer to the second level, or its attached texture block number has changed from number 26 in the previous frame to number 34, this is also considered a semantic binding offset. In addition, the continuity of anchor point morphological features needs to be checked, such as whether the boundary tension of the contour point set breaks within three frames. If the continuity is broken, it is considered a structural deviation even if the coordinates do not exceed the limits. After all the above checks are completed, all anchor points that exceed the tolerance range or whose structural integrity is damaged are marked as "deviation exceeds the limit" and the semantic update operation of the current frame for that anchor point is frozen, and the next repair process begins.
[0032] For anchor points in a state of excessive deviation, a texture bounce correction mechanism is activated to restore their spatial position and semantic state to the phase reference trajectory. Texture bounce does not directly return the anchor point to the position of the previous frame, but rather calculates the bounce path based on the anchor point position data in three frame snapshots (the current frame and the two previous frames), the texture attachment area, the binding layer order, and the phase reference angle. Taking the center anchor point on a character's forehead as an example, its coordinates are (123, 64, 38) in Frame_1050, (124, 66, 39) in Frame_1051, and abnormally jump to (130, 70, 42) in Frame_1052. When executing the bounce mechanism, a weighted average of the coordinates of the three frames is first performed to obtain a stable bounce reference position (125.6, 66.3, 39.6). Then, the centroid positions of the texture map regions in the three frames and the contrasting texture brightness distribution are extracted to generate the bounce texture mapping path. Finally, the coordinates of the bounce result anchor point are adjusted to (126, 66, 40), and its texture attachment block is switched from the incorrect 45th region back to the correct 27th region. This operation not only corrects the coordinate position but also restores the layer priority order and binding attributes of the semantic anchor points. After texture bounce is executed, the layer structure of the current frame needs to be re-verified and sorted to prevent layer rendering order disorder caused by anchor point state changes, especially the risk of foreground anchor points occluding and misaligning.
[0033] After completing the texture bounce correction, all data of the corrected anchor points are synchronously updated to the current frame mapping snapshot and written into the cross-frame tracing chain to construct a continuous evolution trajectory of semantic states. This tracing chain records the following: current frame anchor point identifier, position before bounce, position after bounce, bounce-based frame sequence table, comparison information of texture blocks before and after correction, phase reference drift angle, and whether layer rebinding was achieved during this bounce. Taking the center anchor point of a character's mouth as an example, after texture drift occurs in Frame_1053, triggering texture bounce, its snapshot structure adds the following fields: "Original position (127,62,41), position after bounce (124,60,40), correction reference frame sequence: 1051, 1052, original layer priority 3, corrected priority 1, texture block adjusted from 36 to 22." This data is inserted into the anchor point tracing chain, forming a three-node continuous chain with the previous frame state. If an anchor point experiences more than two texture bounces within 5 frames, the tracing chain will automatically trigger a status warning flag, increase the layer binding priority of that anchor point in subsequent frames, expand its tolerance corridor volume, and enhance its semantic freeze threshold to prevent further unstable offsets. This mechanism ensures that semantic anchor points do not suffer from continuous rendering errors due to minor drifts or short-term misjudgments during multi-frame continuous rendering and rapid scene switching, thereby improving the overall stability and accuracy of 3D semantic mapping.
[0034] For semantic anomalies remaining after bias correction, semantic consistency adjudication is performed based on residual analysis. Mismatch source points are identified through multi-view comparison and skeleton topological constraints, and local semantic remapping is performed after identification. After completing texture bounce correction, some semantic anchors may still exhibit inaccurate spatial positioning, incomplete texture binding, or incorrect layer affiliation. To further correct these residual anomalies, semantic consistency adjudication needs to be performed based on residual analysis, and mismatch source points should be identified by combining multi-view data and skeleton topology relationships. Subsequently, precise semantic remapping is implemented within a local area. The specific steps are as follows: Residual analysis is performed on anchor points that have undergone texture bounce processing. By calculating the numerical differences in key attributes before and after bounce, potential semantic anomalies are identified. Taking the anchor point at the left elbow of a character as an example, assuming its spatial coordinates before bounce are (134, 60, 42), the binding layer is the second level of the foreground layer, the texture area is the pattern area in the middle of the left sleeve of the clothing, and the phase reference angle is 11.6°; after bounce, the coordinates become (129, 58, 40), the layer belongs to the first level of the foreground layer, the texture binding area is updated to the edge area of the pattern on the left shoulder, and the phase reference angle is 15.2°. Multiple residual operations are performed on these data to calculate the 3D coordinate difference vector as (-5, -2, -2), the layer weight change level as 1, the phase offset as 3.6°, and the texture matching area overlap as 62%. If any of these residual values exceeds the set stability threshold, such as the spatial offset exceeding 5 units or the texture overlap being less than 70%, the anchor point is marked as "semantically unstable" and enters the consistency adjudication process. Residual calculation is not limited to static attributes but also includes trajectory extensibility, i.e., whether the movement trend of the anchor point remains consistent across three consecutive frames. If the anchor point moves in a straight line in frames 1050, 1051, and 1052, but deflects abnormally in frame 1053, then a dynamic behavior residual is constituted. Through the above analysis, anchor points that appear to have been repaired on the surface but actually have a fundamental risk of misalignment can be accurately identified.
[0035] For anchor points identified as having semantic anomalies, multi-view consistency comparison analysis is performed, and the source of the mismatch is spatially located by combining its node relationship in the skeleton topology. The specific method is as follows: Three fixed viewpoints are selected in the current scene for data sampling, including a frontal view, a left-front tilted view (offset angle +30°), and a right-rear-downward view (offset angle -45°). Image feature extraction is performed on the anchor point's projection position, occlusion level, texture clarity, and contour boundary in each of the three viewpoints. For example, in the frontal view, the elbow anchor point has a texture clarity of 98% and a boundary integrity of 95%; in the left-front view, the texture clarity drops to 76%, the boundary is bent, and it deviates from its preset skeleton line segment; while in the right-rear-downward view, the anchor point's position shifts upward, deviating from the preset elbow contour trajectory by more than 8 pixels. At this point, it is preliminarily determined that the anchor point has a view consistency problem. Further analysis of the angle between this anchor point and the skeletal segment formed by the shoulder and wrist anchor points revealed that the angle should be 120° but was measured at 163°, indicating a conflict between its spatial positioning and skeletal logic. This confirmed that the anchor point was the source of the mismatch. Through multi-view morphological comparison and skeletal topological angle verification, it was ensured that the identified mismatch source point not only exhibited visual deviation but also constituted a semantic error at the structural logic level.
[0036] After confirming the source of the mismatch, semantic remapping is performed on its local area, which includes three operations: spatial position correction, texture region rebinding, and layer priority adjustment. First, the average coordinates of the anchor point in frames 1050, 1051, and 1052 are extracted from the three perspectives, assumed to be (127.8, 59.3, 40.1), and the coordinates are corrected based on this. If the anchor point position in the current frame is (129, 58, 40), it is adjusted to (127, 59, 40), and the boundary contour line segments are supplemented by spatial interpolation to ensure that it forms a continuous closed structure on the surface of the 3D model. Second, the texture belonging region of the anchor point is re-evaluated. The image patch with the highest matching degree with the position is selected through texture gradient analysis. For example, among the five available sleeve patterns, the 32nd patch with the most consistent texture direction and the closest average color value is selected as the new binding region, and the brightness alignment and deformation compensation of the projection parameters are completed. Finally, the layer priority of the anchor point is reordered, adjusting the layer order value according to its depth position and spatial centroid, for example, promoting it from layer 2 to layer 1, so that it maintains the correct occlusion relationship in the foreground layer of the character. After completing these three operations, a new mapping snapshot is immediately generated to replace the old state of the anchor point in the original snapshot, and marked as "local remapping completed", providing stable input for behavior prediction in subsequent frames.
[0037] The remapping status of anchor points is written into the traceability chain structure, and local consistency verification is performed in the current frame. This traceability chain records the vector difference between the original and corrected coordinates, the original texture number and the new bound texture number, the reference frame number used for remapping, and details of layer priority changes. For example, the traceability chain entry for the left elbow anchor point is: "Original coordinates (129, 58, 40), new coordinates (127, 59, 40), texture changed from 27 to 32, reference frames 1050~1052, layer order changed from 2 to 1." Furthermore, in the current frame, the angle verification of the two skeleton anchor points (shoulder and wrist) connected to this anchor point needs to be performed. If the connection angles at both ends conform to the set topological angle range (e.g., 108° to 132°), the entire skeleton segment is marked as "reconstruction successful"; otherwise, the remapping of other related anchor points needs to continue. Finally, an integrity check is performed on all anchor points that have completed remapping in this frame to ensure there are no texture breaks, occlusion anomalies, or binding detachments, confirming that the local area mapping structure has been restored to stability.
[0038] After the local semantic remapping is completed, an adaptive adversarial verification loop driven by human-computer resonance is constructed based on the user's gaze trajectory and interactive gestures. The semantic weights are adjusted in real time, and the label state is dynamically controlled through the label freezing and unfreezing mechanism to form a semantically stable closed loop in the cross-scene switching process. After completing local semantic remapping, it is necessary to respond in real time to changes in the user's attention area. Based on the user's gaze trajectory and hand movements as input, a semantic state adversarial verification closed loop driven by human-computer resonance is constructed. By adjusting semantic weights and implementing label freezing and unfreezing control, semantic consistency and response accuracy during cross-scene switching are ensured. This process can include the following steps in sequence: An interaction intent map is constructed by collecting user gaze behavior and hand gesture data to identify the user's main focus areas and their operational intentions in the current 3D scene. The specific data collection method is as follows: The user's eye-tracking device is synchronized pixel-level with the 3D display content. Each frame records the coordinates of the gaze point in 3D space as a 3D vector (X, Y, Z) and the duration of gaze at that location. If the dwell time at that location exceeds 700 milliseconds for three consecutive frames, the system marks it as the user's current main focus area. Simultaneously, the system collects the user's hand gesture trajectories in space, such as raising, pointing, pinching, and rotating. Each gesture is converted into a spatial motion path and gesture recognition label using an optical calibration device. For example, if a user pinches while gazing at a character's head anchor point, it indicates a desire to zoom in on that area; if they press down, it indicates an intention to zoom out on the displayed content. Based on this, the coordinate overlap analysis of the user's gaze point and the gesture's point of action is performed. When the overlapping area exceeds 80% projection overlap, it is confirmed as the current main interaction focus area. This interactive focus will become the core target area for subsequent semantic weight calculation and tag state regulation, providing accurate behavioral basis for the dynamic scheduling of semantic tags.
[0039] Based on the identified main interactive focus area, a real-time semantic weight assignment process is initiated to calculate the degree of influence of user behavior on each semantic anchor point, forming a hierarchical semantic response structure. First, the position, layer information, tag category, and currently bound texture number of all semantic anchor points in the current frame are extracted, and their spatial distance to the main interactive focus area is calculated. Anchor points with a distance of less than 15 pixels are designated as Level 1 anchor points and assigned the highest weight; those with a distance between 15 and 45 pixels are designated as Level 2 anchor points and assigned medium weight; the remaining anchor points are classified as Level 3 anchor points and assigned the lowest weight. Taking the left eye area as an example, if the user's gaze center is located on the surface of the left eyeball and their finger is pointing towards the left eyebrow area, then the eyeball anchor point and the brow bone anchor point are classified as Level 1 weight anchor points with a weight factor of 1.0, indicating the highest semantic stability priority in the current frame. The weight of each anchor point not only affects its refresh rate and texture loading priority during rendering but also directly determines its tag inheritance priority during scene transitions. The calculation and updating of semantic weights are performed in real time in each frame, and the status is recorded in conjunction with the frame sequence index number to support the subsequent tag freezing judgment logic.
[0040] After the weight structure is established, a label freeze operation is performed on semantic anchors whose weight factors exceed a set threshold to prevent label misalignment, layer degradation, or texture jumps during multi-scene transitions, viewpoint changes, or background interference. The freezing mechanism is specifically divided as follows: Full freeze applies to primary weight anchors, such as key interactive parts like eyeballs, corners of the mouth, and fingertips. Fully frozen anchors cannot change their bound labels, spatial coordinates, texture styles, or layer numbers in the next five frames, even if the background environment or the user's viewing angle changes; Flexible freeze applies to secondary weight anchors, such as facial contours, shoulders, and clothing edges. It allows texture updates and brightness adjustments within the original layer, but cannot be transferred to other layers or have its label category changed. Once the freeze state is set, a freeze timestamp and thawing trigger conditions are assigned to the anchor, such as "eyes leaving for more than 1 second" or "gesture direction reversed 180 degrees." In addition, each freeze action is synchronously recorded in the anchor's semantic structure and affects its dynamic attribute response capability in subsequent frames. The freezing mechanism ensures that the core tags within the user's focus area remain unchanged during periods of high interaction intensity, thereby maintaining the stability of the semantic scene and the immersive experience.
[0041] After the freeze state is set and executed, a defreezing mechanism is triggered based on changes in user behavior. Simultaneously, all state changes during the freeze-defreeze cycle are incorporated into the semantic tag management chain structure, forming a multi-frame, traceable, and controllable semantic closed-loop structure. Specifically, in each frame, the system determines whether the defreezing conditions are met, including a gaze offset exceeding a set angle (e.g., 25°), a gaze duration of less than 300 milliseconds, and a gesture target point being more than 50 pixels away from the original anchor point. Once the defreezing conditions are met, the anchor point will perform a tag state restoration operation, regaining its dynamic update capability and re-entering the semantic weight evaluation process. This freeze-defreeze control flow is not only applicable to anchor point behavior control within a single scene but also supports a cross-scene inheritance mechanism. For example, if a user freezes the character's eye anchor point in exhibition hall A and then jumps to exhibition hall B, the anchor point remains frozen until the user explicitly switches their focus or cancels the interaction. To ensure the complete preservation of frozen data, the system generates a freeze record chain for each anchor point, including the freeze start frame number, freeze duration in frames, thaw trigger event identifier, and fields affected by the freeze state (such as texture locking, layer preservation, and tag locking). During cross-scene transitions, the freeze chain structure of each anchor point is actively activated to inherit the semantic weights and freeze parameters established in the previous scene, thereby achieving consistent continuity between semantic behavior and user intent. Through this mechanism, the entire traditional art-based 3D content maintains a stable tag structure, consistent rendering performance, and natural user interaction response in scenes with rapid transitions and frequent interactions, thus constructing a complete, continuous, and disturbance-resistant semantically stable closed loop.
[0042] This invention achieves precise correspondence between semantic tags and 3D spatial structures by constructing verifiable semantic anchor signatures and spatiotemporal constraint baselines. The introduction of millisecond-level warm-up frames and the generation of phase references ensure temporal coherence and responsiveness during semantic mapping. Layered semantic feature binding and a mapping snapshot tracing mechanism ensure logical isolation and visual consistency between character features and background materials. Cross-frame tolerance corridors and texture bounce technology enable highly sensitive correction of dynamic offsets. Residual analysis and skeleton topological constraints achieve accurate identification of mismatched source points and local semantic remapping, effectively repairing visual disturbances. Finally, by combining a resonance feedback mechanism of user gaze trajectory and interactive gestures, an adaptive adversarial verification closed loop is established, dynamically adjusting semantic weights and achieving stable inheritance and seamless transition of semantic tags across multiple scenarios through a freeze-thaw mechanism. The overall solution not only enhances the artistic fidelity and spatial expressiveness of traditional 3D art conversion but also significantly improves the stability and continuity of immersive user interaction in virtual exhibitions.
[0043] The foregoing has only described certain exemplary embodiments of the present invention by way of illustration. Undoubtedly, those skilled in the art can modify the described embodiments in various ways without departing from the spirit and scope of the present invention. Therefore, the foregoing drawings and descriptions are illustrative in nature and should not be construed as limiting the scope of protection of the claims of the present invention.
Claims
1. A method for converting traditional art content into 3D based on artificial intelligence, characterized in that: Includes the following steps: Construct a cross-scene semantic anchor point map to generate verifiable semantic anchor point signatures for character features and background materials in traditional art images, and simultaneously establish a spatiotemporal constraint baseline based on spatial coordinates and time series. Before scene switching, a millisecond-level warm-up frame is loaded based on the spatiotemporal constraint baseline, and temporal calibration is performed on the semantic anchor signature to output a phase reference for controlling the subsequent mapping process; Based on phase reference, perform hierarchical binding of semantic features, bind character features to the foreground layer, bind background material to the environment layer, and generate a mapping snapshot of the current semantic state; A cross-frame tolerance corridor is established based on the mapping snapshot. Deviation detection is performed by comparing the actual position of the semantic target with the snapshot record frame by frame. If the detected deviation exceeds the tolerance range, the texture bounce mechanism is triggered to correct the mapping offset and maintain the stability of the layered binding. For semantic anomalies remaining after bias correction, semantic consistency adjudication is performed based on residual analysis. Mismatch source points are identified through multi-view comparison and skeleton topological constraints, and local semantic remapping is performed after identification. After local semantic remapping is completed, an adaptive adversarial verification loop driven by human-computer resonance is constructed based on the user's gaze trajectory and interactive gestures. Semantic weights are adjusted in real time, and the label state is dynamically controlled through a label freezing and unfreezing mechanism.
2. The method for converting traditional art content into 3D based on artificial intelligence according to claim 1, characterized in that, The steps for establishing a spatiotemporal constrained baseline based on spatial coordinates and time series are as follows: Frame by frame, traditional art images are analyzed to extract significant expressive areas of human features and background materials, and human features and background materials are semantically labeled and semantically encoded respectively. Based on spatial location, boundary contour, feature texture, color level and semantic encoding, semantic anchor signatures of human features and background material are generated, and the semantic anchor signatures are bound to the spatial location of the image. By combining semantic anchor signatures with time series and three-dimensional spatial coordinates, the positional offset, morphological changes and spatial trajectory of semantic anchors in the frame sequence are constructed to form a spatiotemporal constraint baseline. The semantic anchor graph, semantic anchor signature, and spatiotemporal constraint baseline are uniformly organized to establish a semantic anchor graph set for multi-scenario loading, which is used for semantic label consistency mapping and spatial position correction during scene switching.
3. The method for converting traditional art content into 3D based on artificial intelligence according to claim 2, characterized in that, The phase reference output steps used to control the subsequent mapping process are as follows: High frame rate image sequences with a loading time range of 15 to 40 milliseconds are used as warm-up frames to extract semantic anchor point state data of character features and background materials; The constructed spatiotemporal constraint baseline is used to perform frame-level comparison of each semantic anchor point in the preheating frame, and to perform temporal series calibration of positional deviation, morphological changes and texture direction. Based on the motion trend, edge frequency, texture gradient and spatial displacement of each anchor point, state phase parameters are generated, and a stable phase reference for each anchor point in the time window is output. Based on the stable phase reference, the spatial location, layer number, depth sorting and label weight of the semantic anchor in the new scene are initialized and configured to complete the mapping preparation before scene switching.
4. The method for converting traditional art content into 3D based on artificial intelligence according to claim 3, characterized in that, The steps for generating a snapshot of the current semantic state are as follows: Based on the dynamic trajectory and spatial behavior trend of semantic anchors in the warm-up frame, and combined with phase reference values, layer positioning preparation is carried out to distinguish between foreground layer and environment layer to bind candidate anchors. Based on the layer positioning results, perform a layered binding operation of semantic anchor points, bind the character features to the foreground layer, bind the background material to the environment layer, and set the layer number and layer sort number; Based on the bound semantic anchor state, generate a semantic state mapping snapshot containing 3D coordinates, layer number, sort number, texture index, phase reference value and timestamp; The current mapping snapshot is combined with the previous frame snapshot to construct a time-series snapshot chain, recording the state change trajectory of each anchor point, forming a dynamic data chain that supports semantic state tracing.
5. The method for converting traditional art content into 3D based on artificial intelligence according to claim 4, characterized in that, The phase reference value in the semantic state mapping snapshot is used to control the rendering priority of anchor points in the layer, and the layer sort number of each anchor point is dynamically adjusted according to the stability of the phase reference value to enhance the visual hierarchy stability of character features and background materials during multi-scene switching.
6. The method for converting traditional art content into 3D based on artificial intelligence according to claim 4, characterized in that, If the detection deviation exceeds the tolerance range, the texture bounce mechanism is triggered to correct the mapping offset and maintain the stability of the layered binding. The steps are as follows: Based on the three-dimensional spatial position and phase reference value of the semantic anchor point in the previous frame's mapping snapshot, a spatial tolerance corridor with defined three-axis boundaries is constructed. Extract the actual 3D position of the semantic anchor point in the current frame and compare it with the tolerance corridor axis by axis difference to determine whether there is a deviation exceeding the limit. For anchor points with excessive deviation, the bounce path is calculated based on three-frame snapshot data, and a texture bounce correction operation is performed. The corrected anchor state is written to the current frame mapping snapshot and updated to the cross-frame tracing chain to achieve dynamic stability of the semantic state.
7. The method for converting traditional art content into 3D based on artificial intelligence according to claim 6, characterized in that, The steps for local semantic remapping are as follows: Based on the anchor point state after texture bounce correction, perform residual analysis to determine whether the differences between the three-dimensional coordinates, texture binding area and phase reference value exceed the stability threshold. For semantic anchors that are determined to be abnormal, the semantic anchors that constitute mismatch source points are identified by comparing multi-view image features and calculating the angle between the skeleton topology. Perform local semantic remapping operations on the spatial location, texture attribution, and layer priority of the identified mismatch source points to generate a new mapping snapshot; The remapped anchor state is written into the traceability chain and local consistency verification is performed to ensure structural integrity and semantic binding continuity.
8. The method for converting traditional art content into 3D based on artificial intelligence according to claim 7, characterized in that, After the local semantic remapping is completed, the following steps are performed to construct a semantically stable closed loop: Collect user eye movements and interactive gestures to identify the main interaction focus area; Calculate the weight level of semantic anchor points based on the main interactive focus area, and establish a semantic response structure; For semantic anchors whose weight factors exceed a set threshold, a freeze operation is performed, and the freeze timestamp and unfreeze conditions are set. Unfreeze the system when there is a gaze shift, insufficient gaze duration, or a change in interaction action. Record the change in frozen state information into the semantic tag management chain to achieve continuity and stability control of semantic behavior.