Intelligent method and system for generating intangible cultural heritage collection model based on gan style constraint mechanism
By using a GAN-based method for generating intangible cultural heritage (ICH) collection models, combined with generative adversarial networks (GANs) and augmented reality (AR) technologies, the problems of insufficient style control and disconnect between ICH collection models and user profiles were solved. This approach enables precise control of artistic style and personalized output, thereby improving the efficiency of digital preservation and the user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-11
- Publication Date
- 2026-03-31
AI Technical Summary
In the digital preservation of intangible cultural heritage, existing technologies lack refined style control in the generated models, user profiles are disconnected from model generation, augmented reality displays are not intelligent enough, and it is difficult to achieve personalization and dynamic interaction.
A method for generating intangible cultural heritage collection models based on GAN style constraint mechanism is adopted. By receiving the original data of intangible cultural heritage collections and user behavior information, user profiles are generated. Combined with generative adversarial networks and augmented reality technology, precise control of artistic style and personalized output can be achieved.
It has achieved precise control over the artistic style of intangible cultural heritage collection models, improved the real-time linkage between user profiles and model generation, enhanced the dynamic interactive capabilities of augmented reality displays, and improved the efficiency of digital modeling and user experience.
Smart Images

Figure CN121095469B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of digital preservation and display technology of intangible cultural heritage, and in particular to a method and system for generating intangible cultural heritage collection models based on GAN style constraint mechanism. Background Technology
[0002] The digital preservation and display of intangible cultural heritage is an important research direction in the field of cultural transmission. Current technologies face the following bottlenecks: In digital modeling, traditional 3D scanning and modeling software relies on manual operation, resulting in cumbersome and inefficient processes that fail to meet the demands of large-scale digitization of intangible cultural heritage collections. While generative adversarial networks (GANs) have been applied to content generation, existing solutions lack effective style constraint mechanisms, making it difficult for generated intangible cultural heritage models to accurately reflect specific cultural style characteristics. Regarding user personalization technology, existing user profiling systems are mostly used in e-commerce recommendation scenarios; their static update mechanisms cannot reflect real-time changes in user behavior and are not deeply integrated with the model generation process. Although augmented reality (AR) display technology can achieve the fusion of virtual and real digital content, existing solutions offer limited content and lack dynamic interactive capabilities based on real-time user behavior.
[0003] The core problems with existing technologies include: the lack of a refined style control mechanism in the generated models, making it difficult to guarantee the accuracy of the artistic style of intangible cultural heritage collections; the disconnect between user profile data and the model generation process, making it impossible to achieve truly personalized output; and the insufficient intelligence of augmented reality display systems, making it difficult to adjust the displayed content based on real-time user behavior. These problems severely restrict the efficiency and display effect of digital protection of intangible cultural heritage, and there is an urgent need for a comprehensive solution that can integrate style constraints, behavioral analysis, and dynamic display.
[0004] To address the aforementioned issues, existing technologies urgently need improvement. Summary of the Invention
[0005] The purpose of this application is to provide a method for generating intangible cultural heritage collection models and an augmented reality display system based on the GAN style constraint mechanism. It has the advantages of achieving precise control of the artistic style of intangible cultural heritage collections, improving the real-time linkage between user profiles and model generation, and enhancing the dynamic interactive capabilities of augmented reality display.
[0006] Firstly, the method for generating intangible cultural heritage collection models based on GAN style constraint mechanism provided in this application adopts the following technical solution:
[0007] A method for generating intangible cultural heritage collection models based on GAN style constraint mechanism includes:
[0008] Receive basic raw data of intangible cultural heritage collections, including image materials and text descriptions;
[0009] Obtain user behavior information, which includes user interaction events, browsing history, or operation records on the digital platform;
[0010] Generating user profile data based on user behavior information includes: extracting key features from the behavior information and constructing a behavior feature vector; calculating feature values, including a first feature value and an average feature value, to quantify behavior patterns; and establishing distinctive labels, which are generated based on multi-dimensional feature information classification.
[0011] Integrate user profile data with basic raw data as input for generating intangible cultural heritage collection models;
[0012] A generative adversarial network is used to generate a 3D model of an intangible cultural heritage collection. The generator receives style constraint parameters, which are used to generate a style encoding vector through contrastive learning. The style encoding vector is used to control the artistic style of the intangible cultural heritage collection model.
[0013] The discriminator optimizes model realism based on adversarial training, and the loss function includes content preservation and style alignment terms.
[0014] The generation weights are adjusted by using feature values from user profile data to optimize the model generation process;
[0015] Use augmented reality exhibition applications to create augmented reality markers and unique identifiers for exhibition targets in real-world scenes;
[0016] When a user's captured image of the current environment matches a pre-stored image of a real scene, the model of the intangible cultural heritage artifact is displayed at the exhibition target location by recognizing the unique identifier of the augmented reality marker.
[0017] Optionally, generating user profile data based on user behavior information includes:
[0018] Acquire user behavior information and preprocess the behavior information to extract key features, including click frequency, session duration or interaction depth;
[0019] Construct behavioral feature vectors and use a pre-trained encoder to transform key features into numerical vector representations. The encoder is based on a neural network model to ensure consistent vector dimensions.
[0020] The eigenvalues are calculated, including a first eigenvalue and an average eigenvalue. The first eigenvalue is obtained by calculating the vector magnitude and reflects the overall strength of the behavior. The average eigenvalue is calculated by calculating the vector average value and measures the stability of the behavior.
[0021] Establish iconic labels, classify feature values based on clustering algorithms, generate labels, and optimize label accuracy through weight allocation.
[0022] Optionally, the generation of the style constraint parameters is achieved through contrastive learning, including:
[0023] Collect target style samples and negative samples, wherein the target style samples include intangible cultural heritage art reference images or text descriptions;
[0024] Embedsion vectors of samples are extracted using a style encoder, which is built on a convolutional neural network;
[0025] By comparing the loss function, the similarity between positive and negative samples is calculated, which brings samples of the same style closer together and widens the distance between samples of different styles.
[0026] A style encoding vector is generated, which quantifies artistic style attributes and is used to initialize the input layer of the generator.
[0027] Optionally, the training of the generative adversarial network includes:
[0028] The generator receives style-encoded vectors and ensemble data as input and outputs a 3D model. The generator structure includes an encoder-decoder architecture, where the encoder processes the fusion features of image materials and text descriptions, and the decoder generates model point cloud data.
[0029] The discriminator optimizes the model's realism based on adversarial training. In addition to content preservation and style alignment terms, the loss function also includes a spatiotemporal consistency term to ensure the model's stability in dynamic scenarios.
[0030] The training process employs an alternating optimization strategy, iteratively updating the generator and discriminator parameters until the model converges to the Nash equilibrium point.
[0031] Optionally, adjusting the generated weights based on the feature values of the user profile data includes:
[0032] After calculating the feature values of the user behavior feature vector, the feature values are normalized to ensure that the values are within a preset range.
[0033] A weighted summation algorithm is used to map feature values to the input layer of the generator, where the weight coefficients are dynamically adjusted based on the iconic labels;
[0034] The adjustment process includes real-time monitoring of feature value changes and recalculating trigger weights to maintain generation stability when behavioral patterns change abruptly.
[0035] After integrating and adjusting the results, weights are generated to control model complexity or level of detail, thereby improving the quality of personalized output.
[0036] Optionally, creating augmented reality tags and unique identifiers includes:
[0037] Identify the environmental location information of real-world scenes and obtain latitude and longitude coordinates and 3D point cloud data through GPS sensors or visual positioning systems;
[0038] Assign a unique identifier to each exhibition target, and use a hash algorithm to generate the identifier to ensure global uniqueness and tamper-proofness;
[0039] An environmental location association between augmented reality markers and intangible cultural heritage collection models is established. This association is stored in a distributed file system network module and a blockchain module is used to store summary information to ensure security.
[0040] The marker deployment supports multiple forms, including QR codes, image markers, or virtual anchors, to adapt to different exhibition scenarios.
[0041] Optionally, the matching of the current environment image with the pre-stored real scene image includes:
[0042] Extract the location information and environmental information of the current environment image. The location information includes GPS data of the shooting device, and the environmental information includes lighting conditions and object outline features.
[0043] Real-time user location is performed, and the location results are verified to be consistent with the information of the current environment image. Verification methods include feature point matching or deep learning model comparison.
[0044] When the consistency verification passes, the pre-stored real scene image is retrieved from the distributed file system network module. The pre-stored image includes a high-resolution panoramic image or a three-dimensional scene model.
[0045] The matching process employs a multi-scale feature extraction algorithm to ensure robustness under different viewpoints and lighting conditions.
[0046] Secondly, this application provides a system for generating intangible cultural heritage collection models based on a GAN style constraint mechanism, including:
[0047] The data acquisition module is used to receive basic raw data of intangible cultural heritage collections, including image materials and text descriptions;
[0048] The behavior information acquisition module is used to acquire user behavior information, which includes user interaction events, browsing history, or operation records on the digital platform.
[0049] The profile generation module is used to generate user profile data based on user behavior information, including: extracting key features from the behavior information and constructing a behavior feature vector; calculating feature values, which include a first feature value and an average feature value, to quantify behavior patterns; and establishing distinctive labels, which are generated based on multi-dimensional feature information classification.
[0050] The integration module is used to integrate user profile data with basic raw data as input for generating intangible cultural heritage collection models;
[0051] The model generation module is used to generate a 3D model of intangible cultural heritage collections using a generative adversarial network. The generator receives style constraint parameters, which generate style encoding vectors through contrastive learning. The style encoding vectors are used to control the artistic style of the intangible cultural heritage collection model.
[0052] The discriminator module is used to optimize the model's realism based on adversarial training. The loss function includes a content preservation term and a style alignment term.
[0053] The optimization module is used to adjust the generation weights based on the feature values of user profile data in order to optimize the model generation process;
[0054] The representation module is used to create augmented reality tags and unique identifiers for exhibition targets in real-world scenes using augmented reality exhibition applications;
[0055] The output module is used to display models of intangible cultural heritage artifacts at the exhibition target location by recognizing the unique identifier of the augmented reality marker when the current environment image captured by the user matches a pre-stored real scene image.
[0056] Thirdly, this application provides a computer device, the device comprising: a memory and a processor, wherein the processor, when executing computer instructions stored in the memory, performs the method described above.
[0057] Fourthly, this application provides a computer-readable storage medium including instructions that, when executed on a computer, cause the computer to perform the method described above.
[0058] In summary, this application integrates user profile data with intangible cultural heritage basic data to drive the generation of adversarial networks, combines style constraint mechanisms to achieve precise control of artistic style, and utilizes augmented reality technology to achieve dynamic scene matching and display. It has the advantages of improving the efficiency of digital modeling of intangible cultural heritage, enhancing user interaction experience, and ensuring the accuracy of cultural inheritance. Attached Figure Description
[0059] Figure 1 This is a schematic diagram of the computer device structure of the hardware operating environment involved in the embodiments of this application;
[0060] Figure 2 This is a flowchart illustrating the first embodiment of the intangible cultural heritage collection model generation method based on the GAN style constraint mechanism of this application;
[0061] Figure 3 This is a structural block diagram of the first embodiment of the intangible cultural heritage collection model generation system based on the GAN style constraint mechanism of this application. Detailed Implementation
[0062] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0063] Reference Figure 1 , Figure 1 This is a schematic diagram of the computer device structure of the hardware operating environment involved in the embodiments of this application.
[0064] like Figure 1 As shown, the computer device may include: a processor 1001, such as a central processing unit (CPU), a communication bus 1002, a user interface 1003, a network interface 1004, and a memory 1005. The communication bus 1002 is used to enable communication between these components. The user interface 1003 may include a display screen or an input unit such as a keyboard; optionally, the user interface 1003 may also include a standard wired interface or a wireless interface. The network interface 1004 may optionally include a standard wired interface or a wireless interface (such as a Wireless-Fidelity (Wi-Fi) interface). The memory 1005 may be high-speed random access memory (RAM) or stable non-volatile memory (NVM), such as a disk drive. The memory 1005 may also optionally be a storage device independent of the aforementioned processor 1001.
[0065] Those skilled in the art will understand that Figure 1 The structure shown does not constitute a limitation on the computer device and may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0066] like Figure 1 As shown, the memory 1005, which serves as a storage medium, may include an operating system, a network communication module, a user interface module, and a non-heritage collection model generation program based on a GAN-style constraint mechanism.
[0067] exist Figure 1In the computer device shown, the network interface 1004 is mainly used for data communication with the network server; the user interface 1003 is mainly used for data interaction with the user; the processor 1001 and the memory 1005 in this application can be set in the computer device. The computer device calls the intangible cultural heritage collection model generation program based on the GAN style constraint mechanism stored in the memory 1005 through the processor 1001, and executes the intangible cultural heritage collection model generation method based on the GAN style constraint mechanism provided in the embodiment of this application.
[0068] This application provides a method for generating intangible cultural heritage collection models based on a GAN style constraint mechanism, referring to... Figure 2 , Figure 2 This is a flowchart illustrating the first embodiment of the intangible cultural heritage collection model generation method based on the GAN style constraint mechanism of this application.
[0069] In this embodiment, the method for generating intangible cultural heritage collection models based on the GAN style constraint mechanism includes the following steps:
[0070] Step S10: Receive basic raw data of intangible cultural heritage collections, including image materials and text descriptions;
[0071] Step S20: Obtain user behavior information, which includes user interaction events, browsing history, or operation records on the digital platform;
[0072] Step S30: Generate user profile data based on user behavior information, including: extracting key features from the behavior information and constructing a behavior feature vector; calculating feature values, the feature values including a first feature value and an average feature value, used to quantify behavior patterns; and establishing distinctive labels, the labels being generated based on multi-dimensional feature information classification.
[0073] Step S40: Integrate user profile data with basic raw data as input for generating intangible cultural heritage collection models;
[0074] Step S50: Generative adversarial network is used to generate a 3D model of intangible cultural heritage collection, wherein the generator receives style constraint parameters, the style constraint parameters generate a style encoding vector through contrastive learning, and the style encoding vector is used to control the artistic style of the intangible cultural heritage collection model;
[0075] Step S60: The discriminator optimizes the model's realism based on adversarial training, and the loss function includes a content preservation term and a style alignment term;
[0076] Step S70: Adjust the generation weights using the feature values of the user profile data to optimize the model generation process;
[0077] Step S80: Use the augmented reality exhibition application to create augmented reality tags and unique identifiers for exhibition targets in the real scene;
[0078] Step S90: When the current environment image captured by the user matches the pre-stored real scene image, the model of the intangible cultural heritage collection is displayed at the exhibition target location by identifying the unique identifier of the augmented reality marker.
[0079] In traditional, existing digitization systems for intangible cultural heritage artifacts, generative adversarial networks (GANs) suffer from insufficient alignment between encoded vectors and target styles in the art style control dimension, resulting in the generated models failing to accurately reflect the diversity of intangible cultural heritage. The low coupling between user profile data and the model generation process, coupled with the inability of static feature value update mechanisms to capture real-time changes in user behavior, leads to personalized outputs deviating from actual needs. Augmented reality (AR) display modules rely on fixed scene matching logic and lack dynamic environment awareness, causing a disconnect between model display and real-time user interaction.
[0080] For example, in convolutional neural network-based style encoder applications, target style samples are only used to generate embedding vectors through single-layer feature extraction without establishing cross-layer semantic associations. This leads to a mismatch between the texture details of the 3D model output by the generator and the characteristics of intangible cultural heritage traditional crafts. User behavior feature vectors are directly input into the generator after generating labels through offline clustering without a real-time weight adjustment mechanism. When the frequency of user interaction changes abruptly during a session, the complexity of the generated model deviates from the user's current preferences. Augmented reality tag deployment uses a static GPS coordinate matching strategy without integrating multi-scale feature extraction from a visual positioning system. When lighting conditions change or the viewing angle shifts by more than 15 degrees, the coordinates of the model's displayed position drift.
[0081] If the above problems are not addressed, the artistic style of the generated model will be unable to adapt to the technical characteristics of different intangible cultural heritage schools, reducing the cultural value of digital reproduction. Delays in updating user profile data will lead to inaccurate weight allocation in the generation process, causing abrupt changes in the model's detailed levels during continuous user interaction, disrupting the continuity of the user experience. Positional drift in augmented reality displays will cause misalignment between virtual and real elements, resulting in logical chaos in the exhibition space and ultimately affecting the effectiveness of the digital transmission of intangible cultural heritage.
[0082] To address the aforementioned issues, this embodiment first tackles the problem of insufficient style control in generative models by exploring how to establish a more precise style constraint mechanism. Traditional style encoders generate embedding vectors through single-layer feature extraction, resulting in a lack of cross-layer semantic association. To resolve this, this embodiment considers introducing a contrastive learning framework to optimize the generation process of style encoding vectors by calculating the similarity between positive and negative samples, ensuring the discriminative power of features from different intangible cultural heritage schools in the vector space.
[0083] To address the issue of low coupling between user profile data and the generation process, existing methods directly input offline clustering labels into the generator, failing to respond to real-time behavioral changes. This embodiment proposes designing a dynamic weight adjustment mechanism at the generator input layer. By monitoring feature value changes in real time and triggering weight recalculation, the complexity of the generated model remains synchronized with the user's current behavioral patterns.
[0084] To address the issue of positional drift in AR displays, traditional GPS coordinate matching strategies are prone to failure when the viewpoint shifts. This embodiment attempts to integrate a visual positioning system with a multi-scale feature extraction algorithm. By jointly analyzing ambient lighting conditions and object contour features, the robustness of scene matching is improved. At the same time, a hash algorithm is used to generate tamper-proof identifiers to ensure the stability of the relationship between virtual and real spaces.
[0085] To address this, this embodiment proposes a method for generating intangible cultural heritage (ICH) collection models based on a GAN style constraint mechanism. The method includes: receiving basic raw data of ICH collections, including image materials and text descriptions; acquiring user behavior information, including user interaction events, browsing history, or operation records on digital platforms; generating user profile data based on the user behavior information, including extracting key features from the behavior information, constructing a behavior feature vector, calculating feature values (including a first feature value and an average feature value) to quantify behavior patterns, establishing a distinctive label, and generating the label based on multi-dimensional feature information classification; and integrating the user profile data with the basic raw data as input for generating the ICH collection model. The process involves: using a generative adversarial network (GAN) to generate 3D models of intangible cultural heritage (ICH) artifacts. The generator receives style constraint parameters, which are then used to generate style encoding vectors through contrastive learning. These style encoding vectors control the artistic style of the ICH artifact models. A discriminator optimizes the model's realism based on adversarial training, with a loss function that includes content preservation and style alignment terms. The generation weights are adjusted using feature values from user profile data to optimize the model generation process. Augmented reality (AR) tags and unique identifiers are created for exhibition targets in real-world scenes using an AR exhibition application. When a user-captured image of the current environment matches a pre-stored image of the real-world scene, the ICH artifact model is displayed at the exhibition target location by recognizing the unique identifier of the AR tag.
[0086] The basic raw data refers to the image materials and text descriptions of intangible cultural heritage (ICH) artifacts. This can be achieved by capturing images with a high-resolution camera and combining them with manually annotated text descriptions, providing the visual and semantic information required for the generative model. User behavior information refers to user interaction events, browsing history, or operation records on digital platforms. This can be achieved by using log collection tools or event tracking technology to record user operation data in real time, capturing user preferences to drive personalized generation. User profile data refers to key feature vectors, feature values, and distinctive labels extracted based on user behavior information. This can be achieved by using a pre-trained encoder to convert behavioral data into numerical vectors and combining them with clustering algorithms to generate labels, quantifying user behavior patterns and dynamically adapting to the generation process. Generative Adversarial Networks (GANs) are deep learning models containing a generator and a discriminator. This can be achieved by using an encoder-decoder architecture to process the fusion features of images and text and generate 3D point cloud data, used to automatically construct digital models of ICH artifacts. Style constraint parameters refer to style encoding vectors generated through contrastive learning. This can be achieved by using a convolutional neural network to extract the embedding vectors of target style samples and combining them with a contrastive loss function to optimize distance relationships, controlling the alignment of the generative model's artistic style with the ICH cultural characteristics. Augmented reality markers and unique identifiers are virtual markers bound to real-world scenes. Specifically, they can be generated using hash algorithms and combined with GPS or visual positioning systems to establish environmental associations, triggering model displays in specific exhibition scenarios. Current environment image matching involves comparing user-captured images with pre-stored scene images, using multi-scale feature extraction algorithms combined with deep learning models for robustness verification, ensuring the accuracy and stability of the augmented reality display.
[0087] The core innovation of this embodiment lies in dynamically integrating user profile data into the input layer of the generative adversarial network, combining it with style encoding vectors generated by contrastive learning to achieve personalized adaptation of the artistic style of intangible cultural heritage collection models, and constructing a closed loop from data generation to scene display through augmented reality tagging and real-time environment matching technology.
[0088] The working process and principle of this embodiment are as follows: First, basic raw data of intangible cultural heritage collections is received, including image materials and text descriptions. User interaction events, browsing history, or operation records on digital platforms are acquired. User profile data is generated based on this behavioral information, key features are extracted to construct behavioral feature vectors, and the first feature value and average feature value are calculated to quantify behavioral patterns. Iconic labels based on multi-dimensional feature information classification are established. The user profile data is integrated with the basic raw data as input to generate the intangible cultural heritage collection model.
[0089] A generative adversarial network (GAN) is used to generate 3D models of intangible cultural heritage artifacts. The generator receives style encoding vectors generated through contrastive learning as style constraint parameters to control the model's artistic style. The discriminator optimizes the model's realism through adversarial training, and the loss function includes content preservation and style alignment terms. The generation weights are adjusted using feature values from user profile data to optimize the model generation process.
[0090] Use an augmented reality exhibition application to create augmented reality markers and unique identifiers for exhibition targets in real-world scenes. When a user-captured image of the current environment matches a pre-stored image of the real-world scene, a model of an intangible cultural heritage artifact is displayed at the exhibition target by recognizing the unique identifier of the augmented reality marker.
[0091] This solution integrates user profile data, style constraint parameters, and augmented reality display technology to achieve personalized generation and interactive display of intangible cultural heritage collection models. Style encoding vectors ensure that the model's artistic style matches the characteristics of intangible cultural heritage, dynamic integration of user profile data enhances personalization, and augmented reality display technology provides an immersive experience.
[0092] In practice, the system receives basic raw data from intangible cultural heritage collections, including high-resolution image materials and detailed text descriptions. It also acquires user behavior information such as clicks, browsing, and favorites on digital platforms. Key features such as click frequency, session duration, and interaction depth are extracted from this behavioral information to construct behavioral feature vectors. Feature values are calculated: the first feature value is obtained through vector magnitude, reflecting the overall intensity of the behavior; the average feature value is obtained through vector average, measuring the stability of the behavior. Finally, a clustering algorithm is used to classify the feature values and generate distinctive labels.
[0093] This study integrates user profile data and basic raw data as input. A generative adversarial network (GAN) is used to generate 3D models of intangible cultural heritage artifacts. The generator receives style constraint parameters and generates style encoding vectors through a contrastive learning framework. Target style samples and negative samples are collected, and a convolutional neural network-based style encoder extracts the embedding vectors of the samples. The similarity between positive and negative samples is calculated using a contrastive loss function to optimize the generation of style encoding vectors.
[0094] The discriminator optimizes model realism through adversarial training, and its loss function includes content preservation and style alignment terms. Feature values from user profile data are normalized and then mapped to the generator input layer using a weighted summation algorithm, dynamically adjusting the generation weights.
[0095] Augmented reality (AR) tags and unique identifiers are created using an AR exhibition application. Environmental location information of the exhibition targets is acquired via GPS sensors or a visual positioning system. Unique identifiers are generated using a hash algorithm. An association is established between the AR tags and the environmental location of the intangible cultural heritage artifact models, and stored in a distributed file system network module.
[0096] When a user captures an image of the current environment, location and environmental information are extracted. Real-time positioning is performed and its consistency with the current environment image information is verified. Upon successful verification, a pre-stored real-scene image is retrieved from a distributed file system. A multi-scale feature extraction algorithm is used for matching to ensure robustness under different viewing angles and lighting conditions. After successful matching, a model of the intangible cultural heritage artifact is displayed at the exhibition target location by identifying the unique identifier of the augmented reality marker.
[0097] Through the above-described scheme, this embodiment achieves precise style control and personalized generation of intangible cultural heritage (ICH) collection models. The style constraint mechanism ensures a high degree of match between the artistic style of the generated model and the characteristics of ICH culture, enhancing the cultural value of digital reproduction. The dynamic integration of user profile data and the real-time weight adjustment mechanism enable the generated model to accurately adapt to the user's current preferences, improving the quality of personalized output. Improvements in augmented reality display technology solve the problem of model display position drift, improving the accuracy of virtual-real fusion and providing users with an immersive ICH collection display experience. These technological innovations collectively promote the effectiveness of digital ICH transmission and enhance the user experience.
[0098] In some of the solutions described above in this embodiment, there are problems such as inaccurate extraction of key features, a single method of feature value quantification, and insufficient accuracy of label classification during the user profile generation process. This results in user behavior patterns not being effectively converted into personalized parameters of the generated model, affecting the adaptability of the intangible cultural heritage collection model.
[0099] This embodiment further proposes a method for generating user profile data based on user behavior information, including: acquiring user behavior information and preprocessing the behavior information to extract key features, including click frequency, session duration, or interaction depth; constructing behavior feature vectors, using a pre-trained encoder to convert key features into numerical vector representations, the encoder being implemented based on a neural network model to ensure consistent vector dimensions; calculating feature values, including a first feature value and an average feature value, the first feature value being obtained through vector magnitude calculation to reflect the overall intensity of the behavior, and the average feature value being obtained through vector average calculation to measure the stability of the behavior; establishing iconic labels, classifying feature values based on a clustering algorithm, generating labels, and optimizing label accuracy through weight allocation.
[0100] The preprocessing of behavioral information employs data cleaning and feature selection techniques, extracting key features by filtering noisy data and redundant information. The pre-trained encoder uses a multilayer perceptron structure; the input layer receives standardized feature data, the hidden layer performs a non-linear transformation using the ReLU activation function, and the output layer generates a fixed-dimensional vector. The first eigenvalue is calculated using the Euclidean norm formula, and the average eigenvalue is calculated using the arithmetic mean of each dimension of the vector. The clustering algorithm uses the K-means algorithm, generating multiple cluster centers based on the eigenvalue distribution, with each cluster corresponding to a unique label. Weight allocation dynamically adjusts the label confidence by calculating the distance between samples within a cluster and the cluster center.
[0101] Specifically, after preprocessing, key features of user behavior information are input into a pre-trained encoder and transformed into 512-dimensional vectors. Consistency in vector dimensionality ensures comparability in subsequent calculations. The first feature value is a scalar value in the range of 0-1 obtained by calculating the vector magnitude, used to quantify user activity. The average feature value is a stability index obtained by calculating the mean of each dimension of the vector, with numerical fluctuations controlled within ±0.2. A clustering algorithm performs unsupervised classification of the feature values, generating two categories of labels: interest preferences and usage habits. Each label is assigned an initial weight coefficient of 0.5. During the weight allocation phase, the label weights are dynamically adjusted based on the temporal changes in user behavior data. When the feature value deviates from the historical average by more than 15% over three consecutive session cycles, a weight update mechanism is triggered to recalculate the matching degree between the label and the current behavior pattern, controlling the matching error to below 5%. Through this process, user profile data is dynamically updated, providing a basis for real-time parameter adjustment for the generated model.
[0102] In practice, user behavior information is acquired and preprocessed to extract key features. These key features include click frequency, session duration, and interaction depth. Click frequency is calculated by recording the number of clicks a user makes within a specific time period. Session duration is calculated from the time difference between each login and logout. Interaction depth is comprehensively evaluated based on the depth of the page the user navigates and the time spent on that page.
[0103] Construct behavioral feature vectors. Use a pre-trained encoder to transform key features into numerical vector representations. The encoder is implemented based on a multilayer perceptron neural network model. The input layer receives key feature data, the hidden layer uses the ReLU activation function, and the output layer generates vectors of fixed dimensions. Ensure consistent vector dimensions for easier subsequent processing.
[0104] The eigenvalues are calculated, including the first eigenvalue and the average eigenvalue. The first eigenvalue is obtained through the vector magnitude and reflects the overall strength of the behavior. Specifically, it is calculated by taking the square root of the sum of the squares of the vector components. The average eigenvalue is obtained through the vector average and measures the stability of the behavior. It is calculated by summing the vector components and then dividing by the number of components.
[0105] Establish key labels. Classify feature values using the K-means clustering algorithm. First, set the number of cluster centers, then iteratively calculate the distance from each sample to the cluster center until convergence. Generate labels and optimize label accuracy through weight allocation. The weight allocation uses the information gain method, calculating the contribution of each feature to the classification result and assigning higher weights to features with higher contributions.
[0106] Through the above technical solution, this embodiment achieves efficient processing and feature extraction of user behavior information. By constructing behavioral feature vectors, calculating feature values, and establishing distinctive labels, accurate user profile data is provided for the subsequent generation of personalized intangible cultural heritage collection models. This method improves the accuracy and real-time performance of user profiles, helps generate intangible cultural heritage collection models that better match user preferences, and enhances the user experience and the personalization of model generation.
[0107] In some of the above-mentioned schemes in this embodiment, when generating intangible cultural heritage collection models, the generative adversarial network lacks an effective multi-dimensional style alignment mechanism in the generation process of style encoding vectors, resulting in a mismatch between artistic styles and intangible cultural heritage characteristics, and failing to accurately reflect diversity.
[0108] This embodiment further proposes that the generation of style constraint parameters is achieved through contrastive learning, including: collecting target style samples and negative samples, where the target style samples include intangible cultural heritage art reference images or text descriptions; extracting the embedding vectors of the samples using a style encoder, which is constructed based on a convolutional neural network; calculating the similarity between positive and negative samples through a contrastive loss function, bringing samples of the same style closer together and distancing samples of different styles further apart; and generating style encoding vectors, which quantify the art style attributes and are used to initialize the input layer of the generator.
[0109] When collecting samples of the target style, the intangible cultural heritage art reference images need to cover different intangible cultural heritage schools or craft categories, and the text descriptions need to include keywords related to materials, patterns, or techniques. Negative samples are selected from unrelated art categories that are significantly different from the target style. The style encoder uses a pre-trained convolutional neural network, whose convolutional layers extract local texture and global structural features of the image, and fully connected layers output fixed-dimensional embedding vectors. The contrastive loss function uses cosine similarity to measure the distance between samples, optimizing the encoder parameters by maximizing the similarity of positive sample pairs and minimizing the similarity of negative sample pairs. The style encoding vector is normalized and mapped to the input layer of the generator, controlling the brushstroke, color, or shape features during the model generation process.
[0110] Specifically, during the generator initialization phase, the target style samples and negative samples are processed by the style encoder to obtain corresponding sets of embedding vectors. A contrastive loss function is used to calculate the similarity scores between positive and negative sample pairs. Gradient descent is then used to update the encoder parameters, ensuring that samples of the same style cluster together in the vector space, while samples of different styles are kept apart. After multiple rounds of iterative training, the embedding vectors output by the style encoder accurately represent the multi-dimensional attributes of intangible cultural heritage art. When the generator receives the style-encoded vectors, the decoder generates a 3D model with specific stylistic features based on the quantified style attributes in the vectors, such as the openwork structure of paper-cutting art or the texture details of ceramic glaze. Through a contrastive learning mechanism, the style-encoded vectors and the generator's parameter updates form a closed-loop feedback loop, ensuring that the artistic style of the generated model is strictly aligned with the input intangible cultural heritage features.
[0111] In practice, the generation of style constraint parameters is achieved through contrastive learning, including the following steps:
[0112] Collect target style samples and negative samples. Target style samples include reference images or text descriptions of intangible cultural heritage art. For example, collect 100 images of traditional blue and white porcelain as positive samples, and 100 images of modern ceramics as negative samples.
[0113] An embedding vector of the samples is extracted using a style encoder. The style encoder is built on a convolutional neural network. Specifically, a ResNet-50 is used as the backbone network, and the last fully connected layer outputs a 256-dimensional feature vector.
[0114] The similarity between positive and negative samples is calculated using a contrastive loss function, bringing samples of the same style closer together and distancing samples of different styles further apart. The contrastive loss function used is InfoNCE loss, with a temperature parameter set to 0.07. In each mini-batch, 16 pairs of positive samples and 48 pairs of negative samples are randomly selected for training.
[0115] A style encoding vector is generated, which quantifies artistic style attributes and is used to initialize the input layer of the generator. The trained style encoder is applied to the target intangible cultural heritage image to obtain a 256-dimensional style encoding vector. This vector is fed into the conditional normalization layer of the generator to control style information during the generation process.
[0116] Through the above technical solution, this embodiment achieves precise quantification and control of the artistic style of intangible cultural heritage collections. Consequently, the generated 3D model accurately captures and represents the unique artistic style of the intangible cultural heritage collections, avoiding style deviation and distortion. Furthermore, the contrastive learning method improves the discriminative ability of style encoding, enabling the generated model to distinguish subtle style differences between different intangible cultural heritage categories. This style constraint mechanism significantly enhances the authenticity and cultural value of the generated intangible cultural heritage collection model, providing strong support for the digital protection of intangible cultural heritage.
[0117] In some of the solutions described above in this embodiment, the 3D model generated by the generative adversarial network during training lacks spatiotemporal consistency in dynamic scenes, resulting in model distortion or positional shift during augmented reality display, which affects the user experience.
[0118] This embodiment further proposes that the training of a generative adversarial network includes a generator that receives style-encoded vectors and ensemble data as input and outputs a 3D model. The generator structure includes an encoder-decoder architecture, where the encoder processes the fusion features of image materials and text descriptions, and the decoder generates model point cloud data. The discriminator optimizes the model's realism based on adversarial training. In addition to content preservation and style alignment terms, the loss function also includes a spatiotemporal consistency term to ensure the model's stability in dynamic scenes. The training process adopts an alternating optimization strategy, iteratively updating the parameters of the generator and discriminator until the model converges to the Nash equilibrium point.
[0119] The encoder employs a multimodal fusion module to process image and text features, such as aligning visual and semantic information through a cross-modal attention mechanism, resulting in a fused feature vector dimension of 512. The decoder consists of a five-layer fully connected network with 256, 128, 64, 32, and 16 neurons per layer, outputting point cloud data containing 5000 3D coordinate points. The discriminator's spatiotemporal consistency term is achieved by calculating the model differences across consecutive time steps and the geometric consistency error after spatial transformation. For example, an optical flow estimation module is used to predict the motion trajectories of adjacent frames, constraining the point cloud displacement error to be less than 0.1 mm. An alternating optimization strategy sets the training step size ratio of the generator to the discriminator to be 3:1, with an initial learning rate of 0.0002, decaying by 10% every 100 iterations.
[0120] Specifically, the encoder performs convolutional operations on the image material to extract texture features, and simultaneously processes the text description using a bidirectional LSTM to obtain semantic features. These features are then concatenated and input into a fully connected layer to generate a 512-dimensional fusion vector. The decoder multiplies the fusion vector element-wise with the style encoding vector, and gradually reduces the dimensionality through a fully connected layer to generate point cloud data. The output layer uses the Tanh activation function to constrain the coordinate values within the range [-1, 1]. The discriminator receives both the generated and real point clouds, and introduces a spatiotemporal consistency loss term when calculating the adversarial loss. Specifically, it extracts multi-scale spatiotemporal features through a 3D convolutional network and calculates the root mean square value of the optical flow error between adjacent frame point clouds. When the error exceeds a threshold, a penalty term is added. During training, the discriminator updates its parameters every three times, and the Adam optimizer maintains the balance of the parameter update direction. The discriminator is considered to have reached Nash equilibrium when the discriminator's loss value fluctuates by less than 5% for ten consecutive times.
[0121] In practice, training a generative adversarial network includes the following steps:
[0122] The generator receives style-encoded vectors and ensemble data as input and outputs a 3D model. The generator structure employs an encoder-decoder architecture. The encoder processes the fused features of image materials and text descriptions, while the decoder generates the model's point cloud data. Specifically, the encoder consists of multiple layers of convolutional neural networks used to extract high-dimensional features from the input data. The decoder consists of deconvolutional layers and fully connected layers, progressively mapping the features back to 3D point cloud coordinates.
[0123] The discriminator optimizes model realism through adversarial training. In addition to content preservation and style alignment terms, the loss function includes a spatiotemporal consistency term to ensure model stability in dynamic scenes. The content preservation term uses L1 distance to measure the structural similarity between the generated model and the real model. The style alignment term uses a Gram matrix to calculate the matching degree of style features. The spatiotemporal consistency term constrains the continuity of the generation process by comparing model differences at adjacent time steps.
[0124] The training process employs an alternating optimization strategy, iteratively updating the generator and discriminator parameters until the model converges to a Nash equilibrium. Specifically, each training epoch contains multiple alternating updates of the generator and discriminator. During generator updates, the discriminator parameters are fixed to minimize the generation loss; during discriminator updates, the generator parameters are fixed to maximize the discrimination accuracy. Through repeated iterations, the two networks gradually reach a dynamic equilibrium.
[0125] Through the above technical solution, this embodiment achieves high-quality generation of 3D models of intangible cultural heritage collections. The encoder-decoder structure of the generator effectively integrates multimodal input information, enhancing the model's expressiveness. The discriminator's multiple loss functions comprehensively constrain the generation process, ensuring the model's realism and stability. The alternating optimization strategy ensures the convergence of adversarial training. Therefore, this solution can generate 3D models of intangible cultural heritage collections with refined artistic style, stable structure, and rich personalized characteristics, providing strong support for digital display and preservation.
[0126] In some of the solutions described above in this embodiment, when the feature values of user profile data are directly used to adjust the generation weights, the significant differences in the numerical ranges of different user behavior feature vectors lead to inconsistent data distribution in the generator's input layer, affecting the stability of model generation. Furthermore, static weight allocation cannot adapt to dynamic changes in behavioral patterns, resulting in a decline in the quality of personalized output.
[0127] This embodiment further proposes to normalize the feature values after calculating the feature values of the user behavior feature vector to ensure that the values are within a preset range; to map the feature values to the input layer of the generator using a weighted summation algorithm, wherein the weight coefficients are dynamically adjusted based on the iconic labels; the adjustment process includes real-time monitoring of feature value changes, and recalculating the trigger weights to maintain generation stability when the behavior pattern changes abruptly; after integrating the adjustment results, the generated weights are used to control the model complexity or level of detail and improve the quality of personalized output.
[0128] The normalization process employs a max-min scaling method to linearly transform feature values to the [0,1] interval, eliminating dimensional differences between different users. In the weighted summation algorithm, weight coefficients are dynamically assigned based on the user's interest category corresponding to the iconic tag; for example, the historical tag corresponds to a weight coefficient of 0.6, and the art tag corresponds to a weight coefficient of 0.4. Real-time monitoring is implemented through a sliding window mechanism with a window length of 10 seconds. When the standard deviation of feature values within the window exceeds a threshold of 0.15, weight recalculation is triggered. When generating weights to control model complexity, weight values are mapped to the number of channels in the generator's convolutional layers; higher weights result in more channels and richer model detail.
[0129] Specifically, the eigenvalue normalization process uses the formula ,in These are the original eigenvalues. and These are preset minimum and maximum boundary values to ensure that the feature values of all users are on a uniform scale. During the weighted summation process, each iconic label is associated with a weight coefficient, which is automatically adjusted according to the label's update frequency. For example, the coefficient increases by 0.05 for each additional label update, with an upper limit of 1.0. The real-time monitoring module detects abrupt changes in feature values through time series analysis. When a sudden change is detected, the current generation process is immediately interrupted, and the weight calculation module is invoked to regenerate the coefficient matrix. The generated weights are applied to the decoder layer of the generator, and the weight values are fused into the feature map through matrix multiplication. For example, the weight matrix is used to perform a Hadamard product with the parameter tensor of the point cloud generation layer, thereby controlling the resolution and surface subdivision of the generated model. This process dynamically adjusts the network capacity, enabling the generated model to maintain basic structural stability while enhancing local detail representation based on real-time user behavior.
[0130] A weighted summation algorithm is used to map the normalized feature values to the input layer of the generator. The formula for calculating the weighted summation is: ,in For the first 1 eigenvalue, These are the corresponding weighting coefficients. The weighting coefficients are dynamically adjusted based on the iconic labels and normalized using the softmax function. , For tags The score is determined by the importance of the label.
[0131] The adjustment process includes real-time monitoring of feature value changes. The system calculates the rate of change of feature values every 5 minutes. When the rate of change of any feature value exceeds a preset threshold of 20%, a recalculation of the weights is triggered. The recalculation uses a sliding window method, employing data from the most recent 30 minutes.
[0132] After integrating and adjusting the results, weights are generated to control model complexity. Specifically, the larger the weight value, the richer the details of the generated model. For example, when the weight value is greater than 0.8, the number of point clouds in the model increases by 50%, and the texture resolution doubles.
[0133] Through the above technical solution, this embodiment achieves a dynamic correlation between user behavior data and model generation, improving the personalization of intangible cultural heritage collection models. Because feature values are updated in real time and mapped to generation weights, the model output can quickly adapt to changes in user preferences. Simultaneously, the label-based weight adjustment mechanism enhances the interpretability and controllability of the generation process. Furthermore, by controlling model complexity, computational resource utilization efficiency is optimized while ensuring personalization.
[0134] In some of the solutions described above in this embodiment, there is a risk of identifier duplication when creating augmented reality markers, and the environmental location association is easily tampered with, resulting in inaccurate matching between the displayed content and the real scene, which affects the user experience.
[0135] This embodiment further proposes a method for creating augmented reality markers and unique identifiers, including: identifying the environmental location information of the real scene, obtaining latitude and longitude coordinates and three-dimensional point cloud data through GPS sensors or visual positioning systems; assigning unique identifiers to exhibition targets, generating identifiers using hash algorithms; establishing an environmental location association between augmented reality markers and intangible cultural heritage collection models, storing the association in a distributed file system network module, and storing summary information through a blockchain module; marker deployment supports multiple forms, including QR codes, image markers, or virtual anchors.
[0136] The environmental location information is obtained through GPS sensors to acquire latitude and longitude coordinates, or through a visual positioning system to extract 3D point cloud data, achieving centimeter-level positioning accuracy. The hash algorithm uses SHA-256 to generate 128-bit identifiers, ensuring global uniqueness. The distributed file system network module uses the IPFS protocol to store association relationships, and the blockchain module uses Ethereum smart contracts to store hash digests, preventing data tampering. The marking format is selected according to the scenario requirements: QR codes are suitable for printing scenarios, image markers are adapted to dynamic lighting environments, and virtual anchors are used for textureless areas.
[0137] Specifically, the visual positioning system captures scene feature points using multiple cameras, generates 3D point cloud data, and aligns it with a pre-stored scene model, controlling the positioning error within ±2 cm. A hash algorithm encrypts the attribute data of the exhibition target, generates a unique identifier, and binds it to latitude and longitude coordinates. When storing associations, the IPFS protocol encapsulates coordinates, identifiers, and model index information into a CID hash value, and an Ethereum smart contract records the CID and generates a transaction certificate. When a user triggers an AR display, the system verifies the integrity of the CID through the blockchain; if the hash value matches, the associated model is retrieved. When deployed with multiple marking methods, virtual anchors construct a spatial map based on SLAM technology, generating persistent anchors in unmarked areas, adapting to various exhibition scenarios such as museums and outdoor spaces.
[0138] In practice, identifying the environmental location information of a real-world scene involves acquiring latitude and longitude coordinates and 3D point cloud data using GPS sensors or a visual positioning system. The GPS sensor employs a high-precision dual-frequency receiver, achieving centimeter-level positioning accuracy. The visual positioning system uses a depth camera to collect 3D environmental information and combines it with SLAM algorithms to achieve centimeter-level positioning.
[0139] A unique identifier is assigned to each exhibition target, generated using the SHA-256 hash algorithm to create a 32-byte identifier. The identifier is generated by combining the exhibition target's location coordinates, a timestamp, and a random number, ensuring global uniqueness. The hash value is encrypted using digital signature technology to prevent tampering.
[0140] Establish environmental location associations between augmented reality markers and intangible cultural heritage artifact models. Data of these associations is stored in the IPFS distributed file system network, employing content addressing for efficient retrieval. Hash digests of the associations are stored on the Ethereum blockchain, and smart contracts are used for data verification and access control.
[0141] Tag deployment supports multiple formats. QR code tags use DataMatrix encoding and support a resolution of 256x256 pixels. Image tags use feature point extraction algorithms such as SIFT to generate feature descriptors. Virtual anchors are based on visual SLAM technology, achieving spatial localization through feature matching and pose estimation.
[0142] Through the above technical solutions, this embodiment achieves high-precision positioning and unique identification of augmented reality markers, improving the accuracy of displaying intangible cultural heritage collection models in real-world scenarios. The application of distributed storage and blockchain technology enhances data security and traceability. Diverse marker deployment methods adapt to different exhibition environments, improving system applicability and user experience.
[0143] In some of the solutions described above in this embodiment, there is a problem that the matching robustness between pre-stored real scene images and real-time captured environment images by the user is insufficient, resulting in matching failure when there is complex lighting or changes in viewing angle, which affects the accuracy of augmented reality display.
[0144] This embodiment further proposes a matching method between the current environment image and a pre-stored real scene image, including: extracting the location information and environmental information of the current environment image, where the location information includes GPS data from the shooting device and the environmental information includes lighting conditions and object contour features; performing real-time positioning of the user and verifying the consistency between the positioning result and the information of the current environment image, with verification methods including feature point matching or deep learning model comparison; when the consistency verification passes, retrieving the pre-stored real scene image from the distributed file system network module, where the pre-stored image includes a high-resolution panoramic image or a 3D scene model; and employing a multi-scale feature extraction algorithm in the matching process to ensure robustness under different viewing angles and lighting conditions.
[0145] Location information is extracted by obtaining latitude and longitude coordinates through GPS sensors and capturing 3D point cloud data using a visual positioning system to achieve spatial positioning. Illumination conditions in the environment are obtained by acquiring color temperature and brightness parameters through image sensors, and object contour features are extracted using edge detection algorithms to extract geometric structures. Real-time positioning verification uses feature point matching, extracting key points and calculating descriptor similarity through the SIFT algorithm, or extracting image embedding vectors through a pre-trained convolutional neural network and performing cosine similarity comparison. When retrieving pre-stored images, the distributed file system network module quickly retrieves high-resolution panoramic images or 3D model data of the corresponding scene based on location hash values. Multi-scale feature extraction algorithms construct Gaussian pyramids to decompose the image, extracting local features at different resolution levels and weighted fusion to enhance adaptability to changes in viewpoint and illumination.
[0146] Specifically, after a user captures an image of the current environment, the device simultaneously acquires GPS coordinates and 3D point cloud data. It then analyzes ambient lighting parameters using an image sensor and generates object contour features through edge detection. During real-time positioning, the GPS coordinates are spatially aligned with the visual positioning results. If the deviation exceeds a threshold, feature point matching or deep learning model comparison is triggered to ensure consistency between the positioning result and the image information. After successful verification, a high-resolution panoramic image or 3D model of the corresponding scene is retrieved from distributed storage based on the location hash value as a matching benchmark. During the matching process, a multi-scale feature extraction algorithm is used to extract features from the current image and pre-stored images in a layered manner. By weighted fusion of feature responses at different scales, perspective changes and lighting interference are eliminated, resulting in a robust matching result. Thus, through multi-dimensional data fusion and layered feature optimization, the problem of insufficient image matching accuracy in complex environments is solved, improving the stability of augmented reality displays.
[0147] In practice, the location and environmental information of the current environment image are extracted. Location information includes GPS data from the capturing device, obtained through the device's built-in GPS module, which provides latitude and longitude coordinates. Environmental information includes lighting conditions and object contour features; image processing algorithms are used to extract the image's brightness histogram and edge detection results.
[0148] The system performs real-time user location tracking and verifies the consistency between the location results and information from the current environmental image. Real-time location tracking employs WiFi fingerprinting technology, comparing the strength of surrounding WiFi signals with a pre-stored database. Consistency verification methods include feature point matching and deep learning model comparison. Feature point matching uses the SIFT algorithm to extract image feature points and calculates the Euclidean distance between them for matching. Deep learning model comparison uses a pre-trained convolutional neural network to extract image feature vectors and calculates cosine similarity for comparison.
[0149] Once the consistency verification passes, pre-stored real-world scene images are retrieved from the distributed file system network module. These pre-stored images include high-resolution panoramic images and 3D scene models. The high-resolution panoramic images are stored in an equidistant cylindrical projection format with a resolution of 8K. The 3D scene models are stored using point cloud data format, containing color and depth information.
[0150] The matching process employs a multi-scale feature extraction algorithm to ensure robustness under different viewpoints and lighting conditions. This multi-scale feature extraction algorithm includes Gaussian pyramids and Scale Invariant Feature Transform (SIFT). Gaussian pyramids generate images at different scales by performing multiple downsampling and Gaussian blurring operations. The SIFT algorithm detects and describes local features at different scales, generating 128-dimensional feature vectors. By comparing the feature matching results at different scales, the optimal scale and viewpoint parameters are selected for the best match.
[0151] Through the above technical solutions, this embodiment achieves accurate matching between the current environmental image and pre-stored real-world scene images. The use of multi-dimensional information verification and multi-scale feature extraction improves the accuracy and robustness of the matching. Real-time positioning and consistency verification ensure the timeliness and reliability of the matching results. Using a distributed file system to store and retrieve high-quality pre-stored images provides a foundation for subsequent augmented reality displays. The multi-scale feature extraction algorithm overcomes the matching difficulties under different viewpoints and lighting conditions, enhancing the system's adaptability. These technical means work together to create conditions for the accurate positioning and display of intangible cultural heritage collection models, improving the user experience of augmented reality displays.
[0152] In some of the solutions described above in this embodiment, the matching process between the current environment image and the pre-stored real scene image has insufficient robustness. In particular, under different viewing angles and lighting conditions, traditional matching methods are prone to misidentification or matching failure, affecting the accuracy of augmented reality display and user experience.
[0153] This embodiment further proposes a matching method between the current environment image and a pre-stored real scene image, including: extracting the location information and environmental information of the current environment image, where the location information includes GPS data from the shooting device and the environmental information includes lighting conditions and object contour features; performing real-time positioning of the user and verifying the consistency between the positioning result and the information of the current environment image, with verification methods including feature point matching or deep learning model comparison; when the consistency verification passes, retrieving the pre-stored real scene image from the distributed file system network module, where the pre-stored image includes a high-resolution panoramic image or a 3D scene model; and employing a multi-scale feature extraction algorithm in the matching process to ensure robustness under different viewing angles and lighting conditions.
[0154] Location information is acquired through GPS sensors or visual positioning systems, while environmental information is obtained by parsing raw data collected by image sensors. During real-time location verification, key points are extracted using SIFT or ORB algorithms for feature point matching, and similarity calculations are performed using pre-trained convolutional neural networks for deep learning model comparison. Multi-scale feature extraction algorithms construct an image pyramid structure to extract local and global features at different resolution levels, and incorporate an adaptive thresholding mechanism to filter noise interference.
[0155] Specifically, after the location and environmental information of the shooting device are collected synchronously, data preprocessing is performed through a distributed file system network module. In the real-time positioning verification phase, a feature point matching algorithm extracts geometric features from the current image and pre-stored images for spatial alignment, and a deep learning model compares the features by calculating the cosine similarity of the feature embedding vectors to determine consistency. When verification is successful, a high-resolution panoramic image or 3D scene model is loaded into the augmented reality engine. During the matching process, a multi-scale feature extraction algorithm integrates texture and shape information at different scales through a hierarchical feature fusion mechanism, and combines illumination-invariant descriptors to eliminate the influence of ambient light changes, ensuring the stability of the matching results under different viewpoints.
[0156] As a preferred embodiment, the solution of this embodiment is implemented as follows: When a user uses a mobile device to photograph the blue and white porcelain display case in the museum exhibition hall, the current latitude and longitude coordinates are first obtained through the device's built-in GPS module. Simultaneously, an image recognition algorithm is used to extract the outline features of objects in the scene, including the geometric shape of the display case's edges and the layout of the exhibits. The light intensity in the environmental information is collected in real time by the device's light sensor, generating an environmental parameter matrix containing brightness and color temperature values.
[0157] The user's real-time location data is spatially mapped to the 3D coordinates of a pre-stored scene. The ORB feature point matching algorithm is used to calculate the similarity between the corner features of the display case in the current image and the pre-stored high-resolution panoramic image. During the verification process, a pre-trained convolutional neural network model is used to compare the scene texture features. When the matching similarity exceeds a preset threshold, the corresponding 3D scene model data is loaded from the distributed storage system.
[0158] The matching algorithm further employs a multi-scale pyramid structure to process the image, extracting SIFT feature descriptors at different resolution levels and removing outlier matching points using the RANSAC algorithm. For local feature distortions caused by changes in viewing angle, spatial correction is performed using an affine transformation matrix to ensure stable feature associations across different viewing angles in the exhibition hall.
[0159] Through the above technical solution, this embodiment effectively solves the matching failure problem caused by environmental changes in existing AR display systems. By integrating multi-dimensional environmental parameters and a dynamic feature extraction mechanism, it improves the accuracy and robustness of real-world scene image matching. This method can stably trigger the display of intangible cultural heritage collection models in complex lighting conditions and multi-view observation scenarios, avoiding the susceptibility to interference inherent in traditional single-feature matching methods. Simultaneously, a distributed data retrieval mechanism ensures efficient loading of large-scale scene models.
[0160] In some of the solutions described above in this embodiment, the matching process between the pre-stored real scene image and the user's current environment image is easily affected by changes in viewing angle or differences in lighting conditions, leading to matching failure or incorrect triggering of augmented reality display, which affects the user experience.
[0161] This embodiment further proposes a matching method between the current environment image and a pre-stored real scene image, including: extracting the location information and environmental information of the current environment image, where the location information includes GPS data from the shooting device and the environmental information includes lighting conditions and object contour features; performing real-time positioning of the user and verifying the consistency between the positioning result and the information of the current environment image, the verification method including feature point matching or deep learning model comparison; when the consistency verification passes, retrieving the pre-stored real scene image from the distributed file system network module, the pre-stored image including a high-resolution panoramic image or a 3D scene model; the matching process uses a multi-scale feature extraction algorithm to ensure robustness under different viewpoints and lighting conditions.
[0162] Location information is acquired via GPS sensors or a visual positioning system; ambient lighting conditions are measured using a color temperature sensor; and object contour features are extracted using an edge detection algorithm. During real-time positioning, a fusion positioning technique combines GPS data and visual feature point coordinates to generate three-dimensional spatial coordinates. For consistency verification, a convolutional neural network is used to extract feature vectors from the current image and pre-stored images, and a cosine similarity threshold is calculated. A multi-scale feature extraction algorithm uses Gaussian pyramid decomposition to extract key points at different resolution levels and performs cross-scale matching using feature descriptors.
[0163] Specifically, when a user triggers an augmented reality display, the device first acquires the latitude and longitude coordinates and lighting parameters of the current environment using its built-in sensors, and simultaneously uses the Canny operator to extract the edge contours of objects in the scene. The real-time positioning module weighted and fused the GPS coordinates with the local coordinates output by the visual SLAM system to generate accurate 3D positioning results. In the consistency verification phase, a pre-trained ResNet model extracts global features from the image and compares them with the panoramic image features stored in a distributed file system. When the similarity exceeds a preset threshold, image retrieval is triggered. During the matching process, a multi-scale pyramid structure is used to downsample the image, capturing global structural features at low-resolution layers and extracting detailed texture features at high-resolution layers. Bidirectional feature mapping eliminates deformation interference caused by viewpoint differences, ultimately achieving stable matching across lighting and viewpoint conditions.
[0164] As a preferred embodiment, the specific implementation of this embodiment is as follows: In a real exhibition scene, environmental location information is acquired through a visual positioning system that collects 3D point cloud data. This system includes a binocular camera array and an inertial measurement unit (IMU) to capture the geometric structure and spatial coordinates of the scene. The location coordinates of the exhibition target are updated in real-time using a SLAM algorithm and input into a hash function to generate a unique 128-bit identifier. The hash function uses the SHA-256 algorithm to ensure the global uniqueness of the identifier. The association between augmented reality markers and intangible cultural heritage collection models is stored in the IPFS distributed file system as key-value pairs. The key is the identifier hash value, and the value is the model file index path. The blockchain module uses an Ethereum smart contract to encrypt and store the summary information of the association. The summary information includes a timestamp and digital signature to prevent data tampering. The deployment method of the augmented reality markers is dynamically selected according to scene requirements. If the lighting conditions of the exhibition scene are complex, high-contrast QR codes are used as markers; if the scene contains rich texture features, SIFT feature point matching image markers are used; if display in an unmarked environment is required, virtual anchor points are generated using the ARKit framework.
[0165] Through the above technical solutions, this embodiment achieves high-precision binding between augmented reality markers and real-world scenes, solving the problems of existing AR displays relying on fixed pre-stored scenes and lacking environmental adaptability. Hash algorithms and blockchain technology ensure the uniqueness of identifiers and data integrity, preventing unauthorized tampering of displayed content. Distributed storage and multi-form marker deployment mechanisms enable the intangible cultural heritage collection models to be stably triggered in exhibition scenarios with different lighting conditions and device types, improving the consistency and interactivity of the user experience.
[0166] In some of the solutions described above in this embodiment, during the matching process between the current environment image and the pre-stored real scene image, the matching accuracy is insufficient due to changes in viewing angle or differences in lighting conditions, which affects the accuracy and stability of augmented reality display.
[0167] This embodiment further proposes a matching method between the current environment image and a pre-stored real scene image, including: extracting the location information and environmental information of the current environment image, where the location information includes GPS data from the shooting device and the environmental information includes lighting conditions and object contour features; performing real-time positioning of the user and verifying the consistency between the positioning result and the information of the current environment image, with verification methods including feature point matching or deep learning model comparison; when the consistency verification passes, retrieving the pre-stored real scene image from the distributed file system network module, where the pre-stored image includes a high-resolution panoramic image or a 3D scene model; and employing a multi-scale feature extraction algorithm in the matching process to ensure robustness under different viewing angles and lighting conditions.
[0168] Specifically, location information is extracted by obtaining latitude and longitude coordinates through GPS sensors, and environmental information is obtained by analyzing light intensity and object contours through image sensors. Real-time positioning combines GPS data with the output of the visual positioning system to generate the user's three-dimensional coordinates. Consistency verification uses a feature point matching algorithm to extract key points and calculate similarity, or performs image comparison through a pre-trained convolutional neural network. Pre-stored real scene images are stored in a distributed file system network module, supporting high-concurrency access. Multi-scale feature extraction algorithms extract local and global features through convolutional kernels of different resolutions to enhance matching robustness.
[0169] Specifically, after the imaging device captures images of the current environment, the system simultaneously acquires GPS data and image sensor data, extracting latitude and longitude coordinates, light intensity, and object contour features. The real-time positioning module integrates GPS positioning results with 3D point cloud data from the visual positioning system to generate the user's precise location. In the consistency verification phase, if feature point matching is used, SIFT or ORB feature points are extracted from the current image and pre-stored images, and the matching degree is calculated; if a deep learning model is used, feature vectors are extracted through a pre-trained ResNet network, and cosine similarity is calculated. When the matching degree exceeds a preset threshold, the corresponding high-resolution panoramic image or 3D scene model is retrieved from the distributed file system network module. A multi-scale feature extraction algorithm is applied during the matching process, extracting low-frequency and high-frequency information from the image through convolutional layers of different scales, and combining illumination invariance transformation to handle differences in ambient lighting, ensuring accurate matching even with changes in viewing angle or fluctuations in lighting. This improves the recognition accuracy and display stability of augmented reality markers, avoiding model display errors or delays caused by environmental interference.
[0170] As a preferred embodiment, the specific implementation of this embodiment is as follows: The geographic coordinates and visual feature points of the current environment image captured by the user equipment are extracted, wherein the geographic coordinates are obtained through a GPS module, and the visual feature points are extracted using the ORB algorithm; the extracted visual feature points are input into a pre-trained ResNet-50 model, outputting a 256-dimensional feature vector as an environmental information descriptor; the user equipment is located in real time, and the device pose is calculated using a visual inertial odometry system to generate six-degree-of-freedom spatial coordinates; the spatial coordinates and the environmental information descriptor are input into a feature matching module, and the cosine similarity algorithm is used to calculate the current environmental coordinates. The similarity score between the current environment image and the pre-stored real scene image is calculated. When the similarity score exceeds a preset threshold, the distributed file system network module is triggered to call the pre-stored high-resolution panoramic image, which contains a multi-level 3D scene model. A multi-scale feature space is constructed using a Gaussian pyramid, and SIFT feature descriptors are extracted at the 32×32, 64×64, and 128×128 pixel levels. The feature points of the current environment image and the pre-stored panoramic image are aligned using a bidirectional nearest neighbor matching algorithm, and mismatched points are removed using the RANSAC algorithm. After matching is completed, the augmented reality engine is activated to load the intangible cultural heritage collection model associated with the scene.
[0171] Through the above technical solution, this embodiment solves the problem of insufficient robustness of pre-stored scene matching in existing AR display technologies. It can achieve high-precision image matching under different lighting conditions, viewpoint shifts and partial occlusion, effectively improve the accuracy of augmented reality marker triggering, ensure the stable display of intangible cultural heritage collection models in dynamic environments, and avoid matching failures caused by resolution differences through multi-scale feature extraction.
[0172] Furthermore, this application also proposes a computer-readable storage medium storing a program for generating intangible cultural heritage collection models based on a GAN style constraint mechanism. When the program for generating intangible cultural heritage collection models based on a GAN style constraint mechanism is executed by a processor, it implements the steps of the method for generating intangible cultural heritage collection models based on a GAN style constraint mechanism as described above.
[0173] Reference Figure 3 , Figure 3 This is a structural block diagram of the first embodiment of the intangible cultural heritage collection model generation system based on the GAN style constraint mechanism of this application.
[0174] like Figure 3 As shown in the embodiments of this application, the intangible cultural heritage collection model generation system based on GAN style constraint mechanism includes:
[0175] Data acquisition module 10 is used to receive basic raw data of intangible cultural heritage collections, including image materials and text descriptions;
[0176] The behavior information acquisition module 20 is used to acquire user behavior information, which includes user interaction events, browsing history or operation records on the digital platform.
[0177] The profile generation module 30 is used to generate user profile data based on user behavior information, including: extracting key features from the behavior information and constructing a behavior feature vector; calculating feature values, the feature values including a first feature value and an average feature value, used to quantify behavior patterns; and establishing a distinctive label, the label being generated based on multi-dimensional feature information classification.
[0178] Integration module 40 is used to integrate user profile data and basic raw data as input for generating intangible cultural heritage collection models;
[0179] The model generation module 50 is used to generate a 3D model of intangible cultural heritage collections using a generative adversarial network. The generator receives style constraint parameters, which generate style encoding vectors through contrastive learning. The style encoding vectors are used to control the artistic style of the intangible cultural heritage collection model.
[0180] Discriminator module 60 is used to optimize the model's realism based on adversarial training. The loss function includes a content preservation term and a style alignment term.
[0181] Optimization module 70 is used to adjust the generation weights based on the feature values of user profile data in order to optimize the model generation process;
[0182] Module 80 is used to create augmented reality tags and unique identifiers for exhibition targets in real-world scenes using an augmented reality exhibition application;
[0183] The output module 90 is used to display models of intangible cultural heritage artifacts at the exhibition target location by recognizing the unique identifier of the augmented reality marker when the current environment image captured by the user matches the pre-stored real scene image.
[0184] It should be understood that the above are merely illustrative examples and do not constitute any limitation on the technical solution of this application. In specific applications, those skilled in the art can make settings as needed, and this application does not impose any restrictions on this.
[0185] This embodiment integrates user profile data and intangible cultural heritage basic data to drive the generation of adversarial networks, combines style constraint mechanisms to achieve precise control of artistic style, and utilizes augmented reality technology to achieve dynamic scene matching and display. It has the advantages of improving the efficiency of digital modeling of intangible cultural heritage, enhancing user interaction experience, and ensuring the accuracy of cultural inheritance.
[0186] It should be noted that the workflow described above is merely illustrative and does not limit the scope of protection of this application. In practical applications, those skilled in the art can select some or all of it to achieve the purpose of this embodiment according to actual needs, and no restrictions are imposed here.
[0187] In addition, for technical details not described in detail in this embodiment, please refer to the method for generating intangible cultural heritage collection models based on GAN style constraint mechanism provided in any embodiment of this application, which will not be repeated here.
[0188] Furthermore, it should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or system that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or system. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or system that includes that element.
[0189] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0190] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as read-only memory (ROM) / RAM, magnetic disk, optical disk), and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods of the various embodiments of this application. The above are only preferred embodiments of this application and do not limit the patent scope of this application. All equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.
Claims
1. A method for generating a non-heritage artifact model based on a GAN style constraint mechanism, characterized in that, The method comprises the following steps: Receiving basic raw data of non-heritage collections, which includes image materials and text descriptions; Obtaining user behavior information, which includes user interaction events, browsing history or operation records on digital platforms; Generating user portrait data based on user behavior information, including: extracting key features from behavior information, constructing behavior feature vectors; calculating feature values, including first feature values and average feature values, for quantifying behavior patterns; establishing iconic labels based on multi-dimensional feature information classification; Integrating user portrait data and basic raw data as input for generating non-heritage collection models; Using a generative adversarial network to generate a three-dimensional model of non-heritage collections, wherein the generator receives style constraint parameters, which generate style encoding vectors through contrastive learning, and the style encoding vectors are used to control the artistic style of the non-heritage collection model; The discriminator optimizes the authenticity of the model based on adversarial training, and the loss function contains a content preservation term and a style alignment term; Adjust the generation weight through the feature value of the user portrait data to optimize the model generation process; Using an augmented reality exhibition application to create augmented reality markers and unique identifiers for exhibition targets in real scenes; When the current environment image taken by the user matches the pre-stored real scene image, display the non-heritage collection model at the exhibition target by identifying the unique identifier of the augmented reality marker; Wherein, the generation of user portrait data based on user behavior information comprises: Obtain user behavior information and preprocess the behavior information to extract key features, including click frequency, session length or interaction depth; Construct a behavior feature vector and convert the key features into a numerical vector representation using a pre-trained encoder, where the encoder is implemented based on a neural network model to ensure consistent vector dimensions; Calculate the feature values, including the first feature values and the average feature values, the first feature values are obtained by vector length calculation, reflecting the overall strength of the behavior, and the average feature values are obtained by vector average value calculation, measuring the behavior stability; Establish iconic labels based on clustering algorithm to classify feature values, generate labels and optimize label accuracy through weight distribution.
2. The method of claim 1, wherein, The generation of the style constraint parameter is realized through contrastive learning, including: Collect target style samples and negative samples, including non-heritage art reference images or text descriptions; Use a style encoder to extract embedding vectors from the samples, which is based on a convolutional neural network; Calculate the similarity of positive and negative samples through a contrastive loss function, which shortens the distance between samples of the same style and lengthens the distance between samples of different styles; Generate a style encoding vector that quantifies artistic style attributes and is used to initialize the input layer of the generator.
3. The method of claim 1, wherein, The training of the generative adversarial network includes: The generator receives the style encoding vector and integrated data as input and outputs a three-dimensional model, wherein the generator structure contains an encoder-decoder architecture, the encoder processes the fusion features of image materials and text descriptions, and the decoder generates model point cloud data; The discriminator optimizes the model authenticity based on adversarial training, and the loss function includes a spatiotemporal consistency term in addition to the content preservation term and the style alignment term, ensuring the stability of the model in dynamic scenes. The training process adopts an alternating optimization strategy, iteratively updating the generator and discriminator parameters until the model converges to a Nash equilibrium point.
4. The method of claim 1, wherein, The generation of the weight through the feature value adjustment of the user portrait data includes: After calculating the feature values of the user behavior feature vector, the feature values are normalized to ensure that the values are within a predetermined range; The feature values are mapped to the input layer of the generator using a weighted sum algorithm, where the weight coefficients are dynamically adjusted based on the landmark labels; The adjustment process includes real-time monitoring of feature value changes, and when the behavior pattern changes abruptly, the weight is recalculated to maintain stability in generation; After integrating the adjustment results, the generated weight is used to control the model complexity or detail level, improving the quality of personalized output.
5. The method of claim 1, wherein, The creation of the augmented reality marker and the unique identifier includes: Identify the environmental location information of the real scene, obtain the latitude and longitude coordinates and three-dimensional point cloud data through GPS sensors or visual positioning systems; Assign a unique identifier to the exhibition target, generate the identifier using a hash algorithm to ensure global uniqueness and tamper resistance; Establish the environmental location association relationship between the augmented reality marker and the intangible cultural heritage collection model, and store the association relationship in the distributed file system network module, and store the summary information through the blockchain module to ensure security; The marker deployment supports multiple forms, including two-dimensional codes, image markers, or virtual anchor points, which are suitable for different exhibition scenarios.
6. The method of claim 1, wherein, The matching of the current environment image with the pre-stored real scene image includes: Extract the location information and environmental information of the current environment image, including the GPS data of the shooting device and the environmental information including the lighting conditions and object contour features; Real-time positioning of the user, consistency verification of the positioning result with the information of the current environment image, verification methods include feature point matching or deep learning model comparison; When the consistency verification is passed, the pre-stored real scene image is retrieved from the distributed file system network module, which includes high-resolution panoramic images or three-dimensional scene models; The matching process uses a multi-scale feature extraction algorithm to ensure robustness under different viewing angles and lighting conditions.
7. A non-heritage collection model generation system based on a GAN style constraint mechanism, characterized by, The method of claim 1 includes: A data acquisition module for receiving basic raw data of intangible cultural heritage collections, including image materials and text descriptions; A behavior information acquisition module for acquiring user behavior information, including user interaction events, browsing history, or operation records on digital platforms; A portrait generation module for generating user portrait data based on user behavior information, including: extracting key features from behavior information, constructing a behavior feature vector; calculating feature values, including first feature values and average feature values, for quantifying behavior patterns; and establishing landmark labels based on multi-dimensional feature information classification; An integration module for integrating user portrait data and basic raw data as input for generating intangible cultural heritage collection models. The model generation module is configured to generate a three-dimensional model of the non-culture relic by using a generative adversarial network, wherein a generator receives a style constraint parameter, the style constraint parameter generates a style encoding vector by contrast learning, and the style encoding vector is used to control an artistic style of the non-culture relic model. The discriminator module is configured to optimize the authenticity of the model based on adversarial training of a discriminator, and a loss function includes a content preservation term and a style alignment term. The optimization module is configured to adjust the generation weight by using the feature value of the user portrait data to optimize the model generation process. The representation module is configured to create an augmented reality marker and a unique identifier for an exhibition target in a real scene by using an augmented reality exhibition application. The output module is configured to display the non-culture relic model at the exhibition target by identifying the unique identifier of the augmented reality marker when a current environment image captured by the user matches a pre-stored real scene image.
8. A computer device, comprising: The device comprises a memory and a processor, and the processor executes the method according to any one of claims 1-6 when running computer instructions stored in the memory.
9. A computer-readable storage medium, characterized in that, The instructions, when executed on a computer, cause the computer to perform the method according to any one of claims 1-6.
Citation Information
Patent Citations
Intangible cultural heritage personalized interaction method and system based on virtual reality
CN120255692A
Virtual reality interaction method and system applied to classical famous picture display
CN120428866A