Figure and creation perception system based on virtuality and reality

Through a cultural and creative perception system combining virtual and reality, the images and structural characteristics of real cultural and creative products are collected, and personalized cultural and creative products are generated based on user interest areas, which solves the problem of lack of deep perception of user interests in the existing technology, and realizes the intelligent generation of personalized cultural and creative products and the in-depth expression of cultural content.

CN120428893AInactive Publication Date: 2025-08-05JIANGMEN POLYTECHNIC
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510865634.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2025-08-05
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing cultural and creative design and display methods lack deep perception and structured expression of user interests, and it is difficult to achieve personalized customization. There is a lack of semantic bridging between real cultural and creative products and virtual cultural content, and there is a lack of cultural content customization mechanism driven by user interaction behavior.

Method used

Based on virtual and reality, the cultural and creative perception system is used to collect product images and structural features through the real cultural and creative display module, combine users' interest areas in the virtual cultural environment, and use deep learning generation models to generate personalized cultural and creative product images, including the real cultural and creative display module, virtual cultural and cultural and creative display module, user interest perception module and cultural and creative generation module.

Benefits of technology

It realizes the automated output of personalized cultural and creative products, improves user interaction experience and cultural identity, enhances the participation, customization and intelligence level of the design process of cultural and creative products, and solves the problems of pattern map distortion and structural misalignment in the existing technology.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120428893A_ABST
    Figure CN120428893A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of virtual display, and provides a virtuality and reality-based cultural creativity perception system, which comprises the steps of displaying at least one preset cultural creativity product, and collecting image data and appearance structure characteristics of the at least one preset cultural creativity product; displaying the virtual culture content corresponding to the at least one preset cultural and creative product; identifying a region of interest of the user in the virtual culture content based on a gazing track, a staying duration and an interaction behavior of the user, and extracting a content label and image information of the region of interest; determining at least one content image based on the content label and the image information; and jointly inputting the image data and the appearance structure features of the at least one preset cultural and creative product and the at least one content image into a cultural and creative generation model based on deep learning, and generating at least one personalized cultural and creative product image containing user interest content.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of virtual display, and specifically relates to a cultural and creative perception system based on virtuality and reality. Background Art

[0002] With the continued development of the cultural and creative industries and the growing consumer demand for personalized cultural expression and immersive experiences, traditional cultural and creative product design and display methods face numerous challenges. Existing cultural and creative designs primarily rely on the combination of flat patterns and style transfer, relying heavily on the designer's experience for pattern conception and manual editing. This approach lacks a deep understanding of user interests and support for structured expression, making efficient personalized customization difficult. Furthermore, during display, the physical form of cultural and creative products is often disconnected from the cultural meaning they convey. Users can only indirectly understand through graphic descriptions or audio explanations, lacking a visual and interactive understanding of cultural background, symbolic meaning, and historical origins. This, in turn, reduces the depth of cultural cognition and the formation of emotional resonance.

[0003] On the other hand, in recent years, virtual reality (VR), augmented reality (AR), and multimodal perception technologies have seen initial application in the field of cultural display. Some cultural museums and exhibition platforms have attempted to digitally display information related to cultural relics and artworks. However, such applications are often limited to surface rendering of real-world images or three-dimensional models, failing to effectively integrate structured cultural semantic maps and lacking mechanisms for cultural content customization driven by user interaction. Therefore, current technologies still face key bottlenecks in the "user-cultural content-personalized product" connection, including a lack of semantic bridging, insufficient utilization of behavioral feedback, and poor adaptability of generative models. This makes it difficult to meet the demand for integrated design and display of large-scale personalized cultural and creative products. Summary of the Invention

[0004] In order to solve the problems in the prior art, the present invention provides a cultural and creative perception system based on virtuality and reality, comprising: A real-life cultural and creative display module, configured to display at least one preset cultural and creative product and collect image data and appearance and structural features of the at least one preset cultural and creative product; a virtual cultural display module, configured to display virtual cultural content corresponding to the at least one preset cultural and creative product, wherein the virtual cultural content includes one or more of the historical background, symbolic meaning, application scenarios, and extended interpretation of cultural elements of the cultural and creative product; A user interest perception module is used to identify user interest areas in the virtual cultural content based on the user's gaze trajectory, dwell time, and interactive behavior, and to extract content labels and image information of the interest areas; a content extraction module, configured to determine at least one content image based on the content tag and the image information, wherein the content image includes one or more combinations of cultural patterns, totem symbols, and story elements; The cultural and creative generation module is used to input the image data and appearance structure features of the at least one preset cultural and creative product and the at least one content image into the cultural and creative generation model based on deep learning to generate at least one personalized cultural and creative product image containing content of user interest.

[0005] Furthermore, the real cultural and creative display module includes: The display control module is used to display cultural and creative products through display racks, support platforms or electric rotating devices, and adjust the angle of the booth and the intensity of the lights to ensure that the cultural and creative products are displayed stably at multiple viewing angles and under different lighting conditions; The image acquisition module is used to collect image data of cultural and creative products through multiple sets of high-resolution image sensors, and optimize image quality by combining automatic exposure and image contrast enhancement algorithms; The structural scanning module is used to obtain the three-dimensional structural information of cultural and creative products based on structured light scanning, laser ranging or multi-view reconstruction, and output standardized geometric model data.

[0006] Furthermore, the virtual culture display module includes: The semantic association module is used to label the historical background, symbolic meaning or cultural allusions of each cultural and creative product with structured tags and establish a mapping relationship with the corresponding cultural entities in the knowledge graph; A virtual scene construction module is used to automatically generate virtual cultural display content based on the tags and superimpose it on the user's field of view using augmented reality or virtual reality technology; The multimodal presentation module is used to simultaneously display cultural information to users in the form of images, voice, text and animation, realizing immersive interactive expression of cultural content.

[0007] Furthermore, the user interest perception module includes: The eye tracking acquisition module is used to record the user's gaze path in the virtual cultural content through an eye tracking device and convert it into a coordinate trajectory; The behavior recording module is used to record the user's interactive behaviors such as clicking, zooming in, and asking questions by voice during the virtual display process; The interest intensity calculation module is used to calculate the user's interest score for different cultural regions based on the gaze trajectory, stay duration and number of behaviors, and use this score for subsequent content region identification.

[0008] Furthermore, the content extraction module includes: A label matching module is used to retrieve corresponding cultural labels according to the coordinate position of the region of interest, including pattern names, totem semantics and story scene information; Image positioning module, used to extract the corresponding cultural element image area in the virtual cultural display interface through edge detection and region segmentation algorithm; The image standardization module is used to unify the format, normalize the size and remove the background of the extracted image to generate a high-quality standard input image.

[0009] Furthermore, the cultural creation generation module includes: Input encoding module, used to encode cultural and creative product image data into texture feature vectors, structural features into shape parameter vectors, and content images into cultural semantic vectors; The feature fusion module is used to fuse the three types of vectors into a unified input vector through splicing and dimensionality reduction, and use it as the input of the cultural and creative generation model.

[0010] Furthermore, the cultural creation generation module also includes: The style mapping fusion module is used to semantically match and spatially map the style features of user content images according to the structural area of cultural and creative products, and realize the fusion embedding of content images and product structures through scaling, rotation and alignment mechanisms.

[0011] Furthermore, the cultural creation generation module further includes: The cultural and creative generation model module includes a semantic-geometric coupling layer for determining the main pattern distribution of cultural content in the product structure based on the fusion vector; The symbolic content refinement layer is used to add cultural totems, legendary elements and match the visual style according to the user's interests; The image self-adjustment layer is used to adjust the pattern density, edge continuity and style coordination of the output image through a multi-objective optimization algorithm.

[0012] Furthermore, the cultural creation generation model module is constructed based on deep learning, and its training process includes: Collect product images, structural models, and cultural element images to form a training data set; The model is trained based on a joint loss function, which includes image reconstruction loss, style matching loss and structural consistency loss, and uses perceptual loss to enhance the expressiveness of cultural elements at the semantic layer.

[0013] Furthermore, after the personalized cultural and creative product image is output, it is also used to generate a three-dimensional printing model or a pattern silk-screen template, so as to realize the manufacturing and presentation of user-customized products in real space.

[0014] The present invention provides a virtual and real-world cultural and creative perception system, which has the following beneficial effects: This invention combines image data of real-world cultural and creative products with structural features, users' areas of interest in virtual cultural environments, and the cultural image content to construct a complete user-perception-driven generation mechanism, enabling the automated output of personalized cultural and creative images. This process not only reflects users' actual cultural preferences but also effectively improves the match between the generated results and their aesthetic intent.

[0015] The generative model proposed in the present invention has a triple structure of semantic-geometric coupling, symbolic content refinement and overall image self-adjustment. It can spatially adapt and style-fuse cultural patterns according to the actual form of the product, ensuring that the generated pattern is highly consistent with the physical surface of the product in visual structure, thereby solving the problems of pattern mapping distortion, structural dislocation and so on in the existing technology.

[0016] Through the fusion of reality and virtuality, this invention establishes a complete closed-loop process from product display to cultural perception and then to customized generation. It not only enhances the user's interactive experience and cultural identity, but also significantly improves the participation, customization and intelligence of the cultural and creative product design process. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0018] Figure 1 It is a system diagram of the method of the present invention. DETAILED DESCRIPTION

[0019] Below, the invention is preferably described with reference to the accompanying drawings and specific embodiments.

[0020] This embodiment solves the above problem through the following steps: In one embodiment, reference Figure 1 The present invention provides a cultural and creative perception system based on virtuality and reality. It is a cultural and creative perception system based on the deep integration of virtuality and reality. The system displays real cultural and creative products and presents virtual content of cultural background synchronously. It combines the user's behavioral feedback and interest perception during the immersive experience process, extracts user preference characteristics, and generates personalized cultural and creative product images based on preset cultural and creative products and cultural elements of interest to the user, thereby realizing the personalized expression of cultural content, intelligent reconstruction and virtual-real linkage, and enhancing the participation, emotion and customization value of cultural and creative products. The system specifically includes the following modules: The real cultural and creative display module is used to display at least one preset cultural and creative product and collect image data and appearance structure features of the at least one preset cultural and creative product.

[0021] To effectively connect real-world cultural and creative products with virtual cultural content within the cultural and creative perception system, the present invention incorporates a real-world cultural and creative display module for structured display of physical cultural and creative products and for collecting characteristic information such as their visual appearance and three-dimensional configuration. This module not only enables users to directly interact with and observe cultural and creative products in real space, but also provides fundamental data support for subsequent virtual content construction, user interest extraction, and personalized cultural and creative content generation.

[0022] In order to achieve the above functions, the real cultural and creative display module specifically includes the following sub-modules: The display control module is used to display pre-set cultural and creative products in the form of a display stand, support platform, or intelligent lighting platform. This module can be connected to the Internet of Things control system to adjust the rotation angle of the display stand, the intensity of the display light, and the color temperature, ensuring that the appearance details of the cultural and creative products can be presented under different viewing angles and lighting conditions. Preferably, the display platform uses a 360-degree rotating structure driven by a programmable motor to ensure that the image acquisition module can capture a complete image sequence.

[0023] The image acquisition module, used to capture static image data of cultural and creative products, integrates multiple industrial-grade cameras with a resolution of at least 12 megapixels and is equipped with a multi-angle light source system to reduce the effects of shadows and reflections. The images captured by this module are automatically numbered and calibrated to construct a pattern feature map.

[0024] The 3D structure scanning module is used to obtain the geometric appearance and structural features of cultural and creative products. Optional implementation solutions include: building a depth map model based on structured light scanning technology, acquiring 3D point cloud data using a laser ranging system, or building a high-precision 3D mesh model using multi-view stereo vision reconstruction technology. Preferably, structured light scanning technology is combined with a color-texture fusion algorithm to simultaneously restore the product's shape and surface texture.

[0025] The data preprocessing module filters, aligns, and standardizes the image and structural data. Image data is contrast-enhanced through bilateral filtering and gamma correction, while structural data is reconstructed into three dimensions using point cloud denoising and normal vector estimation algorithms. The resulting standardized data is packaged and transmitted to the subsequent semantic recognition and generation modules.

[0026] The virtual cultural display module is used to display the virtual cultural content corresponding to the at least one preset cultural and creative product, and the virtual cultural content includes one or more of the historical background, symbolic meaning, application scenarios, and extended interpretation of cultural elements of the cultural and creative product.

[0027] To enable users to gain a deeper understanding of the cultural meaning, historical origins, and symbolic value of real-world cultural and creative products while interacting with them, this invention incorporates a virtual cultural display module. This module aims to convey the intangible cultural connotations behind these products in a three-dimensional manner through visual, auditory, and verbal means, by constructing immersive virtual scenes or multimodal content presentations. This module not only enhances users' cultural perception but also provides key semantic input for subsequent interest identification and personalized creative generation, serving as a bridge between real-world display and intelligent generation.

[0028] In the present invention, virtual cultural content refers to cultural information expressed in digital form, including but not limited to the historical background, cultural symbolic meaning, usage scenarios, stories and legends, social functions or totemic meanings of cultural and creative products. Historical background refers to relevant information about the region, era, historical events or figures to which the cultural and creative product belongs. Symbolic meaning refers to the cultural symbols, spiritual values or social cognitive meanings contained in the cultural and creative product. Application scenarios refer to the actual usage scenarios of the product in traditional or modern society, such as festivals, weddings or daily life. The extended interpretation of cultural elements refers to the semantic interpretation and cultural mapping of external cultural carriers such as product patterns, materials, and shapes.

[0029] In order to achieve the above goals, the virtual cultural display module specifically includes the following sub-modules: The cultural semantic association module is used to map and associate preset cultural and creative products with their cultural semantic content. This association is based on a structured cultural knowledge graph stored in the product database. Graph nodes include product name, place of origin, symbols, and lineage, while edges represent cultural logical relationships such as "originating from," "evolved from," "used for," and "symbolize." Optional implementations include: constructing a cultural triple graph based on RDF (Resource Description Framework) or a cultural label propagation algorithm based on a graph neural network. Preferably, a product semantic graph database constructed by fusing manual annotation with a pre-trained model is used to improve cultural matching accuracy.

[0030] A virtual scene rendering module is used to generate a virtual cultural display scene based on the aforementioned semantic content. The scene can be a static 3D space, a dynamic image animation, a VR interactive environment, or an AR overlay view. Optional implementations include: 3D scene construction based on Unity or Unreal Engine rendering engines, and lightweight web rendering based on WebGL and Three.js. Preferably, for mobile users, a lightweight augmented reality presentation using ARCore or ARKit is used to reduce device dependency.

[0031] The multimodal presentation module is used to present the virtual cultural content to users in various formats, including images, text, voice, and video. This module includes a voice playback unit, a subtitle generation unit, and a text-and-graphics display unit. The voice playback unit supports dynamic explanations based on text-to-speech synthesis. The subtitle generation unit automatically generates and displays language and text synchronously by binding to semantic content nodes. The text-and-graphics display unit automatically pops up a text-and-graphics explanation panel of relevant cultural elements when the user gazes at a specific area.

[0032] The contextual interaction control module monitors user interactions with virtual cultural content and dynamically controls the progress of scene content, including operations such as perspective switching, content expansion, playback pause, and hotspot layer activation. This module uses a depth camera, gyroscope sensor, or eye tracking device to detect the user's head rotation and gaze angle, triggering virtual content stage switching or layer expansion.

[0033] Through the virtual cultural display module, users can experience cultural and creative products while understanding their deeper cultural connotations in a structured and perceptible manner, thereby establishing a continuous path from perceptual cognition to rational cultural construction. This module organically integrates cultural knowledge graphs, multimodal content generation, and interactive presentation mechanisms. This not only enhances the cultural depth and interactivity of cultural and creative displays, but also provides a clear source of cultural input for personalized content generation, elevating cultural and creative design from "style preference" to "cultural resonance."

[0034] For example, in the "Chenpi Pottery Jar" exhibition area of a cultural and creative exhibition hall, the virtual cultural display module, based on the pattern and shape of the pottery jar, calls on the node information of "Late Qing Dynasty Export Pottery", "Cantonese Red Glaze Bottle", "Nanyang Overseas Chinese Gift Culture" and other related knowledge graphs to construct a multi-dimensional scene that includes the gifting of overseas Chinese letters, the placement context of family temple offerings, and festival usage customs. After putting on AR glasses, users can watch the historical re-enactment of the pottery jar in the Jiangmen street market during the Qing Dynasty, while listening to the Cantonese oral explanation of folk customs such as "Chenpi is used as medicine" and "Pottery is a ritual vessel". When the user looks at the cloud pattern on the pottery jar, the layer automatically expands, and the virtual letter animation and audio playback content related to the "Overseas Chinese Letter Heritage" pop up.

[0035] The user interest perception module is used to identify the area of user interest in the virtual cultural content based on the user's gaze trajectory, stay duration and interactive behavior, and extract the content label and image information of the area.

[0036] To ensure that the display of virtual cultural content goes beyond information transmission and enables the identification of cultural interests based on user behavioral characteristics, thus providing precise cultural factor input for the subsequent generation of personalized cultural and creative products, the present invention incorporates a user interest perception module. This module collects multi-dimensional perceptual data such as user gaze trajectory, dwell time, and interactive behavior during the virtual display process, analyzes the intensity of the user's attention to specific cultural elements, identifies cultural areas of interest, and extracts relevant semantic tags and image content. This module serves as the core bridge for the "user perception - cultural positioning - generation trigger" process.

[0037] In the present invention, gaze trajectory refers to the path that the user's eyes follow in the virtual display interface, which can be obtained by an eye tracking device and reflects the flow of the user's visual attention. Dwell time refers to the time that the user gazes at or interacts with a certain cultural content area, which is usually counted in milliseconds. Interactive behavior includes active feedback from users on content through clicks, voice questions, gesture triggers, etc. Area of interest refers to a local area in the virtual content where the user shows high attention. Content label refers to the semantic annotation information associated with the area of interest, including cultural themes, symbolic meanings, craft styles, etc. Image information refers to the image data segment of the area of interest, which is used as input for subsequent image generation models.

[0038] In order to achieve the above functions, the user interest perception module specifically includes the following submodules: The eye tracking acquisition module is used to collect the user's eye tracking information when watching virtual cultural content. This module is implemented by integrating an eye tracking device based on infrared reflection or camera modeling, mapping the user's gaze coordinates at each moment to the spatial coordinate system of the virtual display interface. Preferably, binocular synchronous tracking technology and a 9-point calibration algorithm are used to improve the trajectory accuracy.

[0039] The behavior monitoring module is used to record the user's interactive behavior information in the virtual interface. The interactions include clicking hotspots, asking questions by voice, dragging the interface, and selecting by gaze. The behavior data is collected uniformly through event listeners and assigned timestamps and spatial annotations for subsequent correlation analysis.

[0040] The interest intensity calculation module is used to evaluate the interest intensity value of each cultural area based on the gaze trajectory and interaction behavior. The calculation formula of the interest intensity S is: ; Where T represents the length of time the user stays in the area in seconds, C represents the number of interactions in the area, D represents the gaze concentration (such as the pixel density indicator in the gaze heat map), and α, β, and γ are system-preset weighting coefficients that reflect the weights of different behaviors in interest modeling.

[0041] The interest region identification module is used to set the threshold based on the interest intensity calculation results , all areas with interest intensity greater than the threshold are selected as the user's area of interest, and their spatial coordinate bounding boxes are generated as the positioning basis for content labeling and image extraction.

[0042] The content label extraction module is used to search for the cultural labels of the corresponding area in the semantic database of virtual content based on the identified coordinates of the area of interest, including semantic information such as pattern name, source allusion, material meaning, totem affiliation, etc., and output a structured label set.

[0043] The image slice extraction module is used to crop the image content corresponding to the area of interest from the virtual display layer to generate image data of standardized size, which is used as the input image sample of the subsequent cultural and creative generation model. Preferably, the output size is set to 256×256 pixels and stored in PNG format.

[0044] By incorporating a user interest perception module, the present invention can automatically identify users' key interests within virtual cultural content without relying on active user input, and convert behavioral data into structured interest semantic information. This module not only maps user behavior to cultural semantics but also enhances the content relevance and individual expression capabilities of subsequent cultural creation processes, effectively addressing the problem of traditional cultural creation personalized generation lacking support from deep user interest data.

[0045] For example, when a user views the virtual cultural content of a "tangerine peel pot" through an AR device, their gaze lingers for a long time within the "cloud pattern," "Qiaopi letters," and "tangerine peel drying scene." The behavior monitoring module records that the user asks two voice questions about the "Qiaopi letters" and zooms in on the "cloud pattern." The interest intensity calculation module determines that the "cloud pattern" area has the highest interest value. The interest region identification module extracts the area corresponding to this pattern. The content label extraction module outputs cultural semantic labels such as "auspicious clouds," "safe travel," and "Qiaopi token." The image slice extraction module then captures image fragments of this pattern.

[0046] The content extraction module is used to determine at least one content image based on the content tag and image information, where the content image includes one or more combinations of cultural patterns, totem symbols, and story elements.

[0047] To transform user interests expressed in virtual cultural displays into high-quality input for generating personalized cultural and creative products, the present invention incorporates a content extraction module. This module, based on the content tags and image information identified by the user interest perception module, further selects image segments with clear cultural expression value. This module not only extracts and binds cultural semantics to visual data but also, through multi-level screening and synthesis, ensures that the extracted content and images effectively represent the user's cultural preferences, providing a clearly structured and semantically controllable input source for the subsequent cultural and creative generation module.

[0048] To achieve the above objectives, the content extraction module specifically includes the following submodules: The tag matching and screening module is used to select high-priority cultural elements that match the user's interest tags from the content tag library. The system uses a pre-set tag priority database and classifies them according to historical frequency, cultural weight, and image quality. Preferably, the system establishes a three-tiered tag hierarchy, with the first-tier tags representing core cultural elements, the second-tier representing auxiliary symbolic content, and the third-tier representing extended patterns or background styles. Image areas corresponding to the first-tier and second-tier tags are prioritized for extraction.

[0049] The image localization and boundary recognition module locates the image region corresponding to the target content label within the original virtual display image and performs boundary enhancement and region segmentation using a convolutional image processing algorithm. Optional implementations include contour extraction based on the Canny edge detection algorithm, foreground separation based on the GrabCut algorithm, or precise region extraction based on a semantic segmentation neural network model (such as U-Net). Preferably, a joint determination is made using a combination of user gaze hotspots and label matching regions to ensure that the extracted region maximally overlaps with the user's interest.

[0050] The image normalization module is used to normalize the size of the extracted image content, unify the format, and clean up the background, generating an image input format that is suitable for the subsequent generative model requirements. Preferably, the image is resized to 256×256 pixels and saved in PNG format with a transparent background. Bilateral filtering and histogram equalization are used to improve image quality, ensuring smooth edges and clear details.

[0051] The image combination module is used to execute a cultural element combination strategy to generate a composite content image when there are multiple interest content images. The combination strategy includes layer overlay, edge fusion, theme arrangement, etc. Preferably, it is based on the image feature vector similarity measurement function.

[0052] The content extraction module in this invention achieves an effective transition from interest identification to the extraction of high-value cultural visual content through semantically driven image positioning, structured image cropping, and a standardized output process. Compared to the coarse-grained input method of directly using user-browsed images, this module provides clearly structured, semantically focused, and quality-controlled content image data. This significantly improves the controllability, cultural fit, and visual expressiveness of subsequent cultural and creative generation models, effectively supporting the transition of cultural and creative products from "understanding users" to "customized content."

[0053] For example, a user focuses on the cloud pattern on a "Chenpi Pottery Jar" in a virtual display environment. The user interest perception module identifies the corresponding tags as "auspicious clouds" and "overseas Chinese remittance flow." After the content extraction module is activated, the tag matching and filtering module prioritizes the "cloud pattern main image" as the primary tag area.

[0054] The cultural and creative generation module is used to input the image data and appearance structure features of the at least one preset cultural and creative product and the at least one content image into the cultural and creative generation model based on deep learning to generate at least one personalized cultural and creative product image containing content of user interest.

[0055] To effectively integrate user interests expressed in virtual cultural content into the structural foundation of real-world cultural and creative products, and generate personalized cultural and creative product images with clear user characteristics, cultural orientation, and aesthetic expression, the present invention incorporates a cultural and creative product generation module. This module constructs a joint input vector based on the user's cultural content image and the image and structural features of the target cultural and creative product itself. This vector is then fed into a custom-built generation model. By integrating semantic consistency, structural adaptability, and aesthetic integrity, the module outputs a personalized image that reflects the user's cultural preferences and adapts to the original product's style, thereby enhancing the exclusivity and dissemination of cultural and creative products.

[0056] In order to achieve the above objectives, the cultural creation generation module may include the following submodules: The input construction module is used to encode the preset cultural and creative product image data, appearance structural features, and user interest content images into a unified feature representation. Texture and color features are extracted from the image data using a multi-channel convolution-based encoding network. Structural features are encoded using a coordinate transformation-based morphological vector encoding method to generate a 3D geometric description. The content images are then transformed into cultural vectors using a semantic embedding model. Together, these three components form the input combination of the generative model.

[0057] The style mapping fusion module is used to fuse the content image and product structure in terms of style, scale and visual position in a unified embedding space. Preferably, the guided mapping matrix M is used to represent the pattern projection weight; the formula is: ; in, is the semantic tag of user content images, The element represents the Cultural elements, Product structure Semantic labels for regions, is the semantic matching function between the two, is the fusion energy function, which is used to control the degree of natural edge transition. The fusion result is used as the input benchmark for the next stage of model generation.

[0058] The cultural and creative generation model module is used to generate the final personalized image. The model consists of three core layers: The first layer is the semantic-geometric coupling layer, which inputs the joint semantic and structural vectors to perform the main structure layout of the pattern; The second layer is the symbolic content refinement layer, which adds elements such as cultural totems and story images, and embeds them by scaling, rotating and embedding according to the corresponding style relationships; The third layer is the overall image self-adjustment layer, which controls local density, edge continuity and overall style consistency through a multi-objective optimization function.

[0059] This model does not rely on traditional image stitching or a single style transfer method. Instead, it is based on the geometric semantic partitioning of the product and guides the adaptive embedding of content images through a semantic resonance mechanism. It is suitable for the surface graphic redesign task of complex structural products and is self-explanatory and transparent in generation.

[0060] The image output and verification module formats, adjusts resolution, and verifies usability of generated images. Output formats include PNG, SVG, or UV maps to accommodate different manufacturing processes. Verification steps include checking pattern boundary closure, alignment with the original product structure, and scoring user semantic preference matching.

[0061] Furthermore, after confirming that it meets the output standards, the system will use the image as a personalized cultural and creative product pattern plan, supporting users to export the image or enter the next customization step, such as 3D printing, manual molding or digital silk screen printing.

[0062] The cultural and creative generation module of this invention introduces a joint modeling mechanism for product structure, cultural content, and user preferences, proposing a novel generation model that combines semantic matching and structural adaptation. This effectively closes the logical loop from cultural cognition-driven to personalized product expression. This generation approach transcends the limitations of existing pattern splicing and texture mapping, proactively generating innovative cultural and creative images with artistic style, structural compatibility, and cultural expression based on user cultural resonance. This significantly enhances the intelligence, cultural depth, and user exclusivity of cultural and creative design.

[0063] For example, when browsing the virtual cultural extension of a tangerine peel pot, users showed significant interest in the "cloud totem," "Qiaopi symbols," and "Guangcai red glaze" tones. The cultural and creative generation module encodes the pottery's image texture and structural features, while also performing cultural semantic annotation and geometric adjustments on the user-selected composite image of "postal route graphics + auspicious clouds." The SSC-CNet model automatically determines that the pottery's belly and mouth areas are suitable for embedding cruise ship images and auspicious cloud patterns, respectively. Through a semantically driven pattern layout and contour alignment mechanism, it generates a customized pottery image that combines a red glaze base, light gold cloud patterns, and a postal totem. The image is then output as a high-precision UV map for subsequent customized printing or pottery mold carving, seamlessly integrating cultural perception with personalized creation.

[0064] In a further implementation, the cultural creation generation model module specifically includes: The semantic-geometric coupling layer maps the semantic vectors of user-interested images with the geometric structural vectors of cultural and creative products, establishing the primary structural layout of cultural patterns on the product's surface. This layer embeds the semantic encoding of the content into the product's structural grid, determining the projection areas appropriate for each cultural factor and completing the basic spatial arrangement.

[0065] This layer receives two input vectors, one of which is the semantic feature vector extracted from the content image , and the second is the three-dimensional structure vector of the product appearance structure , combining the two to construct a coupled representation , the specific calculation formula is: ; Among them, Concat represents the vector concatenation operation, and MLP represents the multi-layer perceptron network used for compression and nonlinear expression of structural information.

[0066] The coupled representation is input into the pattern partitioning and layout module, which automatically divides the cultural elements into matching areas in the product surface structure, such as the belly, bottleneck, and lid of the pottery jar, and sets the initial geometric bounding box position and orientation angle.

[0067] This layer achieves high-level adaptation between cultural semantic content and product geometric structure, ensuring a reasonable distribution of the main structure of the generated pattern and avoiding embedding cultural patterns into inappropriate areas of the product. It also provides a spatial distribution basis for subsequent detail filling and enhances the coordination between the pattern layout and the physical space of the product.

[0068] The symbolic content refinement layer builds upon the primary structural layout by detailing and refining the symbolic cultural elements within the pattern, such as totems, characters, and story points. This layer achieves a deeper expression of semantic consistency through multi-scale scaling, angle rotation, style matching, and structural embedding algorithms.

[0069] This layer targets each semantic element , based on its semantic category (such as "dragon pattern", "postmark symbol", "auspicious cloud"), call the corresponding texture attribute vector in the style parameter library , and perform the transformation operation: Scaling function , control the pattern size to match the target area ; Rotation function , control the unity of visual direction; Style mapping function , complete pattern style generation and morphology matching.

[0070] The above results are projected into the structural boundary to form a composite layer. Multiple layers can be blended according to transparency and edge blur parameters to achieve visual fusion between the cultural totem and the carrier surface material.

[0071] This layer embeds user-interested symbolic elements into the main pattern without compromising structural integrity, achieving a continuous expression from "semantic recognition" to "visual restoration." This adjustable transformation strategy enhances the adaptability of pattern details and the expressive tension of cultural elements, helping to create a recognizable visual center with cultural depth.

[0072] The overall image self-adjustment layer is used to globally optimize the overall visual structure of the image after the fused pattern is generated. This includes controlling pattern density distribution, boundary continuity, color consistency, and overall product style coordination. This layer uses a multi-objective loss function to construct the optimization objective and uses an iterative feedback mechanism to adjust the local and overall state of the generated image.

[0073] This layer introduces multiple discriminant indicators, including: Local density index , evaluate the pixel concentration of the pattern area; Boundary continuity index , judge the smoothness of the connection between the pattern edge and the product surface; Style consistency index , used to detect whether the overall style of the generated image matches the style of the original product.

[0074] The total loss function is constructed as follows: ; in, It is a weight coefficient. Different values can be set according to product characteristics to achieve visual adjustment effects.

[0075] The model optimizes the weight updates of the first two layers through back-propagation to achieve overall optimization of the visual generation results.

[0076] This layer achieves global consistency adjustment of the generated image at the end of the model, improving the integrity, style uniformity and product adaptability of the pattern at the perceptual level, effectively avoiding common defects such as "splicing traces", "color faults" and "crowded arrangement", so that personalized images meet the artistic standards and user acceptability required for industrial production.

[0077] Through a three-layer collaborative structure, the cultural and creative generation model not only achieves coupled input of semantics and structure, but also completes the detailed filling of symbolic content and the optimization of the final image style. This structure does not rely on traditional style transfer or image collage mechanisms, but is instead constructed based on the linkage of cultural logic, structural space, and visual harmony logic. It possesses stronger cultural expression capabilities, higher image generation quality, and stronger product adaptability, making it suitable for the rapid customization and intelligent design of personalized cultural and creative products.

[0078] Furthermore, the training process of the generative model includes: The dataset includes: Images and three-dimensional structures of cultural and creative products (e.g., images of pottery + 3D models); Matching cultural element images (totems, patterns, symbols, etc.); Annotated cultural semantic labels (such as "dragon pattern", "Qiaopi", "auspicious clouds"); The target patterns manually generated by the experts after fusion are used as supervision targets.

[0079] Adopt an end-to-end training strategy; It is preferred to use the Adam optimizer with an initial learning rate of 0.0001; A layer-by-layer pre-training + joint fine-tuning approach is adopted, first training the semantic-structural coupling layer, then refining the embedding layer, and finally overall tuning.

[0080] After each round of training, candidate generated images are output; expert scoring + cultural recognition network are used to assist in judging image quality; the validation set uses indicators such as cultural pattern distribution, pattern density, and structural boundary overlap to evaluate generation quality.

[0081] The prior art mentioned in the background technology section and the specific embodiments section of the present invention can be regarded as part of the present invention and used to understand the meaning of some technical features or parameters.

Claims

1. A cultural and creative perception system based on virtuality and reality, characterized by: The system includes the following modules: A real-life cultural and creative display module, configured to display at least one preset cultural and creative product and collect image data and appearance and structural features of the at least one preset cultural and creative product; a virtual cultural display module, configured to display virtual cultural content corresponding to the at least one preset cultural and creative product, wherein the virtual cultural content includes one or more of the historical background, symbolic meaning, application scenarios, and extended interpretation of cultural elements of the cultural and creative product; A user interest perception module is used to identify user interest areas in the virtual cultural content based on the user's gaze trajectory, dwell time, and interactive behavior, and to extract content labels and image information of the interest areas; a content extraction module, configured to determine at least one content image based on the content tag and the image information, wherein the content image includes one or more combinations of cultural patterns, totem symbols, and story elements; The cultural and creative generation module is used to input the image data and appearance structure features of the at least one preset cultural and creative product and the at least one content image into the cultural and creative generation model based on deep learning to generate at least one personalized cultural and creative product image containing content of user interest.

2. The cultural and creative perception system based on virtuality and reality according to claim 1 is characterized in that: The reality cultural and creative display module includes: The display control module is used to display cultural and creative products through display racks, support platforms or electric rotating devices, and adjust the angle of the booth and the intensity of the lights to ensure that the cultural and creative products are displayed stably at multiple viewing angles and under different lighting conditions; The image acquisition module is used to collect image data of cultural and creative products through multiple sets of high-resolution image sensors, and optimize image quality by combining automatic exposure and image contrast enhancement algorithms; The structural scanning module is used to obtain the three-dimensional structural information of cultural and creative products based on structured light scanning, laser ranging or multi-view reconstruction, and output standardized geometric model data.

3. The cultural and creative perception system based on virtuality and reality according to claim 1 is characterized in that: The virtual culture display module includes: The semantic association module is used to label the historical background, symbolic meaning or cultural allusions of each cultural and creative product with structured tags and establish a mapping relationship with the corresponding cultural entities in the knowledge graph; A virtual scene construction module is used to automatically generate virtual cultural display content based on the tags and superimpose it on the user's field of view using augmented reality or virtual reality technology; The multimodal presentation module is used to simultaneously display cultural information to users in the form of images, voice, text and animation, realizing immersive interactive expression of cultural content.

4. The cultural and creative perception system based on virtuality and reality according to claim 1 is characterized in that: The user interest perception module includes: The gaze trajectory acquisition module is used to record the user's gaze path in the virtual cultural content through an eye tracking device and convert it into a coordinate trajectory; The behavior recording module is used to record the user's interactive behaviors such as clicking, zooming in, and asking questions by voice during the virtual display process; The interest intensity calculation module is used to calculate the user's interest score for different cultural regions based on the gaze trajectory, stay duration and number of behaviors, and use this score for subsequent content region identification.

5. The cultural and creative perception system based on virtuality and reality according to claim 1 is characterized in that: The content extraction module includes: A label matching module is used to retrieve corresponding cultural labels according to the coordinate position of the region of interest, including pattern names, totem semantics and story scene information; Image positioning module, used to extract the corresponding cultural element image area in the virtual cultural display interface through edge detection and region segmentation algorithm; The image standardization module is used to unify the format, normalize the size and remove the background of the extracted image to generate a high-quality standard input image.

6. The cultural and creative perception system based on virtuality and reality according to claim 1 is characterized in that: The cultural creation generation module includes: Input encoding module, used to encode cultural and creative product image data into texture feature vectors, structural features into shape parameter vectors, and content images into cultural semantic vectors; The feature fusion module is used to fuse the three types of vectors into a unified input vector through splicing and dimensionality reduction, and use it as the input of the cultural and creative generation model.

7. The cultural and creative perception system based on virtuality and reality according to claim 1 is characterized in that: The cultural creation generation module also includes: The style mapping fusion module is used to semantically match and spatially map the style features of user content images according to the structural area of cultural and creative products, and realize the fusion embedding of content images and product structures through scaling, rotation and alignment mechanisms.

8. The cultural and creative perception system based on virtuality and reality according to claim 1 is characterized in that: The cultural creation generation module further includes: The cultural and creative generation model module includes a semantic-geometric coupling layer for determining the main pattern distribution of cultural content in the product structure based on the fusion vector; The symbolic content refinement layer is used to add cultural totems, legendary elements and match the visual style according to the user's interests; The image self-adjustment layer is used to adjust the pattern density, edge continuity and style coordination of the output image through a multi-objective optimization algorithm.

9. The cultural and creative perception system based on virtuality and reality according to claim 8 is characterized in that: The cultural creation generation model module is built based on deep learning, and its training process includes: Collect product images, structural models, and cultural element images to form a training data set; The model is trained based on a joint loss function, which includes image reconstruction loss, style matching loss and structural consistency loss, and uses perceptual loss to enhance the expressiveness of cultural elements at the semantic layer.

10. The cultural and creative perception system based on virtuality and reality according to claim 1 is characterized in that: After the personalized cultural and creative product image is output, it is also used to generate a three-dimensional printing model or a pattern silk-screen template, thereby realizing the manufacture and presentation of user-customized products in real space.

Citation Information

Cited By

  • Block chain-based scenic spot digital text and creative interaction method

    CN121413747A