Lake xiang wood carving image generation method, device and equipment based on LoRA model and storage medium
By constructing a feature library of Hunan wood carving production process and training the LoRA model in layers, and combining it with text prompts to generate Hunan wood carving images, the problems of low generation efficiency and inaccurate capture of process logic in existing technologies are solved, and efficient process restoration and scene adaptation are achieved.
Patent Information
- Application Number
- CN202511645495.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-11
- Publication Date
- 2026-02-06
- Estimated Expiration
- 2045-11-11
AI Technical Summary
Existing technologies struggle to generate images of Hunan wood carvings that combine high fidelity to the craftsmanship with scene adaptability. High-precision scanning equipment is costly and inefficient, and general generation models cannot accurately capture the unique craftsmanship logic of Hunan wood carvings.
A feature library of Hunan wood carving production process was constructed. The initial LoRA model was trained in layers according to the hierarchical logic of texture layer-carving layer-pattern layer, and then fused with the pre-trained latent diffusion model in layers. The Hunan wood carving image was generated by parsing the text prompt words input by the user.
It has achieved the generation of Hunan wood carving images that combine high fidelity to the craftsmanship and adaptability to different scenes, improving generation efficiency and accurately capturing craftsmanship features.
Smart Images

Figure CN121095382B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of process digitization and image generation, in particular to a method and device for generating images of Hunan wood carving based on a LoRA model, equipment and a storage medium. BACKGROUND
[0002] With the promotion of digital protection of intangible cultural heritage and digital transformation of cultural and creative industry, as a traditional craft carrying regional folk culture, the accurate generation of digital images of Hunan wood carving has become a core demand of technical application. From a technical point of view, the exclusive process features of Hunan wood carving need to be reproduced through digital means, including the texture parameters (texture direction, pore density, annual ring spacing) of wood (such as camphor wood and nanmu), the knife work techniques (direction of operation of relief, openwork and sculpture, carving depth, spacing of knife marks), and the morphological parameters and combination logic of traditional patterns (dragon and phoenix pattern, cloud and thunder pattern, and fret pattern). At the same time, technical adaptation of these process features to folk scenes such as marriage, sacrifice and festivals in terms of size, light and shadow and space for placement is also needed. In addition, the digital image generation technology also needs to meet the high-fidelity requirements of intangible cultural heritage archives and the flexible iteration requirements of cultural and creative design, and provide technical support for subsequent digital display and virtual restoration.
[0003] Currently, there are two types of technical solutions for the generation of digital images of traditional crafts: one is a technical path based on high-precision scanning and image processing, which collects images and three-dimensional data of wood carving objects through three-dimensional laser scanners and high-resolution cameras, and uses image repair algorithms and texture mapping techniques for post-processing to output digital images; the other is to use general generation models such as generative adversarial networks (GAN) to trigger image generation by inputting text prompts.
[0004] The technical solution based on high-precision scanning and image processing is limited by the stock of physical objects, and cannot generate non-existent wood carving styles through technical means. The cost of scanning equipment is high, the data collection period is long, and the post-processing of images relies on manual intervention, which is low in technical efficiency and difficult to adapt to diversified design requirements. The technical solution based on general generation models cannot accurately capture the adaptive relationship between wood texture and knife marks, and the corresponding rules between pattern structure and school style, as the pre-trained model does not deeply learn the exclusive process logic of Hunan wood carving, resulting in low process restoration of the generated images. Therefore, how to generate Hunan wood carving images with process restoration and scene adaptability has become a problem to be solved. SUMMARY
[0005] The present application aims to provide a method and device for generating images of Hunan wood carving based on a LoRA model, equipment and a storage medium, which aims to solve the technical problem of how to generate Hunan wood carving images with process restoration and scene adaptability.
[0006] To achieve the above object, the application provides a Lake Xiang wood carving image generation method based on a LoRA model, which comprises the following steps:
[0007] Constructing a Lake Xiang wood carving production process feature library according to a scanned image of a real Lake Xiang wood carving, wood texture data and a production process video;
[0008] According to the Lake Xiang wood carving production process feature library, the initial LoRA model is trained in layers according to the hierarchical logic of the texture layer, the tooling layer and the pattern layer to obtain a layered LoRA model;
[0009] The layered LoRA model is fused with a pre-trained latent diffusion model in layers to obtain a Lake Xiang wood carving image generation model;
[0010] Based on the keyword system of the Lake Xiang wood carving production process feature library, a text prompt word input by a user containing a wood type and a folk scene is analyzed to obtain a wood carving type, wood texture parameters, folk pattern requirements and scene adaptation features;
[0011] The wood carving type, the wood texture parameters, the folk pattern requirements and the scene adaptation features are input into the Lake Xiang wood carving image generation model to obtain a target Lake Xiang wood carving image.
[0012] In an embodiment, the layered LoRA model comprises a texture layer, a tooling layer and a pattern layer;
[0013] The step of fusing the layered LoRA model with a pre-trained latent diffusion model in layers to obtain a Lake Xiang wood carving image generation model comprises the following steps:
[0014] A pre-trained latent diffusion model is obtained, the latent diffusion model comprising a feature encoding module, a diffusion sampling module and an image decoding module, the feature encoding module comprising a texture generation submodule, the diffusion sampling module comprising a detail rendering submodule, and the image decoding module comprising a composition generation submodule;
[0015] According to the interface parameters of the texture generation submodule, the detail rendering submodule and the composition generation submodule, a mapping relationship table of the main module and the submodules is established;
[0016] According to the mapping relationship table, the texture layer is connected to the texture generation submodule, the tooling layer is connected to the detail rendering submodule, and the pattern layer is connected to the composition generation submodule to obtain an initial fusion model;
[0017] An initial fusion weight is determined according to a preset Lake Xiang wood carving process priority, and a hierarchical fusion unit is constructed according to the initial fusion weight;
[0018] A reference fusion model is obtained based on the initial fusion model and the hierarchical fusion unit;
[0019] A sample dataset is selected from the feature library of Hunan wood carving production process, and the reference fusion model is jointly fine-tuned based on the sample dataset to obtain the Hunan wood carving image generation model.
[0020] In one embodiment, the step of selecting a sample dataset from the feature library of the Hunan wood carving production process and jointly fine-tuning the reference fusion model based on the sample dataset to obtain the Hunan wood carving image generation model includes:
[0021] A sample dataset is selected from the Hunan wood carving production process feature library, which includes a pattern rule sub-library, a wood texture sub-library, and a carving feature sub-library. The sample dataset includes the texture data of the wood texture sub-library, the operation data of the carving feature sub-library, the structural data of the pattern rule sub-library, and the corresponding standard wood carving image.
[0022] The image data in the sample dataset is converted into a feature matrix of a preset size, and the process data is converted into a vector to obtain a fine-tuned dataset.
[0023] The fine-tuned dataset is input into the reference fusion model, and the cross-layer collaborative loss between the texture layer and the texture generation submodule, the knife-cutting layer and the detail rendering submodule, and the pattern layer and the composition generation submodule is calculated.
[0024] Based on the cross-layer collaborative loss, the fusion weights of each docking layer are adjusted through the hierarchical fusion unit, and the matching degree of the process features between the generated image and the standard wood carving image is calculated using the verification sample calculation model.
[0025] When the matching degree of the process features is greater than the preset matching degree threshold, the Hunan wood carving image generation model is obtained.
[0026] In one embodiment, the Hunan wood carving production process feature library includes a pattern rule sub-library, a wood texture sub-library, and a knife work feature sub-library;
[0027] The steps of training the initial LoRA model hierarchically according to the feature library of Hunan wood carving production process, following the hierarchical logic of texture layer - carving layer - pattern layer, to obtain the hierarchical LoRA model include:
[0028] The LoRA model is initialized according to the preset learning rate and preset batch size to obtain the initial LoRA model;
[0029] Based on the wood texture sub-library, the texture layer of the initial LoRA model is trained by feature fitting to obtain the texture layer feature weights;
[0030] Based on the texture layer feature weight, the tooling layer of the initial LoRA model is linked and trained according to the tooling feature sub-library, to obtain a tooling layer feature weight;
[0031] In combination with the texture layer feature weight and the tooling layer feature weight, the pattern layer of the initial LoRA model is cooperatively trained according to the pattern rule sub-library, to obtain a pattern layer feature weight;
[0032] Based on a hierarchical fusion loss function, a deviation value of output features of the texture layer, the tooling layer and the pattern layer from standard features in the Hunan wood carving process feature library is calculated, and a weight proportion of each level is allocated according to the deviation value;
[0033] Based on the weight proportion, a matching degree of output features of the initial LoRA model from the standard features is calculated;
[0034] If the matching degree is less than a preset matching threshold, a learning rate of the initial LoRA model is adjusted by a preset adjustment amplitude, and the steps of hierarchical training and weight allocation are repeated until the matching degree is greater than or equal to the preset matching threshold, to obtain a hierarchical LoRA model.
[0035] In an embodiment, the step of constructing the Hunan wood carving process feature library according to the real object scanning graph, wood texture data and manufacturing process video of the Hunan wood carving includes:
[0036] The real object scanning graph, wood texture data and manufacturing process video of the Hunan wood carving are collected, and denoising and standardization processing are performed;
[0037] The pattern structure features are extracted from the processed real object scanning graph, the texture parameter features are extracted from the processed wood texture data, and the tooling operation features are extracted from the processed manufacturing process video;
[0038] The pattern structure features, the texture parameter features and the tooling operation features are bound according to process association logic to obtain target feature data, and corresponding folk scene labels are labeled for the target feature data;
[0039] Based on the target feature data and the folk scene labels, a pattern rule sub-library, a wood texture sub-library, a tooling feature sub-library and a scene adaptation sub-library are constructed;
[0040] The wood texture sub-library, the tooling feature sub-library, the pattern rule sub-library and the scene adaptation sub-library are integrated, and a cross-sub-library retrieval index is established, to obtain a Hunan wood carving process feature library.
[0041] In an embodiment, the keyword system based on the characteristic library of the Hu Xiang wood carving production process is used to analyze the text prompt word input by the user, which contains wood type and folk scene, to obtain wood carving type, wood texture parameter, folk pattern demand, and scene adaptation characteristic.
[0042] The text prompt word input by the user is segmented, to obtain wood type keyword, folk scene keyword, use keyword, and meaning keyword.
[0043] The keyword system of the characteristic library of the Hu Xiang wood carving production process is associated and matched with the wood type keyword, the folk scene keyword, the use keyword, and the meaning keyword, to obtain wood type matching item, scene matching item, use matching item, and meaning matching item.
[0044] According to the wood type matching item, the corresponding wood texture parameter is determined from the wood texture sub-library.
[0045] According to the use matching item and the scene matching item, the wood carving type is determined from the scene adaptation sub-library.
[0046] According to the scene matching item and the meaning matching item, the folk pattern demand is determined from the pattern rule sub-library.
[0047] According to the scene matching item, the corresponding scene adaptation characteristic is determined from the scene adaptation sub-library.
[0048] In an embodiment, the Hu Xiang wood carving image generation model includes a feature encoding module, a diffusion sampling module, and an image decoding module.
[0049] The wood carving type, the wood texture parameter, the folk pattern demand, and the scene adaptation characteristic are input into the Hu Xiang wood carving image generation model, to obtain a target Hu Xiang wood carving image.
[0050] The wood carving type, the wood texture parameter, the folk pattern demand, and the scene adaptation characteristic are converted into a feature vector recognizable by the Hu Xiang wood carving image generation model.
[0051] The feature vector is input into the feature encoding module for encoding processing, to obtain an encoded feature.
[0052] The encoded feature is input into the diffusion sampling module, to obtain sampling feature data.
[0053] The sampling feature data is input into the image decoding module, to obtain pixel-level image data.
[0054] The pixel-level image data is subjected to format standardization processing to obtain a target Hunan wood carving image.
[0055] In addition, to achieve the above-mentioned purpose, the present application also proposes a Hunan wood carving image generation device based on a LoRA model, which comprises:
[0056] A feature library construction module is configured to construct a Hunan wood carving manufacturing process feature library according to a scanning image of a real Hunan wood carving, wood texture data and manufacturing process videos;
[0057] A hierarchical training module is configured to perform hierarchical training on an initial LoRA model according to the Hunan wood carving manufacturing process feature library, according to the hierarchical logic of the texture layer-chisel work layer-pattern layer, to obtain a hierarchical LoRA model;
[0058] A model fusion module is configured to perform hierarchical fusion of the hierarchical LoRA model and a pre-trained latent diffusion model to obtain a Hunan wood carving image generation model;
[0059] A prompt word analysis module is configured to analyze text prompt words input by a user, which contain wood types and folk scenes, based on a keyword system of the Hunan wood carving manufacturing process feature library, to obtain wood carving types, wood texture parameters, folk pattern requirements and scene adaptation features;
[0060] An image generation module is configured to input the wood carving types, the wood texture parameters, the folk pattern requirements and the scene adaptation features into the Hunan wood carving image generation model to obtain a target Hunan wood carving image.
[0061] In addition, to achieve the above-mentioned purpose, the present application also proposes a Hunan wood carving image generation device based on a LoRA model, which comprises: a memory, a processor and a computer program stored on the memory and executable on the processor, the computer program being configured to implement the steps of the above-mentioned Hunan wood carving image generation method based on a LoRA model.
[0062] In addition, to achieve the above-mentioned purpose, the present application also proposes a storage medium, which is a computer-readable storage medium, and the storage medium stores a computer program, and the computer program is executed by a processor to implement the steps of the above-mentioned Hunan wood carving image generation method based on a LoRA model.
[0063] In addition, to achieve the above-mentioned purpose, the present application also provides a computer program product, which comprises a computer program, and the computer program is executed by a processor to implement the steps of the above-mentioned Hunan wood carving image generation method based on a LoRA model.
[0064] The one or more technical solutions provided in the application have at least the following technical effects:
[0065] First, according to the physical scanning diagram of Hunan wood carving, the wood texture data and the production process video, a Hunan wood carving production process feature library is constructed to provide core data support for subsequent links. Then, according to the feature library, the initial LoRA model is trained in layers according to the hierarchical logic of the texture layer, the tooling layer and the pattern layer, to obtain a layered LoRA model, so that the model can accurately master the core features of Hunan wood carving technology. Then, the layered LoRA model is fused with the pre-trained latent diffusion model to obtain a Hunan wood carving image generation model, which takes into account the generation efficiency and the professional nature of the technology. After that, based on the keyword system of the feature library, the text prompt words input by the user containing wood types and folk scenes are analyzed to obtain wood carving types, wood texture parameters, folk pattern requirements and scene adaptation features, realizing the accurate conversion of user requirements. Finally, these structured features are input into the Hunan wood carving image generation model to obtain the target Hunan wood carving image. The application can generate Hunan wood carving images with process restoration and scene adaptability. BRIEF DESCRIPTION OF DRAWINGS
[0066] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the application and serve to explain the principles of the application together with the specification.
[0067] In order to more clearly illustrate the technical solutions in the embodiments of the application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, for those skilled in the art, other drawings can also be obtained based on these drawings without creative labor.
[0068] Figure 1 A process schematic diagram is provided for the first embodiment of the application, i.e., the Hunan wood carving image generation method based on the LoRA model.
[0069] Figure 2 A process schematic diagram is provided for the second embodiment of the application, i.e., the Hunan wood carving image generation method based on the LoRA model.
[0070] Figure 3 A module structure schematic diagram of the Hunan wood carving image generation device based on the LoRA model is provided for the embodiment of the application.
[0071] Figure 4 A device structure schematic diagram of the hardware running environment involved in the Hunan wood carving image generation method based on the LoRA model is provided for the embodiment of the application.
[0072] The purpose of the application, functional characteristics and advantages will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION
[0073] It should be understood that the specific embodiments described herein are merely intended to explain the technical solutions of the present application, and are not intended to limit the present application.
[0074] In order to better understand the technical solutions of the present application, the following will be described in detail in combination with the drawings of the specification and specific embodiments.
[0075] It should be noted that the execution subject of the embodiments of the present application can be a computing service device with data processing, network communication and program running functions, such as a tablet computer, a personal computer, a mobile phone, etc., or an electronic device, an image generation system, etc. capable of realizing the above functions. The embodiments of the present application and the following embodiments will be described below taking the image generation system as an example.
[0076] Based on this, the present application provides a LoRA model-based image generation method for Hunan wood carving, which is described in detail below with reference to Figure 1 , Figure 1 The flowchart of the first embodiment of the LoRA model-based image generation method for Hunan wood carving of the present application is shown in the figure.
[0077] In the present embodiment, the LoRA model-based image generation method for Hunan wood carving comprises steps S10-S50:
[0078] Step S10, constructing a Hunan wood carving production process feature library according to the scanning map of the real object, the wood texture data and the production process video of Hunan wood carving.
[0079] It should be noted that the wood texture data refers to the texture and material related data presented by the wood used for Hunan wood carving at the micro and macro levels, including texture direction, pore density, annual ring spacing, material hardness and other parameters, which are used to accurately describe the surface texture and inherent material characteristics of the wood. The production process video refers to the video material recording the whole process operation of Hunan wood carving from preparation, design, carving to polishing, focusing on capturing the knife work techniques (such as knife direction, carving depth), tool use and process operation details, providing intuitive basis for extracting knife work features. The Hunan wood carving production process feature library refers to a knowledge set about the modeling rules, material performance, process logic and style elements of Hunan wood carving, which is systematically extracted and structured stored based on multi-modal data such as scanning map of real object, wood texture data and production process video, aiming to provide semantic support with process authenticity for subsequent image generation, style transfer or digital restoration tasks.
[0080] As an example, the step of constructing the Huxiang wood carving production process feature library according to the scanned image of the Huxiang wood carving, the wood texture data and the production process video includes: collecting the scanned image of the Huxiang wood carving, the wood texture data and the production process video, and performing denoising and standardization processing; extracting pattern structure features from the processed scanned image, extracting texture parameter features from the processed wood texture data, and extracting knife operation features from the processed production process video; binding the pattern structure features, the texture parameter features and the knife operation features according to process association logic to obtain target feature data, and labeling corresponding folk scene tags for the target feature data; based on the target feature data and the folk scene tags, constructing a pattern rule sub-library, a wood texture sub-library, a knife feature sub-library and a scene adaptation sub-library; integrating the wood texture sub-library, the knife feature sub-library, the pattern rule sub-library and the scene adaptation sub-library, and establishing cross-sub-library retrieval indexes to obtain the Huxiang wood carving production process feature library.
[0081] The pattern structure features refer to the feature information extracted from the scanned image of the Huxiang wood carving, reflecting the core attributes of traditional patterns, including the shape parameters (such as line curvature, pattern symmetry) of the patterns, the combination logic (such as the arrangement method of primary and secondary patterns) and the implied association relationship. The texture parameter features refer to the key parameters extracted from the wood texture data, describing the texture and material characteristics of the wood used for Huxiang wood carving, including texture direction, pore density, annual ring spacing, material hardness, etc. The knife operation features refer to the operation-related features extracted from the production process video, reflecting the carving process of Huxiang wood carving, including the direction of knife movement, carving depth, knife movement frequency, knife mark spacing and the type of wood suitable for carving. The process association logic refers to the internal adaptation rules of each process link in the production of Huxiang wood carving, including the adaptation rules of wood texture and knife technique, and the matching rules of knife operation and pattern structure, which are used to ensure the rationality of feature binding. The target feature data refers to the complete feature set formed by structurally binding the pattern structure features, the texture parameter features and the knife operation features according to the process association logic, containing the core feature information of the whole wood carving process.
[0082] The folk scene tag refers to the classification mark corresponding to the application scene and use demand of the target feature data label, including folk scene categories such as marriage, sacrifice, and festival, and display space information such as central hall and ancestral hall. The pattern rule sub-library refers to a special sub-library constructed based on the pattern structure features in the target feature data, which stores pattern-related rules and parameters classified by Huxiang wood carving schools (such as Xiangdong, Xiangxi, and Xiangzhong) and pattern types (such as dragon and phoenix patterns and cloud and thunder patterns). The wood texture sub-library refers to a special sub-library constructed based on the texture parameter features in the target feature data, which stores wood texture and material parameter information classified by wood types (such as camphor wood and nanmu). The tool worker feature sub-library refers to a special sub-library constructed based on the tool worker operation features in the target feature data, which is associated with corresponding wood types and stores operation parameters and process standards of various tool worker techniques. The scene adaptation sub-library refers to a special sub-library constructed based on the folk scene tag, which stores scene adaptation information such as wood carving size parameters, process types, display requirements, and light and shadow adaptation parameters classified by scene types. The cross-sub-library retrieval index refers to an index structure established for fast association query between various special sub-libraries, which establishes feature mapping relationships between sub-libraries through association items such as wood types and scene tags, and supports joint retrieval of multi-dimensional process features.
[0083] Firstly, the image generation system collects the physical scanning image of Huxiang wood carving, wood texture data and production process video, and carries out denoising processing (such as removing reflection spots in the scanning image and screen jitter in the video) and standardization processing (such as unifying the image resolution to a preset value and the video frame rate to a preset value) on these data, eliminating interference information and unifying data format, providing high-quality data basis for subsequent feature extraction. Then, the image generation system extracts pattern structure features from the processed physical scanning image through image segmentation and feature recognition algorithm (such as mask R-CNN (Mask Region-based Convolutional Neural Network)); extracts texture parameter features from the processed wood texture data through texture analysis algorithm (such as gray level co-occurrence matrix); extracts knife operation features (such as tracking the inter-frame knife trajectory to obtain the knife direction and knife mark spacing by using the optical flow method (Optical Flow), and calculating the carving depth by combining the OpenPose (Open Source Pose Estimation Library) to capture the hand action amplitude) from the processed production process video through video frame sequence analysis and motion capture technology, which is to extract the key features reflecting the core technology of Huxiang wood carving from the original data. Next, the image generation system structures and binds the extracted pattern structure features, texture parameter features and knife operation features according to the process correlation logic, forms target feature data containing complete process information, and labels the target feature data with corresponding folk scene labels according to the application scene of the wood carving, which is to ensure the process rationality of the features and associate the scene information to meet the subsequent scene adaptation requirements. After that, the image generation system constructs pattern rule sub-library, wood texture sub-library, knife feature sub-library and scene adaptation sub-library based on the target feature data and folk scene labels, which is to structure the feature data according to the process dimension for accurate retrieval and calling. Finally, the image generation system integrates the above four sub-libraries and establishes cross-sub-library retrieval index to realize the rapid correlation query of process features between different sub-libraries, and finally obtains the Huxiang wood carving production process feature library.
[0084] As an example, the step of binding the pattern structure features, the texture parameter features, and the tooling operation features by process association logic to obtain target feature data and labeling corresponding folk scene tags for the target feature data includes: calling process association logic from process rules of the Hunan wood carving production process feature library, the process association logic including wood texture and tooling adaptation rules and tooling and pattern matching rules; according to the process association logic, corresponding associating the pattern structure features, the texture parameter features, and the tooling operation features belonging to the same wood carving object to form a feature association group; performing consistency checking on the feature association group to verify whether the texture parameter features and the tooling operation features conform to the wood texture and tooling adaptation rules and whether the tooling operation features and the pattern structure features conform to the tooling and pattern matching rules; integrating the feature association group that passes the checking into target feature data; extracting folk scene information corresponding to the target feature data from scene records of the Hunan wood carving production process feature library, and labeling folk scene tags for the target feature data according to the folk scene information.
[0085] The process rules refer to a core rule set stored in the Hunan wood carving production process feature library (embedded in the association fields of the pattern rule sub-library, the wood texture sub-library, the tooling feature sub-library, and the scene adaptation sub-library), guiding the orderly development of each link of the Hunan wood carving production, covering adaptation and matching standards of wood, tooling, pattern, scene, and other dimensions, and providing process basis for feature binding. The wood texture and tooling adaptation rules refer to rules that explicitly define the adaptation relationship between the texture parameter features of the wood used for the Hunan wood carving and the corresponding tooling operation features, for example, fine and interlaced wood texture adapts to shallow relief, tooling with slow frequency, and rough texture adapts to deep relief, tooling with large force. The tooling and pattern matching rules refer to rules that standardize the corresponding relationship between the tooling operation features and the pattern structure features of the Hunan wood carving, for example, continuous back pattern adapts to tooling along the grain with short tool marks, and complex dragon and phoenix pattern adapts to a combination of openwork and relief with multi-directional tooling. The feature association group refers to a temporary feature combination formed by corresponding associating the pattern structure features, the texture parameter features, and the tooling operation features belonging to the same wood carving object according to the process association logic, containing complete process feature association information of the wood carving. The folk scene information refers to application scene related information corresponding to the target feature data stored in the scene records of the Hunan wood carving production process feature library, including folk scene types, display spaces, and use purposes.
[0086] Firstly, the image generation system locates the associated fields by sub-library indexing, extracts two types of rules by query statements (such as "SELECT adaptation knife work rule FROM wood texture sub-library WHERE wood type = camphor wood" and "SELECT matching knife work rule FROM pattern rule sub-library WHERE pattern type = back pattern"), integrates them into process association logic, and clarifies the feature binding basis. Secondly, the image generation system inputs the wood carving unique identifier (such as scan number MX-001), filters three types of features marked with the identifier, maps and binds the corresponding fields (such as texture parameter field and knife work operation field) one by one according to the "wood - knife - pattern" dimension, forms a feature association group. Then, the image generation system extracts parameters such as texture density, carving depth, and pattern line complexity in the association group, and compares them one by one with the threshold values (such as fine texture adaptation carving depth ≤ preset value) in the process association logic to verify whether the parameters meet the two types of rules. After that, the image generation system sorts the association groups that pass the verification according to the process dimension, unifies the field format (such as converting to a preset vector format), and integrates them into target feature data. Finally, the image generation system extracts the core keywords (such as "back pattern + middle hall size") of the target feature data, matches them with the scene record keywords of the scene adaptation sub-library using the TF-IDF algorithm, and selects the information with the highest matching degree to assign a standardized label (such as "sacrifice - temple placement") to the target feature data.
[0087] As an example, the step of constructing the pattern rule sub-library, wood texture sub-library, knife work feature sub-library, and scene adaptation sub-library based on the target feature data and the folk scene label includes: structurally disassembling the target feature data to separate independent pattern structure features, texture parameter features, and knife work operation features; extracting pattern shape parameters, combination logic, and implied association information from the pattern structure features, classifying and organizing them according to the Hunan wood carving school and pattern type, and constructing the pattern rule sub-library; extracting texture direction, pore density, annual ring spacing, and material hardness parameters from the texture parameter features, storing them according to the wood species, and constructing the wood texture sub-library; extracting knife direction, carving depth, and knife mark spacing from the knife work operation features, associating them with the corresponding wood types, and constructing the knife work feature sub-library; based on the folk scene label, extracting size parameters, process types, placement space requirements, and light and shadow adaptation parameters from the target feature data, classifying them according to the scene type, and constructing the scene adaptation sub-library.
[0088] Pattern form parameters refer to specific indexes describing the shape characteristics of Hunan wood carving patterns, including line curvature, pattern symmetry, etc. Combination logic and implied association information refer to the arrangement of Hunan wood carving patterns (such as the combination of primary and secondary patterns) and the corresponding relationship between patterns and folk implied meanings. Hunan wood carving schools refer to regional divisions of Hunan wood carving craft schools, including Xiangdong, Xiangxi, and Xiangzhong schools. Pattern types refer to traditional pattern categories in Hunan wood carving with fixed styles and structures, including frets, cloud patterns, dragon and phoenix patterns, etc. Grain direction refers to the extension direction of wood grain (such as longitudinal, staggered), pore density refers to the distribution density of pores inside the wood, annual ring spacing refers to the distance between annual rings, and material hardness parameter refers to the hardness index of the wood. Knife direction refers to the moving direction of the knife during carving (such as with the grain, across the grain), carving depth refers to the depth of the knife into the wood, and knife mark spacing refers to the distance between adjacent knife marks. Size parameters refer to the length, width, and thickness data of wood carvings, process types refer to the manufacturing process of wood carvings (such as relief, openwork, and sculpture), display space requirements refer to the suitable display location of wood carvings (such as the main hall, ancestral hall, and wedding room), and light and shadow adaptation parameters refer to the light and shadow adjustment related data of wood carvings suitable for the scene. Scene type refers to the classification of folk application scenarios of Hunan wood carving, including marriage, sacrifice, and festival, etc.
[0089] First, the image generation system uses a data parsing tool to split the target feature data according to preset fields such as "pattern structure," "texture parameters," and "carving operation." This separates the pattern structure features, texture parameter features, and carving operation features into three independent datasets. This ensures that features from different process dimensions are independent, facilitating targeted processing later. Second, from the separated pattern structure features, a feature extraction tool filters out pattern morphology parameters, combination logic, and symbolic association information. A secondary category directory, "Hunan Woodcarving School - Pattern Type" (e.g., "Eastern Hunan School - meander pattern"), is created in the database. Corresponding information is stored in this directory and indexed, constructing a pattern rule sub-library. This allows pattern features to be stored systematically according to process attributes, improving retrieval efficiency. Third, parameters such as texture direction and porosity are extracted from the texture parameter features. Dedicated storage folders are created in the database for each type of wood, such as camphor wood and nanmu wood. The corresponding wood parameter files are stored in these folders and labeled with attributes, constructing a wood texture sub-library. This directly associates wood features with specific wood types, facilitating subsequent matching of carving techniques. Next, parameters such as knife direction and carving depth are extracted from the knife-working operation features. Using database field association functionality, each knife-working parameter is bound to its corresponding wood type (e.g., camphor wood). These parameters are then categorized and stored in a designated path after being sorted by knife-working technique, constructing a knife-working feature sub-library. This step ensures that the knife-working parameters are compatible with the wood type and conform to the process logic. Finally, based on folk custom scene tags, information such as scene-adaptive size parameters is extracted from the target feature data. Subdirectories are created in the database according to scene types such as weddings and sacrifices. Parameter files for the corresponding scenes are placed in these subdirectories and associated with them, constructing a scene adaptation sub-library. This achieves a precise correspondence between scene features and application scenarios, meeting the needs of subsequent scene-based generation.
[0090] Step S20: Based on the feature library of Hunan wood carving production process, the initial LoRA model is trained in layers according to the hierarchical logic of texture layer-carving layer-pattern layer to obtain a layered LoRA model.
[0091] It should be noted that the initial LoRA model refers to a basic LoRA (Low-Rank Adaptation) model built on a general generative model without specific training on the characteristics of Hunan wood carving techniques. The hierarchical LoRA model refers to a model obtained by training the initial LoRA model in layers according to the hierarchical logic of texture layer, carving technique layer, and pattern layer. The texture layer focuses on learning the texture characteristics of wood, the carving technique layer focuses on learning the carving technique characteristics, and the pattern layer focuses on learning the pattern structure characteristics. Each layer maintains a hierarchical association with the Hunan wood carving techniques, which can accurately capture the features of different craft levels and collaboratively generate Hunan wood carving images that conform to traditional craft specifications.
[0092] As an example, the characteristic library of the Hu Xiang wood carving production process includes a pattern rule sub-library, a wood texture sub-library, and a knife work characteristic sub-library; the step of training the initial LoRA model in layers according to the characteristic library of the Hu Xiang wood carving production process in a hierarchical logic of a texture layer-knife work layer-pattern layer includes: initializing the LoRA model according to a preset learning rate and a preset batch size to obtain an initial LoRA model; fitting and training the texture layer of the initial LoRA model according to the wood texture sub-library to obtain texture layer characteristic weights; based on the texture layer characteristic weights, carrying out linkage training on the knife work layer of the initial LoRA model according to the knife work characteristic sub-library to obtain knife work layer characteristic weights; combining the texture layer characteristic weights and the knife work layer characteristic weights, carrying out collaborative training on the pattern layer of the initial LoRA model according to the pattern rule sub-library to obtain pattern layer characteristic weights; based on a hierarchical fusion loss function, calculating the deviation value of the output features of the texture layer, the knife work layer, and the pattern layer from the standard features in the characteristic library of the Hu Xiang wood carving production process, and assigning the weight proportion of each hierarchical layer according to the deviation value; based on the weight proportion, calculating the matching degree of the output features of the initial LoRA model and the standard features; if the matching degree is less than a preset matching threshold, adjusting the learning rate of the initial LoRA model by a preset adjustment amplitude, and repeating the steps of hierarchical training and weight assignment until the matching degree is greater than or equal to the preset matching threshold, to obtain a hierarchical LoRA model.
[0093] The preset learning rate and the preset batch size refer to fixed training parameters set when initializing the LoRA model, wherein the preset learning rate is 0.0008, which is used to control the step length of model parameter updating; the preset batch size is 16, which is used to specify the number of data samples input into the model at each training iteration. The texture layer feature weight refers to the parameter weight representing the importance of wood texture characteristics obtained after the texture layer of the initial LoRA model is fitted and trained by the wood texture feature sub-library, which is used to strengthen the learning effect of the model on wood texture characteristics. The tool worker layer feature weight refers to the parameter weight reflecting the influence degree of tool worker operation characteristics obtained after the tool worker layer of the initial LoRA model is jointly trained by the tool worker feature sub-library based on the texture layer feature weight, which ensures the adaptability of tool worker characteristics and texture characteristics. The pattern layer feature weight refers to the parameter weight reflecting the importance of pattern structure characteristics obtained after the pattern layer of the initial LoRA model is cooperatively trained by the pattern rule sub-library in combination with the texture layer feature weight and the tool worker layer feature weight, which realizes the collaborative matching of pattern characteristics and the characteristics of the previous two layers. The hierarchical fusion loss function is a composite loss function specially designed for the hierarchical training of the "texture layer-tool worker layer-pattern layer", and the core logic is to first split and calculate the independent loss of each level, then weightedly fuse according to the preset weight proportion, and finally output the global total loss, so as to accurately constrain the fitting degree of the characteristics of each level of the model and the standard characteristics. The weight proportion refers to the contribution proportion of the texture layer, the tool worker layer and the pattern layer in the model respectively, which is determined according to the deviation value of the output characteristics of each level and the standard characteristics, and the preset texture layer is 0.3, the tool worker layer is 0.4, and the pattern layer is 0.3, which is used to balance the influence of each level on the model output. The matching degree refers to the fitting degree of the output characteristics of the initial LoRA model and the standard characteristics in the Hu Xiang wood carving process feature library calculated based on the weight proportion, which is the core index for judging the training effect of the model. The preset matching threshold refers to the critical value for judging whether the model training meets the requirements, which is preset as 0.9, and when the matching degree of the model output characteristics and the standard characteristics reaches or exceeds the value, it indicates that the model training meets the requirements. The preset adjustment amplitude refers to the fixed step length of adjusting the learning rate in the model training, which is preset as 0.0001, which is used to dynamically optimize the learning rate according to the change trend (low decline rate or rising) of the deviation value, so as to ensure the training efficiency and effect.
[0094] Firstly, the image generation system inputs the preset learning rate and the preset batch size into the parameter initialization module of the LoRA model to assign initial values (such as random low-rank matrices for weight matrices) to the parameters of each layer of the model, obtaining an initial LoRA model. Secondly, taking the characteristic data such as the texture direction and pore density in the wood texture sub-library as input, the output of the initial LoRA model texture layer is compared with the standard texture characteristics in the sub-library, and the texture layer parameters (such as the element values of the low-rank matrix) are adjusted through the back propagation algorithm until the output error is lower than the preset value, obtaining the texture layer feature weight. The purpose is to let the model accurately learn the basic physical characteristics of wood first. Then, based on the determined texture layer feature weight, the data associated with the wood type in the tool feature sub-library, such as the tool direction and carving depth, are input into the model. During training, the tool layer parameters are forced to update in linkage with the texture layer weight (such as adjusting the tool layer parameters by a fixed proportion when the texture layer output changes), so that the tool features adapt to the corresponding wood texture, obtaining the tool layer feature weight, and ensuring the process adaptability of tool and texture. Next, integrate the texture layer and tool layer feature weights, input the pattern shape parameters and combination logic data in the pattern rule sub-library, and train the pattern layer to make its parameters respond to the changes of the weights of the previous two layers (such as the carving depth output by the tool layer determines the thickness parameter of the pattern line), so that the pattern features form a logical closed loop with the texture and tool, obtaining the pattern layer feature weight, and realizing the collaborative matching of the three layers of features. Then, call the hierarchical fusion loss function to calculate the difference values (such as mean square error) of the output of the texture layer and the standard texture characteristics, the output of the tool layer and the standard tool characteristics, and the output of the pattern layer and the standard pattern characteristics as the deviation values of each layer, and then allocate weight proportions according to the deviation value size (increase the weight of the layer with large deviation to strengthen the training), so as to balance the learning priority of each layer. Based on the allocated weight proportion, the fit degree (such as cosine similarity) of the output of each layer and the standard characteristics is weighted and summed according to the weight, obtaining the overall matching degree of the initial LoRA model, and judging the training effect of the model. Finally, if the matching degree is less than the preset matching threshold, check the deviation value trend - if the decline rate is slow, increase the learning rate by the preset adjustment amplitude to speed up parameter updating; if the deviation value rises, reduce the learning rate by the amplitude to avoid parameter shock, and then repeat the steps of layered training and weight allocation until the matching degree meets the standard, obtaining the layered LoRA model, and ensuring that the model can accurately generate the characteristics of Hunan woodcarving that meet the process specifications.
[0095] Step S30, the layered LoRA model is fused with the pre-trained latent diffusion model to obtain a Hunan woodcarving image generation model.
[0096] It should be noted that the pre-trained latent diffusion model refers to a latent diffusion model (Latent Diffusion Model) pre-trained on large-scale general image data (such as natural scenes, various artistic images) and having basic image generation capability, which has mastered the core logic of image pixel distribution, feature extraction and reconstruction, but has not integrated the exclusive process features such as texture, knife work and pattern of Hunan wood carving, and is a basic trunk model for subsequent hierarchical fusion.
[0097] It can be understood that first, the system clearly defines the hierarchical correspondence between the texture layer, the knife work layer, and the pattern layer of the LoRA model and the hierarchical correspondence of the pre-trained latent diffusion model U-Net structure (the texture layer corresponds to the early convolution block of U-Net, the knife work layer corresponds to the middle attention layer, and the pattern layer corresponds to the late upsampling block). This is done to accurately embed the process features into the generation framework. Second, the feature weights of each level of the hierarchical LoRA model are injected into the corresponding level module of the pre-trained latent diffusion model. When injecting, the backbone parameters of the latent diffusion model are frozen, and only the low-rank adapter parameters of the LoRA model are enabled for operation to avoid damaging the basic generation capability of the pre-trained model. Finally, the parameter linkage configuration is completed through the model fusion interface (such as making the LoRA adapter and the attention mechanism of the latent diffusion model respond to the input instructions cooperatively) to enable the pre-trained latent diffusion model to obtain the exclusive generation capability of the texture, knife work, and pattern of Hunan wood carving, and ultimately obtain the Hunan wood carving image generation model.
[0098] In step S40, based on the keyword system of the Hunan wood carving production process feature library, the text prompt words input by the user containing wood types and folk scene are analyzed to obtain wood carving types, wood texture parameters, folk pattern requirements, and scene adaptation features.
[0099] It should be noted that the keyword system is a structured keyword set combed in the Hunan wood carving production process feature library, covering wood type, wood carving type, folk scene, pattern type, process parameter and other core dimensions, providing a unified matching standard for analyzing user demand. Wood type refers to the type of wood specified in the user demand for making wood carving, such as camphor wood, nanmu, catalpa wood, cypress wood, etc., which is the core element of determining the texture characteristics of wood carving. Folk scene refers to the application folk scene of wood carving in user demand, such as marriage, sacrifice, festival, etc., which is directly related to the scene adaptation characteristics and implied expression of wood carving. Text prompt word is a natural language description input by the user, which contains the related demand of wood carving production. The core information covers wood type, folk scene, and can also include style, purpose, placement position and other supplementary requirements. Wood carving type refers to the specific category of wood carving parsed from the text prompt word, such as hanging screen, ornament, decorative component, etc. Wood texture parameter refers to the core texture index corresponding to the wood type, such as texture direction, pore density, annual ring spacing, material hardness, etc. Folk pattern demand refers to the pattern style and implied direction that the user expects the wood carving to present, such as Hunan traditional auspicious pattern, school characteristic pattern, etc. Scene adaptation characteristics refer to specific parameters that adapt to the folk scene, such as size, process type, placement space requirement, light and shadow adaptation parameter, etc.
[0100] As an example, the keyword system based on the Hunan wood carving production process feature library analyzes the text prompt word input by the user, which contains wood type and folk scene, to obtain the steps of wood carving type, wood texture parameter, folk pattern demand and scene adaptation characteristics, including: performing word segmentation processing on the text prompt word input by the user, which contains wood type and folk scene, to obtain wood type keyword, folk scene keyword, purpose keyword and implied keyword; associating and matching the keyword system of the Hunan wood carving production process feature library with the wood type keyword, the folk scene keyword, the purpose keyword and the implied keyword respectively to obtain wood type matching item, scene matching item, purpose matching item and implied matching item; determining the corresponding wood texture parameter from the wood texture sub-library according to the wood type matching item; determining the wood carving type from the scene adaptation sub-library according to the purpose matching item and the scene matching item; determining the folk pattern demand from the pattern rule sub-library according to the scene matching item and the implied matching item; determining the corresponding scene adaptation characteristics from the scene adaptation sub-library according to the scene matching item.
[0101] The wood type keyword refers to the wood species related words (such as "cinnamomum camphora" and "phoebe") extracted after the text prompt word input by the user is processed by word segmentation; the folk scene keyword refers to the folk scene related words (such as "marriage scene" and "sacrifice scene") extracted after word segmentation; the use keyword refers to the wood carving use related words (such as "screen" and "ornament") extracted after word segmentation; the implication keyword refers to the wood carving carried implication related words (such as "auspicious implication" and "fortune and luck implication") extracted after word segmentation. The wood type matching item is the structured result (such as "cinnamomum camphora-wood ID: XM001-attribute: soft and fine texture") containing "wood name + feature library corresponding ID + core attribute" obtained by comparing the wood type keyword with the "wood type keyword table" (containing wood name, classification label and attribute identification) in the keyword system, which is directly used to associate the wood texture sub-library extraction parameters. The scene matching item is the standardized result (such as "marriage scene-region: Xiangxi-subdivision: new house decoration") of "scene name + regional attribute + scene subdivision" obtained by matching the folk scene keyword with the "scene type keyword table" (containing scene name, regional label and use scene subdivision) in the feature library, which is used as the core basis for subsequent extraction of scene adaptation features and pattern requirements. The use matching item is the accurate result (such as "screen-function: decoration-use mode: hanging type") of "category name + function attribute + use mode" obtained by matching the use keyword with the "use type keyword table" (containing wood carving category name, use mode and function label) in the feature library, which is used to determine the specific wood carving type from the scene adaptation sub-library. The implication matching item is the clear result (such as "auspicious implication-pattern direction: dragon and phoenix pattern / cloud pattern- connotation: celebration and blessing") of "implication theme + pattern direction + folk connotation" obtained by matching the implication keyword with the "implication type keyword table" (containing implication theme, corresponding pattern label and folk connotation) in the feature library, which is directly used to screen the corresponding folk pattern requirements from the pattern rule sub-library.
[0102] Firstly, the system splits the text prompt input by the user word by word, and then filters the keywords according to the "wood- scene-purpose-meaning" dimension of the preset characteristic library keyword system (eliminates meaningless auxiliary words such as "de" and "yong"), and accurately extracts the wood type keywords, folk scene keywords, purpose keywords and meaning keywords; Secondly, the wood, scene, purpose and meaning four special keyword tables are called out from the characteristic library keyword system, and are compared with the four types of keywords extracted respectively through "accurate matching + fuzzy association" (for example, "fragrant camphor wood" accurately matches the wood table, and "marriage" fuzzy associates the "marriage scene" in the scene table), and outputs the wood type matching items, scene matching items, purpose matching items and meaning matching items containing "keywords + characteristic library unique ID + core attributes", to ensure that the user's demand is uniquely associated with the characteristic library data; Then, the characteristic library ID in the wood type matching item is used as an index to execute a retrieval instruction in the wood texture sub-library, and the texture direction, pore density, annual ring spacing and other wood texture parameters corresponding to the ID are directly called to realize the rapid positioning of wood characteristics; Next, the "category name" (such as "hanging screen") in the purpose matching item and the "scene subdivision" (such as "new house decoration") in the scene matching item are used as double filtering conditions to execute multi-condition queries in the scene adaptation sub-library, and the wood carving types that meet both the purpose and the scene are filtered out to ensure that the wood carving category is highly adapted to the demand; Then, the "scene attribute" (such as "Xiangxi marriage") in the scene matching item and the "pattern direction" (such as "good luck") in the meaning matching item are used as associated fields to perform a joint query in the pattern rule sub-library, and the folk pattern demand that meets the scene folk custom and meaning demand is matched to make the pattern expression fit the core appeal; Finally, the "scene name" (such as "Xiangxi marriage scene") in the scene matching item is used as a retrieval condition to directly extract the corresponding size parameters, space requirements, light and shadow adaptation parameters and other scene adaptation characteristics in the scene adaptation sub-library, completing the whole-process analysis of the user's demand from natural language to structured process parameters.
[0103] Step S50, inputting the wood carving type, the wood texture parameters, the folk pattern demand and the scene adaptation characteristics into the Huxiang wood carving image generation model to obtain a target Huxiang wood carving image.
[0104] It should be noted that the target Huxiang wood carving image refers to a Huxiang wood carving image generated by the model, which meets the requirements of wood texture characteristics, knife operation specifications, folk pattern style and scene adaptation, and accurately corresponds to the wood type and folk scene demand input by the user.
[0105] As an example, the Hunan wood carving image generation model includes a feature encoding module, a diffusion sampling module, and an image decoding module; the step of inputting the wood carving type, the wood texture parameter, the folk pattern requirement, and the scene adaptation feature into the Hunan wood carving image generation model to obtain a target Hunan wood carving image includes: converting the wood carving type, the wood texture parameter, the folk pattern requirement, and the scene adaptation feature into a feature vector recognizable by the Hunan wood carving image generation model; inputting the feature vector into the feature encoding module for encoding processing to obtain an encoded feature; inputting the encoded feature into the diffusion sampling module to obtain sampling feature data; inputting the sampling feature data into the image decoding module to obtain pixel-level image data; and performing format standardization processing on the pixel-level image data to obtain a target Hunan wood carving image.
[0106] The feature encoding module is a component in the Hunan wood carving image generation model that specifically processes input features, composed of a fully connected layer and a Transformer encoder. It is used to convert different types of features (such as the wood carving type "hanging screen" and the texture parameter "pore density 0.7") into a unified numerical format, and through an internal attention mechanism, it binds related features together (such as "marriage scene" and "auspicious pattern" are associated). This allows the model to understand the relationship between these features. The diffusion sampling module is a component that generates core features of the model's image, based on the logic of "gradually restoring an image from noise". It first takes a bunch of random noise data, then combines the wood carving information in the encoded features, and step by step removes the noise - each step uses the model's network to determine "where should the wood grain be" and "where should there be a knife mark". Finally, it obtains intermediate data containing specific wood carving features (such as where the wood texture is and where the carved pattern lines are). The image decoding module is a component that converts abstract features into actual images, composed of a transpose convolution layer and an upsampling layer. It enlarges the small size features output by the diffusion sampling module step by step, and through detail processing, it changes features such as "wood grain direction" and "knife mark depth" into specific pixels (such as dark pixels representing knife marks and light pixels representing wood original color), and finally outputs image detail data that can be directly seen. The feature vector is a string of numbers that the model can calculate by converting various types of input information, such as: the wood carving type "hanging screen" is encoded as [1,0,0] (representing a hanging screen rather than a display); the texture parameter "texture direction 45 degrees" of "camphor wood" is converted to a number like 0.45; "auspicious meaning" corresponds to [0.8,0.2] (representing a bias towards auspicious patterns). These numbers are concatenated in order to form a long string of numbers (such as [1,0,0,0.45,0.8,0.2,...]), which is the feature vector.
[0107] The coding feature is a high-dimensional digital string (such as a string of 512 numbers) obtained after the feature vector is processed by the feature coding module. It is not just a simple concatenation of numbers, but also contains the association between features - for example, the texture parameters of "camphor wood" and the size requirements of "marriage scene" will be bound together, so that the model knows "how big this wood should be in this scene". The sampling feature data is the three-dimensional data (such as 64x64x512 cubic data) containing all the key features of the wood carving output by the diffusion sampling module. Among them, 64x64 represents the approximate outline size of the image, and 512 represents the detailed features of each position (such as the numerical value of a certain point is high, which represents a deep incision mark; the numerical value of a certain area changes regularly, which represents the wood grain direction). The pixel-level image data is the original image data output by the image decoding module, which can be directly displayed as a two-dimensional pixel matrix (such as a 512x512x3 matrix). Among them, "512x512" is the size of the image, and "3" represents the RGB three colors; the numerical value of each position (such as [220, 180, 150]) corresponds to the color of a pixel point, and these points are spliced together to form the wood carving image that can be seen (such as the wood color background and the dark pattern lines).
[0108] Firstly, the wood carving type is converted to a numerical label by one-hot encoding, the wood texture parameters are standardized to floating point values according to the threshold of the feature library, and the folk pattern demand and scene adaptation features are respectively mapped to preset numerical coding. Then, these numerical features are spliced into a one-dimensional vector in a fixed order to obtain a feature vector that can be recognized by the Hunan wood carving image generation model. This is done to convert non-numerical requirements into structured data that the model can calculate. Secondly, the feature vector is input into the feature encoding module. The module maps the vector dimension uniformly to 512 dimensions through a fully connected layer, and then strengthens the association between features (such as the semantic binding of wood carving type and scene adaptation features) through the self-attention mechanism of the Transformer encoder to obtain encoded features, realizing the unified fusion and semantic enhancement of scattered features. Then, the encoded features are input into the diffusion sampling module. The module first initializes a 64x64 random Gaussian noise vector, and then uses the encoded features as a conditional constraint to perform 50-step iterative denoising through the U-Net network (each step predicts noise and subtracts it, while calling the process feature weight adjustment of the hierarchical LoRA model to adjust the denoising direction). The sampling feature data containing wood carving process details is obtained, and the feature distribution conforming to the process specification is constructed in the latent space. Next, the sampling feature data is input into the image decoding module. The module gradually enlarges the feature map size through 3 layers of transpose convolution (from 64x64→128x128→256x256→512x512), and uses PixelShuffle (pixel reorganization) upsampling to extract wood texture, pattern lines and other detailed features, obtaining a two-dimensional pixel matrix in RGB format, i.e. pixel-level image data, completing the conversion of abstract features to concrete pixels. Finally, the pixel-level image data is standardized to compress the pixel value to the range of 0-255, adjust the image resolution to 1024x1024 and convert it to PNG format to obtain the target Hunan wood carving image.
[0109] The embodiment provides a method for generating a Hunan wood carving image based on a LoRA model. First, a feature database of a Hunan wood carving production process is constructed according to a scanning diagram of a real Hunan wood carving, wood texture data and a production process video, to provide core data support for subsequent links. Then, hierarchical training is performed on an initial LoRA model according to the feature database in a hierarchical logic of a texture layer, a tooling layer and a pattern layer, to obtain a hierarchical LoRA model, so that the model accurately masters core features of a Hunan wood carving process. Then, the hierarchical LoRA model is fused with a pre-trained latent diffusion model to obtain a Hunan wood carving image generation model, which takes into account generation efficiency and process professionalism. Then, a keyword system of the feature database is used to analyze a text prompt word input by a user and containing a wood type and a folk scene, to obtain a wood carving type, wood texture parameters, folk pattern requirements and scene adaptation features, and to accurately convert user requirements. Finally, the structured features are input into the Hunan wood carving image generation model to obtain a target Hunan wood carving image. The embodiment can generate a Hunan wood carving image with process restoration and scene adaptability.
[0110] Based on the first embodiment of the present application, the same or similar contents as the above-mentioned first embodiment can be referred to the above description, and the following will not be repeated. On this basis, please refer to Figure 2 , Figure 2 The flowchart of the second embodiment of the method for generating a Hunan wood carving image based on a LoRA model of the present application is shown in the figure. The hierarchical LoRA model includes a texture layer, a tooling layer and a pattern layer. The step S30 of the method for generating a Hunan wood carving image based on a LoRA model includes steps S31-S36.
[0111] In step S31, a pre-trained latent diffusion model is obtained. The latent diffusion model includes a feature encoding module, a diffusion sampling module and an image decoding module. The feature encoding module includes a texture generation submodule. The diffusion sampling module includes a detail rendering submodule. The image decoding module includes a composition generation submodule.
[0112] It should be noted that in the hierarchical LoRA model, the texture layer is a layer that directly acts on the texture generation submodule. Its structure is designed as a low-rank matrix adapter (A / B matrix). By freezing the original model weight, only the A matrix (dimension reduction) and the B matrix (dimension increase) are trained to capture the physical characteristics of the wood. For example, for parameters such as wood grain direction and pore density, the A matrix of the texture layer compresses the input high-dimensional features to a low-rank space (such as rank r=8), and the B matrix maps them back to the original dimension and adds them to the original model output, to realize accurate modeling of the wood texture. This structure maintains the generality of the original model while achieving specialized optimization of wood texture with a small amount of parameters (about 0.1% of the total fine-tuning).
[0113] The tooling layer is a level embedded in the detail rendering submodule, and its core is the LoRA adapter for convolutional layers and attention layers, which learns the "tooling" and "cutting" techniques corresponding to the tool mark features by training A / B matrices. For example, in the residual block of U-Net, the A matrix of the tooling layer projects the input features into a low-dimensional space (such as r=16), capturing tool mark features such as tool mark depth and spacing, and the B matrix injects these features into the diffusion denoising process, dynamically adjusting the sampling path to render tooling details. This design allows the model to optimize the weight distribution of tooling features through hierarchical fusion loss functions without modifying the original diffusion process.
[0114] The pattern layer is a level integrated into the composition generation submodule, and its structure uses an attention mechanism enhanced LoRA adapter to model the morphological combination logic of traditional patterns by training A / B matrices. For example, in the cross-attention layer of the decoder, the A matrix of the pattern layer compresses the input pattern parameters (such as the fret and entwining pattern) into a low-rank space, and the B matrix combines scene adaptation features (such as the placement angle and size ratio) to generate a composition layout that conforms to folk customs and implications through a multi-head attention mechanism. This hierarchical structure decouples global patterns from fine-grained information, ensuring consistent pattern style while improving scene adaptation flexibility.
[0115] The texture generation submodule, as the core component of the latent diffusion model feature encoding module, first receives the low-rank feature output of the texture layer, and then converts the wood parameters into texture feature encodings that can be processed by the model through the cascaded structure of convolutional neural networks (CNN) and LoRA adapters. For example, after the input wood texture parameters (such as annual ring spacing and material hardness) are extracted by CNN to obtain basic features, they are then upgraded by the B matrix of the texture layer, and finally a latent representation containing wood texture is generated, providing an initial condition for the subsequent diffusion process.
[0116] The detail rendering submodule is a key node in the diffusion sampling module that implements tooling layer feature injection, and its structure includes parallel branches of cross-attention layers and LoRA adapters. For example, in the middle block of U-Net, the A matrix of the tooling layer projects the tool mark features into a low-dimensional space, and the cross-attention mechanism fuses these features with conditional information such as text embedding, and then the B matrix injects the fused features into the diffusion denoising step, dynamically adjusting the noise prediction to enhance the realism of tooling marks. This design achieves fine rendering of process details through hierarchical feature fusion without increasing computational complexity.
[0117] The composition generation submodule is the final output layer of the image decoding module integrating the pattern layer and the scene adaptation feature, and the structure adopts a hybrid architecture of a Transformer decoder and a LoRA adapter. For example, the B matrix of the pattern layer inputs the pattern feature and the scene adaptation parameter (such as the spatial scale of the folk scene) into the Transformer decoder to generate a global composition layout through a multi-head attention mechanism, and then the global composition layout is converted into a final image through upsampling and convolution operation. This structure realizes the collaborative optimization of “pattern-scene” by decoupling the pattern style and the scene requirement, and ensures that the generated image not only retains the characteristics of traditional crafts, but also meets the visual requirements of modern application scenes.
[0118] In step S32, a mapping relationship table of the main module and the submodules is established according to the interface parameters of the texture generation submodule, the detail rendering submodule, and the composition generation submodule.
[0119] It should be noted that the interface parameters include the dimension of the input feature, the format (such as tensor / vector) of the output feature, and the interaction protocol (such as TCP / IP) of data transmission, which are the basis for smooth data flow between modules. The mapping relationship table refers to a structured table that records the main module (feature encoding module, diffusion sampling module, image decoding module) and the corresponding submodule (texture generation submodule, detail rendering submodule, composition generation submodule), as well as the interface parameters (input feature dimension, output feature format, data interaction protocol) adapted by the two, which clearly defines the correspondence and docking rules of the main-submodule, ensuring conflict-free interaction between modules.
[0120] It can be understood that first, the main-submodule correspondence relationship of the feature encoding module corresponding to the texture generation submodule, the diffusion sampling module corresponding to the detail rendering submodule, and the image decoding module corresponding to the composition generation submodule is defined; then, the interface parameters of the texture generation submodule, the detail rendering submodule, and the composition generation submodule are extracted respectively; finally, according to the column items of “main module name-submodule name-input feature dimension-output feature format-data interaction protocol”, the above correspondence and interface parameters are sequentially filled into the table to form the mapping relationship table of the main module and the submodule.
[0121] In step S33, according to the mapping relationship table, the texture layer is docked with the texture generation submodule, the knife work layer is docked with the detail rendering submodule, and the pattern layer is docked with the composition generation submodule to obtain an initial fusion model.
[0122] It should be noted that the initial fusion model refers to the model formed after the texture layer is adaptively connected to the texture generation submodule, the tooling layer is connected to the detail rendering submodule, and the pattern layer is connected to the composition generation submodule according to the mapping relationship table, which has the ability of basic process feature injection. It realizes the preliminary integration of the process features of the layered LoRA model and the potential diffusion model generation architecture, can preliminarily combine wood texture, tooling details, and pattern composition features to carry out image generation, but has not yet undergone parameter fine-tuning and optimization, and is the prototype of the Hunan wood carving image generation model.
[0123] It can be understood that first, the image generation system reads the mapping relationship table to obtain the input feature dimension, output feature format, and data interaction protocol of the texture generation submodule, then automatically adjusts the output feature dimension and format of the texture layer according to these parameters, and establishes a two-way data transmission channel between the texture layer and the texture generation submodule through a preset interface adaptation program. This is done to ensure that the wood texture features of the texture layer can be accurately transmitted to the texture generation submodule. Second, the image generation system extracts the interface parameters of the detail rendering submodule from the mapping relationship table, adjusts the feature output parameters of the tooling layer accordingly, and connects the tooling layer to the detail rendering submodule through a module binding instruction to ensure that the tool mark process features of the tooling layer are transmitted according to the protocol. This is done to ensure that the tooling details can be accurately processed by the detail rendering submodule. Finally, the image generation system adjusts the output features of the pattern layer to match the interaction requirements by referring to the interface configuration of the composition generation submodule in the mapping relationship table, and completes the connection between the pattern layer and the composition generation submodule through an interface association program. At this time, the texture layer, the tooling layer, and the pattern layer are connected to the corresponding submodules, and the image generation system integrates these connection relationships to obtain the initial fusion model. This is done to realize the preliminary integration of the layered process features and the generation model submodules.
[0124] In step S34, the initial fusion weight is determined according to the preset Hunan wood carving process priority, and a hierarchical fusion unit is constructed according to the initial fusion weight.
[0125] It should be noted that the preset Hunan wood carving process priority refers to the importance ranking of different core features in the preset Hunan wood carving process, which is tooling features > pattern features > texture features in this embodiment. The initial fusion weight is a quantitative weight value allocated to each layer connection according to the preset Hunan wood carving process priority. In this embodiment, the texture layer is 0.25, the tooling layer is 0.45, and the pattern layer is 0.3, which is used to measure the influence of the features of the texture layer, the tooling layer, and the pattern layer on the final generation result in model fusion. The hierarchical fusion unit is a model component realized by a fully connected layer + weighted summation logic, and its core function is to perform weighted calculation and merging on the three types of connected features (texture, tooling, and pattern) according to the initial fusion weight, and to preferentially strengthen the influence of high-priority process features.
[0126] It can be understood that a full connection layer including an input layer, a calculation layer and an output layer is first built as a basic framework; then a weighted summation logic is embedded in the calculation layer, three ports of the input layer are respectively designated as the texture features after the texture layer and the texture generation submodule are connected, the tooling features after the tooling layer and the detail rendering submodule are connected, and the pattern features after the pattern layer and the composition generation submodule are connected, and initial fusion weights are respectively bound to the three input ports, so that each feature input is automatically multiplied by the corresponding weight; finally, the output layer is set to add the three weighted features to obtain the integrated unified features and output, thereby completing the construction of the hierarchical fusion unit.
[0127] In step S35, a reference fusion model is obtained according to the initial fusion model and the hierarchical fusion unit.
[0128] As an example, the step of obtaining the reference fusion model according to the initial fusion model and the hierarchical fusion unit includes: connecting the feature input ports of the hierarchical fusion unit with the feature output ports of the texture layer-texture generation submodule, the tooling layer-detail rendering submodule and the pattern layer-composition generation submodule in the initial fusion model one by one; and connecting the weighted feature output port of the hierarchical fusion unit with the feature integration input end of the initial fusion model to obtain the reference fusion model.
[0129] The feature input port is a special interface provided on the hierarchical fusion unit, used to receive the texture features, the tooling features and the pattern features output by the texture layer-texture generation submodule, the tooling layer-detail rendering submodule and the pattern layer-composition generation submodule in the initial fusion model respectively, and is a channel for the three types of process features to enter the hierarchical fusion unit. The weighted feature output port is an interface on the hierarchical fusion unit for outputting the processed features, specifically used to transmit the unified features after the hierarchical fusion unit is weighted and integrated according to the initial fusion weights to the corresponding interface of the initial fusion model, and is an output channel for the integrated features. The feature integration input end is a preset interface on the initial fusion model, specially used to receive the weighted integrated features transmitted by the weighted feature output port of the hierarchical fusion unit, and provides the integrated process feature data for the subsequent image generation process of the initial fusion model. The reference fusion model is a model with process feature weighted integration capability formed by one-to-one connection of the feature input ports of the hierarchical fusion unit with the three types of feature output ports of the initial fusion model, and connection of the weighted feature output port of the hierarchical fusion unit with the feature integration input end of the initial fusion model, and is a basic model for further optimization.
[0130] In step S36, a sample data set is selected from the lake Xiang wood carving production process feature library, and the reference fusion model is jointly fine-tuned according to the sample data set to obtain a lake Xiang wood carving image generation model.
[0131] As an example, the step of selecting a sample data set from the lake Xiang wood carving production process feature library and jointly fine-tuning the reference fusion model according to the sample data set to obtain a lake Xiang wood carving image generation model includes: selecting a sample data set from the lake Xiang wood carving production process feature library, the lake Xiang wood carving production process feature library including a pattern rule sub-library, a wood texture sub-library, and a knife work feature sub-library, the sample data set including texture data of the wood texture sub-library, operation data of the knife work feature sub-library, structure data of the pattern rule sub-library, and a corresponding wood carving standard image; converting image data in the sample data set into a feature matrix of a preset size, and converting process data into a vector to obtain a fine-tuning data set; inputting the fine-tuning data set into the reference fusion model and calculating cross-layer cooperative losses of the texture layer and the texture generation submodule, the knife work layer and the detail rendering submodule, and the pattern layer and the composition generation submodule; based on the cross-layer cooperative losses, adjusting fusion weights of each interfacing layer through the hierarchical fusion unit, and using a verification sample to calculate a process feature matching degree between a model generated image and the wood carving standard image; and when the process feature matching degree is greater than a preset matching degree threshold, a lake Xiang wood carving image generation model is obtained.
[0132] The sample dataset is a collection of data selected from the characteristic library of Hu Xiang wood carving production process for model fine-tuning. The texture data is the relevant data recorded in the wood texture sub-library that reflects the physical texture of the wood used in Hu Xiang wood carving, including wood grain direction, pore density, annual ring spacing, material glossiness, and other information that can reflect the wood texture characteristics. The operation data is the relevant data recorded in the knife worker characteristic sub-library that reflects the carving process of Hu Xiang wood carving, including the way of using knives such as punching and cutting, and the characteristics of knife marks such as depth, spacing, direction, and carving intensity. The structure data is the relevant data recorded in the pattern rule sub-library that reflects the characteristics of Hu Xiang wood carving patterns, including the shape, line layout, element combination logic, and meaning association method of traditional auspicious patterns and school characteristic patterns. The wood carving standard image is a clear image that corresponds to the process data (texture, operation, structure data) in the sample dataset and meets the traditional process specifications of Hu Xiang wood carving, serving as a benchmark for comparing the results of model generation. The image data is the visual data in the sample dataset, specifically the original image information corresponding to the wood carving standard image, which is the basic data for the model to learn visual features. The preset size is a unified image size standard set before model training, used to normalize different specifications of image data and ensure consistency in data format for input models. The feature matrix is a structured data form formed by feature extraction algorithms after adjusting the image data to the preset size, carrying the visual features of the image (such as color, contour, texture details) in matrix form, facilitating model calculation and processing. The process data is the non-visual data in the sample dataset that is directly related to the process of Hu Xiang wood carving, including texture data, operation data, and structure data, which are the core data for the model to learn process features. The fine-tuning dataset is a standardized data set formed after the sample dataset is converted, i.e., image data is converted into feature matrices of the preset size and process data is converted into feature vectors, which is specifically used for fine-tuning the reference fusion model. The cross-layer collaborative loss is a loss indicator that measures the matching degree of feature transmission and collaboration between the texture layer and the texture generation sub-module, the knife worker layer and the detail rendering sub-module, and the pattern layer and the composition generation sub-module in the reference fusion model. The higher the value, the greater the cross-layer coordination deviation. The fusion weight of each interface layer is the weighted proportion of the three groups of interface layers, i.e., the texture layer-texture generation sub-module, the knife worker layer-detail rendering sub-module, and the pattern layer-composition generation sub-module, which is used to adjust the influence of each group of process features on the generation result. The verification sample is a part of data from the sample dataset that is not involved in model training, including corresponding feature matrices, feature vectors, and wood carving standard images, which is used to test the generation effect of the fine-tuned model. The process feature matching degree is a quantitative indicator that measures the degree of agreement between the model-generated image and the wood carving standard image in core process features (knife details, pattern structure, wood texture), represented by a value between 0 and 1, with a higher value indicating better agreement.The preset matching degree threshold is a critical value for judging whether the model is qualified (0.9 in this embodiment). When the matching degree of the process features of the generated image and the wood carving standard image exceeds this value, it indicates that the model meets the process restoration requirements.
[0133] First, the image generation system selects the texture data of the wood grain texture sub-library, the operation data of the knife work feature sub-library, the structure data of the pattern rule sub-library, and the corresponding wood carving standard image from the lake Xiang wood carving process feature library to form a sample data set. Second, the image generation system crops and scales the wood carving standard image in the sample data set according to the preset size, and converts it into a feature matrix through a convolution feature extraction algorithm. Meanwhile, the texture, operation, and structure process data are converted into feature vectors through encoding mapping to form a fine-tuning data set. This is done to unify the data format for efficient model processing. Third, the image generation system inputs the fine-tuning data set into the reference fusion model and uses the mean square error (MSE) loss function to calculate the differences between the output features of the texture layer-texture generation sub-module, the knife work layer-detail rendering sub-module, and the pattern layer-composition generation sub-module and the sample standard features. The sum of the three types of difference values is the cross-layer collaborative loss. This is done to quantify the matching deviation of cross-layer feature transmission. Fourth, the image generation system adjusts the fusion weights of each interface layer based on the cross-layer collaborative loss. If the loss value of a certain interface layer is higher than the preset threshold 0.35, its corresponding weight is reduced by 10-20%. The remaining layers are fine-tuned according to the loss trend, while ensuring that the sum of the weights of the three is 1. This is done to optimize the cross-layer collaborative effect and reduce the deviation. Fifth, the image generation system selects a verification sample from the sample data set that is not involved in training and inputs it into the model. The cosine similarity algorithm is used to calculate the feature vector similarity of the generated image and the wood carving standard image in terms of knife work, pattern, and texture features. The average value of the three is taken as the process feature matching degree. This is done to objectively measure the degree of agreement between the model generation effect and the standard process. Finally, when the process feature matching degree is greater than the preset matching degree threshold 0.9, the image generation system stops fine-tuning and determines the current model as the lake Xiang wood carving image generation model. This is done to ensure that the model meets the restoration requirements of the lake Xiang wood carving process.
[0134] The embodiment first acquires a latent diffusion model containing a texture generation submodule, a detail rendering submodule and a composition generation submodule, relies on the mature architecture of the latent diffusion model to reuse the base image generation capability, reduces the model construction cost and ensures the generation stability; secondly, the interface parameters of each submodule are extracted, the association relationship between the main module and the submodules is combed and a mapping relationship table is established, the connection rules are clarified to avoid data conflicts and improve the subsequent fusion efficiency; then, the texture layer, the tooling layer and the pattern layer are respectively interfaced with the texture generation submodule, the detail rendering submodule and the composition generation submodule according to the mapping relationship table, an initial fusion model is obtained, the precise binding of the process features and the generation key links is realized, and the original model architecture is kept complete; then, the initial fusion weights are distributed according to the preset priority, a hierarchical fusion unit containing weighted summation logic is constructed, the importance of the core process is highlighted, and the multi-dimensional features are sequentially fused; then, the hierarchical fusion unit is connected with the initial fusion model, a complete process feature processing link is built, and a reference fusion model is obtained; finally, sample data sets containing texture data, operation data, structure data and standard wood carving image are selected from the Hunan wood carving process feature database, are standardized and input into the reference fusion model, the fusion weights are adjusted by calculating the cross-layer collaborative loss, the process feature matching degree is calculated by using the verification sample until the standard is reached, the Hunan wood carving image generation model is obtained, the process restoration degree is accurately improved, and it is ensured that the generated image meets the process specification.
[0135] It should be noted that the above examples are only used for understanding the present application and do not constitute a limitation on the LoRA model-based Hunan wood carving image generation method of the present application. More forms of simple transformation based on this technical concept are within the protection scope of the present application.
[0136] The present application also provides a LoRA model-based Hunan wood carving image generation device, please refer to Figure 3 , the LoRA model-based Hunan wood carving image generation device comprises:
[0137] The feature database construction module 10 is configured to construct a Hunan wood carving process feature database according to a Hunan wood carving real object scan image, wood texture data and a manufacturing procedure video;
[0138] The hierarchical training module 20 is configured to perform hierarchical training on an initial LoRA model according to the Hunan wood carving process feature database according to the hierarchical logic of the texture layer-tooling layer-pattern layer, and obtain a hierarchical LoRA model;
[0139] The model fusion module 30 is configured to perform hierarchical fusion on the hierarchical LoRA model and a pre-trained latent diffusion model, and obtain a Hunan wood carving image generation model;
[0140] The prompt word analysis module 40 is configured to analyze the text prompt word input by the user based on the keyword system of the characteristic library of the Hu Xiang wood carving process, and obtain the wood carving type, wood texture parameters, folk pattern demand and scene adaptation characteristics.
[0141] The image generation module 50 is configured to input the wood carving type, the wood texture parameters, the folk pattern demand and the scene adaptation characteristics into the Hu Xiang wood carving image generation model to obtain a target Hu Xiang wood carving image.
[0142] The device for generating a Hu Xiang wood carving image based on a LoRA model provided in the present application adopts the method for generating a Hu Xiang wood carving image based on a LoRA model in the above embodiment, and can solve the technical problem of how to generate a Hu Xiang wood carving image with process restoration and scene adaptation. Compared with the prior art, the device for generating a Hu Xiang wood carving image based on a LoRA model provided in the present application has the same beneficial effects as the method for generating a Hu Xiang wood carving image based on a LoRA model provided in the above embodiment, and other technical features in the device for generating a Hu Xiang wood carving image based on a LoRA model are the same as the features disclosed in the above embodiment, which will not be described here.
[0143] The present application provides a device for generating a Hu Xiang wood carving image based on a LoRA model, which comprises at least one processor and a memory in communication connection with the at least one processor. The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method for generating a Hu Xiang wood carving image based on a LoRA model in the above embodiment one.
[0144] Reference will now be made to the following description Figure 4 which shows a structural schematic diagram of a device for generating a Hu Xiang wood carving image based on a LoRA model suitable for implementing the embodiments of the present application. The device for generating a Hu Xiang wood carving image based on a LoRA model in the embodiments of the present application can include but is not limited to mobile terminals such as mobile phones, notebook computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals) and the like, as well as fixed terminals such as digital TVs, desktop computers and the like. Figure 4 The device for generating a Hu Xiang wood carving image based on a LoRA model shown is only an example, and should not impose any limitation on the functions and use range of the embodiments of the present application.
[0145] As shown in Figure 4 The LoRA model-based Hunan wood carving image generation device can include a processing device 1001 (e.g., a central processing unit, a graphics processing unit, etc.) that can perform various appropriate actions and processes according to a program stored in a ROM (Read Only Memory) 1002 or a program loaded from a storage device 1003 into a RAM (Random Access Memory) 1004. In the RAM 1004, various programs and data required for the operation of the LoRA model-based Hunan wood carving image generation device are also stored. The processing device 1001, the ROM 1002, and the RAM 1004 are connected to each other through a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems can be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touch pad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, an LCD (Liquid Crystal Display), a speaker, a vibrator, etc.; the storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 can allow the LoRA model-based Hunan wood carving image generation device to communicate wirelessly or by wire with other devices to exchange data. Although the LoRA model-based Hunan wood carving image generation device with various systems is shown in the figure, it should be understood that all the systems shown are not required to be implemented or possessed. More or fewer systems can be alternatively implemented or possessed.
[0146] In particular, the processes described above with reference to the flowcharts can be implemented as a computer software program according to embodiments of the present disclosure. For example, embodiments of the present disclosure include a computer program product comprising a computer program carried on a computer readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network through the communication device, or installed from the storage device 1003, or installed from the ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the methods of embodiments of the present disclosure are performed.
[0147] The lake Xiang wood carving image generation device based on the LoRA model provided by the application adopts the lake Xiang wood carving image generation method based on the LoRA model in the above embodiment, and can solve the technical problem of how to generate a lake Xiang wood carving image with process restoration and scene adaptability. Compared with the prior art, the lake Xiang wood carving image generation device based on the LoRA model provided by the application has the same beneficial effects as the lake Xiang wood carving image generation method based on the LoRA model provided by the above embodiment, and other technical features in the lake Xiang wood carving image generation device based on the LoRA model are the same as the features disclosed in the previous embodiment method, which will not be repeated here.
[0148] It should be understood that parts of the present application can be realized by hardware, software, firmware or a combination thereof. In the description of the above embodiments, specific features, structures, materials or characteristics can be combined in any one or more embodiments or examples in a suitable manner.
[0149] The above is merely a specific implementation of the present application, but the protection scope of the present application is not limited thereto, and any person skilled in the art can easily think of changes or replacements within the technical scope disclosed by the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
[0150] The present application provides a computer readable storage medium having computer readable program instructions (i.e. computer programs) stored thereon, the computer readable program instructions being used to execute the lake Xiang wood carving image generation method based on the LoRA model in the above embodiment.
[0151] The computer readable storage medium provided in the application may be, for example, a U disk, but is not limited to an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, system or device, or any combination of the above. More specific examples of the computer readable storage medium may include, but are not limited to, an electrical connection with one or more conductive wires, a portable computer disk, a hard disk, a RAM (Random Access Memory), a ROM (Read Only Memory), an EPROM (Erasable Programmable Read Only Memory or flash memory), an optical fiber, a CD-ROM (CD-Read Only Memory), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present embodiment, the computer readable storage medium may be any tangible medium containing or storing a program that can be used or combined with an instruction execution system, system or device. The program code contained on the computer readable storage medium can be transmitted by any suitable medium, including but not limited to electrical wires, optical cables, RF (Radio Frequency), etc., or any suitable combination of the above.
[0152] The above computer readable storage medium may be contained in the LoRA model-based Huxiang wood carving image generation device; or may exist separately without being assembled into the LoRA model-based Huxiang wood carving image generation device.
[0153] The above computer readable storage medium carries one or more programs, which, when executed by the LoRA model-based Huxiang wood carving image generation device, cause the LoRA model-based Huxiang wood carving image generation device to: construct a Huxiang wood carving production process feature library according to a real object scanning graph of a Huxiang wood carving, wood texture data and a production process video; perform hierarchical training on an initial LoRA model according to the Huxiang wood carving production process feature library, in a hierarchical logic of a texture layer, a tooling layer and a pattern layer, to obtain a hierarchical LoRA model; perform hierarchical fusion of the hierarchical LoRA model and a pre-trained latent diffusion model to obtain a Huxiang wood carving image generation model; analyze a text prompt word input by a user and containing a wood type and a folk scene to obtain a wood carving type, wood texture parameters, folk pattern requirements and scene adaptation features, based on a keyword system of the Huxiang wood carving production process feature library; and input the wood carving type, the wood texture parameters, the folk pattern requirements and the scene adaptation features into the Huxiang wood carving image generation model to obtain a target Huxiang wood carving image.
[0154] Computer program code for carrying out operations of the present application can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).
[0155] The computer program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other devices to cause a series of operational steps to be performed on the computer, other programmable apparatus or other devices to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0156] The modules involved in the embodiments of the present application can be implemented in the form of software or in the form of hardware. In some cases, the name of the module does not constitute a limitation on the module itself.
[0157] The readable storage medium provided by the present application is a computer readable storage medium, which stores computer readable program instructions (i.e., computer programs) for executing the above-mentioned LoRA model-based Huxiang wood carving image generation method, and can solve the technical problem of how to generate Huxiang wood carving images with process restoration and scene adaptability. Compared with the prior art, the computer readable storage medium provided by the present application has the same beneficial effects as the LoRA model-based Huxiang wood carving image generation method provided by the above-mentioned embodiments, and will not be described here.
[0158] The application also provides a computer program product comprising a computer program which, when executed by a processor, implements the steps of the method for generating a Lake Xiang Woodcarving image based on a LoRA model as described above.
[0159] The computer program product provided by the application can solve the technical problem of how to generate a Lake Xiang Woodcarving image with process restoration and scene adaptability. Compared with the prior art, the beneficial effects of the computer program product provided by the application are the same as those of the method for generating a Lake Xiang Woodcarving image based on a LoRA model provided by the above-mentioned embodiments, and are not described here.
[0160] The above only describes some embodiments of the application, and does not limit the patent scope of the application. Any equivalent structural transformation, direct / indirect application in other related technical fields, or direct / indirect application in other related technical fields based on the technical concept of the application and the content of the specification and drawings are included in the patent protection scope of the application.
Claims
1. A method for generating images of Hunan wood carvings based on the LoRA model, characterized in that, The method includes: A feature library of Hunan wood carving production process was constructed based on scanned images of actual Hunan wood carvings, wood texture data, and videos of the production process. Based on the feature library of Hunan wood carving production process, the initial LoRA model was trained in layers according to the hierarchical logic of texture layer - carving layer - pattern layer to obtain a layered LoRA model. The hierarchical LoRA model is fused with the pre-trained latent diffusion model to obtain the Hunan wood carving image generation model. Based on the keyword system of the Hunan wood carving production process feature library, the text prompts input by the user, which include wood type and folk scene, are parsed to obtain wood carving type, wood texture parameters, folk pattern requirements and scene adaptation features. The wood carving type, the wood texture parameters, the folk pattern requirements, and the scene adaptation features are input into the Hunan wood carving image generation model to obtain the target Hunan wood carving image. The layered LoRA model includes a texture layer, a knife-cut layer, and a pattern layer; The step of fusing the hierarchical LoRA model with the pre-trained latent diffusion model to obtain the Hunan wood carving image generation model includes: A pre-trained latent diffusion model is obtained, which includes a feature encoding module, a diffusion sampling module, and an image decoding module. The feature encoding module includes a texture generation submodule, the diffusion sampling module includes a detail rendering submodule, and the image decoding module includes a composition generation submodule. Based on the interface parameters of the texture generation submodule, the detail rendering submodule, and the composition generation submodule, establish a mapping relationship table between the main module and the submodules; According to the mapping table, the texture layer is connected to the texture generation submodule, the knife-cutting layer is connected to the detail rendering submodule, and the pattern layer is connected to the composition generation submodule to obtain the initial fusion model; The initial fusion weight is determined based on the preset priority of Hunan wood carving techniques, and a hierarchical fusion unit is constructed based on the initial fusion weight; A reference fusion model is obtained based on the initial fusion model and the hierarchical fusion unit; A sample dataset is selected from the feature library of Hunan wood carving production process, and the reference fusion model is jointly fine-tuned based on the sample dataset to obtain the Hunan wood carving image generation model.
2. The method as described in claim 1, characterized in that, The steps of selecting a sample dataset from the feature library of Hunan wood carving production process and jointly fine-tuning the reference fusion model based on the sample dataset to obtain the Hunan wood carving image generation model include: A sample dataset is selected from the Hunan wood carving production process feature library, which includes a pattern rule sub-library, a wood texture sub-library, and a carving feature sub-library. The sample dataset includes the texture data of the wood texture sub-library, the operation data of the carving feature sub-library, the structural data of the pattern rule sub-library, and the corresponding standard wood carving image. The image data in the sample dataset is converted into a feature matrix of a preset size, and the process data is converted into a vector to obtain a fine-tuned dataset. The fine-tuned dataset is input into the reference fusion model, and the cross-layer collaborative loss between the texture layer and the texture generation submodule, the knife-cutting layer and the detail rendering submodule, and the pattern layer and the composition generation submodule is calculated. Based on the cross-layer collaborative loss, the fusion weights of each docking layer are adjusted through the hierarchical fusion unit, and the matching degree of the process features between the generated image and the standard wood carving image is calculated using the verification sample calculation model. When the matching degree of the process features is greater than the preset matching degree threshold, the Hunan wood carving image generation model is obtained.
3. The method as described in claim 1, characterized in that, The Hunan wood carving production process feature library includes a pattern rule sub-library, a wood texture sub-library, and a knife work feature sub-library. The steps of training the initial LoRA model hierarchically according to the feature library of Hunan wood carving production process, following the hierarchical logic of texture layer - carving layer - pattern layer, to obtain the hierarchical LoRA model include: The LoRA model is initialized according to the preset learning rate and preset batch size to obtain the initial LoRA model; Based on the wood texture sub-library, the texture layer of the initial LoRA model is trained by feature fitting to obtain the texture layer feature weights; Based on the texture layer feature weights, the knife-cutting layer of the initial LoRA model is trained in conjunction with the knife-cutting feature sub-library to obtain the knife-cutting layer feature weights. By combining the feature weights of the texture layer and the feature weights of the knife-cutting layer, and co-training the pattern layer of the initial LoRA model according to the pattern rule sub-library, the feature weights of the pattern layer are obtained. Based on the hierarchical fusion loss function, the deviation values between the output features of the texture layer, the carving layer, and the pattern layer and the standard features in the feature library of Hunan wood carving production process are calculated, and the weight ratio of each layer is allocated according to the deviation values. The matching degree between the output features of the initial LoRA model and the standard features is calculated based on the weight ratio; If the matching degree is less than the preset matching threshold, the learning rate of the initial LoRA model is adjusted by a preset adjustment range, and the steps of hierarchical training and weight allocation are repeated until the matching degree is greater than or equal to the preset matching threshold, thus obtaining a hierarchical LoRA model.
4. The method as described in claim 1, characterized in that, The steps for constructing a feature library of Hunan wood carving production process based on scanned images of actual Hunan wood carvings, wood texture data, and videos of the production process include: Collect scanned images of Hunan wood carvings, data on wood texture, and videos of the production process, and perform noise reduction and standardization processing. Extract pattern structure features from the processed physical scan image, extract texture parameter features from the processed wood texture data, and extract knife operation features from the processed production process video; The pattern structure features, texture parameter features, and knife operation features are bound together according to the process association logic to obtain target feature data, and the target feature data is labeled with corresponding folk scene tags. Based on the target feature data and the folk scene tags, a pattern rule sub-library, a wood texture sub-library, a knife work feature sub-library, and a scene adaptation sub-library are constructed. By integrating the wood texture sub-library, the knife work feature sub-library, the pattern rule sub-library, and the scene adaptation sub-library, and establishing a cross-sub-library retrieval index, a feature library of Hunan wood carving production process is obtained.
5. The method as described in claim 4, characterized in that, The steps of parsing user-inputted text prompts containing wood type and folk scene based on the keyword system of the Hunan wood carving production process feature library to obtain wood carving type, wood texture parameters, folk pattern requirements, and scene adaptation features include: The text prompts input by the user, which contain information about wood type and folk scene, are segmented to obtain keywords for wood type, folk scene, use, and symbolism. The keyword system of the Hunan wood carving production process feature library is matched with the keywords of wood type, folk scene, use and meaning respectively to obtain wood type matching items, scene matching items, use matching items and meaning matching items; Based on the wood type matching item, the corresponding wood texture parameters are determined from the wood texture sub-library; Based on the usage matching item and the scene matching item, determine the wood carving type from the scene adaptation sub-library; Based on the scene matching items and the meaning matching items, the folk pattern requirements are determined from the pattern rule sub-library; Based on the scene matching item, the corresponding scene adaptation feature is determined from the scene adaptation sub-library.
6. The method according to any one of claims 1 to 5, characterized in that, The Hunan wood carving image generation model includes a feature encoding module, a diffusion sampling module, and an image decoding module; The step of inputting the wood carving type, the wood texture parameters, the folk pattern requirements, and the scene adaptation features into the Hunan wood carving image generation model to obtain the target Hunan wood carving image includes: The wood carving type, the wood texture parameters, the folk pattern requirements, and the scene adaptation features are converted into feature vectors that can be recognized by the Hunan wood carving image generation model. The feature vector is input into the feature encoding module for encoding processing to obtain the encoded features; The encoded features are input into the diffusion sampling module to obtain sampling feature data; The sampled feature data is input into the image decoding module to obtain pixel-level image data; The pixel-level image data is processed to standardize the format to obtain the target Hunan wood carving image.
7. A device for generating images of Hunan wood carvings based on the LoRA model, characterized in that, The device includes: The feature library construction module is used to build a feature library of Hunan wood carving production process based on the scanned images of actual Hunan wood carvings, wood texture data, and videos of the production process. The layered training module is used to train the initial LoRA model in layers according to the feature library of Hunan wood carving production process, following the hierarchical logic of texture layer-carving layer-pattern layer, to obtain the layered LoRA model. The model fusion module is used to perform layered fusion of the layered LoRA model and the pre-trained latent diffusion model to obtain a Hunan wood carving image generation model. The layered LoRA model includes a texture layer, a carving layer, and a pattern layer. The step of performing layered fusion of the layered LoRA model and the pre-trained latent diffusion model to obtain the Hunan wood carving image generation model includes: obtaining the pre-trained latent diffusion model, which includes a feature encoding module, a diffusion sampling module, and an image decoding module. The feature encoding module includes a texture generation submodule, the diffusion sampling module includes a detail rendering submodule, and the image decoding module includes a composition generation submodule. Based on the texture generation submodule and the detail rendering submodule... The interface parameters of the block and the composition generation submodule are used to establish a mapping table between the main module and the submodules. According to the mapping table, the texture layer is connected to the texture generation submodule, the carving layer is connected to the detail rendering submodule, and the pattern layer is connected to the composition generation submodule to obtain an initial fusion model. The initial fusion weight is determined according to the preset priority of Hunan wood carving process, and a hierarchical fusion unit is constructed according to the initial fusion weight. A reference fusion model is obtained according to the initial fusion model and the hierarchical fusion unit. A sample dataset is selected from the feature library of Hunan wood carving production process, and the reference fusion model is jointly fine-tuned according to the sample dataset to obtain the Hunan wood carving image generation model. The prompt word parsing module is used to parse the text prompt words containing wood type and folk scene based on the keyword system of the Hunan wood carving production process feature library, and obtain the wood carving type, wood texture parameters, folk pattern requirements and scene adaptation features. The image generation module is used to input the wood carving type, the wood texture parameters, the folk pattern requirements, and the scene adaptation features into the Hunan wood carving image generation model to obtain the target Hunan wood carving image.
8. A device for generating images of Hunan wood carvings based on the LoRA model, characterized in that, The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the LoRA-based image generation method for Hunan wood carving as described in any one of claims 1 to 6.
9. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the Hunan wood carving image generation method based on the LoRA model as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Modal language model image editing technology fusing non-perpetual culture elements
CN120125946A
Ru porcelain image generation method fusing Ru porcelain knowledge graph and fine tuning control
CN120451310A