An artware auxiliary manufacturing method and system based on a multi-modal process knowledge base
Patent Information
- Application Number
- CN202611040794.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-14
- Publication Date
- 2026-09-29
AI Technical Summary
这一过程中,设计图样到实际制作之间常需要多次人工修改和实体试制,每一轮修改均涉及重新绘制、重新制作骨架和重新彩绘,周期较长且材料损耗较大
本发明提供的一种基于多模态工艺知识库的工艺品辅助制作方法及系统,将工艺品制作中的造型结构、配色规则、纹样布置、材料工艺和制作约束转化为可检索和可迭代优化的工艺数据,从而提高从需求输入到制作图样输出的效率。传统制作方案依赖人工经验进行反复绘制和试制,难以在早期快速判断结构比例、纹样密度和配色方案是否适合制作。本发明通过工艺特征建模使工艺品的不同部分(例如龙头、龙身、灯节)、纹样、材料和工序信息能够被结构化调用,通过结构控制条件生成使候选图样能够遵循参考轮廓、部件比例和构图布局,降低生成图像中常见的结构混乱问题,通过工艺一致性评价和反馈优化机制根据可制造性评分调整提示词、约束提示、检索权重和结构控制强度,减少无效方案,提高工艺品打样前的方案筛选效率。
Smart Images

Figure CN122840892A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of computer-aided design and artificial intelligence-generated content technology, and particularly relates to a method and system for assisting in the production of handicrafts based on a multimodal process knowledge base. Background Technology
[0002] As generative artificial intelligence technology is increasingly applied in image design and product appearance generation, its potential in assisting with craft production is gradually emerging. Existing generative tools are typically trained on massive amounts of general image data, capable of generating visually appealing craft renderings based on text prompts, providing designers with preliminary aesthetic references. However, these general generative models primarily focus on the aesthetic quality and visual effects of images during design, with their training objectives and evaluation metrics revolving around visual realism, lacking modeling and constraints on the inherent laws governing the specific craft production process. Therefore, when such tools are directly applied to assist in the design of traditional handicrafts with complex structures and rigorous production processes, such as the "Chengnan Dragon Lantern," the generated images, while visually appealing, often suffer from several problems that are disconnected from actual production. Specifically, these problems manifest as follows: unreasonable component proportions, such as the dragon head being too large or too small relative to the dragon body, leading to visual imbalance and making it impossible to produce according to the drawing; unusable lantern section arrangements, such as the spacing, number, or arrangement of lantern sections not meeting the actual requirements for skeletal assembly; excessively dense pattern distribution, exceeding the limits of fineness in hand-painted or carved designs, making it difficult to reproduce on a physical surface; difficulty in implementing color schemes, as the selected colors cannot achieve the expected effect or have insufficient colorfastness on specific materials (such as paper, bamboo strips, and fabric); and inconvenient skeletal structure production, as the generated shapes lack consideration for internal support structures, resulting in beautiful shapes but an inability to construct a solid skeletal framework.
[0003] The production of "Chengnan Dragon Lanterns" involves several closely related stages, including dragon head design, body segmentation, skeletal structure, color matching, pattern arrangement, and assembly sequence. Each stage is subject to strict technological constraints and sequential dependencies. Traditional production relies primarily on the experience of artisans, who must use their accumulated practical experience to conceptualize the shapes of each component and coordinate their proportions. They then gradually approach the desired effect through hand-drawn designs and repeated trials. This process often requires multiple manual modifications and physical trials between the design and actual production. Each round of modifications involves redrawing, remaking the skeleton, and repainting, resulting in a lengthy process and significant material waste. Due to the lack of intelligent auxiliary tools capable of understanding and applying specific technological rules, artisans struggle to quickly assess the feasibility of structural proportions, pattern density, color schemes, and assembly relationships during the digital design phase. This necessitates extensive adjustments during the physical production stage, further increasing trial production costs and time. To address the problem that existing generative design tools are insufficient for practical manufacturing, it is necessary to propose an auxiliary manufacturing method that combines process knowledge, structural control, and manufacturability evaluation. This method would enable the system to not only generate appearance images but also output structural drawings, pattern areas, color matching rules, and constraint specifications related to the manufacturing process, thereby improving the efficiency and stability of craft design and manufacturing. Summary of the Invention
[0004] To address the aforementioned technical problems, this invention proposes a method and system for assisting in the production of handicrafts based on a multimodal process knowledge base, thereby resolving the issues present in the prior art.
[0005] Firstly, to achieve the above objectives, the present invention provides a method for assisting in the production of handicrafts based on a multimodal process knowledge base, comprising: S1. Establish a multimodal process dataset, which includes structured image data, text data, and video frame data, and establish the association between the same process object in different modal data. S2. Based on the multimodal process dataset, construct several process feature units. Each process feature unit contains structured process attributes and corresponding image vector representations, and a process knowledge base is composed of several process feature units. S3. Receive and parse the production requirements, parse the production requirements into a structured task description, and generate a query vector corresponding to the structured task description through a multimodal aligned text encoder. S4. Based on the query vector and the structured task description, retrieve matching process feature units from the process knowledge base to obtain a target process feature unit set; S5. Generate positive prompts, constraint prompts, and structural control conditions based on the target process feature unit set and the structured task description; S6. Call the generation model to generate candidate craft patterns based on the positive prompts, the constraint prompts, and the structural control conditions; S7. Conduct a process consistency evaluation on the candidate craft designs. The process consistency evaluation includes evaluation of task matching degree, style consistency, structural rationality, process manufacturability, and compliance with manufacturing constraints. S8. When the process consistency evaluation result does not meet the preset requirements, adjust the search weight, prompt word content, constraint prompt word strength, structural control strength or domain adaptation weight, and repeat S3 to S8; when the process consistency evaluation result meets the preset requirements, output the candidate craft pattern and the corresponding manufacturing auxiliary information.
[0006] Optionally, the process of building a multimodal process dataset includes: The process involves acquiring photos of the finished product, including physical images, partial structural images, pattern images, process flow text, material descriptions, and production records. The acquired images are then processed for deduplication, clarity filtering, and size standardization. Selected images are labeled with objects; when the craft is a dragon lantern, the labeled objects include the dragon head, body, lantern sections, dragon scales, auspicious cloud patterns, flame patterns, and skeletal nodes. Each labeled object is assigned a category, location, size, color, and pattern attribute. The acquired text is cleaned, its terminology standardized, and segmented into paragraphs. The resulting text is labeled with its shape structure, component proportions, materials and processes, color rules, pattern structure, assembly steps, and production constraints. The acquired video is processed by frame extraction to obtain image frames of the production process, retaining timestamps and process names.
[0007] Optionally, the process of constructing process feature units includes: performing object detection and image segmentation on the images in the multimodal process dataset; when the craft is a dragon lantern craft, obtaining the dragon head region, dragon body region, lantern section region, and pattern region, as well as the corresponding bounding boxes and segmentation masks; performing edge detection on the segmented target regions to extract the outer contour, curvature direction, and structural key points; performing color space conversion and color clustering on the segmented target regions to extract the primary hue, secondary color, and color ratio; converting the processed image into an image vector representation using an image encoder; extracting structured process attributes from the text in the multimodal process dataset using a large language model, the structured process attributes including shape structure, color rules, pattern structure, material process, production steps, and production constraints; associating and binding the image vector representation corresponding to the same process object with the structured process attributes to form process feature units.
[0008] Optionally, the process of parsing production requirements includes: parsing the user-inputted natural language production requirements into a structured task description using a large language model, wherein the structured task description includes fields such as product type, size specifications, media form, style requirements, and production constraints; converting the structured task description into a query vector using a text encoder; and normalizing the query vector to obtain a query vector for subsequent retrieval.
[0009] Optionally, the retrieval and matching process includes: calculating the similarity between the query vector and the image vectors of each process feature unit in the process knowledge base to obtain a vector similarity score; performing rule matching between the structured task description and the structured process attributes of each process feature unit, calculating the proportion of the number of matching items to the total number of rule matching items, and obtaining a rule matching score; performing a weighted sum based on the vector similarity score, the rule matching score, and a preset priority weight to obtain a comprehensive score for each process feature unit; selecting a preset number of process feature units in descending order of comprehensive score to obtain a candidate process feature set; and filtering the candidate process feature set according to the manufacturing constraints in the structured task description, eliminating candidates that do not meet the size specifications, material limitations, or processing methods to obtain a target process feature unit set.
[0010] Optionally, the process of generating positive prompts, constraint prompts, and structural control conditions includes: extracting the shape structure, color rules, pattern structure, material process, production steps, and production constraints from the structured process attributes of the target process feature unit set, and combining this with the product type, size specifications, and media form in the structured task description to generate positive prompts and constraint prompts through a preset template; performing target detection and image segmentation on the reference image corresponding to the target process feature unit set to obtain the bounding boxes and segmentation masks of each target region; performing edge detection or contour extraction on the segmented target regions to generate edge maps and contour maps; generating a layout structure diagram based on the product type and size specifications in the structured task description; and the edge map, contour map, segmentation map, layout map, or pattern region map constituting the structural control conditions.
[0011] Optionally, the process of calling the generative model to generate candidate craft patterns includes: inputting the positive prompts and the constraint prompts into a text encoder to obtain positive text condition vectors and constraint text condition vectors; inputting the structural control conditions into a structural control network to extract structural control features; injecting the structural control features into the intermediate layer of the diffusion model to participate in the noise prediction process; loading the domain model weights trained by the low-rank adaptation method to constrain the generation space of the diffusion model to the style space of the target craft; the diffusion model performing multi-step denoising based on the initial noise, the positive text condition vectors, the constraint text condition vectors, and the structural control features to obtain a latent space representation; and decoding the latent space representation into candidate craft patterns in the image space using a decoder.
[0012] Secondly, the present invention also provides a craft-assisted production system based on a multimodal process knowledge base, for implementing a craft-assisted production method based on a multimodal process knowledge base. The system includes: a data acquisition module for establishing a multimodal process dataset, the dataset containing structured image data, text data and video frame data, and establishing the association relationship between the same process object in different modal data. The process feature modeling module is used to construct several process feature units based on the multimodal process dataset. Each process feature unit contains structured process attributes and corresponding image vector representations, and the process feature units together form a process knowledge base. The task parsing module is used to receive and parse production requirements, and generate a structured task description and corresponding query vectors. The process retrieval module is used to retrieve matching process feature units from the process knowledge base based on the query vector and the structured task description, and obtain a set of target process feature units. The condition generation module is used to generate positive prompts, constraint prompts, and structural control conditions based on the target process feature unit set and the structured task description. The image generation module is used to call the generation model and generate candidate craft patterns based on the positive prompts, the constraint prompts, and the structural control conditions. The evaluation feedback module is used to evaluate the consistency of the process of the candidate craft images. When the evaluation result does not meet the requirements, the generation parameters are adjusted and the condition generation module and the image generation module are controlled to repeat the corresponding operation.
[0013] Thirdly, the present invention also provides a computer terminal device, comprising: One or more processors; A memory, coupled to the processor, for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the steps of the craft-aided manufacturing method based on a multimodal process knowledge base in the first aspect described above.
[0014] Fourthly, the present invention also provides a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, it implements the steps of the craft-aided manufacturing method based on a multimodal process knowledge base in the first aspect described above.
[0015] Compared with the prior art, the present invention has the following advantages and technical effects: This invention provides a method and system for assisting in the production of handicrafts based on a multimodal process knowledge base. It transforms the shape, structure, color scheme, pattern arrangement, material processes, and production constraints in handicraft production into searchable and iteratively optimized process data, thereby improving the efficiency from requirement input to production drawing output. Traditional production schemes rely on repeated drawing and trial production based on manual experience, making it difficult to quickly determine whether the structural proportions, pattern density, and color schemes are suitable for production in the early stages. This invention, through process feature modeling, enables the structured retrieval of information on different parts of the handicraft (e.g., dragon head, dragon body, lantern sections), patterns, materials, and processes. By generating structural control conditions, candidate drawings can follow reference contours, component proportions, and composition layouts, reducing common structural confusion problems in generated images. Through process consistency evaluation and feedback optimization mechanisms, prompts, constraint prompts, search weights, and structural control strength are adjusted based on manufacturability scores, reducing invalid schemes and improving the efficiency of scheme screening before handicraft prototyping. Attached Figure Description
[0016] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an undue limitation of the invention. In the drawings: Figure 1 This is a flowchart illustrating a method for assisting in the production of handicrafts based on a multimodal process knowledge base, according to an embodiment of the present invention. Figure 2 This is a schematic diagram of the structure of a craft-aided production system based on a multimodal process knowledge base, according to an embodiment of the present invention. Detailed Implementation
[0017] It should be noted that, unless otherwise specified, the embodiments and features described in the present invention can be combined with each other. The present invention will now be described in detail with reference to the accompanying drawings and embodiments.
[0018] It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.
[0019] Example 1 like Figure 1 As shown, this embodiment provides a method for assisting in the production of handicrafts based on a multimodal process knowledge base, including: S1. Establish a multimodal process dataset, which includes structured image data, text data, and video frame data, and establish the association between the same process object in different modal data. S2. Based on the multimodal process dataset, construct several process feature units. Each process feature unit contains structured process attributes and corresponding image vector representations, and a process knowledge base is composed of several process feature units. S3. Receive and parse the production requirements, and generate a structured task description and corresponding query vector; S4. Based on the query vector and the structured task description, retrieve matching process feature units from the process knowledge base to obtain a target process feature unit set; S5. Generate positive prompts, constraint prompts, and structural control conditions based on the target process feature unit set and the structured task description; S6. Call the generation model to generate candidate craft patterns based on the positive prompts, the constraint prompts, and the structural control conditions; S7. Conduct a process consistency evaluation on the candidate craft designs. The process consistency evaluation includes evaluation of task matching degree, style consistency, structural rationality, process manufacturability, and compliance with manufacturing constraints. S8. When the process consistency evaluation result does not meet the preset requirements, adjust the search weight, prompt word content, constraint prompt word strength, structural control strength or domain adaptation weight, and repeat S3 to S8; when the process consistency evaluation result meets the preset requirements, output the candidate craft pattern and the corresponding manufacturing auxiliary information.
[0020] Furthermore, the process of establishing a multimodal process dataset includes: Acquire physical photos, partial structural photos, pattern photos, process flow text, material descriptions, and production records related to the production of dragon lantern handicrafts; perform deduplication, clarity filtering, and size standardization on the acquired images; annotate the filtered images, including the dragon head, dragon body, lantern sections, dragon scales, auspicious cloud patterns, flame patterns, and skeleton nodes, and add category, location, size, color, and pattern attributes to each annotated object; clean the acquired text, standardize terminology, and segment paragraphs, and annotate the shape structure, component proportions, material processes, color rules, pattern structure, assembly steps, and production constraint fields; extract frames from the acquired video to obtain image frames of the production process and retain timestamps and process names.
[0021] Specifically, the implementation process of this embodiment includes: Step 1: First, acquire multi-source data related to the production of "Chengnan Dragon Lantern" handicrafts, including images, text, videos, production records, and process descriptions. Then, clean, label, and structure these data to form a multimodal process dataset. Specifically, this embodiment collects photos of target handicrafts such as actual dragon lanterns, partial photos of the dragon head, photos of the dragon body and lantern sections, partial photos of patterns, photos of the production process, images of existing handicraft examples, process flow text, material descriptions, size records, and experience descriptions from production personnel. The images are then deduplicated, filtered for clarity, standardized in size, and labeled to obtain image data including the dragon head, dragon body, lantern sections, patterns, skeleton, and finished product scene.
[0022] Image deduplication can be implemented using perceptual hashing algorithms, including average hashing, difference hashing, or perceptual hashing. The distance between image hash values is calculated to determine if images are duplicates. In a specific embodiment, open-source tools are used to implement near-duplicate image detection. Sharpness filtering employs a blur detection method based on the variance of the Laplacian operator. The high-frequency information intensity of the image is calculated, and when the image sharpness score is below a preset threshold, the image is marked as a low-quality image.
[0023] The Laplace variance method is written as: ; in, Indicates the input image. This represents the result after applying the Laplacian operator to the image. This represents the variance. When the variance is less than a preset threshold, the image is considered blurry and is discarded.
[0024] Size unification is achieved using image processing libraries such as OpenCV or Pillow. In this embodiment, the image is scaled to a preset size according to the requirements of subsequent model input, and the integrity of the main body area is maintained through proportional scaling, center cropping, or edge padding. Object annotation is achieved by combining model-assisted annotation with manual correction. Initial annotation results are first automatically generated using the closed-source large model ChatGPT, and then manually corrected. For images that cannot be annotated using a large model, this embodiment uses annotation tools to perform mask annotation, bounding box annotation, or attribute annotation on the dragon head, dragon body, lantern joints, dragon scales, auspicious cloud patterns, flame patterns, skeleton nodes, and finished product areas in the image, and adds category, position, size, color, pattern, and production description to each annotated object.
[0025] This embodiment also collects text data related to the production process, including process step descriptions, material usage records, size ratio records, color matching descriptions, pattern drawing descriptions, assembly sequence descriptions, and descriptions of existing design schemes. The text is cleaned, terminology is standardized, paragraphs are segmented, and fields are annotated to highlight aspects such as shape structure, component proportions, material processes, color rules, pattern structures, assembly steps, and production constraints. For video data, this embodiment obtains production process image frames through frame extraction or keyframe extraction, retaining timestamps, process names, component status, and action information. After processing, this embodiment establishes unified numbering and association relationships for images, text, video frames, and process descriptions corresponding to the same process object, forming a multimodal process dataset. The multimodal process dataset includes original material data, cleaned image data, structural control data, text process data, structured annotation data, vector representation data, and production case data; among which, structural control data includes edge maps, outline maps, line drawings, segmentation maps, and layout reference maps; and structured annotation data includes shape structure, size ratios, material processes, color rules, pattern structures, production steps, and production constraints.
[0026] Furthermore, the process of constructing process feature units includes: performing target detection and image segmentation on the images in the multimodal process dataset to obtain the dragon head region, dragon body region, lantern section region, and pattern region, as well as the corresponding bounding boxes and segmentation masks; performing edge detection on the segmented target regions to extract the outer contour, curvature direction, and structural key points; performing color space conversion and color clustering on the segmented target regions to extract the primary hue, secondary color, and color ratio; converting the processed image into an image vector representation using an image encoder; extracting structured process attributes from the text in the multimodal process dataset using a large language model, wherein the structured process attributes include shape structure, color rules, pattern structure, material process, production steps, and production constraints; and associating and binding the image vector representation corresponding to the same process object with the structured process attributes to form process feature units.
[0027] Specifically, the implementation process of this embodiment includes: Step 2: Using a combination of rule-based annotation, image feature extraction, and text field extraction, the production data of the "Chengnan Dragon Lantern" handicrafts is abstracted into process feature units. For image data, this embodiment extracts visual and structural features such as the dragon head, body, scales, lantern sections, decorative patterns, main color, auxiliary color, skeleton position, and component proportions through methods such as object detection, image segmentation, edge extraction, color clustering, and pattern recognition. These features are then converted into image vector representations for subsequent similarity retrieval and matching. For text and production record data, this embodiment extracts structured process attributes, including shape structure, color rules, pattern structure, material technology, process requirements, and production constraints, using natural language processing models or large language models.
[0028] For image data, the YOLO series object detection model is first used to identify the dragon head, body, lantern sections, tail, skeletal nodes, pattern areas, and finished product main body areas in the image, obtaining the category label, confidence score, and bounding box coordinates for each target object. For the detected target areas, this embodiment further uses the Segment Anything Model image segmentation model to extract target masks to obtain more accurate object contours. After object detection and image segmentation, this embodiment obtains images of the dragon head area, dragon body area, lantern section area, local pattern images, and corresponding bounding boxes, segmentation masks, and object category information.
[0029] For the dragon head, body, and lantern sections, this embodiment further extracts morphological and structural features. Canny edge detection is used to process the segmented target areas, obtaining the dragon head's outer contour, the dragon body's curvature, the lantern section arrangement, and local decorative contours. The contour extraction results can be represented as the set of boundary points of the target area, contour curve, length-to-width ratio, area ratio, degree of curvature, local corner points, and structural key points. For the dragon body or lantern section structure, this embodiment statistically analyzes the number of lantern sections, the spacing between lantern sections, the dragon body length ratio, and the dragon body curvature direction based on the target detection results, thereby forming visual features to describe the structure of the craftwork.
[0030] For color features, this embodiment performs color space conversion and color clustering processing on the image or target region. This embodiment converts the RGB image to HSV or Lab color space and uses K-means clustering, color histogram statistics, or dominant color extraction algorithms to obtain information such as the dominant hue, secondary colors, color proportions, and color saturation in the image. For pattern features, this embodiment first locates the pattern region using YOLO series object detection or Segment Anything Model segmentation methods, and then determines the pattern category using a texture feature extraction model or image classification model, obtaining the pattern category, pattern distribution location, pattern density, pattern repetition direction, and spatial relationship between the pattern and the main structure.
[0031] For each image or each target object, this embodiment converts it into an image vector representation using an image encoder. The image encoder can employ a visual Transformer encoding model. The image vector representation is as follows: ; in, Indicates the first i An image or image region, Indicates an image encoder. This represents the image vector. This image vector is used for subsequent similarity retrieval and matching with the task query vector or reference image vector.
[0032] For text and production record data, this embodiment first performs text preprocessing, including noise reduction, sentence segmentation, paragraphing, terminology standardization, duplicate content removal, and preliminary keyword screening. Subsequently, this embodiment uses a large language model or rule template to extract information from the text content. Extraction targets include shape structure, color rules, pattern structure, material processes, production steps, assembly relationships, and production constraints. For example, this embodiment extracts fields such as "shape structure = open-mouthed front-view dragon head," "main color = red," "pattern structure = continuous scales," and "production constraints = avoid overly dense patterns" from the statement "the dragon head skeleton must maintain a forward-viewing open mouth shape, the main color is red, the scales are continuously distributed along the dragon body, and overly dense patterns should not be used during production to avoid affecting painting efficiency."
[0033] For the same process object, this embodiment associates and binds its corresponding image vector representation with structured process attributes to construct a process feature unit. The process feature unit is represented as... ,in Indicates structured process attributes, This represents the vector representation of the encoded image data. Structured process attributes. Represented as ,in Indicates the shape and structure. Representing color rules, Indicates the pattern structure. Indicates material processing, Indicates the production steps, This indicates the constraints to be created. To ensure the accuracy of the structured attributes, this embodiment introduces a manual correction mechanism, where process engineers or designers review the automatically extracted fields. The reviewed fields are written into the process knowledge base, and fields that fail the review are modified, deleted, or marked as low confidence.
[0034] Furthermore, the process of parsing production requirements includes: parsing the user-inputted natural language production requirements into a structured task description using a large language model, wherein the structured task description includes fields such as product type, size specifications, media form, style requirements, and production constraints; converting the structured task description into a query vector using a text encoder; and normalizing the query vector to obtain a query vector for subsequent retrieval.
[0035] Specifically, the implementation process of this embodiment includes: Step 3: Perform structured parsing of the user-inputted craft production requirements to obtain a task description that includes product type, target purpose, size specifications, medium form, process complexity, material preferences, and production constraints. After the user inputs their production requirements, this embodiment first parses the natural language requirements into a structured task description using rule templates or a large language model: ; Wherein, Product indicates product type, Use indicates intended use, Size indicates size specifications, Medium indicates output medium or carrier, Style indicates style requirements, ProcessLevel indicates process complexity level, and Constraints indicate design and manufacturing constraints. Subsequently, this embodiment uses a text encoder to... Convert to query vector q: .
[0036] Furthermore, the retrieval and matching process includes: calculating the similarity between the query vector and the image vectors of each process feature unit in the process knowledge base to obtain a vector similarity score; performing rule matching between the structured task description and the structured process attributes of each process feature unit, calculating the proportion of the number of matching items to the total number of rule matching items, and obtaining a rule matching score; performing a weighted sum based on the vector similarity score, the rule matching score, and a preset priority weight to obtain a comprehensive score for each process feature unit; selecting a preset number of process feature units in descending order of comprehensive score to obtain a candidate process feature set; and filtering the candidate process feature set according to the manufacturing constraints in the structured task description, eliminating candidates that do not meet the size specifications, material limitations, or processing methods to obtain a target process feature unit set.
[0037] Specifically, the implementation process of this embodiment includes: Step 3 also includes: In this embodiment, the query vector q is compared with the image vector representation of each process feature unit in the knowledge base. Matching is performed, and a comprehensive similarity score is calculated by combining structured process attributes and manual priority. Each process feature unit is represented as... The comprehensive similarity scoring function is as follows: ; in, The vector similarity between the query vector and the process feature image vector can be calculated using cosine similarity. ; This indicates the degree of rule matching between the structured task description and the process feature attributes, such as whether dimensions, uses, component types, color schemes, materials, and manufacturing constraints are consistent, and is represented as: ; in, M The number of rule matches. Indicates the first m The rule is checked for a match; if it matches, the value is 1, and if it doesn't match, the value is 0. Different weights can also be set according to the importance of the rule.
[0038] This indicates the priority set manually or confirmed by the engineer. , , These represent the weights of vector similarity, process attribute matching, and manual priority in the retrieval process, respectively.
[0039] Calculations yielded all process feature units. Then, in this embodiment, the candidate process features are sorted from high to low according to their scores to obtain a set of candidate process features: ; in, TopK This indicates that the highest score is selected. K Each process feature unit.
[0040] Subsequently, this embodiment performs a secondary screening of the candidate set based on the manufacturing task constraints, eliminating candidates that do not meet the size specifications, material limitations, processing methods, pattern complexity, or assembly requirements, thereby obtaining the target process feature unit set. The screening process can be represented as follows: ; in, This represents the final target process feature unit. The structured process attributes in the target process feature unit will be used for subsequent prompt generation, constraint prompt generation, structural control condition generation, and manufacturing scheme evaluation.
[0041] Furthermore, the process of generating positive prompts, constraint prompts, and structural control conditions includes: extracting the shape structure, color rules, pattern structure, material process, production steps, and production constraints from the structured process attributes of the target process feature unit set, and combining this with the product type, size specifications, and media form in the structured task description to generate positive prompts and constraint prompts through a preset template; performing target detection and image segmentation on the reference image corresponding to the target process feature unit set to obtain the bounding boxes and segmentation masks of each target region; performing edge detection or contour extraction on the segmented target regions to generate edge maps and contour maps; generating a layout structure diagram based on the product type and size specifications in the structured task description; and the edge map, contour map, segmentation map, layout map, or pattern region map constitute the structural control conditions.
[0042] Specifically, the implementation process of this embodiment includes: Step 4: Read the target process feature unit set obtained in Step 3, and automatically generate positive prompts, constraint prompts, and structural control conditions based on the structured process attributes, reference images, and task analysis results. These are used to guide the subsequent generation model to output candidate solutions that can be used for craft production. In this embodiment, the shape structure, color rules, pattern structure, material process, production steps, and production constraints are extracted from the structured process attributes. Combined with the product type, size specifications, medium form, and process complexity in the task analysis results, positive prompts are generated by rewriting using preset templates or large language models. The shape structure is transformed into a main description, the color rules into a color scheme description, the pattern structure into a decorative description, the material process into a production process description, and the size and purpose into a product specification description.
[0043] During the constraint prompt word generation process, this embodiment reads the manufacturing constraints from the structured process attributes. Based on the size and material limitations in the task analysis results, constraint warning words are generated to suppress unmanufacturable or incompatible results. These constraint warning words mainly include issues such as incorrect structural proportions, unreasonable lamp section arrangement, excessively dense patterns, colors not matching material properties, inability to assemble the skeleton, and inability to manufacture certain components. Positive warning words are represented as: ; Among them, Template( ) represents the prompt word template function.
[0044] The constraint prompt is represented as: ; in, Rule ( ) represents the constraint mapping function.
[0045] The comprehensive prompt words are represented as follows: ; in, Indicates the strength of the constraint prompt.
[0046] In the process of generating structural control conditions, this embodiment does not directly use the original images in the knowledge base as structural control inputs. Instead, it extracts structural information from retrieved reference images or automatically generates structural information based on structured process attributes and manufacturing tasks. This embodiment can identify the dragon head, dragon body, lantern sections, patterns, and skeleton regions in the reference image using a target detection model to obtain bounding boxes. A target mask is obtained using an image segmentation model, and edge maps or contour maps are generated using edge detection or contour extraction methods. Alternatively, a layout structure diagram can be generated based on product type, size specifications, and manufacturing scenario. The aforementioned edge maps, contour maps, segmentation maps, layout diagrams, or pattern region maps collectively constitute the structural control conditions. It is used to constrain the main shape, spatial layout, and pattern distribution of the subsequently generated image.
[0047] Furthermore, the process of calling the generative model to generate candidate craft patterns includes: inputting the positive prompts and the constraint prompts into a text encoder to obtain positive text condition vectors and constraint text condition vectors; inputting the structural control conditions into a structural control network to extract structural control features; injecting the structural control features into the intermediate layer of the diffusion model to participate in the noise prediction process; loading the domain model weights trained by the low-rank adaptation method to constrain the generation space of the diffusion model to the style space of the target craft; the diffusion model performing multi-step denoising based on the initial noise, the positive text condition vectors, the constraint text condition vectors, and the structural control features to obtain a latent space representation; and decoding the latent space representation into candidate craft patterns in the image space using a decoder.
[0048] Specifically, the implementation process of this embodiment includes: Step 5: In the generation stage, this embodiment uses the StableDiffusion image generation model based on a diffusion model to generate candidate designs for the crafts. The input to this generation model includes positive prompts. Constraint prompts Random noise Structural control conditions and domain adaptation weight L The output is the target image that meets the requirements of the production task. I Its generation process is represented as follows: ; In this embodiment, positive prompts and constraint prompts are first input into a text encoder to obtain positive text condition vectors and constraint text condition vectors, respectively. Positive text conditions are used to describe the desired shape, color scheme, pattern, material, and purpose of the craft item, while constraint text conditions are used to suppress undesirable content such as structures that are impossible to manufacture, patterns that are too dense, unreasonable skeletons, or local parts that are difficult to process.
[0049] In the denoising process, this embodiment employs a stepwise denoising mechanism using a diffusion model, starting from initial noise and gradually recovering the latent space representation of the target image through multiple time steps. To ensure the generated result conforms to the style and production characteristics of the target handicraft, this embodiment introduces a domain model trained using the low-rank adaptation method LoRA on top of the basic diffusion model. During the LoRA training phase, images of the dragon lantern handicraft and their accompanying descriptive text are used as training data to learn low-rank parameters for the UNet network and an optional text encoder in the basic diffusion model. For the original weight matrix W, LoRA learns low-rank increments. ,make ,and ,in and It is a low-rank matrix. The role of LoRA is to constrain the general visual generation capability of the base model to the style space of the target dragon lantern craft, making it easier for the generative model to generate candidate patterns with the target shape, color scheme and pattern combination.
[0050] To ensure the rationality of the generated results in terms of structural layout and component proportions, this embodiment further introduces a structural control network, ControlNet. The input to the structural control network is the structural control conditions generated in step 4. The structural control conditions can be edge maps, contour maps, layout maps, semantic segmentation maps, or pattern region maps. In this embodiment, the structural control conditions are first input into the structural control network to extract structural features: ; in, Represents the structural control network. Indicates the first The structural control features extracted at each time step are then injected into the intermediate layer of the basic diffusion model to influence the noise prediction results, thereby ensuring that the generated image maintains consistency with the structural control conditions in terms of subject position, contour shape, light cluster arrangement, and pattern distribution. The noise prediction after incorporating structural control is expressed as follows: ; in, This indicates the main generator network with LoRA weights loaded. Indicates structural control strength. This is achieved through adjustment. This embodiment controls the degree to which the generated image conforms to structural conditions. When When the image is larger, the generated image adheres more strictly to the reference contour or layout. When the size is smaller, the generated images have greater freedom and creativity.
[0051] After multiple denoising steps, this embodiment obtains the final representation in the latent space, and decodes the latent variable into a candidate craft pattern in the image space through a decoder.
[0052] After step 5, this embodiment obtains multiple candidate craft design images while retaining the corresponding prompts, constraint prompts, structural control conditions, LoRA weights, control strengths, and random seeds. These candidate results can be used not only for visual display but also as a basis for subsequent generation of patterns, design layouts, color schemes, and process descriptions.
[0053] Further, in step 6: During the result evaluation stage, this embodiment evaluates the process consistency and manufacturability of the generated candidate craft designs. The evaluation dimensions include task matching, style consistency, structural rationality, process manufacturability, and compliance with manufacturing constraints. The evaluation function can be expressed as: ; in, This indicates whether the candidate image matches the product type, purpose, and size specifications entered by the user. This indicates whether the candidate image matches the style characteristics of the target craft item. Indicate whether the dragon head, body, lantern sections, patterned areas, and composition layout are reasonable. This indicates whether the materials, component proportions, pattern density, and assembly relationships are suitable for manufacturing. This indicates whether the production constraints have been violated. α, β, γ, δ, and μ are weighting coefficients that are adjusted according to the task requirements.
[0054] This embodiment calculates the task matching degree between the generated image and the task description using a text-image similarity model, calculates the style consistency between the generated image and the target process feature reference image using an image similarity model, determines whether the generated result meets the structural requirements of the dragon head, dragon body, lantern sections, and pattern areas using a structural detection method, and checks whether the generated result violates manufacturing constraints using a process rule base. When the overall score is lower than a preset threshold, this embodiment adjusts the prompt words, constraint prompt words, knowledge retrieval weights, LoRA weights, and structural control strength based on feedback from process engineers or designers, and executes the generation process again until the generated result meets the preset process consistency and manufacturability requirements.
[0055] Example 2 like Figure 2 As shown, based on the same general inventive concept, this invention also provides a craft-aided production system based on a multimodal process knowledge base. The craft-aided production system based on a multimodal process knowledge base provided by this invention will be described below. The craft-aided production system based on a multimodal process knowledge base described below can be referred to in conjunction with the craft-aided production method based on a multimodal process knowledge base described above. The system includes: The data acquisition module is used to establish a multimodal process dataset, which includes structured image data, text data, and video frame data, and to establish the correlation between the same process object in different modal data. The process feature modeling module is used to construct several process feature units based on the multimodal process dataset. Each process feature unit contains structured process attributes and corresponding image vector representations, and the process feature units together form a process knowledge base. The task parsing module is used to receive and parse production requirements, and generate a structured task description and corresponding query vectors. The process retrieval module is used to retrieve matching process feature units from the process knowledge base based on the query vector and the structured task description, and obtain a set of target process feature units. The condition generation module is used to generate positive prompts, constraint prompts, and structural control conditions based on the target process feature unit set and the structured task description. The image generation module is used to call the generation model and generate candidate craft patterns based on the positive prompts, the constraint prompts, and the structural control conditions. The evaluation feedback module is used to evaluate the consistency of the process of the candidate craft images. When the evaluation result does not meet the requirements, the generation parameters are adjusted and the condition generation module and the image generation module are controlled to repeat the corresponding operation.
[0056] Example 3 (equivalent to Example 2 above) In this embodiment, a computer terminal device is provided, including: One or more processors; A memory, coupled to the processor, for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the steps of the above-described craft-aided manufacturing method based on a multimodal process knowledge base.
[0057] In this embodiment, a computer-readable storage medium is also provided, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the above-described method for assisting in the production of handicrafts based on a multimodal process knowledge base.
[0058] In the technical solution proposed in this invention, the appearance drawing of the craft can also be directly generated using only a general image generation model. However, this method lacks process knowledge retrieval and manufacturing constraints, easily resulting in unreasonable structural proportions or difficulties in manufacturing. Alternatively, the craft can be manufactured entirely through manual drawing and prototyping, but this method is time-consuming, has high trial-and-error costs, and makes it difficult to quickly generate multiple comparable solutions. Therefore, within the applicant's knowledge, the solution proposed in this invention, based on a multimodal process knowledge base, structural control generation, and manufacturability feedback optimization, has significant advantages in improving the efficiency of manufacturing solution generation, reducing trial production costs, and improving the usability of drawings.
[0059] This invention provides a method and system for assisting in the production of handicrafts based on a multimodal process knowledge base. It transforms the shape, structure, color scheme, pattern arrangement, material processes, and production constraints in handicraft production into searchable and iteratively optimized process data, thereby improving the efficiency from requirement input to production drawing output. Traditional production schemes rely on repeated drawing and trial production based on manual experience, making it difficult to quickly determine whether the structural proportions, pattern density, and color schemes are suitable for production in the early stages. This invention enables the structured retrieval of information on the dragon head, body, lantern sections, patterns, materials, and processes through process feature modeling. It generates structural control conditions to ensure that candidate drawings follow reference contours, component proportions, and composition layouts, reducing common structural confusion problems in generated images. Through process consistency evaluation and feedback optimization mechanisms, it adjusts prompts, constraint prompts, search weights, and structural control strength based on manufacturability scores, reducing invalid schemes and improving the efficiency of scheme screening before handicraft prototyping.
[0060] The above are merely preferred embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A method for assisting in the production of handicrafts based on a multimodal process knowledge base, characterized in that, include: S1. Establish a multimodal process dataset, which includes structured image data, text data, and video frame data, and establish the association between the same process object in different modal data. S2. Based on the multimodal process dataset, construct several process feature units. Each process feature unit contains structured process attributes and corresponding image vector representations, and a process knowledge base is composed of several process feature units. S3. Receive and parse the production requirements, parse the production requirements into a structured task description, and generate a query vector corresponding to the structured task description through a multimodal aligned text encoder. S4. Based on the query vector and the structured task description, retrieve matching process feature units from the process knowledge base to obtain a target process feature unit set; S5. Generate positive prompts, constraint prompts, and structural control conditions based on the target process feature unit set and the structured task description; S6. Call the generation model to generate candidate craft patterns based on the positive prompts, the constraint prompts, and the structural control conditions; S7. Conduct a process consistency evaluation on the candidate craft designs. The process consistency evaluation includes evaluation of task matching degree, style consistency, structural rationality, process manufacturability, and compliance with manufacturing constraints. S8. When the process consistency evaluation result does not meet the preset requirements, adjust the search weight, prompt word content, constraint prompt word strength, structural control strength or domain adaptation weight, and repeat S3 to S8; when the process consistency evaluation result meets the preset requirements, output the candidate craft pattern and the corresponding manufacturing auxiliary information.
2. The method according to claim 1, characterized in that, The process of establishing a multimodal process dataset includes: The process involves acquiring photos of the finished product, including physical images, partial structural images, pattern images, process flow text, material descriptions, and production records. The acquired images are then processed for deduplication, clarity filtering, and size standardization. Selected images are labeled with objects; when the craft is a dragon lantern, the labeled objects include the dragon head, body, lantern sections, dragon scales, auspicious cloud patterns, flame patterns, and skeletal nodes. Each labeled object is assigned a category, location, size, color, and pattern attribute. The acquired text is cleaned, its terminology standardized, and segmented into paragraphs. The resulting text is labeled with its shape structure, component proportions, materials and processes, color rules, pattern structure, assembly steps, and production constraints. The acquired video is processed by frame extraction to obtain image frames of the production process, retaining timestamps and process names.
3. The method according to claim 1, characterized in that, The process of constructing process feature units includes: performing object detection and image segmentation on the images in the multimodal process dataset; when the craft is a dragon lantern craft, obtaining the dragon head region, dragon body region, lantern section region, and pattern region, as well as the corresponding bounding boxes and segmentation masks; performing edge detection on the segmented target regions to extract the outer contour, curvature direction, and structural key points; performing color space conversion and color clustering on the segmented target regions to extract the primary hue, secondary color, and color ratio; converting the processed image into an image vector representation using an image encoder; extracting structured process attributes from the text in the multimodal process dataset using a large language model, the structured process attributes including shape structure, color rules, pattern structure, material process, production steps, and production constraints; and associating and binding the image vector representation corresponding to the same process object with the structured process attributes to form process feature units.
4. The method according to claim 1, characterized in that, The process of parsing production requirements includes: parsing the user's input natural language production requirements into a structured task description using a large language model, wherein the structured task description includes fields such as product type, size specifications, media form, style requirements, and production constraints; converting the structured task description into a query vector using a text encoder; and normalizing the query vector to obtain a query vector for subsequent retrieval.
5. The method according to claim 1, characterized in that, The retrieval and matching process includes: calculating the similarity between the query vector and the image vectors of each process feature unit in the process knowledge base to obtain a vector similarity score; performing rule matching between the structured task description and the structured process attributes of each process feature unit, calculating the proportion of the number of matching items to the total number of rule matching items, and obtaining a rule matching score; performing a weighted sum based on the vector similarity score, the rule matching score, and a preset priority weight to obtain a comprehensive score for each process feature unit; selecting a preset number of process feature units in descending order of comprehensive score to obtain a candidate process feature set; and filtering the candidate process feature set according to the manufacturing constraints in the structured task description, eliminating candidates that do not meet the size specifications, material limitations, or processing methods to obtain a target process feature unit set.
6. The method according to claim 1, characterized in that, The process of generating positive prompts, constraint prompts, and structural control conditions includes: extracting shape structure, color rules, pattern structure, material process, production steps, and production constraints from the structured process attributes of the target process feature unit set, and combining this with the product type, size specifications, and media form in the structured task description to generate positive prompts and constraint prompts through a preset template; performing target detection and image segmentation on the reference image corresponding to the target process feature unit set to obtain the bounding boxes and segmentation masks of each target region; performing edge detection or contour extraction on the segmented target regions to generate edge maps and contour maps; generating a layout structure diagram based on the product type and size specifications in the structured task description; and the edge map, contour map, segmentation map, layout map, or pattern region map constitute the structural control conditions.
7. The method according to claim 1, characterized in that, The process of generating candidate craft patterns using a generative model includes: inputting the positive prompts and constraint prompts into a text encoder to obtain positive text condition vectors and constraint text condition vectors; inputting the structural control conditions into a structural control network to extract structural control features; injecting the structural control features into the intermediate layer of the diffusion model to participate in the noise prediction process; loading the domain model weights trained using a low-rank adaptation method to constrain the generation space of the diffusion model to the style space of the target craft; the diffusion model performing multi-step denoising based on the initial noise, the positive text condition vectors, the constraint text condition vectors, and the structural control features to obtain a latent space representation; and decoding the latent space representation into candidate craft patterns in the image space using a decoder.
8. A craft-aided production system based on a multimodal process knowledge base, characterized in that, The system for implementing the method of any one of claims 1-7 comprises: The data acquisition module is used to establish a multimodal process dataset, which includes structured image data, text data, and video frame data, and to establish the correlation between the same process object in different modal data. The process feature modeling module is used to construct several process feature units based on the multimodal process dataset. Each process feature unit contains structured process attributes and corresponding image vector representations, and the process feature units together form a process knowledge base. The task parsing module is used to receive and parse production requirements, and generate a structured task description and corresponding query vectors. The process retrieval module is used to retrieve matching process feature units from the process knowledge base based on the query vector and the structured task description, and obtain a set of target process feature units. The condition generation module is used to generate positive prompts, constraint prompts, and structural control conditions based on the target process feature unit set and the structured task description. The image generation module is used to call the generation model and generate candidate craft patterns based on the positive prompts, the constraint prompts, and the structural control conditions. The evaluation feedback module is used to evaluate the consistency of the process of the candidate craft images. When the evaluation result does not meet the requirements, the generation parameters are adjusted and the condition generation module and the image generation module are controlled to repeat the corresponding operation.
9. A computer terminal device, characterized in that, include: One or more processors; A memory, coupled to the processor, for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors perform the steps of the method as described in any one of claims 1-7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1-7.