A jewelry 3D model design method and system
By collecting data from multiple sources and using multi-task deep learning models, a cross-modal association database is constructed to achieve objective quantitative evaluation of jewelry design. This solves the problem of relying on subjective experience in traditional design and improves design quality and efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2026-03-17
AI Technical Summary
In traditional jewelry design, aesthetic evaluation relies on the designer's subjective experience and lacks objective standards. This results in design quality being highly dependent on individual ability, a high rework rate, an inability to accurately predict consumer preferences, a high return rate after product launch, and damage to the brand image.
Multimodal data of the entire life cycle of jewelry design is acquired through a multi-source data acquisition platform. Three-dimensional feature matrices are extracted using computer vision and natural language processing technologies. A cross-modal association database is constructed. A multi-task deep learning model is used to generate a multimodal aesthetic feature database. A multi-task AI aesthetic evaluation model is trained to conduct quantitative evaluation of 3D designs.
It enables objective and quantitative evaluation of jewelry design, improves design quality and efficiency, reduces rework rates, increases design efficiency for both designers and non-designer users, and enhances the objectivity and accuracy of the evaluation.
Smart Images

Figure CN120763533B_ABST
Abstract
Description
Technical Field
[0001] This invention generally relates to the field of artificial intelligence technology. More specifically, this invention relates to a method and system for designing 3D models of jewelry. Background Technology
[0002] Jewelry design is a complex process that integrates aesthetic creativity, craftsmanship, and user needs. Its core objective is to achieve a unity of visual appeal and functional wearing through the combination of form, materials, and color. In the traditional design process, designers must complete all stages from conception and sketching to 3D modeling. Aesthetic evaluation is a crucial step in determining the success or failure of a design—including judgments on core dimensions such as symmetry, proportional harmony, and stylistic fit. However, these evaluations have long relied on the designer's subjective experience and aesthetic preferences, lacking quantifiable objective standards, resulting in design quality being highly dependent on individual ability. Although modern CAD tools (such as Rhino and Matrix) have achieved parametric geometric modeling, the "good or bad" judgment at the aesthetic level still remains at the stage of manual review.
[0003] With the upgrading of consumption, the jewelry market is shifting from standardized mass production to personalized customization, and users' aesthetic requirements for design are showing a diversified trend (such as minimalism, Chinese trend style, retro deconstructionism, and other niche demands). The shortcomings of the traditional manual evaluation model are further amplified. The high rework rate caused by inconsistent aesthetic evaluation standards leads to high design costs and extended cycles. In addition, evaluations relying on human experience cannot accurately predict consumer preferences, resulting in high return rates after product launch and damage to brand image.
[0004] In view of this, there is an urgent need to provide a 3D model design method for jewelry, so as to establish an objective and quantitative evaluation system through digital and intelligent means, and transform implicit aesthetic experience into calculable and optimizable explicit knowledge. Summary of the Invention
[0005] In order to at least solve one or more of the technical problems mentioned above, the present invention proposes a method for designing 3D models of jewelry in several aspects.
[0006] In a first aspect, the present invention provides a method for designing 3D models of jewelry, characterized by comprising: acquiring historical jewelry design full-modal data; using computer vision and natural language processing technologies to extract and annotate multi-dimensional features from the full-modal data to construct a three-dimensional feature matrix; constructing a cross-modal association database based on the three-dimensional feature matrix through cross-modal structured annotation; generating a multi-modal aesthetic feature database using the cross-modal association database and a multi-task deep learning model; training a multi-task AI aesthetic evaluation model integrating geometry, style, and user preferences using the multi-modal aesthetic feature database; acquiring jewelry 3D design data through a graphical interface, and using the multi-task AI evaluation model to perform quantitative evaluation of the 3D design to generate a design evaluation result.
[0007] In some embodiments, acquiring historical jewelry design full-modal data includes: acquiring multimodal data covering the entire lifecycle of jewelry design through a multi-source data acquisition platform; wherein the multi-source data acquisition platform includes a design end, a user end, and a production end; the full-modal data includes 2D renderings, 3D model files, designer ratings, and user purchase feedback.
[0008] In some embodiments, the three-dimensional feature matrix includes a geometric feature dimension, a style semantic dimension, and a sentiment preference dimension.
[0009] In some embodiments, a cross-modal association database is constructed based on the three-dimensional feature matrix through cross-modal structured annotation, including: constructing a cross-modal indexing system using knowledge graph technology, and performing structured annotation of quantified features through the cross-modal indexing system combined with weakly supervised learning.
[0010] In some embodiments, the use of knowledge graph technology to construct a cross-modal indexing system includes: structurally associating the 2D images and 3D model data of each design case with their geometric features, style tags, and user sentiment scores to construct a multimodal training dataset that can be recognized by AI models.
[0011] In some embodiments, generating a multimodal aesthetic feature database using a cross-modal association database and a multi-task deep learning model includes: generating multimodal aesthetic feature data containing quantitative aesthetic scores, defect heatmaps, and style matching degrees using a cross-modal association database and a multi-task deep learning model.
[0012] In some embodiments, the multi-task AI aesthetic evaluation model includes: evaluating and generating a three-dimensional quantitative aesthetic score, a defect location heatmap, and a style matching report based on voxelized data of a 3D design model, 2D multi-view images, style keywords and design parameters input by the user.
[0013] In some embodiments, the design method further includes: generating multiple versions of aesthetic enhancement schemes for the 3D design based on the design evaluation results, combined with user preferences and process constraints, through a conditional generative adversarial network; and dynamically optimizing the design knowledge base and AI aesthetic evaluation model by recording user decisions through reinforcement learning.
[0014] In a second aspect, the present invention provides a 3D model design system for jewelry, comprising: a data acquisition module for acquiring historical jewelry design data across all modalities; a feature matrix generation module for extracting and annotating multi-dimensional features from the all-modal data using computer vision and natural language processing technologies to construct a three-dimensional feature matrix; an association database generation module for constructing a cross-modal association database based on the three-dimensional feature matrix through cross-modal structured annotation; an aesthetic feature database generation module for generating a multi-modal aesthetic feature database using the cross-modal association database and a multi-task deep learning model; an AI aesthetic evaluation model training module for training a multi-task AI aesthetic evaluation model integrating geometry, style, and user preferences using the multi-modal aesthetic feature database; and a design evaluation module for acquiring 3D jewelry design data through a graphical interface and using the multi-task AI evaluation model to perform quantitative evaluation of the 3D design to generate design evaluation results.
[0015] The jewelry 3D model design method described above utilizes multimodal aesthetic feature data for training and allows users to dynamically adjust aesthetic weights, balancing the rigor of professional design with consumers' personalized preferences. This enables non-designer users to complete high-quality designs through an interactive interface, improving the efficiency of producing satisfactory solutions. Furthermore, by constructing an AI aesthetic evaluation model, the objectivity of the evaluation is significantly improved, enhancing design efficiency and quality. Attached Figure Description
[0016] The above and other objects, features, and advantages of exemplary embodiments of the present invention will become readily apparent upon reading the following detailed description with reference to the accompanying drawings. In the drawings, several embodiments of the invention are illustrated by way of example and not limitation, and like or corresponding reference numerals denote like or corresponding parts, wherein:
[0017] Figure 1 An exemplary flowchart of a jewelry 3D model design method according to some embodiments of the present invention is shown;
[0018] Figure 2 An exemplary structural block diagram of a jewelry 3D model design system according to other embodiments of the present invention is shown. Detailed Implementation
[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0020] The specific embodiments of the present invention will now be described in detail with reference to the accompanying drawings.
[0021] Figure 1 An exemplary flowchart of a jewelry 3D model design method 100 according to some embodiments of the present invention is shown.
[0022] like Figure 1 As shown, in step S101 of method 100, historical jewelry design full modal data is obtained.
[0023] In one embodiment, acquiring historical jewelry design full-modal data includes obtaining multi-modal data covering the entire lifecycle of jewelry design through a multi-source data acquisition platform. This multi-source data acquisition platform includes design, user, and production endpoints. The full-modal data includes 2D renderings, 3D model files, designer ratings, and user purchase feedback. The process of acquiring data through the multi-source data acquisition platform is described in detail below.
[0024] Multimodal data covering the entire lifecycle of jewelry design is acquired through a multi-source data acquisition platform, specifically including:
[0025] Design data mainly includes 2D renderings (such as JPEG / PNG format, resolution ≥1024×1024), 3D model files (such as STL / OBJ / B-rep format, accuracy ≤0.01mm), CAD engineering drawings (DWG format), and design specification texts (such as style descriptions, material parameters, process requirements, etc.).
[0026] User-side data mainly includes user reviews on e-commerce platforms, purchase records, personalized customization requests (such as "preferring rose gold material" or "requesting emerald inlay"), and wearing size data (such as ring size and wrist circumference).
[0027] Production-side data mainly includes 3D printing process parameters (such as support structure design and printing layer thickness), precious metal loss rate, and finished product quality inspection reports (such as surface roughness and inlay firmness).
[0028] In terms of technical implementation, a raw dataset containing 100,000+ samples can be built through web crawling (targeting publicly available design cases), integration with enterprise CRM systems (to obtain user-customized data), and industrial sensor data acquisition (to obtain production data). The data storage adopts a distributed file system (HDFS) and supports TB-level data expansion.
[0029] Furthermore, in step S102: computer vision and natural language processing technologies are used to extract and associate multi-dimensional features from the full-modal data to construct a three-dimensional feature matrix.
[0030] In one embodiment, the three-dimensional feature matrix may include geometric feature dimensions, style semantic dimensions, and sentiment preference dimensions. The specific three-dimensional feature extraction process is as follows:
[0031] Geometric feature extraction (CV technology): The 3D model is analyzed using the OpenCASCADE library to calculate the number of symmetry axes, the fit of the golden ratio (the Fibonacci ratio of the main stone position to the contour), and the uniformity of surface curvature (using the Laplacian operator for detection, with a standard deviation of curvature <0.05mm). -1 ), volume-to-surface-area ratio (to assess the rationality of material usage); perform edge detection (Canny operator) on 2D rendered images, and extract contour complexity (such as perimeter / area ratio) and element distribution uniformity (which can be calculated through the gray-level co-occurrence matrix).
[0032] Style semantic extraction (NLP technology): The design specification text is parsed using a BERT pre-trained model to extract style keywords (such as "Baroque" and "Cyberpunk"), which are then mapped to a pre-defined style feature library (e.g., containing 10 basic styles and 20 sub-styles, with each style defining 10+ geometric / visual features. For example, "minimalist style" requires ≤3 types of elements and a line simplicity of >80%). A "style-feature" mapping table is then constructed; for example, the "New Chinese style" is associated with feature tags such as "meander pattern," "chalcedony material," and "symmetrical openwork structure."
[0033] Sentiment Preference Extraction (NLP + Machine Learning): LSTM models are used to perform sentiment analysis on user reviews. For example, five core sentiment dimensions such as "refinement," "uniqueness," and "value for money" can be extracted, outputting a sentiment score of 1-5. Then, purchase data is combined to calculate preference weights (e.g., the "comfort" weight is automatically increased by 20% for designs with high repurchase rates).
[0034] In the specific technical implementation process, each design sample can generate a feature vector V = [G, S, E] with 100+ dimensions, where geometric features G occupy 40 dimensions (including axis of symmetry, curvature, etc.), style semantics S occupy 30 dimensions (including style tags, element types, etc.), and sentiment preference E occupy 30 dimensions (including sentiment score, purchase conversion rate, etc.), forming a standardized three-dimensional feature matrix.
[0035] The following example illustrates the feature extraction process. For instance, the process of collecting dimensional features from the full modality data of a vintage-style ring is as follows:
[0036] Geometric feature dimension: The curvature of the ring holder surface is obtained through 3D scanning (e.g., average curvature of 0.08mm). -1 ), the golden ratio of the main stone to the ring band (0.618:1, 92% fit), and the number of symmetry axes (2 vertical symmetry axes).
[0037] Style semantic dimension: NLP analysis of design specifications extracts "Baroque" style keywords and associates them with preset features (such as carving density > 40 pieces / square centimeter, surface curvature > 70°).
[0038] Sentiment preference dimension: Crawling user reviews such as "strong retro feel" and "exquisite details" are converted into quantitative indicators (such as detail density of 35% and historical purchase conversion rate of 18%) through sentiment analysis.
[0039] Further, the process proceeds to step S103: Based on the three-dimensional feature matrix, a cross-modal association database is constructed through cross-modal structured annotation.
[0040] In one embodiment, based on the three-dimensional feature matrix, a cross-modal associated database is constructed through cross-modal structured annotation, including: constructing a cross-modal indexing system using knowledge graph technology. Specifically, this may include structurally associating the 2D images and 3D model data of each design case with their geometric features, style tags, and user sentiment ratings to construct a multimodal training dataset that can be recognized by AI models.
[0041] Specifically, in the construction of the cross-modal indexing system, knowledge graph technology is used to label data as triples <object, feature type, feature value>, such as <ring A, number of symmetry axes, 2>, <pencil B, style, modern minimalist>, <earring C, sentiment score, 4.5>. Secondly, cross-modal associations are established (such as the positive correlation between "high curvature surface" and "refinement score > 4 points", with a confidence level > 85%).
[0042] Furthermore, a cross-modal indexing system, combined with weakly supervised learning, enables structured annotation of quantified features. Structured annotation tools can be used to annotate quantified features. The visual annotation platform supports designers manually annotating difficult-to-quantify features (such as "visual center of gravity balance"), and automatically completes missing annotations using weakly supervised learning (based on a rule engine). Using annotation platform tools significantly improves annotation efficiency compared to purely manual annotation. Secondly, a domain expert verification mechanism is introduced, using a Kappa coefficient (>0.85) to ensure annotation consistency, and a high-quality cross-modal relational database (storage capacity ≥500GB, supporting second-level full-text search and multi-dimensional filtering) is constructed.
[0043] For example, in one embodiment, the 3D model (STL), 2D rendering (JPEG), and user reviews (such as "delicate carving") of a "retro style ring" are labeled as the same feature group, and the geometric feature "carving density 38%" is associated with the emotional score "refinement 4.2 points".
[0044] In terms of technical implementation, the database can adopt a hybrid architecture of graph database (such as Neo4j) and relational database (such as PostgreSQL), supporting efficient storage and related queries of geometric features (numerical), style tags (textual), and sentiment scores (vector).
[0045] The embodiments provided by the present invention described above can improve data integrity by 80% compared with traditional solutions by constructing a three-dimensional feature matrix covering three major dimensions and 100+ sub-indicators, thus realizing the transformation from "experience-based judgment" to "data definition".
[0046] Cross-modal association databases support multi-dimensional retrieval (such as filtering models by "golden ratio > 0.9 and user satisfaction rate > 90%)", which can improve design reference efficiency by 300%.
[0047] By using structured tools to improve data annotation efficiency, the data preparation cycle for model training can be shortened from 3 months to 2 weeks. During the design iteration process, the aesthetic compliance rate of AI-generated solutions has increased from 50% to 85%, and the time taken for a typical piece of jewelry from "parameter input to final model" has been reduced to 2-3 hours, which significantly improves design efficiency compared to traditional manual design.
[0048] Further, the process proceeds to step S104: using a cross-modal association database, a multimodal aesthetic feature database is generated through a multi-task deep learning model.
[0049] In one embodiment, generating a multimodal aesthetic feature database using a cross-modal association database and a multi-task deep learning model includes: generating multimodal aesthetic feature data containing quantitative aesthetic scores, defect heatmaps, and style matching degrees using a cross-modal association database and a multi-task deep learning model.
[0050] The structure of a multi-task deep learning model can adopt a hybrid architecture of Transformer + 3D CNN + Graph Neural Network (GNN), which may include 3D CNN branches, Transformer branches, and GNN branches (for inter-modal dependency modeling).
[0051] 3D CNN branch: Input a voxelized 3D model (resolution 128×128×128) and extract geometric features, such as symmetry and proportionality. In one implementation scenario, it is implemented as follows:
[0052] Symmetry score calculation (0-10 points): The formula is as follows:
[0053] In the above formula, D sym This represents the average Euclidean distance between mirror pairs of points along the model's axis of symmetry. Specifically, the axis of symmetry can be determined using Principal Component Analysis (PCA), and the coordinates of the mirror pairs of points can be extracted (p...). i ,p i '),calculate
[0054] D max This indicates the maximum permissible symmetrical distance for jewelry of the same category, such as ring category D. max =1.5mm, Pendant Category D max =2.0mm, which can be statistically analyzed based on historical high-quality designs.
[0055] Proportional consistency score (0-10 points), formula:
[0056] In the above formula, parameter φ represents the actual proportional parameters (such as the height of the main stone / width of the ring, the length of the bracket / total height of the pendant). Parameter φ gold This represents the target value for the golden ratio (0.618 or its reciprocal 1.618, automatically matched based on the design elements). φ thres This represents the proportion tolerance threshold (the standard deviation of the proportion of the top 20% of high-quality samples in the same design, which can be dynamically adjusted).
[0057] The Transformer branch processes design specification text and user comments to generate style semantic vectors. For example, the semantic embedding vector for "retro style". In an example scenario of style semantic vector generation, the specific implementation is as follows:
[0058] The technical implementation of style keyword encoding can use a BERT-based pre-trained model, with the input design specification text T = {t1, t2, ..., t...} n}, output character-level embedding h i Global semantic vectors s are generated through pooling layers. text =CLS(h1,h2,...,h n This allows us to construct a style feature library S = {s1, s2, ..., s}. 30 (30-dimensional style tags, such as "retro" corresponding to a combination of features like carving density and surface curvature). Then, the matching degree between the text and the style library is calculated using cosine similarity.
[0059] The formula for weighting user comments based on sentiment can be: (Comment sentiment dimension k). Wherein, parameter w k The weights represent the emotional dimensions (trained using historical purchase data, such as the weight of "refined" which automatically increases with repurchase rate). LSTM(·) represents the LSTM encoding of the five emotional dimensions, including "refined" and "uniqueness," and outputs a 10-dimensional emotional vector.
[0060] GNN branch: Modeling cross-modal feature associations. For example, the weight of the influence of "gold material + complex carving" on the difficulty of the craftsmanship, and the weight association between "high curvature surface" and "exquisite feel".
[0061] In the definition of Graph Attention (GAT), nodes represent geometric features (curvature c, number of elements n), style tags, and process parameters (print layer thickness l, number of tessellation claws m), expressed by the formula v. i =[g i ,s i ,p i ].
[0062] First, the node features are initialized. This is done by inputting the cross-modal node feature matrix V = {v1, v2, ..., v...}. N Node types (in heterogeneous graph scenarios) can be geometric feature nodes (e.g., "high curvature surface," "symmetrical structure"), style semantic nodes (e.g., "retro style," "gold material," "complex carving"), or user preference nodes (e.g., "refined look," "craftsmanship difficulty" ratings). Feature dimensions: F represents the original feature dimension (such as the curvature value of geometric features or the semantic vector of text embedding).
[0063] Secondly, cross-modal feature linear transformation is performed. This is done by sharing the weight matrix. Mapping features from different modalities to a unified space: h i =Wv i , Here, parameter W represents the cross-modal feature alignment matrix, enabling geometric, semantic, and preference features to interact (e.g., embedding the text "gold material" and mapping the geometric feature "metallic luster" to the same space). F' represents the transformed feature dimension (usually set to 256 or 512, adjusted according to model complexity).
[0064] Then, the attention coefficient is calculated. For node i's neighboring nodes j∈N... i Calculate the cross-modal association weight α ij The formula is:
[0065] The parameters obtained in the above formula This represents the attention weight vector, which captures the relevance of node pairs (i,j) (e.g., the correlation strength between "complex carving" and "craft difficulty"). This represents the feature concatenation operation (fusing cross-modal features of nodes i and j). N i This represents the set of neighbors of node i (e.g., the neighbors of the node "gold material" include related nodes such as "complex carving" and "high craftsmanship difficulty").
[0066] The following is a scenario example: if i is the "complex carving" node and j is the "craftsmanship difficulty" node, α ij This indicates the weight of the influence of the carving complexity on the difficulty of the craftsmanship.
[0067] Finally, cross-modal feature aggregation aggregates neighbor features through attention weights to generate a new feature h for node i. i ′: In the above formula, the parameter σ represents the activation function (such as ReLU, which enhances the nonlinear expression). Its function is to aggregate the features of its neighbors "highly reflective geometric surfaces" and "retro carvings" for the "gold material" node, and output a comprehensive feature that integrates material and process.
[0068] Then, the multi-head attention mechanism enhances robustness by using K independent attention heads, and the results are spliced or averaged: Where K represents the number of attention heads (usually K=8), different heads can capture different types of associations (e.g. head 1 focuses on "material-craftsmanship" associations, head 2 focuses on "geometry-style" associations).
[0069] Edges represent the correlation between features (e.g., a positive correlation edge between "high curvature c>0.1" and "refinement score>4", with weights initialized to historical co-occurrence frequencies).
[0070] The following example further illustrates the innovative aspects of GAT in the context of jewelry design in this invention.
[0071] In heterogeneous graph modeling of cross-modal associations, node definitions include:
[0072] Geometric modal nodes: curvature c i Symmetry index s i Volume ratio v i (Quantify the geometric features of the 3D model).
[0073] Semantic modal node: Material embedding m j (e.g., One-Hot encoding or Word2Vec vectors for "gold" and "diamond"), style tags j (e.g., text embedding of words like "retro" or "minimalist").
[0074] Preference modality node: User rating r k (Such as "refinement" and "uniqueness" ratings, normalized to [0,1]).
[0075] Edge definitions include:
[0076] Geometric-semantic edges: such as the positive correlation edge between "high curvature surface" and "refinement" (weight +0.8).
[0077] Semantic-preference edges: such as the association edge between "retro style" and "user age 30+" (weights are obtained from historical data statistics).
[0078] In the constraint optimization of attention weights, for the process constraints in jewelry design (such as the feasibility of matching "complex carvings" with "gold material"), a prior knowledge regularization term is introduced into the attention coefficient: (c i,j (Compliance) This indicates an indicator function; if node pair (i,j) satisfies process constraints (e.g., compliance with "diamond setting" and "metal prong setting structure"), then... Otherwise, it is 0. λ represents the constraint strength coefficient (empirical value set to 0.5), which enhances the weight of compliance association.
[0079] Then, in the feature fusion with 3D CNN and Transformer, the cross-modal correlation features h output by GAT are... i ′ and geometric features g extracted by 3D CNN k The semantic vectors s generated by Transformer m splicing:
[0080] The purpose of the above fusion is to provide multi-task models with comprehensive features that integrate geometric structure, semantic description, and cross-modal association, thereby improving the accuracy of quantitative scoring (such as symmetry and style purity).
[0081] Through the embodiments provided by the present invention, the cross-modal correlation features output by GAT will be used as input to a multi-task deep learning model to support core functions such as quantitative aesthetic scoring, defect localization, and style matching degree calculation, thereby improving the intelligence level of jewelry 3D design.
[0082] Multi-task training objectives can be categorized as follows.
[0083] Regression task: Output three sub-dimension quantitative scores (symmetry, proportionality, style purity, range 0-10, mean squared error MSE<0.1).
[0084] Regression task loss (symmetry / proportion / style purity) Wherein, parameter λ i This indicates the task weight (dynamically adjusted; for example, the weight of style purity is set to 0.4 in brand design and 0.2 in personalized customization). This indicates the model output score. This indicates that the experts have labeled the true values.
[0085] Location task: Generate a defect heatmap (based on U-Net network, locate visual imbalance areas with pixel-level accuracy IOU>0.9, such as areas with a center of gravity offset>5% marked in red).
[0086] Formula L for location task loss (defect heatmap generation) seg = α·DiceLoss + (1-α)·CrossEntropyLoss. This formula combines Dice loss (to address class imbalance) and cross-entropy loss to improve the detection accuracy of small defect areas (such as centroid offset points). α represents the balance coefficient (dynamically adjusted according to the proportion of defect areas, default 0.7).
[0087] Classification task: Calculate style matching degree (classification accuracy of 10 basic styles > 95%, output cosine similarity vector, threshold ≥ 0.85 is used to determine style compliance).
[0088] The formula for the classification task loss (style matching degree) function can be L. cls = -log(Softmax(StyleSim) global StyleSim global =0.6·StyleSim text +0.4·StyleSim 3D (Style matching between integrated text and 3D model).
[0089] Additionally, a joint loss function can be included, with the formula L. total =L reg +β·L seg +γ·Lcls +λ adv ·L adv Adversarial loss L is introduced through Domain Adversarial Transfer (DATL). adv By using a gradient reversal layer to obfuscate the distribution of labeled and unlabeled data, the generalization ability of niche styles can be improved.
[0090] In one embodiment, generating a multimodal aesthetic feature database using a cross-modal association database and a multi-task deep learning model further includes: introducing domain adversarial transfer learning through domain enhancement techniques to increase the number of scarce style samples in the multimodal aesthetic feature database.
[0091] In the process of introducing Domain Adversarial Transfer Learning (DATL), we can use rare style samples (such as Art Nouveau and Bohemian style) labeled by 5,000+ professional designers to improve the model's evaluation accuracy for niche styles (reducing the scoring error by 40%).
[0092] In the specific technical implementation process, model training can adopt the PyTorch framework, and the distributed training cluster (8 GPUs) supports batch processing. The single sample feature generation time is <200ms, generating a multimodal aesthetic feature database containing 200,000+ samples.
[0093] The following example further illustrates multimodal data-driven intelligent design and optimization, which is achieved by fusing multimodal data for data-driven processing.
[0094] First, the initial model is generated. After inputting user parameters, CGAN retrieves similar style samples from a cross-modal database (such as the Top 100 cases of the "modern minimalist" style) and extracts average geometric features (bracket width 2.5mm, main stone ratio 1.5:1) as the generation prior to ensure that the initial model conforms to historical high-quality design paradigms.
[0095] Secondly, dynamic weight evaluation is performed. The evaluation engine calls the multimodal aesthetic feature database in real time. For example, when calculating "style purity", it not only analyzes geometric elements, but also associates with the user's historical preference data (such as if "minimalist style" accounts for 60% of the user's past purchase records, then the weight of this dimension will be automatically increased).
[0096] Finally, human-machine co-evolution is implemented. The 3D feature matrix is synchronously updated based on user decision-making actions (such as adjusting the position of the main stone). For example, after a certain eccentric adjustment, the system records the correlation feature of "asymmetric design + high sentiment score," feeding back into model training and improving the ability to generate asymmetric styles. This, in turn, increases the adoption rate of this type of design in subsequent schemes.
[0097] Through the embodiments provided by the present invention, the consistency between quantitative aesthetic scoring and expert scoring reaches 92%, the defect heatmap positioning error is <0.1mm, and the style matching degree detection time is <50ms / time, effectively improving key indicators compared to the single-modal model. It supports cross-validation of "material-structure-style-preference" (e.g., automatically identifying the process risks of the "platinum material + complex carving" combination and prompting 3D printing support structure design), improving the accuracy of solution feasibility assessment.
[0098] Furthermore, by accumulating over 200,000 design features in a multimodal aesthetic feature database, reusable "digital design assets" can be formed, thereby shortening the training cycle for new designers. For example, correlation analysis revealed that the proportion of the combination "asymmetrical design + micro-inlay technique" increased from 3% to 15% in historical data, indicating an emerging customer preference. Therefore, this can help companies capture market trends and shorten new product development cycles.
[0099] Further, the process moves to step S105: using a multimodal aesthetic feature database, train a multi-task AI aesthetic evaluation model that integrates geometry, style, and user preferences;
[0100] In one embodiment, a multi-task AI aesthetic evaluation model may include: evaluating and generating a three-dimensional quantitative aesthetic score, a defect location heatmap, and a style matching report based on voxelized data of a 3D design model, 2D multi-view images, user-input style keywords, and design parameters.
[0101] The following examples describe in detail the process of implementing the multi-task AI aesthetic evaluation model technology.
[0102] I. Model Input and Feature Fusion Process
[0103] Input layer design: The model receives multimodal input data, including:
[0104] 3D voxel data: Voxelized data of a 3D model with a resolution of 128×128×128, denoted as... Used to extract geometric features.
[0105] 2D multi-view image: three images, namely the front view, side view, and top view, denoted as I. front ,I side , (H and W represent the image height and width, and 3 represents the RGB channel).
[0106] User input data: Style keyword text sequence T = {t1, t2, ..., t} n} and the design parameter vector P = [p1, p2, ..., p m ], where p iIt indicates parameters such as material and gemstone specifications.
[0107] The feature fusion formula can be: fusing features from different modalities into a unified feature vector F. fusion :
[0108] F fusion =Concat(F 3D ,F 2D ,F text ,F param )
[0109] Wherein: F 3D The geometric feature vector extracted by 3D CNN, with dimension d. 3D ;F 2D : 2D image feature vectors extracted by ResNet, with dimension d 2D ;F text : BERT-encoded style keyword semantic vectors, with dimension d text ;F param : Embedded vector of design parameters, with dimension d param Concat is a concatenation operation.
[0110] This process has the following technical effects:
[0111] Multimodal information complementarity: By integrating 3D geometric data (voxels), 2D visual features (multi-view images), semantic text (style keywords), and parametric design requirements, it covers all dimensions of design information and can solve the one-sidedness of single-modal evaluation (such as the inability to capture the semantic connotation of "retro style" by relying solely on geometric data).
[0112] Feature space alignment: By using cross-modal splicing and attention mechanisms, “material text description” and “metal reflective geometric features” are mapped to a unified semantic space, which can improve feature interaction capabilities (e.g., the correlation strength between “platinum” text embedding and “low reflectivity surface” feature is increased by 40%).
[0113] Enhanced input flexibility: Supports parametric input, sketch recognition, and import of existing models, compatible with different design habits, which can lower the barrier to entry for users (input efficiency can be improved by 50% for non-professional users).
[0114] II. Three-Dimensional Quantitative Aesthetic Scoring Calculation Process
[0115] 1. Symmetry score (S) sym , 0-10 points) formula:
[0116]
[0117] Where N: the number of corresponding point pairs on both sides of the symmetry axis of the 3D model; d id: Euclidean distance between the i-th pair of points; max The maximum permissible symmetry distance threshold for similar jewelry designs is determined based on statistical data of historical high-quality designs, such as rings. max =1.2mm.
[0118] 2. Proportional consistency score (S) ratio (0-10 points) The formula is:
[0119] Where a and b represent the dimensions of key geometric elements in the 3D model (such as the length of the main stone and the width of the ring band); The golden ratio is 0.618; δ: the proportional tolerance coefficient, which is dynamically adjusted according to the design type and determined by the proportional standard deviation of similar designs.
[0120] 3. Style Purity Rating (S) style (0-10 points) The formula is:
[0121] In the above formula, k: the number of style features; w i The weight of the i-th style feature is determined based on domain knowledge and data statistics; f i : The i-th style feature vector extracted from the design; s i : The i-th standard feature vector of the target style; Sim(·): Cosine similarity function, used to calculate the similarity between the design features and the standard style features.
[0122] This process has the following technical effects:
[0123] Symmetry scoring: Symmetry is quantified by the distance between geometric points, reducing the standard deviation of the score from ±1.8 in manual assessment to ±0.35, eliminating subjective differences among designers (e.g., the consistency of the assessment of a "biaxially symmetrical ring" can be increased from 65% to 92%). It automatically identifies issues such as center of gravity shift and structural imbalance, with a positioning error of <0.1mm, improving efficiency by 80% compared to traditional CAD measurement.
[0124] Proportional Harmony Scoring: Based on the golden ratio and dynamic tolerance thresholds, it ensures that the design conforms to aesthetic principles (e.g., the compliance rate of the main stone and ring size ratio can be increased from 8% to 91%). It supports custom proportion parameters (e.g., automatically adjusting the tolerance when users prefer "exaggerated proportions"), balancing standardization and personalization needs.
[0125] Style purity score: Based on style feature weights and cosine similarity, it quantifies the fit between design elements and the target style (e.g., the accuracy of compliance detection for "minimalist" elements can be increased from 70% to 95%). It automatically filters out elements with conflicting styles (e.g., the detection rate of redundant carvings in "minimalist design" can reach 98%), reducing manual screening time.
[0126] III. Defect Location Heatmap Generation: The U-Net network architecture is used for defect location, and an attention mechanism is introduced to enhance the detection of key areas.
[0127] The loss function formula is: L seg =α×DiceLoss+(1-α)×FocalLoss.
[0128] In the above formula, α is a weighting coefficient used to balance the two losses, with a dynamic value range of [0.3, 0.7].
[0129] DiceLoss: The dice loss function, used to handle class imbalance problems. The formula is:
[0130] Where p(x,y) is the pixel probability predicted by the model, and g(x,y) is the true label pixel value.
[0131] FocalLoss: The focal loss function reduces the weight of simpler samples. The formula is:
[0132] FocalLoss = -(1 - p(x,y)) γ ×log(p(x,y)), where γ is the focusing parameter, empirically taken as 2.
[0133] This process has the following technical effects:
[0134] Pixel-level precision inspection: Combining DiceLoss and FocalLoss, small defects (e.g., 0.2mm) can be detected. 2 The detection rate of curvature abnormal regions can be increased from 60% to 89%, and the IOU (Intersection over Union) reaches 0.92, which is 35% higher than that of traditional edge detection algorithms.
[0135] Real-time visual feedback: Defective areas are highlighted in red and improvement suggestions are overlaid (such as "center of gravity is 0.3mm to the left, it is recommended to move the secondary stone to the right"). The time for designers to locate problems can be reduced from 30 minutes / model to 5 minutes / model.
[0136] Process risk prediction: By linking geometric features with process constraints (such as the thermal region of fracture risk corresponding to "narrow neck structure"), manufacturability issues can be predicted in advance, reducing the rework rate on the production side by 65%.
[0137] IV. The process of calculating style matching degree is as follows:
[0138] Style feature extraction: The BERT model is used to encode style keywords to obtain the style semantic vector S. target Simultaneously, style feature vector S is extracted from the design. design .
[0139] Cosine similarity calculation: Among them, S target ·S design The dot product of two vectors; ||S target ‖、‖S design ‖: The L2 norms of the two vectors, respectively.
[0140] Feature difference analysis: Calculate the differences between design features and target style features across various dimensions. The formula can be: ΔF i =|f i,design -f i,target |, where F i f represents the i-th style feature dimension. i,design and f i,target These are the feature values of the design and target styles in this dimension, respectively.
[0141] This process has the following technical effects:
[0142] Semantic style quantization: Through BERT encoding and cosine similarity, numerical matching of styles such as "retro" and "cyberpunk" is achieved (the accuracy of classification of 10 basic styles can reach 96.2%, and the recognition rate of niche styles can be improved by 28%).
[0143] Feature differences can be explained: output difference values in various dimensions (such as "carving density is insufficient -15%" and "curved surface curvature deviation +8°") to guide designers to make precise adjustments (the number of style matching iterations can be reduced from 5 times / model to 2 times / model).
[0144] Brand compliance management: A predefined brand style feature library (e.g., "Brand A requires symmetry ≥ 8 points") automatically blocks non-compliant designs, and the brand standard compliance rate can be increased to 100%.
[0145] V. The data augmentation and contrastive learning process is as follows.
[0146] 1. Data augmentation includes:
[0147] 3D model enhancement: rotation (angle range [-15°, 15°]), scaling (scale range [0.9, 1.1]), and local deformation (randomly changing the position of some voxels).
[0148] 2D image enhancement: horizontal flip, vertical flip, add Gaussian noise (mean 0, standard deviation 0.05).
[0149] 2. Comparative learning includes:
[0150] Constructing the contrastive learning loss function L contrast :
[0151]
[0152] In the above formula, F i : Anchor point sample feature vector; Feature vectors of positive samples (similar high-scoring designs); Negative sample (different class or low-scoring design) feature vector; K: number of negative samples; τ: temperature parameter, controlling the tightness of the feature distribution, empirically taken as 0.1.
[0153] This process has the following technical effects:
[0154] Data diversity expansion: 3D geometry augmentation and 2D image augmentation generate 500,000+ samples, which can improve the model's generalization ability (32% reduction in evaluation error under rare angle views).
[0155] Feature space clustering optimization: Contrastive learning reduces the intra-class distance of features in the "high-scoring design" by 30% and increases the inter-class distance of features in the "low-scoring design" by 45%, thereby improving the model's discriminative power (the accuracy of positive and negative sample classification can be increased from 82% to 94%).
[0156] Overfit suppression: By mining hard samples and comparing them with negative samples, the risk of overfitting the model on small datasets (such as only 1000 labeled samples) can be reduced by 50%.
[0157] VI. Transfer Learning Strategies
[0158] 1. Pre-training phase: The model is pre-trained on the ImageNet (image classification) and ShapeNet (3D shape classification) datasets to learn general visual and geometric feature representations.
[0159] 2. Fine-tuning phase: Fine-tuning is performed on jewelry-related data using an adaptive learning rate adjustment strategy.
[0160] In the above formula, lr0: initial learning rate; t: current training round number; T: total training round number; β: decay coefficient, empirically taken as 0.9.
[0161] This process has the following technical effects:
[0162] Training efficiency leap: Based on ImageNet / ShapeNet pre-trained model initialization, the training cycle in the jewelry field can be shortened from 8 weeks to 4 weeks, and the consumption of computing resources is reduced by 50%.
[0163] Enhanced generalization ability: F1-score > 0.92 on industry-standard test sets (such as Jewelry Design Benchmark), an improvement of 18% compared to training from scratch.
[0164] Cross-domain knowledge reuse: The transfer of general visual features (such as edge detection and shape recognition) and geometric priors (such as symmetry rules) can improve the model's adaptability to novel designs (such as irregular gemstone inlay) by 40%.
[0165] The aforementioned multi-task AI aesthetic evaluation model, through the comprehensiveness of multimodal feature fusion, the objectivity of quantitative scoring, the accuracy of defect detection, the intelligence of style matching, and the efficiency of data utilization, constructs a digital evaluation system for jewelry design aesthetics. Based on multiple experiments, it can achieve:
[0166] Evaluation efficiency: Single model evaluation time is less than 200ms, which is more than 99% more efficient than manual evaluation.
[0167] Design quality: The aesthetic score pass rate increased from 50% to 85%, and the proportion of designs with a style matching degree of ≥0.85 increased by 60%.
[0168] Innovation empowers: It supports users with "zero experience" to complete professional-level aesthetic designs, and can increase the output efficiency of solutions for non-designer users by 300%;
[0169] Knowledge Accumulation: Accumulated evaluation data feeds back into the model, forming a closed loop of "design-evaluation-optimization", which can shorten the adaptation cycle of emerging styles from 3 months to real-time response.
[0170] Therefore, through the above formulas, algorithms, and technologies, the multi-task AI aesthetic evaluation model can effectively integrate multiple factors such as geometry, style, and user preferences to achieve high-precision aesthetic scoring, defect localization, and style matching, significantly improving the intelligence level and design quality of 3D jewelry design.
[0171] Based on the above description of the multi-task AI aesthetic evaluation model, those skilled in the art can understand that by inputting voxelized data of a 3D model (128×128×128), 2D multi-view images (front view / side view / top view), user-input style keywords and design parameters, a three-dimensional quantitative score (symmetry, proportional harmony, style purity), a defect location heatmap (marking areas with abnormal curvature, center of gravity shift, etc.), and a style matching report (cosine similarity with the target style, and feature difference analysis) can be output.
[0172] For example, geometric transformations (rotation, scaling, local deformation) can be applied to 3D models, and data augmentation (flipping, adding noise) can be performed on 2D images to generate over 500,000 augmented samples, thereby improving the model's generalization ability. A contrastive learning mechanism is introduced to force the model to distinguish the feature differences between "high-scoring designs" and "low-scoring designs," optimizing the feature space clustering effect (reducing intra-class distance by 30%).
[0173] In technical implementation, a transfer learning strategy can be adopted. The pre-trained model is initialized on the ImageNet and ShapeNet datasets and then fine-tuned for jewelry data. This can shorten the training cycle by 50% and make the evaluation model achieve an F1-score of >0.92 on the industry standard test set.
[0174] Further, the process proceeds to step S106: acquire 3D design data of jewelry through a graphical interface, and use a multi-task AI evaluation model to conduct a quantitative evaluation of the 3D design in order to generate design evaluation results.
[0175] In this embodiment, the model can be generated through user-parameterized input (such as style, material, gemstone specifications, and wearing size), sketch recognition (such as parsing hand-drawn sketches based on the YOLO model and converting them into geometric constraints), and importing existing models (such as being compatible with mainstream CAD formats). Then, using a Conditional Generative Adversarial Network (CGAN), combined with user parameters and high-quality design priors from a cross-modal database (such as the average geometric parameters of historically high-scoring models of the same style), an STL format base model is generated (which can achieve a generation time of <2 minutes and an accuracy of 0.01mm).
[0176] Furthermore, in one embodiment, the design method provided by the present invention may further include: based on the design evaluation results, combined with user preferences and process constraints (such as the minimum wall thickness of 0.8 mm for 3D printing and the minimum strength requirement of the inlay claw), generating multiple versions of aesthetic enhancement schemes for the 3D design through a conditional generative adversarial network.
[0177] In the specific implementation process, the conditional input layer design can be carried out first. The input vector includes: basic parameters (material, gemstone specifications, size of the wearing part); current defect vector (e.g., [symmetry 4.2, proportional harmony = 5.5]); user preference vector (personalized weights trained through historical interaction data, such as "preference for asymmetrical design" weight > 0.7); and process constraint vector (e.g., "minimum number of prongs = 4", "minimum thickness of precious metal = 0.5mm").
[0178] Secondly, a geometric optimization branch (3D-GAN) is implemented using a dual-branch GAN architecture: This branch addresses structural defects (such as center of gravity shift or curvature abrupt changes) by generating geometric parameter adjustment schemes (e.g., modifying the coordinates of the secondary stone layout or adjusting the bracket curvature parameters), and outputs a local modification file in STL format (supporting Boolean operations with the original model). A style optimization branch (image GAN) addresses style conflicts (e.g., a "minimalist style" containing too many decorative elements) by matching and replacing parts from a pre-defined element library (1000+ parametric modules, such as "geometric blocks" and "smooth metal strips"), generating a 2D rendering for user preview.
[0179] Finally, through multi-scheme screening and visualization, 5-8 candidate schemes are generated each time. The Pareto optimal solution is then selected among three objectives: "aesthetic improvement", "process feasibility", and "modification complexity" using a non-dominated sorting genetic algorithm (NSGA-II). Each scheme includes a visual difference analysis (such as heatmap comparisons of visual focus distribution and material reflectivity differences between the original and optimized schemes). Users can click to synchronize the results to the 3D modeling interface with a single click, supporting real-time preview of the modified effects.
[0180] In the technical implementation process, the geometry optimization branch adopts Progressive Growing 3D-GAN to improve the generation quality of complex structures. The style optimization branch can be based on StyleGAN2, supporting consistent lighting and shadow rendering after element replacement. According to tests, the solution generation time is less than 30 seconds per round.
[0181] Furthermore, by using reinforcement learning to record user decisions, the design knowledge base and AI aesthetic evaluation model are dynamically optimized. This allows the recording of user decision-making behavior through reinforcement learning to dynamically optimize the AI evaluation model and feed back into the design knowledge base, forming a self-evolving closed loop.
[0182] In recording user decisions, user preference modeling is required. First, the user interaction trajectory (parameter adjustment range, solution adoption / rejection history, weight modification records) is recorded through a decision tracking system, generating a preference trajectory matrix T = [t_1, t_2, ..., t_n], where t_i is the weight adjustment vector of the i-th interaction.
[0183] Secondly, a deep Q-network (DQN) is used to model user preferences. When the number of similar decisions is ≥50, the model is fine-tuned (e.g., if the user continues to prefer low symmetry design, the system automatically reduces the "symmetry" weight threshold by 20% and increases the probability of generating asymmetric elements to 60%).
[0184] Furthermore, a dynamic design knowledge base iteration is implemented. First, parametric modules are accumulated, with frequently adopted schemes (adoption rate > 60%) broken down into reusable parametric design modules (such as an "eccentric elliptical main stone + single-row micro-pave" structure), stored in a component library, and supported for one-click access (reducing repetitive design time by 40%). Second, federated learning updates are performed. Anonymized feature vectors from multiple users are periodically aggregated (only hashed feature summaries are uploaded), and the global evaluation model is updated through federated learning. This protects user privacy while improving the model's adaptability to emerging trends (the accuracy of niche style evaluation increases by 10%-15% after every 10,000 interactions).
[0185] During the technical implementation process, the preference learning module adopts an online learning mechanism, generating an incremental model update package every 100 interactions. The knowledge base management system supports version control, records the module's evolution history, and ensures design traceability.
[0186] In summary, the embodiments provided by this invention, through the deep integration of full-modal data acquisition, cross-modal correlation modeling, and multi-task intelligent generation, construct a closed loop of "data-model-application" for jewelry design. By transforming jewelry design from "experience-driven" to "data-driven," it solves the core problems of subjective evaluation, data fragmentation, and inefficient iteration in traditional design, providing the industry with a new intelligent design paradigm that is quantifiable, traceable, and evolvable.
[0187] Figure 2 An exemplary structural block diagram of a jewelry 3D model design system 200 according to other embodiments of the present invention is shown.
[0188] like Figure 2 As shown, system 200 includes: a data acquisition module 201 for acquiring historical jewelry design data across all modalities; a feature matrix generation module 202 for using computer vision and natural language processing technologies to extract and annotate multi-dimensional features from the data across all modalities to construct a three-dimensional feature matrix; a relational database generation module 203 for constructing a cross-modal relational database based on the three-dimensional feature matrix through cross-modal structured annotation; an aesthetic feature database generation module 204 for generating a multi-modal aesthetic feature database using the cross-modal relational database and a multi-task deep learning model; an AI aesthetic evaluation model training module 205 for training a multi-task AI aesthetic evaluation model that integrates geometry, style, and user preferences using the multi-modal aesthetic feature database; and a design evaluation module 206 for acquiring 3D jewelry design data through a graphical interface, using a multi-task AI evaluation model to perform quantitative evaluation of the 3D design, and generating design evaluation results.
[0189] Through the jewelry 3D model design system provided in the above embodiments of the present invention, it can be understood that system 200 is a specific implementation of the aforementioned method 100 embodiment. Therefore, the features of the embodiments described in method 100 above can be similarly applied here.
[0190] While numerous embodiments of the invention have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Many modifications, alterations, and alternatives will occur to those skilled in the art without departing from the spirit and essence of the invention. It should be understood that various alternatives to the embodiments of the invention described herein may be employed in the practice of the invention. The appended claims are intended to define the scope of protection of the invention and therefore cover equivalents or alternatives within the scope of these claims.
Claims
1. A jewelry 3D model designing method, characterized by, The method comprises the following steps: S1, acquiring historical jewelry design full-modal data; S2, using computer vision and natural language processing technology to perform multi-dimensional feature extraction and correlation labeling on the full-modal data to construct a three-dimensional feature matrix; S3, according to the three-dimensional feature matrix, constructing a cross-modal correlation database through cross-modal structured labeling; S4, using the cross-modal correlation database, generating a multi-modal aesthetic feature database through a multi-task deep learning model; The method of generating a multi-modal aesthetic feature database through a multi-task deep learning model using a cross-modal correlation database comprises: generating multi-modal aesthetic feature data including quantitative aesthetic scores, defect heat maps, and style matching degrees through a multi-task deep learning model using a cross-modal correlation database; the structure of the multi-task deep learning model adopts a hybrid architecture of Transformer, 3DCNN, and graph neural network; specifically as follows: In the definition of the graph structure attention mechanism, the node represents the geometric feature, the style label, and the process parameter, and the formula is represented as ; First, the node features are initialized by inputting the cross-modal node feature matrix ; feature dimension: , F is the original feature dimension; Secondly, cross-modal feature linear transformation is performed, and different modal features are mapped to a unified space through a shared weight matrix will be mapped to a unified space: , F' represents the feature dimension after transformation; Then, the attention coefficient is calculated, and the neighbor nodes of the node i , the cross-modal correlation weight is calculated , the formula is: , ; The parameters in the above formula denotes the attention weight vector, capturing the relevance of node pairs . denotes the feature concatenation operation, denotes the neighbor set of node i; i is the "complex carving" node, j is the "process difficulty" node, represents the influence weight of carving complexity on process difficulty; Finally, the cross-modal feature aggregation aggregates neighbor features through attention weights to generate new features of node i : , in the above formula, the parameter represents the activation function, and the output is the comprehensive feature of the fusion material and process; the multi-head attention mechanism enhances the robustness through K independent attention heads, and the results are spliced or averaged: , where K represents the number of attention heads, and different heads capture different types of associations; Introducing priori knowledge regular term in attention coefficient: ; wherein, represents an indicator function, if the node pair satisfies the process constraint, then , otherwise 0; represents the constraint strength coefficient; S5, using the multi-modal aesthetic feature database to train a multi-task AI aesthetic evaluation model that integrates geometry, style, and user preferences; S6, acquiring jewelry 3D design data through a graphical interface, using the multi-task AI evaluation model to perform quantitative evaluation of the 3D design, and generating design evaluation results.
2. The method of claim 1, wherein, The method of acquiring historical jewelry design full-modal data comprises: Acquiring multi-modal data covering the entire life cycle of jewelry design through a multi-source data collection platform; wherein, The multi-source data collection platform comprises a design end, a user end, and a production end; The full-modal data includes 2D rendering images, 3D model files, designer scores, and user purchase feedback.
3. The method of claim 1, wherein, The three-dimensional feature matrix includes geometry feature dimensions, style semantic dimensions, and emotional preference dimensions.
4. The method of claim 1, wherein, According to the three-dimensional feature matrix, constructing a cross-modal correlation database through cross-modal structured labeling comprises: Using knowledge graph technology to construct a cross-modal index system, Through the cross-modal index system, combined with weak supervision learning, performing structured labeling of quantitative features.
5. The method of claim 4, wherein, The method of constructing a cross-modal index system using knowledge graph technology comprises: structurally associating the 2D images and 3D model data of each design case with its geometric features, style labels, and user emotional scores to construct a multi-modal training data set that can be recognized by an AI model.
6. The method of claim 1, wherein, The method of generating a multi-modal aesthetic feature database through a multi-task deep learning model using a cross-modal correlation database further comprises: Using domain adversarial transfer learning through domain enhancement technology to increase the rare style samples in the multi-modal aesthetic feature database.
7. The method of claim 1, wherein, The multi-task AI aesthetic evaluation model comprises: According to the voxelized data of the 3D design model, the 2D multi-view images, the user input style keywords and design parameters, generating three-dimensional quantitative aesthetic scores, defect positioning heat maps, and style matching degree reports.
8. The method of claim 1, wherein, The design method further comprises: According to the design evaluation results, combining user preferences and process constraints, generating multi-version aesthetic enhancement schemes for 3D designs through conditional generative adversarial networks; Through reinforcement learning to record user decisions, dynamically optimizing the design knowledge base and the AI aesthetic evaluation model.
9. A jewelry 3D model designing system, which implements the method as claimed in claim 1, characterized in that, The method comprises the following steps: The data acquisition module is configured to acquire historical jewelry design full-modal data. The feature matrix generation module is configured to perform multi-dimensional feature extraction and correlation labeling on the full-modal data by using computer vision and natural language processing technology, so as to construct a three-dimensional feature matrix. The correlation database generation module is configured to construct a cross-modal correlation database by cross-modal structured labeling according to the three-dimensional feature matrix. The aesthetic feature database generation module is configured to generate a multi-modal aesthetic feature database by a multi-task deep learning model using the cross-modal correlation database. The AI aesthetic evaluation model training module is configured to train a multi-task AI aesthetic evaluation model that fuses geometry, style and user preferences by using the multi-modal aesthetic feature database. The design evaluation module is configured to acquire jewelry 3D design data through a graphical interface, perform quantitative evaluation of the 3D design by using the multi-task AI evaluation model, and generate a design evaluation result.
Citation Information
Patent Citations
Virtual twin product aesthetic and cultural style evaluation system and method based on big data
CN118246990A
Ai-based product recommendation method, apparatus, and system for jewelry products rendered using 3D drawings
KR102811342B1