Jewelry 3D model design method and system
By building a cross-modal association database and a multi-task AI aesthetic evaluation model, the problem of inconsistent aesthetic evaluation in traditional jewelry design is solved, efficient, accurate and personalized quantitative evaluation of jewelry design is achieved, and design quality and efficiency are improved.
Patent Information
- Application Number
- CN202510893717.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2045-06-30
AI Technical Summary
Traditional jewelry design relies on manual evaluation and inconsistent aesthetic standards, resulting in high rework rates, high design costs and long cycles. It is unable to accurately predict consumer preferences and affects brand image.
Through multi-source data collection, computer vision and natural language processing technology, we extract the multi-dimensional features of jewelry design, build a cross-modal correlation database, use a multi-task deep learning model to generate a multi-modal aesthetic feature database, train a multi-task AI aesthetic evaluation model, and conduct quantitative evaluation of 3D design.
It achieves objective quantitative evaluation of jewelry design, improves design efficiency and quality, reduces rework rate, improves design efficiency and satisfaction of designers and non-designer users, and enhances the objectivity and accuracy of aesthetic evaluation.
Smart Images

Figure CN120763533A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application generally relates to the field of artificial intelligence. More specifically, the present application relates to a jewelry 3D model design method and system. BACKGROUND
[0002] Jewelry design is a complex process that combines aesthetic creativity, technical craftsmanship, and user needs. The core goal is to achieve the unity of visual aesthetics and wearing function through the combination of form, material, and color. In the traditional design process, designers need to complete the whole link from concept conception, sketch drawing to 3D modeling. Aesthetic evaluation is the key link to determine the success or failure of the design, including the judgment of symmetry, proportion coordination, style fit, and other core dimensions. However, these evaluations have long relied on the subjective experience and aesthetic preferences of designers, lacking quantifiable objective standards, resulting in a high dependence on individual ability for design quality. Although modern CAD tools (such as Rhino, Matrix) have achieved parameterization of geometric modeling, the "good or bad judgment" in the aesthetic level still remains in the stage of artificial review.
[0003] With the upgrading of consumption, the jewelry market has shifted from standardized mass production to personalized customization, and users' aesthetic requirements for design have shown a trend of diversification (such as minimalist, national style, and retro deconstructionism). The defects of traditional manual evaluation mode are further magnified. High rework rate caused by non-uniform aesthetic evaluation standards leads to high design cost and prolonged cycle. In addition, the evaluation relying on manual experience cannot accurately predict consumer preferences, which also causes high return rate after product launch and damages brand image.
[0004] Therefore, there is an urgent need to provide a jewelry 3D model design method to establish an objective and quantitative evaluation system through digital and intelligent means, and to convert implicit aesthetic experience into explicit knowledge that can be calculated and optimized. SUMMARY
[0005] To solve at least one or more of the above-mentioned technical problems, the present application proposes a jewelry 3D model design method in multiple aspects.
[0006] In a first aspect, the present application provides a jewelry 3D model design method, characterized in that it comprises: obtaining historical jewelry design full-modal data; using computer vision and natural language processing technology to perform multi-dimensional feature extraction and associated labeling on the full-modal data to construct a three-dimensional feature matrix; according to the three-dimensional feature matrix, constructing a cross-modal associated database through cross-modal structured labeling; using the cross-modal associated database, generating a multi-modal aesthetic feature database through a multi-task deep learning model; using the multi-modal aesthetic feature database, training a multi-task AI aesthetic evaluation model that integrates geometry, style, and user preferences; obtaining jewelry 3D design data through a graphical interface, using the multi-task AI evaluation model to perform quantitative evaluation of 3D design, and generating design evaluation results.
[0007] In some embodiments, the obtaining historical jewelry design full-modal data comprises: obtaining multi-modal data covering the entire life cycle of jewelry design through a multi-source data collection platform; wherein the multi-source data collection platform comprises a design end, a user end, and a production end; and the full-modal data comprises 2D rendering images, 3D model files, designer scores, and user purchase feedback.
[0008] In some embodiments, the three-dimensional feature matrix comprises geometry feature dimensions, style semantic dimensions, and emotional preference dimensions.
[0009] In some embodiments, according to the three-dimensional feature matrix, a cross-modal associated database is constructed through cross-modal structured labeling, which comprises: using knowledge graph technology to construct a cross-modal index system, and using the cross-modal index system to perform quantitative feature structured labeling through weakly supervised learning.
[0010] In some embodiments, the use of knowledge graph technology to construct a cross-modal index system comprises: structurally associating the 2D images and 3D model data of each design case with its geometric features, style labels, and user emotional scores to construct a multi-modal training data set that can be recognized by an AI model.
[0011] In some embodiments, the use of the cross-modal associated database to generate a multi-modal aesthetic feature database through a multi-task deep learning model comprises: using the cross-modal associated database to generate multi-modal aesthetic feature data including quantitative aesthetic scores, defect heat maps, and style matching degrees through a multi-task deep learning model.
[0012] In some embodiments, the multi-task AI aesthetic evaluation model comprises: generating three-dimensional quantitative aesthetic scores, defect positioning heat maps, and style matching degree reports based on voxelized data of a 3D design model, 2D multi-view images, and user input style keywords and design parameters.
[0013] In some embodiments, the design method further comprises: generating multiple versions of aesthetic enhancement schemes for the 3D design by a conditional generative adversarial network according to the design evaluation results, in combination with user preferences and process constraints; and dynamically optimizing a design knowledge base and an AI aesthetic evaluation model by recording user decisions through reinforcement learning.
[0014] In a second aspect, the present application provides a jewelry 3D model design system, comprising: a data acquisition module for acquiring historical jewelry design full-modal data; a feature matrix generation module for performing multi-dimensional feature extraction and association labeling on the full-modal data by computer vision and natural language processing technology to construct a three-dimensional feature matrix; an associated database generation module for constructing a cross-modal associated database by cross-modal structured labeling according to the three-dimensional feature matrix; an aesthetic feature database generation module for generating a multi-modal aesthetic feature database by a multi-task deep learning model using the cross-modal associated database; an AI aesthetic evaluation model training module for training a multi-task AI aesthetic evaluation model that integrates geometry, style and user preferences using the multi-modal aesthetic feature database; and a design evaluation module for obtaining jewelry 3D design data through a graphical interface, performing quantitative evaluation of the 3D design using the multi-task AI evaluation model, and generating design evaluation results.
[0015] By the jewelry 3D model design method provided above, the embodiments of the present application are trained by multi-modal aesthetic feature data, can also support dynamic adjustment of aesthetic weight by users, and take into account the rigor of professional design and the individualized preferences of consumers. Thus, non-designer users can complete high-quality design through an interactive interface, and the efficiency of outputting satisfactory schemes is improved. Furthermore, by constructing an AI aesthetic evaluation model, the evaluation objectivity is significantly improved, and the design efficiency and design quality are improved. BRIEF DESCRIPTION OF DRAWINGS
[0016] The above and other objects, features and advantages of the exemplary embodiments of the present application will be more apparent from the following detailed description taken in conjunction with the accompanying drawings, in which:
[0017] Figure 1 An exemplary flowchart of a jewelry 3D model design method according to some embodiments of the present application is shown;
[0018] Figure 2 An exemplary structural block diagram of a jewelry 3D model design system according to some other embodiments of the present application is shown. DETAILED DESCRIPTION
[0019] With reference to the accompanying drawings, the technical solutions in the embodiments of the present application will be described clearly and completely. Obviously, the described embodiments are only some, but not all, of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts should fall into the scope of the present application.
[0020] The specific implementation of the present application will be described in detail below with reference to the accompanying drawings.
[0021] Figure 1 An exemplary flowchart of a jewelry 3D model design method 100 according to some embodiments of the present application is shown.
[0022] As Figure 1 shown, in flow step S101 of the method 100, historical jewelry design full-modal data is acquired.
[0023] In one embodiment, acquiring historical jewelry design full-modal data includes acquiring multi-modal data covering the full life cycle of jewelry design through a multi-source data acquisition platform. The multi-source data acquisition platform includes a design end, a user end and a production end. The full-modal data includes 2D rendering images, 3D model files, designer scores, user purchase feedback, etc. The process of acquiring data by the multi-source data acquisition platform will be described in detail below.
[0024] Acquiring multi-modal data covering the full life cycle of jewelry design through a multi-source data acquisition platform specifically includes:
[0025] Design end data: mainly includes 2D rendering images (such as JPEG / PNG format, resolution ≥ 1024x1024), 3D model files (such as STL / OBJ / B-rep format, accuracy ≤ 0.01mm), CAD engineering drawings (DWG format), design description text (such as including style description, material parameters, process requirements, etc.), etc.
[0026] User end data: mainly includes e-commerce platform user comments, purchase records, personalized customization requirements (such as "prefer rose gold material" "require inlay emerald"), wearing size data (such as ring size, wrist circumference), etc.
[0027] Production end data: mainly includes 3D printing process parameters (such as support structure design, printing layer thickness), precious metal loss rate, finished product quality inspection report (such as surface roughness, inlay firmness), etc.
[0028] In terms of technical implementation, a network crawler (for public design cases), enterprise CRM system integration (to obtain user customization data), and industrial sensor collection (to obtain production data) can be used to build an original dataset containing 100,000+ samples. The data storage uses a distributed file system (HDFS) to support TB-level data expansion.
[0029] Further, in step S102: multi-dimensional feature extraction and associated labeling of the full modal data are performed using computer vision and natural language processing techniques to construct a three-dimensional feature matrix.
[0030] In one embodiment, the three-dimensional feature matrix can include geometric feature dimensions, style semantic dimensions, and emotional preference dimensions. The specific three-dimensional feature extraction process is as follows:
[0031] Geometric feature extraction (CV technology): OpenCASCADE library is used to analyze 3D models, calculate the number of symmetry axes, golden section proportion fitting degree (Fibonacci proportion of main stone position and contour), surface curvature uniformity (Laplacian operator detection, curvature standard deviation <0.05 mm -1 ), volume-surface area ratio (to evaluate the rationality of material usage); edge detection (Canny operator) is performed on 2D rendering images to extract contour complexity (such as perimeter / area ratio) and element distribution uniformity (which can be calculated by a gray level co-occurrence matrix).
[0032] Style semantic extraction (NLP technology): BERT pre-training model is used to analyze design description text to extract style keywords (such as "Baroque" and "Cyberpunk"), which are mapped to a pre-set style feature library (such as containing 10 basic styles, 20 sub-styles, and 10+ geometric / visual features for each style. For example, "minimalist style" requires ≤3 element types and >80% line simplicity).
[0033] Emotional preference extraction (NLP + machine learning): LSTM model is used for sentiment analysis of user reviews, such as extracting 5 core emotional dimensions: "delicacy", "uniqueness", "value for money", etc., and outputting a sentiment score of 1-5. Then, combined with purchase data, the preference weight is calculated (such as automatically increasing the "comfort" weight of high-frequency repeat purchase designs by 20%).
[0034] Specifically, in the technical implementation process, each design sample can be used to generate a 100+-dimensional feature vector V = [G, S, E], where the geometric feature G occupies 40 dimensions (including symmetry axis, curvature, etc.), the style semantics S occupies 30 dimensions (including style labels, element types, etc.), and the emotional preference E occupies 30 dimensions (including emotional scores, purchase conversion rates, etc.), forming a standardized three-dimensional feature matrix.
[0035] The following describes the feature extraction process in detail through an example. For example, the process of collecting dimensional features of full modal data of a vintage-style ring is as follows:
[0036] Geometric feature dimensions: 3D scanning is used to obtain the curvature of the ring setting surface (e.g. average curvature 0.08mm -1 ), the golden ratio of the main stone and the ring (0.618:1 fit 92%), and the number of symmetry axes (2 vertical symmetry axes).
[0037] Style semantic dimension: NLP analyzes the design description to extract "Baroque" style keywords and associates them with preset features (such as carving density >40 pieces / square centimeter and surface curvature >70°).
[0038] Emotional preference dimension: crawl user comments such as "strong retro feeling" and "exquisite details", and convert them into quantitative indicators through sentiment analysis (such as detail density 35% and historical purchase conversion rate 18%).
[0039] Furthermore, the process proceeds to step S103: constructing a cross-modal association database through cross-modal structured annotation based on the three-dimensional feature matrix.
[0040] In one embodiment, a cross-modal association database is constructed based on the three-dimensional feature matrix through cross-modal structured annotation, including: using knowledge graph technology to build a cross-modal indexing system. Specifically, this may include structurally associating each design case's 2D image and 3D model data with its geometric features, style labels, and user sentiment scores to create a multimodal training dataset that can be recognized by the AI model.
[0041] Specifically, in building a cross-modal indexing system, we use knowledge graph technology to annotate data into triples: <object, feature type, feature value>. For example, <ring A, number of axes of symmetry, 2>, <pendant B, style, modern minimalist>, and <earring C, sentiment score, 4.5>. Next, we establish cross-modal associations (e.g., a positive correlation between "high curvature surface" and "refinement score > 4" with a confidence level > 85%).
[0042] Further, through the cross-modal index system, combined with weakly supervised learning, the structured labeling of quantitative features is carried out. In the structured labeling, the structured labeling tool can be used for labeling the quantitative features. In the visual labeling platform, the designer can manually label the difficult quantitative features (such as “visual center of gravity balance degree”), and the weakly supervised learning (based on rule engine) is used to automatically complete the missing labeling. Through the labeling platform tool, the labeling efficiency can be greatly improved compared with pure manual labeling. Secondly, the field expert verification mechanism is introduced, and the Kappa coefficient (>0.85) is used to ensure the labeling consistency, and a high-quality cross-modal correlation database (storage capacity ≥500GB, supporting second-level full-text retrieval and multi-dimensional filtering) is constructed.
[0043] For example, in one embodiment, the 3D model (STL) of the “vintage style ring”, the 2D rendering graph (JPEG), and the user praise (such as “delicate carving”) are labeled as the same feature group, and the geometric feature “carving density 38%” is associated with the emotional score “delicacy 4.2 points”.
[0044] In the technical implementation, the database can adopt a hybrid architecture of a graph database (such as Neo4j) and a relational database (such as PostgreSQL), which supports efficient storage and associated query of geometric features (numeric type), style labels (text type), and emotional scores (vector type).
[0045] The above embodiments provided by the present application can improve the data integrity by 80% compared with the traditional scheme by constructing a three-dimensional feature matrix covering three dimensions and 100+ sub-indicators, and realize the transition from “experience judgment” to “data definition”.
[0046] The cross-modal correlation database supports multi-dimensional retrieval (such as filtering models according to “golden section ratio >0.9 and user praise rate >90%”), and the design reference efficiency can be improved by 300%.
[0047] By using the structured tool in labeling to improve the data labeling efficiency, the model training data preparation period can be shortened from 3 months to 2 weeks. In the design iteration process, the aesthetic standard rate of the AI generated scheme is improved from 50% to 85%, and the whole process time of the typical jewelry from “parameter input-final model” is compressed to 2-3 hours, which greatly improves the design efficiency compared with the traditional manual design.
[0048] Further, the process proceeds to step S104: using the cross-modal correlation database, a multi-task deep learning model is used to generate a multi-modal aesthetic feature database.
[0049] In one embodiment, the cross-modal association database is used to generate a multimodal aesthetic feature database through a multi-task deep learning model, including: using the cross-modal association database to generate multimodal aesthetic feature data including quantitative aesthetic scores, defect heat maps, and style matching degrees through a multi-task deep learning model.
[0050] The structure of the multi-task deep learning model can adopt a hybrid architecture of Transformer+3D CNN+Graph Neural Network (GNN), which can include 3D CNN branch, Transformer branch, and GNN branch (inter-modal dependency modeling).
[0051] 3D CNN branch: Input voxelized 3D model (resolution 128×128×128) and extract geometric features such as symmetry and proportional coordination. In one embodiment, this is achieved as follows:
[0052] Symmetry score calculation (0-10 points), the formula can be:
[0053] In the above formula, D sym It represents the average Euclidean distance of the mirror point pairs along the symmetry axis of the model. Specifically, the symmetry axis can be determined by principal component analysis PCA, and the coordinates of the symmetric point pairs (p i ,p i '),calculate
[0054] D max Indicates the maximum allowable symmetry distance of jewelry of the same category, such as D for rings max =1.5mm, Pendant type D max =2.0mm, which can be statistically analyzed based on historical high-quality designs.
[0055] Proportion coordination score (0-10 points), the formula is:
[0056] In the above formula, the parameter φ represents the actual proportion parameters (such as the height of the main stone / the width of the ring, the length of the bracket / the total height of the pendant). gold Indicates the target value of the golden ratio (0.618 or its reciprocal 1.618, automatically matched according to the design elements). thres Represents the proportion tolerance threshold (the proportion standard deviation of the top 20% high-quality samples of the same design is taken, which can be adjusted dynamically).
[0057] The Transformer branch processes the design description text and user comments to generate a style semantic vector. For example, a semantic embedding vector for "retro style" is generated. In an example scenario, the specific implementation of the style semantic vector generation is as follows:
[0058] The technical implementation of style keyword encoding can use the BERT-base pre-training model, input design description text T = {t1, t2, ..., t n}, output word-level embedding h i , generate the global semantic vector s through the pooling layer text =CLS(h1,h2,...,h n ). Thus, the style feature library S={s1,s2,...,s 30}(30-dimensional style labels, such as "retro" corresponding to a combination of features such as carving density and surface curvature). Then, the cosine similarity is used to calculate the matching degree between the text and the style library:
[0059] The user review sentiment weighting formula can be: (Comment sentiment dimension k). Among them, parameter w k represents the weight of the sentiment dimension (trained using historical purchase data, for example, the weight of "refinement" automatically increases with repurchase rate). LSTM(·) represents the LSTM encoding of five sentiment dimensions, including "refinement" and "uniqueness," and outputs a 10-dimensional sentiment vector.
[0060] GNN branch: Modeling cross-modal feature associations, such as the weight of the influence of "gold material + complex carving" on the difficulty of craftsmanship, and the weight association between "high curvature surface" and "sophistication".
[0061] In the definition of the graph attention mechanism (GAT), nodes represent geometric features (curvature c, number of elements n), style labels (style), and process parameters (printing layer thickness l, number of inlay claws m). The formula is expressed as v i =[g i ,s i ,p i ].
[0062] First, the node features are initialized. By inputting the cross-modal node feature matrix V={v1,v2,...,v N Node types (heterogeneous graph scenarios) can be geometric feature nodes (such as "high curvature surface" and "symmetrical structure"), style semantic nodes (such as "retro style", "gold material", and "complex carving"), and user preference nodes (such as "refinement" and "craftsmanship difficulty" ratings). Feature dimensions: F is the original feature dimension (such as the curvature value of geometric features and the semantic vector of text embedding).
[0063] Secondly, perform cross-modal feature linear transformation. By sharing the weight matrix Mapping different modal features to a unified space: h i =Wv i , Where W represents the cross-modal feature alignment matrix, making geometric, semantic, and preference features interactive (e.g., mapping the "golden material" text embedding to the same space as the "metallic luster" geometric feature). F' represents the transformed feature dimension (usually set to 256 or 512, adjusted according to model complexity).
[0064] Then, attention coefficient calculation is performed. For the neighbor node j ∈ N i of node i, the cross-modal correlation weight α ij is calculated, with the formula:
[0065] In the above formula, the parameter represents the attention weight vector, capturing the relevance of node pair (i, j) (e.g., the correlation strength of "complex carving" and "process difficulty"). represents the feature concatenation operation (fusing the cross-modal features of nodes i and j). N i represents the neighbor set of node i (e.g., the neighbors of the "golden material" node include associated nodes such as "complex carving" and "high process difficulty").
[0066] For example, if i is the "complex carving" node and j is the "process difficulty" node, α ij represents the influence weight of carving complexity on process difficulty.
[0067] Finally, cross-modal feature aggregation aggregates neighbor features through attention weights to generate the new feature h i ′ of node i: In the above formula, the parameter σ represents the activation function (e.g., ReLU, enhancing nonlinear expression). The role is to aggregate the features of its neighbors "high-reflective geometric surface" and "retro carving" for the "golden material" node, and output the integrated features of material and process.
[0068] Then, the multi-head attention mechanism (Multi-Head Attention) enhances robustness through K independent attention heads, with the result being concatenated or averaged: Where K represents the number of attention heads (usually set K = 8), and different heads can capture different types of associations (e.g., head 1 focuses on "material-process" associations, and head 2 focuses on "geometry-style" associations).
[0069] Edge represents the association between features (e.g., a positive correlation edge between "high curvature c > 0.1" and "delicate feeling score > 4", with the weight initialized as the historical co-occurrence frequency).
[0070] The following example further illustrates the innovation of GAT in the jewelry design scenario in the present application.
[0071] In the modeling of cross-modal associations in heterogeneous graphs, node definitions include:
[0072] Geometric modality node: curvature c i , symmetry index s i , volume ratio v i (Quantify 3D model geometric features).
[0073] Semantic modality node: material embedding m j (e.g. One-Hot encoding or Word2Vec vector of “gold”, “diamond”), style label t j (e.g. text embedding of “retro”, “minimalist”).
[0074] Preference modality node: user rating r k (e.g. rating of “delicacy”, “uniqueness”, normalized to [0, 1]).
[0075] Edge definition includes:
[0076] Geometric-semantic edge: e.g. positive correlation edge (weight +0.8) between “high-curvature surface” and “delicacy”.
[0077] Semantic-preference edge: e.g. association edge (weight obtained from historical data statistics) between “retro style” and “user age 30+”.
[0078] In the constraint optimization of attention weights, for the process constraints in jewelry design (e.g. “complex carving” needs to match the processing feasibility of “gold material”), prior knowledge regularization terms are introduced in the attention coefficients: (c i,j Compliance). Where, indicates the indicator function, if the node pair (i, j) satisfies the process constraint (e.g. “diamond inlay” and “metal claw inlay structure” compliance), then otherwise 0. λ represents the constraint strength coefficient (the empirical value is set to 0.5), which enhances the weight of the compliance association.
[0079] Then in the feature fusion with 3D CNN, Transformer, the cross-modal association features h i ′ output by GAT are concatenated with the geometric features g k extracted by 3D CNN and the semantic vectors s m generated by Transformer:
[0080] The role of the above fusion is to provide a multi-task model with integrated features that fuse geometric structure, semantic description, and cross-modal association, thereby improving the accuracy of quantitative ratings (e.g. symmetry, style purity).
[0081] Through the above embodiments of the present application, the cross-modal correlation features output by the GAT will serve as the input of the multi-task deep learning model, supporting core functions such as quantitative aesthetic score, defect positioning, and style matching degree calculation, and improving the intelligent level of jewelry 3D design.
[0082] The multi-task training target can be divided into the following types.
[0083] Regression task: output 3 sub-dimension quantitative scores (symmetry, proportion coordination, and style purity, range 0-10, mean square error MSE <0.1).
[0084] Regression task loss (symmetry / proportion / style purity) wherein the parameter λ i represents the task weight (dynamically adjusted, such as setting the weight of style purity to 0.4 in brand design and 0.2 in personalized customization); represents the model output score, and represents the expert annotation true value.
[0085] Positioning task: generate a defect heat map (based on a U-Net network, locate the visual imbalance area, pixel-level accuracy IOU > 0.9, such as marking red for the area with a center of gravity offset > 5%).
[0086] Positioning task loss (defect heat map generation) formula L seg = α·DiceLoss + (1-α)·CrossEntropyLoss. This formula can combine Dice loss (to handle class imbalance) and cross-entropy loss to improve the detection accuracy of small defect areas (such as the center of gravity offset point). α represents the balance coefficient (dynamically adjusted according to the proportion of the defect area, default 0.7).
[0087] Classification task: calculate the style matching degree (10 basic style classification accuracy > 95%, output a cosine similarity vector, and a threshold ≥ 0.85 is determined as style compliance).
[0088] The classification task loss (style matching degree) function formula can be L cls =-log(Softmax(StyleSim global )), where StyleSim global = 0.6·StyleSim text + 0.4·StyleSim 3D (fusion of text and 3D model style matching degree).
[0089] In addition, a joint loss function can also be included, with the formula L total =L reg + β·L seg + γ·Lcls +λ adv ·L adv , introducing an adversarial loss L by Domain Adversarial Transfer Learning (DATL) adv , confusing the distribution discrimination between labeled data and unlabeled data by Gradient Reversal Layer, improving the generalization ability of rare style.
[0090] In one embodiment, the use of cross-modal association database, through multi-task deep learning model, generates multi-modal aesthetic feature database, further comprising: introducing domain adversarial transfer learning by domain enhancement technology, increasing the rare style sample of multi-modal aesthetic feature database.
[0091] In the process of introducing domain adversarial transfer learning (DATL), 5000+ professional designer labeled rare style samples (such as Art Nouveau, Bohemian style) can be used to improve the evaluation accuracy of the model for rare styles (score error reduced by 40%).
[0092] Specifically in the process of technical implementation, model training can use PyTorch framework, distributed training cluster (8 card GPU) supports batch processing, single sample feature generation time <200ms, and multi-modal aesthetic feature database containing 200,000+ samples is generated.
[0093] The following further illustrates the multi-modal data driven intelligent design and optimization through an embodiment, which is driven by data through the fusion of full modal data.
[0094] First, initial model generation. After inputting user parameters, CGAN retrieves similar style samples (such as "modern minimalist" style Top100 cases) from the cross-modal database, extracts average geometric features (bracket width 2.5mm, main stone ratio 1.5:1) as generation prior, to ensure that the initial model conforms to the historical high-quality design paradigm.
[0095] Second, dynamic weight evaluation. The evaluation engine calls the multi-modal aesthetic feature database in real time, for example, when calculating "style purity", not only analyzes geometric elements, but also associates user historical preference data (such as the user's past purchase record "minimalist style" accounts for 60%, which automatically increases the weight of this dimension).
[0096] Finally, human-computer co-evolution. Through user decision-making behavior (such as adjusting the position of the main stone), the three-dimensional feature matrix is updated synchronously, for example, after a certain eccentric adjustment, the system records the "asymmetric design + high emotional score" associated features, which in turn benefits model training and improves the generation ability of asymmetric style. Thus, the adoption rate of such designs is improved in subsequent schemes.
[0097] The embodiments of the present invention demonstrated a 92% consistency between quantitative aesthetic scores and expert ratings, a defect heatmap positioning error of <0.1mm, and a style matching detection time of <50ms per test, significantly improving key metrics compared to single-modal models. This system also supports cross-validation across "material-structure-style-preference" criteria (e.g., automatically identifying process risks associated with a "platinum material + complex engraving" combination and providing guidance on 3D printing support structure design), improving the accuracy of feasibility assessments.
[0098] Furthermore, the multimodal aesthetic feature database has accumulated over 200,000 design features, forming reusable "digital design assets" that can shorten the training cycle for new designers. For example, correlation analysis revealed that the proportion of the "asymmetric design + micro-inlay" combination in historical data has increased from 3% to 15%, indicating an emerging customer preference. This can help companies capture market trends and shorten new product development cycles.
[0099] Furthermore, the process proceeds to step S105: using the multimodal aesthetic feature database, training a multi-task AI aesthetic evaluation model that integrates geometry, style, and user preferences;
[0100] In one embodiment, the multi-task AI aesthetic evaluation model may include: evaluating and generating a three-dimensional quantitative aesthetic score, a defect location heat map, and a style matching report based on voxelized data of the 3D design model, 2D multi-view images, and style keywords and design parameters input by the user.
[0101] The following describes in detail the process of implementing the multi-task AI aesthetic evaluation model technology through examples.
[0102] 1. Model input and feature fusion process
[0103] Input layer design: The model receives multimodal input data, including:
[0104] 3D voxel data: 3D model voxel data with a resolution of 128×128×128, denoted as Used to extract geometric features.
[0105] 2D multi-view images: front view, side view, and top view, respectively denoted as I front ,I side , (H and W are the image height and width, and 3 represents the RGB channels).
[0106] User input data: style keyword text sequence T = {t1, t2, ..., t n} and design parameter vector P=[p1,p2,...,p m ], where p iRepresent the parameters such as material, gem specifications, etc.
[0107] The feature fusion formula can be: different modal features are fused into a unified feature vector F fusion :
[0108] F fusion =Concat(F 3D ,F 2D ,F text ,F param )
[0109] Where: F 3D : 3D CNN extracted geometric feature vector, dimension d 3D ; F 2D : ResNet extracted 2D image feature vector, dimension d 2D ; F text : BERT encoded style keyword semantic vector, dimension d text ; F param : embedding vector of design parameters, dimension d param ; Concat is a concatenation operation.
[0110] This process has the following technical effects:
[0111] Multi-modal information complementarity: fusion of 3D geometric data (voxel), 2D visual features (multi-view image), semantic text (style keyword) and parameterized design requirements, covering design full-dimensional information, which can solve the one-sidedness of single modal evaluation (such as geometric data cannot capture the semantic connotation of "retro style").
[0112] Feature space alignment: through cross-modal splicing and attention mechanism, "material text description" and "metal reflective geometric feature" are mapped to a unified semantic space, which can improve the feature interaction ability (such as "platinum" text embedding and "low reflectivity surface" feature correlation strength is improved by 40%).
[0113] Input flexibility enhancement: support parameterized input, sketch recognition, existing model import, compatible with different design habits, which can reduce the user threshold (non-professional user input efficiency can be improved by 50%).
[0114] Two and three-dimensional quantitative aesthetic score calculation process
[0115] 1. Symmetry score (S sym , 0-10 points) formula:
[0116]
[0117] Where, N: the number of corresponding point pairs on both sides of the symmetry axis of the 3D model; d iEuclidean distance between the ith pair of points; d max Maximum symmetry distance threshold allowed for similar jewelry design, calculated from historical high-quality design data, e.g., ring class d max = 1.2 mm.
[0118] 2. Proportionality score (S ratio , 0-10) formula:
[0119] where a, b: dimensions of key geometric elements in the 3D model (e.g., main stone length and ring band width); Golden ratio 0.618; δ: proportion tolerance coefficient, dynamically adjusted according to design type, determined by the proportion standard deviation of similar designs.
[0120] 3. Style purity score (S style , 0-10) formula:
[0121] In the above formula, k: number of style features; w i : weight of the ith style feature, determined according to field knowledge and data statistics; f i : ith style feature vector extracted from the design; s i : ith standard feature vector of the target style; Sim(·): cosine similarity function, used to calculate the similarity between the design feature and the standard style feature.
[0122] This process has the following technical effects:
[0123] Symmetry score: Quantify symmetry through geometric point pair distance, score standard deviation can be reduced from ±1.8 to ±0.35, eliminating designer subjective differences (e.g., "double-axis symmetric ring" evaluation consistency can be improved from 65% to 92%). Automatically identify problems such as center of gravity shift and structural imbalance, with positioning error <0.1 mm, improving efficiency by 80% compared to traditional CAD measurement.
[0124] Proportionality score: Based on the golden ratio and dynamic tolerance threshold, ensure that the design conforms to aesthetic rules (e.g., main stone and ring band proportion compliance rate can be improved from 8% to 91%). Support custom proportion parameters (e.g., automatically adjust tolerance when user prefers "exaggerated proportion"), balancing standardization and individualization needs.
[0125] Style purity score: Based on style feature weight and cosine similarity, quantify the fit between design elements and target style (e.g., "minimalist style" element compliance detection accuracy can be improved from 70% to 95%). Automatically filter style conflict elements (e.g., "minimalist design" redundant carving detection rate can reach 98%), reducing manual screening time.
[0126] 3. Defect Localization Heatmap Generation The U-Net network architecture is used for defect localization, and the attention mechanism is introduced to enhance key area detection.
[0127] The loss function formula is: L seg =α×DiceLoss+(1-α)×FocalLoss.
[0128] In the above formula, α is a weighting coefficient used to balance the two losses, with a dynamic value range of [0.3, 0.7].
[0129] DiceLoss: Dice loss function, used to deal with category imbalance problems, the formula is:
[0130] Where p(x,y) is the pixel probability predicted by the model, and g(x,y) is the true label pixel value.
[0131] FocalLoss: focal loss function, which reduces the weight of simple samples. The formula is:
[0132] FocalLoss=-(1-p(x,y)) γ ×log(p(x,y)), where γ is the focusing parameter, which is empirically taken as 2.
[0133] This process has the following technical effects:
[0134] Pixel-level precision detection: Combining DiceLoss and FocalLoss, small defects (such as 0.2mm 2 The detection rate of abnormal curvature areas can be increased from 60% to 89%, and the IOU (intersection over union) reaches 0.92, which is 35% higher than the traditional edge detection algorithm.
[0135] Real-time visual feedback: Defective areas are marked in red and overlaid with improvement suggestions (e.g., "The center of gravity is 0.3mm to the left; it is recommended to move the auxiliary stone to the right"). Designers can reduce the time it takes to locate problems from 30 minutes per model to 5 minutes per model.
[0136] Process risk prediction: By associating geometric features with process constraints (such as the "narrow neck structure" corresponding to the thermal area with fracture risk), early warning of manufacturability issues can be provided, reducing the rework rate on the production side by 65%.
[0137] 4. The calculation process of style matching is as follows
[0138] Style feature extraction: Use the BERT model to encode style keywords and obtain the style semantic vector S target , and extract the style feature vector S from the design design .
[0139] Cosine similarity calculation: Where S target · S design : dot product of two vectors; ‖S target ‖, ‖S design ‖: L2 norm of two vectors respectively.
[0140] Feature difference analysis: Calculate the difference between the design feature and the target style feature in each dimension. The formula can be: ΔF i = |f i,design -f i,target |, where F i represents the i-th style feature dimension, f i,design and f i,target are the feature values of the design and target style in this dimension respectively.
[0141] This process has the following technical effects:
[0142] Semantic-level style quantification: Through BERT encoding and cosine similarity, realize the numerical matching of "retro" "cyberpunk" and other styles (the accuracy rate of 10 basic style classification can reach 96.2%, and the recognition rate of rare styles can be improved by 28%).
[0143] Feature difference is interpretable: output the difference value of each dimension (such as "carving density insufficient-15%" "curved arc deviation +8°"), guide the designer to accurately adjust (the number of style matching iterations can be reduced from 5 times / model to 2 times / model).
[0144] Brand compliance management: Predefine brand style feature library (such as "Brand A requires symmetry ≥8 points"), automatically intercept non-compliant designs, and brand standard execution rate can be improved to 100%.
[0145] Five, data enhancement and contrast learning process as follows.
[0146] 1. Data enhancement includes:
[0147] 3D model enhancement: rotation (angle range [-15°, 15°]), scaling (scale range [0.9, 1.1]), local deformation (randomly change part of the voxel position).
[0148] 2D image enhancement: horizontal flip, vertical flip, add Gaussian noise (mean 0, standard deviation 0.05).
[0149] 2. Contrast learning includes:
[0150] Construct contrast learning loss function L contrast :
[0151]
[0152] In the above formula, F i : anchor point sample feature vector; Positive sample (similar high-scoring design) feature vector; Negative sample (different class or low-scoring design) feature vector; K: number of negative samples; τ: temperature parameter, which controls the tightness of feature distribution, with an empirical value of 0.1.
[0153] This process has the following technical effects:
[0154] Data diversity expansion: 3D geometry enhancement and 2D image augmentation generate more than 500,000 samples, which can improve the model's generalization ability (evaluation error under rare angle views is reduced by 32%).
[0155] Feature space clustering optimization: Contrastive learning reduces the intra-class distance of features for "high-scoring designs" by 30% and increases the inter-class distance for "low-scoring designs" by 45%, improving the model's discriminative power (the accuracy of positive and negative sample classification can be increased from 82% to 94%).
[0156] Overfitting suppression: By mining difficult samples and comparing them with negative samples, the risk of overfitting of the model on small data sets (such as only 1,000 labeled samples) can be reduced by 50%.
[0157] 6. Transfer Learning Strategy
[0158] 1. Pre-training stage: Pre-train the model on the ImageNet (image classification) and ShapeNet (3D shape classification) datasets to learn general visual features and geometric feature representations.
[0159] 2. Fine-tuning stage: Fine-tuning is performed on jewelry data, using an adaptive learning rate adjustment strategy:
[0160] In the above formula, lr0: initial learning rate; t: current training round number; T: total training round number; β: decay coefficient, empirically set to 0.9.
[0161] This process has the following technical effects:
[0162] Improved training efficiency: Based on ImageNet / ShapeNet pre-trained model initialization, the jewelry training cycle can be shortened from 8 weeks to 4 weeks, reducing computing resource consumption by 50%.
[0163] Enhanced generalization ability: F1-score>0.92 on industry-standard test sets (such as JewelryDesignBenchmark), an 18% improvement over training from scratch.
[0164] Cross-domain knowledge reuse: The migration of common visual features (such as edge detection and shape recognition) and geometric priors (such as symmetry rules) can increase the model's adaptability to new designs (such as special-shaped gemstone inlays) by 40%.
[0165] This multi-task AI aesthetic evaluation model, through comprehensive multimodal feature fusion, objective quantitative scoring, precise defect detection, intelligent style matching, and efficient data utilization, has established a digital evaluation system for jewelry design aesthetics. Multiple experiments have shown that it can achieve the following:
[0166] Evaluation efficiency: Single model evaluation takes less than 200ms, which is more than 99% more efficient than manual evaluation.
[0167] Design quality: The aesthetic score compliance rate increased from 50% to 85%, and the proportion of designs with a style matching degree ≥ 0.85 increased by 60%.
[0168] Innovation empowerment: Support users with no basic knowledge to complete professional-level aesthetic design, and the output efficiency of non-designer users can be increased by 300%;
[0169] Knowledge accumulation: Accumulated evaluation data feeds back into the model to form a closed loop of "design-evaluation-optimization". The adaptation cycle for emerging styles can be shortened from 3 months to real-time response.
[0170] Therefore, through the above-mentioned formula algorithms and technical implementation, the multi-task AI aesthetic evaluation model can effectively integrate multiple factors such as geometry, style, user preferences, etc., to achieve high-precision aesthetic scoring, defect location and style matching, significantly improving the intelligence level and design quality of jewelry 3D design.
[0171] Through the above description of the multi-task AI aesthetic evaluation model, those skilled in the art can understand that by inputting voxelized data (128×128×128) of the 3D model, 2D multi-view images (front view / side view / top view), and user-entered style keywords and design parameters, it is possible to output a three-dimensional quantitative score (symmetry, proportion coordination, style purity), a defect location heat map (marking areas such as curvature anomalies and center of gravity offset), and a style matching report (cosine similarity with the target style, and feature difference analysis).
[0172] For example, it can perform geometric transformations (rotation, scaling, and local deformation) on 3D models and data augmentation (flipping and noise addition) on 2D images, generating over 500,000 augmented samples and improving model generalization. A contrastive learning mechanism is introduced to force the model to distinguish between the features of "high-scoring designs" and "low-scoring designs," optimizing feature space clustering (reducing intra-class distances by 30%).
[0173] In technical implementation, a transfer learning strategy can be adopted. The pre-trained model is initialized on the ImageNet and ShapeNet datasets, and then fine-tuned for jewelry field data. This can shorten the training cycle by 50%, making the evaluation model's F1-score > 0.92 on the industry standard test set.
[0174] Further, the process proceeds to step S106: obtaining jewelry 3D design data through a graphical interface, and using a multi-task AI evaluation model to perform a quantitative evaluation of the 3D design to generate a design evaluation result.
[0175] In an embodiment, the system can be implemented through user parameterized input (e.g., style, material, gemstone specifications, and wearing size), sketch recognition (e.g., parsing hand-drawn sketches based on the YOLO model and converting them into geometric constraints), and importing existing models (e.g., compatible with mainstream CAD formats). Then, through a conditional generative adversarial network (CGAN), the system combines user parameters with high-quality design priors from a cross-modal database (e.g., the average geometric parameters of historically high-scoring models of similar styles) to generate a basic STL-formatted model (with a generation time of less than 2 minutes and an accuracy of 0.01mm).
[0176] Furthermore, in one embodiment, the design method provided by the present invention may also include: based on the design evaluation results, combined with user preferences and process constraints (such as the minimum wall thickness of 0.8mm for 3D printing and the minimum strength requirement for inlay claws), generating multiple versions of aesthetic enhancement solutions for the 3D design through a conditional generative adversarial network.
[0177] During implementation, the first step is to design a conditional input layer. The input vector includes: basic parameters (material, gemstone specifications, and wearer size); the current defect vector (e.g., [Symmetry 4.2, Proportion Harmony = 5.5]); the user preference vector (personalized weights trained using historical interaction data, such as "prefer asymmetric design" with a weight > 0.7); and the process constraint vector (e.g., "minimum number of prongs = 4," "minimum precious metal thickness = 0.5mm").
[0178] Secondly, a geometric optimization branch (3D-GAN) is performed through a dual-branch GAN architecture: for structural defects (such as center of gravity offset and curvature mutation), geometric parameter adjustment plans are generated (such as modifying the layout coordinates of the auxiliary stone and adjusting the bracket curvature parameters), and local modification files in STL format are output (supporting Boolean operations with the original model); and a style optimization branch (image GAN): for style conflicts (such as "minimalist style" containing too many decorative elements), replacement parts are matched from the preset element library (1000+ parametric modules, such as "geometric blocks" and "glossy metal strips") to generate 2D renderings for users to preview.
[0179] Finally, through multi-scheme screening and visualization, 5-8 candidate schemes are generated each time, and the Pareto optimal solution is screened among the three targets of "aesthetic improvement amplitude", "process feasibility" and "modification complexity" by non-dominated sorting genetic algorithm (NSGA-II). The scheme is accompanied by visual difference analysis (such as heat map comparison of visual focus distribution and material reflection difference between the original scheme and the optimized scheme), and the user can click to synchronize to the 3D modeling interface, supporting real-time preview of modification effect.
[0180] In the process of technical implementation, the geometry optimization branch adopts Progressive Growing 3D-GAN to improve the generation quality of complex structures. The style optimization branch can be based on StyleGAN2, supporting consistent rendering of light and shadow after element replacement. After testing, the scheme generation time is <30 seconds / round.
[0181] Further, through reinforcement learning to record user decision-making, dynamically optimize the design knowledge base and AI aesthetic evaluation model, so that the user decision-making behavior is recorded through reinforcement learning, the AI evaluation model is dynamically optimized, and the design knowledge base is fed back, forming a self-evolution closed loop.
[0182] In recording user decisions, user preference modeling is needed. First, record user interaction trajectory (parameter adjustment amplitude, scheme adoption / rejection history, weight modification record) through decision tracking system, generate preference trajectory matrix T = [t_1, t_2,..., t_n], where t_i is the weight adjustment vector of the i-th interaction.
[0183] Second, use deep Q network (DQN) to model user preferences. When the number of similar decisions is ≥50 times, trigger model fine-tuning (such as the user's continuous preference for low symmetry design, the system automatically reduces the "symmetry" weight threshold by 20%, and increases the probability of generating asymmetric elements to 60%).
[0184] Further, dynamic design knowledge base iteration is carried out. First, parameterized module sedimentation is carried out, and high-frequency adopted schemes (adoption rate >60%) are disassembled into reusable parameterized design modules (such as "eccentric ellipse main stone + single row micro inlay" structure), and stored in the component library, supporting one-key calling (reducing repeated design time by 40%). Second, federal learning update is carried out. Regularly aggregate multi-user desensitized feature vectors (only upload hash processed feature abstract), update global evaluation model through federal learning, protect user privacy while improving the adaptability of the model to emerging trends (every 10,000 interactions, the accuracy of rare style evaluation increases by 10%-15%).
[0185] In the process of technical implementation, the preference learning module adopts online learning mechanism, and generates a model incremental update package every 100 interactions. The knowledge base management system supports version control, records module evolution history, and ensures design traceability.
[0186] In summary, the above-mentioned embodiments of the present invention, through the deep integration of omnimodal data acquisition, cross-modal correlation modeling, and multi-task intelligent generation, establish a closed "data-model-application" loop for jewelry design. By transforming jewelry design from "experience-driven" to "data-driven," they address the core issues of subjective evaluation, data fragmentation, and inefficient iteration in traditional design, providing the industry with a new paradigm for intelligent design that is quantifiable, traceable, and adaptable.
[0187] Figure 2 FIG. 2 shows an exemplary structural block diagram of a jewelry 3D model design system 200 according to some other embodiments of the present invention.
[0188] like Figure 2 As shown, system 200 includes: a data acquisition module 201 for acquiring historical omnimodal data on jewelry designs; a feature matrix generation module 202 for performing multi-dimensional feature extraction and association annotation on the omnimodal data using computer vision and natural language processing techniques to construct a three-dimensional feature matrix; a correlation database generation module 203 for constructing a cross-modal correlation database based on the three-dimensional feature matrix through cross-modal structured annotation; an aesthetic feature database generation module 204 for generating a multimodal aesthetic feature database using the cross-modal correlation database and a multi-task deep learning model; an AI aesthetic evaluation model training module 205 for training a multi-task AI aesthetic evaluation model that integrates geometry, style, and user preferences using the multi-modal aesthetic feature database; and a design evaluation module 206 for acquiring 3D jewelry design data through a graphical interface and performing quantitative evaluation of the 3D design using a multi-task AI evaluation model to generate design evaluation results.
[0189] Through the jewelry 3D model design system provided by the embodiment of the present invention, it can be understood that the system 200 is a specific implementation of the embodiment of the aforementioned method 100. Therefore, the embodiment features described in the method 100 above can be similarly applied here.
[0190] Although a number of embodiments of the present invention have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Those skilled in the art may devise numerous modifications, variations, and alternatives without departing from the concept and spirit of the present invention. It should be understood that in practicing the present invention, various alternatives to the embodiments of the present invention described herein may be employed. The appended claims are intended to define the scope of the present invention and therefore cover equivalents or alternatives within the scope of these claims.
Claims
1. A method for designing a 3D model of jewelry, characterized in that: include: Obtain full modality data of historical jewelry designs; Using computer vision and natural language processing technology to perform multi-dimensional feature extraction and association annotation on the full modality data to construct a three-dimensional feature matrix; According to the three-dimensional feature matrix, a cross-modal association database is constructed through cross-modal structured annotation; Using a cross-modal correlation database and a multi-task deep learning model, a multimodal aesthetic feature database is generated; Leveraging a multimodal aesthetic feature database, we train a multi-task AI aesthetic assessment model that integrates geometry, style, and user preferences. The 3D design data of jewelry is obtained through a graphical interface, and a multi-task AI evaluation model is used to perform quantitative evaluation of the 3D design to generate design evaluation results.
2. The method according to claim 1, characterized in that The acquisition of historical jewelry design full modality data includes: Through the multi-source data collection platform, multi-modal data covering the entire life cycle of jewelry design is obtained; among them, The multi-source data acquisition platform includes a design end, a user end, and a production end; The full modal data includes 2D renderings, 3D model files, designer ratings, and user purchase feedback.
3. The method according to claim 1, characterized in that The three-dimensional feature matrix includes a geometric feature dimension, a style semantic dimension, and an emotional preference dimension.
4. The method according to claim 1, wherein Based on the three-dimensional feature matrix, a cross-modal association database is constructed through cross-modal structured annotation, including: Using knowledge graph technology to build a cross-modal indexing system, Through a cross-modal indexing system combined with weakly supervised learning, structured annotation of quantitative features is performed.
5. The method according to claim 1, wherein The knowledge graph technology is used to build a cross-modal indexing system, including: structurally associating the 2D images and 3D model data of each design case with its geometric features, style labels, and user sentiment scores to build a multimodal training dataset that can be recognized by the AI model.
6. The method according to claim 1, characterized in that The cross-modal association database is used to generate a multimodal aesthetic feature database through a multi-task deep learning model, including: Utilizing a cross-modal association database, a multi-task deep learning model is used to generate multimodal aesthetic feature data including quantitative aesthetic scores, defect heat maps, and style matching.
7. The method according to claim 1, characterized in that The method of generating a multimodal aesthetic feature database by using a cross-modal association database and a multi-task deep learning model further includes: Through domain enhancement technology, domain adversarial transfer learning is utilized to increase the scarce style samples of the multimodal aesthetic feature database.
8. The method according to claim 1, characterized in that The multi-task AI aesthetic evaluation model includes: Based on the voxelized data of the 3D design model, 2D multi-view images, and user-entered style keywords and design parameters, the evaluation generates a three-dimensional quantitative aesthetic score, a defect location heat map, and a style matching report.
9. The method according to claim 1, characterized in that The design method further includes: Based on the design evaluation results, combined with user preferences and process constraints, a conditional generative adversarial network is used to generate multiple versions of aesthetic enhancement solutions for the 3D design. Through reinforcement learning, user decisions are recorded and the design knowledge base and AI aesthetic evaluation model are dynamically optimized.
10. A jewelry 3D model design system, characterized in that: include: Data acquisition module: used to obtain full-modal data of historical jewelry designs; Feature matrix generation module: used to extract and associate multi-dimensional features of the full modality data using computer vision and natural language processing technology to construct a three-dimensional feature matrix; A relational database generation module is configured to construct a cross-modal relational database based on the three-dimensional feature matrix through cross-modal structured annotation; Aesthetic feature database generation module: used to generate a multimodal aesthetic feature database using a cross-modal association database and a multi-task deep learning model; AI aesthetic evaluation model training module: used to train a multi-task AI aesthetic evaluation model that integrates geometry, style, and user preferences using a multimodal aesthetic feature database; Design evaluation module: used to obtain jewelry 3D design data through a graphical interface, and use a multi-task AI evaluation model to perform quantitative evaluation of 3D design to generate design evaluation results.
Citation Information
Patent Citations
Image aesthetics quality evaluation method based on cross-modal collaborative reasoning
CN112580636A
Multi-modal data driven generation type fashion compatible costume design method and system
CN117951763A
Virtual twin product aesthetic and cultural style evaluation system and method based on big data
CN118246990A
Ai-based product recommendation method, apparatus, and system for jewelry products rendered using 3D drawings
KR102811342B1
System of generating multi-style learning tutorials based on learning preference evaluation
US20250148929A1
Cited By
Three-dimensional model establishment method applied to plane art
CN121259246A