A Smart Design Method and System for Slippers Based on AIGC
By constructing a semantic-geometry-image joint embedding space model and using an iterative optimization method, the problems of ambiguous target expression and deviation in generated results in AIGC slipper design were solved. This achieved precise alignment between the generated results and the design intent, as well as efficient iteration, ensuring that the generated images are adapted to the actual shoe last structure.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- JIEYANG VOCATIONAL & TECH COLLEGE
- Filing Date
- 2026-05-11
- Publication Date
- 2026-07-31
AI Technical Summary
Existing AIGC-based slipper design methods suffer from problems such as vague design goal expression, large deviation between generated results and design intent, and low iteration efficiency. They also lack a quantitative measurement mechanism between generated images and design goals, making it difficult to generate design renderings that fit the actual shoe last structure.
A semantic-geometric-image joint embedding space model is constructed. The encoder is trained through triple contrastive learning to generate structured prompt word templates. The semantic discriminator, geometric discriminator and aesthetic discriminator are used for iterative optimization to ensure that the generated image fits the shoe last geometry until convergence.
It achieves precise alignment between the generated results and the design intent, improves the practicality and iteration efficiency of the design, ensures that the generated design renderings can be directly adapted to the actual shoe last structure, and supports efficient closed-loop optimization.
Smart Images

Figure CN122490615A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of artificial intelligence-aided design technology, specifically relating to an AIGC-based intelligent design method and system for slippers. Background Technology
[0002] With the rapid development of AI-generated content (AIGC) technology, especially the deep integration of diffusion models and large language models, generating product design renderings based on text prompts has become an important technological direction in the field of industrial design. In the field of slipper design, designers typically need to perform structural design based on specific shoe lasts (i.e., foot molds used in shoemaking) to ensure that the final product achieves a balance between aesthetic expression and wearing comfort.
[0003] However, existing AIGC-based slipper design methods suffer from the following technical shortcomings: ① Vague design goals. User-inputted text descriptions (such as "casual style slippers" or "sports sandals") are often semantically ambiguous, making it difficult to accurately translate them into design constraints that the generative model can understand. This results in high randomness and poor controllability of the generated results. ② Large deviation between generated results and design intent. Existing methods lack a quantitative measurement mechanism for the deviation between the generated image and the design goal, particularly lacking the ability to verify geometric constraints such as slipper structure and last fit. This results in visually appealing design renderings that fail to adapt to the actual shoe last structure, making them unsuitable for production. ③ Low iteration efficiency. Existing methods often rely on trial and error by manually adjusting prompts, lacking a mathematically guided closed-loop optimization mechanism. The iteration path is random, and the convergence speed is slow, failing to meet the needs of efficient design.
[0004] Therefore, there is an urgent need for a smart slipper design method that can integrate the geometric constraints of the shoe last into the generation process, establish a semantic-geometric-image joint space, and achieve closed-loop iterative optimization. Summary of the Invention
[0005] The purpose of this invention is to provide an AIGC-based intelligent design method and system for slippers, which solves the technical problems of vague target expression, large deviation between generated results and design intent, and low iteration efficiency in the existing AIGC slipper design.
[0006] To achieve the above objectives, the present invention adopts the following technical solution: A smart slipper design method based on AIGC includes: Step 1: Construct a joint embedding space model containing a semantic encoder, a geometric encoder, and an image encoder; Step 2: Obtain the user's design target text and the 3D file of the shoe last, extract the geometric vector of the shoe last and encode the target text into a semantic vector, and fuse them to obtain the constraint target vector; Step 3: The large language model generates a structured prompt word template containing fixed semantic slots and dynamic geometric parameter slots based on the constraint target vector; Step 4: Input the prompt word template into the generative AIGC model to generate a slipper design rendering. Use a semantic discriminator, a geometric discriminator, and an aesthetic discriminator to output semantic fit, geometric fit, and aesthetic score for the rendering. The generated image vector is obtained by the image encoder. Based on the distance between the generated image vector and the constraint target vector in the joint embedding space and combined with the score, construct a fitting loss and iteratively update the prompt word template until convergence. Step 5: Geometrically register the converged rendering with the 3D shoe last file and generate a matching 3D slipper design file.
[0007] Furthermore, the joint embedding space model is trained using triple contrastive learning, wherein the triple includes a design semantic description, a corresponding shoe last geometric vector, and a slipper design image that matches the semantic description and the shoe last.
[0008] Furthermore, the extraction of the shoe last geometric vector includes: performing mesh sampling on the shoe last 3D file, extracting principal component analysis coefficients, radial basis function descriptors, and curvature distribution of key sections, and splicing them to form a geometric feature vector.
[0009] Furthermore, the structured prompt word template includes at least one dynamic geometric parameter slot, the value of which is associated with at least one dimension of the shoe last geometric vector and allows for continuous updates during iterative optimization.
[0010] Furthermore, the geometric discriminator includes a geometric constraint calculation module, which is used to extract the slipper outline or implicit geometric features from the slipper design rendering and perform structural fit calculation with the shoe last geometric vector to obtain a geometric fit score.
[0011] This invention also provides an AIGC-based intelligent slipper design system, applied to the method, comprising: A joint embedding space building block is used to build, train, and store semantic encoders, geometric encoders, and image encoders. The shoe last parsing and target fusion module is used to receive shoe last 3D files and design target text, extract shoe last geometric vectors and semantic vectors and fuse them to generate constraint target vectors; The parameterized prompt word engine module contains a large language model for generating and iteratively updating structured prompt word templates containing fixed semantic slots and dynamic geometric parameter slots; The iterative optimization control module is used to input the prompt word template into the generative AIGC model to generate a slipper design rendering. It uses a semantic discriminator, a geometric discriminator, and an aesthetic discriminator to output semantic fit, geometric fit, and aesthetic score for the rendering. The generated image vector is obtained by the image encoder. Based on the distance between the generated image vector and the constraint target vector in the joint embedding space and combined with the score, a fitting loss is constructed. The prompt word template is iteratively updated until convergence. The 3D fitting modeling module is used to geometrically register the converged rendering with the 3D shoe last file and output a matching 3D slipper design file.
[0012] In summary, due to the adoption of the above technical solution, the beneficial effects of the present invention are: 1. This invention improves design practicality by introducing shoe last geometric priors. It uses shoe last 3D files as design input and extracts key structural features through a geometric encoder to ensure that the generated design renderings have inherent consistency with the final wearable structure. It constructs a semantic-geometric-image joint embedding space to achieve accurate cross-modal alignment. Through triplet contrastive learning, it maps semantic descriptions, shoe last geometric vectors, and design images to the same latent space, establishing a mathematical relationship between the three and providing a unified metric for subsequent closed-loop iterations. Attached Figure Description
[0013] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0014] Figure 1 The diagram illustrates the steps of an AIGC-based intelligent slipper design method according to the present invention. Detailed Implementation
[0015] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0016] Example 1, such as Figure 1 The AIGC-based smart slipper design method shown includes the following steps: Step 1: Construct a joint embedding space model containing a semantic encoder, a geometric encoder, and an image encoder; By constructing a training dataset Complete data preparation, including To design semantic descriptive text, For the corresponding shoe last 3D file (preferably OBJ or STL format). Design renderings (RGB images) of slippers that match this semantic description and shoe last; in Obtained from the traceable "design requirements / design brief" in the product design process of footwear companies. It can be directly exported from the enterprise's shoe last CAD library, or obtained from the 3D scanning and reconstruction of physical shoe lasts. .
[0017] After data preparation is complete, a semantic encoder is constructed. Geometric encoder With image encoder Preferably, a semantic encoder Pre-trained large language models (such as BERT or RoBERTa) are used as the backbone network to output semantic features; a geometric encoder is also employed. The 3D file of the shoe last is preprocessed and subjected to uniform mesh sampling (preferably sampling to M=1024 vertices). Then, PointNet++ or DGCNN is used to extract global geometric features. Principal component analysis coefficients and radial basis function descriptors can be further concatenated to enhance the geometric representation, forming the shoe last geometric vector. (In this embodiment, p=256 is preferred); Image encoder ViT-B / 16 or ResNet-50 are preferred for outputting image features.
[0018] The outputs of the three encoders are mapped to a unified latent space through a linear projection layer. and order , , During training, the triplet contrastive learning loss function is preferred. ,in The boundary threshold is set to ensure that the distance between the semantic vector and the last vector is greater than the distance between the semantic vector and the matching image vector under the same design objective, thereby establishing a triangular relationship between "semantics-last-image". The optimized trainer is Adam, and the initial learning rate is preferably... The preferred batch size is 64, and the preferred number of training rounds is 50. After training, the three encoders and their projection layers together constitute a joint embedding space model.
[0019] Step 2: Obtain the user's design target text and the 3D file of the shoe last, extract the geometric vector of the shoe last and encode the target text into a semantic vector, and fuse them to obtain the constraint target vector; The slipper design goal description text receives user input. and the 3D file of the shoe last as the design basis. , The data is obtained by the user directly inputting natural language through the interactive interface; alternatively, the system also provides a structured form for the user to select fields such as "style / material / color scheme / scene / structural elements," and the system then transcribes the structured fields into natural language text using the same template as in the training phase. Further, optionally, if the user inputs via voice, the voice is first transcribed into text before being used. The preferred method for extracting the geometric vectors of the shoe last includes: performing mesh sampling on the input 3D file of the shoe last, extracting its principal component analysis coefficients (e.g., taking the first 20 principal component coefficients), radial basis function descriptors (e.g., calculating the radial distance and normal angle of 200 key points uniformly selected on the surface of the shoe last to form a 200-dimensional descriptor), and curvature distribution of key sections (e.g., calculating the curvature distribution statistics of 10 sections equidistantly cut along the length of the shoe last to form a 30-dimensional feature), and then stitching the above features together to form... Then Input the trained geometric encoder To obtain the geometric embedding vector At the same time, using semantic encoders Will Transform into semantic vectors The semantic vector and the geometric embedding vector are fused to generate the constrained target vector. The fusion function is selected as a weighted fusion based on an attention mechanism; in this embodiment, a cross-attention mechanism is preferred. ; in, For a learnable projection matrix, This allows the constraint target vector to dynamically adjust the weight of attention to last features based on semantic content.
[0020] Step 3: The large language model generates a structured prompt word template containing fixed semantic slots and dynamic geometric parameter slots based on the constraint target vector; Prompt word generator driven by large language model According to the constrained target vector and shoe last geometry vector Generate structured prompt templates Its structure satisfies To ensure that the values of the dynamic geometric parameters slots are consistent with... It is associated with at least one dimension and can be continuously updated.
[0021] Since the input to large language models is usually text, the preferred engineering implementation is as follows: First and Decode into readable structured fields, then hand them over to... Generate the final template text. Optionally, this decoding can be performed in either of the following ways: Retrieval-based decoding: in the joint embedding space Retrieve the most similar training samples High-frequency semantic tags / keywords are extracted as candidates for fixed semantic slots, and combined with the user's original... Perform consistency filtering; Mapping decoding: from Dimensions or combinations of dimensions related to last length, last circumference, and heel convexity are selected, and after linear / nonlinear mapping and normalization to a preset range, dynamic geometric parameter slots are formed. The initial values correspond to the last length ratio coefficient, last circumference scaling coefficient, and heel convexity adjustment coefficient. In this embodiment, the value range is limited to [0.8, 1.2]. Subsequently, the fixed semantic slots are... Composition of structured text input In one example, the output structured prompt template This can be represented as "<Style: Sporty and Casual><Upper: Mesh Fabric><Last: Material: EVA Color scheme: Black / White>, ensuring that the template can be used directly by the AIGC model and is easy to continuously update ϕ in subsequent iterations.
[0022] Step 4: Input the prompt word template into the generative AIGC model to generate a slipper design rendering. Use a semantic discriminator, a geometric discriminator, and an aesthetic discriminator to output semantic fit, geometric fit, and aesthetic score for the rendering. The generated image vector is obtained by the image encoder. Based on the distance between the generated image vector and the constraint target vector in the joint embedding space and combined with the score, construct a fitting loss and iteratively update the prompt word template until convergence. Structured prompt template Input the AIGC model to generate the initial slipper design rendering. The AIGC model is preferably a latent diffusion model, and the image generated in the t-th subsequent generation is denoted as... Using a set of multimodal discriminators to The evaluation should include at least a semantic discriminator. Geometric discriminant and aesthetic discriminant Output semantic fit scores respectively Geometric fit score And aesthetic rating At the same time, using an image encoder Map the generated image to a generated image vector. .
[0023] Among them, geometric discriminator Preferably, a geometric constraint calculation module is included to extract the slipper outline or implicit geometric features from the generated image and compare them with the input shoe last geometric vector. Perform structural fit calculation: Preferably, for the generated image Semantic segmentation is performed to obtain a binary mask of the slipper region, and edge detection is used to extract the outer contour point set. Further optimization involves establishing a regression model that maps shoe last geometric vectors to standard contours. ,in Given an ideal contour point set, calculate the average contour distance. Then, the distance is normalized to obtain the geometric fit score. ,in The scaling factor is set to [value] in this embodiment. .
[0024] The embedding space distance between the generated image vector and the constraint target vector is calculated in the joint embedding space, and a multimodal design target fitting loss is constructed by combining three scores. ,in , For embedding spatial distance items , For the reason and Derived rating loss item , The weighting coefficients are used to minimize the fitting loss of the multimodal design objective. To achieve this, dynamically update the structured prompt template. For the dynamic geometric parameter part and semantic description part, repeat step four until the convergence condition is met. Preferably, the dynamic update adopts a hybrid optimization strategy: gradient-based continuous optimization is used for the dynamic geometric parameter slots, specifically gradient descent update. ,in The learning rate and gradient optimization are obtained through backpropagation of the AIGC model's decoder. For discrete semantic slots (such as style, facet, and material), discrete optimization based on policy gradient can be used. The action space is defined as the set of candidate words for each slot, and the policy network... The current state s (which includes the generated image features and the current loss value) is used as input to select the output probability, and the policy parameters are updated through the REINFORCE algorithm.
[0025] The convergence condition can be set as follows: in multiple consecutive iterations The rate of change is lower than the preset threshold Or reach the preset maximum number of iterations In this embodiment, it is preferred to set , And calculate after each iteration ,when The iteration terminates when the condition is met three times consecutively, and the final result is output. .
[0026] Step 5: Geometrically register the converged rendering with the 3D shoe last file and generate a matching 3D slipper design file.
[0027] The final slipper design rendering after iterative convergence. Input shoe last 3D file Perform geometric registration and thickness mapping to generate a 3D design file for slippers that matches the shoe last.
[0028] Preferably, for Normal map decomposition was performed, and the sole thickness field was separated using a pre-trained decoupling network. With help in the field ; using the curved surface of the shoe last Using the reference (the shoe last surface is represented as a surface related to the geometry of the shoe last 3D file), the thickness field is offset along the surface normal n. With envelope field Then and Mesh fusion is performed to generate a complete 3D model of the slippers, where n represents the surface of the shoe last. The surface normal vectors at each point (usually unit normal vectors) are output as STL or STEP format slipper 3D design files, thereby achieving automated conversion from semantic intent to 3D engineering files that match the shoe last structure.
[0029] Example 2: A smart slipper design system based on AIGC, specifically including the following: A joint embedding space building block is used to build, train, and store semantic encoders, geometric encoders, and image encoders. The shoe last parsing and target fusion module is used to receive shoe last 3D files and design target text, extract shoe last geometric vectors and semantic vectors and fuse them to generate constraint target vectors; The parameterized prompt word engine module contains a large language model for generating and iteratively updating structured prompt word templates containing fixed semantic slots and dynamic geometric parameter slots; The iterative optimization control module is used to input the prompt word template into the generative AIGC model to generate a slipper design rendering. It uses a semantic discriminator, a geometric discriminator, and an aesthetic discriminator to output semantic fit, geometric fit, and aesthetic score for the rendering. The generated image vector is obtained by the image encoder. Based on the distance between the generated image vector and the constraint target vector in the joint embedding space and combined with the score, a fitting loss is constructed. The prompt word template is iteratively updated until convergence. The 3D fitting modeling module is used to geometrically register the converged rendering with the 3D shoe last file and output a matching 3D slipper design file.
[0030] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
[0031] The preferred embodiments of the present invention disclosed above are merely illustrative of the invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the invention to specific implementations. Clearly, many modifications and variations can be made based on the content of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the invention, thereby enabling those skilled in the art to better understand and utilize the invention. The invention is limited only by the claims and their full scope and equivalents.
Claims
1. A smart design method for slippers based on AIGC, characterized in that, include: Step 1: Construct a joint embedding space model containing a semantic encoder, a geometric encoder, and an image encoder; Step 2: Obtain the user's design target text and the 3D file of the shoe last, extract the geometric vector of the shoe last and encode the target text into a semantic vector, and fuse them to obtain the constraint target vector; Step 3: The large language model generates a structured prompt word template containing fixed semantic slots and dynamic geometric parameter slots based on the constraint target vector; Step 4: Input the prompt word template into the generative AIGC model to generate a slipper design rendering. Use a semantic discriminator, a geometric discriminator, and an aesthetic discriminator to output semantic fit, geometric fit, and aesthetic score for the rendering. The generated image vector is obtained by the image encoder. Based on the distance between the generated image vector and the constraint target vector in the joint embedding space and combined with the score, construct a fitting loss and iteratively update the prompt word template until convergence. Step 5: Geometrically register the converged rendering with the 3D shoe last file and generate a matching 3D slipper design file.
2. The AIGC-based intelligent slipper design method according to claim 1, characterized in that, The joint embedding space model is trained using triple contrastive learning, wherein the triple includes a design semantic description, a corresponding shoe last geometric vector, and a slipper design image that matches the semantic description and the shoe last.
3. The AIGC-based intelligent slipper design method according to claim 1, characterized in that, The extraction of the shoe last geometric vector includes: performing mesh sampling on the shoe last 3D file, extracting principal component analysis coefficients, radial basis function descriptors, and curvature distribution of key sections, and splicing them to form a geometric feature vector.
4. The AIGC-based intelligent slipper design method according to claim 1, characterized in that, The structured prompt word template includes at least one dynamic geometric parameter slot, the value of which is associated with at least one dimension of the shoe last geometric vector and allows for continuous updates during iterative optimization.
5. The AIGC-based intelligent slipper design method according to claim 1, characterized in that, The geometric discriminator includes a geometric constraint calculation module, which is used to extract the slipper outline or implicit geometric features from the slipper design rendering and perform structural fit calculation with the shoe last geometric vector to obtain a geometric fit score.
6. A smart slipper design system based on AIGC, applied to the method described in any one of claims 1-5, characterized in that, include: A joint embedding space building block is used to build, train, and store semantic encoders, geometric encoders, and image encoders. The shoe last parsing and target fusion module is used to receive shoe last 3D files and design target text, extract shoe last geometric vectors and semantic vectors and fuse them to generate constraint target vectors; The parameterized prompt word engine module contains a large language model for generating and iteratively updating structured prompt word templates containing fixed semantic slots and dynamic geometric parameter slots; The iterative optimization control module is used to input the prompt word template into the generative AIGC model to generate a slipper design rendering. It uses a semantic discriminator, a geometric discriminator, and an aesthetic discriminator to output semantic fit, geometric fit, and aesthetic score for the rendering. The generated image vector is obtained by the image encoder. Based on the distance between the generated image vector and the constraint target vector in the joint embedding space and combined with the score, a fitting loss is constructed. The prompt word template is iteratively updated until convergence. The 3D fitting modeling module is used to geometrically register the converged rendering with the 3D shoe last file and output a matching 3D slipper design file.