A shoe and shoe style generation method based on multi-source data fusion and visual style

By using multi-source data fusion and visual style methods, the problem of unstable mapping of style information to component-level spatial style constraints in footwear style generation is solved, achieving style consistency and clear spatial constraints among components, and ensuring structural stability and controllability of the generated results.

CN122133340APending Publication Date: 2026-06-02HANGZHOU HUILIMA ARTIFICIAL INTELLIGENCE TECHNOLOGY CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HANGZHOU HUILIMA ARTIFICIAL INTELLIGENCE TECHNOLOGY CO LTD
Filing Date
2026-03-05
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

In the current technology for footwear style generation, style information is difficult to be stably mapped into component-level spatial style constraints, and the generated results lack a directional correction mechanism linked with component mask verification, resulting in insufficient coordination and consistency between components.

Method used

By collecting historical footwear images and texts, standardizing the images and texts, generating traceability identifiers, forming a trend sample set, and performing multimodal alignment, a style primitive library and style syntax diagram are constructed to obtain component masks and component style vectors. Spatial style planning diagrams and prompt diagrams are constructed to generate component blueprints. Visual features are extracted using component masks, and similarity is calculated as a consistency index. Local redrawing and verification are performed to achieve stable mapping and directional correction.

Benefits of technology

It achieves a stable mapping of multi-source graphic style information to component-level spatial style constraints, improving the consistency of styles among components and the clarity of spatial constraints, as well as the structural stability and controllability of the generated results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122133340A_ABST
    Figure CN122133340A_ABST
Patent Text Reader

Abstract

This invention discloses a footwear style generation method based on multi-source data fusion and visual style, relating to the field of computer vision technology. The method includes: performing multimodal alignment on a trend sample set; constructing a style primitive library and a style syntax graph; obtaining component masks and component style vectors; constructing a spatial style planning graph and a cue graph to form a component blueprint; generating a structural control graph based on the component blueprint; generating candidate graphs through controlled diffusion under the constraints of the structural control graph, component mask, and cue graph; extracting component visual features based on the component mask and calculating the similarity with the component style vectors as a component consistency index; selecting the candidate graph with the highest component consistency index and locally redrawing the component with the lowest consistency index to obtain a partially redrawn style image; performing semantic segmentation on the partially redrawn style image to obtain a prediction mask, which is then verified against the component blueprint; and outputting the final style image and the prediction mask.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer vision technology, and in particular to a method for generating footwear styles based on multi-source data fusion and visual patterns. Background Technology

[0002] Footwear style generation and digital design technology is an important direction for the integrated development of footwear industrial design, computer vision, and image processing. With the increasing iteration frequency of footwear products, the segmentation of consumption scenarios, and the acceleration of online content dissemination, enterprises are increasingly relying on digital methods to collect style references, analyze styles, and generate solutions. Existing technologies typically build a design data foundation based on product images and text, brand historical style data, publicly available design materials, and market communication content. Through image preprocessing, text cleaning, feature representation, style retrieval, similar style analysis, and condition generation, an auxiliary design process is formed. Related methods can establish style descriptions in dimensions such as color matching, texture, outline, and decoration, and generate style diagrams by combining component partitioning, template constraints, or interactive editing for trend analysis, series extension, solution comparison, and design communication. In recent years, the application of image-text joint modeling and generative image technology has further improved the efficiency of footwear design expression and digital collaboration capabilities.

[0003] However, conventional methods still have two limitations in key technical aspects. The style information is mainly represented by the overall style, which is difficult to stably map into the spatial style constraints of components such as the upper body, side panels, and outsole. This results in insufficient coordination and consistency between components. The verification and correction of the generated results mostly adopts the post-screening method, which lacks a directional correction mechanism linked with the verification of component masks, affecting the stability of the results and the efficiency of iteration. Summary of the Invention

[0004] In view of the aforementioned existing problems, the present invention is proposed.

[0005] Therefore, this invention provides a footwear style generation method based on multi-source data fusion and visual style, which solves the problems that multi-source graphic style information is difficult to stably map into component-level spatial style constraints and that the generated results lack a directional correction mechanism linked with component mask verification.

[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution: This invention provides a method for generating footwear styles based on multi-source data fusion and visual style, comprising: Collect historical images and texts of footwear, standardize the images and texts, generate traceability identifiers, and form a trend sample set.

[0007] Multimodal alignment is performed on the trend sample set, a style primitive library and style syntax diagram are constructed, component masks and component style vectors are obtained, a spatial style planning diagram and a hint diagram are constructed, and a component blueprint is formed.

[0008] Based on the component blueprint, a structural control diagram is generated. Under the constraints of the structural control diagram, component mask, and hint diagram, candidate diagrams are generated through controlled diffusion. The visual features of the components are extracted from the component mask and the similarity is calculated with the component style vector as the component consistency index. The candidate diagram with the largest component consistency index is selected and the component with the smallest component consistency index is locally redrawn to obtain the style diagram after local redrawing.

[0009] Semantic segmentation is performed on the partially redrawn style drawing to obtain a prediction mask, which is then compared with the component blueprint to output the final style drawing and prediction mask.

[0010] As a preferred embodiment of the footwear style generation method based on multi-source data fusion and visual style described in this invention, the steps of collecting historical footwear images and text, and standardizing the images and text are as follows: collecting footwear images and text and writing each collection record into the source field; extracting frames from the video content at equal intervals; performing grayscale conversion and edge extraction on each frame; performing a closing operation on the edges and filling holes to obtain a closed foreground region; taking the closed foreground region with the largest area as the shoe body region and calculating the minimum bounding rectangle of the shoe body region; taking the minimum bounding rectangle as the bounding region of the shoe body; taking the intersection of the two diagonals as the center of the bounding region; and taking the frame with the smallest distance from the center of the bounding region to the geometric center of the image as the footwear image.

[0011] Extract the foreground boundary of the shoe body area to form the outer contour of the sole. Based on the shoe body area, unify the toe direction, crop and expand, scale and fill the square canvas, and save the color space without loss. Concatenate the text in the order of title, description and tag, and perform noise reduction and synonym normalization to obtain the normalized image and normalized text.

[0012] As a preferred embodiment of the footwear style generation method based on multi-source data fusion and visual style described in this invention, the specific steps for generating traceability identifiers and forming a trend sample set are as follows: extracting keyword sequences from the normalized text according to the footwear type terminology, component terminology, appearance and decoration terminology, color scheme terminology, and scene terminology; generating a source summary, image summary, and text summary for each collection record based on the source field, normalized image, and normalized text; and generating a traceability identifier.

[0013] Based on the source identification, duplicates are removed, and the normalized image, normalized text, source identification, and source field are written into the trend sample set. For duplicate records, the collection record with the higher resolution of the normalized image is retained first.

[0014] As a preferred embodiment of the footwear style generation method based on multi-source data fusion and visual style described in this invention, the steps of performing multimodal alignment on the trend sample set and constructing a style primitive library and style syntax graph are as follows: mapping the normalized image and normalized text of each record in the trend sample set to the same style attribute dimension; extracting image-side style vectors and text-side style vectors based on binning tables, appearance and decoration word lists and scene word lists; generating style representations by taking the largest value of each item according to the same index; and binding them with source identifiers to form a trend sample set with style representations.

[0015] Clustering is performed on the trend sample set with pattern representations, using real samples as representative points to form a pattern primitive library. Each record is assigned to the set of representative points most similar to the pattern representation. For each set of representative points, the real record with the highest average cosine similarity to the other pattern representations in the set is selected as the representative sample of the set. A pattern syntax graph is constructed based on the pattern representations and the pattern primitive library. The pattern primitive set is selected by sorting by similarity. The co-occurrence intensity is determined by the ratio of the number of records in the same record pattern primitive set where two pattern primitives appear simultaneously to the number of records where at least one pattern primitive appears.

[0016] As a preferred embodiment of the footwear style generation method based on multi-source data fusion and visual style described in this invention, the specific steps for obtaining component masks and component style vectors are as follows: for the normalized image in the trend sample set, extract the outer contour of the shoe body based on the foreground boundary of the shoe body region and calculate the outer circumscribed region of the shoe body; perform connected component decomposition and boundary tracking within the shoe body region with the outer circumscribed region of the shoe body as a constraint to obtain candidate component regions; and group them into the main body of the upper, side plate, toe, heel, outsole, logo and shoelace according to relative position and geometric shape to obtain a set of binary component masks of the same size.

[0017] For representative samples of each style primitive in the style primitive library, crop the component region image based on the component mask coverage area and set the non-mask-covered pixels as a solid color background. Extract and normalize the component visual features according to the binning table, generate the component style vector of each component, and bind and save it with the source identifier of the representative sample.

[0018] As a preferred embodiment of the footwear style generation method based on multi-source data fusion and visual style described in this invention, the specific steps of constructing a spatial style planning diagram and a prompt diagram to form a component blueprint are as follows: determining the style primitives of the main body of the shoe upper in the style primitive library according to the standardized text and keyword sequence of the style to be generated, and selecting style primitives with greater co-occurrence intensity with the style primitives of the main body of the shoe upper in the style syntax diagram for the remaining components in the component order.

[0019] Based on the component mask set, the component style vector is written into the component region according to the component to which the pixel belongs, forming a spatial style planning map with the same resolution as the normalized image. The component region image of the sample represented by the selected style primitive is filled according to the shape of the component mask to form a hint map. The component mask, component style vector, spatial style planning map and hint map are bound together to obtain the component blueprint.

[0020] As a preferred embodiment of the footwear style generation method based on multi-source data fusion and visual style described in this invention, the step of extracting visual features of components based on component masks and calculating similarity with component style vectors as component consistency index specifically involves: extracting the outer contour line of the shoe body, the outer contour line of the shoe sole, and the component boundary line based on the component mask set in the component blueprint and superimposing them to generate a black and white binary structure control map.

[0021] After aligning the structural control chart, component mask set, and prompt image, the controlled diffusion is used as the input for morphological constraints, spatial constraints, and appearance guidance. A candidate image set is generated based on a random seed generated by the traceability identifier. The component region image is cropped according to the component mask of the candidate image set and the component visual features are extracted. The cosine similarity between the component visual features and the component style vector is calculated as the component consistency. The minimum value of the consistency of all components in each candidate image is used as the component consistency index. The candidate image with the largest component consistency index is selected as the optimal candidate image, and the component with the smallest component consistency in the optimal candidate image is selected as the weakest component.

[0022] As a preferred embodiment of the footwear style generation method based on multi-source data fusion and visual style described in this invention, the steps of taking the candidate image with the largest component consistency index and performing local redrawing on the component with the smallest component consistency index to obtain the style image after local redrawing are as follows: taking the optimal candidate image as the base image, defining the local redrawing area with the local redrawing area mask after expanding the component mask of the weakest component, and performing controlled diffusion local redrawing based on a random seed generated by the traceability identifier and the name of the weakest component under the constraints of the structural control chart, the component mask set, and the prompt image, generating a set of local redrawing candidate images, repeatedly calculating the component consistency index on the set of local redrawing candidate images and selecting the local redrawing candidate image with the largest component consistency index, fusing it with the optimal candidate image according to the local redrawing area mask to obtain the style image after local redrawing, and binding and saving it with the component blueprint.

[0023] As a preferred embodiment of the footwear style generation method based on multi-source data fusion and visual style described in this invention, the semantic segmentation of the locally redrawn style image specifically involves: unifying the direction of the locally redrawn style image according to the outer contour of the sole and the direction of the toe; generating candidate component regions by grayscale conversion, edge extraction, closing operation, hole filling, shoe body region extraction, connected component decomposition, and boundary tracking; and merging them to obtain a set of binary prediction masks of the same size.

[0024] As a preferred embodiment of the footwear style generation method based on multi-source data fusion and visual style described in this invention, the specific steps of obtaining the prediction mask and verifying it with the component blueprint are as follows: calculating the mask consistency degree based on the binary prediction mask set and the component mask set, taking the component with the smallest mask consistency degree as the verification component, performing consistency verification on all components, if the prediction mask set is completely consistent with the component mask set of the component blueprint, then the style image after partial redrawing is determined as the final style image; otherwise, correction is performed and semantic segmentation is repeated, and the style image after correction is determined as the final style image.

[0025] The beneficial effects of this invention are as follows: by constructing a spatial style planning diagram and a prompt diagram and forming a component blueprint, a stable mapping of multi-source graphic style information to component-level spatial style constraints is achieved, thus achieving style consistency and clear spatial constraints among components. By extracting the visual features of components based on the component mask and calculating the similarity with the component style vector as the component consistency index, the candidate image with the largest component consistency index is selected and the component with the smallest component consistency index is locally redrawn, thereby achieving linkage between generation and directional correction, thus achieving structural stability and controllable results. Attached Figure Description

[0026] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0027] Figure 1 This is a flowchart of a footwear style generation method based on multi-source data fusion and visual style.

[0028] Figure 2 Flowchart for constructing the trend sample set.

[0029] Figure 3 A flowchart for generating style primitives and component masks.

[0030] Figure 4 Optimize the flowchart for style drawing generation and review.

[0031] Figure 5 A comparison chart showing the impact of grid granularity on boundary crossing rate in spatial pattern planning.

[0032] Figure 6 This is a diagram showing the relationship between structural stability and style consistency. Detailed Implementation

[0033] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0034] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.

[0035] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.

[0036] Reference Figures 1-6 This is one embodiment of the present invention, which provides a method for generating footwear styles based on multi-source data fusion and visual styles, including the following steps: S1. Collect historical footwear images and text, standardize the images and text, generate traceability identifiers, and form a trend sample set.

[0037] Footwear content is retrieved sequentially from e-commerce, social media, competitor product release channels, and brand historical product databases, and footwear images and text are saved one by one. The text includes title, description, and tag information.

[0038] When the footwear content is a video, multiple frames are extracted from the video at equal time intervals. Each frame is then converted to grayscale and edge-trimmed. A closing operation is performed on the edges, and hole filling is performed to obtain a closed foreground region. Hole filling continues until there are no holes inside the closed region. Within the closed foreground region, closed foreground regions that are in contact with the image boundary are first removed. The closed foreground region with the largest area is taken as the footwear region, and the minimum bounding rectangle of the footwear region is calculated. The intersection of the two diagonals of the minimum bounding rectangle is taken as the center of the bounding region. The distance from the center of the bounding region to the geometric center of the image is calculated, and the frame with the smallest distance is taken as the footwear image. The geometric center of the image is the pixel position formed by the midpoint of the image width direction and the midpoint of the image height direction. If there are two frames with the smallest area, the frame with the largest footwear region area is taken as the footwear image, and the text corresponding to the current frame is saved.

[0039] Write a source field to each collection record. The source field includes the source platform name, page link, content identifier, publication time, and collection time. Then bind and save the source field with the current collection record.

[0040] Each frame of the footwear image is converted to grayscale and edge extracted. The edges are closed and holes are filled to obtain a closed foreground region. The closed foreground region with the largest area is taken as the footwear region. The minimum bounding rectangle of the footwear region is calculated as the bounding region of the footwear. If multiple footwear regions exist in the same footwear image, resulting in multiple closed foreground regions side by side, the distance from the center of each bounding region to the geometric center of the image is calculated. The footwear region with the smallest distance and the largest bounding region area is taken as the target footwear region. The image is cropped according to the bounding region of the target footwear region. The image is expanded outward by the same pixel distance around the bounding region to avoid cutting off the outer edge of the sole, resulting in the cropped footwear image.

[0041] The same pixel distance is the outward pixel distance, which is determined by rounding the product of the number of pixels on the short side of the smallest outer rectangle of the area to be expanded and the outward expansion ratio, and the rounded result is limited to the minimum outward pixel distance and the maximum outward pixel distance.

[0042] Extract the foreground boundary of the shoe body area to form the outer contour of the sole. The outer contour of the sole is obtained by taking the set of boundary points with the largest vertical coordinate in the foreground boundary and connecting them according to the horizontal coordinate. Compare the sharpness of the curvature at both ends along the length of the shoe body, take the sharper end as the direction of the toe, and rotate the toe to a uniform direction.

[0043] The cropped shoe image is scaled proportionally and filled onto a square canvas. The filled area uses a solid color background, and the canvas size and filling method are kept consistent to obtain a normalized image.

[0044] Standardized images are saved using an image encoding format that is lossless and uses a uniform color space, ensuring that the saved parameters remain unchanged during the acquisition process.

[0045] For each text entry, the title, description, and tags are concatenated into a single text segment in sequence, with the title first, the description centered, and the tags last, and connected by a whitespace separator.

[0046] Remove emoticons, advertising slogans, irrelevant link fragments, and repetitive phrases; unify synonyms into a unified expression; unify consecutive whitespace into a single whitespace; unify line breaks into a single line break representation; and unify capitalization and punctuation formats to obtain standardized text.

[0047] Synonym unification is performed based on the synonym mapping table, which is a mapping table that covers high-frequency words related to shoe type, components, appearance, decoration, color scheme and scene.

[0048] The keyword sequences related to shoe type, components, appearance, decoration, color scheme, and scene are retained from the normalized text and saved together with the complete normalized text. Keyword extraction only retains terms that match the shoe type vocabulary, component vocabulary, appearance and decoration vocabulary, color scheme vocabulary, and scene vocabulary to obtain the keyword sequence. The normalized text is encoded in a unified character encoding format.

[0049] The shoe type vocabulary, component vocabulary, appearance and decoration vocabulary, color scheme vocabulary, and scene vocabulary are constructed from historical shoe images and text and source fields. The collected titles, descriptions, tag information, source platform category attribute fields, and brand historical model naming information are segmented and phrases are extracted. After removing irrelevant terms, they are categorized by shoe type, component, appearance and decoration, color scheme, and scene category. The categorization results are normalized by synonym expression and duplicate terms are merged. They are then sorted by frequency of occurrence. If the frequency of occurrence is the same, they are sorted by the time of collection to form the corresponding vocabulary.

[0050] For each collected record, the source field is concatenated to form a source summary. After the normalized image is saved as an image file in a lossless format, the entire content of the image file is read and processed using a publicly available standard digest algorithm (such as SHA-256) to obtain an image summary. After the normalized text is saved as a text file using a unified character encoding, the entire content of the text file is read and processed using a publicly available standard digest algorithm to obtain a text summary.

[0051] The source summary, image summary, and text summary are concatenated in sequence, and the source identifier is obtained by using a publicly available standard summary algorithm. The same source identifier is generated for the same normalized image and normalized text under the same source field. The concatenation order is source field first, image summary in the middle, and text summary last.

[0052] If the traceability identifier already exists, it is determined to be a duplicate. If the traceability identifier does not exist, the normalized image, normalized text, traceability identifier, and source field are written into the trend sample set. When the same traceability identifier corresponds to multiple duplicate collection records, the collection record with the higher resolution of the normalized image is retained, and the source field is appended to the traceability list of the traceability identifier. When the normalized image resolution of duplicate collection records is the same, the collection record with the earlier collection time is retained, and the remaining source fields are appended to the traceability list.

[0053] The trend sample set includes normalized images, normalized text, traceability identifiers, and traceability lists.

[0054] S2. Perform multimodal alignment on the trend sample set, construct a style primitive library and style syntax diagram, obtain component mask and component style vector, construct a spatial style planning diagram and hint diagram, and compose a component blueprint.

[0055] The normalized image and normalized text of each record in the trend sample set are mapped to the same style attribute dimension. The style attribute dimension includes color attribute, texture attribute, outline attribute, decoration attribute, and scene attribute. The index terms of scene attribute and decoration attribute come from the scene vocabulary and appearance and decoration vocabulary, respectively. The index terms of color attribute, texture attribute, and outline attribute are given by the binning table, which is the color binning table, texture binning table, and outline binning table, respectively, so that each record is filled with the style representation.

[0056] The color binning table is generated by statistical analysis of the color distribution of the shoe body region in the trend sample set; the texture binning table is generated by statistical analysis of the frequency of local binary texture patterns in the shoe body region in the trend sample set; and the contour binning table is generated by statistical analysis of the gradient direction distribution in the shoe body region in the trend sample set.

[0057] Image style features are extracted from the normalized image and filled into the style attribute dimension. For color attributes, the shoe area is transformed to the same color space as the unified color space. The color distribution of the shoe area is statistically analyzed according to the color binning table and normalized to obtain color features. For texture attributes, the shoe area is grayscaled. The distribution of local binary texture patterns is statistically analyzed according to the texture binning table and normalized to obtain texture features. For contour attributes, the edges of the shoe area are extracted, and the gradient direction distribution is statistically analyzed according to the contour binning table and normalized to obtain contour features. For decorative attributes, the shoe area is divided into regular grids. The edge density within each grid is statistically analyzed, and decorative features are written according to the decorative index items corresponding to the appearance and decorative vocabulary to obtain decorative feature vectors. For scene attributes, no new scene information is added on the image side; the scene attribute remains a zero vector on the image side and is filled by the text side.

[0058] Each decoration index entry in the appearance and decoration vocabulary is bound to a set of regular grid positions. When writing decoration features, the edge density of the regular grid positions is averaged and written into the decoration index entry.

[0059] The color scheme attribute features, texture attribute features, outline attribute features, decoration attribute features, and scene attribute features are concatenated into the image-side style vector of the current record, and the image-side style vector is uniformly normalized and aligned with the text-side in terms of numerical scale.

[0060] Text style features are extracted from normalized text and keyword sequences and populated into the style attribute dimension. If the keyword belongs to the color scheme vocabulary, it is recorded in the corresponding color scheme bin. If the keyword belongs to the appearance and decoration vocabulary, it is recorded in the corresponding decoration index item. If the keyword belongs to the scene vocabulary, it is recorded in the corresponding scene index item. Texture and outline attributes are populated on the text side by vocabulary matching. When the keyword matches a texture-related term, it is written to the corresponding texture bin. When the keyword matches an outline-related term, it is written to the corresponding outline bin. If no match is found, the value remains zero.

[0061] The text-side statistical results are concatenated into a text-side style vector in an index order that is completely consistent with that of the image side, and the text-side style vector is uniformly normalized.

[0062] The image-side style vector and the text-side style vector are aligned item by item according to the same index, and the style representation of the current record is obtained by taking the largest value of each item. The fused style representation is bound to the source identifier and written back to the trend sample set to obtain the trend sample set with style representation.

[0063] Clustering is performed on the trend sample set with pattern representation, using real samples as representative points to form a pattern primitive library. In the cluster initialization phase, the record with the earlier collection time is used as the initial representative point, and the remaining records are used as candidate representative points. The representative point set is supplemented in order of less similarity to the selected representative point set until the number of pattern primitives is reached. If there is a tie for the smallest number, the record with the earlier collection time in the traceability list is selected as the new representative point. In the cluster iteration phase, each record is assigned to the set of representative points that are most similar to the pattern representation. In the representative point update phase, for each representative point set, the real record with the largest average cosine similarity to the other pattern representations in the set is selected as the representative sample of the set. If there is a tie, the real record with the earlier collection time is selected.

[0064] "Less similarity" means that the minimum cosine similarity between the candidate representative point style representation and the selected representative point style representation is used as the priority selection criterion.

[0065] The number of style primitives refers to how many reusable style primitives extracted from trend samples are in the style primitive library.

[0066] A style syntax graph is constructed based on a style primitive library. For each record in the trend sample set, the style representation is compared with the style representations of representative samples in the style primitive library using cosine similarity. The style primitives with the highest similarity ranking and the number of candidates are selected to form the style primitive set for the current record. The style primitive set is then bound and saved with the current record's source identifier. If the similarity ranking is tied at the boundary, the style primitive set is prioritized based on the earlier collection time of the representative sample. For any two style primitives, the number of records in the style primitive set where both style primitives appear simultaneously in the same record in the trend sample set is counted. Then, the number of records in the style primitive set where at least one of the two style primitives appears in the same record is counted. The ratio of the number of records in the style primitive set where both style primitives appear simultaneously in the same record to the number of records in the style primitive set where at least one of the two style primitives appears in the same record is taken as the co-occurrence strength of the two primitives. The co-occurrence strength is used as the degree of association between the two primitives in the style syntax graph. The style syntax graph and the style primitive library are bound and saved together in the source list.

[0067] The number of candidates refers to the integer obtained by taking the square root of the total number of style primitives in the style primitive library and rounding it up.

[0068] For each normalized image in the trend sample set, a component mask is generated. The component dimensions are: upper body, side panel, toe, heel, outsole, logo, and shoelaces. The outer contour of the shoe body is extracted based on the foreground boundary of the shoe body region, and the inscribed region of the shoe body is calculated. Connected component decomposition and boundary tracking are performed within the shoe body region using the inscribed region of the shoe body as a constraint to obtain candidate component regions. The candidate component regions are then merged into upper body, side panel, toe, heel, outsole, logo, and shoelaces in order of relative position and geometric shape. The merged result is converted into a binary image of the same size as the normalized image. The component mask set is used, and each pixel belongs to only one component mask. When there is boundary overlap between candidate component areas, the areas that are closer to the outer contour of the shoe body and the outer contour of the sole are prioritized to be grouped into the outsole. The remaining areas are then grouped into the toe, heel, upper body and side panels in order of their relative position to the toe direction. The candidate component area with the smallest area located in the surface area of ​​the upper body is grouped into the identifier. The candidate component area that is long and thin and spans the upper body is grouped into the shoelace. The obtained component mask set is bound and saved with the traceability identifier of the image.

[0069] "Closer to the outer contour of the shoe body" and "closer to the outer contour of the sole" refer to using the smaller minimum distance from the boundary of the candidate component area to the corresponding outer contour boundary as the criterion.

[0070] "Crossing the main body of the shoe upper" refers to using the overlap between the candidate component area and multiple separate locations within the main body of the shoe upper as the criterion for judgment.

[0071] The term "slender strip" refers to the condition where the ratio of the longest side to the shortest side of the smallest bounding rectangle of the candidate component region is the largest within the candidate component region.

[0072] For each style primitive in the style primitive library, a normalized image of the representative sample and a set of component masks are taken. For each component mask-covered area, the component region image is obtained by cropping the minimum bounding rectangle of the component mask-covered area. The pixels not covered by the mask are set to a solid color background. The component visual features are extracted from the component region image. The component visual features are formed by splicing the color histogram, the orientation gradient histogram, and the local texture statistics. The color histogram is statistically analyzed according to the color matching binning table, the orientation gradient histogram is statistically analyzed according to the contour binning table, and the local texture statistics are statistically analyzed according to the texture binning table. The splicing results are uniformly normalized. The normalized component visual features are used as the component style vector of the style primitive on the current component and bound to the representative sample traceability identifier for storage.

[0073] When constructing component blueprints for the style to be generated, after determining the selected style primitives for each component according to the style syntax diagram, the component style vectors corresponding to each component are read, and the component style vectors are written into the corresponding component regions according to the component to which the pixel belongs, based on the component mask set, forming a spatial style planning diagram with the same resolution as the normalized image. Each pixel position is only written with the component style vector corresponding to its own component. The process of determining the selected style primitives for each component is as follows: based on the normalized text and keyword sequence of the style to be generated, the style primitives most similar to the normalized text in the style primitive library are selected as the style primitives of the upper body. Then, in the order of components, the style primitives with greater co-occurrence intensity with the style primitives of the upper body in the style syntax diagram are selected for the remaining components. If there is a tie, the style primitive that matches the keyword sequence of the style to be generated better is selected. If they are still tied, the style primitive with an earlier acquisition time is selected. The corresponding component region image of the sample represented by the selected style primitive is used as a texture block and filled into the corresponding position according to the shape of the component mask to form a prompt image. The component mask, component style vector, spatial style planning diagram and prompt image are bound together to obtain the component blueprint.

[0074] Figure 5 Used to characterize the effect of spatial style planning diagrams on component-level spatial style constraints. Figure 5 The horizontal axis represents the grid granularity of the spatial style planning diagram. The grid granularity refers to the number of grid divisions along one side of the normalized image canvas. Specifically, under a uniform canvas size, the normalized image canvas is divided into the same number of grid intervals along both the width and height directions to form a regular grid. A larger number of grid divisions in one side indicates a smaller coverage area for each individual grid and a finer spatial style planning diagram; conversely, a smaller number of grid divisions in one side indicates a larger coverage area for each individual grid and a coarser spatial style planning diagram. Comparisons under different grid granularity conditions are performed after the structural control diagram, hint diagram, and component mask set have been normalized to the same canvas size. This ensures that changes in the grid granularity of the spatial style planning diagram only reflect changes in the fineness of component-level spatial style constraints and are not affected by differences in canvas size. Figure 5 The ordinate represents the mean out-of-bounds rate. The style images to be analyzed are grouped by component type to obtain a predicted mask set. This predicted mask set is then matched one-to-one with the component mask sets in the component blueprints, categorized by upper body, side panel, toe, heel, outsole, logo, and laces. For any component, the out-of-bounds rate is defined as the proportion of pixels in the component's predicted mask that fall outside the corresponding component blueprint's component mask, relative to the total number of pixels in the component's predicted mask. The mean out-of-bounds rate for a single style image is defined as the arithmetic mean of the out-of-bounds rates for all valid components. Figure 5 The average boundary violation rate shown is the sample average of the average boundary violation rates of multiple style drawings to be statistically analyzed under the same grid granularity. Figure 5 The curves in the diagram correspond to four processing groups: no-part blueprint, part-part blueprint, part consistency index optimization, and local redrawing of the weakest part. Figure 5 As can be seen, with the increase of the grid granularity of the spatial style planning diagram, the average out-of-bounds rate of each group decreases overall. The processing group using component blueprints is lower than the processing group without component blueprints, and the processing group with the weakest component local redrawing is the lowest. This indicates that by constructing spatial style planning diagrams and prompt diagrams and forming component blueprints, multi-source graphic style information is mapped to component-level spatial style constraints more stably, reducing style content out-of-bounds across components. At the same time, the candidate diagram selection based on the component consistency index and the local redrawing of the weakest component further suppress local out-of-bounds, demonstrating the improvement of the present invention in terms of spatial constraint clarity.

[0075] S3. Generate a structural control diagram based on the component blueprint. Under the constraints of the structural control diagram, component mask, and hint diagram, generate candidate diagrams through controlled diffusion. Extract the visual features of the components based on the component mask and calculate the similarity with the component style vector as the component consistency index. Select the candidate diagram with the largest component consistency index and perform local redrawing on the component with the smallest component consistency index to obtain the style diagram after local redrawing.

[0076] The component mask set is read from the component blueprint, and the union of the component masks is taken to obtain the shoe body region mask. The outer contour of the shoe body is extracted based on the shoe body region mask. Boundary tracing is performed on the foreground boundary of the shoe body region mask to obtain the outer contour line of the shoe body. Boundary tracing adopts the eight-neighbor connectivity boundary tracing rule and outputs a contour line with a single pixel width. If multiple closed contours are obtained by boundary tracing, the closed contour with the largest area is selected as the outer contour line of the shoe body. The set of boundary points with the largest vertical coordinate in the foreground boundary of the shoe body region is selected and connected according to the horizontal coordinate to form the outer contour line of the sole. The component parts are extracted based on the component mask set. For each component mask, boundary tracing is performed to obtain the boundary lines of each component. All component boundary lines are merged into a component boundary line layer. When different component boundary lines overlap at the same pixel position, the current pixel position is retained as the boundary line pixel to ensure that the boundary line is closed. The outer contour line of the shoe body, the outer contour line of the shoe sole, and the component boundary line layers are superimposed on a canvas of the same size as the normalized image to obtain a structural control map. The structural control map is saved in a lossless image format and bound to the traceability identifier. The structural control map is a black and white binary image, with the contour line and boundary line pixels set to white and the remaining pixels set to black.

[0077] The system reads the hint image and component mask set from the same component blueprint and aligns them with the structural control diagram at the same resolution. If there are resolution differences between the hint image, component mask set, and structural control diagram, they are scaled proportionally and filled into a square canvas. After normalizing to the same canvas size, the system inputs the controlled diffusion algorithm. The structural control diagram is used as the morphological constraint input, the component mask as the spatial constraint input, and the hint image as the appearance guidance input. The system then enters the controlled diffusion generation stage to obtain a candidate image set. The random seed of the controlled diffusion is obtained by converting the source identifier corresponding to the component blueprint using a publicly available standard digest algorithm. The same source identifier corresponds to the same random seed, and the same component blueprint corresponds to the same candidate image set.

[0078] Controlled diffusion generation refers to initializing a random pixel image of the entire canvas with a random seed as the initial image, performing iterative denoising and updating. During iterative denoising and updating, the structure control map, component mask, and cue map are simultaneously input as conditions, so that the initial image gradually converges into candidate images. In each iteration of denoising and updating, only the output image of the current iteration is updated without changing the structure control map, component mask, and cue map.

[0079] For each candidate image in the candidate image set, crop the component region image part by part according to the component mask set, and set the non-mask-covered pixels as a solid color background. Extract the component visual features for each component region image. For each component in the same candidate image, use the cosine similarity between the current component visual features and the current component style vector as the consistency of the current component. Use the minimum consistency of all components in the current candidate image as the component consistency index of the current candidate image. Select the candidate image with the largest component consistency index as the optimal candidate image, and select the component with the smallest component consistency as the weakest component. The expression is: ; ; ; ; in, The candidate image number is Component number is The component consistency value, Indicates the candidate image number. Indicates the component number. The candidate image number is And the component number is The visual characteristics of the components Indicates the component number is component style vectors, The candidate image number is The component consistency index, Represents a set of components. Represents the set of candidate graphs. Indicates the optimal candidate graph number. Indicates the number of the weakest component. The component number in the optimal candidate graph is... The component consistency value.

[0080] When the consistency index of components is tied for the highest, the candidate graph generated earlier in the candidate graph set is selected as the optimal candidate graph. When the weakest component is tied for the lowest, the components are compared in the following order: upper body, side plate, toe, heel, outsole, logo, and laces, and the component that appears first is selected as the weakest component.

[0081] The optimal candidate image is used as the base image for local redrawing. The component mask corresponding to the weakest component is taken as the initial local redrawing region mask. The initial local redrawing region mask is expanded outward by the same pixel distance to obtain the local redrawing region mask. The base image is kept unchanged outside the local redrawing region mask. During local redrawing, the same structural control map is used as the shape constraint input, the same prompt image is used as the appearance guidance input, and the same set of component masks is used for spatial constraints to ensure that the weakest component after local redrawing is still consistent with the outer contour of the shoe body, the outer contour of the shoe sole, and the component boundary line, and is consistent with the component texture and color scheme expressed by the prompt image.

[0082] Under constraints, controlled diffusion local redrawing is performed to generate a set of candidate images for local redrawing. The random seed for local redrawing is formed by concatenating the source identifier and the name of the weakest component in sequence with a separator to form a seed input string. After converting the seed input string into a byte sequence using a unified character encoding, a digest byte sequence is obtained through a public standard digest algorithm. The first four bytes of the digest byte sequence are converted into big-endian unsigned integers to obtain the random seed for local redrawing, which remains unchanged during the local redrawing process. The controlled diffusion of local redrawing is to initialize a random pixel image with the random seed as the initial region within the mask coverage area of ​​the local redrawing region, and keep the pixels of the optimal candidate image unchanged outside the mask coverage area of ​​the local redrawing region. In the iterative denoising update, the structure control map, component mask, and prompt map are continuously input as conditions.

[0083] The component consistency index is repeatedly calculated for each candidate image in the partial redrawing candidate image set. The candidate image with the largest component consistency index is taken as the style image after partial redrawing. The candidate images and the best candidate image are merged according to the partial redrawing area mask to obtain the style image after partial redrawing. The style image after partial redrawing is bound and saved with the component blueprint. Pixels within the area covered by the partial redrawing area mask are taken from the candidate images, and pixels outside the area covered by the partial redrawing area mask are taken from the best candidate image. When the component consistency index of the candidate images is tied for the largest, the candidate image with the smallest candidate number in the partial redrawing candidate image set is taken as the style image after partial redrawing. The candidate numbers are sequentially increased from the starting number according to the generation order.

[0084] S4. Perform semantic segmentation on the partially redrawn style drawing to obtain the prediction mask and verify it with the component blueprint. Output the final style drawing and prediction mask.

[0085] Read the partially redrawn style drawing and the component mask set from the component blueprint bound to it. Extract the outer contour of the sole and determine the direction of the toe on the partially redrawn style drawing to unify the direction, rotating the toe to align with the component blueprint. Figure 1 To achieve a unified direction, the style image after direction unification is converted to grayscale and edge extraction is performed. Closure operations are applied to the edges, and hole filling is performed to obtain a closed foreground region. After removing closed foreground regions that touch the image boundary, the largest closed foreground region is selected as the shoe body region. The inscribed region of the shoe body is calculated, and connected component decomposition and boundary tracing are performed within the shoe body region using the inscribed region as a constraint to obtain candidate component regions. These candidate component regions are then sequentially merged into the main upper, side panels, toe, heel, outsole, logo, and laces. The merged result is then converted into a binary image of the same size as the style image. The set of binary component masks is used, and each pixel belongs to only one prediction mask. When there is boundary overlap between candidate component regions, the regions that are closer to the outer contour of the shoe body and the outer contour of the sole are prioritized to be merged into the outsole. The remaining regions are merged into the toe, heel, upper body and side plate in order of their relative position with respect to the toe direction. The candidate component regions with the smallest area and located in the surface area of ​​the upper body are merged into the identifier. The candidate component regions that are long and thin and cross the upper body are merged into the shoelaces. The resulting set of binary component masks is used as the prediction mask and bound to the style drawing after local redrawing.

[0086] Read the component mask set from the component blueprint, and match it with the predicted mask according to the component diameter of the upper body, side panel, toe, heel, outsole, logo, and shoelaces. Calculate the mask consistency for each component, and select the component with the lowest mask consistency as the verification component. The expression is: ; ; in, This indicates the verification component number with the lowest mask consistency. Indicates part number For mask consistency, The part number in the component blueprint is Component mask, The component number obtained from semantic segmentation is: The prediction mask.

[0087] When the mask consistency is the lowest and there is a tie, the components are compared in the following order: upper body, side panel, toe, heel, outsole, logo, and laces. The component that appears first is taken as the verification component.

[0088] Figure 6 The horizontal axis represents the minimum mask consistency (i.e., the minimum value among the mask consistency of each component, used to characterize the lower bound of structural stability), and the vertical axis represents the minimum component consistency index (i.e., the component consistency index obtained by extracting the visual features of the component according to the component mask and calculating the similarity with the component style vector, used to characterize the lower bound of style consistency). Figure 6 The different colored scatter points correspond to G0 blueprints without parts, G1 blueprints with parts, G3 part consistency index optimization, and G5 local redrawing of the weakest part, respectively. The straight lines of each color are the fitting lines for the corresponding groups. Figure 6 As can be seen, all fitted lines have positive slopes, indicating that as the minimum mask consistency improves, the minimum component consistency index also increases synchronously. This means that structural stability and style consistency exhibit a synergistic improvement relationship in this invention. The G3 and G5 sample points are more concentrated in the upper right region, indicating that after optimizing the component consistency index and locally redrawing the weakest component, generation and directional correction work in tandem. This ensures that the locally redrawn style drawing maintains the component blueprint constraints while also considering structural stability and style consistency, achieving controllable results. Figure 6 The graph is a scatter plot of fitted lines. The peak value is not used as the main criterion for interpretation. If the high-density clustering area of ​​the point cloud is used as the density peak for observation, the density peaks of G3 and G5 move to the upper right, which further illustrates that the present invention has a more obvious effect on the simultaneous improvement of structural stability and pattern consistency.

[0089] A pixel-level consistency check is performed on all components. If the predicted mask set is completely consistent with the component mask set of the component blueprint pixel by pixel, the locally redrawn style image is determined as the final style image, and the final style image and the predicted mask are output. If any component's predicted mask is not completely consistent with the component blueprint mask, a correction is performed. The locally redrawn style image is used as the base image, and the component blueprint mask corresponding to the checked component is used as the initial locally redrawn region mask. The locally redrawn region mask is obtained by expanding outward by the same pixel distance. The base image remains unchanged outside the locally redrawn region mask. At the same time, the generated structural control bound to the component blueprint is read. The process involves creating and prompting diagrams, and then performing controlled diffusion local redrawing under the constraints of structural control charts, component masks, and prompting diagrams. This generates a set of correction candidate diagrams, with candidate numbers increasing sequentially from the starting number. The corrected style diagram is selected based on the component consistency index. If multiple correction candidate diagrams have the same maximum component consistency index, the one with the smallest candidate number is chosen as the corrected style diagram. The corrected style diagram undergoes repeated semantic segmentation to obtain a new prediction mask, and this corrected style diagram is then determined as the final style diagram. The final style diagram and the new prediction mask are then output.

[0090] In summary, this invention achieves stable mapping of multi-source graphic style information to component-level spatial style constraints by constructing spatial style planning diagrams and prompt diagrams to form component blueprints, thus achieving style consistency and clear spatial constraints among components. By extracting component visual features based on component masks and calculating similarity with component style vectors as component consistency index, the candidate image with the largest component consistency index is selected and the component with the smallest component consistency index is locally redrawn, thereby achieving linkage between generation and directional correction, and achieving structural stability and controllable results.

[0091] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A method for generating footwear styles based on multi-source data fusion and visual style, characterized in that: include, Collect historical footwear images and text, standardize the images and text, generate traceability identifiers, and form a trend sample set; Multimodal alignment of the trend sample set is performed to construct a style primitive library and style syntax graph, obtain component mask and component style vector, construct spatial style planning graph and hint graph and compose component blueprint; Based on the component blueprint, a structural control diagram is generated. Under the constraints of the structural control diagram, component mask, and prompt diagram, candidate diagrams are generated through controlled diffusion. The visual features of the components are extracted from the component mask and the similarity is calculated with the component style vector as the component consistency index. The candidate diagram with the largest component consistency index is selected and the component with the smallest component consistency index is locally redrawn to obtain the style diagram after local redrawing. Semantic segmentation is performed on the partially redrawn style drawing to obtain a prediction mask, which is then compared with the component blueprint to output the final style drawing and prediction mask.

2. The footwear style generation method based on multi-source data fusion and visual style as described in claim 1, characterized in that: The specific steps for collecting historical footwear images and text, and standardizing the images and text, are as follows: Collect footwear images and text and write each collection record into the source field. Extract frames from the video content at equal intervals. Perform grayscale conversion and edge extraction on each frame image. Perform closing operation on the edges and fill holes to obtain a closed foreground area. Take the closed foreground area with the largest area as the footwear area and calculate the minimum bounding rectangle of the footwear area. Take the minimum bounding rectangle as the bounding area of ​​the footwear. Take the intersection of the two diagonals as the center of the bounding area. Take the frame with the smallest distance from the center of the bounding area to the geometric center of the image as the footwear image. Extract the foreground boundary of the shoe body area to form the outer contour of the sole. Based on the shoe body area, unify the toe direction, crop and expand, scale and fill the square canvas, and save the color space without loss. Concatenate the text in the order of title, description and tag, and perform noise reduction and synonym normalization to obtain the normalized image and normalized text.

3. The footwear style generation method based on multi-source data fusion and visual style as described in claim 2, characterized in that: The specific steps for generating traceability identifiers and forming a trend sample set are as follows: Keyword sequences are extracted from the normalized text according to the shoe type thesaurus, component thesaurus, appearance and decoration thesaurus, color matching thesaurus, and scene thesaurus. Based on the source field, normalized image, and normalized text, a source summary, image summary, and text summary are generated for each collection record, and a source traceability identifier is generated. Based on the source identification, duplicates are removed, and the normalized image, normalized text, source identification, and source field are written into the trend sample set. For duplicate records, the collection record with the higher resolution of the normalized image is retained first.

4. The footwear style generation method based on multi-source data fusion and visual style as described in claim 1, characterized in that: The specific steps for performing multimodal alignment on the trend sample set and constructing a style primitive library and style syntax graph are as follows: The normalized image and normalized text of each record in the trend sample set are mapped to the same style attribute dimension. Based on the binning table, appearance and decoration vocabulary and scene vocabulary, image-side style vector and text-side style vector are extracted. The style representation is generated by taking the largest value of each item according to the same index and binding it with the source identifier to form a trend sample set with style representation. Clustering is performed on the trend sample set with pattern representations, using real samples as representative points to form a pattern primitive library. Each record is assigned to the set of representative points most similar to the pattern representation. For each set of representative points, the real record with the highest average cosine similarity to the other pattern representations in the set is selected as the representative sample of the set. A pattern syntax graph is constructed based on the pattern representations and the pattern primitive library. The pattern primitive set is selected by sorting by similarity. The co-occurrence intensity is determined by the ratio of the number of records in the same record pattern primitive set where two pattern primitives appear simultaneously to the number of records where at least one pattern primitive appears.

5. The footwear style generation method based on multi-source data fusion and visual style as described in claim 4, characterized in that: The specific steps for obtaining the component mask and component style vector are as follows: For the normalized images in the trend sample set, the outer contour of the shoe body is extracted based on the foreground boundary of the shoe body region and the inscribed region of the shoe body is calculated. With the inscribed region of the shoe body as a constraint, connected component decomposition and boundary tracking are performed within the shoe body region to obtain candidate component regions. These regions are then grouped into the main body of the upper, side panels, toe, heel, outsole, logo, and shoelaces according to their relative positions and geometric shapes, resulting in a set of binary component masks of the same size. For representative samples of each style primitive in the style primitive library, crop the component region image based on the component mask coverage area and set the non-mask-covered pixels as a solid color background. Extract and normalize the component visual features according to the binning table, generate the component style vector of each component, and bind and save it with the source identifier of the representative sample.

6. The footwear style generation method based on multi-source data fusion and visual style as described in claim 5, characterized in that: The specific steps for constructing the spatial style planning diagram and prompt diagram to form the component blueprint are as follows: Based on the normalized text and keyword sequence of the style to be generated, determine the style primitives of the main body of the shoe upper in the style primitive library, and select the style primitives with greater co-occurrence intensity with the style primitives of the main body of the shoe upper in the style syntax diagram for the remaining parts in the order of the parts. Based on the component mask set, the component style vector is written into the component region according to the component to which the pixel belongs, forming a spatial style planning map with the same resolution as the normalized image. The component region image of the sample represented by the selected style primitive is filled according to the shape of the component mask to form a hint map. The component mask, component style vector, spatial style planning map and hint map are bound together to obtain the component blueprint.

7. The footwear style generation method based on multi-source data fusion and visual style as described in claim 1, characterized in that: The specific steps for extracting visual features of components based on component masks and calculating similarity with component style vectors as a component consistency index are as follows: Based on the component mask set in the component blueprint, the outer contour line of the shoe body, the outer contour line of the shoe sole, and the component boundary line are extracted and superimposed to generate a black and white binary structure control map. After aligning the structural control chart, component mask set, and prompt image, the controlled diffusion is used as the input for morphological constraints, spatial constraints, and appearance guidance. A candidate image set is generated based on a random seed generated by the traceability identifier. The component region image is cropped according to the component mask of the candidate image set and the component visual features are extracted. The cosine similarity between the component visual features and the component style vector is calculated as the component consistency. The minimum value of the consistency of all components in each candidate image is used as the component consistency index. The candidate image with the largest component consistency index is selected as the optimal candidate image, and the component with the smallest component consistency in the optimal candidate image is selected as the weakest component.

8. The footwear style generation method based on multi-source data fusion and visual style as described in claim 7, characterized in that: The process involves selecting the candidate image with the highest component consistency index and partially redrawing the component with the lowest consistency index to obtain the partially redrawn style image. The specific steps are as follows: Using the optimal candidate image as the base image, the local redrawing area is defined by the local redrawing area mask after expanding the component mask of the weakest component. Under the constraints of the structural control chart, component mask set, and prompt image, controlled diffusion local redrawing is performed based on a random seed generated by the traceability identifier and the name of the weakest component, generating a set of local redrawing candidate images. The component consistency index is repeatedly calculated for the set of local redrawing candidate images, and the local redrawing candidate image with the largest component consistency index is selected. It is then fused with the optimal candidate image according to the local redrawing area mask to obtain the style image after local redrawing, and it is bound and saved with the component blueprint.

9. The footwear style generation method based on multi-source data fusion and visual style as described in claim 1, characterized in that: The specific steps for semantic segmentation of the partially redrawn style drawing are as follows: After partially redrawing the style image, the direction of the shoe sole and the shoe toe is unified again. Candidate component regions are generated by grayscale conversion, edge extraction, closing operation, hole filling, shoe body region extraction, connected component decomposition and boundary tracking. These regions are then merged to obtain a set of binary prediction masks of the same size.

10. The footwear style generation method based on multi-source data fusion and visual style as described in claim 9, characterized in that: The specific steps for obtaining the predicted mask and verifying it with the component blueprint are as follows: The mask consistency is calculated based on the binary prediction mask set and the component mask set. The component with the smallest mask consistency is selected as the verification component. All components are verified for consistency. If the prediction mask set is completely consistent with the component mask set of the component blueprint, the style drawing after partial redrawing is determined as the final style drawing. Otherwise, correction is performed and semantic segmentation is repeated. The style drawing after correction is determined as the final style drawing.