Fabric matching method and system based on multi-modal information matching
By combining semantic structuring enhancement driven by a terminology vocabulary of the fabric industry with an improved CLIP model and image preprocessing, the problem of matching accuracy between fabric images and text descriptions was solved, achieving high-precision fabric localization and semantic consistency alignment between images and text under complex conditions.
Patent Information
- Application Number
- CN202511051916.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-29
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2045-07-29
AI Technical Summary
Existing technologies suffer from incomplete semantic coverage, inaccurate texture representation, and low matching accuracy in matching fabric images with text descriptions. In particular, it is difficult to achieve accurate fabric positioning and semantic consistency matching between images and text under complex shooting environments and non-standard descriptions.
A multimodal information matching method is adopted, which uses semantic structuring enhancement processing driven by a terminology vocabulary of the fabric industry, an improved CLIP text encoding model, image preprocessing, and cross-modal semantic alignment matching to achieve high-precision alignment between fabric images and text descriptions.
It improves the accuracy and response speed of fabric retrieval, and enhances the fabric matching accuracy and robustness in open description contexts, especially under complex textures and non-standard description conditions.
Smart Images

Figure CN120976585A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of intelligent matching recognition of fabrics, and in particular to a fabric matching method and system based on multi-modal information matching. BACKGROUND
[0002] At present, with the wide application of multi-modal learning models (such as CLIP) in image-text retrieval and content understanding, the industry has begun to try to introduce them into the matching task between fabric images and text descriptions to assist in clothing selection, design reference and supply chain fabric positioning. However, the existing methods still have key technical deficiencies, especially in processing fabric-level semantic matching, there are problems such as incomplete semantic coverage, inaccurate texture representation, and low matching accuracy, which are difficult to directly meet the needs of the industry end for fast and accurate selection.
[0003] For example, general image-text matching models are usually trained on open domain datasets, and the label space is difficult to accurately express professional fabric terms such as "acetic acid drape", "cationic blended twill", and "water-washed silk", which makes the model unable to fully identify the material properties, organizational structure and process characteristics expressed in user input. At the same time, existing methods generally ignore the complexity of the shooting environment of fabric images, such as wrinkles, reflections, background interference and other factors, which seriously dilutes the response strength of key texture information in the visual semantic vector in the image, resulting in unstable image-text matching results, especially in long-tail semantic description or cross-brand retrieval scenarios.
[0004] The existing technology cannot fully meet the needs of designers and buyers to quickly locate the corresponding fabric based on natural language description and further associate it to the Taobao, Douyin, and social platform best-selling styles. Therefore, there is an urgent need for a multi-modal matching method that can achieve high-precision and robust semantic alignment between fabric images and text descriptions in the case of insufficient industry semantic space or non-ideal image quality, in order to improve the matching accuracy, response speed and actual conversion rate of the intelligent selection system. SUMMARY
[0005] In view of the above technical deficiencies, the purpose of the present application is to provide a fabric matching method based on multi-modal information matching, which aims to solve the technical problems that the existing technology uses simple text retrieval or single-channel feature matching method, especially under the conditions of non-standard user description, complex fabric image texture or repeated structure, which cannot realize accurate fabric positioning and image-text semantic consistency matching.
[0006] To solve the above technical problems, the present application adopts the following technical scheme: the present application provides a fabric matching method based on multi-modal information matching,
[0007] The fabric matching method based on multi-modal information matching comprises:
[0008] Step S10: based on the preset fabric industry term vocabulary W, the original text T input by the user is subjected to semantic structured enhancement processing, and an enhanced description text is output ; wherein the semantic structured enhancement processing includes adopting a multi-level semantic regularization recognition mechanism and a semantic insertion splicing strategy based on weight prompt embedding;
[0009] Step S20: input the enhanced description text into an improved CLIP text encoding model, the improved CLIP text encoding model outputs a first text semantic vector, and adopts a trainable vector projection mechanism based on semantic category guidance to extract a second text semantic vector related to the fabric structure attribute from the first text semantic vector;
[0010] Step S30: obtain an original fabric image from a preset background database, perform foreground mask extraction processing, multi-channel texture fidelity enhancement processing and structure equalization normalization processing on the original fabric image, and output preprocessed image data;
[0011] Step S40: perform texture saliency region division on the preprocessed image data, and generate an image global semantic vector V and a texture region sub-semantic vector B based on the region division result;
[0012] Step S50: based on the image global semantic vector V and the texture region sub-semantic vector B generated in step S40, perform a cross-modal semantic alignment matching task in combination with the second text semantic vector, and output a final matching fabric result based on a multi-factor joint sorting mechanism.
[0013] Preferably, in step S10, the fabric industry term vocabulary W includes N hierarchical fabric terms, including a material subset W1, an organization texture subset W2, a pattern subset W3, and a style modification subset W4; in step S10, based on the preset fabric industry term vocabulary W, the original text T input by the user is subjected to semantic structured enhancement, and an enhanced description text is output. The steps specifically include: in the original text T, a multi-level semantic regularization recognition mechanism is adopted to perform term matching, a term sequence of the hierarchical fabric term is identified, an industry annotation expansion is performed based on the term sequence of the hierarchical fabric term through a preset word meaning mapping table, and a connection is constructed on the annotation expansion result by adopting a semantic insertion splicing strategy based on weight prompt embedding, and an enhanced description text is output.
[0014] Preferably, in step S20, the improved CLIP text encoding model specifically includes: an input layer for receiving the enhanced description text ; a semantic structure prompt embedding layer for introducing a guide prompt vector representing fabric term classification; a multi-scale attention focusing layer for improving the model's attention response capability to fabric keywords at different levels; a domain feature adaptation enhancement layer for fusing the semantic context distribution pattern obtained by fine-tuning the fabric industry corpus; and an output layer for outputting a first text semantic vector;
[0015] The second text semantic vector is used for subsequent semantic alignment with the texture salient region extracted from the image. The extraction method of the second text semantic vector specifically includes: performing weight learning and semantic clustering processing on each dimension in the first text semantic vector by setting a vector projection structure for organizational structure attribute recognition, identifying and extracting a sub-semantic feature part related to fabric organization and texture morphology, and performing vectorization processing on the sub-semantic feature part to obtain the second text semantic vector.
[0016] Preferably, in step S30, the original fabric image is obtained from the preset background database, and the original fabric image is subjected to foreground mask extraction processing, multi-channel texture fidelity enhancement processing and structure uniformity normalization processing, and the step of outputting the preprocessed image data specifically includes:
[0017] Step S301: The original image data is called from the preset background database in the form of an open API using the https protocol;
[0018] Step S302: Perform pixel-level region segmentation on the original image data based on the preset weakly supervised training type semantic segmentation model to generate a fabric region mask image, and perform masking processing on the original image data using the fabric region mask image to remove background area interference and obtain optimized image data;
[0019] Step S303: Decompose the optimized image data into two types of sub-channels, i.e., a brightness channel and a texture direction channel, and perform local structure clarity enhancement and direction consistency enhancement on each sub-channel, respectively, to obtain a reconstructed enhanced image;
[0020] Step S304: Detect high-frequency abnormal regions in the reconstructed enhanced image, and perform edge smoothing processing on wrinkle shadows, light overexposure and reflection regions to retain real texture information, and output preprocessed image data.
[0021] Preferably, in step S40, the preprocessed image data is subjected to texture saliency region modeling operation, and the step of generating an image global semantic vector and a texture region sub-semantic vector based on the modeling result specifically includes:
[0022] The preprocessed image data is divided into N local image regions based on the adaptive image division mechanism of the perceived texture response, a local feature vector set of each region is extracted, and a local feature vector of each region i is generated based on the local feature vector set ;
[0023] According to the local feature vector of each region Calculate its texture saliency score ;
[0024] Based on the texture saliency score of all regions, the local feature vector is weighted and fused to obtain the image global semantic vector V. ;
[0025] All regions are arranged in descending order according to the texture saliency score, and the first K regions are selected to extract the corresponding local feature vectors, and then spliced or averaged to generate the texture region sub-semantic vector B.
[0026] Preferably, in step S40, the texture saliency score Defined by the following formula:
[0027] ;
[0028] Wherein, Indicates the response energy obtained after the local feature vector Multi-scale multi-direction Gabor filtering processing, used to measure the texture directionality of the region; Indicates the two-norm of the image gradient, reflecting the structural change intensity of the local region; Indicates the information entropy of the local pixel gray scale distribution, used to measure the image complexity or texture density of the region; α, β, γ are saliency weighting coefficients, corresponding to the relative weights of the above three dimensions respectively, satisfying α+β+γ= 1.
[0029] Preferably, in step S50, based on the image global semantic vector V and the texture region sub-semantic vector B generated in step S40, the second text semantic vector is combined to perform a cross-modal semantic alignment matching task, and a multi-factor joint sorting mechanism is used to output the final matching fabric result. The steps specifically include:
[0030] The cosine similarity between the second text semantic vector and the image global semantic vector V is calculated to obtain the global semantic matching score S1 between the text and the image, which is used to measure the semantic consistency between the text description and the overall style of the image.
[0031] The cosine similarity between the second text semantic vector and each sub-vector in the texture region sub-semantic vector B is calculated, and the local similarity scores are weighted and fused based on the preset Top-K aggregation rule to obtain the local texture semantic matching score S2, which is used to measure the fine-grained semantic association degree between the text description and the key texture structure of the image.
[0032] An inter-region texture semantic discreteness indicator S3 is calculated based on the semantic included angle or distance between all sub-vectors in the texture region sub-semantic vector B, and the inter-region texture semantic discreteness indicator S3 is used to reflect the diversity distribution of the texture structure of the image;
[0033] A joint matching score function is constructed based on the global image-text semantic matching score S1, the local texture semantic matching score S2 and the inter-region texture semantic discreteness indicator S3, the joint matching score function outputs a joint matching score result, the original fabric image in step S30 is sorted according to the joint matching score result, and a final matching fabric result is output.
[0034] The application also provides a fabric matching system based on multi-modal information matching, which comprises:
[0035] A text semantic enhancement module is configured to perform semantic structured enhancement processing on the original text T input by the user based on a preset fabric industry term table W, and output an enhanced description text ; wherein the semantic structured enhancement processing comprises a multi-level semantic regularization recognition mechanism and a semantic insertion splicing strategy based on weight prompt embedding;
[0036] A semantic vector extraction module for organizational attributes is configured to input the enhanced description text into an improved CLIP text encoding model, the improved CLIP text encoding model outputs a first text semantic vector, and a trainable vector projection mechanism based on semantic category guidance is used to extract a second text semantic vector related to the fabric organizational structure attribute from the first text semantic vector;
[0037] A fabric image preprocessing module is configured to obtain an original fabric image from a preset background database, perform foreground mask extraction processing, multi-channel texture fidelity enhancement processing and structure equalization normalization processing on the original fabric image, and output preprocessed image data;
[0038] A texture saliency modeling module is configured to perform texture saliency region division on the preprocessed image data, and generate an image global semantic vector V and a texture region sub-semantic vector B based on the region division result;
[0039] A cross-modal alignment and sorting output module is configured to perform a cross-modal semantic alignment matching task based on the image global semantic vector V and the texture region sub-semantic vector B generated in step S40, in combination with the second text semantic vector, and output a final matching fabric result based on a multi-factor joint sorting mechanism.
[0040] The application further provides a fabric matching device based on multi-modal information matching, comprising a memory, a processor, and a fabric matching program based on multi-modal information matching stored on the memory and capable of running on the processor, which implements the fabric matching method based on multi-modal information matching when executed by the processor.
[0041] The application further provides a computer program product comprising a fabric matching program based on multi-modal information matching, which implements the fabric matching method based on multi-modal information matching when executed by a processor.
[0042] The application has the beneficial effect that, compared with the prior art which adopts simple text retrieval or single-channel feature matching mode, especially under the condition that the user description is not standard, the fabric image texture is complex, or the structure is repeated, the technical problem of precise fabric positioning and matching of text and image semantics consistency cannot be solved. Since the application introduces an industry ontology vocabulary driven semantic enhancement mechanism and a saliency guided cross-modal matching architecture, it realizes multi-granularity modeling and fine-granularity semantic alignment of fabric semantic features, and improves the fabric retrieval accuracy in an open description context. BRIEF DESCRIPTION OF DRAWINGS
[0043] In order to more clearly illustrate the technical solutions in the embodiments of the application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or the prior art description. Obviously, the drawings in the following description only constitute some embodiments of the application, and for those skilled in the art, other drawings can also be obtained without creative labor.
[0044] Figure 1 The flowchart of the first embodiment of the fabric matching method based on multi-modal information matching of the application.
[0045] Figure 2 The device schematic diagram of the fabric matching method based on multi-modal information matching of the application. DETAILED DESCRIPTION
[0046] The technical solutions in the embodiments of the application will be described clearly and completely below with reference to the drawings in the embodiments of the application. Obviously, the described embodiments only constitute some embodiments of the application, rather than all the embodiments. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative labor fall within the protection scope of the application.
[0047] Embodiment one: as Figure 1As shown, it is a flowchart of the first embodiment of the fabric matching method based on multi-modal information matching of the present application, and the first embodiment of the fabric matching method based on multi-modal information matching of the present application is proposed.
[0048] In the first embodiment, the fabric matching method based on multi-modal information matching comprises:
[0049] Step S10: based on the preset fabric industry term table W, the original text T input by the user is subjected to semantic structured enhancement processing, and an enhanced description text T* is output ; wherein the semantic structured enhancement processing comprises a multi-level semantic regularization recognition mechanism and a semantic insertion splicing strategy based on weight prompt embedding;
[0050] It should be noted that the "multi-level semantic regularization recognition mechanism" refers to a combined processing process of part-of-speech analysis, hierarchical classification and semantic constraint recognition of the original text T based on the fabric industry term table W, including but not limited to independent extraction and clustering modeling of fabric material type words (such as "acetic acid" and "nylon"), organizational structure words (such as "twill" and "double yarn"), and process treatment words (such as "washing" and "sanding"); and the "semantic insertion splicing strategy based on weight prompt embedding" refers to setting weights according to the semantic importance of the above identified core words and guiding through a prompt template (prompt), embedding missing key terms or modifiers into the original text according to the set rules, and constructing a structured enhanced description text T*.
[0051] It can be understood that this structured enhancement processing can effectively solve the problems of missing industry terms, semantic ambiguity or non-standard expression in user natural language description, so that the input text is more consistent with the semantic alignment requirements in the fabric image-text matching task, thereby improving the perception ability of the subsequent coding model to fine-grained semantics.
[0052] It should be understood that, compared with the traditional CLIP model which directly processes the original text, the present application introduces an explicit semantic regularization structure and a prompt enhancement mechanism, which significantly improves the ability of the text embedding vector to distinguish industry-specific attributes without changing the original model structure, especially when long-tail terms (such as "nylon and light weight") are involved.
[0053] For example, in a group of original texts containing natural descriptions "thick hand feeling, suitable for skirts, slightly wrinkled chiffon", the traditional CLIP model fails to identify the corresponding "high twist chiffon" semantic label, while the present application generates "thick chiffon fabric with high twist yarn organization, with natural wrinkle feeling, suitable for skirt scene" after structured enhancement, and the Top-1 retrieval accuracy of the corresponding fabric image is improved from 68.4% to 91.2%, significantly improving the professionalism and commercial practicality of the matching result.
[0054] Step S20: input the enhanced description text into the improved CLIP text encoding model, the improved CLIP text encoding model outputs a first text semantic vector, and a trainable vector projection mechanism based on semantic category guidance is used to extract a second text semantic vector related to the fabric structure attribute from the first text semantic vector;
[0055] It should be noted that the "trainable vector projection mechanism based on semantic category guidance" refers to that in the first text semantic vector output by the original CLIP text encoder, a plurality of semantic projection subspaces are constructed by guiding labels using a set of semantic categories predefined by fabric industry knowledge (such as "structure", "material composition", "functional characteristics", "application scenarios", etc.), and a soft constraint projection is performed on the first vector to output one or more sub-vectors related to one or more semantic dimensions, and the second text semantic vector corresponding to the "structure attribute" represents the embedding information of the description in the structure dimension.
[0056] It can be understood that this mechanism adaptively models the contribution strength of each dimension information in the first semantic vector to the "structure" category semantic through a learning weight matrix or attention parameter, thereby mapping the abstract semantic embedding to an industry-dimension interpretable and function-specific sub-vector space, so that the subsequent image matching task not only "knows what you said", but also "knows what you said".
[0057] It should be understood that compared with the single semantic vector uniformly generated by the traditional CLIP model, the relative weights of "material" and "structure" in the description cannot be distinguished, and matching drift often occurs when multiple attributes are mixed in the description. The present application effectively separates and purifies the target attribute dimension semantic through the guided semantic decomposition mechanism, so that the matching process has stronger structure selectivity and target orientation, and is especially suitable for understanding and extracting of commodity descriptions with multiple attribute compressed expressions (such as "cotton and linen twill elastic cloth").
[0058] For example, the user input description is "soft twill cotton cloth, suitable for casual suits", the structured text is generated after enhancement, the first text semantic vector is output by CLIP, and the "structure" sub-vector is projected through the mechanism, which focuses on the "twill" attribute expression, and the vector is matched with the "twill weaving features" in the image. The comparative experiment shows that after using the guided semantic projection mechanism of the present application, the Top-1 matching accuracy related to the structure is improved from 74.2% to 92.3%, which significantly improves the dimension accuracy and controllability of image-text alignment.
[0059] Step S30: Obtain the original fabric image from the preset background database, perform foreground mask extraction processing, multi-channel texture fidelity enhancement processing and structure balance normalization processing on the original fabric image, and output preprocessed image data;
[0060] It should be noted that in step S30, the steps of obtaining the original fabric image from the preset background database, performing foreground mask extraction processing, multi-channel texture fidelity enhancement processing and structure balance normalization processing on the original fabric image, and outputting preprocessed image data, specifically include: step S301: using https protocol to call original image data from the preset background database in the form of open API; step S302: performing pixel-level region segmentation on the original image data based on the preset weakly supervised training semantic segmentation model to generate a fabric region mask image, and performing masking processing on the original image data using the fabric region mask image to remove background area interference and obtain optimized image data; step S303: decomposing the optimized image data into two types of sub-channels, i.e. brightness channel and texture direction channel, and respectively performing local structure definition enhancement and direction consistency enhancement to obtain a reconstructed enhanced image; step S304: detecting high-frequency abnormal regions in the reconstructed enhanced image, and performing edge smoothing processing on wrinkle shadows, light overexposure and reflection regions to retain real texture information, and outputting preprocessed image data. Foreground mask extraction processing refers to extracting the main fabric region and excluding background interference through a lightweight segmentation model (such as improved U-Net or SAM) combined with edge enhancement and color clustering; multi-channel texture fidelity enhancement processing refers to enhancing texture contrast and directionality in RGB, LAB and grayscale spaces respectively through frequency domain enhancement and Retinex light normalization to retain local changes of fabric features; structure balance normalization processing refers to performing balance on image gray mean, contrast and sharpness based on the distribution of the whole image gray structure tensor to make the image maintain comparability and expression consistency under different shooting conditions.
[0061] It can be understood that the core goal of this three-stage processing flow is to maximize the expression of significant textures in the image that are strongly related to fabric attributes without introducing false features, while suppressing the negative interference of factors such as light changes, wrinkle shadows, scanning backgrounds, and fabric edge interference on subsequent semantic encoders.
[0062] It should be understood that traditional image enhancement operations (such as global histogram stretching or sharpening filtering) are not specific to different fabric types and are prone to overfitting to certain image styles, thereby losing generality. The enhancement strategy of the present application emphasizes the combination of semantic relevance and visual consistency, and enhances the robustness through multi-channel fusion, so that the image encoding vector has high stability under different light, shooting distance and roll fabric wrinkle conditions, significantly improving the domain adaptation ability and anti-drift ability of the image-text matching model.
[0063] For example, facing the original drawing of "white pigment color cotton and linen cloth", the background in the original image is a desktop texture and the cloth corners are rolled up. Without processing, the CLIP model extracts vectors mainly in the "desktop wood grain + cloth projection" area, and the matching result is often mistaken as "wood grain jacquard". After S30 processing, the background is successfully removed, the main texture direction of the fabric is enhanced, and after local brightness normalization, the model can accurately identify its "plain weave and white color" characteristics, and the matching accuracy is improved by about 19.6%, and the interference of "false detection of high-contrast decorative patterns" is significantly reduced.
[0064] Step S40: performing texture saliency region division on the preprocessed image data, and generating an image global semantic vector V and a texture region sub-semantic vector B based on the region division result;
[0065] It should be noted that the "texture saliency region division" of step S40 is not simply dividing the image into regular grid regions, but is based on adaptive partitioning of the image content. The purpose is to use the local performance of the image in terms of texture directionality, detail complexity and structure gradient to determine which regions have more semantic expression ability. This division strategy combines the statistics of multiple visual factors, and gives each local region a "texture saliency score". This score is used to measure whether the region has strong organization, typical structure or recognition value. This division is not an isolated image segmentation, but provides structural guidance for the subsequent semantic vector construction process.
[0066] It can be understood that this step is to solve the problem of missing main texture expression caused by the generalization method of "average processing" and "full image perception" of original CLIP and other multi-modal models. Due to the presence of shooting defects such as wrinkles, light spots, rolled edges, and uneven illumination in actual fabric images, if the entire image is input into the visual encoder with equal weight, the proportion of the main texture signal in the encoding vector will be significantly reduced, causing the core semantics of the matching to drift. The present application introduces a structured saliency region recognition mechanism before image encoding, ensuring that the main texture region can obtain a higher weight, which not only improves the resolution of image semantics, but also enhances the sensitivity of the model in the matching task.
[0067] It should be understood that, compared with the method of directly using ResNet or ViT to extract an image vector and then calculating the cosine similarity with the text vector, the saliency region division mechanism introduced in this step has two key advantages: (1) it constructs a dual-channel expression structure of "global semantic vector" and "texture region sub-semantic vector", the former ensures the expression of overall fabric style and tone, and the latter focuses on key texture details; (2) it actively excludes irrelevant or even interfering regions (such as background tablecloth, shadow, corner wrinkle, etc.) in the image through the structure modeling process, thereby significantly enhancing the focus and matching robustness of semantic alignment.
[0068] For example, for a "diamond jacquard tencel" image taken in natural light with slight wrinkles, the traditional method may tend to match the "light color highlight plain cloth" due to the large weight of the highlight reflection area in the lower right corner of the image in the visual vector. In this step, through the saliency score, the central region is identified as having strong directional texture, rich structure gradient change and high local entropy, so the region is extracted as a high saliency region and a sub-semantic vector is constructed. In the subsequent alignment stage, this texture sub-semantic vector becomes dominant in matching, successfully retrieving the same type of tencel fabric with "diamond texture", improving search accuracy and enhancing user experience.
[0069] Step S50: Based on the image global semantic vector V and the texture region sub-semantic vector B generated in step S40, a cross-modal semantic alignment matching task is performed in combination with the second text semantic vector, and a final matching fabric result is output based on a multi-factor joint ranking mechanism.
[0070] It should be noted that the "cross-modal semantic alignment matching task" refers to the joint alignment of visual semantics and text semantics in a multi-scale semantic space by calculating the embedding space distance between the image global semantic vector V, the texture region sub-semantic vector B and the second text semantic vector. This process does not simply use cosine similarity as the only measurement criterion, but constructs a multi-factor scoring function that integrates three core indicators: semantic relevance, texture saliency consistency and structure feature coverage, and adjusts the parameters to achieve adaptive optimization of the ranking results.
[0071] It can be understood that the core of the design of this step is to solve the following two problems: (1) the general CLIP often lacks the ability to perceive details in the alignment task, and it is difficult to distinguish the structural differences between "plain pure cotton" and "twill pure cotton"; (2) although the texture information in the fabric image is rich, if the texture area is not given enough weight in the matching score function, the problem of "correct style but wrong details" in the matching result is easy to appear. The application introduces an image sub-semantic vector B to participate in the scoring process, and specially strengthens the influence of structural details on the alignment result, effectively alleviating the industry-level pain point of "semantic alignment but physical inconsistency".
[0072] For example, for the user input "twill with crisp feeling and water-washed tencel fabric", the traditional CLIP method is likely to recommend a soft fabric with similar color tone but plain weave, ignoring keywords such as "twill" and "crisp feeling" with fine-grained structure. The method introduces a second text vector T2 and an image B alignment mechanism in the Score function, which improves the ranking weight of the "twill" dimension in the final result, so that the returned result preferentially contains fabric images with "twill direction" and "crisp texture". The actual measurement shows that the human eye structure consistency score of the top-3 recommended results is improved from 63.5% of the original model to 92.1%, which has a significant effect of improving the practical value of matching.
[0073] Embodiment two: In addition, the fabric matching system based on multi-modal information matching provided by the application adopts the fabric matching method based on multi-modal information matching in the above embodiment, which can solve the technical problem of the fabric matching based on multi-modal information matching. Compared with the prior art, the fabric matching system based on multi-modal information matching provided by the application has the same beneficial effects as the fabric matching method based on multi-modal information matching provided by the above embodiment, and other technical features in the fabric matching system based on multi-modal information matching are the same as the features disclosed in the above embodiment method, which will not be repeated here.
[0074] Embodiment three: The application provides a fabric matching device based on multi-modal information matching, please refer to Figure 2An apparatus for fabric matching based on multi-modal information matching includes at least one processor; and a memory connected to the at least one processor in communication; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method for fabric matching based on multi-modal information matching of the above-mentioned embodiment one. The apparatus for fabric matching based on multi-modal information matching of the embodiments of the present application can include, but is not limited to, mobile terminals such as mobile phones, notebook computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), car terminals (for example, car navigation terminals), and the like, and fixed terminals such as digital TVs, desktop computers, and the like. The apparatus for fabric matching based on multi-modal information matching is merely an example, and should not impose any limitation on the functions and use range of the embodiments of the present application. The apparatus for fabric matching based on multi-modal information matching can include a processing device 1001 which can perform various appropriate actions and processes according to programs stored in a read-only memory 1002 or loaded into a random access memory 1004 from a storage device 1003. In the random access memory 1004, various programs and data required for the operation of the apparatus for fabric matching based on multi-modal information matching are also stored. The processing device 1001, the read-only memory 1002, and the random access memory 1004 are connected to each other through a bus 1005. An I / O interface 1006 is also connected to the bus. In general, the following systems can be connected to the I / O interface 1006: input devices 1007 including, for example, touch screens, touch pads, keyboards, mice, image sensors, microphones, accelerometers, gyroscopes, and the like; output devices 1008 including, for example, liquid crystal displays (LCDs), speakers, vibrators, and the like; the storage device 1003; and a communication device 1009. The communication device 1009 can allow the apparatus for fabric matching based on multi-modal information matching to communicate wirelessly or wiredly with other devices to exchange data. Although the apparatus for fabric matching based on multi-modal information matching having various systems is shown in the drawing, it should be understood that all of the shown systems are not required to be implemented or provided. More or less systems can be alternatively implemented or provided.
[0075] Embodiment Four: The present application also provides a computer program product comprising a computer program which, when executed by a processor, implements the steps of a fabric matching method based on multi-modal information matching as described above. The computer program product provided by the present application can solve the technical problem of fabric matching based on multi-modal information matching. Compared with the prior art, the beneficial effects of the computer program product provided by the present application are the same as those of the fabric matching method based on multi-modal information matching provided by the above-described embodiments, and are not described here in detail.
[0076] In particular, according to the embodiments disclosed by the present application, the processes described above with reference to the flowcharts can be implemented as a computer software program. For example, the embodiments disclosed by the present application include a computer program product comprising a computer program carried on a computer readable medium, the computer program comprising program code for executing the method shown in the flowchart. In such embodiments, the computer program can be downloaded and installed from a network through a communication device, or installed from a storage device 1003, or installed from a read-only memory 1002. When the computer program is executed by a processing device 1001, the above-mentioned functions defined in the method of the embodiments disclosed by the present application are executed.
[0077] It should be understood that various parts of the present application can be realized in hardware, software, firmware, or a combination thereof. In the description of the above-described embodiments, specific features, structures, materials or characteristics can be combined in any appropriate manner in any one or more embodiments or examples.
[0078] Obviously, those skilled in the art can make various modifications and variations to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application also intends to include these modifications and variations.
Claims
1. A fabric matching method based on multimodal information matching, characterized in that, The methods include: Step S10: Based on the preset fabric industry terminology glossary W, perform semantic structuring enhancement processing on the user-input raw text T, and output enhanced descriptive text. Among them, semantic structuring enhancement processing includes adopting a multi-level semantic regularization recognition mechanism and a semantic insertion and splicing strategy based on weighted hint embedding; Step S20: Enhance the description text The input is fed into the improved CLIP text encoding model, which outputs a first text semantic vector and uses a semantic category-guided trainable vector projection mechanism to extract a second text semantic vector related to the fabric structure attributes from the first text semantic vector. Step S30: Obtain the original fabric image from the preset background database, perform foreground mask extraction, multi-channel texture fidelity enhancement and structural balance normalization on the original fabric image, and output the preprocessed image data. Step S40: Perform texture saliency region segmentation on the preprocessed image data, and generate the global semantic vector V and texture region sub-semantic vector B based on the region segmentation results; Step S50: Based on the global semantic vector V of the image and the sub-semantic vector B of the texture region generated in step S40, perform a cross-modal semantic alignment matching task in combination with the second text semantic vector, and output the final matching fabric result based on the multi-factor joint ranking mechanism.
2. The fabric matching method based on multimodal information matching as described in claim 1, characterized in that, In step S10, the fabric industry terminology glossary W includes N hierarchical fabric terms, including a material subset W1, a texture subset W2, a pattern subset W3, and a style modification subset W4. In step S10, based on the preset fabric industry terminology glossary W, the user-inputted raw text T is semantically structured and enhanced, and an enhanced description is output. The steps specifically include: performing term matching in the original text T using a multi-level semantic regularization recognition mechanism to identify the term sequence of layered fabric terms; performing industry annotation expansion based on the term sequence of layered fabric terms through a preset word meaning mapping table; and constructing connections for the annotation expansion results using a semantic insertion and splicing strategy based on weighted hint embedding to output enhanced descriptive text. .
3. The fabric matching method based on multimodal information matching as described in claim 1, characterized in that, In step S20, the improved CLIP text encoding model specifically includes: an input layer for receiving enhanced descriptive text. The semantic structure hint embedding layer is used to introduce guiding hint vectors representing fabric term classification; the multi-scale attention focusing layer is used to improve the model's attention response capability to fabric keywords at different levels; the domain feature adaptation enhancement layer is used to integrate the semantic context distribution pattern obtained by fine-tuning the fabric industry corpus; and the output layer is used to output the first text semantic vector. The second text semantic vector is used for subsequent semantic alignment with the texture salient regions extracted from the image. The extraction method of the second text semantic vector specifically includes: by setting a vector projection structure for organizational structure attribute recognition, weight learning and semantic clustering are performed on each dimension of the first text semantic vector to identify and extract the sub-semantic features related to fabric organization and texture morphology, and the sub-semantic features are vectorized to obtain the second text semantic vector.
4. The fabric matching method based on multimodal information matching as described in claim 1, characterized in that, Step S30 involves obtaining the original fabric image from a preset background database, performing foreground mask extraction, multi-channel texture fidelity enhancement, and structural equalization normalization on the original fabric image, and outputting the preprocessed image data. Specifically, this includes: Step S301: Retrieve raw image data from a pre-set backend database using the HTTPS protocol and an open API. Step S302: Based on the preset weakly supervised training semantic segmentation model, the original image data is segmented into pixels at the region level to generate a fabric region mask map. The fabric region mask map is then used to perform masking processing on the original image data to remove background interference and obtain optimized image data. Step S303: Decompose the optimized image data into two sub-channels: brightness channel and texture direction channel. Improve the local structural clarity and enhance the direction consistency of the sub-channels respectively to obtain the reconstructed and enhanced image. Step S304: Detect high-frequency abnormal regions in the reconstructed and enhanced image, and perform edge smoothing on wrinkle shadows, overexposed lighting, and reflection areas to preserve real texture information and output preprocessed image data.
5. The fabric matching method based on multimodal information matching as described in claim 1, characterized in that, Step S40, which involves performing texture saliency region modeling on the preprocessed image data and generating global semantic vectors and texture region sub-semantic vectors based on the modeling results, specifically includes: The adaptive image segmentation mechanism based on perceptual texture response divides the preprocessed image data into N local image regions, extracts the local feature vector set of each region, and generates the local feature vector of each region i based on the local feature vector set. ; Based on the local feature vectors of each region Calculate its texture saliency score ; Based on the texture saliency scores of all regions, local feature vectors Weighted fusion is performed to obtain the global semantic vector V of the image; All regions are arranged in descending order of texture saliency score. The top K regions are selected to extract their corresponding local feature vectors, which are then concatenated or averaged to generate the texture region sub-semantic vector B.
6. The fabric matching method based on multimodal information matching as described in claim 5, characterized in that, In step S40, texture saliency score Defined by the following formula: ; in, Represents the local eigenvectors The response energy obtained after multi-scale, multi-directional Gabor filtering is used to measure the texture directionality of the region. It represents the L2 norm of the image gradient, reflecting the intensity of structural changes in local regions; The information entropy represents the local pixel grayscale distribution and is used to measure the image complexity or texture density of the region; α, β, and γ are significance weighting coefficients, which correspond to the relative weights of the three dimensions mentioned above, satisfying α+β+γ=1.
7. The fabric matching method based on multimodal information matching as described in claim 1, characterized in that, Step S50, which involves performing a cross-modal semantic alignment matching task based on the image global semantic vector V and texture region sub-semantic vector B generated in step S40, combined with the second text semantic vector, and outputting the final matching fabric result based on a multi-factor joint ranking mechanism, specifically includes: Cosine similarity calculation is performed on the second text semantic vector and the image global semantic vector V to obtain the image-text global semantic matching score S1. The image-text global semantic matching score S1 is used to measure the semantic consistency between the text description and the overall style of the image. The cosine similarity is calculated between the semantic vector of the second text and each sub-vector of the texture region sub-semantic vector B. The local similarity scores are then weighted and fused based on the preset Top-K aggregation rule to obtain the local texture semantic matching score S2. The local texture semantic matching score S2 is used to measure the fine-grained semantic association between the text description and the key texture structure of the image. Based on the semantic angle or distance between all sub-vectors in the texture region sub-semantic vector B, the texture semantic discreteness index S3 between regions is calculated. The texture semantic discreteness index S3 between regions is used to reflect the diversity distribution of image texture structure. A joint matching score function is constructed based on the global semantic matching score S1, the local texture semantic matching score S2, and the inter-regional texture semantic discreteness index S3. The joint matching score function outputs the joint matching score result. The original fabric images in step S30 are sorted according to the joint matching score result, and the final matching fabric result is output.
8. A fabric matching system based on multimodal information matching, applied to the fabric matching method based on multimodal information matching according to any one of claims 1 to 7, characterized in that, The fabric matching system based on multimodal information matching includes: The text semantic enhancement module is used to perform semantic structuring enhancement on the user-input raw text T based on a preset fabric industry terminology lexicon W, and output enhanced descriptive text. Among them, semantic structuring enhancement processing includes adopting a multi-level semantic regularization recognition mechanism and a semantic insertion and splicing strategy based on weighted hint embedding; A semantic vector extraction module for organizational attributes is used to enhance descriptive text. The input is fed into the improved CLIP text encoding model, which outputs a first text semantic vector and uses a semantic category-guided trainable vector projection mechanism to extract a second text semantic vector related to the fabric structure attributes from the first text semantic vector. The fabric image preprocessing module is used to obtain the original fabric image from the preset background database, perform foreground mask extraction, multi-channel texture fidelity enhancement and structural balance normalization on the original fabric image, and output preprocessed image data. The texture saliency modeling module is used to perform texture saliency region segmentation on preprocessed image data and generate a global semantic vector V and texture region sub-semantic vector B based on the region segmentation results. The cross-modal alignment and sorting output module is used to perform cross-modal semantic alignment matching task based on the image global semantic vector V and texture region sub-semantic vector B generated in step S40, combined with the second text semantic vector, and output the final matching fabric result based on the multi-factor joint sorting mechanism.
9. A fabric matching device based on multimodal information matching, characterized in that, The fabric matching device based on multimodal information matching includes: a memory, a processor, and a fabric matching program based on multimodal information matching stored in the memory and executable on the processor. When the fabric matching program based on multimodal information matching is executed by the processor, it implements a fabric matching method based on multimodal information matching according to any one of claims 1 to 7.
10. A computer program product, characterized in that, The computer program product includes a fabric matching program based on multimodal information matching, which, when executed by a processor, implements a fabric matching method based on multimodal information matching as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Virtual fitting method and device
CN119444349A
Knowledge guidance-based camouflage target detection method and system
CN120070855A
Fine-grained costume image retrieval method and device based on large language model common knowledge injection
CN120196777A
Image relevance to search queries based on unstructured data analytics
US20160063096A1